Differential privacy-based multi-agent task unloading and resource optimization method and system
By introducing differential privacy protection mechanisms and information sharing mechanisms in the optimization of multi-agent task offloading and resource allocation, the problems of privacy leakage, inefficient communication efficiency and insufficient global optimization capabilities are solved, and efficient and secure multi-agent collaborative optimization is achieved.
Patent Information
- Application Number
- CN202510311997.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems such as privacy leakage, inefficient communication and insufficient global optimization capabilities in multi-agent task offloading and resource allocation optimization.
Multi-agent task offloading and resource optimization methods based on differential privacy protection are adopted. By introducing Gaussian noise in message transmission, sensitive data is ensured not directly leaked, and efficient transmission and accurate processing of global information is achieved through the information sharing mechanism between agents.
It effectively protects private information between agents, improves communication efficiency and global optimization capabilities, and improves the collaboration performance and task processing efficiency of multi-agent systems.
Smart Images

Figure CN120151944A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of multi-agent reinforcement learning, and particularly relates to a multi-agent task offloading and resource optimization method and system based on differential privacy. Background Technique
[0002] Mobile Edge Computing (MEC) is a new type of computing architecture that provides low-latency, high-bandwidth, and efficient computing services for mobile devices (MDs) by sinking computing resources to the network edge. A typical MEC system includes mobile devices, base stations, and edge servers. Mobile devices are responsible for generating computing tasks, which can be either completed locally or offloaded to the edge server (ES) for processing via a wireless link. The edge server embedded in the base station provides computing resources for mobile devices, effectively alleviating the computing pressure on mobile devices. With the rapid development of 5G networks and Internet of Things technologies, MEC has demonstrated extensive application value in multiple fields. For example, in the field of intelligent transportation, it supports real-time task processing in vehicle-to-everything (V2X), such as vehicle path planning, collision warning, and optimization of autonomous driving decisions; in industrial Internet of Things, it is applied to the status monitoring, fault diagnosis, and predictive maintenance of production equipment; in online real-time games, it provides multi-player low-latency interaction support; in the field of high-definition video processing, it realizes real-time video coding and content distribution optimization; in virtual reality / augmented reality (VR / AR), it supports complex rendering tasks and multi-person immersive experiences. In addition, MEC also plays an important role in scenarios such as smart cities, healthcare, and public safety, improving system performance and resource utilization efficiency through task offloading and resource allocation (TOTPA) optimization to meet the requirements of complex applications for high computing power, low latency, and real-time response. As one of the core technologies of MEC, TOTPA helps to improve the overall service quality and collaboration efficiency in complex multi-task environments by optimizing task offloading decisions and resource allocation.
[0003] In an MEC system, task processing includes steps such as task generation, local processing, or edge offloading. Task information includes data size, CPU cycle count, and maximum tolerance delay, etc. Local processing has the characteristics of low transmission delay but high computing energy consumption, while task offloading transmits tasks to the edge server for processing via a wireless link, which can reduce the local computing pressure but may increase transmission delay and energy consumption. Finally, the edge server schedules computing resources according to the received task characteristics, completes task processing, and returns the results to the mobile device.
[0004] Task Offloading Decision and Resource Allocation Optimization (TOTPA) is the core research problem of the MEC system. The optimization goal is to minimize the task processing cost by jointly optimizing the task offloading decision and transmission power allocation, subject to the constraints of latency, energy consumption, and task queue stability. Specifically, the task processing cost is usually defined as the weighted sum of latency and energy consumption, and the constraints include that the task completion time is less than the maximum tolerable delay, the computing and transmission energy consumption of the mobile device do not exceed the budget, and the stability of the task queue.
[0005] For the TOTPA problem, existing methods are mainly divided into two categories. The first category is the method based on traditional mathematical programming, which realizes TOTPA by establishing a mathematical model and solving the optimization problem. For example, in "Joint task offload-ing andcaching for massive mimo-aided multi-tier computing networks", a non-convex power allocation problem is transformed into a linear optimization problem and the Lagrangian partial relaxation method is adopted to relax the task offloading and caching constraints, and a dual problem is proposed; "Task offloading in hybrid intelligent reflecting surfaceand massive mimo relay networks" designs an alternating optimization algorithm based on the difference convex optimization framework to optimize the power allocation. However, with the expansion of the MEC network scale, the computational complexity of this kind of method is high and the real-time performance is poor, making it difficult to adapt to the dynamic and complex multi-task environment.
[0006] The second category is the method based on deep reinforcement learning (DRL). Through interacting with the environment, DRL can learn an approximately optimal TOTPA policy, which is especially suitable for complex task scenarios. For example, the method in "Deep q-network based resource allocation for uav-assisted ultra-dense networks" optimizes the transmission power allocation of a single agent based on the deep Q-network (DQN), and the method in "Multi-agent drl for task offloading and resource allocation in multi-uav enabled iot edge network" optimizes the long-term task cost in the multi-MD cooperation scenario based on multi-agent deep reinforcement learning (such as MADDPG). However, these methods still have limitations in practical applications. On the one hand, due to the lack of information sharing among agents, the multi-agent method cannot achieve the collaborative optimization of global task offloading and resource allocation; on the other hand, even if the multi-agent method realizes information sharing through the communication mechanism, the existing communication schemes have significant deficiencies in privacy protection.
[0007] In the problem of multi-agent task offloading and resource allocation optimization, communication efficiency and privacy protection are the core challenges affecting the collaborative performance of multi-agents. Task offloading requires agents to share task status, resource usage information, and offloading decisions in real time, while resource allocation depends on the efficient transmission and accurate processing of global information. However, the current communication and privacy protection mechanisms have obvious deficiencies in these aspects.
[0008] In the TOTPA method based on multi-agent reinforcement learning, agents usually need to share key information such as task status, channel quality, and resource usage, and this information may contain sensitive data (such as task queue length, user location information, transmission power, etc.). Current communication-assisted methods such as "Learning multiagent communication with backpropagation" and "Targeted multi-agent communication" respectively propose a differentiable communication framework, which supports end-to-end communication learning among multi-agents and uses the attention mechanism to extract useful information from communication messages to enhance communication efficiency. However, these methods do not protect the privacy of shared data, and sensitive information is easily leaked during transmission. Especially in the multi-task competition scenario, the risk of privacy leakage is further aggravated. This communication mechanism lacking privacy protection not only limits the applicability of the method, but also may have a negative impact on the accuracy of task offloading and resource allocation decisions.
[0009] Regarding the privacy protection issue, the current solution "Privacy-preserving q-learning with functional noise in continuous spaces" reduces the risk of sensitive information leakage by introducing random perturbations into the reward signal. However, this method has poor applicability in multi-task scenarios. Especially in the case where task offloading and resource allocation require real-time decisions, random perturbations may lead to delays in offloading tasks and biases in resource allocation. For example, task offloading depends on accurate task queue lengths and network transmission states. However, due to the introduction of noise, this information may be distorted, thereby affecting the effectiveness of offloading decisions and increasing energy consumption or latency.
[0010] In addition, although the existing centralized training and decentralized execution framework allows for global information sharing during the training phase, during the execution phase, due to relying only on the local observations of agents, it is difficult to achieve global optimal offloading and resource allocation in multi-task competition scenarios. For example, when multiple agents simultaneously compete for limited transmission bandwidth or edge server resources, the local observation limitation causes some agents to be overloaded while the resource utilization of other agents is relatively low, ultimately leading to unbalanced or conflicting resource allocation and reducing the overall efficiency of task offloading and resource allocation in the system. Summary of the Invention
[0011] In view of the deficiencies of the prior art, the present invention proposes a multi-agent task offloading and resource optimization method and system based on differential privacy protection, aiming to solve the problems of privacy leakage, low communication efficiency, and insufficient global optimization ability in the prior art, and improve the collaboration performance and global optimization effect of the multi-agent system.
[0012] The first aspect of the present invention provides a multi-agent task offloading and resource optimization method based on differential privacy, including the following steps:
[0013] Define the mobile edge computing environment, and construct the mobility model of mobile devices, the task model of mobile devices, the propagation model between the base station and mobile devices, the communication model of mobile devices, the local computing model of mobile devices, the edge computing model of the base station, and the computing task cost model in the mobile edge computing environment;
[0014] Based on the mobility model of mobile devices, the task model of mobile devices, the propagation model between the base station and mobile devices, the communication model of mobile devices, the local computing model of mobile devices, the edge computing model of the base station, and the computing task cost model in the mobile edge computing environment, establish the objective function and its constraints for task offloading and resource allocation;
[0015] Each mobile device in the mobile edge computing environment is regarded as an agent, and the task offloading and resource allocation problem is formulated as a Markov decision process and this Markov decision process is defined;
[0016] According to the defined Markov decision process and the objective function and its constraints of task offloading and resource allocation, the decision-making strategy of the agent is optimized to obtain the trained agent;
[0017] Each trained agent outputs an action according to the state of the current agent by using its own Actor network, and the mobile device performs task offloading and resource allocation according to the output action.
[0018] Furthermore, the mobile edge computing environment includes several base stations and N mobile devices, and the set of mobile devices is represented as MD s ={1, 2, …, n, …, N}, where n is a mobile device and N is the number of mobile devices. Each mobile device randomly moves within the range set by one of the base stations. Each mobile device is equipped with a computing unit, a transmitting antenna, and a battery. The computing frequency of the computing unit is f MD , the maximum transmitting power of the transmitting antenna is P MD , and the energy budget of the battery is E MD ; The base stations are deployed with edge servers having a computing frequency of f ES . Time is evenly divided into T discrete time slots, and the set of time slots is represented as {1, 2, …, t, …, T}, where the length of each time slot t ∈ T is ζ seconds; It is set that each mobile device generates a computing task at each time slot, and a first-in-first-out computing task queue is deployed in the mobile device and the base station.
[0019] Furthermore, the movement model of the mobile device is:
[0020]
[0021] where is the moving speed of mobile device n at time slot t + 1, is the moving speed of mobile device n at time slot t, Δv is the speed change amount of the mobile device within time slot t, V max is the maximum moving speed of the mobile device, is the moving direction of the mobile device at time slot t + 1, is the moving direction of the mobile device at time slot t, Δθ is the direction change amount of the mobile device within time slot t, is the abscissa of mobile device n at time slot t, is the abscissa of mobile device n at time slot t + 1, is the ordinate of mobile device n at time slot t, is the vertical coordinate of mobile device n at time slot t + 1;
[0022] The task model of the mobile device is:
[0023]
[0024] where, represents the computing task generated by mobile device n at time slot t, represents the data size of the computing task generated by mobile device n at time slot t, ω represents the number of CPU cycles required to process each bit of data, and Γ represents the maximum tolerable delay of the computing task;
[0025] The propagation model between the base station and the mobile device is:
[0026]
[0027] where, represents the channel gain between mobile device n and the base station at time slot t, represents small-scale Rayleigh fading, represents large-scale fading, which includes geometric fading and shadow fading:
[0028]
[0029] where, represents the distance between the base station and mobile device n at time slot t, z is a log-normal random variable representing shadow fading;
[0030] The communication model of the mobile device is:
[0031]
[0032] where, represents the signal-to-interference-plus-noise ratio of mobile device n at time slot t, represents the transmission power of mobile device n at time slot t, represents whether the computing task generated by mobile device n at time slot t is executed locally or offloaded to the edge server, represents that the computing task is offloaded to the edge server for execution, represents that the computing task is executed locally; σ 2 is the noise power; represents the data transmission rate of mobile device n at time slot t, and W is the bandwidth of the wireless channel;
[0033] The local computing model of the mobile device is:
[0034]
[0035]
[0036] ξ = 10 -27 (f MD ) 2 (13)
[0037] Among them, represents the local computing delay when the computing task generated by mobile device n at time slot t is executed locally, ζ represents the length of each time slot, represents the number of computing tasks in the local computing task queue of mobile device n at time slot t - 1, represents the number of computing tasks in the local computing task queue of mobile device n at time slot t, represents the energy consumption of local computing of mobile device n at time slot t, and ξ is the computing energy consumption coefficient;
[0038] When the computing task is offloaded to the edge server for computing, the edge computing model of the base station is:
[0039]
[0040] Among them, n’ represents a mobile device different from mobile device n, represents the transmission delay when the computing task is offloaded from mobile device n to the base station at time slot t, represents the energy consumption when the computing task is offloaded from mobile device n to the base station at time slot t; represents the computing delay of the computing task generated by mobile device n at time slot t on the edge server, represents the number of computing tasks in the computing task queue on the edge server at time slot t;
[0041] The computing task cost model is:
[0042]
[0043] Among them, is the computing task cost; ω t and ω e are the weight coefficients of time delay and energy consumption respectively.
[0044] Furthermore, the objective function of task offloading and resource allocation is:
[0045]
[0046] Among them, P represents the objective function, C avg represents the weighted average cost of time delay and energy consumption of the computing task;
[0047] The constraints of the objective function of task offloading and resource allocation are:
[0048]
[0049]
[0050] Among them, s.t. represents the constraint condition, sup represents the supremum, T represents the total time period, E represents the mathematical expectation, is the available energy budget of mobile device n at the end of time slot t, and is the available energy budget of mobile device n at the end of time slot t - 1.
[0051] Furthermore, formulating the task offloading and resource allocation problem as a Markov decision process and defining the specific Markov decision process is as follows:
[0052] Define the global state as S t , including the states of all agents. The state of each agent in the global state is s t , including the channel gain between the mobile device and the base station, the computing tasks generated by the mobile device in each time slot, the available energy budget of the mobile device, the number of computing tasks in the local computing task queue of the mobile device, the bandwidth of the wireless channel, the number of computing tasks in the task queue on the edge server, and the computing frequency of the edge server. Therefore, the state of each agent at time slot t is expressed as:
[0053]
[0054] Define the joint action A t , including the actions of all agents. The action of each agent in the joint action is a t , and the action defines the decision that the agent needs to execute in each time slot, including the task offloading decision and the transmission power allocation. The task offloading decision determines whether the task is processed locally or offloaded to the edge server for processing. Therefore, the action of each agent at time slot t is expressed as:
[0055]
[0056] Define the reward function to measure the quality of each action. Since the goal is to minimize the task delay and energy consumption cost, the reward function is:
[0057]
[0058] The negative sign of the reward function indicates that the optimization goal is to maximize the reward, that is, to minimize the total task cost;
[0059] Define the state transition probability P(s t+1 |st , a t ), the state transition probability describes that after the agent executes the action a t , from the current state s t transfers to the next state s t .
[0060] Further, optimizing the decision-making strategy of the agent according to the defined Markov decision process and the objective function and its constraints of task offloading and resource allocation to obtain the trained agent is specifically as follows:
[0061] S1: For each agent, randomly initialize the weights θ Q of the Critic network Q(s, a|θ Q ) and the weights θ μ of the Actor network μ(s|θ μ ), where s represents the state of the agent and a represents the action executed by the agent in the current state;
[0062] S2: Initialize two target networks Q'(s, a) and μ'(s), and make the weights of the two target networks equal to the weights of the Critic network and the Actor network respectively, that is, θ Q‘ ← θ Q and θ μ‘ ← θ μ , where θ Q‘ is the weight of the target network Q'(s, a), and θ μ‘ is the weight of the target network μ'(s);
[0063] S3: Initialize the experience replay pool R for storing the experience data (s t , a t , r t , s t+1 ) of environmental interaction, where r t is the reward function;
[0064] S4: Initialize the random process M for action exploration;
[0065] S5: Set the total number of cycles Ep of loop training, the maximum time slot T in each cycle of loop training, the current number of cycles of loop training e = 0, and the current time slot t = 0;
[0066] S6: At each time slot t, obtain the current state s t of the agent and input it into the Actor network to generate the original action;
[0067] a t = μ(s t |θ μ ) (29)
[0068] S7: Incorporate a privacy protection mechanism and inject Gaussian noise into action a using a message encoding function t , generate a protected action and send it to other agents;
[0069] a t ′ = μ(s t |θ μ ) + N t (30)
[0070] where a t ′ represents the protected action and N t represents Gaussian noise;
[0071] The variance of the Gaussian noise is:
[0072]
[0073] where represents the variance of the Gaussian noise, γ 1 and γ 2 are the sampling rates, C is the L 2 norm of the message encoding function, δ is used to control the noise amplitude; ε n is the privacy budget, used to control the degree of privacy protection; β represents a trade-off parameter;
[0074] S8: The current agent receives the protected actions sent by other agents and decodes the protected actions to obtain the decoded actions of other agents;
[0075] S8.1: The current agent receives a set of noisy messages m (-n)n , which includes the protected actions sent by other agents;
[0076] S8.2: Decode the received set of noisy messages using a decoding function to obtain a decoded set of messages, which includes the decoded actions of other agents;
[0077] q n = f n (m (-n)n ) (32)
[0078] where q n represents the decoded set of messages and f n represents the decoding function;
[0079] S8.3: Optimize the parameters of the decoding function using the gradient chain rule;
[0080] The optimization formula of the gradient chain rule is:
[0081]
[0082] Among them, is the weight of the Actor network of agent n, is the weight of the Actor network of the gradient, that is, the direction of policy update; is the expected value of the reward function; τ is the trajectory; A represents the joint action; E τ,s,A is the expectation under the trajectory τ, state s and joint action A; π n represents the policy of the Actor network of agent n, represents the expectation under the policy π n ; f n (q n |m (-n)n ) is the decoding function; is the partial derivative of q n ; π n (a n |s,q n ) The policy for the agent to select an action in the state s; Q π (A,s) is the action-state value function, and the expected return when the joint action A and state s are given under the policy π;
[0083] S9: Each agent executes the protected action a t , obtains the environmental feedback, including the reward r t and the next state s t+1 , and stores the experience data (s t ,a t ,r t ,s t+1 ) into the experience replay pool;
[0084] S10: Update the time slot t = t + 1;
[0085] S11: Determine whether t < T is satisfied. If so, execute S6; otherwise, execute S12;
[0086] S12: Update the number of cycles e of the loop training to e = e + 1;
[0087] S13; Randomly sample L pieces of experience data from the experience pool;
[0088] S14: Update the Critic network using the randomly sampled L pieces of experience data;
[0089] S15: Update the Actor network using the randomly sampled N pieces of experience data;
[0090] S16: Synchronize the weights of the two target networks using the soft update method;
[0091] S17: Determine if e < Ep. If yes, return to S4; otherwise, the training ends and the trained agent is obtained.
[0092] The second aspect of the present invention provides a multi-agent task offloading and resource optimization system based on differential privacy for implementing the multi-agent task offloading and resource optimization method based on differential privacy, including:
[0093] Mobile devices, used to generate computing tasks and perform task offloading and resource allocation according to the task offloading and resource allocation strategy obtained by the multi-agent task offloading and resource allocation module;
[0094] Base stations, used to execute the computing tasks generated by mobile devices using the deployed edge servers;
[0095] The multi-agent task offloading and resource allocation module regards each mobile device as an agent. The agent uses its own Actor network to output actions based on the current state to obtain the task offloading and resource allocation strategy.
[0096] The third aspect of the present invention provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the multi-agent task offloading and resource optimization method based on differential privacy are executed.
[0097] The fourth aspect of the present invention provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is run by a processor, the steps of the multi-agent task offloading and resource optimization method based on differential privacy are executed.
[0098] Compared with the prior art, the beneficial effects of the present invention are:
[0099] In view of the characteristics that task offloading and resource allocation scenarios in multi-agent systems require frequent sharing of task status, queue information, and resource usage, the present invention proposes a differential privacy protection mechanism based on Gaussian noise. Since task information usually involves sensitive data of devices (such as task queue length, transmission power, etc.), direct transmission is likely to lead to privacy leakage. The present invention ensures that sensitive data is not directly leaked during the collaboration process by introducing a differential privacy protection mechanism in message transmission. The sender uses an optimized Gaussian noise addition strategy to dynamically adjust the noise intensity using the privacy budget parameter, achieving a balance between privacy protection and information effectiveness, ensuring the effectiveness of key data while guaranteeing privacy security, thereby supporting collaborative optimization of task offloading and resource allocation decisions among agents. The receiver uses the gradient chain rule to parse and recover the noise-added information, minimizing the impact of noise on optimizing task offloading and resource allocation decisions, so as to ensure that agents make optimal decisions based on complete and effective information.
[0100] In addition, in response to the problem of insufficient local observations of agents during the execution phase, the present patent proposes a multi-agent real-time communication mechanism, breaking through the local observation limitations in the execution phase of traditional multi-agent systems. In traditional methods, due to local observation limitations, agents often have difficulty grasping global information, resulting in unbalanced resource allocation and low collaboration efficiency. Through the information sharing mechanism among agents, the receiver can integrate the global task queue, resource status, and channel conditions in real time, thereby optimizing the collaborative effect of task offloading and resource allocation, enhancing the global optimization effect in task offloading and resource allocation decisions, avoiding problems of over-concentration or shortage of resources, and significantly improving the overall task processing efficiency of the system.
[0101] Through the above improvements, the present invention realizes the organic combination of privacy protection and communication, providing an efficient and reliable solution for the global optimization of multi-agent task offloading and resource allocation. While enhancing the privacy security of communication, the present invention effectively balances privacy protection and global optimization performance, and is applicable to complex and dynamic multi-task scenarios. Brief Description of the Drawings
[0102] Figure 1 It is a schematic diagram of the system network architecture for multi-agent task offloading and resource allocation in an embodiment of the present invention;
[0103] Figure 2 It is a schematic diagram of the optimization process of the Actor-Critic network in an embodiment of the present invention;
[0104] Figure 3 It is a schematic diagram of the differential privacy communication structure in an embodiment of the present invention;
[0105] Figure 4 It is a graph of the average cost results with different privacy budgets in an embodiment of the present invention;
[0106] Figure 5 This is the average reward result graph with different privacy budgets in the embodiments of the present invention. Detailed implementation manners
[0107] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.
[0108] This embodiment provides a multi-agent task offloading and resource optimization method based on differential privacy, including the following steps:
[0109] Step 1: Define a mobile edge computing (MEC) environment, and construct a mobility model of mobile devices, a task model of mobile devices, a propagation model between a base station and mobile devices, a communication model of mobile devices, a local computing model of mobile devices, an edge computing model of the base station, and a computing task cost model in the mobile edge computing (MEC) environment;
[0110] As Figure 1 shown, the present invention considers a typical mobile edge computing (MEC) environment composed of three base stations and N mobile devices (MDs). The set of mobile devices is denoted as MD s = {1, 2,..., n,..., N +, where n is a mobile device and N is the number of mobile devices. Each mobile device randomly moves around one of the base stations (BSs). Each mobile device is equipped with a computing unit, a transmitting antenna, and a battery. The computing frequency of the computing unit is f MD , the maximum transmitting power P of the transmitting antenna MD , and the energy budget E of the battery MD ; The base station deploys an edge server (ES) with a computing frequency f ES which is usually more powerful than the local computing unit of each mobile device; Time is uniformly divided into T discrete time slots, and the set of time slots is denoted as {1, 2,..., t,..., T +. The length of each time slot t ∈ T is ζ seconds; It is assumed that each mobile device generates a computing task at each time slot. In the case of computing task failure and discard under heavy load, a first-in-first-out (FIFO) computing task queue is deployed in the mobile device and the base station;
[0111] The mobility model of the mobile device represents the moving speed, moving direction, and post-movement position of the mobile device around the base station in each time slot. The moving speed and moving direction of the mobile device follow the Haas mobility model, which is expressed as:
[0112]
[0113] where is the moving speed of mobile device n at time slot t + 1, Let \(v_{n}(t)\) be the moving speed of mobile device \(n\) at time slot \(t\), \(\Delta v\) be the speed change of the mobile device within time slot \(t\), and \(V\) max be the maximum moving speed of the mobile device, \(\theta_{n}(t + 1)\) be the moving direction of the mobile device at time slot \(t+1\), \(\theta_{n}(t)\) be the moving direction of the mobile device at time slot \(t\), and \(\Delta\theta\) be the direction change of the mobile device within time slot \(t\), \(x_{n}(t)\) be the abscissa of mobile device \(n\) at time slot \(t\), \(x_{n}(t + 1)\) be the abscissa of mobile device \(n\) at time slot \(t+1\), \(y_{n}(t)\) be the ordinate of mobile device \(n\) at time slot \(t\), \(y_{n}(t + 1)\) be the ordinate of mobile device \(n\) at time slot \(t+1\);
[0114] The task model of the mobile device represents the computing tasks generated by the mobile device in each time slot:
[0115]
[0116] where \(T_{n}(t)\) represents the computing tasks generated by mobile device \(n\) at time slot \(t\), \(L_{n}(t)\) represents the data size of the computing tasks generated by mobile device \(n\) at time slot \(t\), \(\omega\) represents the number of CPU cycles required to process each bit of data, and \(\Gamma\) represents the maximum tolerable delay of the computing tasks;
[0117] The propagation model between the base station and the mobile device represents the channel gain between the base station and the mobile device:
[0118]
[0119] where \(h_{n}(t)\) represents the channel gain between mobile device \(n\) and the base station at time slot \(t\), \(h_{f}\) represents small-scale Rayleigh fading, \(h_{l}\) represents large-scale fading, which includes geometric fading and shadow fading:
[0120]
[0121] where \(d_{n}(t)\) represents the distance between the base station and mobile device \(n\) at time slot \(t\), and \(z\) is a log-normal random variable representing shadow fading;
[0122] The communication model of the mobile device is:
[0123]
[0124] where denotes the signal-to-interference-plus-noise ratio (SINR) of mobile device n at time slot t, denotes the transmission power of mobile device n at time slot t, represents whether the computing task generated by mobile device n at time slot t is executed locally or offloaded to the edge server, denotes that the computing task is offloaded to the edge server for execution, denotes that the computing task is executed locally; σ 2 is the noise power; denotes the data transmission rate of mobile device n at time slot t, where W is the bandwidth of the wireless channel;
[0125] The local computing model of the mobile device is:
[0126]
[0127] where, denotes the local computing delay when the computing task generated by mobile device n at time slot t is executed locally ζ represents the length of each time slot, denotes the number of computing tasks in the local computing task queue of mobile device n at time slot t-1, denotes the number of computing tasks in the local computing task queue of mobile device n at time slot t, denotes the energy consumption of the local computing of mobile device n at time slot t, where ξ is the computing energy consumption coefficient;
[0128] When the computing task is offloaded to the edge server for computing The edge computing model of the base station is:
[0129]
[0130] where n' represents a mobile device different from mobile device n, denotes the transmission delay when the computing task is offloaded from mobile device n to the base station at time slot t, denotes the energy consumption when the computing task is offloaded from mobile device n to the base station at time slot t; denotes the computing delay of the computing task generated by mobile device n at time slot t on the edge server ES, denotes the number of computing tasks in the computing task queue on the edge server at time slot t. Since the amount of returned result data is relatively small compared to the uploaded data volume, the return delay and energy consumption are ignored;
[0131] The computing task cost model is:
[0132]
[0133] Among them, for calculating the task cost, which is the weighted sum of energy consumption and latency; ω t and ω e are the weight coefficients of latency and energy consumption respectively;
[0134] Step 2: According to the mobility model of mobile devices, the task model of mobile devices, the propagation model between the base station and mobile devices, the communication model of mobile devices, the local computing model of mobile devices, the edge computing model of the base station, and the computing task cost model in the mobile edge computing (MEC) environment, establish the objective function and its constraints for task offloading and resource allocation;
[0135] The objective function for task offloading and resource allocation is:
[0136]
[0137] Among them, P represents the objective function, and C avg represents the weighted average cost of latency and energy consumption of the computing task;
[0138] The constraints of the objective function for task offloading and resource allocation are:
[0139]
[0140] Among them, s.t. represents the constraint condition, sup represents the supremum, T represents the total time period, E represents the mathematical expectation, is the available energy budget of mobile device n at the end of time slot t, and the available energy budget of mobile device n at the end of time slot t - 1; Equation (20) constrains that each computing task is either completely processed locally or completely offloaded to the edge server ES for processing; (21) constrains that the transmission power of each mobile device shall not exceed its maximum power limit; (22) constrains that at any time, the remaining energy of the mobile device must be non - negative; (23) constrains that the total delay of the task (including the local or offloaded processing latency) must be less than the allowed maximum latency Γ; (24) constrains that the expected value of the computing task queue length of the mobile device remains finite (i.e., the stability of the mobile device task queue); (25) constrains that the expected value of the computing task queue length of the base station should also remain finite (i.e., the stability of the base station task queue);
[0141] It can be seen that the objective function (19) is a non-convex function. In theory, the optimal solution to the problem can be obtained by using the exhaustive method, but its high complexity is unacceptable in reality. Therefore, next, the present invention will adopt the multi-agent reinforcement learning (MADRL) method to solve the above joint optimization problem, and use the multi-agent deep deterministic policy gradient (MADDPG) algorithm to seek an approximate optimal solution;
[0142] Step 3: Regard each mobile device in the mobile edge computing (MEC) environment as an agent, formulate the task offloading and resource allocation problem as a Markov decision process, and define this Markov decision process;
[0143] The present invention models the joint optimization problem of agent task offloading and resource allocation as a partially observable Markov decision process, and proposes a multi-agent reinforcement learning algorithm based on differential privacy communication to solve it. Each mobile device is regarded as an agent, and they will make decisions based on the surrounding environmental state and the information of other mobile devices, and continuously learn and update the policy. The entire learning process is carried out using a centralized training and distributed execution architecture.
[0144] The Markov decision process is a discrete-time stochastic process used to describe the probability distribution of future decisions and state transitions of a system given the current state and the observed event sequence. It is modeled based on the Markov chain, where each state is related to its previous state and the currently observed information. In the Markov decision process, the decision maker faces a series of decision problems, each of which is based on the current state and past observation results. In the above model, the agent is the decision maker, and each decision it makes is related to the state of the observed environment. However, the agent is limited by its own physical conditions, and the environment it can observe is limited. Therefore, the present invention models the above joint optimization problem as a partially observable Markov decision process;
[0145] Define the global state as S t , which includes the states of all agents. The state of each agent in the global state is s t , which includes the channel gain between the mobile device and the base station, the computing tasks generated by the mobile device in each time slot, the available energy budget of the mobile device, the number of computing tasks in the local computing task queue of the mobile device, the bandwidth of the wireless channel, the number of computing tasks in the task queue on the edge server, and the computing frequency of the edge server. Therefore, the state of each agent at time slot t is expressed as:
[0146]
[0147] Define the joint action A t, including the actions of all agents, and the action of each agent in the joint action is a t , the action defines the decision that the agent needs to execute in each time slot, including the task offloading decision and the transmission power allocation. The task offloading decision determines whether the task is processed locally or offloaded to the edge server for processing. Therefore, the action of each agent at time slot t is expressed as:
[0148]
[0149] Define the reward function used to measure the quality of each action. Since the goal is to minimize the task latency and energy consumption cost, the reward function is:
[0150]
[0151] The negative sign of the reward function indicates that the optimization goal is to maximize the reward, that is, to minimize the total task cost;
[0152] Define the state transition probability P(s t+1 |s t ,a t ), the state transition probability describes the probability that the agent transfers from the current state s t to the next state s t after executing the action a t . The main influencing factors include: 1) Computation task queue update: If the computation task is offloaded, the length of the computation task queue at the base station becomes longer. If the computation task is processed locally, the length of the computation task queue at the mobile device decreases; 2) Channel dynamics: The channel gain changes as the mobile device MD moves; 3) Resource consumption: The allocation of transmission power affects the energy state of the mobile device MD;
[0153] Step 4: According to the defined Markov decision process and the objective function and its constraints of task offloading and resource allocation, optimize the decision-making strategy of the agent to obtain the trained agent;
[0154] In the training stage, the present invention adopts an experience replay mechanism (Replay Buffer) to improve the efficiency of learning the optimal strategy. The observation data of all agents are stored in an experience replay pool (replay buffer D), which is located on the macro base station (MBS). Each observation data contains an information tuple of state, action, reward, and next state. During each training process, the present invention updates the neural network by randomly sampling a small batch of observation data from the experience replay pool. In this way, the correlation between data is reduced and the training efficiency is improved.
[0155] By defining the above Markov decision process (MDP), the problem is transformed into optimizing the policy π(a|s) to maximize the long-term cumulative reward where γ ∈ (0, 1) is the discount factor, and R t is the system revenue obtained after adopting a specific offloading and resource allocation strategy in time slot t, that is, the immediate reward obtained in time slot t, so as to find the optimal solution in the process of task offloading and resource allocation. In this step, the policy optimization method of reinforcement learning is used. By continuously updating the Actor network and the Critic network, the optimal policy is gradually approximated, and finally the efficient decision-making of task offloading and resource allocation in the multi-agent system is realized;
[0156] As Figure 2 shown, it includes the following steps:
[0157] Step 4.1: For each agent, randomly initialize the weights θ Q of the Critic network Q(s,a|θ Q ) and the weights θ μ of the Actor network μ(s|θ μ ), where s represents the state of the agent, such as the task queue length, computing resources, channel state, etc., and a represents the action executed by the agent in the current state;
[0158] Step 4.2: Initialize two target networks Q'(s,a) and μ'(s), and make the weights of the two target networks equal to the weights of the Critic network and the Actor network respectively, that is, θ Q‘ ←θ Q and θ μ‘ ←θ μ , where θ Q‘ is the weight of the target network Q'(s,a), and θ μ‘ is the weight of the target network μ'(s);
[0159] Step 4.3: Initialize the experience replay pool R for storing the experience data (s t ,a t ,r t ,s t+1 ) of environmental interaction, where r t is the reward function;
[0160] Step 4.4: Initialize the random process Μ for action exploration;
[0161] Step 4.5: Set the number of epochs Ep for the total loop training, the maximum time slot T in each loop training epoch, the current loop training epoch number e = 0, and the current time slot t = 0;
[0162] Step 4.6: At each time slot t, obtain the current state s of the agentt and input it into the Actor network to generate the original action;
[0163] a t = μ(s t |θ μ ) (29)
[0164] Step 4.7: As shown in Figure 3 , add a privacy protection mechanism, and use the message encoding function to inject Gaussian noise into the action a t , generate the protected action and send it to other agents;
[0165] a t ' = μ(s t |θ μ ) + N t (30)
[0166] where a t ' represents the protected action, and N t represents Gaussian noise;
[0167] To effectively protect the sensitive information between multiple agents during the communication process and ensure the effectiveness of task offloading and resource allocation decisions, the present invention introduces a privacy protection mechanism based on Gaussian noise. Specifically, when generating a message, the message sender combines the key task information with Gaussian noise, and by dynamically adjusting the noise intensity (based on the privacy budget parameter), it protects privacy while retaining the information related to task offloading and resource allocation decisions. Add Gaussian noise to the generated message, and the Gaussian noise follows where is the noise variance, which is determined by the privacy budget parameter and, and the variance of the Gaussian noise is:
[0168]
[0169] where represents the variance of the Gaussian noise, γ 1 and γ 2 are the sampling rates, C is the L 2 norm of the message encoding function, δ is a measure of the maximum impact of the input data on the output message, used to control the noise amplitude to ensure privacy protection while maintaining the effectiveness of the message; ε n is the privacy budget, used to control the degree of privacy protection, and ε nThe smaller it is, the stronger the protection, to ensure the effective protection of the privacy of sensitive data and avoid leakage, but the intensity (variance) of the noise will also increase; according to the adjustment of the privacy budget, the noise intensity can balance between privacy protection and information effectiveness to ensure the privacy of sensitive data; β represents a trade-off parameter used to balance the norm of the message encoding function and the influence of the noise;
[0170] Step 4.8: The current agent receives the protected actions sent by other agents and decodes the protected actions to obtain the decoded actions of other agents;
[0171] Step 4.8.1: As Figure 3 shown, the current agent receives a set of noisy messages m (-n)n , which includes the protected actions sent by other agents;
[0172] Step 4.8.2: Use the decoding function to decode the received set of noisy messages to obtain a set of decoded messages, which includes the decoded actions of other agents;
[0173] q n = f n (m (-n)n ) (32)
[0174] where q n represents a set of decoded messages, and f n represents the decoding function;
[0175] Step 4.8.3: Use the gradient chain rule to optimize the parameters of the decoding function;
[0176] The optimization formula of the gradient chain rule is:
[0177]
[0178] where is the weight of the Actor network of agent n, is the gradient of the weight of the Actor network, that is, the direction of policy update; is the expected value of the reward function; τ is the trajectory, which refers to a sequence of states, actions, rewards, and next states generated by the agent during the interaction with the environment; A represents the joint action; E τ,s,A is the expectation under the trajectory τ, state s, and joint action A; π n represents the policy of the Actor network of agent n, represents the expectation under the policy π n ; f n (q n |m (-n)n) is the decoding function; is for q n to take the partial derivative; π n (a n |s, q n ) The policy for the agent to select an action in state s; Q π (A, s) is the action-state value function, which is the expected return given the joint action A and state s under the policy π;
[0179] The message receiver optimizes the noisy message based on the gradient chain rule. By reconstructing key data and noise reduction strategies, it attempts to restore the effectiveness of the task state information as much as possible, thereby improving communication efficiency and multi-agent collaboration performance.
[0180] Step 4.9: Each agent executes the protected action a t , and obtains environmental feedback, including the reward r t and the next state s t+1 , and stores the experience data (s t , a t , r t , s t+1 ) into the experience replay pool;
[0181] Step 4.10: Update the time slot t = t + 1;
[0182] Step 4.11: Determine whether t < T is satisfied. If so, execute Step 4.6; otherwise, execute Step 4.12;
[0183] Step 4.12: Update the number of cycles e of the loop training by e = e + 1;
[0184] Step 4.13; Randomly sample L pieces of experience data from the experience pool;
[0185] Step 4.14: Update the Critic network using the randomly sampled L pieces of experience data;
[0186] Calculate the target value using the target network:
[0187] y n = r t + γQ′(s t+1 , μ′(s t+1 |θ μ′ )|θ Q′ ) (34)
[0188] where y n represents the target value, which is the target estimated based on the current reward r t and the Q value calculated through the target network; r t represents the reward obtained when executing the action a in the state s t ;t The reward obtained by the subsequent agent; γ represents the discount factor; Q′(s t+1 , μ′(s t+1 |θ μ′ )|θ Q′ ) represents the Q-value estimation of the Critic network for the next state s t+1 and the action a obtained through the Actor network t+1 ;
[0189] The Critic network is updated by minimizing the loss function of the Critic network:
[0190]
[0191] where L represents the loss function of the Critic network, and represents the squared difference between the estimated target value y n and the Q-value output by the current Critic network; Q(s t , a t |θ Q ) represents the Q-value output by the Critic network when taking the action a in the state s t ; t ;
[0192] Step 4.15: Update the Actor network using N pieces of empirically sampled data;
[0193] Optimize the policy of the Actor network using the gradient chain rule:
[0194]
[0195] where represents the derivative of the loss function J with respect to the weights θ μ of the Actor network, that is, the gradient of the Actor-Critic network parameters; represents the gradient of the Q-value function of the Critic network with respect to the action a, indicating the rate of change of the value of the action evaluated by the Critic with respect to the action; μ(s t ) and represent the action output by the Actor network in the state s t ; represents the gradient of the policy function μ of the Actor network with respect to the parameter θ t in the state s μ ; μ(s|θ μ ) represents the action output by the Actor network in the state s; is the noise compensation term. During the optimization process, the receiving end analyzes the message with added noise and uses the statistical information of the noise (such as the mean and variance) to recover the key information, further improving the effect of policy update;
[0196] Step 4.16: Synchronize the weights of the two target networks in a soft update manner;
[0197] θ Q’ ′←τθ Q +(1 - τ)θ Q′ (37)
[0198] θ μ’ ′←τθ μ +(1 - τ)θ μ′ (38)
[0199] where, θ Q' ′ and θμ ' ′ represent the weights of the updated target network, τ represents the weight coefficient of soft update, τ∈(0, 1);
[0200] Step 4.17: Judge whether e < Ep. If yes, return to Step 4.4; otherwise, the training ends and the trained agent is obtained.
[0201] Step 5: Each trained agent outputs an action according to the current state of the agent using its own Actor network, and the mobile device performs task offloading and resource allocation according to the output action.
[0202] To verify the effectiveness of the multi - agent task offloading and resource allocation optimization method based on differential privacy protection proposed in the present invention, multiple groups of experiments were carried out and the results were compared with the standard MADDPG algorithm. The experimental results show that the method of the present invention can significantly reduce the task cost while ensuring privacy security, and improve the reward value composed of the weighted sum of delay and energy consumption. This indicates that the present invention not only effectively optimizes the task cost, but also further improves the task offloading efficiency and the rationality of resource allocation, fully meeting the performance requirements of the multi - agent cooperation system.
[0203] From Figure 4 the average task cost result graph with different privacy budgets, it can be seen that the average task cost shows regular changes with the change of the privacy budget. When the privacy budget is small (such as ε = 0.05), due to the small intensity of the added Gaussian noise, the task cost is high, and the average task cost at stability is about 2.5, but it is still lower than the cost of the standard MADDPG algorithm. This shows that even under relatively strict privacy protection conditions, the method proposed in the present invention can still maintain a low task cost.
[0204] When the privacy budget is further reduced (ε = 0.05), the average task cost is reduced by about 20%, but at this time, a poor optimization effect is shown, which reflects a certain trade-off between privacy protection and performance optimization. When the privacy budget ε = 0.1, the method proposed in the present invention shows the best performance. At this time, the noise intensity gradually decreases and the optimization effect is significant, indicating that an appropriate privacy budget can balance privacy protection and performance optimization.
[0205] Furthermore, Figure 5 the average reward results in show a significant impact of the privacy budget on the reward performance of policy learning. When the privacy budget is small (such as ε = 0.05), the average reward is low, and the stable value is about -1×10^3, indicating that the noise brought by privacy protection has caused a certain interference to policy learning. However, as the privacy budget increases (such as ε = 0.1), the average reward increases significantly and finally stabilizes at about -3×10^3, which indicates that the present invention can effectively enhance the system performance while protecting privacy. When the privacy budget ε = 0.2, the average reward begins to decrease, which further confirms the need to find a balance between privacy protection and system performance.
[0206] In summary, the experimental results fully prove that the method proposed in the present invention can achieve a dynamic balance between privacy protection and optimized performance by flexibly adjusting the privacy budget parameter ε while protecting privacy. When the privacy budget is small (ε = 0.05), more attention is paid to privacy protection, although this may lead to a slight decrease in performance; while when the privacy budget is appropriate (ε = 0.1), the method proposed in the present invention can achieve the optimal optimization effect and show high flexibility and robustness.
[0207] This embodiment provides a multi-agent task offloading and resource optimization system based on differential privacy for implementing the multi-agent task offloading and resource optimization method based on differential privacy, including:
[0208] Mobile devices, which are used to generate computing tasks and perform task offloading and resource allocation according to the task offloading and resource allocation strategy obtained by the multi-agent task offloading and resource allocation module;
[0209] Base stations, which are used to execute the computing tasks generated by mobile devices by using the deployed edge servers;
[0210] The multi-agent task offloading and resource allocation module regards each mobile device as an agent. The agent uses its own Actor network to output actions according to the current state to obtain the task offloading and resource allocation strategy.
[0211] This embodiment provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the multi-agent task offloading and resource optimization method based on differential privacy are executed.
[0212] This embodiment provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is run by a processor, the steps of the multi-agent task offloading and resource optimization method based on differential privacy as described are executed.
Claims
1. A multi-agent task offloading and resource optimization method based on differential privacy, characterized in that: It includes the following steps: Define a mobile edge computing environment, and construct a mobility model of mobile devices, a task model of mobile devices, a propagation model between a base station and mobile devices, a communication model of mobile devices, a local computing model of mobile devices, an edge computing model of the base station, and a computing task cost model in the mobile edge computing environment; According to the mobility model of mobile devices, the task model of mobile devices, the propagation model between the base station and mobile devices, the communication model of mobile devices, the local computing model of mobile devices, the edge computing model of the base station, and the computing task cost model in the mobile edge computing environment, establish the objective function of task offloading and resource allocation and its constraints; Regard each mobile device in the mobile edge computing environment as an agent, formulate the task offloading and resource allocation problem as a Markov decision process and define the Markov decision process; According to the defined Markov decision process and the objective function of task offloading and resource allocation and its constraints, optimize the decision-making strategy of the agent to obtain a trained agent; Each trained agent uses its own Actor network to output actions according to the current state of the agent, and the mobile device performs task offloading and resource allocation according to the output actions.
2. The multi-agent task offloading and resource optimization method based on differential privacy according to claim 1 is characterized in that: The mobile edge computing environment includes several base stations and N mobile devices, where the set of mobile devices is represented by MD s ={1,2,…,n,…,N}, n is a mobile device, N is the number of mobile devices, each mobile device randomly moves within the range set by one of the base stations, each mobile device is equipped with a computing unit, a transmitting antenna and a battery, and the computing frequency of the computing unit is f MD , the maximum transmitting power P of the transmitting antenna MD , the battery energy budget E MD ; Base station deployment has a calculation frequency f ES edge server; time is evenly divided into T discrete time slots, the set of time slots is represented by {1,2,…,t,…,T}, and the length of each time slot t∈T is ζ seconds; it is set that each mobile device generates a computing task at each time slot, and a first-in-first-out computing task queue is deployed in the mobile devices and base stations.
3. The multi-agent task offloading and resource optimization method based on differential privacy according to claim 1 is characterized in that: The mobility model of the mobile device is: in, is the moving speed of mobile device n at time slot t+1, is the moving speed of mobile device n at time slot t, Δv is the speed change of mobile device in time slot t, V max is the maximum moving speed of the mobile device, is the moving direction of the mobile device at time slot t+1, is the moving direction of the mobile device at time slot t, Δθ is the direction change of the mobile device in time slot t, is the horizontal coordinate of mobile device n at time slot t, is the horizontal coordinate of mobile device n at time slot t+1, is the ordinate of mobile device n at time slot t, is the ordinate of mobile device n at time slot t+1; The task model of the mobile device is: in, represents the computing task generated by mobile device n at time slot t, represents the data size of the computing task generated by mobile device n at time slot t, ω represents the number of CPU cycles required to process each bit of data, and Γ represents the maximum tolerable delay of the computing task; The propagation model between the base station and the mobile device is: in, represents the channel gain between mobile device n and the base station at time slot t, represents the small-scale Rayleigh weakening, Represents large-scale attenuation, which includes both geometric attenuation and shadow attenuation: in, represents the distance between the base station and mobile device n at time slot t, z is a log-normal random variable, representing shadow attenuation; The communication model of the mobile device is: in, represents the signal-to-interference-noise ratio of mobile device n at time slot t, represents the transmission power of mobile device n at time slot t, represents whether the computing task generated by mobile device n at time slot t is executed locally or offloaded to the edge server for execution, Indicates that the computing task is offloaded to the edge server for execution. Indicates that the computing task is executed locally; σ 2 is the noise power; represents the data transmission rate of mobile device n at time slot t, W is the bandwidth of the wireless channel; The local computing model of the mobile device is: ξ=10 -27 (f MD ) 2 (13) in, represents the local computational delay of the computational task generated by mobile device n at time slot t when it is executed locally, ζ represents the length of each time slot, represents the number of computing tasks in the local computing task queue of mobile device n at time slot t-1, represents the number of computing tasks in the local computing task queue of mobile device n at time slot t, represents the energy consumption of local computation of mobile device n at time slot t, ξ is the computation energy consumption coefficient; When the computing task is offloaded to the edge server for computing, the edge computing model of the base station is: Wherein n' represents a mobile device different from mobile device n, represents the transmission delay of the computing task offloaded from mobile device n to the base station at time slot t, represents the energy consumption of offloading the computing task from mobile device n to the base station at time slot t; represents the computational delay of the computational task generated by mobile device n at time slot t on the edge server, represents the number of computing tasks in the computing task queue on the edge server at time slot t; The computing task cost model is: in, To calculate the task cost; t and ω e are the weight coefficients of delay and energy consumption respectively.
4. The multi-agent task offloading and resource optimization method based on differential privacy according to claim 3 is characterized in that: The objective function of task offloading and resource allocation is: Among them, P represents the objective function, C avg Represents the weighted average cost of latency and energy consumption of computing tasks; The constraints of the objective function of task offloading and resource allocation are: Among them, st represents the constraint condition, sup represents the supremum, T represents the total time period, and E represents the mathematical expectation. is the available energy budget of mobile device n at the end of time slot t, and Available energy budget of mobile device n at the end of time slot t-1.
5. The multi-agent task offloading and resource optimization method based on differential privacy according to claim 3 is characterized in that: The specific process of formulating the task offloading and resource allocation problem as a Markov decision process and defining the Markov decision process is: Define the global state as S t , including the states of all agents, the state of each agent in the global state is s t , including the channel gain between the mobile device and the base station, the computing tasks generated by the mobile device in each time slot, the available energy budget of the mobile device, the number of computing tasks in the local computing task queue of the mobile device, the bandwidth of the wireless channel, the number of computing tasks in the task queue on the edge server, and the computing frequency of the edge server. Therefore, the state of each agent in time slot t is expressed as: Define joint action A t , including the actions of all agents, the action of each agent in the joint action is a t , the action defines the decision that the agent needs to perform in each time slot, including task offloading decision and transmission power allocation. The task offloading decision determines whether the task is processed locally or offloaded to the edge server for processing. Therefore, the action of each agent in time slot t is expressed as: Defining the reward function It is used to measure the pros and cons of each action. Since the goal is to minimize the delay and energy consumption cost of the task, the reward function is: The negative sign of the reward function indicates that the optimization goal is to maximize the reward, that is, to minimize the total task cost; Define the state transition probability P(s t+1 |s t ,a t ), the state transition probability describes the agent's t After that, from the current state s t Transition to next state s t probability.
6. The multi-agent task offloading and resource optimization method based on differential privacy according to claim 5 is characterized in that: The specific process of optimizing the decision-making strategy of the agent according to the defined Markov decision process and the objective function of task offloading and resource allocation and its constraints to obtain a trained agent is: S1: For each agent, randomly initialize the Critic network Q(s,a|θ Q )’s weight θ Q and Actor network μ(s|θ μ )’s weight θ μ , where s represents the state of the agent, and a represents the action performed by the agent in the current state; S2: Initialize two target networks Q'(s,a) and μ'(s), so that the weights of the two target networks are equal to the weights of the Critic network and the Actor network, that is, θ Q′ ←θ Q and θ μ′ ←θ μ , where θ Q′ is the weight of the target network Q'(s,a), θ μ′ is the weight of the target network μ'(s); S3: Initialize the experience replay pool R, which is used to store the experience data of the environment interaction (s t ,a t ,r t ,s t+1 ), where r t is the reward function; S4: Initialize the random process M for action exploration; S5: Set the number of epochs Ep for the total loop training, the maximum time slot T in each loop training epoch, the current loop training epoch number e = 0, and the current time slot t = 0; S6: At each time slot t, obtain the current state s of the agent t And input into the Actor network to generate original actions; a t =μ(s t |θ μ ) (29) S7: Add a privacy protection mechanism and use the message encoding function to inject Gaussian noise into action a t , generate protected actions and send them to other agents; a t ′=μ(s t |θ μ )+N t (30) Among them, a t ′ indicates a protected action, N t represents Gaussian noise; The variance of the Gaussian noise is: in, represents the variance of Gaussian noise, γ1 and γ2 are sampling rates, C is the L2 norm of the message coding function, α = log(δ -1 ) / ε n ,δ is used to control the noise amplitude;ε n is the privacy budget, which is used to control the degree of privacy protection; β represents a trade-off parameter; S8: The current agent receives the protected actions sent by other agents and decodes the protected actions to obtain the decoded actions of other agents; S8.1: The current agent receives a set of noisy messages m (-n)n , which includes protected actions sent by other agents; S8.2: Use the decoding function to decode a set of noisy messages received to obtain a set of decoded messages, including the decoded actions of other agents; q n =f n (m (-n)n ) (32) Among them, q n Represents a set of decoded messages, f n represents the decoding function; S8.3: Optimize the parameters of the decoding function using the gradient chain rule; The optimization formula of the gradient chain rule is: in, is the weight of the Actor network of agent n, The weight of the Actor network The gradient of , that is, the direction of strategy update; is the expected value of the reward function; τ is the trajectory; A represents the joint action; E τ,s,A is the expectation under trajectory τ, state s and joint action A; π n represents the strategy of the Actor network of agent n, In the strategy π n expectations under n (q n |m (-n)n ) is a decoding function; For q n Find the partial derivative; π n (a n |s,q n ) The strategy of the agent to choose an action in state s; Q π (A, s) is the action-state value function, which is the expected return given the joint action A and state s under the strategy π; S9: Each agent performs a protected action a t , get environmental feedback, including reward r t and the next state s t+1 , and the empirical data (s t ,a t ,r t ,s t+1 ) is stored in the experience replay pool; S10: Update the time slot t = t + 1; S11: Determine whether t < T is satisfied. If so, execute S6; otherwise, execute S12; S12: Update the loop training epoch number e = e + 1; S13; Randomly sample L pieces of experience data from the experience pool; S14: Update the Critic network using L pieces of empirical data sampled randomly; S15: Update the Actor network using N pieces of empirical data sampled randomly; S16: Synchronize the weights of the two target networks in a soft update manner; S17: Judge whether e < Ep. If so, return to S4. Otherwise, the training ends and an agent that has completed training is obtained.
7. A multi-agent task offloading and resource optimization system based on differential privacy, used to implement the multi-agent task offloading and resource optimization method based on differential privacy according to any one of claims 1 to 6, characterized in that: Comprising: A mobile device, configured to generate computing tasks and perform task offloading and resource allocation according to the task offloading and resource allocation strategy obtained by the multi-agent task offloading and resource allocation module; A base station, configured to execute the computing tasks generated by the mobile device by using the deployed edge server; A multi-agent task offloading and resource allocation module, which regards each mobile device as an agent. The agent uses its own Actor network to output actions based on the current state, so as to obtain the task offloading and resource allocation strategy.
8. An electronic device, characterized in that: Comprising: A processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the electronic device runs, the processor communicates with the memory through the bus. When the machine-readable instructions are executed by the processor, the steps of the multi-agent task offloading and resource optimization method according to any one of claims 1-6 are executed.
9. A computer-readable storage medium, characterized in that: A computer program is stored in the computer-readable storage medium. When the computer program is run by a processor, the steps of the multi-agent task offloading and resource optimization method according to any one of claims 1-6 are executed.