Task offloading decision method based on mec and anonymous data volume

By constructing a task offloading decision method based on MEC and anonymized data volume, and combining the vehicle-to-everything (V2X) communication network model and deep Q-network model, the problems of task offloading latency and privacy protection in the V2X system are solved, and a balance between resource allocation and privacy protection is achieved.

CN115766764BActive Publication Date: 2026-01-16JIANGSU HENGTONG DIGITAL INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211181766.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-27
Publication Date
2026-01-16
Estimated Expiration
2042-09-27

AI Technical Summary

Technical Problem

Vehicle-to-everything (V2X) systems face challenges such as extended task unloading times and user privacy protection, especially when MEC servers are vulnerable to attacks, leading to the potential leakage of user privacy data.

Method used

A MEC computing resource allocation model based on vehicle users and roadside units is constructed. Combining an anonymization and confusion model and a deep Q-network model, resource allocation and privacy protection are achieved by minimizing the latency optimization function P.

Benefits of technology

This system achieves resource allocation within the vehicle networking system while protecting user privacy, reducing task unloading latency, and enhancing system security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115766764B_ABST
    Figure CN115766764B_ABST
Patent Text Reader

Abstract

The application discloses a task offloading decision method based on MEC and anonymous data volume, and comprises the following steps: constructing a vehicle networking communication network model for MEC computing resource allocation with a vehicle user and a roadside unit as a basic unit; constructing an anonymous confusion model based on a protection strategy in the vehicle networking system, and outputting determined data volume in a probabilistic manner; constructing a time delay minimization optimization function P according to the vehicle networking communication network model and the anonymous confusion model; and adopting a deep Q network model to offload and distribute the computing task generated by the vehicle user in the vehicle networking system according to the time delay minimization optimization function P. The application can realize resource allocation and user privacy protection of the vehicle networking system.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of task offloading decision in vehicle Internet of Things system, and particularly relates to a task offloading decision method based on MEC and anonymous data volume in vehicle Internet of Things system. BACKGROUND

[0002] With the rapid development of Internet of Things (IoT) technology and new generation wireless technology, the environment we live in is more intelligent and controllable. Vehicle Internet of Things is the application of Internet of Things in vehicle terminal. At present, the technologies such as assisted driving, real-time download of high-precision map and remote vehicle control are all emerging technologies based on vehicle Internet of Things system. Under this background, the traditional cloud computing mode obviously cannot provide the super data volume, ultra-low latency and ultra-reliability computing services required by vehicle Internet of Things system.

[0003] Mobile Edge Computing (MEC) is widely used in vehicle Internet of Things system as a new emerging information computing service. The basic framework of MEC is usually a three-layer structure of cloud-edge-end. In vehicle Internet of Things system, MEC servers are usually deployed in Road Side Unit (RSU) which is the edge layer. By distributing and offloading the computing tasks in traditional services that need to be transmitted to the cloud center computing center to the MEC server at the edge of the user terminal, the latency is reduced to meet the computing service requirements of ultra-low latency and ultra-reliability required by vehicle Internet of Things.

[0004] How to shorten the latency in task offloading in vehicle Internet of Things system has always been a hot research topic. The previous researches often only consider offloading decision centering on latency or energy consumption. However, data privacy protection in vehicle Internet of Things task offloading is also a problem that cannot be ignored. Under the huge data base, the edge server is vulnerable and untrustworthy to vehicle users. For example, autonomous driving technology often collects huge vehicle terminal information. If a malicious attacker attacks the MEC server, the personal privacy data of the user is likely to be leaked. When the malicious attacker obtains the offloading decision of the vehicle user or even multiple offloading data, after multiple simulation learning, the personal privacy data such as the usage habit of the vehicle user and the location information of the vehicle will be leaked. Therefore, the development of a good offloading strategy in vehicle Internet of Things system should not only consider the allocation of resources, but also consider the protection of user privacy. SUMMARY

[0005] The technical problem to be solved by the present application is to provide a task offloading decision method based on MEC and anonymous data volume, which can realize the allocation of resources and the protection of user privacy in vehicle Internet of Things system.

[0006] In order to solve the above technical problems, the application provides a task offloading decision method based on MEC and anonymous data volume, comprising the following steps: constructing a vehicle networking communication network model for MEC computing resource allocation based on vehicle users and roadside units as basic units; constructing an anonymous confusion model based on protection strategies in the vehicle networking system, and outputting the determined data volume in a probabilistic manner; constructing a delay minimization optimization function P according to the vehicle networking communication network model and the anonymous confusion model; and adopting a deep Q network model to offload and distribute the computing tasks generated by the vehicle users in the vehicle networking system according to the delay minimization optimization function P.

[0007] As preferred, the vehicle networking communication network model comprises a local computing model carried on the vehicle users and an MEC computing model carried on the roadside units, in each time stamp t, the local computing model completes the calculation of the task volume, and the MEC computing model completes the calculation of the task volume, wherein, the computing task volume generated for the vehicle users, and γ(t) is the proportion of the computing volume unloaded to the roadside units in the total task volume in the vehicle networking system, and γ(t) ∈ [0, 1].

[0008] As preferred, the vehicle networking communication network model comprises a vehicle user layer and an edge device layer; the vehicle user layer comprises N vehicle users, and N is a value of {1, 2,..., n}; and the edge device layer is composed of a plurality of roadside units on which MEC is deployed, and the roadside unit set is defined as M, and M is a value of {1, 2,..., m}.

[0009] As preferred, the delay calculation method of the vehicle networking communication network model is as follows:

[0010] In each time stamp t, there are task volumes that are calculated locally, and ρ represents the average number of execution cycles of the vehicle end, and f n is the computing capability of the nth vehicle, and the computing task generated by the nth vehicle has a computing delay when the computing task is generated locally:

[0011]

[0012] In each time stamp t, there are task volumes that are offloaded to the MEC server of the roadside unit for calculation, and the vehicle user accesses the MEC server in an FDMA manner; w n,m represents the channel bandwidth obtained by the nth vehicle user when accessing the mth roadside unit, p n represents the transmission power of the vehicle n, and g n,mdenotes the average channel gain from vehicle n to the road side unit, N0 denotes the background noise power, and the communication speed between the vehicle user and the road side unit is given by the Shannon formula as follows:

[0013]

[0014] the delay T generated by the task amount of the vehicle user MEC the transmission delay and the MEC computing delay , and the calculation formula is as follows:

[0015]

[0016] wherein ρ represents the average number of execution cycles at the vehicle end, f mec is the average computing capacity of the MEC server;

[0017] Then, in the process of calculating the task offloading site, assuming that the local and MEC computing services are performed simultaneously, the delay generated is the greater value between the local computing delay and the MEC computing delay:

[0018]

[0019] As preferred, the calculation formula of the delay minimization optimization function P is as follows:

[0020]

[0021] s.t.

[0022]

[0023]

[0024] C3: γ(t) ∈ [0, 1]

[0025] wherein C1 is the constraint of the privacy degree, C2 represents that the channel bandwidth available to the vehicle user in the entire system cannot exceed the total bandwidth of the system, and C3 is the offloading decision variable.

[0026] As preferred, the construction method of the deep Q network model is as follows:

[0027] According to the delay minimization optimization function P, the state space S, the action space A, and the model with the delay as the reward function Reward are constructed;

[0028] The state space S of the specific deep Q network model is composed of the computing task amount generated by the vehicle user wireless communication environment and the privacy degree constraint constant θ t , that is

[0029] The action space A of a specific deep Q-network model is composed of the proportion γ(t) of the unloading decision, i.e., A = {γ(t)};

[0030] The specific reward function formula for the deep Q-network model is as follows:

[0031]

[0032] A deep reinforcement learning model is constructed based on the state space S, action space A, and a reward function with time delay. The formula for calculating the target Q value of the target network is as follows:

[0033]

[0034] in, Represents the current s t The reward value for action in state a, where λ represents the reward decay factor, and s t+1 A state represents a state s in the state space. t The next state, a′ represents the next action of action a in the current action space A, ω - For the target neural network parameters;

[0035] Meanwhile, the loss function L(ω) of the deep Q-network model is defined as follows:

[0036] L(ω)=E[Q * -Q(s, α; ω)] 2

[0037] Where s represents the current state of the state space, a represents the current action in the state space, and ω represents the parameters of the evaluation neural network;

[0038] Update the evaluation neural network parameter ω using the SGD method:

[0039]

[0040] Where α represents the learning rate constant;

[0041] The target neural network parameter ω is updated by the neural network parameter iteration update frequency Z. - That is, after Z iterations, ω is assigned to ω. - .

[0042] As a preferred embodiment, the method for constructing the unloading decision algorithm based on the deep Q-network model is as follows:

[0043] S41, Randomly initialize the target neural network and its parameters ω - Evaluate the neural network parameters ω and the experience pool D, and set the number of iterations k;

[0044] S42, for any state s t , the action a in the action space is selected by the greedy algorithm, the action a is executed, and a new state s is obtained t+1 , and the reward Reward is obtained; (s t , a, Reward, s t+1 ) is stored in the experience pool;

[0045] S43, the loss function L(ω) is calculated by selecting a random (s t , a, Reward, s t+1 ) from the experience pool D and updating the target network, and the evaluation network parameter ω is updated by the SGD method;

[0046] S44, whether the number of iterations is greater than k, otherwise return to step S42;

[0047] S45, the number of iterations reaches k times, and the action under the current state is the optimal action, and the optimal action a is output.

[0048] As preferred, the method for outputting the determined data amount in a probabilistic manner by the anonymous confusion model is: probabilizing the data amount data, and outputting the determined amount in a probabilistic manner.

[0049] As preferred, the data amount confusion formula is:

[0050]

[0051] Wherein, ∈ is a privacy budget, is a real data amount, is a confusion data amount after anonymous confusion processing; and a privacy degree represents the degree of privacy protection of the system.

[0052] As preferred, at each timestamp t, any adjacent data sets And The data amount confusion formula satisfies the definition of differential privacy, that is:

[0053]

[0054] Wherein, is an adjacent data set, the sensitivity Noise satisfies distribution.

[0055] Compared with the prior art, the beneficial effects of the present application are:

[0056] The present invention first constructs a vehicle-to-everything (V2X) communication network model based on intelligent vehicle-road cooperation, with vehicle users and roadside units (RSUs) as the basic units for MEC computing resource allocation. Next, addressing the privacy risks inherent in V2X systems, an anonymization and obfuscation model based on differential privacy technology is constructed, outputting deterministic quantities as probabilistic values ​​to achieve obfuscation. Then, based on the V2X communication network model and the anonymization and obfuscation model, a latency minimization optimization function P is constructed. Finally, the latency minimization optimization function P is solved, a deep Q-neural network model is constructed, and through multiple iterative calculations, the optimal offloading decision is obtained, achieving user privacy protection while realizing resource allocation within the V2X system. Attached Figure Description

[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the technical description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0058] Figure 1 This is a flowchart of a task offloading decision method based on MEC and anonymous data volume in a vehicle networking system according to a preferred embodiment of the present invention;

[0059] Figure 2 This is a schematic diagram of a vehicle user computing task offloading model based on MEC and anonymous data volume in a preferred embodiment of the vehicle networking system of the present invention;

[0060] Figure 3 This is a flowchart of the unloading decision algorithm based on a deep Q-network model in a preferred embodiment of the present invention;

[0061] Figure 4 This is the pseudocode for the unloading decision algorithm based on a deep Q-network model in a preferred embodiment of the present invention. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] Example

[0064] Reference Figures 1-3 As shown, this invention discloses a task offloading decision method based on MEC and anonymized data volume, including the following steps:

[0065] S1, constructing a vehicle networking communication network model based on MEC computing resource allocation;

[0066] S2, constructing an anonymous confusion model based on a privacy protection strategy in the vehicle networking system;

[0067] S3, constructing a delay minimization optimization function P according to the vehicle networking communication network model and the anonymous confusion model;

[0068] S4, using a deep Q network model to appropriately offload and distribute the computing tasks generated by the vehicle users in the vehicle networking system according to the delay minimization optimization function P.

[0069] The whole method first constructs a vehicle networking communication network model based on intelligent vehicle-road cooperation, taking vehicle users and roadside units (RSUs) as the basic units of MEC computing resource allocation. Next, in view of the risk of privacy leakage in the vehicle networking system, an anonymous confusion model based on differential privacy technology is constructed to output the determined quantity in a probabilistic manner, thereby achieving the purpose of confusion. Then, a delay minimization optimization function P is constructed according to the vehicle networking communication network model and the anonymous confusion model. Finally, the delay minimization optimization function P is solved, a deep Q neural network model is constructed, and multiple iterations are calculated to finally obtain the optimal offloading decision.

[0070] Specifically, the above vehicle networking communication network model includes a vehicle user layer and an edge device layer. The vehicle user layer includes N vehicle users {12,..., n}. The edge device layer is composed of roadside units (RoadSide Unit, RSU) deployed with MEC. The roadside unit set is defined as M, M = {1, 2,..., m}. The vehicle local and MEC server provide computing services in the vehicle networking system. The computing tasks generated by the users are divisible, so the vehicle users can offload part of the computing tasks to the MEC server to complete to reduce the corresponding delay of the service.

[0071] Specifically: in each timestamp t, the amount of computing tasks generated by the vehicle users is In order to better offload the computing task amount, an offloading ratio γ(t) is defined to represent the proportion of the computing amount offloaded to the roadside unit in the total task amount in the vehicle networking system, γ(t) ∈ [0, 1]; that is, the vehicle users offload the computing tasks of the task amount to the roadside unit MEC server for completion, and the remaining computing tasks of the task amount are directly calculated locally; the whole system computing model is divided into a local computing model and an MEC computing model:

[0072] (1) Local computing model

[0073] In each timestamp t, there are The amount of tasks is calculated locally, and p represents the average number of execution cycles at the vehicle end, f n (CPU cycles / s) is the computing power of the nth vehicle, and the computing task generated by the nth vehicle has a local generation computing delay of:

[0074]

[0075] (2) MEC computing model

[0076] In each timestamp t, there are The amount of tasks will be offloaded to the MEC server of the roadside unit to complete the calculation, where the user accesses the MEC server in the frequency division multiple access (FDMA) manner; w n,m represents the channel bandwidth obtained by the nth vehicle user accessing the mth roadside unit, p n represents the transmission power of vehicle n, g n,m represents the average channel gain from vehicle n to the roadside unit, and N0 represents the background noise power. According to the Shannon formula, the communication speed of the vehicle user and the roadside unit is:

[0077]

[0078] The delay T MEC generated by the amount of tasks is composed of the transmission delay and the MEC computing delay , and the calculation formula is:

[0079]

[0080] where p represents the average number of execution cycles at the vehicle end, f mec (CPU cycles / s) is the average computing power of the MEC server;

[0081] Then, during the calculation task offloading process, assuming that local and MEC computing services are performed simultaneously, the delay generated is necessarily the greater value between the local computing delay and the MEC computing delay:

[0082]

[0083] In addition, the privacy leakage risk in the Internet of Vehicles (IoV) can be identified by a privacy threat model, which states that there is a risk of privacy leakage in the IoV system: when the vehicle is close to the roadside unit and the wireless communication channel is good, the vehicle user tends to adopt an offloading strategy of γ(t)=1, that is, offloading all tasks to the MEC server to minimize the computation latency as much as possible; while attackers can study sensitive information such as vehicle location information by analyzing the offloading ratio parameters between users and vehicles.

[0084] To address this threat, the following solution is employed: The vehicle-to-everything (V2X) model is equipped with an intelligent information obfuscation device. This device anonymizes and obfuscates the data generated by vehicle users, thereby obfuscating the unloading ratio. The function of the intelligent information obfuscation device is to probabilize the data volume, outputting deterministic quantities as probabilistic values, thus achieving the purpose of obfuscation. The data volume obfuscation formula is as follows:

[0085]

[0086] Where ∈ represents the privacy budget. This is the actual amount of data. It is the amount of obfuscated data after anonymization and obfuscation; and it defines the level of privacy. This represents the degree of privacy protection afforded to the system. The data volume obfuscation formula represents the amount of data obfuscation when the actual data volume is... At that time, the output obfuscation level is The probability of.

[0087] For each timestamp t, the above data obfuscation formula satisfies ∈-differential privacy, the proof of which is as follows:

[0088]

[0089] in, For adjacent datasets, sensitivity Noise satisfies distributed;

[0090] like Figure 2 The diagram illustrates a scenario of a vehicle-to-everything (V2X) system after privacy protection processing. Assuming identical wireless communication, the Roadside Unit (RSU) deploys an MEC server. Within a certain timestamp, vehicles A, B, and C within its coverage area generate the same computational task. A and C are equidistant from the RSU, and theoretically, their unloading decisions should be similar. However, in reality, their unloading decisions differ significantly. Vehicle B, being closer to the RSU, tends to make an unloading decision with γ(t) = 1; however, in practice, this decision differs considerably from the actual unloading decision with γ(t) = 1.

[0091] In step S3, based on the vehicle-to-everything (V2X) communication network model and the anonymization and obfuscation model, the latency minimization optimization function P is constructed as follows:

[0092]

[0093] st

[0094]

[0095]

[0096] C3: γ(t)∈[0,1]

[0097] Where C1 represents the privacy constraint, C2 represents that the channel bandwidth available to vehicle users in the entire system cannot exceed the total system bandwidth, and C3 represents the offloading decision variable.

[0098] In step S4, the wireless communication environment and different data sizes are understood. and the privacy constraint constant θ t After considering these three prerequisites, first consider the amount of data. Perform obfuscation, and then adjust the amount of obfuscated data accordingly. The time delay minimization optimization function P is used, and the deep Q-network model selects an appropriate unloading decision γ(t) to... Data of varying size is offloaded to the MEC server of the nearest roadside unit; the remaining data... The computation of data of varying sizes is performed locally; therefore, the aforementioned wireless communication environment Different data sizes and the privacy constraint constant θ t It is a key factor in determining the uninstallation decision. In a particular uninstallation process, the system considers the wireless communication environment and the amount of data... and the privacy constraint constant θ t The most suitable unloading decision γ(t) was selected, the latency cost of the entire system was calculated, and the privacy constraint constant θ for the next unloading by the vehicle user was updated. t+1 ;

[0099] Construct a state space S, an action space A, and a model with delay as the reward function Reward based on the delay minimization optimization function P;

[0100] The state space S of the specific deep Q-network model is generated by the vehicle user and the amount of computational data. Wireless communication environment and the privacy constraint constant θ t Composition, that is

[0101] The action space A of a specific deep Q-network model is composed of unloading decisions γ(t), i.e., A = {γ(t)};

[0102] The specific reward function formula for the deep Q-network model is as follows:

[0103]

[0104] A deep reinforcement learning model is constructed based on the state space S, action space A, and a reward function with time delay. The formula for calculating the target Q value of the target network is as follows:

[0105]

[0106] in, Represents the current (s) t The reward value for state a (action), where λ represents the reward decay factor, s t+1 A state represents a state s in the state space. t The next state, a′ represents the next action of action a in the current action space A, ω - For the target neural network parameters;

[0107] Meanwhile, the loss function L(ω) of the deep Q-network model is defined as follows:

[0108] L(ω)=E[Q * -Q(s, α; ω)] 2

[0109] Where s represents the current state of the state space, a represents the current action in the state space, and ω represents the parameters of the evaluation neural network;

[0110] Update the evaluation neural network parameter ω using the SGD (Stochastic Gradient Descent) method:

[0111]

[0112] Where α represents the learning rate constant;

[0113] The target neural network parameter ω is updated by the neural network parameter iteration update frequency Z. - That is, after Z iterations, ω is assigned to ω. - .

[0114] like Figure 3 As shown, the steps of the unloading decision algorithm based on the deep Q-network model are as follows:

[0115] S41, Randomly initialize the target neural network and its parameters ω - Evaluate the neural network parameters ω and the experience pool D, and set the number of iterations k;

[0116] S42, selecting an action a in the action space by the greedy algorithm, performing the action a to obtain a new state s t t+1 and storing (s, a, Reward, s) into the experience pool; t t+1

[0117] S43, calculating the loss function L(ω) by selecting a random (s, a, Reward, s) from the experience pool D and updating the target network, and updating the evaluation network parameter ω by the SGD method; t t+1

[0118] S44, whether the iteration number is greater than k, if not, returning to step S42;

[0119] S45, the iteration number reaching k times, the action under the current state being the optimal action, and outputting the optimal action a.

[0120] The above description of disclosed embodiments enables those skilled in the art to carry out or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.​​​​​

Claims

1.A method for task offloading decision based on MEC and anonymous data volume, characterized in that, The method comprises the following steps: A vehicle networking communication network model for MEC computing resource allocation is constructed based on vehicle users and roadside units as basic units; An anonymous confusion model based on protection strategies in the vehicle networking system is constructed, and a determined data amount is output in a probabilistic manner; According to a vehicle networking communication network model and an anonymous confusion model, a delay minimization optimization function is constructed ; Adopt deep Q network model, according to delay minimization optimization function The computing tasks generated by the vehicle users in the vehicle networking system are offloaded and distributed; The method for calculating the delay of the vehicle networking communication network model is: in each timestamp , the task amount of is calculated locally, the average number of execution cycles of the vehicle end is represented, the computing capacity of the first vehicle, and the computing task generated by the first vehicle produces a calculation delay locally. At each timestamp In the middle, there is The workload will be offloaded to the MEC server of the roadside unit for calculation, and the vehicle user will access the MEC server using FDMA. Indicates the first The vehicle user accessed the first The channel bandwidth obtained when there are 1 roadside unit Indicates vehicle Transmission power, Indicates vehicle Average channel gain to roadside unit Let the background noise power be represented. According to Shannon's formula, the communication speed between the vehicle user and the roadside unit is: the amount of tasks of the MEC server by the transmission delay and the MEC computing delay , and the calculation formula is: wherein, represents the average number of execution cycles of the vehicle side, is the average computing power of the MEC server; In the process of computing task offloading, assuming that local computing and MEC computing services are performed simultaneously, the time delay is the greater value between the local computing time delay and the MEC computing time delay: ; The latency minimization optimization function The calculation formula is as follows: wherein, is a privacy constraint, represents that the channel bandwidth available to the vehicle users within the entire system cannot exceed the total bandwidth of the system, is an offloading decision variable; The method for the anonymous confusion model to output a determined data amount in a probabilistic manner is to probabilize the data amount data and output the determined amount in a probabilistic manner; The data volume confusion formula is: Wherein, is the privacy budget, is the real data volume, is the confusion data volume after the anonymous confusion processing; and the privacy degree is defined, which represents the degree of privacy protection of the system; At each timestamp , any adjacent data set and , the data volume confusion formula meets the definition of differential privacy, that is: wherein, is the sensitivity , the noise satisfies distribution. 2.The MEC and anonymous data volume based task offloading decision method of claim 1, wherein, The vehicle networking communication network model comprises a local computing model carried on a vehicle user and an MEC computing model carried on a roadside unit, and the local computing model completes the calculation of the task amount at each time stamp , and the MEC computing model completes the calculation of the task amount, wherein the task amount generated for the vehicle user, the proportion of the computing amount unloaded to the roadside unit in the total task amount in the vehicle networking system, and . 3.The MEC and anonymous data volume based task offloading decision method of claim 2, wherein, The vehicle networking communication network model comprises a vehicle user layer and an edge device layer; the vehicle user layer comprises one vehicle user, , and the value of ; the edge device layer is composed of a plurality of roadside units in which MEC is deployed, and the roadside unit set is defined as , , and the value of . 4.The method of claim 1, wherein, The construction method of the deep Q network model is: According to the time delay minimization optimization function constructing a state space S, an action space A, and a model with a time delay as a reward function Reward; The state space S of a particular deep Q-network model is generated by a vehicle user computing task volume , a wireless communication environment , and a privacy degree constraint constant , that is ; The action space A of the specific deep Q-network model is constituted by the proportion of the unloading decision , i.e. ; The reward function Reward formula of the specific deep Q network model is: A deep reinforcement learning model is constructed on the basis of the state space S, the action space A, and the time delay as the reward function Reward. A target Q value of a target network is calculated according to the following formula: wherein, represents a current state of an action, represents a reward decay factor, state represents a next state of a state in a state space, represents a next action of an action in a current action space A, are target neural network parameters; At the same time, define the loss function of the deep Q network model : wherein, representing a current state of the state space, representing a current action in the state space, for evaluating the neural network parameters; Updating evaluation neural network parameters by SGD method : wherein, represents a learning rate constant; updating the target neural network parameters by a neural network parameter iteration update frequency Z , i.e. after Z rounds of iteration is assigned to . 5.The method of claim 4, wherein, The construction method of the offloading decision algorithm based on the deep Q network model is: S41, randomly initialize the target neural network and parameters , evaluate the neural network parameters and the experience pool D, set the number of iterations ; S42, for any state select an action in the action space by a greedy algorithm perform the action get a new state and a reward Reward; store in the experience pool; S43, updating the target network by selecting a random sample from the experience pool D and updating the target network, computing the loss function updating the evaluation network parameters by the SGD method ; S44, whether the iteration number is greater than , otherwise return to step S42; S45, the iteration number reaches action in the current state is the optimal action, and the optimal action is output .

Citation Information

Patent Citations

  • Internet of Vehicles calculation unloading method and system based on multi-objective reinforcement learning

    CN113961204A

  • Joint task unloading and resource allocation method, device and equipment for heterogeneous Internet of Vehicles

    CN114845272A