Internet of vehicles security computing offloading and resource allocation method, computer device and terminal

By employing a multi-agent decision-making method based on deep reinforcement learning, this study addresses the high latency and multiple eavesdropper security issues in dynamic scenarios within the Internet of Vehicles (IoV), enabling inter-vehicle collaboration and secure computation offloading, thus meeting the high reliability and low latency requirements of IoV.

CN114827947BActive Publication Date: 2026-01-16XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210253563.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2026-01-16
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

Existing technologies are ill-suited to the complex and dynamic computing offloading scenarios of vehicle-to-everything (V2X) networks, and cannot meet the demands for high-reliability, high-speed data transmission services. Furthermore, traditional methods cannot address the security issues caused by multiple dynamic eavesdroppers.

Method used

A multi-agent decision-making method based on deep reinforcement learning is adopted, which combines frequency band selection and edge server selection. The DDQN algorithm is used to train vehicles to select reasonable strategies, thereby reducing service latency and improving safety.

Benefits of technology

It enables vehicle-to-vehicle collaboration in dynamic vehicle-to-everything (V2X) scenarios, meeting the communication requirements of ultra-low latency, high security, and high reliability, and adapting to complex multi-user environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114827947B_ABST
    Figure CN114827947B_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of Internet of Vehicles edge computing, and discloses an Internet of Vehicles security computing offloading and resource allocation method, a computer device and a terminal. The present application can overcome the uncertainty caused by multiple eavesdroppers and vehicle movement, reduce service delay while considering physical layer security to ensure communication security. First, the optimization problem is modeled as a multi-agent sequential decision problem, and a reinforcement learning method is used to solve it. Since the DQN (Deep Q learning) method has overestimation problem, it will overestimate the Q value size and reduce the performance. Therefore, the DDQN (Dueling deep Q learning) method is used to train the multi-agent model. The queuing theory is used to model the dynamic process of the vehicle, making the scene closer to the actual scene. This method enables users to select a reasonable strategy to minimize the maximum delay among all vehicles.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of Internet of Vehicles edge computing, and particularly relates to a method for secure computing offloading and resource allocation in Internet of Vehicles, a computer device and a terminal. BACKGROUND

[0002] With the continuous progress of technology and the increasing demand, the application of big data in Internet of Vehicles promotes vehicles to generate more and more delay-sensitive tasks to support new services including traffic flow prediction, and there are two methods to solve this problem: one is to enhance the computing power of the on-board chip so that it can handle these tasks. The other is to use mobile edge computing technology to handle tasks. Mobile edge computing technology uses wireless access network to provide users with the required services and cloud computing functions nearby, thus creating a communication service environment with high performance, low delay and high bandwidth. Mobile edge computing technology can effectively solve the problem of insufficient computing power of vehicles, but due to the open nature of the wireless channel, the computing offloading process has the risk of information leakage, and physical layer security technology protects the privacy of users by utilizing the characteristics of the wireless channel, such as signal processing, channel coding, multi-antenna modulation, etc. With the application of big data in Internet of Vehicles scenarios, the contradiction between massive data transmission and limited spectrum resources is increasingly prominent, and emerging spectrum sharing technology can significantly improve the utilization rate of spectrum while ensuring normal communication needs of users, saving spectrum resources.

[0003] In existing research, there is currently no research that combines physical layer security technology and spectrum sharing technology in vehicle edge computing networks. On the one hand, the rapid change of network topology structure caused by the high-speed movement of vehicles makes it difficult for previous schemes to make quick decisions, and on the other hand, considering the security scheme of vehicle edge computing network makes it difficult to meet the requirement of ultra-low delay. Due to the high-speed movement of vehicles and the presence of more eavesdroppers, the Internet of Vehicles scenario is relatively complex, and traditional mathematical optimization methods are difficult to adapt to the dynamic and complex Internet of Vehicles scenario, so it is necessary to use the decision-making and learning ability of deep reinforcement learning (DRL) to solve complex dynamic optimization problems. Based on this, a transmission scheme (joint secure offloading and resource allocation, SoRA) based on deep reinforcement learning (DRL) is designed for the multi-user communication scenario of Internet of Vehicles, which quickly adapts to complex and dynamic communication environments and can maximize the reduction of service delay while ensuring the communication security of individual users.

[0004] In actual vehicle networking multi-user service scenarios, multiple users may compete for the same high-quality spectrum resource or the same edge server computing resource, which will cause a competitive game problem. Therefore, how to organically combine the frequency band selection and the edge server selection, reduce the overall delay of the service while considering the vehicle power, how to adapt to the rapid changes of the dynamic scene in the vehicle networking and solve the multi-eavesdropper security problem, and meet the service demand in the dynamic scene of the vehicle networking are problems to be solved in the development of vehicle networking communication computing offloading technology. The disadvantage of using mathematical derivation to ensure physical layer security research is that it can be carried out in a static scene, and cannot adapt to the high-speed dynamic scene in the vehicle networking. The existing research on the physical layer security of the vehicle networking still stays in considering only one static eavesdropper, but in actual scenarios, there are often many eavesdroppers. Once the above problems are solved, the security of communication in vehicle edge computing can be ensured.

[0005] Through the above analysis, the problems and defects of the prior art are:

[0006] (1) The traditional optimization method is difficult to adapt to the complex dynamic vehicle networking computing offloading scene, and cannot meet the service demand of high reliability and high-speed data transmission.

[0007] (2) Due to the selfish preference of the vehicle node, part of the vehicles will prefer to minimize their service delay, but will cause the service delay of the remaining vehicles to exceed the tolerable range, resulting in the decline of the overall performance of the system.

[0008] (3) The existing technology only considers a single static eavesdropper, and in actual scenarios, there are bound to be multiple dynamic eavesdroppers, so it is difficult to ensure the security of the computing offloading process. SUMMARY

[0009] In view of the problems existing in the prior art, the present application provides a vehicle networking secure computing offloading and resource allocation method, computer equipment and terminal.

[0010] The present application is implemented as follows: a vehicle networking secure computing offloading and resource allocation method, which can effectively break through the limitation of static scene and realize real-time decision of vehicle networking secure computing offloading. First, the dynamic process of the vehicle is modeled by using queuing theory, and there are multiple dynamic eavesdroppers in the scene. Secondly, the optimization problem is modeled as a multi-agent sequential decision problem, and the reinforcement learning method of DDQN is used for multi-agent training and solution, so that the user can select a reasonable strategy to minimize the maximum service delay of all vehicles while performing secure offloading. The present application actively promotes the cooperation between the network nodes of the vehicle networking, meets the communication demand of ultra-low delay, high security and high reliability, and makes it adapt to the dynamic vehicle networking scene. Further, the vehicle networking secure computing offloading and resource allocation method comprises the following steps:

[0011] First, a single base station of vehicle networking communication scene is constructed, and the base station is connected to the edge server to provide computing offloading service; a dynamic vehicle networking scene is built for the subsequent modeling and analysis.

[0012] Second, the transmission process of different links is modeled, and the communication process is modeled; the communication channel used in the subsequent use of the application is laid.

[0013] Third, Wyner's eavesdropping coding scheme is used, and the optimization target is modeled; the eavesdropping rate of the eavesdropper is calculated, and the model training is laid.

[0014] Fourth, the base station obtains the state information at the current time by interacting with the surrounding environment information, including the information of the target vehicle, including the vehicle speed, position coordinates, current state, frequency band allocation information and resource allocation information on the edge server as the state input of deep reinforcement learning, wherein the deep reinforcement learning uses the DDQN algorithm; the state space for the subsequent training of the agent is determined.

[0015] Fifth, based on the current state information, the vehicle selects the corresponding action; the action of the current state is power selection, frequency band selection and edge server computing resource block selection; the action space of the agent is determined.

[0016] Sixth, according to the model and strategy constructed in the second step, the reward mechanism and the structure of the neural network are designed; the reward mechanism is designed to make the user vehicles in the system better cooperate to minimize the maximum delay.

[0017] Seventh, the DDQN neural network in the sixth step is used to extract the input features of the current state, fit the Q function, obtain the Q values of different actions under various input states, select the action under the current state according to the ∈-greedy strategy, and train and update the neural network parameters combined with the reward mechanism in the sixth step, mainly update the neural network parameters; update the neural network parameters.

[0018] Eighth, the trained DDQN network is used, the state information of the current environment is taken as the state input, the Q value sequence of the corresponding action under the current state is output, and the action with the maximum Q value is taken as the strategy of power selection, frequency band selection and edge server computing resource selection of the target vehicle under the current state; the convergence and convergence time of the model are guaranteed.

[0019] Further, the process of the first step is as follows: the arrival process of vehicles is modeled by using queuing theory, the time interval t of vehicle arrival obeys negative exponential distribution, and the probability density function is as follows:

[0020]

[0021] where λ is the vehicle arrival rate and t is the time interval of vehicle arrival.

[0022] Further, the second step is as follows:

[0023] 2.1 In the communication process, the channel gain g between the transmitting end and the receiving end k is composed of large-scale fading α k and small-scale fading h k :

[0024] g k = α k h k ;

[0025] 2.2 Large-scale fading h k is composed of path loss and shadow fading, the path loss of V2V is divided into LOS and NLOS cases, in the LOS case:

[0026]

[0027] where f c is the carrier frequency, d is the distance, d BP is the effective distance, h0 and h1 are the heights of different vehicles, and the path loss in the NLOS case is:

[0028] PL Nlos (d1,d2)=PL los (d1)+20-12.5n j +10n j log 10 d2+3log 10 (f c / 5)

[0029] where n j = max(2.8-0.0024d1, 1.84), d1 and d2 represent the length and width of each road grid in the Manhattan grid layout;

[0030] Shadow fading of V2V:

[0031]

[0032] where D is the updated distance matrix, D corr is 10 on general urban roads, N S (n) is an M x M matrix, which is a normal distribution matrix with an expected value of 0 and a variance of 1;

[0033] The path loss PL of V2I V2I = a + blog 10R, where R denotes the distance between the vehicle and the base station, a, b are path loss parameters related to the scenario; Shadowing fading for V2I:

[0034]

[0035] where D i is the matrix of the distance update for the i-th vehicle user, D corr is 50, R is an M x M matrix with k on the diagonal and k / 2 elsewhere, N i (n) is the M x 1 matrix generated by the i-th vehicle user, which is a normal distribution matrix with expected value 0 and variance 1 ;

[0036] 2.3 The offloading link rate of the k-th vehicle user to the base station through the m-th sub-channel is where W is the channel bandwidth, is the signal-to-noise ratio, and is expressed as:

[0037]

[0038] where denotes the power of the k-th vehicle user to the base station, g k,B [m] denotes the channel gain of the k-th vehicle to the base station in the m-th frequency band, σ 2 denotes the noise, is the interference received by the k-th vehicle user during offloading, denotes the transmission power of the m-th vehicle for V2V communication, g m,B [m] denotes the interference channel gain of the m-th vehicle for V2V communication to V2I communication, p k' [m] = 1 indicates that this frequency band is used, p k' [m] = 0 indicates that this frequency band is not used;

[0039] 2.4 The rate of the n-th eavesdropper when eavesdropping on the k-th vehicle user in the m-th sub-band is denotes:

[0040]

[0041] where denotes the power of the k-th vehicle user, g k,n [m] denotes the channel gain of the k-th vehicle to the eavesdropper in the m-th frequency band, σ 2 denotes the noise, is the interference received during eavesdropping, denotes the transmission power of the m-th vehicle for V2V communication, g m,n[m] represents the channel gain of the mth vehicle to the eavesdropper, ρ k' [m] = 1 means using this frequency band, ρ k' [m] = 0 means not using this frequency band.

[0042] Further, the third step is as follows:

[0043] 3.1 Offloading rate: offloading refers to connecting to an edge server through a base station to provide computing offloading services; secure offloading rate The formula for measuring the rate at which the kth vehicle user offloads data to the base station is wherein, Rk represents the rate at which the kth vehicle user normally offloads data to the base station, i.e., the rate at which data is transmitted from the vehicle user to the base station without considering the influence of the eavesdropper; Rsk represents the rate at which the kth vehicle user data can be obtained by the eavesdropper ve when the eavesdropper exists, Rmax represents the maximum rate that can be obtained by all possible eavesdroppers; by this calculation method, it is ensured that Rsk can reflect the rate that can be actually used for secure transmission in the presence of eavesdropping risk.

[0044] 3.2 Time for the kth vehicle user to transmit to the base station wherein B k represents the size of the computing task, Rk represents the secure offloading rate, and the time for the task to be calculated on the edge computing server:

[0045]

[0046] wherein B k represents the size of the computing task, z k [j] = 1 means that the jth resource block is allocated to the kth vehicle user, z k [j] = 0 means that the jth resource block is not allocated to the kth vehicle user, N c,j represents the total number of edge server processing cores, u E represents the processing rate of each core; then the total latency

[0047] 3.3 Minimizing the maximum service latency among all vehicles, the objective function is:

[0048]

[0049] Subject to:

[0050]

[0051] where N u denotes the total number of served vehicles, N b denotes the edge server resource blocks, N c denotes the total number of processing cores and processing capacity of MEC servers, N p denotes the number of vehicle power options, means the kth vehicle user selects the ith gear power as the transmission power, otherwise C1 ensures that the total number of processing cores does not exceed the number of cores of the edge server, C2, C3, C4 three constraint conditions ensure that each vehicle user can only select one frequency band, one transmission power and one computing resource block, C5 specifies the decision variable of the optimization target as a binary variable.

[0052] Further, the process of the fifth step is as follows:

[0053] 5.1 The action space can be represented by a three-dimensional coordinate, the x-axis represents the frequency band selection, the y-axis represents the vehicle transmission power selection, and the z-axis represents the selection of the computing resource block on the edge server; assuming that there are N a frequency band options, N p vehicle power options, and N b edge server resource block options, then for any vehicle requiring service, the action can be N a ×N b ×N p ;

[0054] 5.2 An ∈-greedy strategy is used to balance the training process and the exploration process, at time t, the base station selects the action with the maximum Q value with a probability of 1-∈, and selects an action from the state space A with a probability of ∈.

[0055] Further, the process of the sixth step is as follows:

[0056] 6.1 The reward is divided into N w grades according to the time of service delay;

[0057] 6.2 When the computing offloading rate is too low, there will be a large delay, and the reward is 0.

[0058] Further, the neural network training process of the seventh step is as follows:

[0059] 7.1 Initialize the environment information and Q network parameters, and generate vehicle running data;

[0060] 7.2 In each training round, update and obtain the current vehicle position and environment state, reset the frequency band power selection, and edge server resource allocation strategy;

[0061] 7.3 Select an action for the target vehicle according to the current state information and the greedy algorithm, that is, a combination scheme of band selection, vehicle power and edge server resource allocation, and update the information of the environment;

[0062] 7.4 Obtain the action combination scheme of all target vehicles, obtain the reward value r c,i and the returned reward value {r t};

[0063] 7.5 Store the state, action, reward and next state at time t as a sample into the experience pool,

[0064] 7.6 When the number of experience pool samples is sufficient, start training the model, randomly extract a small batch of samples (s t ,a t ,r t ,s t+1 ) from the experience pool, train the network parameters, and update the target network weights.

[0065] Another object of the present application is to provide a computer device comprising a memory and a processor, the memory storing a computer program, the computer program being executed by the processor to enable the processor to perform the steps of the vehicle networking security computing offloading and resource allocation method.

[0066] Another object of the present application is to provide an information data processing terminal for executing the steps of the vehicle networking security computing offloading and resource allocation method.

[0067] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solution to be protected by the present application are analyzed from the following aspects:

[0068] First, in view of the technical problems existing in the prior art and the difficulty in solving the problems, the technical solution to be protected by the present application is closely combined with the results and data in the research and development process, and the technical problems solved by the technical solution are analyzed in detail and deeply. Some creative technical effects brought about after solving the problems are described as follows:

[0069] The application can overcome the uncertainty caused by multiple eavesdroppers and vehicle movement, reduce service delay while considering physical layer security to make communication safe. First, the optimization problem is modeled as a multi-agent sequential decision problem, and a reinforcement learning method is used to solve it. Because the DQN (Deep Q learning) method has overestimation problem, it will overestimate the Q value size and reduce the performance. Therefore, the DDQN (Dueling deep Q learning) method is used to train the multi-agent model. The queuing theory is used to model the dynamic process of the vehicle, making the scene closer to the actual scene. This method enables users to choose a reasonable strategy to minimize the maximum delay in all vehicles.

[0070] Second, from the perspective of the product as a whole, the technical effects and advantages of the technical solution to be protected by the application are described as follows: The application studies the service problem of multiple users and multiple eavesdroppers for vehicle computing offloading, and proposes a SoRA strategy based on DRL, which can help vehicles to quickly make the optimal strategy according to the current environment to minimize the service delay. In the model, we consider the high-speed movement characteristics of vehicles, the competition in the process of band selection and edge server resource block selection, and the interference in the multi-user scenario. The simulation results of the model show that the method proposed in the application can reduce the overall delay of vehicle computing offloading service while improving the security of communication, etc.

[0071] Third, as the auxiliary evidence for the creativity of the claims of the application, it is also reflected in the following important aspects:

[0072] The technical solution of the application fills the technical gap in the industry at home and abroad:

[0073] The application proposes a secure computing offloading and resource allocation method, which can effectively break through the limitations of static scenarios and realize real-time decision-making of dynamic Internet of Vehicles. At the same time, the application can solve the problem of increasing network overall delay caused by selfish preference between nodes in the prior art, effectively encourage cooperation between vehicles, minimize the maximum service delay in the network system, and ensure the security of computing offloading while considering multiple dynamic eavesdroppers, meet the communication requirements of ultra-low delay, high reliability and high security in Internet of Vehicles communication, adapt to dynamic and complex Internet of Vehicles communication and edge computing scenarios, fill the gap in the Internet of Vehicles industry at home and abroad, and promote the landing of edge computing services. BRIEF DESCRIPTION OF DRAWINGS

[0074] Figure 1 is a flowchart of the secure computing offloading and resource allocation method for Internet of Vehicles provided by the embodiment of the application.

[0075] Figure 2is an implementation flowchart of a vehicle networking security computing offloading and resource allocation method provided by an embodiment of the present application.

[0076] Figure 3 is a vehicle networking millimeter wave multi-user communication scenario schematic diagram provided by an embodiment of the present application.

[0077] Figure 4 is a DDQN network schematic diagram provided by an embodiment of the present application.

[0078] Figure 5 is a system performance and vehicle performance comparison schematic diagram of different schemes under different traffic patterns provided by an embodiment of the present application.

[0079] Figure 6 is an average connection probability schematic diagram of different schemes under different capacity threshold limits provided by an embodiment of the present application. DETAILED DESCRIPTION

[0080] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0081] I. Explanation of Embodiments. In order to enable those skilled in the art to fully understand how the present application is specifically implemented, this part is an explanation of the embodiments of the technical solutions claimed in the claims.

[0082] As shown in Figure 1 , the vehicle networking security computing offloading and resource allocation method provided by the present application includes the following steps:

[0083] S101: A single base station vehicle networking communication scenario is constructed, and the base station is connected to an edge server to provide computing offloading service;

[0084] S102: The communication process is modeled for the transmission process of different links;

[0085] S103: Wyner's wiretapping coding scheme is used to improve the security of the vehicle networking edge computing network, and the optimization target is modeled;

[0086] S104: The base station obtains the state information at the current time by acting on the surrounding environment information, including the information of the target vehicle (including vehicle speed, position coordinates, current state), frequency band allocation information and resource allocation information on the edge server as the state input of deep reinforcement learning, wherein the deep reinforcement learning uses the DDQN algorithm;

[0087] S105: Based on the current state information, the vehicle selects the corresponding action; the current state action is power selection, frequency band selection and edge server computing resource block selection;

[0088] S106: According to the model and strategy constructed in S102, the reward mechanism and the structure of the neural network are designed;

[0089] S107: The input features of the current state are extracted by the DDQN neural network in S106, the Q function is fitted, the Q values of different actions under various input states are obtained, the action under the current state is selected according to the ∈-greedy strategy, and the neural network parameters are trained and updated in combination with the reward mechanism in S106;

[0090] S108: Using the trained DDQN network, the state information of the current environment is input as the state, and the Q value sequence of the corresponding action under the current state is output, and the action with the maximum Q value is selected as the power selection, frequency band selection and edge server computing resource selection strategy of the target vehicle under the current state.

[0091] The process of step S101 is as follows: the queuing theory is used to model the arrival process of vehicles, the time interval t of vehicle arrival obeys negative exponential distribution, and the probability density function is as follows:

[0092]

[0093] Where λ is the vehicle arrival rate, and t is the time interval of vehicle arrival.

[0094] The process of step S102 is as follows:

[0095] S2.1 In the communication process, the channel gain g between the transmitting end and the receiving end k is composed of large-scale fading α k and small-scale fading h k .

[0096] g k = α k h k ;

[0097] S2.2 Large-scale fading h k is composed of path loss and shadow fading. The path loss of V2V is divided into LOS and NLOS cases. In the LOS case:

[0098]

[0099] Where f c is the carrier frequency, d is the distance, d BP is the effective distance, h0 and h1 are the heights of vehicles, and the path loss in the NLOS case is:

[0100] PL Nlos (d1, d2) = PL los (d1) + 20 - 12.5n j + 10n j log 10 d2 + 3 log 10 (f c / 5)

[0101] where n j = max(2.8 - 0.0024d1, 1.84), d1 and d2 represent the length and width of each road grid in Manhattan grid layout.

[0102] Shadow fading for V2V:

[0103]

[0104] where D is the updated distance matrix, D corr is 10 on general urban roads, N S (n) is an M x M matrix which is a normally distributed matrix with mean 0 and variance 1.

[0105] Path loss for V2I PL V2I = a + blog 10 R, where R represents the distance between the vehicle and the base station, a, b are path loss parameters related to the scenario. Shadow fading for V2I:

[0106]

[0107] where D i represents the matrix of the i-th vehicle user updating distance, D corr is 50, R is an M x M matrix with all diagonal elements being k and other elements being k / 2, N i (n) is the M x 1 matrix generated by the i-th vehicle user, which is a normally distributed matrix with mean 0 and variance 1.

[0108] S2.3 The offloading link rate of the k-th vehicle user to the base station through the m-th sub-channel is where W is the channel bandwidth, is the signal-to-noise ratio represented as:

[0109]

[0110] where represents the power of the k-th vehicle user to the base station, g k,B [m] represents the channel gain of the k-th vehicle to the base station on the m-th frequency band, σ2 Indicates noise. This refers to the interference encountered by the k-th user's vehicle during unloading. g represents the transmission power of the m-th vehicle performing V2V communication. m,B [m] represents the interference channel gain caused by the m-th vehicle's V2V communication to V2I communication, ρ k' [m] = 1 indicates that this frequency band is used, ρ k' [m] = 0 indicates that this frequency band is not used;

[0111] S2.4 The rate at which the nth eavesdropper eavesdrops on the kth vehicle user in the mth sub-band. Represented as

[0112]

[0113] in G represents the power of the k-th vehicle user. k,n [m] represents the channel gain from the k-th vehicle to the eavesdropper in the m-th frequency band, σ 2 Indicates noise. It was interference during the eavesdropping. g represents the transmission power of the m-th vehicle performing V2V communication. m,n [m] represents the channel gain for V2V communication between the m-th vehicle and the eavesdropper, ρ k' [m] = 1 indicates that this frequency band is used, ρ k' [m] = 0 indicates that this frequency band is not used.

[0114] The process in step S103 is as follows:

[0115] S3.1 Secure Offload Rate: Offload refers to the provision of computing offload services through connection to the edge server via the base station; secure offload rate. The rate at which the k-th vehicle user securely offloads data to the base station is measured by the following formula: in, This represents the rate at which the k-th vehicle user normally unloads data to the base station, that is, the rate at which data is transmitted from the vehicle user to the base station without considering the influence of eavesdroppers. This represents the rate at which the eavesdropper ve can obtain data related to the k-th vehicle user, given the presence of an eavesdropper ve. This represents the maximum rate that can be obtained from all possible eavesdroppers; this calculation method ensures that Rsk reflects the actual rate that can be used for secure transmission under the risk of eavesdropping.

[0116] S3.2 Time of transmission from the kth vehicle user to the base station Among them B kdenotes the size of the computing task, denotes the privacy offloading rate. The time of the task computed on the edge computing server:

[0117]

[0118] where B k denotes the size of the computing task, z k [j] = 1 means the jth resource block is allocated to the kth vehicle user for use, z k [j] = 0 means the jth resource block is not allocated to the kth vehicle user for use, N c,j denotes the total number of processing cores of the edge server, u E denotes the processing rate of each core. Then the total latency

[0119] S3.3 Minimize the maximum service latency among all vehicles, the objective function is:

[0120]

[0121] Subject to:

[0122]

[0123] where N u denotes the total number of service vehicles, N b denotes the edge server resource block, N c denotes the total number of processing cores and processing capacity of the MEC server, N p denotes the number of vehicle power options, means the kth vehicle user selects the ith power as the transmission power, otherwise C1 guarantees that the total number of processing cores does not exceed the number of cores of the edge server, C2, C3, C4 three constraint conditions guarantee that each vehicle user can only choose one frequency band, one transmission power and one computing resource block. C5 specifies the decision variable of the optimization target as a binary variable.

[0124] The process of step S105 is as follows:

[0125] S5.1 The action space can be represented by a three-dimensional coordinate, the x-axis represents the frequency band selection, the y-axis represents the vehicle transmission power selection, and the z-axis represents the selection of the computing resource block on the edge server. Let the frequency band selection have N a , the vehicle power selection have N b , and the edge server resource block selection have N p , then for any vehicle that needs service, the action can be N a × N bXN p .

[0126] S5.2, the epsilon-greedy strategy is used to balance the training process and the exploration process. At time t, the base station selects the action with the maximum Q value with a probability of 1-epsilon, and selects an action from the state space A with a probability of epsilon.

[0127] The steps in step S106 are as follows:

[0128] S6.1, the reward is divided into N w grades according to the time of service delay.

[0129] S6.2, when the calculated offloading rate is too low, there will be a large delay, and the reward is 0.

[0130] The neural network training process in step S107 is as follows:

[0131] S7.1, initialize the environment information and Q network parameters, and generate vehicle running data.

[0132] S7.2, in each training round, update and obtain the current vehicle position and environment state, reset the frequency band power selection, and edge server resource allocation strategy.

[0133] S7.3, according to the current state information and the greedy algorithm, select an action for the target vehicle, that is, the combination scheme of frequency band selection, vehicle power and edge server resource allocation, and update the information of the environment.

[0134] S7.4, obtain the action combination scheme of all target vehicles, and then obtain the reward value r c,i related to the capacity and the returned reward value {r t}.

[0135] S7.5, store the state at time t, action, reward and next state as a sample into the experience pool.

[0136] S7.6, when the number of experience pool samples is sufficient, start training the model. Randomly sample a small batch of samples (s t ,a t ,r t ,s t+1 ) from the experience pool, train the network parameters, and update the target network weights.

[0137] II. Application Examples. In order to prove the creativity and technical value of the technical scheme of the application, this part is an application example of the technical scheme of the claims on specific products or related technologies.

[0138] The application is applied and simulated in a dynamic vehicle networking computing offloading scene. An application example considers a two-way crossroads communication system, the time interval of vehicle arrival on each road obeys a negative exponential distribution, the vehicle arrival rate is 0.5, and the vehicle speed is 72 km / h. Therefore, the scene requires the base station to quickly make edge server resource blocks, vehicle transmission power and frequency band block selection according to limited state information. At the same time, the model training and verification analysis proposed in the application are carried out on the application implementation case. Figures 5-6 For the performance analysis diagram of the embodiment, multi-dimensional energy analysis is carried out, and the effectiveness and robustness of the proposed safe computing offloading and resource allocation method are verified, which can significantly improve the overall system performance. This has a profound significance for promoting the development of vehicle networking and edge computing technology. It should be noted that the embodiments of the application can be realized by hardware, software or a combination of software and hardware. The hardware part can be realized by using a special logic chip; the software can be stored in a memory to execute appropriate instructions. The device and its modules of the application can be realized by a hardware circuit such as a very large scale integrated circuit or a gate array, a semiconductor such as a logic chip, a transistor, or a programmable hardware device such as a field programmable gate array, a programmable logic device, etc., or can be realized by software executed by various types of processors, or can be realized by a combination of the above hardware circuit and software, such as firmware.

[0139] III. Evidence of the effects of the embodiments. The embodiments of the application have achieved some positive effects during research and development or use, and indeed have great advantages compared with the prior art. The following content is described in combination with data, graphs and the like during the test process.

[0140] Figure 5 It is a total time delay schematic diagram of the application in different positions. By randomly generating 10 position points, the maximum delay of all vehicles is calculated as the total processing delay under different schemes. It can be found from the figure that the delay performance of the SoRA scheme is much better than that of the local computing scheme and the scheme without frequency band sharing. Since frequency band sharing not only causes interference to the target but also causes interference to the eavesdropper, the eavesdropping rate of the eavesdropper is reduced, and finally the performance of the SoRA scheme is shorter than that of the scheme without sharing. For all random position points, the SoRA scheme is very close to the optimal scheme, compared with the optimal scheme which needs to spend a lot of time to traverse all possibilities, the SoRA strategy based on DRL quickly adapts to the characteristics of the vehicle networking environment, which shows the high efficiency of the scheme.

[0141] Figure 6The average connection probability of different schemes under different capacity threshold limits. By setting different capacity thresholds, the performance of different schemes can be seen. As can be seen from the figure, with the continuous increase of the capacity threshold, the connection probability of the random strategy first sharply decreases and then slowly decreases, the connection probability of the optimal scheme remains unchanged, and the SoRA strategy is better than the non-sharing and random strategy, and is not much different from the optimal strategy.

[0142] It should be noted that the embodiments of the present application can be realized by hardware, software or a combination of software and hardware. The hardware part can be realized by special logic; the software part can be stored in a memory and executed by a suitable instruction execution system, such as a microprocessor or a specially designed hardware. Those skilled in the art can understand that the above-mentioned devices and methods can be realized by computer executable instructions and / or included in processor control code, such as provided on a carrier medium, such as a magnetic disk, CD or DVD-ROM, a programmable memory, such as a read-only memory (firmware), or a data carrier, such as an optical or electronic signal carrier. The devices of the present application and their modules can be realized by hardware circuits, such as very large scale integrated circuits or gate arrays, semiconductors, such as logic chips, transistors, etc., or programmable hardware devices, such as field programmable gate arrays, programmable logic devices, etc., by software executed by various types of processors, or by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0143] The above is only a specific embodiment of the present application, but the protection scope of the present application is not limited thereto, any modification, equivalent replacement and improvement made by those skilled in the art within the technical range disclosed by the present application, as long as it is within the spirit and principle of the present application, should be covered within the protection scope of the present application.

Claims

1. A method for secure computation offloading and resource allocation in Internet of Vehicles, characterized in that, The vehicle networking security computing offloading and resource allocation method first models the dynamic process of vehicles by using queuing theory, and there are multiple dynamic eavesdroppers in the scene; Secondly, the optimization problem is modeled as a multi-agent sequential decision problem, and the reinforcement learning method of DDQN is used for multi-agent training and solution, so that users can choose a reasonable strategy to minimize the maximum service delay in all vehicles while performing security offloading; The vehicle networking security computing offloading and resource allocation method comprises the following steps: First, a single base station vehicle networking communication scene is constructed, and the base station is connected to an edge server to provide computing offloading services; Second, the transmission process of different links is modeled, and the communication process is modeled; Third, Wyner's eavesdropping coding scheme is used, and the optimization target is modeled; Fourth, the base station obtains the state information at the current time through the interaction with the surrounding environment information, including the information of the target vehicle, including vehicle speed, position coordinates, current state, frequency band allocation information and edge server resource allocation information as the state input of deep reinforcement learning, wherein the deep reinforcement learning uses the DDQN algorithm; Fifth, based on the current state information, the vehicle selects the corresponding action; the action of the current state is power selection, frequency band selection and edge server computing resource block selection; Sixth, the reward mechanism and the structure of the neural network are designed according to the model and the strategy constructed in the second step; Seventh, the input features of the current state are extracted by using the DDQN neural network in the sixth step, the Q function is fitted, the Q values of different actions under various input states are obtained, the action under the current state is selected according to the ∈-greedy strategy, and the neural network parameters are trained and updated in combination with the reward mechanism in the sixth step; Eighth, the trained DDQN network is used to input the state information of the current environment as the state input, and output the Q value sequence of the corresponding action under the current state, and the action with the maximum Q value is selected as the strategy of power selection, frequency band selection and edge server computing resource selection of the target vehicle under the current state. The process of the third step is as follows: 3.1 Secrecy offloading rate: offloading refers to connecting to an edge server through a base station to provide a computing offloading service; the secrecy offloading rate The secrecy offloading rate of the kth vehicle user is calculated as follows Wherein, represents the normal offloading rate of the kth vehicle user, that is, the rate at which data is transmitted from the vehicle user to the base station without considering the influence of the eavesdropper; represents the rate related to the data of the kth vehicle user that the eavesdropper ve can obtain when the eavesdropper exists, represents the maximum rate that can be obtained among all possible eavesdroppers; through this calculation, it is ensured that the secrecy offloading rate can reflect the rate that can be actually used for secret transmission in the case where there is a risk of eavesdropping; 3.2 Time of transmission of the kth vehicle user to the base station where B k denotes the size of the computing task, denotes the secret offloading rate, the time of computation of the task on the edge computing server: where B k denotes the size of the computing task, z k [j] = 1 indicates that the jth resource block is allocated to the kth vehicular user for use, z k [j] = 0 indicates that the jth resource block is not allocated to the kth vehicular user for use, N c,j denotes the total number of edge server processing cores, u E denotes the processing rate of each core; then the total latency 3.3 Minimize the maximum service delay in all vehicles, and the objective function is: Subject to: where N u denotes the total number of serving vehicles, N b denotes the edge server resource blocks, N c denotes the total number of processing cores and processing capacity of the MEC server, N p denotes the number of vehicle power options, means that the kth vehicle user selects the ith gear power as the transmission power, otherwise ,N c denotes the total number of processing cores and processing capacity of the MEC server, N u denotes the total number of serving vehicles, C1 guarantees that the total number of processing cores does not exceed the number of cores of the edge server, C2, C3, C4 three constraint conditions guarantee that each vehicle user can only select one frequency band, one transmission power and one computing resource block, C5 specifies the decision variable of the optimization target as a binary variable; The process of the fifth step is as follows: 5.1 The action space can be represented by a three-dimensional coordinate, where the x-axis represents the frequency band selection, the y-axis represents the vehicle transmit power selection, and the z-axis represents the selection of the computing resource block on the edge server; let the frequency band selection have N a , the vehicle power selection have N b , and the edge server resource block selection have N p , then the action for any vehicle requiring service can be N a ×N b ×N p ; 5.2, the ∈-greedy strategy is used to balance the training process and the exploration process, at time t, the base station selects the action with the maximum Q value with a probability of 1-∈, and selects an action from the state space A with a probability of ∈. 2.The method of claim 1, wherein, The process of the first step is as follows: the arrival process of vehicles is modeled by using queuing theory, the time interval t of vehicle arrival obeys negative exponential distribution, and the probability density function is as follows: Where λ is the vehicle arrival rate, and t is the time interval of vehicle arrival. 3.The method of claim 1, wherein, The process of the second step is as follows: 2.1 In the communication process, the channel gain g between the transmitting end and the receiving end k consists of a large-scale fading a k and a small-scale fading h k : g k = a k h k ; 2.2 Large-scale fading h k The path loss of V2V is divided into LOS and NLOS cases, which consists of path loss and shadow fading. In the LOS case: Where f c Let d be the carrier frequency and d be the distance. BP Let h0 and h1 be the vehicle heights, and h0 be the effective distance. The path loss in the NLOS scenario is: PL Nlos (d1, d2) = PL los (d1) + 20 - 12.5n j + 10n j log 10 d2 + 3 log 10 (f c / 5) where n j = max(2.8 - 0.0024d1, 1.84), d1 and d2 represent the length and width of each road grid in the Manhattan grid layout; Shadow fading of V2V: where D is the updated distance matrix, D corr 10 on general urban roads, N S (n) is an M x M matrix that is a normally distributed matrix with an expected value of 0 and a variance of 1 ; Path loss PL for V2I V2I = a + blog 10 R, where R denotes the distance of the vehicle from the base station, a, b are path loss parameters related to the scenario; Shadowing fading for V2I: 2.3 The offloading link rate of the kth vehicle user to the base station through the mth sub-band is where W is the channel bandwidth, The signal-to-noise ratio is denoted as: wherein P k represents the power of the kth vehicle user to the base station, g k,B [m] represents the channel gain of the kth vehicle to the base station on the mth frequency band, σ 2 represents the noise, is the interference suffered by the kth user vehicle in offloading, P m represents the transmit power of the mth vehicle for V2V communication, g m,B [m] represents the interference channel gain of the mth vehicle for V2V communication to V2I communication, p k' [m] = 1 indicates that this frequency band is used, p k' [m] = 0 indicates that this frequency band is not used; 2.4 Rate of the nth eavesdropper eavesdropping on the kth vehicle user in the mth sub-band is represented as: where denotes the power of the kth vehicle user, g k,n [m] denotes the channel gain from the kth vehicle to the eavesdropper in the mth frequency band, σ 2 denotes the noise, is the interference when eavesdropping, denotes the transmit power of the mth vehicle for V2V communication, g m,n [m] denotes the channel gain from the mth vehicle to the eavesdropper for V2V communication, p k' [m] = 1 indicates that this frequency band is used, p k' [m] = 0 indicates that this frequency band is not used.

4. The vehicle networking security computing offloading and resource allocation method of claim 1, wherein the process of the sixth step is as follows: 6.1 Divide the rewards into N levels according to the time of service delay w ; 6.2 When the computing offloading rate is too low, there will be a large delay, and the reward is 0. 5.The method of claim 1, wherein, The neural network training process of the seventh step is as follows: 7.1 Initialize environment information and Q-network parameters, generate vehicle running data; 7.2 In each training round, update and obtain the current vehicle position and environment state, reset the frequency band power selection, edge server resource allocation strategy; 7.3 According to the current state information and the greedy algorithm, select an action for the target vehicle, that is, the combination scheme of frequency band selection, vehicle power and edge server resource allocation, and update the information of the environment; 7.4 Obtain the action combination scheme of all target vehicles, obtain the reward value r related to the capacity c,i and the returned reward value {r t}; 7.5 Store the state, action, reward and next state at time t as a sample into the experience pool, 7.6 When the number of experience pool samples is sufficient, start training the model, randomly draw a small batch of samples (s t ,a t ,r t ,s t+1 ) from the experience pool, train the network parameters, and update the target network weights.

6. A computer device, comprising: The computer device comprises a memory and a processor, the memory stores a computer program, and the computer program is executed by the processor, so that the processor executes the steps of the vehicle Internet of Things security computing offloading and resource allocation method in any one of claims 1-5.

7. An information data processing terminal, characterized by The information data processing terminal is used for executing the steps of the vehicle Internet of Things security computing offloading and resource allocation method in any one of claims 1-5.

Citation Information

Patent Citations

  • Physical layer security resource allocation method in ICV network

    CN112153744A

  • Millimeter wave vehicle networking joint beam allocation and relay selection method

    CN113709701A