Drone-Assisted Resource Allocation Method, Apparatus and Electronic Device for Vehicular Network
The UAV-assisted Internet of Vehicle resource allocation method solves the problems of delay and service quality in vehicle resource allocation by predicting vehicle location and dynamic resource allocation, and improves the performance and stability of on-board communication.
Patent Information
- Application Number
- CN202210612465.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-31
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-05-31
AI Technical Summary
The existing vehicle resource allocation technical solutions ignore the vehicle's mobility and time-varying nature of resource requests, and cannot meet the delay limit and service quality requirements of vehicle tasks.
Through the UAV-assisted Vehicle Network resource allocation method, reinforcement learning and deep deterministic strategy gradient methods are used to predict vehicle locations, establish resource allocation models, and dynamically allocate resources to meet the delay and service quality requirements of vehicle tasks.
It realizes the rational allocation of resources to meet the delay limit and service quality requirements of vehicle tasks while considering vehicle mobility and resource time-varying, and improves the performance and stability of on-board communication.
Smart Images

Figure CN115002725B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of vehicle networking resource allocation, and particularly to a method, device, and electronic device for unmanned aerial vehicle-assisted vehicle networking resource allocation. Background Art
[0002] Vehicle networking is the product of the integration of the Internet and the Internet of Things, providing convenient and diverse services for intelligent transportation. Currently, the two major camps of vehicle networking technologies are the dedicated short range communications (DSRC) led by the United States and the long term evolution for vehicle-to-vehicle communication (LTE-V) system promoted by domestic enterprises. With the rapid development of industrial Internet of Things technology in vehicle networks, data exchange between vehicles, between vehicles and pedestrians, and between vehicles and infrastructure units is becoming increasingly frequent, requiring powerful data processing capabilities. During the process of providing services, it is necessary to continuously process information of surrounding vehicles, and the amount of data is extremely large. Therefore, reasonable vehicle networking resource allocation is crucial for reducing interference, improving network efficiency, and ultimately optimizing wireless communication performance.
[0003] Currently, most of the existing vehicle resource allocation technical solutions ignore the mobility of vehicles and the time-varying nature of resource requests, and cannot meet the delay constraints and quality of service requirements of vehicle tasks. Summary of the Invention
[0004] Embodiments of this application provide a method, device, and electronic device for unmanned aerial vehicle-assisted vehicle networking resource allocation, aiming to solve the technical problem that most of the existing vehicle resource allocation technical solutions ignore the mobility of vehicles and the time-varying nature of resource requests, and cannot meet the delay constraints and quality of service requirements of vehicle tasks.
[0005] In a first aspect, embodiments of this application provide a method for unmanned aerial vehicle-assisted vehicle networking resource allocation, including:
[0006] Predicting the vehicle position at the next moment of the vehicle according to the detected trajectory point data of the vehicle;
[0007] Receiving a task offloading request of the vehicle, and establishing a resource allocation model for vehicle networking resource allocation for the vehicle based on the task offloading request; the task offloading request includes a vehicle association mode, the amount of computing resources required for the task, the amount of data required for the task, and the maximum delay that the task can tolerate; the vehicle association mode includes a vehicle-unmanned aerial vehicle association mode and a vehicle-unmanned aerial vehicle non-association mode;
[0008] Solving the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy.
[0009] In one embodiment, the predicting the vehicle position at the next moment of the vehicle according to the detected trajectory point data of the vehicle includes:
[0010] Determine multiple trajectory point data of the detected vehicle;
[0011] Calculate the speed and acceleration corresponding to the multiple trajectory point data of the vehicle;
[0012] Calculate the distance between the vehicle and the drone, and the vehicle azimuth angle based on the speed and the acceleration of the multiple trajectory point data.
[0013] In one embodiment, the resource allocation model for vehicle Internet of Vehicles resource allocation established based on the task offloading request includes:
[0014] Establish a resource allocation model for vehicle Internet of Vehicles resource allocation based on the task offloading request, vehicle available resources, drone available resources, transmission rate of the uplink between the vehicle and the drone, and transmission rate of the vehicle's own link.
[0015] In one embodiment, solving the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy includes:
[0016] Obtain an optimization problem function based on the resource allocation model;
[0017] Convert the optimization problem function based on the reinforcement learning method;
[0018] Solve the converted result according to the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy.
[0019] In one embodiment, converting the optimization problem function based on the reinforcement learning method includes:
[0020] Convert the optimization problem function into an environmental state space set, an action decision set, and a reward function;
[0021] The environmental state space set includes the amount of computing resources required for the vehicle task, the amount of data required for the task, the maximum latency tolerated by the task, the vehicle position, and the drone position;
[0022] The action decision set includes the vehicle association mode, and the proportion of computing resources and cache resources allocated by the drone to the associated vehicle;
[0023] The reward function is based on the maximum latency tolerated by the vehicle task and the construction of the cache resources possessed by the vehicle.
[0024] In one embodiment, solving the converted result according to the deep deterministic policy gradient method to obtain the vehicle-associated optimal mode and resource allocation strategy includes
[0025] Initializing the network parameters of the deep deterministic policy gradient method, selecting an action decision from the action decision set based on the state of the environmental state space set and executing it to obtain the reward function;
[0026] Training the network of the deep deterministic policy gradient method based on the empirical data as a training set, updating the network parameters, and obtaining the vehicle-associated optimal mode and resource allocation strategy.
[0027] In a second aspect, an embodiment of the present application provides a UAV-assisted vehicle network resource allocation device, including:
[0028] A vehicle position prediction module, configured to predict the vehicle position of the vehicle at the next moment according to the detected trajectory point data of the vehicle;
[0029] A resource allocation model establishment module, configured to receive the task offloading request of the vehicle, and establish a resource allocation model for vehicle network resource allocation based on the task offloading request; the task offloading request includes a vehicle association mode, the amount of computing resources required for the task, the amount of data required for the task, and the maximum delay that the task can tolerate; the vehicle association mode includes a vehicle-UAV association mode and a vehicle-non-UAV association mode;
[0030] A solving module, configured to solve the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the vehicle-associated optimal mode and resource allocation strategy.
[0031] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory storing a computer program, and the processor implements the steps of the UAV-assisted vehicle network resource allocation method described in the first aspect when executing the program.
[0032] In a fourth aspect, an embodiment of the present application provides a non-transitory computer-readable storage medium, on which a computer program is stored, and the computer program implements the steps of the UAV-assisted vehicle network resource allocation method described in the first aspect when executed by a processor.
[0033] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and the computer program implements the steps of the UAV-assisted vehicle network resource allocation method described in the first aspect when executed by a processor.
[0034] The method for resource allocation in a drone-assisted vehicle network provided by the embodiments of this application (invention name) predicts the vehicle position at the next moment based on the detected trajectory point data of the vehicle, thereby considering the mobility of the vehicle and enabling timely interaction of communication data between vehicles. The embodiments of this application receive the task offloading requests of the vehicles and establish a resource allocation model for resource allocation in the vehicle network based on the task offloading requests. The resource allocation model is solved based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy. Therefore, the embodiments of this application consider the time-varying nature of resources, enabling limited resources to be reasonably and dynamically allocated to vehicles requesting resources, thereby meeting the delay constraints and quality of service requirements of vehicle tasks, improving the performance of vehicle-mounted communication, and showing stability and high convergence in the optimization of a series of continuous vehicle action spaces. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] To more clearly illustrate the technical solutions in this application or the prior art, the following briefly introduces the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 is one of the schematic flowcharts of the method for resource allocation in a drone-assisted vehicle network provided by the embodiments of this application;
[0037] Figure 2 is the drone-assisted vehicle network scenario provided by the embodiments of this application;
[0038] Figure 3 is the second schematic flowchart of the method for resource allocation in a drone-assisted vehicle network provided by the embodiments of this application;
[0039] Figure 4 is the third schematic flowchart of the method for resource allocation in a drone-assisted vehicle network provided by the embodiments of this application;
[0040] Figure 5 is the deep reinforcement learning model based on the deep deterministic policy gradient method provided by the embodiments of this application;
[0041] Figure 6 is the schematic flowchart of the algorithm based on the deep deterministic policy gradient method provided by the embodiments of this application;
[0042] Figure 7 is the schematic structural diagram of the device for resource allocation in a drone-assisted vehicle network provided by the embodiments of this application;
[0043] Figure 8This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0044] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0045] Figure 1 It is a method for resource allocation in a vehicle network assisted by a drone. Please refer to Figure 1 An embodiment of the present application provides a method for resource allocation in a vehicle network assisted by a drone, which may include:
[0046] Step 100: Predict the vehicle position of the vehicle at the next moment according to the detected trajectory point data of the vehicle;
[0047] An electronic device predicts the vehicle position of the vehicle at the next moment according to the detected trajectory point data of the vehicle. Among them, the electronic device may be a drone. In an embodiment of the present application, please refer to Figure 2 Figure 2 It represents the scenario of a vehicle network assisted by a drone in an embodiment of the present application. The scenario of a vehicle network assisted by a drone consists of N vehicles in a straight two-way road and M drones deployed on aerial rotors. Effective resource allocation is performed in the vehicle network to maximize the tasks successfully completed by the vehicles and drones. The tasks successfully completed by the vehicles and drones refer to the association between the vehicles and the drones, and the vehicles unload the tasks to the MEC (Multi-access Edge Computing) server of the drones for execution.
[0048] In one embodiment, please refer to Figure 3 Step 100: The predicting the vehicle position of the vehicle at the next moment according to the detected trajectory point data of the vehicle includes:
[0049] Step 110: Determine multiple trajectory point data of the detected vehicle;
[0050] An electronic device determines multiple trajectory point data of the detected vehicle. Specifically, in practice, the multiple trajectory point data of the vehicle can be sensed by the radar devices of multiple different drones. For example, when there are S drones, the S drones can sense three trajectory points of the vehicle through the radar devices. Assume that a certain vehicle is within the coverage range of the S drones, then the three trajectory point sets sensed by the S drones are respectively:
[0051] Set a: {(xn,1, y n,1 ),(x n,2 ,y n,2 ),...,(x n,S ,y n,S )};
[0052] Set b: {(x (n-1),1, y (n-1),1 ),(x (n-1),2, y (n-1),2 ),...,(x (n-1),S, y (n-1),S )};
[0053] Set c: {(x (n-2),1, y (n-2),1 ),(x (n-2),2, y (n-2),2 ),...,(x (n-2),S, y (n-2),S )};
[0054] In the application embodiment, three fused trajectory points of the vehicle are obtained by weighted averaging of three trajectory point data, (x n ,y n ), (x n-1 ,y n-1 ), (x n-2 ,y n-2 ), which is expressed according to the weighted averaging formula as:
[0055]
[0056]
[0057]
[0058] Among them, (x n ,y n ), (x n-1 ,y n-1 ), (x n-2 ,y n_2 ) can be used as the multiple trajectory point data of the detected vehicle in the application embodiment.
[0059] Step 120, calculate the speed and acceleration corresponding to the multiple trajectory point data of the vehicle;
[0060] The electronic device calculates the speed and acceleration corresponding to the multiple trajectory point data of the vehicle. Specifically, in the application embodiment, the calculation of the speed and acceleration corresponding to the multiple trajectory point data of the vehicle is expressed by the following formula:
[0061]
[0062] Among them, v represents speed, a represents acceleration, and ΔT represents the time interval from this moment to the next moment.
[0063] Step 130: Calculate the distance between the vehicle and the drone, and the vehicle azimuth angle according to the speed and the acceleration of the multiple trajectory point data.
[0064] The electronic device calculates the distance between the vehicle and the drone, and the vehicle azimuth angle according to the speed and the acceleration of the multiple trajectory point data. Specifically, considering that the acceleration of the vehicle changes little from the (n - 2)-th time slot to the n-th time slot, that is, a x,n ≈a x,n_1 ≈a x,n-2 , predict the position coordinate parameters of the vehicle at the next moment according to the state information of the vehicle corresponding to the three trajectory point data, including the azimuth angle of the vehicle. Calculate the distance on the x and y axes, and then calculate the distance between the vehicle and the corresponding drone according to the known fixed flight height H of the drone and the Pythagorean theorem. The formula is:
[0065]
[0066] y n+1|n =3y n -3y n-1 +y n_2 ≈3y n -3y n-1 +y n-2 ,
[0067]
[0068]
[0069] Among them, represents the distance component of the vehicle on the x-axis from the n-th time slot to the (n + 1)-th time slot, represents the distance component of the vehicle on the y-axis from the n-th time slot to the (n + 1)-th time slot, H represents the fixed flight height of the drone, represents the straight-line distance between the vehicle and the drone from the n-th time slot to the (n + 1)-th time slot, represents the azimuth angle of the vehicle at the (n + 1)-th time slot.
[0070] In this embodiment, the trajectory point data of the vehicle is detected by using the drone radar device, and the next position of the vehicle is predicted based on the trajectory point data by using the prediction formula, so as to realize the perception of the vehicle position and timely interact the communication data between the vehicles.
[0071] Step 200: Receive the task offloading request of the vehicle, and establish a resource allocation model for vehicle Internet of Vehicles resource allocation based on the task offloading request; the task offloading request includes the vehicle association mode, the amount of computing resources required for the task, the amount of data required for the task, and the maximum latency that the task can tolerate; the vehicle association mode includes the vehicle - drone association mode and the vehicle - non - drone association mode;
[0072] After the drone predicts the next - moment position of the vehicle, the vehicle will randomly generate different computing tasks and send a task offloading request to the drone as needed. The task offloading request includes the vehicle association mode. The vehicle association mode considers b i as a binary variable, b i (t) = {0, 1}, b i = 1 indicates that the vehicle is associated with the drone and offloads the task to the MEC server of the drone for execution; b i = 0 indicates that the vehicle is not associated with the drone and executes the computing task by itself. The task offloading request sent by vehicle i at time t is represented by where are respectively the amount of computing resources required for the task, the amount of data required for the task, and the maximum latency that the task can tolerate.
[0073] The electronic device receives the task offloading request of the vehicle and establishes a resource allocation model for vehicle Internet of Vehicles resource allocation based on the task offloading request.
[0074] In one embodiment, Step 200, the establishing a resource allocation model for vehicle Internet of Vehicles resource allocation based on the task offloading request specifically includes:
[0075] The electronic device establishes a resource allocation model for vehicle Internet of Vehicles resource allocation based on the task offloading request, the vehicle's available resources, the drone's available resources, the transmission rate of the uplink between the vehicle and the drone, and the transmission rate of the vehicle's own link.
[0076] Among them, the vehicle's available resources include the computing resources and cache resources possessed by the vehicle. The computing resources and cache resources of the vehicle that executes the task by itself are represented by where i represents the serial number of the vehicle. The drone's available resources include the available computing and cache resources of the drone. The available computing and cache resources of the drone are respectively represented by where j represents the serial number of the drone. The drone will allocate computing resources and cache resources to the associated vehicles, and the allocation ratios are respectively represented by ; the condition for the drone and the vehicle to successfully complete the task is that the drone's cache resources should be greater than or equal to the amount of data required for the task, that is For vehicle \(i\in N(t)\), the total time from task generation to receiving the processing result is \(T\). i (t) is expressed as:
[0077]
[0078] where \(T\). i (t) is the resource allocation model for vehicle Internet of Vehicles resource allocation, \(e\). j,i (t) is the transmission rate of the vehicle - to - UAV uplink, \(e\). i (t) is the transmission rate of the vehicle's own link.
[0079] In the embodiments of the present application, after the vehicle task is generated, the association mode between the vehicle and the UAV is determined, the dynamic resource allocation ratio of the Internet of Vehicles is reasonably adjusted, and a resource management model is established to maximize the number of tasks successfully completed by the vehicle and the UAV, so as to meet the delay limit and quality - of - service requirements of the vehicle task.
[0080] Step 300: Solve the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and the resource allocation strategy.
[0081] The electronic device solves the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and the resource allocation strategy.
[0082] Specifically, in one embodiment, please refer to Figure 4 , the solving of the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and the resource allocation strategy includes:
[0083] Step 310: Obtain the optimization problem function based on the resource allocation model.
[0084] After the electronic device establishes the resource allocation model, it obtains the optimization problem function \(F\), which is expressed as the following formula:
[0085]
[0086]
[0087] where \(b(t)\) is the vehicle association mode matrix, \(f\). co (t) is the allocation matrix of UAV computing resources, \(f\). ca (t) is the UAV cache resource allocation matrix, \(H(\cdot)\) is the step function, which is 1 when the variable \(\geq0\) and 0 otherwise. That is, for vehicle \(i\) that has been allocated sufficient cache resources and meets the task delay requirements, there is or Its constraint condition is to maximize the utilization of computing resources and cache resources.
[0088] Step 320: Transform the optimization problem function based on the reinforcement learning method;
[0089] The electronic device transforms the optimization problem function based on the reinforcement learning method. Since the optimization problem function F is a non-convex function with high complexity, the embodiment of the present application uses the reinforcement learning method to transform the optimization problem function F.
[0090] Specifically, in one embodiment, step 320, the transformation of the optimization problem function based on the reinforcement learning method includes:
[0091] Transform the optimization problem function into an environmental state space set, an action decision set, and a reward function. The environmental state space set includes the computing resource amount required for the vehicle task, the data amount required for the task, the maximum delay that the task can tolerate, the vehicle position, and the drone position; the action decision set includes the vehicle association mode, and the computing resource ratio and cache resource ratio allocated by the drone to the associated vehicle; the reward function is constructed based on the maximum delay that the vehicle task can tolerate and the cache resources possessed by the vehicle.
[0092] Among them, the environmental state space set S is expressed as the following set:
[0093]
[0094]
[0095] x’1(t),x’2(t),...,x’ M (t),y’1(t),y’2(t),...,y’ M (t),z’1(t),z’2(t),...,z’ M (t)};
[0096] Let the number of vehicles associated with the drone be N’ j (j∈{1,2,…,M}), define the action space A, the drone selects the vehicle association mode and the computing resource ratio and cache resource ratio allocated by the drone to the associated vehicle at time t according to the current policy π, and define the action decision set as a(t), that is:
[0097]
[0098] After executing the action decision a(t) in the environmental state space set s(t), a reward is returned to the drone, defined as the reward function R, expressed as:
[0099]
[0100] The reward function guides the UAV to update its strategy, where two reward elements are defined, expressed as:
[0101]
[0102]
[0103] where are respectively the amount of computing resources required for the task, the amount of data required for the task, and the maximum latency that the task can tolerate. The computing resources and cache resources of the vehicle itself for executing the task are represented by denote.
[0104] Step 330: Solve the converted result according to the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy.
[0105] The electronic device solves the converted result according to the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy.
[0106] The electronic device uses the deep deterministic policy gradient method to solve the converted optimization problem function. The deep reinforcement learning model based on the deep deterministic policy gradient method is as Figure 5 shown.
[0107] In the embodiment of the present application, the UAV (electronic device) is used as an agent, selects an action based on the current state and executes it, obtains a reward function, and updates and selects the optimal strategy through feedback. According to the determined S, A, R, the evaluation function Q is obtained, expressed as
[0108]
[0109] where E represents expectation, γ is the discount factor of r(t), r(t) represents the immediate reward returned to the UAV at time t, which is defined as the average reward of the vehicle, and τ << 1.
[0110] In one embodiment, Step 330, the solving the converted result according to the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy includes
[0111] Step 331: Initialize the network parameters of the deep deterministic policy gradient method, select an action decision from the action decision set based on the state of the environmental state space set and execute it to obtain the reward function;
[0112] Step 332: Train the network of the deep deterministic policy gradient method based on the empirical data as a training set, update the network parameters, and obtain the optimal vehicle association mode and resource allocation strategy.
[0113] Specifically, the DDPG method (i.e., the Deep Deterministic Policy Gradient method) has two networks, an Actor network and a Critic network. The Actor network is used to generate the current policy, and the Critic network is used to evaluate the quality of the policy in the current state. The schematic diagram of the algorithm process based on the Deep Deterministic Policy Gradient method in the embodiments of the present application is as shown in Figure 6 shown below.
[0114] To improve the stability of training, a Target-Actor network and a Target-Critic network are introduced. Step 330 in the embodiments of the present application includes the following specific steps:
[0115] 1) Initialize the Actor network π and the Critic network Q, as well as the network parameters θ π and θ Q ;
[0116] 2) Initialize the Target-Actor network π′ and the Target-Critic network Q′, as well as the network parameters θ π′ and θ Q′
[0117] 3) Initialize the successful experience cache pool R success and the failed experience cache pool R failure .
[0118] 4) For each episode, loop the following steps:
[0119] (1) Select the initial state s1;
[0120] (2) For each step in the episode, loop the following steps:
[0121] ① According to the current input state s t and the Actor network, execute to output the action a t , obtain the immediate reward r t and the next state s t+1 , and further obtain the experience data (s t , a t , r t , s t+1 );
[0122] ② Judge whether this round of learning terminates. If it does not terminate, store the experience data (s t , a t , r t , s t+1 ) into the successful experience cache pool R success , otherwise execute ③;
[0123] ③ Store the empirical data (s t , a t , r t , s t+1 ) into the failure experience cache pool R failure . Take out N success empirical data from R failure and also put them into R failure ;
[0124] ④ Randomly sample m empirical data (s i , a i , r i , S i+1 ) from the two experience pools, where i ≤ m;
[0125] ⑤ Calculate the expected return of the current action through the Target-Critic network:
[0126] y i = r i + γQ′(s i+1 , π′, θ Q′ )
[0127] ⑥ Define the loss function minimized by the Critic network to update the network parameters:
[0128]
[0129] ⑦ Update the Actor network parameters through the following gradient:
[0130]
[0131] ⑧ Update the Target-Actor network and Target-Critic network parameters through the following equations:
[0132]
[0133] (3) End the step loop.
[0134] 5) End the episode loop.
[0135] After the training of the training set is completed, the optimization objective function is solved. The embodiments of the present application obtain the optimal vehicle association mode and resource allocation ratio strategy, achieving the purpose of meeting the vehicle delay limit and service quality requirements, thereby improving the performance of vehicle communication.
[0136] The embodiments of this application use the reinforcement learning method and the Deep Deterministic Policy Gradient (DDPG) method to transform and solve the optimization problem function, effectively making joint decisions on vehicle association patterns and resource allocation in the vehicle network, meeting the delay constraints of vehicles and the requirements of task service quality, improving the performance of vehicle-mounted communication, and showing stability and high convergence in the optimization of a series of continuous vehicle action spaces.
[0137] By predicting the vehicle position at the next moment of the vehicle according to the detected trajectory point data of the vehicle, the mobility of the vehicle is considered to timely interact the communication data between vehicles; the embodiments of this application receive the task offloading request of the vehicle and establish a resource allocation model for vehicle network resource allocation based on the task offloading request; the resource allocation model is solved based on the reinforcement learning method and the Deep Deterministic Policy Gradient method to obtain the optimal vehicle association mode and resource allocation strategy. Therefore, the embodiments of this application consider the time-varying nature of resources, enabling limited resources to be reasonably and dynamically allocated to vehicles requesting resources, thus meeting the delay limit and service quality requirements of vehicle tasks, improving the performance of vehicle-mounted communication, and showing stability and high convergence in the optimization of a series of continuous vehicle action spaces.
[0138] The drone-assisted vehicle network resource allocation device provided by the embodiments of this application will be described below. The drone-assisted vehicle network resource allocation device described below can be mutually referred to with the drone-assisted vehicle network resource allocation method described above.
[0139] Please refer to Figure 7 , the embodiments of this application provide a drone-assisted vehicle network resource allocation device, including:
[0140] A vehicle position prediction module 201, configured to predict the vehicle position at the next moment of the vehicle according to the detected trajectory point data of the vehicle;
[0141] A resource allocation model establishment module 202, configured to receive the task offloading request of the vehicle and establish a resource allocation model for vehicle network resource allocation based on the task offloading request; the task offloading request includes a vehicle association mode, the amount of computing resources required for the task, the amount of data required for the task, and the maximum delay that the task can tolerate; the vehicle association mode includes a vehicle-drone association mode and a vehicle-non-drone association mode;
[0142] A solution module 203, configured to solve the resource allocation model based on the reinforcement learning method and the Deep Deterministic Policy Gradient method to obtain the optimal vehicle association mode and resource allocation strategy.
[0143] The UAV-assisted vehicle networking resource allocation device according to the embodiments of the present application predicts the vehicle position at the next moment of the vehicle based on the detected trajectory point data of the vehicle, thereby considering the mobility of the vehicle and timely performing the interaction of communication data between vehicles; the embodiments of the present application receive the task offloading request of the vehicle, and establish a resource allocation model for vehicle networking resource allocation based on the task offloading request; the resource allocation model is solved based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and the resource allocation strategy. Therefore, the embodiments of the present application consider the time-varying nature of resources, enabling limited resources to be reasonably and dynamically allocated to vehicles requesting resources, thereby meeting the delay constraints and quality of service requirements of vehicle tasks, improving the performance of vehicle-mounted communication, and showing stability and high convergence in the optimization of a series of continuous vehicle action spaces.
[0144] In one embodiment, the vehicle position prediction module includes:
[0145] A trajectory point data determination module, configured to determine a plurality of trajectory point data of the detected vehicle;
[0146] A speed and acceleration calculation module, configured to calculate the speed and acceleration corresponding to the plurality of trajectory point data of the vehicle;
[0147] A position prediction module, configured to calculate the distance between the vehicle and the UAV and the vehicle azimuth angle according to the speed and the acceleration of the plurality of trajectory point data.
[0148] In one embodiment, the resource allocation model establishment module is specifically configured to establish a resource allocation model for vehicle networking resource allocation for the vehicle based on the task offloading request, vehicle available resources, UAV available resources, the transmission rate of the uplink between the vehicle and the UAV, and the transmission rate of the vehicle's own link.
[0149] In one embodiment, the solving module includes:
[0150] An optimization problem function acquisition module, configured to acquire an optimization problem function based on the resource allocation model;
[0151] A conversion module, configured to convert the optimization problem function based on the reinforcement learning method;
[0152] A final solving module, configured to solve the converted result according to the deep deterministic policy gradient method to obtain the optimal vehicle association mode and the resource allocation strategy.
[0153] In one embodiment, the conversion module is specifically configured to convert the optimization problem function into an environmental state space set, an action decision set, and a reward function;
[0154] The environmental state space set includes the amount of computing resources required for vehicle tasks, the amount of data required for tasks, the maximum latency that tasks can tolerate, vehicle positions, and drone positions;
[0155] The action decision set includes vehicle association modes, as well as the proportion of computing resources and the proportion of cache resources allocated by drones to associated vehicles;
[0156] The reward function is constructed based on the maximum latency that vehicle tasks can tolerate and the cache resources available to the vehicles.
[0157] In one embodiment, the final solution module is configured to:
[0158] Initialize the network parameters of the deep deterministic policy gradient method, select an action decision from the action decision set based on the state of the environmental state space set and execute it to obtain the reward function;
[0159] Train the network of the deep deterministic policy gradient method using empirical data as a training set, update the network parameters, and obtain the optimal vehicle association mode and resource allocation strategy.
[0160] Figure 8 An example of the physical structure diagram of an electronic device is shown as Figure 8 shown. The electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 communicate with each other through the communication bus 840. The processor 810 can call a computer program in the memory 830 to execute the steps of the method for resource allocation in a drone-assisted vehicle network, such as: predicting the vehicle position at the next moment of the vehicle according to the detected trajectory point data of the vehicle; receiving the task offloading request of the vehicle, and establishing a resource allocation model for vehicle network resource allocation based on the task offloading request; the task offloading request includes a vehicle association mode, the amount of computing resources required for the task, the amount of data required for the task, and the maximum latency that the task can tolerate; the vehicle association mode includes a vehicle-drone association mode and a vehicle-non-drone association mode; solving the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy.
[0161] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0162] On the other hand, an embodiment of this application also provides a computer program product. The computer program product includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the steps of the method for drone-assisted vehicle networking resource allocation provided in the above-mentioned various embodiments, for example, including: predicting the vehicle position at the next moment of the vehicle based on the detected trajectory point data of the vehicle; receiving the task offloading request of the vehicle, and establishing a resource allocation model for vehicle networking resource allocation for the vehicle based on the task offloading request; the task offloading request includes the vehicle association mode, the amount of computing resources required for the task, the amount of data required for the task, and the maximum delay that the task can tolerate; the vehicle association mode includes the vehicle-unmanned aircraft association mode and the vehicle-non-unmanned aircraft association mode; solving the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and the resource allocation strategy.
[0163] On the other hand, an embodiment of this application also provides a processor-readable storage medium. The processor-readable storage medium stores a computer program. The computer program is used to cause the processor to execute the steps of the method provided in the above-mentioned various embodiments, for example, including: predicting the vehicle position at the next moment of the vehicle based on the detected trajectory point data of the vehicle; receiving the task offloading request of the vehicle, and establishing a resource allocation model for vehicle networking resource allocation for the vehicle based on the task offloading request; the task offloading request includes the vehicle association mode, the amount of computing resources required for the task, the amount of data required for the task, and the maximum delay that the task can tolerate; the vehicle association mode includes the vehicle-unmanned aircraft association mode and the vehicle-non-unmanned aircraft association mode; solving the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and the resource allocation strategy.
[0164] The processor-readable storage medium may be any available medium or data storage device accessible by the processor, including but not limited to magnetic memory (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memory (such as CDs, DVDs, BDs, HVDs, etc.), and semiconductor memory (such as ROM, EPROM, EEPROM, non-volatile memory (NANDFLASH), solid state drives (SSD)), etc.
[0165] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, i.e., they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative effort.
[0166] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical disks, etc., including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0167] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present application.
Claims
1. A method for resource allocation in a drone-assisted vehicle network, characterized in that Including: Predicting the vehicle position at the next moment of the vehicle according to the detected trajectory point data of the vehicle; Receiving the task offloading request of the vehicle, and establishing a resource allocation model for vehicle Internet of Vehicles resource allocation based on the task offloading request; the task offloading request includes the vehicle association mode, the amount of computing resources required for the task, the amount of data required for the task, and the maximum delay that the task can tolerate; the vehicle association mode includes the vehicle - drone association mode and the vehicle - non - drone association mode; Solving the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy; The establishing a resource allocation model for vehicle Internet of Vehicles resource allocation based on the task offloading request includes: Establishing a resource allocation model for vehicle Internet of Vehicles resource allocation based on the task offloading request, vehicle available resources, drone available resources, the transmission rate of the uplink between the vehicle and the drone, and the transmission rate of the vehicle's own link; The resource allocation model is: ; Among them, is the resource allocation model; is the transmission rate of the uplink of the vehicle and the drone; is the transmission rate of the vehicle's own link; the task offloading request is represented by where are respectively the amount of computing resources required for the task, the amount of data required for the task, and the maximum delay that the task can tolerate; represents the serial number of the drone; represents the available computing resources of the drone; represents the proportion of the computing resources that the drone will allocate to the vehicle; the available resources of the vehicle include the computing resources and the cache resources possessed by the vehicle, represents the computing resources possessed by the vehicle, represents the cache resources possessed by the vehicle; represents the vehicle association mode; The solving the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy includes: Obtaining an optimization problem function based on the resource allocation model; The optimization problem function is: ; Among them, represents the vehicle serial number, N(t) represents the number of vehicles at time t; F is the optimization problem function; is the vehicle association mode matrix; is the allocation matrix of the UAV computing resources; is the UAV cache resource allocation matrix; is the step function, when the variable has a value of 1, otherwise 0; represents the proportion of the cache resources that the UAV will allocate to the vehicle; represents the available cache resources of the UAV; Converting the optimization problem function based on the reinforcement learning method; Solving the converted result according to the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy; The converting the optimization problem function based on the reinforcement learning method includes: Converting the optimization problem function into an environmental state space set, an action decision set, and a reward function; The reward function is: ; The reward function is used to guide the UAV to update its strategy, and are two reward elements, where: ; ; The environmental state space set includes the amount of computing resources required for the vehicle task, the amount of data required for the task, the maximum delay that the task can tolerate, the vehicle position and the drone position at the next moment of the vehicle; The action decision set includes the vehicle association mode, and the proportion of computing resources and cache resources allocated by the drone to the associated vehicle; The reward function is constructed based on the maximum delay that the vehicle task can tolerate and the cache resources of the vehicle.
2. The method for drone-assisted vehicle network resource allocation according to claim 1, wherein The predicting the vehicle position at the next moment of the vehicle according to the detected vehicle trajectory point data includes: Determining a plurality of trajectory point data of the detected vehicle; Calculating the speed and acceleration corresponding to the plurality of trajectory point data of the vehicle; Calculating the distance between the vehicle and the drone, and the vehicle azimuth angle according to the speed and the acceleration of the plurality of trajectory point data.
3. The method for drone-assisted vehicle network resource allocation according to claim 1, wherein The solving the converted result according to the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy includes Initializing the network parameters of the deep deterministic policy gradient method, selecting the action decision of the action decision set based on the state of the environmental state space set and executing it to obtain the reward function; Training the network of the deep deterministic policy gradient method based on the empirical data as the training set, updating the network parameters, and obtaining the optimal vehicle association mode and resource allocation strategy.
4. An unmanned aerial vehicle-assisted vehicle-to-everything (V2X) resource allocation device, characterized in that Including: A vehicle position prediction module, configured to predict the vehicle position of the vehicle at the next moment according to the detected trajectory point data of the vehicle; A resource allocation model establishment module, configured to receive the task offloading request of the vehicle, and establish a resource allocation model for vehicle Internet of Vehicles resource allocation based on the task offloading request; the task offloading request includes a vehicle association mode, the amount of computing resources required for the task, the amount of data required for the task, and the maximum latency that the task can tolerate; the vehicle association mode includes a vehicle and drone association mode and a vehicle non - association with drone mode; A solution module, configured to solve the resource allocation model based on the reinforcement learning method and the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy; The resource allocation model establishment module is specifically configured to establish a resource allocation model for vehicle Internet of Vehicles resource allocation for the vehicle based on the task offloading request, vehicle available resources, drone available resources, the transmission rate of the uplink between the vehicle and the drone, and the transmission rate of the vehicle's own link; The resource allocation model is: ; wherein, is the resource allocation model; is the transmission rate of the uplink of the vehicle and the UAV; is the transmission rate of the vehicle's own link; the task offloading request is represented by wherein are respectively the amount of computing resources required for the task, the amount of data required for the task, and the maximum latency tolerated by the task; represents the serial number of the UAV; represents the available computing resources of the UAV; represents the proportion of computing resources that the UAV will allocate to the vehicle; the available resources of the vehicle include the computing resources and the cache resources possessed by the vehicle, represents the computing resources possessed by the vehicle, represents the cache resources possessed by the vehicle; represents the vehicle association mode; The solution module includes: an optimization problem function acquisition module, configured to obtain an optimization problem function based on the resource allocation model; a conversion module, configured to convert the optimization problem function based on the reinforcement learning method; a final solution module, configured to solve the converted result according to the deep deterministic policy gradient method to obtain the optimal vehicle association mode and resource allocation strategy; The optimization problem function is: ; Among them, represents the vehicle serial number, N(t) represents the number of vehicles at time t; F is the optimization problem function; is the vehicle association mode matrix; is the allocation matrix of UAV computing resources; is the UAV cache resource allocation matrix; is the step function, when the variable the value is 1, otherwise 0; represents the proportion of cache resources that the UAV will allocate to the vehicle; represents the available cache resources of the UAV; The conversion module is specifically configured to convert the optimization problem function into an environmental state space set, an action decision set, and a reward function; The reward function is: ; The reward function is used to guide the UAV to update its strategy, and are two reward elements, where: ; ; The environmental state space set includes the amount of computing resources required for the vehicle task, the amount of data required for the task, the maximum latency that the task can tolerate, the vehicle position and drone position of the vehicle at the next moment; the action decision set includes the vehicle association mode, and the proportion of computing resources and cache resources allocated by the drone to the associated vehicle; the reward function is constructed based on the maximum latency that the vehicle task can tolerate and the cache resources possessed by the vehicle.
5. An electronic device, comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the drone - assisted vehicle Internet of Vehicles resource allocation method according to any one of claims 1 to 3 are implemented.
6. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the drone - assisted vehicle Internet of Vehicles resource allocation method according to any one of claims 1 to 3 are implemented.
7. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the drone - assisted vehicle Internet of Vehicles resource allocation method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Unmanned aerial vehicle task unloading method and system based on reinforcement learning in edge calculation
CN111787509A
Vehicle edge computing task unloading method based on depth deterministic strategy
CN113760511A