Real-time traffic data collection method for air-land collaboration based on deep reinforcement learning
By optimizing the drone trajectory, access strategy and bandwidth allocation based on deep reinforcement learning, the problem of resource waste and complexity in data collection in Internet of Vehicles is solved, the success rate of data transmission is improved and the real traffic scenario is adapted.
Patent Information
- Application Number
- CN202310126694.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2043-02-16
AI Technical Summary
The existing Internet of Vehicle data collection schemes are difficult to effectively coordinate drone networks and ground networks, resulting in increased resource waste and complexity, and the assumed traffic model is too simple to meet the needs of the actual situation.
The air-to-land collaborative real-time transportation data collection method based on deep reinforcement learning is adopted to maximize the data transmission success rate by jointly optimizing the drone trajectory, access strategy and communication bandwidth allocation.
It improves the success rate of data transmission in the Internet of Vehicles scenario, optimizes access strategy, bandwidth allocation and drone trajectory, solves the non-convex optimization problem that is difficult to deal with by traditional methods, and considers real road section information and traffic factors.
Smart Images

Figure CN116112897B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of vehicle networking communications, and relates to unmanned aerial vehicle (UAV) technology and deep reinforcement learning (DRL) technology, and specifically to an air-land collaborative real-time traffic data collection method based on deep reinforcement learning. Background Art
[0002] With the development of autonomous driving technology, traffic environment perception technology and vehicle multimedia, heterogeneous vehicle networks have gradually evolved into the Internet of Things (IoV). In the IoV, any user can enjoy a wealth of network services, but such services require the support of large amounts of real-time traffic data. Therefore, how to efficiently collect data from vehicles to network nodes such as road side units (RSUs) has become a problem worth studying.
[0003] Drones are highly maneuverable and flexible devices. In situations where users are densely populated and the number of accesses is huge, drones can effectively collect data in the Internet of Vehicles, greatly expand the service scope of the Internet of Vehicles, and relieve the pressure on other communication equipment.
[0004] This technology of using drones to assist ground communication equipment in data transmission is called air-ground collaborative network; there are already many cases of applying drones to the Internet of Things, Internet of Vehicles and other traditional ground communication networks. The high mobility of drones allows them to be quickly deployed to designated locations, expanding the coverage of the network and improving the communication capabilities of the network.
[0005] However, existing research simply applies drones to various scenarios without organically combining them with the relevant communication equipment that is already available in the scenarios, resulting in a waste of resources.
[0006] Secondly, the high maneuverability of drones not only expands the communication range, but also increases the complexity of the air-ground collaborative network, which makes it very difficult to design a reasonable and efficient Internet of Vehicles data collection solution.
[0007] Therefore, how to coordinate the drone network and the ground network, reasonably allocate limited communication resources, and plan the flight trajectory of the drone are all difficult problems facing the current Internet of Vehicles.
[0008] In many inventions or studies involving the Internet of Vehicles, the assumed traffic models are too simple and do not conform to the actual situation. In many studies involving drone position optimization, the optimal position of the drone can only be given for a single fixed situation. In fact, the flight process of the drone is continuous, and this optimization solution based entirely on discrete states is difficult to be truly applied. Finally, since the relevant optimization problem is a non-convex optimization problem coupled with the time dimension, related work often decomposes the problem into several approximate sub-problems and then solves them separately. The approximate solution obtained in this way is often far from the optimal solution of the problem, and it cannot solve more complex problems.
[0009] In general, only by considering the complete air-land collaborative collection process can real optimization be achieved, which is exactly what is lacking in existing work. In addition, real road section information, traffic light information, and vehicle movement conditions, which are very important factors in real-time traffic, are also ignored by many existing works. Summary of the invention
[0010] The present invention aims at optimizing the air-land collaborative network under the background of real-time traffic, taking into account traffic lights and other traffic factors. In order to improve the data transmission success rate (DTSR), a real-time traffic data collection method for air-land collaboration based on deep reinforcement learning is used. By jointly optimizing the trajectory of drones, access strategy and communication bandwidth allocation, DTSR is maximized under conditional constraints.
[0011] The specific steps of the air-land collaborative real-time traffic data collection method based on deep reinforcement learning are as follows:
[0012] Step 1: Build a real-time traffic-based air-ground collaborative vehicle network data collection architecture that includes a roadside unit (RSU), a UAV, and several vehicles;
[0013] Among them, the roadside unit RSU and the UAV are both the receiving ends of vehicle data;
[0014] Step 2: Each vehicle will generate a data packet of size D to upload in each time slot, and the RSU and UAV jointly collect the data generated by all vehicles;
[0015] Use collection and They represent all vehicles in the scene, the receiver of the data, and all time slots considered, respectively.
[0016] Step 3: For the mth vehicle, the vehicle can only select one receiving end in the same time slot, and set the constraint conditions between the vehicle and the data unloading to the receiving end k;
[0017] The constraints are:
[0018] a m,k (t)∈{0,1}, A binary variable representing the offloading decision. When the mth vehicle needs to offload its data to the receiver k, a m,k (t)=1, otherwise a m,k (t)=0.
[0019] Step 4: When the mth vehicle communicates with the receiving end k, the channel rate R is calculated based on the bandwidth allocation. m,k ;
[0020] The expression is:
[0021]
[0022] where b m,k is the bandwidth allocation factor, satisfying the constraint B k is the available bandwidth; P0, A r , n0 represent the transmission power, antenna gain and noise power density at the receiving end respectively; PL m,k is the average path loss of line-of-sight and non-line-of-sight channels;
[0023] Step 5: When the mth vehicle is moving, the safe speed of the following vehicle is calculated by comprehensively considering the limit of safe speed, vehicle performance, road speed limit and actual operation of the driver;
[0024] Safe speed of the vehicle behind Expressed as
[0025]
[0026] Where w is the maximum deceleration, T r is the driver’s reaction time, T i is the time required for the car to accelerate from rest to maximum speed, v l is the speed of the front car, and S is the distance between the front and rear cars.
[0027] Step 6: Jointly optimize access strategy, bandwidth allocation, and drone trajectory to meet the optimization objective function of maximizing the data transmission success rate;
[0028] The optimization objective function is:
[0029]
[0030]
[0031]
[0032] C3:q u (1) = q init ,
[0033]
[0034]
[0035] ξ is the data transmission success rate, which is calculated by dividing the cumulative number of successfully transmitted data packets by the total number of upload requests; A is the access strategy: B is the bandwidth allocation Q is the trajectory of the drone The horizontal position of the UAV is denoted as q u =(x u ,y u );
[0036] ψ m,k (t)∈{0,1} is a binary variable representing the completion status of data transmission of the mth vehicle. m,k (t)R m,k (t)τ≥D,ψ m,k (t)=1, otherwise ψ m,k (t) = 0; τ is the duration of a time slot; ψ total Indicates the total number of upload requests;
[0037] C1 is the constraint on access policy, C2 is the constraint on bandwidth allocation, and C3 defines the starting position q of the drone. init C4 and C5 respectively ensure that the acceleration and speed of the drone will not exceed the upper limit it can reach; a u and v u They are the acceleration vector and velocity vector of the drone respectively; is the upper limit of the drone’s acceleration; It is the upper speed limit of the drone;
[0038] Step 7: Use the deep reinforcement learning method TD3 to solve the objective function and obtain the data collection with the highest success rate under the coordination of RSU and UAV.
[0039] The TD3 consists of three parts: state space, action space and reward function;
[0040] The state space is:
[0041] s(t)=(q u (t),q m (t),v u (t))
[0042] q u(t) is the position of the UAV changing with time; q m (t) is the vehicle position; v u (t) is the speed of the drone;
[0043] Action Space:
[0044]
[0045] a m,uav (t) is a binary variable representing the mth vehicle unloading data to the receiving end uav; is the flight angle of the drone. (Note: Since the altitude of the drone is fixed, this angle is the angle between the acceleration direction of the drone and the assumed x-axis)
[0046] The reward function is
[0047]
[0048] Is a fixed value reward:
[0049]
[0050] χ c is a constant.
[0051] r vio (t) is the reward corresponding to the penalty for speed and acceleration violations: r vio (t) = ζ ac (t)x ac +ζ v (t)x v
[0052] Among them, when the output action satisfies the acceleration constraint of the drone, ζ ac (t)=1, otherwise ζ ac (t) = 0. ac For fixed rewards;
[0053] When ζ v (t)=0, otherwise ζ v (t) = 1; κ vs =0.05,κ vl =0.95, χ v For fixed rewards.
[0054] The advantages of the present invention are:
[0055] 1. The real-time traffic data collection method for air-land collaboration based on deep reinforcement learning integrates available equipment and collects data by combining static RSU with mobile UAV, thereby improving the success rate of data transmission in the Internet of Vehicles scenario.
[0056] 2. The real-time traffic data collection method for air-land collaboration based on deep reinforcement learning uses real-time traffic scenarios based on real road sections for simulation, which is closer to reality and more practical.
[0057] 3. A real-time traffic data collection method for air-land collaboration based on deep reinforcement learning, which simultaneously optimizes access strategies, bandwidth allocation, and drone trajectories that change over time, effectively improving the success rate of data transmission.
[0058] 4. The air-land collaborative real-time traffic data collection method based on deep reinforcement learning uses TD3 to solve the non-convex optimization problem that is difficult to solve with traditional optimization methods: how to maximize the system's transmission success rate DTSR by optimizing the trajectory of the unmanned aerial vehicle (UAV), the strategy of vehicle access to the UAV and RSU, and the bandwidth allocation of vehicle upload data. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 It is a flow chart of the method for collecting real-time traffic data of air-land collaboration based on deep reinforcement learning of the present invention;
[0060] Figure 2 This is the architecture of air-ground collaborative vehicle networking data collection based on real-time traffic built by the present invention;
[0061] Figure 3 It is a schematic diagram of solving the objective function using the deep reinforcement learning method TD3 adopted in the present invention;
[0062] Figure 4 This is a comparison chart of the cumulative reward value of the present invention compared with the number of training rounds compared with the three existing algorithms;
[0063] Figure 5 It is a comparison chart of the change of DTSR with the number of training rounds compared with the three existing algorithms;
[0064] Figure 6 is a performance comparison chart of the present invention compared with the three existing algorithms when the number of vehicles changes;
[0065] Figure 7 It is a comparison chart of DTSR changes when the data transmitted by the vehicle changes compared with the three existing algorithms. DETAILED DESCRIPTION
[0066] The embodiments of the present invention are described in detail and clearly below with reference to the accompanying drawings.
[0067] The present invention discloses a method for collecting real-time traffic data for air-land collaboration based on deep reinforcement learning. It is a new solution for collecting data jointly by UAV and RSU. It uses drone technology to improve the communication capacity and communication coverage of the existing Internet of Vehicles and improve the throughput of the network. The drone dynamically adjusts its position to better collect the data generated by the moving traffic. The present invention considers a real-time traffic scenario, and the selection of traffic sections is based on real roads. At the same time, it optimizes the access strategy, bandwidth allocation and drone trajectory that change over time to improve DTSR. Finally, TD3 is used to solve the proposed problem. Compared with traditional work, the present invention realizes the optimization of the entire process of air-land collaborative communication completed by UAV and RSU, achieves the purpose of improving the success rate of data collection, has stronger feasibility, and provides a new solution TD3-TAR (joint trajectory design, access decision and resource allocation) in the field of Internet of Vehicles data collection.
[0068] like Figure 1 As shown, the specific steps are as follows:
[0069] Step 1: Build a real-time traffic-based air-ground collaborative vehicle network data collection architecture that includes a roadside unit (RSU), a UAV, and several vehicles;
[0070] like Figure 2 As shown in the figure, the architecture includes a roadside unit RSU and an unmanned aerial vehicle UAV. The RSU and UAV serve as the receiving ends of vehicle data and jointly collect the data generated by the vehicle. The position of the vehicle changes over time, and the vehicle generates a fixed size of data to be uploaded in each time slot.
[0071] Step 2: Each vehicle will generate a data packet of size D to upload in each time slot, and the RSU and UAV jointly collect the data generated by all vehicles;
[0072] Use collection and They represent all vehicles in the scene, the receiver of the data, and all time slots considered, respectively.
[0073] Step 3: For the mth vehicle, the vehicle can only select one receiving end in the same time slot, and set the constraint conditions between the vehicle and the data unloading to the receiving end k;
[0074] The constraints are:
[0075] a m,k (t)∈{0,1}, A binary variable representing the offloading decision. When the mth vehicle needs to offload its data to the receiver k, a m,k (t)=1, otherwise a m,k (t) = 0. Because vehicle m can only select one receiving end in the same time slot, this binary variable must satisfy the condition
[0076] Step 4: When the mth vehicle communicates with the receiving end k, the channel rate R is calculated based on the bandwidth allocation. m,k ;
[0077] The horizontal and vertical positions of the RSU in the three-dimensional scene are represented by q r =(x r ,y r ), h r ; Similarly, the horizontal and vertical positions of the UAV are represented by q u =(x u ,y u ), h u ; The horizontal and vertical positions of the mth vehicle are represented by q m =(x m (t),y m (t)),h m The distance between vehicle m and RSU and UAV is expressed as:
[0078] d m,rsu (t)=||x r -x m (t),y r -y m (t),h r -h m ||
[0079] d m,uav (t)=||x u (t)-x m (t),y u (t)-y m (t),h u -h m ||
[0080] With the support of drone technology, the communication channel is very likely to be a line-of-sight channel (LoS). Depending on different propagation conditions and transmission distances, the probability of the channel being a line-of-sight channel will change. The specific probability formula is as follows:
[0081]
[0082] Where c1 and c2 are constants related to the propagation environment, is the elevation angle, where h = h r-h,h=h u -h.
[0083] m,rsu mm,uav m
[0084] Then, the path loss of the line-of-sight channel and the path loss of the non-line-of-sight channel can be expressed as
[0085]
[0086] in is the inverse of the free space propagation loss, c is the speed of light, and f c is the carrier frequency, β LoS and β NLoS is the attenuation coefficient.
[0087] Furthermore, the average path loss is
[0088]
[0089] The present invention uses orthogonal frequency division multiple access technology, so the interference of adjacent channels can be ignored, and the following channel rate expression is obtained:
[0090]
[0091] where b m,k is the bandwidth allocation factor, satisfying the constraint B k is the available bandwidth; P0, A r , n0 represents the transmission power, antenna gain and noise power density at the receiving end respectively.
[0092] Step 5: When the mth vehicle is moving, the safe speed of the following vehicle is calculated by comprehensively considering the limit of safe speed, vehicle performance, road speed limit and actual operation of the driver;
[0093] In order to make the simulation closer to the real situation, the present invention adopts an improved Krauss following model. If there are only two cars, the safe speed of the following car is expressed as
[0094]
[0095] Among them, v l is the speed of the preceding vehicle, T r is the driver’s reaction time, T i is the time required for the car to accelerate from rest to maximum speed, w is the maximum deceleration, S = x l -x f -l is the distance between the front and rear vehicles, where x l is the position of the front car, x f is the position of the following vehicle and l is the length of the vehicle.
[0096] The speed of the following vehicle is not only limited by the safe speed, but also needs to consider the vehicle performance, road speed limit and the actual operation of the driver. In summary, the speed of the following vehicle can be expressed as:
[0097]
[0098] in is the maximum speed of the car, is the maximum speed limit of the road, v c (t)+a c τ is the maximum speed that the car can reach after a time slot, a c is the acceleration of the car, ∈ v It is a factor used to measure the driver's imperfect operation.
[0099] Step 6: Jointly optimize access strategy, bandwidth allocation, and drone trajectory to meet the optimization objective function of maximizing the data transmission success rate;
[0100] In this scenario, each vehicle generates a data packet of size D in each time slot, and it needs to upload the data packet to the RSU or UAV in the time slot. In order to simplify the optimization problem, the present invention does not consider the time spent on processing data. Specifically: using the binary variable ψ m,k (t)∈{0,1} represents the completion of data transmission of the mth vehicle. If a m,k (t)R m,k (t)τ≥D,ψ m,k (t)=1, otherwise ψ m,k (t) = 0; τ is the duration of a time slot; the total number of upload requests is ψ total =MT; Data transmission success rate (DTSR) ξ can be expressed as the cumulative number of successfully transmitted data packets divided by the total number of upload requests; In order to maximize DTSR, the present invention jointly optimizes the access strategy Bandwidth allocation And drone tracks
[0101] The optimization problem can be summarized as:
[0102]
[0103]
[0104]
[0105] C3:q u (1) = q init ,
[0106]
[0107]
[0108] C1 is the constraint on access policy, C2 is the constraint on bandwidth allocation, and C3 defines the starting position q of the drone. init , C4 and C5 respectively ensure that the acceleration and speed of the drone will not exceed the upper limit it can reach; a u and v u They are the acceleration vector and velocity vector of the drone respectively; is the upper limit of the drone’s acceleration; It is the upper speed limit of the drone;
[0109] Step 7: Use the deep reinforcement learning method TD3 to solve the objective function and obtain the collected data with the highest success rate under the coordination of RSU and UAV.
[0110] Since the above optimization problem is a time-coupled non-convex optimization problem, it is difficult to solve it with traditional methods. Therefore, the present invention uses an efficient deep reinforcement learning method, TD3, to solve it. Compared with traditional methods, TD3 has better perception and generalization capabilities, and it still has better solution capabilities even if the solution space is large.
[0111] TD3 is a method of deep reinforcement learning, which converts the variables to be optimized into the output of the agent, and optimizes the agent to maximize the cumulative reward function directly related to DTSR to solve the original problem. The following specifically introduces how the present invention uses the TD3 algorithm to solve the proposed optimization problem from the key elements of reinforcement learning - state space, action space, reward design and neural network design.
[0112] State Space:
[0113] In order to maximize the DTSRξ mentioned above, the present invention focuses on two aspects of the state. First, DTSR is directly related to the communication rate, which is affected by the distance between the vehicle and the RSU or the UAV. Therefore, the UAV position q that changes with time is u (t) and vehicle position q m (t) is included in the state space. Secondly, the speed of the drone has a great influence on the flight control of the drone, so it is also included in the state space. The final state space is:
[0114] s(t)=(q u (t),q m (t),v u (t))
[0115] q u (t) is the position of the UAV changing with time; q m (t) is the vehicle position; v u (t) is the speed of the drone;
[0116] Action Space:
[0117] The design of rewards can be divided into three parts: first, access strategy a m,k , because of the limitation of C1, we only need to consider a m,rsu and a m,uav One of these two; the second is bandwidth allocation b m,k , which represents the bandwidth ratio provided by RSU or UAV to a vehicle. It should be noted that if a m,k = 0, then b m,k = 0; finally, the UAV speed v related to trajectory design u And the direction of speed Since each time slot τ is very short, the drone approximately moves in a uniformly accelerated straight line within the time slot, and the equation of motion is:
[0118]
[0119] In summary, the action space is:
[0120]
[0121] a m,uav (t) is a binary variable representing the mth vehicle unloading data to the receiving end uav; is the flight angle of the drone. (Note: Since the altitude of the drone is fixed, this angle is the angle between the acceleration direction of the drone and the assumed x-axis)
[0122] Reward function:
[0123] The optimization goal of the present invention is to maximize DTSR. Therefore, when the agent outputs appropriate actions to complete the transmission task well, it is given a fixed value reward.
[0124]
[0125] where χ c is a constant. In addition, in order to punish the speed violation and acceleration violation caused by the action output by the agent, the corresponding reward r is designed vio (t) = ζ ac (t)x ac +ζ v (t)x v , χac is a fixed reward, with a value of -1; when the output action satisfies C4, ζ ac (t)=1, otherwise ζ ac (t)=0. When the fixed value κ vs =0.05,κ vl =0.95, ζ v (t)=0, otherwise ζ v (t) = 1; where χ v It is a fixed reward, and its value is -1.
[0126] In summary, the designed reward function is:
[0127]
[0128] The present invention uses the TD3 algorithm to solve the optimization problem through multiple rounds of iterations to achieve the highest success rate of data collection. Figure 3 As shown:
[0129] First, the network parameters φ, θ1 and θ2 are initialized using the Kaiming method, and the same initial values are assigned to the target network φ * ,θ1 * and θ2 * , and then perform several rounds of iterative updates;
[0130] The specific update method within a round is:
[0131] 1. Initialize the position q of the drone before each round of training m , speed v u And the direction of acceleration
[0132] 2. Get Action
[0133] 3. If the constraint condition C4 can be satisfied, continue to execute. If it is not satisfied, automatically adjust v u (t) and So that C4 can be satisfied;
[0134] 4. Get reward r and next state s′=s(t+1);
[0135] 5. Store (s,a,s′,r) in the experience pool middle;
[0136] 6. Repeat steps 2-5 until the round is over.
[0137] 7. From the experience pool Randomly pick N bJump data.
[0138] 8. Update critic and actor;
[0139] 9. Update the target network.
[0140] The simulation results show:
[0141] The present invention uses OpenStreetMap tool to intercept the road section information of Wudaokou in Beijing as a simulation road section; uses RANDOMTRIPS tool to generate simulated traffic flow, and inputs the simulated road section information and traffic flow information into SUMO simulation software to simulate road vehicles.
[0142] TD3-ARO, TD3-TO and DueDQN-TAR were designed as benchmarks for comparison with the main method TD3-TAR proposed in the present invention. TD3-TAR means using the TD3 method to optimize the UAV trajectory, access strategy and resource allocation (in the present invention, it refers to bandwidth allocation); TD3-ARO means only optimizing the access strategy and bandwidth allocation, and the UAV flight trajectory is fixed; TD3-TO means only optimizing the UAV flight trajectory, and the access strategy and bandwidth allocation are fixed. Among them, the fixed bandwidth allocation is an average allocation, and the fixed flight trajectory is a circle. DueDQN-TAR refers to using the Dueling-DQN method to optimize the UAV trajectory, access strategy and resource allocation (in the present invention, it refers to bandwidth allocation).
[0143] like Figure 4 and Figure 5 As shown, it shows how the cumulative reward value and DTSR change with the number of training rounds. Obviously, the method TD3-TAR proposed in the present invention performs better than the baseline method in both indicators. The reason for the poor performance of TD3-ARO and TD3-TO is that for the proposed problem, the drone trajectory, access strategy and bandwidth allocation are highly coupled, and the lack of any one of them will lead to performance degradation. Although DueDQN-TAR also considers all optimization variables, due to Dueling-DQN is an algorithm that can only output discrete action values, its performance in processing complex problems is weaker than TD3, which can output continuous actions.
[0144] like Figure 6As shown, it shows the performance of different methods when the number of vehicles changes. It can be seen that whether the number of vehicles increases or decreases, the DTSR index of the TD3-TAR algorithm proposed in the present invention is significantly better than other methods, which proves the superiority of TD3-TAR. In addition, when the number of vehicles increases, the DTSR decreases. There are two reasons for this result: first, the communication range of a UAV and an RSU is limited. When the number of vehicles increases, there is a greater probability of exceeding the coverage limit of these two devices. Second, the total bandwidth set in the simulation is fixed. The addition of more vehicles means that the bandwidth resources available to each vehicle will gradually decrease.
[0145] Figure 7 The figure shows the change of DTSR (after convergence) when the data size D generated by the vehicle in each time slot changes. It can be seen that DTSR decreases when D increases. This is because the total bandwidth is limited, and the increase in the amount of data to be transmitted will make the transmission task more difficult to complete. In addition, the DTSR obtained by the TD3-TAR algorithm is much higher than that of other algorithms, which can further illustrate the superiority of the TD3-TAR algorithm.
Claims
1. A real-time traffic data collection method for air-land collaboration based on deep reinforcement learning, characterized in that: The specific steps are as follows: First, a real-time traffic-based air-ground collaborative vehicle network data collection architecture is built, which includes a roadside unit (RSU), a UAV (UAV), and several vehicles. Each vehicle will generate a data packet of size D to upload in each time slot, and the RSU and UAV jointly collect the data generated by all vehicles. Then, for the communication between the mth vehicle and the receiving end k, set the constraint conditions between the vehicle and the data offloading to the receiving end k, and calculate the channel rate R in combination with the bandwidth allocation m,k ; The constraints are: A binary variable representing the offloading decision. When the mth vehicle needs to offload its data to the receiver k, a m,k (t)=1, otherwise a m,k (t) = 0; and They represent all vehicles in the scene, the receiver of the data, and all time slots considered respectively; Channel rate R m,k (t) is expressed as: where b m,k is the bandwidth allocation factor, satisfying the constraint B k is the available bandwidth; P0, A r , n0 represents the transmission power, antenna gain and noise power density at the receiving end respectively; PL m,k is the average path loss of line-of-sight and non-line-of-sight channels; When the mth vehicle is driving, the safe speed of the following vehicle is calculated by comprehensively considering the safety speed limit, vehicle performance, road speed limit and the actual operation of the driver. Expressed as Where w is the maximum deceleration, T r is the driver’s reaction time, T i is the time required for the car to accelerate from rest to maximum speed, v l is the speed of the front car, S is the distance between the front car and the rear car; Finally, the access strategy, bandwidth allocation and UAV trajectory are jointly optimized to build an optimization objective function to maximize the data transmission success rate. The objective function is solved using the deep reinforcement learning method TD3 to obtain the collection data with the highest success rate under the coordination of RSU and UAV. The TD3 consists of three parts: state space, action space and reward function; The state space is: s(t)=(q u (t),q m (t),v u (t)) q u (t) is the position of the UAV changing with time; q m (t) is the vehicle position; v u (t) is the speed of the drone; Action Space: a m,uav (t) is a binary variable representing the mth vehicle unloading data to the receiving end uav; is the flight angle of the drone; The reward function is Is a fixed value reward: χ c is a constant; r vio (t) is the reward corresponding to the penalty for speed and acceleration violations: r vio (t) = ζ ac (t)x ac +ζ v (t)x v Among them, when the output action satisfies the acceleration constraint of the drone, ζ ac (t)=1, otherwise ζ ac (t) = 0; x ac For fixed rewards; When ζ v (t)=0, otherwise ζ v (t) = 1; κ vs and κ vl is a fixed value, χ v For fixed rewards; The optimization objective function is: C3:q u (1)=q init , ξ is the data transmission success rate, which is calculated by dividing the cumulative number of successfully transmitted data packets by the total number of upload requests; A is the access strategy: B is the bandwidth allocation Q is the trajectory of the drone The horizontal position of the UAV is denoted as q u =(x u ,y u ); ψ m,k (t)∈{0,1} is a binary variable representing the completion status of data transmission of the mth vehicle. m,k (t)R m,k (t)τ≥D,ψ m,k (t)=1, otherwise ψ m,k (t) = 0; τ is the duration of a time slot; ψ total Indicates the total number of upload requests; C1 is the constraint on access policy, C2 is the constraint on bandwidth allocation, and C3 defines the starting position q of the drone. init ; C4 and C5 respectively ensure that the acceleration and speed of the drone will not exceed the upper limit it can reach; a u and v u They are the acceleration vector and velocity vector of the drone respectively; is the upper limit of the drone’s acceleration; It is the upper speed limit of the drone.