A method for minimizing age of information for data distribution in internet of vehicles
By establishing an optimization model and training drone trajectory planning in the Internet of Vehicles, the problem of minimizing the information age of drone-assisted vehicle data packet scheduling was solved, achieving more efficient information updates and energy consumption management, and improving the decision-making accuracy and timeliness of the urban traffic management system.
Patent Information
- Application Number
- CN202410943894.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-15
- Publication Date
- 2025-11-07
- Estimated Expiration
- 2044-07-15
AI Technical Summary
In the context of vehicle-to-everything (V2X) communication, existing technologies have failed to effectively address the issues of drone-assisted vehicle data packet scheduling and trajectory planning to minimize information age, thus affecting the timeliness and accuracy of decision-making in urban traffic management systems.
An optimization model is established with the goal of minimizing the average information age of vehicles and the energy consumption of UAVs. The UAVs are trained using Markov processes and near-end policy optimization algorithms to optimize information age updates and trajectory planning, and global decision-making is carried out using base stations.
It reduces the computational and decision-making burden on drones, improves energy efficiency and mission execution, optimizes network resources and energy utilization, reduces information age and energy consumption, and improves system performance.
Smart Images

Figure CN119155805B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of intelligent transportation systems and vehicle networking communication technology, in particular to a vehicle networking data distribution information age minimization method. BACKGROUND
[0002] In the intelligent transportation system, vehicles transmit data through a wireless communication network, and base stations collect and process vehicle information. However, when the base station is located in a relatively remote location in the city, the long-distance transmission between the base station and the vehicle in the city will affect the quality of the data packet, and further affect the decoding and recovery ability of the base station to the data, and ultimately affect the decision of the urban traffic management system.
[0003] The flexibility and reliable line-of-sight characteristics of unmanned aerial vehicles (UAVs) enable them to serve as mobile relay nodes to assist base stations in collecting vehicle information. Unmanned aerial vehicles can dynamically adjust their positions and communication strategies while moving to optimize the efficiency of data transmission. However, in this scenario, the freshness of data packet information (i.e., information age, AoI) is a key indicator, as it directly relates to the timeliness and accuracy of urban traffic decisions.
[0004] So, how should the data packets of walking vehicles be scheduled to ensure the minimum information age when using unmanned aerial vehicles as mobile relay nodes in a vehicle networking environment? When the information age is minimized, how should the trajectory of the unmanned aerial vehicle be planned? There is no direct answer to this problem in the existing technology. SUMMARY
[0005] The technical problem to be solved by the present application is how to achieve the optimal solution of unmanned aerial vehicle-assisted vehicle data packet scheduling and unmanned aerial vehicle trajectory planning in a vehicle networking environment. To overcome the defects of the prior art, the present application provides a vehicle networking data distribution information age minimization method.
[0006] The vehicle networking data distribution information age minimization method provided by the present application comprises the following steps:
[0007] S1: Based on the vehicle networking with unmanned aerial vehicles as mobile relay nodes, an optimization model is established to minimize the average information age of all vehicles in the vehicle networking and the energy consumption of the unmanned aerial vehicle;
[0008] S2: Based on the optimization model, a Markov process is obtained through the interaction of the unmanned aerial vehicle with the constraint environment of the vehicle networking, and the number of cycles of the Markov process is specified;
[0009] S3: Based on the Markov process, the unmanned aerial vehicle is trained through a proximal policy optimization algorithm and reaches the number of cycles to obtain a proximal optimization model;
[0010] S4: controlling the UAV operation in the base station with the proximal optimization model to minimize the average information age of all vehicles and the energy consumption of the UAV.
[0011] The disclosed vehicle networking data distribution information age minimization method can reduce the computational and decision-making burden of the UAV by training a proximal optimization model in advance and using the model to perform information age updating and UAV trajectory selection, so that the UAV can focus mainly on physical movement and communication tasks, thereby improving the energy efficiency and task execution effect of the UAV. In addition, the base station as a central control node can collect and process the state information of the entire network to make a globally optimal decision. Furthermore, by optimizing the information age updating and UAV movement path through the proximal optimization model, network resources and energy can be effectively utilized, unnecessary communication and movement overhead can be reduced, and the energy efficiency of the system can be improved. Not only suitable for vehicle network data distribution in intelligent transportation systems, but also applicable to multiple fields such as UAV swarm cooperation tasks, emergency rescue and disaster management, and has wide application prospect and commercial value.
[0012] In one possible implementation, the step S1 includes the following steps:
[0013] S11: establishing a communication network structure of one base station associated with one UAV, one UAV associated with multiple vehicles, and obtaining the vehicle networking based thereon;
[0014] S12: dividing a specified time into multiple time slots, and obtaining the connection constraints between the UAV and each vehicle in each time slot based on the fact that the change in the UAV position will lead to the change in its coverage range;
[0015] S13: obtaining the channel gain between the UAV and each vehicle in each time slot based on the communication channel characteristics of the vehicle networking;
[0016] S14: obtaining the data transmission rate between the UAV and each vehicle in each time slot based on the channel gain between them obtained in the step S13;
[0017] S15: obtaining the propulsion energy consumption of the UAV in each time slot according to the energy consumption algorithm of the UAV propulsion process, and obtaining the total energy consumption of the UAV according to the fact that the UAV is hovering in a scheduled operation;
[0018] S16: setting information sampling behaviors of all vehicles to follow sampling update in all time slots, and taking the minimum data amount of state update as the minimum data amount required for a successful state update of each vehicle, and then obtaining data amount of state update of each vehicle in each time slot based on the data transmission rate obtained in step S14, and obtaining information age update algorithm of each vehicle according to comparison between data amount of state update of each vehicle in each time slot and the minimum data amount of state update;
[0019] S17: obtaining average energy consumption of the UAV in the specified time based on total energy consumption of the UAV in each time slot obtained in step S15, and obtaining average information age of all vehicles in the specified time based on the information age update algorithm obtained in step S16, and taking the minimum value of the sum of the average energy consumption and the average information age as the target;
[0020] S18: defining a working area of the UAV, and taking the working area and the connection constraint obtained in step S12 as corresponding constraints of the target, and obtaining the optimization model accordingly;
[0021] Compared with the prior art, the above technical solution can provide constraints and targets for optimization of information age and energy consumption reduction based on display constraints of the Internet of Vehicles, thereby providing a guarantee basis for establishment of a Markov process and training of a proximal optimization model accordingly.
[0022] In a possible implementation, the formula of the connection constraint between the UAV and each vehicle obtained in step S12 is:
[0023]
[0024] wherein,
[0025] I is the total number of vehicles;
[0026] The specified time is divided into N time slots, and n=1, 2, …, N represents the number of time slots;
[0027] is a binary variable, represents that the i-th vehicle is dispatched by the UAV, otherwise
[0028] Compared with the prior art, the above technical solution can take into account mobility of the UAV, and provide a constraint for UAV vehicle dispatching according to the fact that change of the UAV position will lead to change of its coverage range.
[0029] In a possible implementation, the formula of the channel gain obtained in step S13 is:
[0030]
[0031] wherein,
[0032] represents the channel gain between the UAV and the vehicle i at time slot n;
[0033] represents the position of the vehicle i at time n;
[0034] represents the two-dimensional coordinate of the UAV;
[0035] q BS = (x BS , y BS ) represents the position of the base station;
[0036] h0represents the channel gain at a reference distance d0= 1 m;
[0037] H represents the flight height of the UAV;
[0038] d i,u (n) represents the distance between the vehicle i and the UAV at time slot n;
[0039] This scheme provides a fast and accurate way for calculating the channel gain between the UAV and the vehicle i at each time slot, thereby achieving accurate acquisition of the information gain.
[0040] In a possible implementation, the formula for the data transmission rate obtained in the step S14 is:
[0041]
[0042] wherein,
[0043] represents the data transmission rate between the UAV and the vehicle i at time slot n;
[0044] B represents the communication bandwidth;
[0045] σ 2 represents the additive white Gaussian noise;
[0046] P represents the transmission power of the vehicle;
[0047] This scheme can achieve accurate acquisition of the data transmission rate between the UAV and the vehicle i at each time slot, thereby providing a guarantee basis for calculation of the data amount for state update of each vehicle at each time slot.
[0048] In a possible implementation, in the step S15, the formula for the energy consumption algorithm of the UAV propulsion process is:
[0049]
[0050] wherein,
[0051] P(v n ) represents the propulsion energy consumption of the UAV in time slot n;
[0052] P0 represents the profile power of the UAV in hover;
[0053] P1 represents the induced power of the UAV in hover;
[0054] v n represents the speed of the UAV in time slot n;
[0055] U tip represents the tip speed of the rotor blade of the UAV;
[0056] v0 represents the average rotor induced speed of the UAV in hover;
[0057] d represents the fuselage drag ratio of the UAV;
[0058] p represents the air density;
[0059] s represents the rotor solidity of the UAV;
[0060] A represents the rotor disc area of the UAV;
[0061] The total energy consumption of the UAV is calculated as:
[0062]
[0063] P h = P0+ P1,
[0064] wherein,
[0065] T s represents the flight time of the UAV;
[0066] T c represents the hover time of the UAV;
[0067] P h represents the power consumption of the UAV in hover;
[0068] This scheme is based on the fact that the main energy consumption of the UAV is composed of propulsion, hover and communication, and the energy consumption of communication is ignored, which reduces the complexity of energy consumption calculation while ensuring accuracy.
[0069] In a possible implementation, in the step S16, the formula of the data amount of the state update of each vehicle in each time slot is:
[0070]
[0071] wherein,
[0072] the data amount of the state update of vehicle i in time slot n;
[0073] delta n representing the length of time slot n;
[0074] The formula of the information age update algorithm of each vehicle is:
[0075]
[0076] wherein,
[0077] representing the information age of vehicle i in time slot n;
[0078] representing the minimum data amount of the state update of vehicle i;
[0079] representing the set of vehicles;
[0080] This scheme introduces the minimum data amount of the state update of vehicle i, and provides an update strategy based on the minimum data amount, thereby ensuring the ordered update of the information age.
[0081] In a possible implementation, the obtained Markov process in the step S2 is expressed as follows:
[0082] The UAV is taken as an agent, the agent cyclically interacts with a constraint environment in which the Internet of Vehicles is located, the number of cyclic interactions is the number of cycles of the Markov process specified in the step S2, and the number of steps performed in each cycle is the same;
[0083] For each step, the agent observes the constraint environment in the current step to obtain an observation state of the current step, and based on the observation state and a strategy of the agent itself, the agent selects a hybrid action from a hybrid action space in the current step;
[0084] The hybrid action space is composed of a discrete action space and a continuous action space;
[0085] The agent selects a hybrid action from the hybrid action space in this step, which means that the agent selects a discrete action from the discrete action space for the vehicle whose information age needs to be updated at present, and selects two continuous actions in the continuous action space for the UAV moving speed and direction;
[0086] After the agent performs the hybrid action in this step, the state of the constrained environment changes, and the agent obtains a corresponding reward value, after which this step ends and the next step begins;
[0087] This scheme realizes the conversion of the optimization model into a Markov process, and provides a model prototype for training the UAV, which can train the agent by using the proximal optimization algorithm, thereby providing a prerequisite guarantee for obtaining the proximal optimization model.
[0088] In a possible implementation, the step S3 includes the following steps:
[0089] S31: According to the requirements of the proximal policy optimization algorithm, a neural network containing a performer network and a critic network is used to initialize all parameters of the neural network, and an experience replay pool of the agent is initialized, and the number of training stages is set to be the same as the number of cycles of the Markov process, and the number of steps performed in each cycle of the training stage is set to be the same as the number of steps set in the Markov process;
[0090] The performer network is composed of a discrete action network module for outputting discrete actions and a continuous action network module for outputting continuous actions;
[0091] S32: In the current step, the agent observes the constrained environment at the beginning of the current step to obtain the current observation state of each time slot, which is expressed as:
[0092]
[0093] In the formula,
[0094] State represents the current observation state;
[0095] represents a set of time slots;
[0096] S33: The current observation state is input into the performer network and the critic network at the same time;
[0097] The intelligent agent scheduling vehicle index of each time slot is output by the discrete action network module of the performer network, and the vehicle represented by the index sends a state update message according to the scheduling vehicle index of each time slot, the state update message contains the information age of the vehicle involved, and the information age of the vehicle involved at the intelligent agent is reset to δ n The information age of the vehicle not sending the state update message is increased by δ n The discrete action of the present step is updated
[0098] The current output value of the continuous action of each time slot is calculated by the continuous action network module of the performer network, which contains the moving direction of the intelligent agent and the moving speed of the intelligent agent, and the intelligent agent executes according to the current output value, and the continuous action of the present step is updated
[0099] After the discrete action and continuous action of the present step are updated, the intelligent agent obtains a reward value based on the constraint environment update of the present step, and obtains the next step observation state of each time slot based on the reward and the output of the critic network
[0100] S34: store the state, action, reward, and next step observation state of each time slot of the present step in the experience replay pool
[0101] S35: determine whether the present step is executed in a cycle, if not, take the next step observation state of each time slot as the current observation state of each time slot and return to execute the step S32
[0102] If yes, execute the next step
[0103] S36: determine whether the number of cycles of the present step reaches the number of cycles or reaches the storage limit of the experience replay pool, if at least one of them is met, obtain the proximal optimization model by performing several times of time iteration based on all data collected by the experience replay pool, and embed the proximal optimization model in the base station, otherwise, take the next step observation state of each time slot as the current observation state of each time slot and return to execute the step S32
[0104] This scheme can optimize in discrete action space and continuous action space at the same time, so that information updating and unmanned aerial vehicle moving path planning are more efficient and accurate. This dual-space optimization method can significantly reduce information age and energy consumption, and improve the overall system performance. BRIEF DESCRIPTION OF DRAWINGS
[0105] Figure 1 A flow chart of a vehicle networking data allocation information age minimization method disclosed in an embodiment of the present application
[0106] Figure 2The step S1 flowchart disclosed in the embodiments of the present application;
[0107] Figure 3 The step S3 flowchart disclosed in the embodiments of the present application;
[0108] Figure 4 The neural network structure involved in the embodiments of the present application. DETAILED DESCRIPTION
[0109] First of all, those skilled in the art should understand that these embodiments are only used to explain the technical principles of the embodiments of the present application, and are not intended to limit the protection scope of the embodiments of the present application. Those skilled in the art can adjust them as needed to adapt to specific application occasions.
[0110] In the embodiments of the present application, unless otherwise explicitly specified and limited, the first feature is "on" or "under" the second feature, which can be that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature can be above or obliquely above the second feature, or only means that the horizontal height of the first feature is higher than that of the second feature. The first feature can be below or obliquely below the second feature, or only means that the horizontal height of the first feature is less than that of the second feature.
[0111] The present application will be further described in detail below in combination with the drawings and specific embodiments.
[0112] Referring to Figures 1-3 The embodiments of the present application disclose a method for minimizing data distribution information age of Internet of Vehicles, referring to Figure 1 The method comprises the following steps:
[0113] S1: An optimization model is established based on the Internet of Vehicles taking a drone as a mobile relay node, aiming to minimize the average information age of all vehicles in the Internet of Vehicles and the energy consumption of the drone.
[0114] Referring to Figure 2 In the present embodiment, step S1 comprises the following steps:
[0115] S11: A communication network structure is established in which a base station is associated with a drone, and the drone is associated with multiple vehicles, and accordingly the Internet of Vehicles is obtained;
[0116] Of course, the specific Internet of Vehicles can be formed by parallel or side-by-side arrangement of the communication network structure. As a specific example, it is assumed that the Internet of Vehicles has one master base station, I vehicles, and one drone. The communication between the drone and the base station or mobile device is realized through a dedicated frequency band to ensure the reliability and stability of the communication.
[0117] S12: dividing the specified time into multiple time slots, and obtaining connection constraints between the UAV and each vehicle in each time slot based on the fact that the UAV position change will cause its coverage change;
[0118] In this embodiment, the formula of the connection constraints between the UAV and each vehicle obtained in step S12 is:
[0119]
[0120] In the formula,
[0121] I is the total number of vehicles;
[0122] The specified time is divided into N time slots, and n = 1, 2, …, N represents the number of time slots;
[0123] is a binary variable, represents that the ith vehicle is dispatched by the UAV, otherwise
[0124] S13: obtaining channel gain between the UAV and each vehicle in each time slot based on the communication channel characteristics of the Internet of Vehicles;
[0125] In this embodiment, the formula of the channel gain obtained in step S13 is:
[0126]
[0127] In the formula,
[0128] represents the channel gain between the UAV and the vehicle i in the time slot n;
[0129] represents the position of the vehicle i at time n;
[0130] represents the two-dimensional coordinates of the UAV;
[0131] q BS = (x BS , y BS ) represents the position of the base station;
[0132] h0 represents the channel gain when the reference distance d0 = 1 m;
[0133] H represents the flight height of the UAV;
[0134] d i,u (n) represents the distance between the vehicle i and the UAV in the time slot n.
[0135] S14: obtaining the data transmission rate between each time slot UAV and each vehicle based on the channel gain between each time slot UAV and each vehicle obtained in step S13;
[0136] In the embodiment, the formula of the data transmission rate obtained in step S14 is:
[0137]
[0138] wherein,
[0139] represents the data transmission rate between the UAV in time slot n and vehicle i;
[0140] B represents the communication bandwidth;
[0141] σ 2 represents the additive white Gaussian noise;
[0142] P represents the transmission power of the vehicle.
[0143] S15: obtaining the propulsion energy consumption of each time slot UAV according to the energy consumption algorithm of the UAV propulsion process, and obtaining the total energy consumption of the UAV according to the fact that the UAV hovers in the scheduling operation;
[0144] In step S15 of the embodiment, the formula of the energy consumption algorithm of the UAV propulsion process is:
[0145]
[0146] wherein,
[0147] P(v n ) represents the propulsion energy consumption of the UAV in time slot n;
[0148] P0 represents the profile power of the UAV when hovering;
[0149] P1 represents the induced power of the UAV when hovering;
[0150] v n represents the speed of the UAV in time slot n;
[0151] U tip represents the tip speed of the rotor blade of the UAV;
[0152] v0 represents the average rotor induced speed of the UAV when hovering;
[0153] d represents the fuselage drag ratio of the UAV;
[0154] ρ represents the air density;
[0155] s represents the rotor solidity of the UAV;
[0156] A represents the rotor disc area of the UAV.
[0157] At the same time, in step S15, the formula of the total energy consumption of the UAV is:
[0158]
[0159] P h = P0+ P1,
[0160] In the formula,
[0161] T s represents the flight time of the UAV;
[0162] T c represents the hovering time of the UAV;
[0163] P h represents the power consumption of the UAV in the hovering state.
[0164] S16: setting the information sampling behavior of all vehicles to follow the sampling update in all time slots, and taking the minimum data amount of state update as the minimum data amount required for a successful state update of each vehicle once, then obtaining the data amount of state update of each vehicle in each time slot based on the data transmission rate obtained in step S14, and obtaining the information age update algorithm of each vehicle according to the comparison between the data amount of state update of each vehicle in each time slot and the minimum data amount of state update;
[0165] Specifically, in step S16 of the embodiment, the formula of the data amount of state update of each vehicle in each time slot is:
[0166]
[0167] In the formula,
[0168] the data amount of state update of vehicle i in time slot n;
[0169] δ n represents the time length of time slot n.
[0170] At the same time, in step S16, the formula of the information age update algorithm of each vehicle is:
[0171]
[0172] In the formula,
[0173] represents the information age of vehicle i in time slot n;
[0174] represents the minimum data amount of state update of vehicle i;
[0175] represent a set of vehicles.
[0176] S17: obtaining the average energy consumption of the UAV at the specified time based on the total energy consumption of the UAV in each time slot obtained in step S15, obtaining the average information age of all vehicles at the specified time based on the information age updating algorithm obtained in step S16, and taking the minimum value of the sum of the average energy consumption and the average information age as the target;
[0177] Specifically, in step S17 of the present embodiment, the average information age of all vehicles in the Internet of Vehicles at the specified time and the average energy consumption of the UAV can be represented as:
[0178]
[0179] wherein,
[0180] ξ and η are positive values representing the weight of the corresponding item in the optimization target.
[0181] S18: defining the working area of the UAV, taking the working area and the connection constraint obtained in step S12 as the corresponding constraint, and obtaining the optimal model accordingly;
[0182] Accordingly, the specific expression of the obtained optimal model is as follows:
[0183]
[0184] wherein,
[0185] X max and Y max represent the maximum planar area range of the UAV activity.
[0186] S2: obtaining the Markov process of the interaction between the UAV and the constraint environment in which the Internet of Vehicles is located based on the optimal model, and defining the number of cycles of the Markov process.
[0187] In the present embodiment, step S2 actually converts the solving process of the above-mentioned optimal model into a Markov process, and the obtained Markov process is expressed as follows:
[0188] The UAV is taken as an intelligent agent, the intelligent agent is provided with a corresponding experience replay pool, the intelligent agent and the constraint environment in which the Internet of Vehicles is located are cyclically interacted, the number of cycles of the cyclic interaction is the number of cycles of the Markov process defined in the above-mentioned step S2, and the number of steps performed in each cycle is the same;
[0189] For each step, the intelligent agent observes the constraint environment at the beginning of the step, obtains the observation state of the current step, and based on the observation state and the policy of the intelligent agent itself, the intelligent agent selects a mixed action from the mixed action space in the step;
[0190] The hybrid action space is composed of a discrete action space and a continuous action space;
[0191] The agent selects a hybrid action from the hybrid action space in this step, which means that the agent selects a discrete action from the discrete action space for the vehicle whose information age needs to be updated, and selects two continuous actions for the speed and direction of the UAV in the continuous action space;
[0192] After the agent performs the hybrid action in this step, the state of the constraint environment changes, and the agent obtains the corresponding reward value, thus this step ends and enters the next step.
[0193] S3: Based on the Markov process, the UAV is trained by the proximal policy optimization algorithm to reach the cycle number, and a proximal optimization model is obtained;
[0194] The proximal policy optimization (PPO) algorithm is a policy-based method that uses two deep networks, and the corresponding network structure diagram is as shown in Figure 4 The first network is an actor network, which is the object of training, and it maintains the current policy π θ (a|s). The actor network takes the state as input and outputs a list of probabilities corresponding to each action in the action space, and these probabilities form a distribution from which actions can be sampled. The second network is a critic network, which is used to “criticize” the current policy of the actor network according to its estimated state value V θ (a|s). After performing an action, the critic network will evaluate whether the action taken has increased or decreased the expected feedback based on the current state.
[0195] Referring to Figure 3 In this embodiment, step S3 includes the following steps:
[0196] S31: According to the requirements of the proximal policy optimization algorithm, a neural network containing an actor network and a critic network is used to initialize all parameters (i.e., network weights) of the neural network, and the experience replay pool of the agent is initialized, and the number of cycles in the training phase is set to the number of cycles of the Markov process, and the number of steps performed in each cycle of the training phase is specified to be the same as the number of steps set by the Markov process;
[0197] The actor network is composed of a discrete action network module for outputting discrete actions and a continuous action network module for outputting continuous actions;
[0198] As a specific example, the number of cycles is set to MaxIter, each cycle has N steps, and each step is an Episode.
[0199] S32: In the current step (Episode), the agent observes the constraint environment at the beginning of the current step, and obtains the current observation state of each time slot, which is expressed as:
[0200]
[0201] In the formula,
[0202] State represents the current observation state;
[0203] represents a set of time slots.
[0204] S33: The current observation state is input into the actor network and the critic network at the same time;
[0205] The discrete action network module of the actor network outputs the agent scheduling vehicle index of each time slot, and according to the scheduling vehicle index of each time slot, the index representing the vehicle sends a state update message, which contains the information age of the vehicle involved, and resets the information age of the vehicle involved at the agent to δ n , and the information age of the vehicle that does not send the state update message increases by δ n , and the discrete action of this step is updated;
[0206] The continuous action network module of the actor network calculates the current output value of the continuous action of each time slot, which contains the moving direction of the agent and the moving speed of the agent, and the agent executes according to the current output value. The continuous action of this step is updated;
[0207] After the discrete action and the continuous action of this step are updated, the agent obtains the reward value based on the updated constraint environment of this step, and obtains the next observation state of each time slot based on the reward and the output of the critic network;
[0208] Referring to Figure 4 , based on the proximal policy optimization algorithm, the mixed action a(n) selected by the agent in the current step is obtained by observing the probability π θ (a(n0)|s(n0)) of the current policy. Then the agent executes the mixed action a(n) in the current step. After the completion of the mixed action, that is, the UAV moves to a new position, the information transmission of vehicle i is completed, the constraint environment is updated, and the agent obtains the reward value r(n) in the current step. Then the agent observes the environment at the end of the current step (i.e., before the next Episode starts) and obtains a new observation state s(n+1).
[0209] Actor-Critic architecture includes a Critic Network and an Actor Network. See Figure 4 As shown, the hidden layers of the Actor Network and the Critic Network are composed of two fully connected layers, and the input layer and the two fully connected layers have 256, 128, 64 neurons, respectively. In this embodiment, the policy and the value function are represented by deep neural networks, and thus, the discrete action policy, the continuous action policy, and the action value function are represented as and V θ (s(n)).
[0210] S34: Store the state, action, reward, and next observation state of each time slot of the current step in the experience replay pool.
[0211] S35: Determine whether a cycle is completed, if not, use the next observation state of each time slot as the current observation state of each time slot and return to execute the step S32.
[0212] If yes, execute the next step.
[0213] S36: Determine whether the number of cycles has reached the number of cycles or the storage limit of the experience replay pool, if at least one of the two conditions is met, obtain the proximal optimization model by performing several time iterations on all the data collected based on the experience replay pool and embed the proximal optimization model in the base station, otherwise, use the next observation state of each time slot as the current observation state of each time slot and return to execute the step S32.
[0214] As a specific example, the storage limit of the experience replay pool can be set to the condition that the number of steps n at this time is divisible by the storage amount L of the experience replay pool, that is, the step S33 is performed once every L time steps. Then, several time iterations are performed on the collected data. In each iteration, the algorithm randomly collects small batches of samples from the experience replay pool, and each batch has B experiences, which are generated by the latest version of the policy and the current state value The advantage estimation value For each batch, the collected state is transmitted into the actor and critic networks to establish a new policy and the corresponding critic target value, respectively.
[0215] In the environment interaction phase of the step S33, the state obtained by the agent is sent to the actor network and the critic network. Then, the discrete action network determines the discrete decision of scheduling the vehicle, and the continuous action network determines the direction and speed of the flight of the unmanned aerial vehicle, and the two form a hybrid action and feedback to the environment.
[0216] An experience store (denoted by D) is used, with S, A and R representing the states, actions and rewards collected in the experience store D. After each reward feedback, D = (s(n), a(n), r(n), s(n+1)) is stored in the experience store D. Once enough samples are collected, the algorithm enters the learning phase.
[0217] First, the advantage estimate is calculated, which is given by:
[0218]
[0219] where,
[0220] N' represents the stopping time of the current learning iteration;
[0221] γ∈[0, 1] represents the discount factor, and in the present example γ = 0.99;
[0222] λ p represents the variance in the model.
[0223] A small batch of experiences is then drawn from the experience store memory. For each batch of experiences, the states in the small batch of experiences are passed to the actor network and critic network to generate a new policy θ and target advantage value, from which the loss function is derived. For discrete actions, the loss function is:
[0224]
[0225] where,
[0226] represents the ratio of the discrete action policy to the old discrete action policy, and
[0227]
[0228] represents the advantage estimate of the corresponding experience sample in the i-th batch;
[0229] The clip function is defined as:
[0230]
[0231] The clip function is added to prevent the update from being too large each time importance sampling is performed. Similarly, the loss function for continuous action policies is:
[0232]
[0233] where,
[0234] a ratio of a representative old continuous action policy to a representative new continuous action policy.
[0235] The critic network's loss function is:
[0236]
[0237] This example uses the Adam optimizer. In addition, the discount factor γ = 0.99, the learning rate of the actor network α = 0.01, the critic network gets a learning rate α = 0.01, and ∈ in the Clip function clip = 0.1.
[0238] S4: In the base station, the unmanned aerial vehicle operation is controlled with a proximal optimization model to minimize the average information age of all vehicles and the energy consumption of the unmanned aerial vehicle.
[0239] Compared with the prior art, the vehicle networking data distribution information age minimization method disclosed in the embodiments of the application can reduce the computational and decision-making burden of the unmanned aerial vehicle by updating the information age at the base station instead of at the unmanned aerial vehicle, so that the unmanned aerial vehicle can mainly focus on physical movement and communication tasks, thereby improving the energy efficiency and task execution effect of the unmanned aerial vehicle. As a central control node, the base station can collect and process the state information of the entire network, thereby making a globally optimal decision. By optimizing the information age update and the unmanned aerial vehicle movement path, network resources and energy can be effectively utilized, unnecessary communication and movement overhead can be reduced, and the energy efficiency of the system can be improved. In addition, optimization is simultaneously performed in the discrete action space and the continuous action space, so that the information update and the unmanned aerial vehicle movement path planning are more efficient and accurate. This dual-space optimization method can significantly reduce the information age (AoI) and energy consumption, and improve the overall system performance.
[0240] The technical solutions of the embodiments of the application are not only applicable to vehicle network data distribution in intelligent transportation systems, but also can be applied to unmanned aerial vehicle swarm cooperation tasks, emergency rescue and disaster management, and many other fields, and have wide application prospects and commercial value.
[0241] In the description of the embodiments of the application, it should be noted that in the description of the application, the terms indicating the direction or position relationship such as "inner" and "outer" are based on the direction or position relationship shown in the drawings, which is only for the convenience of description, and does not indicate or imply that the device or member must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the application.
[0242] In the description of the application, the description of the terms "one embodiment", "some embodiments", "in this embodiment", "specific example", or "some examples" and the like means that the specific features, mechanisms, materials or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the application. In the description, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described can be combined in any suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in the specification and the features of the different embodiments or examples without contradiction.
[0243] The above description is merely a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed in the present application, which should be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for minimizing the age of data allocation information in a vehicle-to-everything (V2X) network, characterized in that, The method comprises the following steps: S1: establishing an optimization model based on a vehicle-to-everything network with a drone as a mobile relay node, the optimization model aiming to minimize the average information age of all vehicles in the vehicle-to-everything network and the energy consumption of the drone; S2: obtaining a Markov process of the drone interacting with a constraint environment where the vehicle-to-everything network is located based on the optimization model, and defining the number of cycles of the Markov process; S3: training the drone by a proximal policy optimization algorithm based on the Markov process and reaching the number of cycles to obtain a proximal optimization model; S4: driving the drone to operate in the base station with the proximal optimization model to minimize the average information age of all vehicles and the energy consumption of the drone; In the step S3, a neural network containing a performer network and a critic network is used to obtain the proximal optimization model by parameter optimization of reinforcement learning. 2.The method of minimizing age of information for data distribution in vehicular networks according to claim 1, wherein, The step S1 comprises the following steps: S11: establishing a communication network structure of a base station associated with a drone, the drone associated with multiple vehicles, and obtaining the vehicle-to-everything network accordingly; S12: dividing a specified time into multiple time slots, and obtaining the connection constraints between the drone and each vehicle in each time slot based on the fact that the coverage of the drone changes due to the change of the position of the drone; S13: obtaining the channel gain between the drone and each vehicle in each time slot based on the communication channel characteristics of the vehicle-to-everything network; S14: obtaining the data transmission rate between the drone and each vehicle in each time slot based on the channel gain obtained in the step S13; S15: obtaining the propulsion energy consumption of the drone in each time slot according to the energy consumption algorithm of the propulsion process of the drone, and obtaining the total energy consumption of the drone according to the fact that the drone hovers in a scheduling operation; S16: setting the information sampling behavior of all vehicles to follow the sampling update in all time slots, taking the minimum data amount of state update as the minimum data amount required for each vehicle to successfully update the state once, and then obtaining the data amount of state update of each vehicle in each time slot based on the data transmission rate obtained in the step S14, and obtaining the information age update algorithm of each vehicle according to the comparison between the data amount of state update of each vehicle in each time slot and the minimum data amount of state update; S17: obtaining the average energy consumption of the drone in the specified time based on the total energy consumption of the drone in each time slot obtained in the step S15, obtaining the average information age of all vehicles in the specified time based on the information age update algorithm obtained in the step S16, and taking the minimum value of the sum of the average energy consumption and the average information age as the target; S18: defining the working area of the drone, taking the working area and the connection constraints obtained in the step S12 as the corresponding constraints of the target, and obtaining the optimization model accordingly. 3.The method of minimizing age of information for data distribution in vehicular networks according to claim 2, wherein, The formula of the connection constraints between the drone and each vehicle obtained in the step S12 is: , wherein, V is the total number of vehicles; The specified time is divided into N time slots, representing the number of time slots; is a binary variable, represents the vehicle is dispatched by the drone, otherwise .
4. The method of minimizing age of information for data distribution in vehicular networks of claim 3, wherein, The formula of the channel gain obtained in the step S13 is: , wherein, representative time slot channel gain between the drone and the vehicle channel gain between the drone and the vehicle representative vehicle at time of day a two-dimensional coordinate representative of the drone; This represents the location of the base station; representative reference distance channel gain at the time This represents the flight altitude of the drone. representative vehicle in a time slot distance from the drone.
5. The method of minimizing age of information for data distribution in vehicular networks of claim 4, wherein, The formula of the data transmission rate obtained in the step S14 is: , wherein, representative time slot the drone and the vehicle data transfer rate Represents communication bandwidth; denotes an additive white Gaussian noise; This represents the transmission power of the vehicle.
6. The method of minimizing age of information for data distribution in vehicular networks of claim 5, wherein, In the step S15, the energy consumption algorithm of the UAV propulsion process is as follows: , In the formula, representative time slot propulsion energy consumption of the drone representing a profile power when the drone is hovering; This represents the induced power when the drone is hovering; representing a speed of the drone at a time slot of the time slot; This represents the tip velocity of the rotor blades of the aforementioned UAV; representing an average rotor induced velocity when the drone is hovering; This represents the airframe drag ratio of the aforementioned drone; Represents air density; This represents the rotor solidity of the aforementioned drone; This represents the area of the rotor disk of the drone. The total energy consumption of the UAV is as follows: , , In the formula, This represents the flight time of the aforementioned drone; This represents the hovering time of the aforementioned drone; This represents the power consumption of the drone while it is hovering.
7. The method of minimizing age of information for data distribution in vehicular networks of claim 6, wherein, In the step S16, the data amount of the state update of each vehicle in each time slot is as follows: , In the formula, representative vehicle in a time slot amount of data for state updates; representative time slot of the time duration; The information age update algorithm of each vehicle is as follows: , In the formula, representative vehicle in a time slot age of information; representative vehicle the state update minimum data amount; A collection representing vehicles.
8. The method of minimizing age of information for data distribution in vehicular networks of claim 7, wherein, The obtained Markov process in the step S2 is as follows: The UAV is taken as an agent, the agent cyclically interacts with the constraint environment in which the Internet of Vehicles is located, the cyclic interaction times are the cycle times of the Markov process defined in the step S2, the step number of each cycle is the same; For each step, the agent observes the constraint environment at the beginning of the step, obtains the observation state of the current step, and selects a hybrid action from the hybrid action space based on the observation state and the strategy of the agent itself; The hybrid action space is composed of a discrete action space and a continuous action space; The selection of the hybrid action by the agent in the step is that the agent selects a discrete action from the discrete action space for the vehicle whose information age needs to be updated, and selects two continuous actions from the continuous action space for the moving speed and direction of the UAV; After the agent executes the hybrid action in the step, the state of the constraint environment changes, and the agent obtains a corresponding reward value, and the step ends and enters the next step.
9. The method of minimizing age of information for data distribution in vehicular networks of claim 8, wherein, The step S3 includes the following steps: S31: according to the requirements of the proximal policy optimization algorithm, a neural network containing a performer network and a critic network is adopted, all parameters of the neural network are initialized, an experience replay pool corresponding to the agent is initialized, the cycle number of the training stage is set as the cycle number of the Markov process, and it is stipulated that the step number of each cycle of the training stage is the same as the step number defined in the Markov process; The performer network is composed of a discrete action network module for outputting a discrete action and a continuous action network module for outputting a continuous action; S32: in the current step, the agent observes the constraint environment at the beginning of the current step, and obtains the current observation state of each time slot; S33: the current observation state is input into the performer network and the critic network at the same time; The intelligent agent dispatches vehicle indexes of each time slot through the discrete action network module of the performer network, and sends state update messages to the vehicles represented by the indexes in turn according to the dispatch vehicle indexes of each time slot, wherein the state update messages contain information ages of the vehicles involved, and the information ages of the vehicles involved at the intelligent agent are reset to while the information ages of the vehicles that do not send state update messages are increased , and the discrete action of this step is updated accordingly; The current output value of the continuous action of each time slot is calculated through the continuous action network module of the performer network, the current output value contains the moving direction of the agent and the moving speed of the agent, and the agent executes in turn according to the current output value, and the continuous action of the step is updated; After the discrete action and the continuous action of the step are updated, the agent obtains a reward value based on the constraint environment update of the step, and obtains the next step observation state of each time slot based on the reward and the output of the critic network; S34: the state, action, reward of the current step, and the next step observation state of each time slot are stored in the experience replay pool. S35: judging whether a current cycle is completed, if not, taking a next observation state of each time slot as a current observation state of each time slot and performing the step S32 again; if yes, performing a next step; S36: judging whether a current cycle number reaches the cycle number or reaches a storage limit of the experience replay pool, if at least one of them is satisfied, obtaining the proximal optimization model by means of several times of duration iteration based on all data collected by the experience replay pool and embedding the proximal optimization model into the base station, otherwise, taking a next observation state of each time slot as a current observation state of each time slot and performing the step S32 again.
Citation Information
Patent Citations
Unmanned aerial vehicle data collection method based on minimized information age
CN110543185A
Unmanned aerial vehicle track adaptive optimization method based on information age
CN115696211A