Collaborative Path Planning and Scheduling Method Based on Blockchain in the Cognitive Vehicular Internet of Things Scenario
The blockchain-enabled path planning and scheduling method addresses the inefficiencies in CIoVs by promoting secure, efficient, and coordinated vehicle collaboration, optimizing traffic and computational load across mixed driving scenarios.
Patent Information
- Application Number
- CN202211569303.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-08
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-12-08
AI Technical Summary
In the cognitive vehicle network scenario, vehicle path planning and scheduling methods cannot effectively achieve global optimization, and the existing technology has failed to effectively solve the problems of computing delay and load balancing, especially in the hybrid driving scenario, the service requirements differences of different types of vehicles have not been fully considered.
The blockchain-based collaborative path planning and scheduling method is adopted, and the vehicles are mapped into virtual nodes in the blockchain network, and the mobile edge computing nodes and roadside units are used for environmental perception and decision-making consensus. Combined with Q-learning reinforcement learning and distributed multi-agent algorithm, collaborative path planning and scheduling between vehicles is realized.
It realizes global optimal path planning and scheduling between vehicles in hybrid driving scenarios, reduces calculation delay and traffic congestion, meets the service needs of different types of vehicles, and reduces calculation complexity.
Smart Images

Figure CN116030623B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Internet of Vehicles, and particularly to a collaborative path planning and scheduling method based on blockchain in a cognitive Internet of Vehicles scenario. Background Art
[0002] Cognitive Internet of Vehicles (CIoVs) introduces a cognitive engine to perceive the traffic environment and network status, and assist in the path planning and scheduling of vehicles. In addition, the large-scale deployment of various advanced in-vehicle sensors and the emergence of connected autonomous driving have brought a large number of computing tasks for decision-making and automation applications. Intelligent vehicles need to efficiently process complex computing tasks to improve traffic efficiency, such as path planning, trajectory tracking, collaborative positioning, environment recognition, etc. At the same time, the path planning and scheduling of vehicles consume a large amount of computing resources. However, vehicles with limited storage and computing resources cannot efficiently process the exploding computing tasks from various intelligent applications. Vehicles need to offload a large number of computing tasks to Mobile Edge Computing Nodes (MECNs) for processing to reduce computing latency.
[0003] In the CIoVs scenario, the intelligent path planning and scheduling of vehicles are crucial for reducing the travel time of vehicles in Intelligent Transportation Systems (ITS) and the task processing latency. Traffic efficiency and computing task processing latency are closely related to traffic conditions and the load of MECNs. The unbalanced distribution of vehicles and computing loads usually leads to road congestion and MECN overload. In addition, due to the selfish behavior of individuals, independent decision-making methods may gather too many vehicles on non-congested sections, causing secondary congestion, thus leading to the opposite result. Collaborative decision-making among vehicles provides a solution for obtaining a globally optimal path planning and scheduling strategy. However, collaboration among vehicles requires vehicles to share a large amount of information. Due to security and privacy concerns, vehicles are reluctant to share personal information, which hinders collaborative decision-making among vehicles and cannot obtain a globally optimal path planning and scheduling strategy in the CIoVs scenario.
[0004] Blockchain has characteristics such as decentralization, immutability, auditability, and anonymity, providing a secure and privacy-protected solution for collaborative optimization among vehicles. In addition, blockchain technology can ensure that the information on the chain is not tampered with, thus providing trustworthy information for vehicle collaboration. Utilizing blockchain technology for secure and trustworthy information sharing among vehicles can promote intelligent collaborative decision-making among intelligent transportation participants. However, considering cost and efficiency issues, blockchain with high computational and communication overhead is difficult to directly apply to CIoVs. Moreover, in the actual traffic environment, connected automated vehicles (CAVs) and connected ordinary vehicles (COVs) with different communication, computational, and control capabilities have different service requirements for computing and transportation services that need to be met.
[0005] In the prior art, for example, 1) the dynamic path planning method for intelligent transportation systems based on data prediction and load balancing (Dynamic Path Planning Algorithms with Load Balancing Based on Data Prediction for Smart Transportation Systems) studied the load balancing problem in intelligent transportation systems, which can dynamically adapt to the traffic environment and avoid urban traffic congestion. It established a prediction model based on historical traffic data and current traffic information, used the K-Nearest Neighbor (KNN) algorithm to predict the average driving speed of road segments, and proposed a path planning algorithm based on data prediction to find the path with the shortest driving time. Additionally, according to the prediction results and the number of concurrent requests for road segments, a load balancing strategy was proposed to obtain the path with the shortest travel time while maintaining global load balancing. Although this method considered the load balancing problem in the path planning process, it lacked the analysis of information interaction among vehicles and the impact of vehicle behavior on the environmental state, which would lead to inconsistency with the actual traffic situation. Moreover, the data prediction-based method requires a large amount of computational resources, easily resulting in computational overload, thus being unable to effectively obtain real-time traffic conditions, and having a high prediction error. The proposed path planning method based on heuristic algorithms needs to traverse all possible results, with a large number of iterations, suffering from problems such as high complexity and slow response speed.
[0006] Another example is 2) An Autonomous Lane-Changing System with Knowledge Accumulation and Transfer Assisted by Vehicular Blockchain. It utilizes the collective intelligence shared by Connected Automated Vehicles (CAVs) to address the problems of limited driving scenarios involved in a single CAV and low efficiency of independent learning methods. Vehicular blockchain is applied to ensure the security and privacy of users and data. The introduction of blockchain can encourage more users to participate in collective learning. A machine learning (ML) model using deep reinforcement learning (DRL) is used for autonomous driving decision-making. The lane-changing problem is modeled as a DRL process, and an autonomous lane-changing strategy is learned through the Deep Deterministic Policy Gradient (DDPG) algorithm. To accelerate the learning process and further reduce the communication burden, the corresponding knowledge is extracted from the ML model as shared privileged information instead of directly sharing the local ML model. This method uses blockchain technology for information and knowledge sharing and the reinforcement learning method for vehicle individual training decisions. However, it only considers the independent decision-making of a single vehicle and does not perform collaborative decision optimization among vehicles. At the same time, vehicle decision-making consumes a large amount of computing resources. Due to the lack of global path planning and scheduling for vehicles, the load of edge computing nodes and the density of vehicles will seriously affect vehicle decision-making latency and traffic efficiency, resulting in high computing latency and travel time.
[0007] Given that current path planning methods focus on reducing travel time, lack joint optimization with computing latency, and ignore the fact that different types of vehicles have different service requirements in the mixed driving scenario of the vehicle network. In addition, single-vehicle decision-making cannot obtain the global optimal strategy. Therefore, it is necessary to establish a safe and efficient vehicle collaboration model for active load balancing to meet the service requirements of different vehicles. Summary of the Invention
[0008] Aiming at the shortcomings of current path planning methods, the present invention proposes a blockchain-based collaborative path planning and scheduling method to actively balance the traffic environment and MECNs in the cognitive vehicle network scenario, thereby alleviating traffic congestion and reducing computing latency.
[0009] To achieve the above object, the present invention provides the following technical solutions:
[0010] A collaborative path planning and scheduling method based on blockchain in the cognitive vehicle-to-everything (V2X) scenario, where each vehicle in the physical space is mapped to a virtual node in the blockchain network, and micro base stations or roadside units are deployed as mobile edge computing nodes on each road section in the urban area and equipped with edge computing servers, and macro base stations are connected to remote cloud servers; the method comprises the following steps:
[0011] S1. Sense the environmental state, including the road traffic state and the load distribution of edge computing nodes;
[0012] S2. For connected autonomous vehicles and connected manned vehicles in the mixed driving scenario, establish Markov decision processes respectively. The vehicles obtain the path planning decision that maximizes their cumulative rewards by using the Q-learning reinforcement learning method according to the sensed environmental state;
[0013] S3. Hash and digest the decision information, sign it with its own private key, and package and upload it to the cognitive engine of the mobile edge computing node for verification to obtain the update of the environmental state and the decision consensus;
[0014] S4. The roadside unit verifies the decision information by using the public key of the uploading node. If the verification passes, the cognitive engine iteratively updates or batch-updates the environmental state according to the decision information of the vehicle and feeds it back to other cooperative vehicles;
[0015] S5. Other cooperative vehicles update the reward function of the Markov decision process according to the updated environmental state, and make distributed cooperative path planning and scheduling decisions;
[0016] S6. Through continuous iteration, the vehicle obtains the globally optimal path planning and scheduling strategy, and the mobile edge computing node packages the decision information block for consensus.
[0017] Furthermore, the vehicle-to-everything (V2X) network uses the PC5 interface for vehicle-to-vehicle and vehicle-to-infrastructure communication, and uses the Uu interface for vehicle-to-network communication. Adjacent roadside units use different frequency bands. The communication rate between the m-th roadside unit and the n-th vehicle it serves is expressed as:
[0018]
[0019] where δ 2 represents the noise power of additive white Gaussian noise with a mean of 0 and a variance of δ 2 , and the fraction in the parentheses represents the signal-to-interference-plus-noise ratio (SINR) of the n-th vehicle served by the m-th roadside unit. represents the interference from other micro base stations in the vehicle-to-everything (V2X) scenario , h n,m is the wireless channel gain between vehicle n and roadside unit m, p n,m is the transmission power from the vehicle to roadside unit m, and Bm The bandwidth is (t), and N n,m (t) represents the number of vehicles served by the roadside unit m.
[0020] Furthermore, the road traffic status in step S1 includes vehicle mobility and traffic flow, where:
[0021] Vehicle mobility represents the driving time of vehicle n on section g:
[0022]
[0023] where L g is the length of section g, v g is the maximum speed limit of section g, N g (t) is the number of vehicles on section g at time t, V g is the estimated speed, N jam is the maximum number of vehicles when the road is congested;
[0024] Traffic flow is expressed as the vehicle flow on section g at time t:
[0025] N g (t) = N g (t - 1) + f in,g (t) - f out,g (t)
[0026] where F g (t) = f in,g (t) - fo ut,g (t) is the change in vehicle flow, f in,g (t) and f out,g (t) are the inflow and outflow vehicle flows respectively, N g (t) ≥ 0.
[0027] Furthermore, the calculation formula for the load distribution of the edge computing node m in step S1 is:
[0028]
[0029] where χ(t) is the proportion of connected autonomous vehicles, N g (t) ≥ 0 is the vehicle flow on section g at time t, J n (t) is the computing task volume of the connected autonomous vehicle n. The total computing delay of all connected autonomous vehicles in the system unloading tasks to the edge computing node for processing is expressed as:
[0030]
[0031] where q is the total number of connected autonomous vehicles, M is the total number of edge computing nodes, T nmis the total time delay for the connected autonomous vehicle n to offload the computing task to the edge computing node m for processing.
[0032] Further, the Markov decision process is established in step S2 as follows:
[0033] 1) Agent: The collaborative agents are the vehicles that learn and explore from the environment, namely the connected autonomous vehicle and the connected manned vehicle;
[0034] 2) State: The state represents the position, category, and environmental state information of each agent. The state of the collaborative agent n is represented as:
[0035] Sn(t) = {positioni(t), classi, Wi(t)} i = 1, …, Na
[0036] where position represents the vehicle position and the selected mobile edge computing node for processing the offloading task, class is the vehicle category, N a is the number of collaborative agents, and W i (t) is the environmental state, which changes dynamically with the actions of the agents;
[0037] 3) Action: At each intersection, there are two discrete actions, southward and westward. The action set of agent n from the starting point to the ending point is The collaborative action decision of multiple agents is represented as:
[0038]
[0039] where H is the number of action steps of each agent, and K Na is the signature of the collaborative agent;
[0040] 4) Environmental state update: Multiple agents update the environmental state W(t) according to the actions Two strategies are adopted, namely iterative update: updating W(t) according to the actions of each agent, and batch update: updating based on the actions of the group; the environmental state update is represented as:
[0041]
[0042] Then, the collaborative agents learn from the updated environmental state and find the optimal decision-making strategy with the maximum reward, so as to maximize the global cumulative reward;
[0043] 5) Reward: Normalize the vectors V g and Load m The rewards for different types of agents i and j are respectively:
[0044]
[0045]
[0046] where λ and β represent weight coefficients respectively, is the penalty factor, and the total cumulative reward of the cooperative agent at time step τ is:
[0047]
[0048] The optimization objective function is to maximize the total cumulative reward obtained by the cooperative agent. The optimal cooperative path planning and scheduling strategy can be obtained by solving the following problem:
[0049]
[0050] Furthermore, the Q-learning reinforcement learning method in step S2 is as follows:
[0051] To obtain the optimal policy for maximum cumulative reward The ε-greedy algorithm is adopted to explore and exploit the action space. The agent randomly selects an action with a probability of ε and selects the maximum Q value Q in the Q-table with a probability of 1 - ε * for the corresponding action. The action selection of the agent is expressed as:
[0052]
[0053] A dynamic decay strategy based on ε-greedy is designed to adjust the exploitation and exploration ratio of the distributed multi-agent reinforcement learning algorithm. The dynamic decay update function of ε-greedy is:
[0054] ε(τ + 1) = ε(τ) × (1 - ε Decay )
[0055] where ε Decay is the decay factor of ε, and the value of ε is updated iteratively in each episode;
[0056] The time difference method is adopted to update the Q value Q(s τ , a τ ). The update strategy of the Q value function is as follows:
[0057]
[0058] where the part in parentheses is the loss function, and the learning rate 0 < α < 1.
[0059] Furthermore, in step S3, the consensus process adopts a vote-based decentralized consensus algorithm and simultaneously introduces proof of assets and proof of reputation to incentivize participants to comply with the consensus rules.
[0060] Furthermore, the consensus process in step S3 is specifically as follows: for each consensus cycle, the authorized mobile edge computing nodes form a group of validators according to the assets; then, the validators vote based on the reputation of the candidates to generate a block packaging group, and then randomly select one from the packaging group by drawing lots as the block producer; in the current consensus cycle, the producer packages the shared decision into a new block with a specific structure, and the producer uses the private key K m pr to sign the new block and broadcast it to the validators to reach a consensus; if more than half of the validators approve the block, the block will be added to the end of the chain.
[0061] Furthermore, the verification process in step S4 uses asymmetric encryption technology and hash functions to verify the identity of the decision and the authenticity of the data and protect the privacy of the participants.
[0062] Furthermore, the decision-making process and the consensus process in step S6 are carried out simultaneously.
[0063] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0064] The present invention considers the scenario of mixed driving of CAVs and COVs, jointly optimizes the vehicle traffic road network and the load balance of edge computing nodes, and proposes a blockchain-based collaboration framework to support the path planning and scheduling of vehicle collaboration. In addition, the present invention establishes a task processing model and analyzes the influencing factors of the computing delay of the proposed collaboration framework. Based on the perceived traffic and network states, a traffic condition and load distribution model is established, the joint optimization problem is modeled, and a distributed reinforcement learning algorithm based on Q-Learning is proposed for collaborative path planning and scheduling to achieve the active load balance of road infrastructure and MECNs, so as to minimize the driving time and computing delay. The present invention meets the different service requirements of different types of vehicles and reduces the computing complexity.
[0065] Compared with the prior art 1, the collaborative path planning and scheduling method based on blockchain in the cognitive vehicle networking scenario proposed by the present invention uses blockchain technology to share individual decision-making information, calculates the impact of vehicle individual decisions on the environment, and has a faster convergence speed and lower computing complexity without sacrificing the path planning and load balance performance.
[0066] Compared with the prior art 2, the collaborative path planning and scheduling method based on blockchain in the cognitive vehicle networking scenario proposed by the present invention extends the decision-making algorithm based on reinforcement learning in the single-agent decision-making scenario to the vehicle collaboration scenario, uses the distributed reinforcement learning algorithm and jointly optimizes the load balance of roads and edge computing nodes, effectively avoiding the aggregation effect caused by vehicle selfish behavior, and realizing the global collaborative path planning and scheduling optimization of road traffic infrastructure and edge computing nodes MECNs in the mixed driving scenario on the premise of ensuring low computational latency and low feedback overhead, which has more practical value. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.
[0068] Figure 1 It is the system architecture diagram of the collaborative path planning and scheduling method based on blockchain in the cognitive vehicle networking scenario provided by the embodiment of the present invention;
[0069] Figure 2 It is the model schematic diagram of the collaborative path planning and scheduling method based on blockchain in the cognitive vehicle networking scenario provided by the embodiment of the present invention;
[0070] Figure 3 It is the road network schematic diagram provided by the embodiment of the present application;
[0071] Figure 4 It is the Markov decision process schematic diagram of the path planning problem provided by the embodiment of the present application;
[0072] Figure 5 It is the process schematic diagram of the collaborative path planning and scheduling method based on blockchain in the cognitive vehicle networking scenario provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0073] The present invention aims at the joint optimization problem of vehicle collaborative decision-making path planning and load balancing in the cognitive Internet of Vehicles scenario, and innovatively designs a collaborative path planning and scheduling method based on blockchain. Vehicles share local path planning and scheduling decisions to support global collaborative decision optimization. In the collaborative path planning and scheduling process, the traffic status and MECNs computing load and other characteristics in the cognitive Internet of Vehicles scenario are analyzed through information interaction to construct a traffic situation model and a computing task processing model. Then, a joint optimization problem is modeled to minimize the driving time and computing delay by finding the global optimal path. To solve this problem, a distributed multi-agent reinforcement learning (DMARL) algorithm based on Q-learning is proposed to obtain the optimal global path planning and scheduling decisions for active load balancing of traffic roads and MECNs, thereby solving the problems of traffic congestion and increased computing delay caused by uneven distribution of vehicles and computing load.
[0074] In order to enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0075] 1 System architecture of the present invention
[0076] Figure 1 The system architecture of the blockchain-enabled collaboration framework for CIoVs scenarios. Each participant in the physical space is mapped as a virtual node in the blockchain network. The vehicle set is represented as The set of MECNs is represented as In mixed driving scenarios, CAVs and COVs have different communication, computing, and control capabilities, and have different driving and computing service requirements. CAVs generate a lot of computing tasks with low latency requirements to support safe connected autonomous driving systems. Meanwhile, manned COVs are more focused on improving driving experience and reducing driving time. The set of CAVs and COVs is represented as As shown in the figure, micro base stations or roadside units (RSUs) are deployed as MECNs on each road segment g in the urban area and equipped with edge computing servers to provide latency-sensitive computing services. CAVs with limited computing resources can offload latency-sensitive and computationally intensive tasks to MECNs. At the same time, macro base stations are connected to remote cloud servers to provide vehicles with information and entertainment services that are not sensitive to delays. Each MECN serves a road segment area and multiple vehicles. Vehicles can be divided into different clusters or clusters as batch collaboration units, and each vehicle is served by one MECN within a time slot within the coverage area of MECNs.
[0077] System Model of the Present Invention
[0078] (1) Blockchain-based Collaboration Model
[0079] As Figure 2 shown, the present invention designs a blockchain-based collaboration model to support auditable, traceable, and trustworthy collaboration path planning and scheduling. Using a consortium blockchain, the consensus process adopts a Voting-based Decentralized Consensus (VDC) algorithm. Considering energy consumption and time efficiency, the consensus algorithm introduces Proof of Asset and Proof of Reputation to incentivize participants to comply with the consensus rules. Asset and reputation are the asset and reputation rewards obtained by nodes for complying with the consensus rules.
[0080] The consensus process is as follows:
[0081] For each consensus cycle, authorized MECNs form a group of verifiers based on assets. Then, the verifiers vote according to the reputation of the candidates to generate a block packaging group, and then randomly select one from the packaging group by lottery as the block producer. As Figure 2 shown, in the current consensus cycle, the producer packs the shared decision into a new block with a specific structure. The producer uses the private key K m pr to sign the new block and broadcast it to the verifiers to reach a consensus. If more than half of the verifiers approve the block, the block will be added to the end of the chain. Once a consensus is reached, the collaboration decision will be permanently and securely stored on the blockchain, which helps with the traceability of traffic events. During the collaboration process, vehicles do not obtain decision information, but adjust their own strategies according to the environmental state changes caused by other vehicles and updated by the cognitive engine.
[0082] In addition, the present invention uses asymmetric encryption technology and hash functions to verify the identity of the decision and the authenticity of the data and protect the privacy of participants.
[0083] Taking vehicle n as an example, vehicle n registers with the blockchain system to obtain a public-private key pair K n pu and K n pr to ensure the legitimacy of the node. The private key K n pr is used to sign the decision digest to ensure the authenticity and legitimacy of the uploaded decision information. Then, the receiver uses the public key K n pu to verify the identity of the sender and verify the decision hash value to ensure that the received message has not been tampered with. Any other node can verify the signature using the public key of the signer.
[0084] In addition, if a traffic accident or incident requires liability determination and investigation, the service requester will pay an appropriate token to access the information on the blockchain.
[0085] The blockchain-enabled collaborative path planning and scheduling process is as follows:
[0086] First, based on the perceived environmental state, the vehicle performs path planning locally using the Q-learning method and then shares the signed decision with the cognitive engine deployed in the nearby MECN.
[0087] Then, after the decision information is verified, the cognitive engine updates the environmental state related to traffic conditions and computational load according to the shared decision information. Other collaborative vehicles dynamically adjust the reward function and make decisions based on the updated environmental state.
[0088] Finally, the vehicle obtains the globally optimal collaborative decision strategy, and the MECNs package them into new blocks and reach a consensus to achieve consistency. The distributed decision-making and VDC consensus processes are carried out simultaneously.
[0089] (2) Communication Model
[0090] The vehicle-to-everything (V2X) network uses the PC5 interface for vehicle-to-vehicle (V2V) and vehicle-to-infrastructure (V2I) communications and the Uu interface for vehicle-to-network (V2N) communications. Assume that the wireless channel gain between vehicle n and RSU m is h n,m , which follows an exponential distribution due to Rayleigh fading, and the transmission power from the vehicle to the RSU is p n,m . Assume the bandwidth is B m (t), and N n,m (t) represents the number of vehicles served by RSU m. The bandwidth B m (t) is dynamically allocated to the served vehicles in time slot t. In addition, to avoid inter-cell interference, adjacent RSUs use different frequency bands. The communication rate between RSU m and the nth served vehicle can be expressed as:
[0091]
[0092] where δ 2 represents a mean of 0 and a variance of δ 2The noise power of additive Gaussian white noise (AGWN). The fraction in the parentheses represents the signal-to-interference-plus-noise ratio (SINR) of the nth vehicle served by the mth RSU. Indicates interference from other micro base stations in the vehicle networking scenario Interference.
[0093] (3) Traffic state model
[0094] 1) Mobility model
[0095] Based on traffic environment perception, a traffic situation model is established. Assume that the number of vehicles on section g at time t is N g (t), then the vehicle density on section g can be deduced as:
[0096]
[0097] where l n is the length of vehicle n, L g is the length of section g, u g is the number of lanes. The estimated speed V g is related to the vehicle density on the road and can be expressed as:
[0098]
[0099] where v g is the maximum speed limit of section g, is the vehicle density when section g is congested. Since the total length of the road is fixed, it can be deduced that N jam is the maximum number of vehicles when the road is congested. Then the driving time of vehicle n on section g can be deduced as:
[0100]
[0101] 2) Traffic flow model
[0102] A traffic flow model is established to adapt to dynamic traffic conditions. The inflow and outflow directly affect the congestion change of the road, which can be sensed through inductive loop, vehicle sensors and cameras. The inflow and outflow vehicle flows of section g at time slot t are f in,g (t) and f out,g (t). The vehicle flow on section g at time t is deduced as:
[0103]
[0104] where N gf(t) ≥ 0. If f in,g (t) < f out,g (t), then the number of vehicles on the road is decreasing, and the road congestion level and the load on MECNs nodes can be reduced. When it continuously decreases and even N g (t) = 0, it will lead to a decrease in the utilization rate of computing resources and road infrastructure, especially during peak traffic hours. On the contrary, if f in,g (t) > f out,g (t), it will increase the load on MECNs and the road. Define the change in traffic flow as F g (t) = f in,g (t) - f out,g (t), where F g (t) > 0 means that the number of vehicles flowing into the section is greater than the outflow, and |F g (t)| represents the number of vehicles added to section g, and vice versa. To obtain a better driving experience and computing performance, COVs and CAVs should be reasonably scheduled without causing congestion and make full use of road and edge computing resources without causing congestion and without exceeding the computing load of MECN.
[0105] 3) Calculation load distribution
[0106] The computing load of MECNs is closely related to the vehicle density and the proportion of CAVs with various computing tasks such as image recognition, target detection, and decision control. Based on traffic condition perception, the load distribution of MECN m is derived as:
[0107]
[0108] where J n (t) is the amount of computing tasks of CAV n, and χ(t) is the proportion of CAVs. As the vehicle density and the proportion of CAVs increase, a large number of computing tasks will be offloaded to the connected MECN, increasing the load on MECN.
[0109] (4) Computing task processing model
[0110] The delay for offloading computing tasks to MECNs for processing includes transmission delay, queuing waiting delay, and computing delay. It depends on the number of tasks, bandwidth, and available computing resources of MECNs, that is, the number of CPU cycles per second of the central processing unit. Assume that CAVn has J different computing tasks, and the data volume of each task is The computing resources required to complete task j are Computing task j needs to be completed within the specified maximum delay . The transmission delay for CAV n to offload computing tasks to MECNm can be expressed as:
[0111]
[0112] where represents the selection factor for CAV n to offload the computing task j to the MECNs, means that CAVn offloads task j, otherwise Similarly, b nm = 1 indicates that CAV n offloads the computing task to MECNm, otherwise b nm = 0.
[0113] If the available computing resources of MECN m are less than the minimum computing resources of task j then task j needs to queue up. The MECN uses a First Input First Output (FIFO) queue to process the arriving tasks. Assume that MECNm has C cap computing power, and there are Φ computing tasks waiting to be processed before task j. The queuing factor of task j can be expressed as:
[0114]
[0115] where is the computing resources occupied by the previous task i, represents the computing resources occupied by MECN m. The queuing delay of task j offloaded from CAVn to MECN m can be expressed as:
[0116]
[0117] where is the computing resources required by the previous task i. Then, the computing delay of CAV n can be derived as:
[0118]
[0119] where is the computing resources allocated to task j, is the computing resources required for CAV n to complete task j, N Φ represents the number of service vehicles corresponding to the previous Φ tasks. The total delay for CAV n to offload the computing task to MECN m for processing is:
[0120]
[0121] The total computing delay of all CAVs offloading tasks can be used as an index to evaluate the computing performance of the present invention, and can be derived as:
[0122]
[0123] where q is the total number of CAVs and M is the total number of MECNs.
[0124] 3 Technical Solution of the Invention
[0125] (1) Construction of Load Balancing Metrics
[0126] Divide the map area into different road segments The vehicle density set of the road segment and the calculation load set of MECNs are respectively and To obtain the load balance of road infrastructure and MECN, define the road load balance metric and the MECNs load balance metric Expressed as:
[0127]
[0128]
[0129] where and are the vehicle density of road segment g and the average vehicle density of all road segments; Load m and are the load of MECN m and the average load of all MECNs respectively. The uniformity of load distribution is inversely proportional to the metric values and The smaller the value, the better the load balancing performance.
[0130] (2) Modeling of Joint Optimization Problem
[0131] Since the MECNs load balance metric directly affects the total calculation delay T m , the present invention models a joint optimization problem to obtain the optimal cooperative path planning and scheduling strategy Ψ = {ψ1, ψ2,..., ψ q , ψ q+1 ,..., ψ N} as follows:
[0132]
[0133]
[0134]
[0135]
[0136]
[0137] where λ1 and λ2 are jointly optimized weights, is the number of CAVs scheduled from road segment g to the adjacent road segment g′ at time t, represents the set of adjacent road segments of road segment g, and the optimization objective is to minimize the road load balancing metric and the total computing delay T m .
[0138] Embodiment
[0139] As Figure 3 shown, the present invention uses the OpenStreetMap open-source map to construct a road network, which is modeled as a grid with vertices and edges. Since the traffic environment state changes dynamically over time and space and has strong short-term correlations, large-scale planning at one time lacks effective foresight of the future environment. Therefore, a 3×3 map network is selected within a limited area and modeled as Figure 3 structure:
[0140] Graph(t) = <V, E, W(t)> (16)
[0141] where V represents vertices, E represents edges, and W(t) is a weight matrix representing the computing load distribution and vehicle density situation. Each edge represents a road segment, and the vertices represent intersections. The weight matrix W(t) = {V g (t), Load m (t)}, characterizes the vehicle driving speed and the computing load of MECNs.
[0142] The present invention models the path planning problem as a Markov Decision Process (MDP). As Figure 4 shown, the numbers of the intersections correspond to {G = 1, D = 2, H = 3, A = 4, E = 5, I = 6, B = 7, F = 8, C = 9}, where node G represents the starting point and node C represents the destination.
[0143] The collaborative agents, states, actions, and rewards of distributed multi-agent reinforcement learning are introduced as follows:
[0144] 1) Agents: The collaborative agents are vehicles that learn and explore from the environment, i.e., CAVs and COVs.
[0145] 2) States: The states represent the position, category, and environmental state information of each agent. The state of collaborative agent n can be expressed as:
[0146]
[0147] where position represents the position of the vehicle and the MECN selected to handle the unloading task, class is the category of the vehicle, and N a is the number of collaborative agents, and W i (t) is the environmental state, which changes dynamically with the actions of the agents.
[0148] 3) Actions: The MDP has two discrete actions of southward and westward at each intersection, namely ("south↓"; "west←"). The action set of agent n from the starting point to the ending point is The collaborative action decision of multiple agents can be expressed as:
[0149]
[0150] where H is the number of action steps of each agent, is the signature of the collaborative agent.
[0151] 4) Environmental state update: Since the decisions of the agents affect the environmental state and the decisions of other agents, the environmental state needs to be updated dynamically according to the action strategy. Therefore, multiple agents update the environmental state W(t) according to the actions of the agents This invention considers two strategies, namely iterative update: updating W(t) according to the actions of each agent; batch update: updating based on the actions of the group. The environmental state update is expressed as:
[0152]
[0153] Then, the collaborative agents learn from the updated environmental state and find the optimal decision-making strategy with the maximum reward, so as to maximize the global cumulative reward.
[0154] 5) Rewards: In the mixed driving scenario, the rewards of different types of vehicles are related to the driving speed V g and the computing load Load of the MECNsm m are relevant. Normalize the vectors V g and Load m , and the rewards of different types of agents i and j are respectively:
[0155]
[0156]
[0157] where λ and β represent weight coefficients respectively, is the penalty factor, and the total cumulative reward of the collaborative agent at time step τ is:
[0158]
[0159] The optimization objective function is to maximize the total cumulative reward obtained by the collaborative agents. Thus, the optimal collaborative path planning and scheduling strategy can be obtained by solving the following problem:
[0160]
[0161] To obtain the optimal strategy for maximum cumulative reward The ε-greedy algorithm is used to explore and exploit the action space. The agent randomly selects an action with probability ε and selects the action with the maximum Q value in the Q-table with probability 1 - ε * for the corresponding action. When the agent is not familiar enough with the environment, it needs more exploration, and ε will decay as learning progresses. The action selection of the agent can be expressed as:
[0162]
[0163] A dynamic decay strategy based on ε-greedy is designed to adjust the exploitation and exploration ratio of the distributed multi-agent reinforcement learning algorithm. As the number of training episodes increases, ε gradually decreases until the training ends or reaches the minimum value. The dynamic decay update function of ε-greedy is:
[0164] ε(τ + 1) = ε(τ) × (1 - ε Decay ) (25)
[0165] where ε Decay is the decay factor of ε. The value of ε is iteratively updated in each episode. This is because the agent usually does not understand the environment at the beginning and needs a higher proportion of random actions to obtain experience.
[0166] Based on the reward from the environmental feedback, Q-learning uses the temporal difference method to update the Q value, Q(s τ , a τ ), and the update strategy of the Q-value function is derived as follows:
[0167]
[0168] where the part in parentheses is the loss function, and the learning rate 0 < α < 1.
[0169] In summary, the method flow of the present invention is as shown in Figure 5As shown in the figure. In the initial stage of the distributed multi-agent reinforcement learning algorithm, the environmental state is first initialized, which includes the road traffic state and the load distribution of edge computing nodes, and hyperparameters are set. For different types of vehicles in the mixed driving scenario, Markov decision processes are respectively established. Then, the agent uses the Q-learning reinforcement learning method to obtain the path planning decision that maximizes its own cumulative reward, and then uploads the decision information to the cognitive engine for environmental state update and decision consensus. The cognitive engine iteratively updates or batch-updates the environmental state according to the agent's decision and feeds it back to the cooperative agent. Then, the cooperative agent makes its own path planning and scheduling decisions according to the updated environmental state to maximize the cumulative reward it obtains. Through continuous iteration, the cooperative agent obtains the globally optimal path planning and scheduling strategy and adds it to the blockchain through consensus, reducing traffic congestion and travel time, and reducing the load and computing delay of edge computing nodes. The present invention meets the different service requirements of different types of vehicles and reduces the computational complexity.
[0170] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features, but these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A collaborative path planning and scheduling method based on blockchain in the cognitive vehicle networking scenario, characterized in that Each vehicle in the physical space is mapped to a virtual node in the blockchain network. Micro base stations or roadside units are deployed as mobile edge computing nodes on each section of the urban area and equipped with edge computing servers. Macro base stations are connected to remote cloud servers. The method comprises the following steps: S1. Sense the environmental state, including the road traffic state and the load distribution of edge computing nodes; S2. For the connected autonomous vehicles and connected manned vehicles in the mixed driving scenario, establish Markov decision processes respectively. The vehicles obtain the path planning decision that maximizes their cumulative rewards by using the Q-learning reinforcement learning method according to the sensed environmental state. The establishment of the Markov decision process in step S2 is as follows: 1) Agent: The cooperative agents are the vehicles that learn and explore from the environment, i.e., connected autonomous vehicles and connected manned vehicles; 2) State: The state represents the position, category and environmental state information of each agent. The state of cooperative agent n is expressed as: where position represents the position of the vehicle and the selected mobile edge computing node for processing the unloading task, class is the category of the vehicle, N a is the number of cooperative agents, W i (t) is the environmental state, which changes dynamically with the actions of the agents; 3) Actions: At each intersection, there are two discrete actions, southward and westward. The action set of agent n from the starting point to the end point is The collaborative action decision of multiple agents is represented as: where H is the number of action steps of each agent, is the signature of the collaborative agent; 4) Environmental state update: Multiple agents update the environmental state W(t) according to their actions There are two strategies for updating the environmental state W(t), namely iterative update, where W(t) is updated based on the actions of each agent, and batch update, where it is updated based on the collective actions of the group. The environmental state update is expressed as: Then, the cooperative agents learn from the updated environmental state and find the optimal decision-making strategy with the maximum reward, so as to maximize the global cumulative reward; 5) Reward: For vector V g and Load m are normalized, and the rewards for different types of agents i and j are respectively: where λ and β represent the weight coefficients respectively, is the penalty factor, and the total cumulative reward of the collaborative agent at time step τ is: The optimization objective function is to maximize the total cumulative reward obtained by the cooperative agents. The optimal cooperative path planning and scheduling strategy can be obtained by solving the following problem: S3. Hash and summarize the decision information, sign it with its own private key, and package and upload it to the cognitive engine of the mobile edge computing node for verification to obtain the update of the environmental state and decision consensus; S4. The roadside unit verifies the decision information by using the public key of the uploading node. If the verification passes, the cognitive engine iteratively updates or batch-updates the environmental state according to the decision information of the vehicle and feeds it back to other cooperative vehicles; S5. Other cooperative vehicles update the reward function of the Markov decision process according to the updated environmental state, and make distributed cooperative path planning and scheduling decisions; S6. Through continuous iteration, the vehicles obtain the global optimal path planning and scheduling strategy, and the mobile edge computing node packages the decision information block for consensus.
2. The collaborative path planning and scheduling method based on blockchain in the cognitive vehicle-to-everything scenario according to claim 1, wherein The vehicle-to-everything (V2X) network uses the PC5 interface for vehicle-to-vehicle and vehicle-to-infrastructure communications, and uses the Uu interface for vehicle-to-network communications. Adjacent roadside units use different frequency bands. The communication rate between roadside unit m and the nth vehicle it serves is expressed as: where δ 2 represents the noise power of additive white Gaussian noise with a mean of 0 and a variance of δ 2 , and the fraction in the parentheses represents the signal-to-interference-plus-noise ratio of the nth vehicle served by the mth roadside unit. represents the interference from other small base stations in the vehicle-to-everything scenario , h n,m is the wireless channel gain between vehicle n and roadside unit m, p n,m is the transmission power from the vehicle to roadside unit m, B m (t) is the bandwidth, and N n,m (t) represents the number of vehicles served by roadside unit m.
3. The collaborative path planning and scheduling method based on blockchain in the cognitive vehicle networking scenario according to claim 1, characterized in that In step S1, the road traffic state includes vehicle mobility and traffic flow, where: Vehicle mobility represents the driving time of vehicle n on section g: where L g is the length of section g, v g is the maximum speed limit of section g, N g (t) is the number of vehicles on section g at time t, V g is the estimated speed, N jam is the maximum number of vehicles when the road is congested; Traffic flow is expressed as the traffic volume on section g at time t: N g N(t) = g N(t - 1)+f in,g N(t)-f out,g N(t) Among them, F g (t) = f in,g (t) - f out,g (t) is the change in traffic flow, f in,g (t) and f out,g (t) are the incoming and outgoing traffic flows respectively, N g (t) ≥ 0.
4. The collaborative path planning and scheduling method based on blockchain in the cognitive vehicle networking scenario according to claim 1, wherein The calculation formula for the load distribution of edge computing node m in step S1 is: where χ(t) is the proportion of connected and autonomous vehicles, and N g (t) ≥ 0 is the traffic flow on road segment g at time t, and J n (t) is the computing task volume of connected and autonomous vehicle n. The total computing delay of all connected and autonomous vehicles in the system unloading tasks to the edge computing node for processing is expressed as: where q is the total number of connected autonomous vehicles, M is the total number of edge computing nodes, and T nm is the total latency for connected autonomous vehicle n to offload its computing tasks to edge computing node m for processing.
5. The collaborative path planning and scheduling method based on blockchain in the cognitive vehicle networking scenario according to claim 1, characterized in that The Q-learning reinforcement learning method in step S2 is as follows: To achieve the maximum cumulative reward and obtain the optimal strategy The ε-greedy algorithm is used to explore and exploit the action space. The agent randomly selects an action with probability ε and selects the action corresponding to the maximum Q value in the Q-table with probability 1-ε * The corresponding action The action selection of the agent is expressed as: Design a dynamic decay strategy based on ε-greedy to adjust the exploitation and exploration ratio of the distributed multi-agent reinforcement learning algorithm. The dynamic decay update function of ε-greedy is: where ε Decay is the decay factor of ε, and the value of ε is iteratively updated in each round; The time-difference method is used to update the Q-value, Q(s τ , a τ ). The update strategy of the Q-value function is as follows: The part in the brackets is the loss function, and the learning rate 0 < α < 1.
6. The collaborative path planning and scheduling method based on blockchain in the cognitive vehicle-to-everything scenario according to claim 1, characterized in that In step S3, the consensus process adopts a voting-based decentralized consensus algorithm and introduces proof of assets and proof of reputation to incentivize participants to abide by the consensus rules.
7. The collaborative path planning and scheduling method based on blockchain in the cognitive vehicle networking scenario according to claim 6, wherein The specific process of step S3 consensus is as follows: For each consensus cycle, the authorized mobile edge computing nodes form a group of validators according to the assets; then, the validators vote based on the reputation of the candidates to generate a block packaging group, and then randomly select one from the packaging group by lottery as the block producer; in the current consensus cycle, the producer packages the shared decision into a new block with a specific structure, and the producer uses the private key to sign the new block and broadcast it to the validators to reach a consensus; if more than half of the validators approve the block, the block will be added to the end of the chain.
8. The collaborative path planning and scheduling method based on blockchain in the cognitive vehicle-to-everything scenario according to claim 1, wherein In step S4, the verification process uses asymmetric encryption technology and hash functions to verify the authenticity of the decision-making identity and data and protect the privacy of participants.
9. The collaborative path planning and scheduling method based on blockchain in the cognitive vehicle-to-everything scenario according to claim 1, wherein In step S6, the decision-making process and the consensus process are carried out simultaneously.
Citation Information
Patent Citations
Vehicle computing task unloading method based on blockchain data sharing
CN112532676A