Method for realizing efficient and reliable data sharing among multiple vehicles in vehicle platooning system
By introducing a trust model and reinforcement learning algorithm to optimize the consensus mechanism in the vehicle platooning system, the reliability and efficiency issues in multi-vehicle data sharing are solved, efficient and reliable data consensus and storage are achieved, and adaptation to complex environmental changes is achieved.
Patent Information
- Application Number
- CN202411358979.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2044-09-27
AI Technical Summary
How to achieve efficient and reliable data sharing among multiple fleets in a vehicle platooning system, solve the problem that existing solutions are difficult to strike a balance between consensus reliability and efficiency, and lack an effective trust management model.
A vehicle platoon trust model is used to evaluate and manage the reliability of the fleet participating in the consensus. Artificial intelligence technology is used to optimize the performance of the consensus process. Data sharing is achieved through a hierarchical structure and blockchain network. The consensus mechanism is optimized by combining trust value calculation and reinforcement learning algorithm.
It achieves efficient and reliable data consensus and storage among multiple fleets, ensures information security and timeliness, dynamically adapts to network changes, and balances performance and reliability.
Smart Images

Figure CN119316819B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for realizing efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system, and belongs to the technical field of vehicle networking, in particular to the technical field of vehicle platooning in the vehicle networking. Background Art
[0002] With the continuous advancement of connected vehicles and edge computing technologies, vehicle platooning systems have become a key research area in intelligent transportation systems. Scientifically designed vehicle platooning systems can improve traffic safety, alleviate road congestion, and reduce vehicle energy consumption. However, the implementation of vehicle platooning systems relies on real-time data communication and information sharing between different fleets. Multiple fleets must reach consensus on shared data to achieve a comprehensive and reliable global traffic perception and facilitate collaborative fleet control. Blockchain technology, which provides secure and reliable data management services and consensus mechanisms for multi-fleet data consensus and storage, can be an effective solution. Fleets upload the real-time data they need to share to a blockchain system for consensus and secure data storage. Furthermore, blockchain systems can be deployed on edge devices, such as road test units (RSUs), to form a blockchain network, addressing the issue of limited vehicle computing and storage resources.
[0003] However, due to the real-time nature of vehicle data and its importance to traffic safety, achieving efficient and reliable data consensus among multiple fleets is a key challenge facing this research effort. While some studies have proposed consensus solutions, these solutions still have certain shortcomings. For example, some solutions guarantee the reliability of consensus but fail to meet efficiency requirements; others focus on addressing performance issues but ignore the security and reliability of the consensus process. Few studies have been able to strike a good balance between the two. Furthermore, ensuring consensus reliability requires an assessment of the trustworthiness and reliability of different fleets, and existing IoV trust solutions lack a trust management model tailored for fleet networks.
[0004] How to effectively solve the above-mentioned problems existing in the current vehicle platooning technology in the Internet of Vehicles and realize efficient and reliable data sharing between multiple fleets in the vehicle platooning system is a technical problem that urgently needs to be solved in the field of vehicle platooning technology in the Internet of Vehicles. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to invent a consensus optimization technology solution for vehicle platooning networks, which uses a vehicle platooning trust model to evaluate and manage the reliability of the fleet participating in the consensus, and uses artificial intelligence technology to optimize the performance of the consensus process, so as to achieve the technical goal of efficient and reliable fleet data consensus.
[0006] To achieve the above-mentioned object, the present invention proposes a method for realizing efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system, the method comprising the following steps:
[0007] (1) The vehicle platooning system is divided into two parts: the vehicle platooning layer and the blockchain network layer. The vehicle platooning layer is a vehicle platooning communication network composed of multiple platoons on the road. Each platoon consists of a platoon leader (PH) and several platoon members (PMs). The blockchain network layer is the core of the entire vehicle platooning system and is a permissioned blockchain network composed of edge devices (RSUs) with sufficient computing and storage resources.
[0008] (2) The vehicle platooning system performs vehicle registration. After each vehicle is authenticated by the traffic management agency, a certificate and a key are issued to complete the vehicle registration and join the vehicle platooning system.
[0009] (3) The vehicle platooning system performs intra-platoon communication. The PH of each platoon is responsible for leading the entire platoon. It needs to periodically communicate with other platoons and edge devices (RSUs) to obtain traffic data and pass real-time safety information and driving strategy instructions to the PMs of the platoon to ensure the safety and consistency of the entire platoon.
[0010] (4) The vehicle platooning system performs inter-team communication, that is, data sharing is achieved between fleets through the blockchain network layer. The sender of the data first uploads the data to the blockchain network layer. After the data reaches a consensus and is stored, other fleets can access the blockchain network layer and obtain the data shared by the sender after obtaining permission.
[0011] (5) The vehicle platooning system performs trust value calculation. The vehicle platooning system calculates the trust value of each vehicle in each platoon, especially the PH. The specific process is as follows: the vehicle platooning layer periodically uploads the external evaluation scores between different platoons and the internal evaluation scores of platoon members to the local RSU, which calculates the trust value of all vehicles according to the set trust value calculation method and saves the results to the blockchain network layer;
[0012] (6) The vehicle platooning system executes a consensus mechanism to achieve data sharing; the blockchain network layer receives data uploaded by all fleets, including shared traffic data and trust value data, and packages the data into blocks, runs the set consensus mechanism, reaches consensus, and completes data storage and sharing.
[0013] In step (5), the calculation method of the PH trust value of the team leader is as follows:
[0014] The time period T is divided into multiple time slots. At time slot t, the trust value of the team leader PH is calculated according to the following formula:
[0015]
[0016] In the above formula, The team leader h at time slot t i The trust value, Indicates the team leader h i The original trust value of The team leader h at time slot t i The evaluation trust value; ξ is the weight parameter, which implements the "slow rise and fast fall" update strategy of the trust value of the team leader PH, that is, when the original trust value is less than the evaluation trust value, ξ is small, so that the trust value of the team leader rises slowly; on the contrary, ξ is large, so that the trust value of the team leader drops rapidly, and the larger the gap between the two, the larger ξ;
[0017] In the above formula, the team leader h at time slot t i The evaluation trust value is calculated according to the following formula:
[0018]
[0019] In the above formula, 0<ω<1 is used as the weight parameter;
[0020] Indicates the vehicle platoon system's response to the platoon leader h at time slot t i The objective evaluation results are not affected by the subjective evaluation results of the current vehicle, and are completely based on historical data. i The behavior in the past time period T is scored; there are two main reference standards, one is h i The interactive performance of h i The former is based on the experience of the team leader in the past time period T. i The number of positive and negative interactions with other vehicles is calculated using the mean of the Beta distribution, which is obtained by counting the number of h in the past time period T. i The number of successes and failures in the PH campaign are also calculated using the mean of the Beta distribution. The sum of the two and divided by two can be used to get the team leader h. i Objective evaluation results
[0021] Indicates the vehicle platoon system's response to the platoon leader h at time slot t i The subjective evaluation results are calculated according to the following formula:
[0022]
[0023] In the above formula, Indicates the inter-team evaluation value, which is determined by the PH of other teams in h iThe usefulness of the information transmitted in the current time slot t is evaluated, with a score range between 0 and 1. The scores of all PHs participating in the evaluation are summarized and the confidence of these PHs is used to make a weighted average of the scores;
[0024] In the above formula, Indicates the evaluation value within the team, the team head h i Internal members of the team are h i The correctness of the instructions issued in the current time slot t and the quality of the shared security information are evaluated. The correctness index ranges from 0 to 1, where 0 means h i Most of the instructions issued are considered to be wrong, 1 means h i Most issued instructions are considered correct, and the information quality index is scored between 0 and 1. Team members are close to the team leader and have a more comprehensive understanding of the team leader. Therefore, two indicators are set to achieve more accurate internal evaluation. The team score given by each member is obtained by multiplying the correctness index and the information quality index. The scores of all members are aggregated and weighted averaged using each member's confidence.
[0025] The confidence level of the vehicle is calculated according to the following formula:
[0026]
[0027] In the above formula, represents the confidence of vehicle v in time slot t in the past T time period, λ>0 is a constant, It represents the average trust value of vehicle v in the past T time period; the higher the cumulative trust value of the vehicle, the higher the confidence level.
[0028] The specific content of the consensus mechanism implemented by the vehicle platooning system in step (6) includes the following steps:
[0029] (61) Filter out the top N with the highest credibility c Instead of having all nodes participate in the consensus, nodes are used as consensus nodes to reduce communication costs. The nodes are the edge devices RSU that constitute the blockchain network layer.
[0030] (62) The master node and block size are determined according to the set rules. The master node generates the block and broadcasts the consensus request to other consensus nodes for verification;
[0031] (63) After the consensus node completes the verification, it broadcasts the audit results again to other consensus nodes except the master node. If the consensus node receives more than 2f audit results, it means that the verification is valid and sends a confirmation request to the master node; where f represents the number of Byzantine nodes;
[0032] (64) If the master node receives more than 2f+1 confirmation results, it means that the verification is successful, consensus is reached, and the broadcast block can be put on the chain.
[0033] The reputation value of the node is calculated according to the following formula:
[0034]
[0035] In the above formula, l i (t') represents node I i The reputation value at time slot t', 0<σ<1 is the weight parameter, s' represents the node I i The number of successful block production within a specified time, f' represents the number of nodes I i The number of failures to produce blocks within a specified time, K represents the number of failures to produce blocks to node I within the t' time slot. i The number of fleets uploading data, Represents the trust value of fleet PH at time slot t'.
[0036] The specific content of the rules for determining the master node and block size in step (62) is: based on the delay of the consensus process, the throughput and reliability of the blockchain network layer, and according to the dynamic changes of the nodes and network characteristics of the blockchain network layer, the key blockchain parameters, namely the master node and block size, are decided through the reinforcement learning algorithm.
[0037] The reinforcement learning algorithm is a diffusion model-based soft actor-critic algorithm, namely the DMSAC algorithm; the soft actor-critic algorithm, namely the SAC algorithm itself is an algorithm based on the maximum entropy theory, with strong exploration efficiency and stability. The DMSAC algorithm combines the generative artificial intelligence model with the reinforcement learning model and introduces the diffusion model into the SAC algorithm, which can enhance the algorithm's exploration, improve sampling efficiency, model more complex environments, and further optimize the performance of SAC.
[0038] The state space, action space and reward function of the DMSAC algorithm are defined as follows:
[0039] State space s t’ : It consists of the computing resources c of the consensus node, the reputation value l and the data transmission rate R between nodes;
[0040] Action space a t’ : Decision variables include masternode selection and block size;
[0041] Reward function r(s t’ ,a t’ ): defined as follows:
[0042] r(s t’ ,a t’ )=(α1D T +α2Ω T )log(1+l p (t'))
[0043] In the above formula, α1 and α2 are constant coefficients, l p (t') represents the master node I p Reputation value;
[0044] T C represents the consensus delay, which is calculated according to the following formula:
[0045] T C =T G +T V
[0046] Where T G Indicates the block generation delay, which is determined by the block size and the computing power of the master node. V Indicates the block verification delay, which mainly includes the audit and communication delay of the consensus process;
[0047] Ω T It represents the blockchain throughput and is calculated according to the following formula:
[0048]
[0049] In the above formula, S B represents the block size, χ represents the average transaction data size;
[0050] The DMSAC algorithm consists of the following four parts:
[0051] Buffer pool: At each decision time slot t', observe the environment to get the current state s t’ And input it into the model, the model calculates the action probability distribution, and samples the action a t’ And execute, the environment feedback reward value r(s t’ ,a t’ ) and transfer to the next state s t’+1 ; After each round of interaction with the environment, the interaction trajectory data (s t’ ,a t’ ,r(s t’ ,a t’ ),s t’+1 ) is added to the buffer pool and used for model parameter update;
[0052] Policy network: The policy network is used to map the state to the action probability distribution. Unlike the SAC algorithm, the policy network is built based on the diffusion model network rather than a simple multi-layer perceptron. It uses the reverse denoising process of the diffusion model to transform the state s t’ After multiple steps of denoising, and each step of denoising transformation obeys Gaussian distribution, we finally get the action probability distribution, from which we sample action a t’ ;
[0053] Dual Value Network: Similar to SAC, the dual value network is used to evaluate the long-term value of different actions in the current state. It consists of cumulative reward returns and entropy values. The purpose of using a dual network is to reduce errors caused by overestimation.
[0054] Target network: The target network is the corresponding target network of the dual value network. It is used to stabilize the training process and update the parameters using a soft update method.
[0055] The training process of the DMSAC algorithm is as follows:
[0056] In each step of each round of training, observe the blockchain network environment status s t’ , input it into the policy network, the policy network outputs the action probability distribution, from which the action a is sampled t ', decide the master node and block size of the consensus in this stage, execute the consensus process, calculate the latency, throughput and master node reliability, input the reward function to get the reward value r(s t’ ,a t’ ), the blockchain network transitions to the next state s t’+1 ; This trajectory data (s t’ ,a t’ ,r(s t’ ,a t’ ),s t’+1 ) is added to the buffer pool; after a certain amount of data has accumulated in the buffer pool, random sampling is performed from it to update the parameters of the value network and the policy network respectively, and finally the parameters of the target network are soft-updated; at this point, one step in a round of training is completed. After each round of training steps is iterated, the blockchain network environment state is reset and the next round is entered, and this cycle repeats until the training round is completed.
[0057] The beneficial effects of the present invention are as follows: the method of the present invention provides an efficient and reliable consensus mechanism for data sharing between multiple fleets in a vehicle platooning system. The secure data sharing between multiple fleets complies with the data sharing logic in conventional Internet of Vehicles scenarios, while taking into account both internal and inter-fleet communication modes, and utilizing edge computing and blockchain technology to ensure secure and reliable data sharing and storage; the fleet trust model proposed by the method of the present invention and the consensus optimization mechanism based on reinforcement learning achieve efficient and reliable data consensus and storage between multiple fleets. The fleet trust model mainly completes the credibility evaluation of fleet vehicles, especially the fleet head, for multi-fleet scenarios, ensuring the reliability of the subjects participating in the consensus. The consensus optimization mechanism based on reinforcement learning can dynamically adapt to environmental changes in the blockchain network, and implement consensus decisions that take into account both performance and reliability. The consensus process under the combination of the two meets the requirements for information security and timeliness in the vehicle network scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 It is a flow chart of a method for realizing efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system proposed by the present invention. DETAILED DESCRIPTION
[0059] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings.
[0060] See also Figure 1 The present invention introduces a method for realizing efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system. The method includes the following steps:
[0061] (1) The vehicle platoon system is divided into two parts: the vehicle platoon layer and the blockchain network layer. The vehicle platoon layer is a vehicle platoon communication network composed of multiple platoons on the road. Each platoon consists of a platoon head (PH) and several platoon members (PMs). The blockchain network layer is the core of the entire vehicle platoon system and is a permissioned blockchain network composed of edge devices (RSUs) with sufficient computing and storage resources.
[0062] (2) The vehicle platooning system performs vehicle registration. After each vehicle is authenticated by the traffic management agency, a certificate and a key are issued to complete the vehicle registration and join the vehicle platooning system.
[0063] (3) The vehicle platooning system performs intra-platoon communication. The PH of each platoon is responsible for leading the entire platoon. It needs to periodically communicate with other platoons and edge devices (RSUs) to obtain traffic data and pass real-time safety information and driving strategy instructions to the PMs of the platoon to ensure the safety and consistency of the entire platoon.
[0064] (4) The vehicle platooning system performs inter-team communication, that is, data sharing is achieved between fleets through the blockchain network layer. The sender of the data first uploads the data to the blockchain network layer. After the data reaches a consensus and is stored, other fleets can access the blockchain network layer and obtain the data shared by the sender after obtaining permission.
[0065] (5) The vehicle platooning system performs trust value calculation. The vehicle platooning system calculates the trust value of each vehicle in each platoon, especially the PH. The specific process is as follows: the vehicle platooning layer periodically uploads the external evaluation scores between different platoons and the internal evaluation scores of platoon members to the local RSU, which calculates the trust value of all vehicles according to the set trust value calculation method and saves the results to the blockchain network layer;
[0066] (6) The vehicle platooning system executes a consensus mechanism to achieve data sharing; the blockchain network layer receives data uploaded by all fleets, including shared traffic data and trust value data, and packages the data into blocks, runs the set consensus mechanism, reaches consensus, and completes data storage and sharing.
[0067] In step (5), the calculation method of the PH trust value of the team leader is as follows:
[0068] When the vehicle platooning system performs trust value calculation, the time period T is divided into multiple time slots. In this embodiment, the time slot value is 30 seconds. At time slot t, the trust value of the platoon leader PH is calculated according to the following formula:
[0069]
[0070] In the above formula, The team leader h at time slot t i The trust value, Indicates the team leader h i The original trust value of The team leader h at time slot t i The evaluation trust value; ξ is the weight parameter, which implements the "slow rise and fast fall" update strategy of the trust value of the team leader PH, that is, when the original trust value is less than the evaluation trust value, ξ is small, so that the trust value of the team leader rises slowly; on the contrary, ξ is large, so that the trust value of the team leader drops rapidly, and the larger the gap between the two, the larger ξ;
[0071] For example, in the experiment of the embodiment, at time slot 1, the original trust value T of the head of the 5th team is cur(5) =0.5, the evaluation trust value after one round of evaluation The weight ξ is calculated to be 0.57. The updated trust value of the team leader 5 at time slot 1 is obtained by solving the above formula:
[0072] In the above formula, the team leader h at time slot t i The evaluation trust value is calculated according to the following formula:
[0073]
[0074] In the above formula, 0<ω<1 is used as a weight parameter, and the value in the embodiment is 0.7;
[0075] Indicates the vehicle platoon system's response to the platoon leader h at time slot t i The objective evaluation results are not affected by the subjective evaluation results of the current vehicle, and are completely based on historical data. i The behavior in the past time period T is scored; there are two main reference standards, one is h i The interactive performance of h i The former is based on the experience of the team leader in the past time period T. i The number of positive and negative interactions with other vehicles is calculated using the mean of the Beta distribution, which is obtained by counting the number of h in the past time period T. i The number of successes and failures in the PH campaign are also calculated using the mean of the Beta distribution. The sum of the two and divided by two can be used to get the team leader h. i Objective evaluation results
[0076] For example, in the experiment of the embodiment, at time slot 1, the interactive performance evaluation of the leader of the 5th team in one round of objective evaluation is 0.34, and the team leader's experience calculation is 0.5. The objective evaluation result of the team leader 5 at time slot 1 is obtained by the above calculation method.
[0077] Indicates the vehicle platoon system's response to the platoon leader h at time slot t i The subjective evaluation results are calculated according to the following formula:
[0078]
[0079] In the above formula, Indicates the inter-team evaluation value, which is determined by the PH of other teams in h i The usefulness of the information transmitted in the current time slot t is evaluated, with a score range between 0 and 1. The scores of all PHs participating in the evaluation are summarized and the confidence of these PHs is used to make a weighted average of the scores;
[0080] In the above formula, Indicates the evaluation value within the team, the team head h iInternal members of the team are h i The correctness of the instructions issued in the current time slot t and the quality of the shared security information are evaluated. The correctness index ranges from 0 to 1, where 0 means h i Most of the instructions issued are considered to be wrong, 1 means h i Most issued instructions are considered correct, and the information quality index is scored between 0 and 1. Team members are close to the team leader and have a more comprehensive understanding of the team leader. Therefore, two indicators are set to achieve more accurate internal evaluation. The team score given by each member is obtained by multiplying the correctness index and the information quality index. The scores of all members are aggregated and weighted averaged using each member's confidence.
[0081] For example: In the experiment of the embodiment, at time slot 1, the leader of the 5th team has an inter-team evaluation value of Team evaluation value Solve the above formula to get the subjective evaluation result of team head 5 at time slot 1
[0082] The confidence level of the vehicle is calculated according to the following formula:
[0083]
[0084] In the above formula, represents the confidence of vehicle v in time slot t in the past T time period, λ>0 is a constant, It represents the average trust value of vehicle v in the past T time period; the higher the cumulative trust value of the vehicle, the higher the confidence level.
[0085] For example, in the experiment of the embodiment, at time slot 1, the leader of the 5th team will recalculate the confidence according to the above formula after completing the above evaluation, where the constant λ is 2, the time period T is 12, and the average trust value is Therefore, the confidence level of team head 5 at time slot 1 is obtained
[0086] The specific content of the consensus mechanism implemented by the vehicle platooning system in step (6) includes the following steps:
[0087] (61) Filter out the top N with the highest credibility c Nodes are used as consensus nodes instead of all nodes participating in the consensus, so as to reduce communication costs; the nodes are the edge devices RSU that constitute the blockchain network layer; in the embodiment, N c The value is 6;
[0088] (62) The master node and block size are determined according to the set rules. The master node generates the block and broadcasts the consensus request to other consensus nodes for verification;
[0089] (63) After the consensus node completes the verification, it broadcasts the audit results again to other consensus nodes except the master node. If the consensus node receives more than 2f audit results, it means that the verification is valid and sends a confirmation request to the master node; where f represents the number of Byzantine nodes; in this embodiment, the value of f is 3;
[0090] (64) If the master node receives more than 2f+1 confirmation results, it means that the verification is successful, consensus is reached, and the broadcast block can be put on the chain.
[0091] The reputation value of the node is calculated according to the following formula:
[0092]
[0093] When the vehicle formation system executes the consensus mechanism, the time period is also divided into multiple time slots. In the embodiment, the time slot value is 60 seconds. In the above formula, i (t') represents node I i The reputation value at time slot t', 0<σ<1 is the weight parameter, s' represents the node I i The number of successful block production within a specified time, f' represents the number of nodes I i The number of failures to produce blocks within a specified time, K represents the number of failures to produce blocks to node I within the t' time slot. i The number of fleets uploading data, Represents the trust value of fleet PH at time slot t'.
[0094] For example: In the experiment of the embodiment, the block production performance of consensus node 2 in the past was 0.5, the number of team leaders of all fleets uploading data to this node in time slot 3 was K=3, the trust mean was 0.57, and the weight parameter σ was 0.3. According to the above formula, the reputation value of node 2 in time slot 3 was l2(3)=0.52.
[0095] The specific content of the rules for determining the master node and block size in step (62) is: based on the delay of the consensus process, the throughput and reliability of the blockchain network layer, and according to the dynamic changes of the nodes and network characteristics of the blockchain network layer, the key blockchain parameters, namely the master node and block size, are decided through the reinforcement learning algorithm.
[0096] The reinforcement learning algorithm is a diffusion model-based soft actor-critic algorithm, namely the DMSAC algorithm; the soft actor-critic algorithm, namely the SAC algorithm itself is an algorithm based on the maximum entropy theory, with strong exploration efficiency and stability. The DMSAC algorithm combines the generative artificial intelligence model with the reinforcement learning model and introduces the diffusion model into the SAC algorithm, which can enhance the algorithm's exploration, improve sampling efficiency, model more complex environments, and further optimize the performance of SAC.
[0097] The state space, action space and reward function of the DMSAC algorithm are defined as follows:
[0098] State space s t’ : It consists of the computing resources c of the consensus node, the reputation value l and the data transmission rate R between nodes;
[0099] Action space a t’ : Decision variables include masternode selection and block size;
[0100] Reward function r(s t’ ,a t’ ): defined as follows:
[0101] r(s t’ ,a t’ )=(α1D T +αt2Ω T )log(1+l p (t'))
[0102] In the above formula, α1 and α2 are constant coefficients. In the embodiment, their values are 300 and 0.1 respectively. p (t') represents the master node I p Reputation value;
[0103] T C represents the consensus delay, which is calculated according to the following formula:
[0104] T C =T G +T V
[0105] Where T G Indicates the block generation delay, which is determined by the block size and the computing power of the master node. V Indicates the block verification delay, which mainly includes the audit and communication delay of the consensus process;
[0106] For example: In the experiment of the embodiment, after the algorithm is optimized, the consensus is executed, and the master node and block size are determined according to the blockchain network status, and the consensus delay is calculated according to the above formula. The block generation delay T in a certain consensus is G =0.31s, block verification delay T G =1.63s, the consensus delay is calculated to be the block generation interval T C =1.94s
[0107] Ω T It represents the blockchain throughput and is calculated according to the following formula:
[0108]
[0109] In the above formula, S B represents the block size, χ represents the average transaction data size;
[0110] For example, in the experiment of the embodiment, the block generation interval is substituted into the above formula to obtain the blockchain throughput, where the block size S is B =0.4MB, average transaction data size χ = 200B, T C =1.94s, the throughput Ω is calculated T =1532trx / s, where trx / s represents the number of transactions packaged per second.
[0111] The DMSAC algorithm consists of the following four parts:
[0112] Buffer pool: At each decision time slot t', observe the environment to get the current state s t’ And input it into the model, the model calculates the action probability distribution, and samples the action a t’ And execute, the environment feedback reward value r(s t’ ,a t’ ) and transfer to the next state s t’+1 ; After each round of interaction with the environment, the interaction trajectory data (s t’ ,a t’ ,r(s t’ ,a t’ ),s t’+1 ) is added to the buffer pool and used for model parameter update;
[0113] Policy network: The policy network is used to map the state to the action probability distribution. Unlike the SAC algorithm, the policy network is built based on the diffusion model network rather than a simple multi-layer perceptron. It uses the reverse denoising process of the diffusion model to transform the state s t’ After multiple steps of denoising, and each step of denoising transformation obeys Gaussian distribution, we finally get the action probability distribution, from which we sample action a t’ ;
[0114] Dual Value Network: Similar to SAC, the dual value network is used to evaluate the long-term value of different actions in the current state. It consists of cumulative reward returns and entropy values. The purpose of using a dual network is to reduce errors caused by overestimation.
[0115] Target network: The target network is the corresponding target network of the dual value network. It is used to stabilize the training process and update the parameters using a soft update method.
[0116] The training process of the DMSAC algorithm is as follows:
[0117] In each step of each round of training, observe the blockchain network environment status s t’ , input it into the policy network, the policy network outputs the action probability distribution, from which the action a is sampled t’ , decide the master node and block size of the consensus in this stage, execute the consensus process, calculate the latency, throughput and master node reliability, input the reward function to get the reward value r(s t’ ,a t’ ), the blockchain network transitions to the next state s t’+1 ; This trajectory data (s t’ ,a t’ ,r(s t’ ,a t’ ),s t’+1 ) is added to the buffer pool; after a certain amount of data has accumulated in the buffer pool, random sampling is performed from it to update the parameters of the value network and the policy network respectively, and finally the parameters of the target network are soft-updated; at this point, one step in a round of training is completed. After each round of training steps is iterated, the blockchain network environment state is reset and the next round is entered, and this cycle repeats until the training round is completed.
[0118] The inventors conducted a large number of experiments on the method proposed in the present invention. The experiments used a Python environment and performed simulations on a Linux server to evaluate the performance of the fleet trust model and the consensus optimization algorithm.
[0119] Fleet trust model evaluation: This evaluation conducted the following experiments, including a comparison of trust value changes between normal PH and unqualified PH, and a comparison of trust value changes between normal PM and malicious PM. The fleet trust model of the present invention was compared with two other trust models (TSL and MWSL) in terms of vehicle trust changes and unqualified vehicle recognition rate. In this series of experiments, the trust model of this solution was able to make accurate trust assessments and outperformed the other two trust models.
[0120] For more information about the TSL model, see S. Zhong, J. Chen and Y. R. Yang, "Sprite: a simple, cheat-proof, credit-based system for mobile ad-hoc networks," IEEE INFOCOM 2003. Twenty-second Annual Joint Conference of the IEEE Computer and Communications Societies (IEEE Cat. No. 03CH37428), San Francisco, CA, USA, 2003, pp. 1987-1997, vol. 3, doi: 10.1109 / INFCOM.2003.1209220.
[0121] For more information about the MWSL model, see Z. Ying, M. Ma, Z. Zhao, X. Liu and J. Ma, "A Reputation-Based Leader Election Scheme for Opportunistic Autonomous Vehicle Platoon," in IEEE Transactions on Vehicular Technology, vol. 71, no. 4, pp. 3519-3532, April 2022, doi:10.1109 / TVT.2021.3106297.
[0122] Consensus Optimization Algorithm Evaluation: This evaluation focused on the performance of the DMSAC algorithm. Experiments included a comparison of the cumulative reward values of the DMSAC algorithm with those of other mainstream reinforcement learning algorithms, a comparison of the cumulative reward values of the DMSAC algorithm after removing the master node decision factor or the block size decision factor, a comparison of how the reward values under these algorithms change with node computing resources and average transaction size, and a comparison of how the reliability of the consensus process changes with the proportion of unqualified PHs. These experiments also demonstrated that the consensus optimization algorithm of this solution outperformed other algorithms, achieving an efficient and reliable consensus process.
[0123] The above experimental results prove that the method proposed in the present invention is effective.
Claims
1. A method for achieving efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system, characterized by: The method comprises the following steps: (1) The vehicle platooning system is divided into two parts: the vehicle platooning layer and the blockchain network layer. The vehicle platooning layer is a vehicle platooning communication network composed of multiple platoons on the road. Each platoon consists of a platoon leader (PH) and several platoon members (PMs). The blockchain network layer is the core of the entire vehicle platooning system and is a permissioned blockchain network composed of edge devices (RSUs) with sufficient computing and storage resources. (2) The vehicle platooning system performs vehicle registration. After each vehicle is authenticated by the traffic management agency, a certificate and a key are issued to complete the vehicle registration and join the vehicle platooning system. (3) The vehicle platooning system performs intra-platoon communication. The PH of each platoon is responsible for leading the entire platoon. It needs to periodically communicate with other platoons and edge devices (RSUs) to obtain traffic data and pass real-time safety information and driving strategy instructions to the PMs of the platoon to ensure the safety and consistency of the entire platoon. (4) The vehicle platooning system performs inter-team communication, that is, data sharing is achieved between fleets through the blockchain network layer. The sender of the data first uploads the data to the blockchain network layer. After the data reaches a consensus and is stored, other fleets can access the blockchain network layer and obtain the data shared by the sender after obtaining permission. (5) The vehicle platooning system performs trust value calculation. The vehicle platooning system calculates the trust value of each vehicle in each platoon, especially the PH. The specific process is as follows: the vehicle platooning layer periodically uploads the external evaluation scores between different platoons and the internal evaluation scores of platoon members to the local RSU, which calculates the trust value of all vehicles according to the set trust value calculation method and saves the results to the blockchain network layer; (6) The vehicle platooning system executes a consensus mechanism to achieve data sharing; the blockchain network layer receives data uploaded by all fleets, including shared traffic data and trust value data, and packages the data into blocks, runs the set consensus mechanism, reaches consensus, and completes data storage and sharing.
2. The method for achieving efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system according to claim 1, characterized in that: In step (5), the calculation method of the PH trust value of the team leader is as follows: The time period T is divided into multiple time slots. At time slot t, the trust value of the team leader PH is calculated according to the following formula: In the above formula, The team leader h at time slot t i The trust value, Indicates the team leader h i The original trust value of The team leader h at time slot t i The evaluated trust value; ξ is a weight parameter, which implements the "slow rise and fast fall" update strategy for the trust value of the team leader PH. That is, when the original trust value is less than the evaluated trust value, ξ is small, causing the trust value of the team leader to rise slowly; conversely, ξ is large, causing the trust value of the team leader to fall rapidly, and the larger the gap between the two, the larger ξ. In the above formula, the team leader h at time slot t i The evaluation trust value is calculated according to the following formula: In the above formula, 0<ω<1 is used as the weight parameter; Indicates the vehicle platoon system's response to the platoon leader h at time slot t i The objective evaluation results are not affected by the subjective evaluation results of the current vehicle, and are completely based on historical data. i The behavior in the past time period T is scored; there are two main reference standards, one is h i The interactive performance of h i The former is based on the experience of the team leader in the past time period T. i The number of positive and negative interactions with other vehicles is calculated using the mean of the Beta distribution, which is obtained by counting the number of h in the past time period T. i The number of successes and failures in the PH campaign are also calculated using the mean of the Beta distribution. The sum of the two and divided by two can be used to get the team leader h. i Objective evaluation results Indicates the vehicle platoon system's response to the platoon leader h at time slot t i The subjective evaluation results are calculated according to the following formula: In the above formula, Indicates the inter-team evaluation value, which is determined by the PH of other teams in h i The usefulness of the information transmitted in the current time slot t is evaluated, with a score range between 0 and 1. The scores of all PHs participating in the evaluation are summarized and the confidence of these PHs is used to make a weighted average of the scores; In the above formula, Indicates the evaluation value within the team, the team head h i Internal members of the team are h i The correctness of the instructions issued in the current time slot t and the quality of the shared security information are evaluated. The correctness index ranges from 0 to 1, where 0 means h i Most of the instructions issued are considered to be wrong, 1 means h i Most of the instructions issued are considered correct, and the information quality index score ranges from 0 to 1. Team members are close to the team leader and have a more comprehensive understanding of the team leader. Therefore, two indicators are set to achieve more accurate internal evaluation; the team score given by each member is obtained by multiplying the correctness index and the information quality index. The scores of all members are summarized and the scores are weighted averaged using the confidence of each member.
3. The method for achieving efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system according to claim 2, characterized in that: The confidence level of the vehicle is calculated according to the following formula: In the above formula, represents the confidence of vehicle v in time slot t in the past T time period, λ>0 is a constant, represents the average trust value of vehicle v in the past T time period; The higher the vehicle's cumulative trust value, the higher the confidence level.
4. The method for achieving efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system according to claim 1, characterized in that: The specific content of the consensus mechanism implemented by the vehicle platooning system in step (6) includes the following steps: (61) Filter out the top N with the highest credibility c Instead of having all nodes participate in the consensus, nodes are used as consensus nodes to reduce communication costs. The nodes are the edge devices RSU that constitute the blockchain network layer. (62) The master node and block size are determined according to the set rules. The master node generates the block and broadcasts the consensus request to other consensus nodes for verification; (63) After the consensus node completes the verification, it broadcasts the audit results again to other consensus nodes except the master node. If the consensus node receives more than 2f audit results, it means that the verification is valid and sends a confirmation request to the master node; where f represents the number of Byzantine nodes; (64) If the master node receives more than 2f+1 confirmation results, it means that the verification is successful, consensus is reached, and the broadcast block can be put on the chain.
5. The method for realizing efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system according to claim 4, characterized in that: The reputation value of the node is calculated according to the following formula: In the above formula, l i (t') represents node I i The reputation value at time slot t', 0<σ<1 is the weight parameter, s' represents the node I i The number of successful block production within a specified time, f' represents the number of nodes I i The number of failures to produce blocks within a specified time, K represents the number of failures to produce blocks to node I within the t' time slot. i The number of fleets uploading data, Represents the trust value of fleet PH at time slot t'.
6. The method for achieving efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system according to claim 4, characterized in that: The specific content of the rules for determining the master node and block size in step (62) is: based on the delay of the consensus process, the throughput and reliability of the blockchain network layer, and according to the dynamic changes of the nodes and network characteristics of the blockchain network layer, the key blockchain parameters, namely the master node and block size, are decided through the reinforcement learning algorithm.
7. The method for achieving efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system according to claim 6, characterized in that: The reinforcement learning algorithm is a diffusion model-based soft actor-critic algorithm, namely the DMSAC algorithm; the soft actor-critic algorithm, namely the SAC algorithm itself is an algorithm based on the maximum entropy theory, with strong exploration efficiency and stability. The DMSAC algorithm combines the generative artificial intelligence model with the reinforcement learning model and introduces the diffusion model into the SAC algorithm, which can enhance the algorithm's exploration, improve sampling efficiency, model more complex environments, and further optimize the performance of SAC.
8. The method for achieving efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system according to claim 7, characterized in that: The state space, action space and reward function of the DMSAC algorithm are defined as follows: State space s t’ : It consists of the computing resources c of the consensus node, the reputation value l and the data transmission rate R between nodes; Action space a t’ : Decision variables include masternode selection and block size; Reward function r(s t’ ,a t’ ): defined as follows: r(s t’ ,a t’ )=(α1D T +α2Ω T )log(1+l p (t’)) In the above formula, α1 and α2 are constant coefficients, l p (t') represents the master node I p Reputation value; T C represents the consensus delay, which is calculated according to the following formula: T C =T G +T V Where T G Indicates the block generation delay, which is determined by the block size and the computing power of the master node. V Indicates the block verification delay, which mainly includes the audit and communication delay of the consensus process; Ω T It represents the blockchain throughput and is calculated according to the following formula: In the above formula, S B represents the block size, and X represents the average transaction data size; The DMSAC algorithm consists of the following four parts: composition: Buffer pool: At each decision time slot t', observe the environment to get the current state s t’ And input it into the model, the model calculates the action probability distribution, and samples the action a t’ And execute, the environment feedback reward value r(s t’ ,a t’ ) and transfer to the next state s t’+1 ; After each round of interaction with the environment, the interaction trajectory data (s t’ ,a t’ ,r(s t’ ,a t’ ),s t’+1 ) is added to the buffer pool and used for model parameter update; Policy network: The policy network is used to map the state to the action probability distribution. Unlike the SAC algorithm, the policy network is built based on the diffusion model network rather than a simple multi-layer perceptron. It uses the reverse denoising process of the diffusion model to transform the state s t’ After multiple steps of denoising, and each step of denoising transformation obeys Gaussian distribution, we finally get the action probability distribution, from which we sample action a t’ ; Dual Value Network: Similar to SAC, the dual value network is used to evaluate the long-term value of different actions in the current state. It consists of cumulative reward returns and entropy values. The purpose of using a dual network is to reduce errors caused by overestimation. Target network: The target network is the corresponding target network of the dual value network. It is used to stabilize the training process and update the parameters using a soft update method.
9. The method for realizing efficient and reliable data sharing among multiple vehicle fleets in a vehicle platooning system according to claim 8, characterized in that: The training process of the DMSAC algorithm is as follows: In each step of each round of training, observe the blockchain network environment status s t’ , input it into the policy network, the policy network outputs the action probability distribution, from which the action a is sampled t’ , decide the master node and block size of the consensus in this stage, execute the consensus process, calculate the latency, throughput and master node reliability, input the reward function to get the reward value r(s t’ ,a t’ ), the blockchain network transitions to the next state s t’+1 ; This trajectory data (s t’ ,a t’ ,r(s t’ ,a t’ ),s t’+1 ) is added to the buffer pool; After a certain amount of data has accumulated in the buffer pool, random sampling is performed from it to update the parameters of the value network and the policy network respectively, and finally the parameters of the target network are soft-updated. At this point, one step in a round of training is completed. After each round of training steps is completed, the blockchain network environment state is reset and the next round begins. This cycle repeats until the training round is completed.