Multi-vehicle intelligent driving cooperative training method based on block chain in Internet of Vehicles

By building a cloud-edge-end three-layer network model and a two-layer blockchain architecture in the Internet of Vehicles environment, combining multi-agent reinforcement learning and dynamic reputation evaluation, the problem of delay optimization and data privacy protection in the coordinated training of intelligent driving in the Internet of Vehicles is solved, and efficient and secure intelligent driving model sharing and collaborative training is achieved.

CN120197730APending Publication Date: 2025-06-24CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202510347996.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

In the Internet of Vehicles environment, intelligent driving collaborative training faces challenges such as delay optimization, local model quality affects global model accuracy, and data privacy protection, which limits the wide application of federated learning in intelligent driving scenarios.

Method used

Build a three-layer network model of cloud-edge-end, combined with a two-layer blockchain architecture, use multi-agent reinforcement learning (MARL) algorithm to optimize delay, dynamic reputation evaluation mechanism to screen high-reputation nodes, and achieve safe and efficient model sharing through an asynchronous federated learning framework.

Benefits of technology

It effectively reduces the total delay of the system for intelligent driving collaborative training, improves the performance and convergence stability of the global model, and ensures the security of data privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197730A_ABST
    Figure CN120197730A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-vehicle intelligent driving cooperative training method based on a block chain in the Internet of Vehicles, and belongs to the technical field of mobile communication. Aiming at the problems of data privacy leakage, too high synchronous federated learning time delay, inaccurate node contribution degree evaluation, insufficient single-chain block chain expandability and the like existing in intelligent driving cooperative training in the Internet of Vehicles, an asynchronous federated learning framework with fusion of a cloud-edge-end three-layer architecture and a double-layer block chain is constructed. A local training strategy is optimized through multi-agent reinforcement learning to minimize the total time delay of a system, a dynamic reputation evaluation mechanism based on training interaction timeliness, confirmed site occupancy and model quality contribution is designed to screen high-reputation nodes, and asynchronous model verification and PBFT main chain global aggregation are realized by adopting a DAG block chain. According to the method, on the premise of guaranteeing data privacy, the model sharing efficiency is improved by more than 30%, the training time delay is reduced by 40%, the convergence stability of a global model is effectively enhanced, and the expansibility bottleneck of a traditional architecture is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of mobile communications, and relates to a multi-vehicle intelligent driving collaborative training method based on blockchain in a vehicle-to-everything (V2X) network. Background Art

[0002] With the rapid development of artificial intelligence technology, its application scope in key tasks of the V2X network has been continuously expanding, covering multiple core fields such as environmental perception, path planning, and vehicle control. The technological progress of connected and autonomous vehicles (CAVs) largely depends on the ability to extract driving-related features from massive amounts of data. However, the scenarios experienced by individual vehicles are limited, and their perception systems and computing resources are difficult to cope with complex traffic environments. In addition, the knowledge obtained by CAVs through local training is often limited to specific scenarios, making it difficult to achieve cross-vehicle collaboration and knowledge transfer. It is worth noting that traditional distributed machine learning methods based on raw data sharing have significant privacy leakage risks, which limits their wide application in the V2X network environment.

[0003] Federated learning (FL) is a decentralized machine learning method that enables multiple vehicles to collaborate in developing models, broadening the scope of learning from various driving environments, improving overall performance, while protecting the privacy and security of local vehicle data. Thanks to the rapid development of blockchain technology, CAVs can integrate FL with blockchain technology to achieve secure interaction in the fields of decision-making and collaborative perception. Blockchain, with its characteristics of decentralization, immutability, and traceability, can provide reliable data storage and verification functions for distributed networks without relying on a third-party trust institution, which enables it to be deeply integrated with the FL framework. The blockchain-enabled FL framework promotes the transformation of the vehicle role from a traditional data collection terminal to an intelligent collaborative entity, providing important technical support for building a safe and efficient intelligent transportation system.

[0004] Although FL shows broad application prospects in the field of intelligent driving, it still faces the following key challenges in actual deployment: First, in the synchronous FL mode, each round of global model aggregation requires waiting for all CAVs to complete the upload of local models. However, the vehicle networking environment is highly dynamic and time-sensitive. The communication delays, computational power differences, and network instabilities of CAVs may lead to model training and aggregation failures. Therefore, how to optimize the training and aggregation mechanisms to reduce the time cost and improve system efficiency is an urgent problem to be solved by FL in the intelligent driving scenario. Second, the quality of local models has a significant impact on the accuracy of the global model. The data collected by CAVs and the models trained by them are usually affected by multiple factors such as the external environment, computational performance, and the enthusiasm of nodes to participate in training, which may lead to errors of varying degrees. If low-quality local models are aggregated, the accuracy of the global model will be severely reduced. Therefore, the Road Side Unit (RSU) or Base Station (BS) needs to comprehensively evaluate the reputation values of CAV nodes and selectively aggregate their local models to optimize the performance of the global model. In addition, in the data sharing framework empowered by the Directed Acyclic Graph (DAG) blockchain, the verification probability of the models shared by CAVs is closely related to their accuracy, and this correlation directly affects the time delay of reaching consensus by corresponding nodes in the DAG structure. Therefore, it is necessary to establish a mathematical model to deeply analyze the delay characteristics, so as to provide a theoretical basis for the design of optimization strategies.

[0005] To address the above issues, the present invention first constructs a cloud-edge-end three-layer network model in the vehicle networking scenario. Second, for the delay optimization problem in intelligent driving collaborative training, a delay optimization algorithm based on Multi-Agent Reinforcement Learning (MARL) is proposed to minimize the total system delay. In addition, a dynamic reputation evaluation mechanism is adopted to optimize the selection of training nodes, comprehensively considering multi-dimensional features such as the Interaction Timeliness (IT), Confirmation of Site Occupancy (CSO), and Model Quality Contribution (MQC) of intelligent connected vehicles in the blockchain training interaction, and screening high-reputation vehicles to participate in federated learning, thereby improving the training efficiency. This solution realizes secure and efficient asynchronous model sharing in the vehicle networking scenario, while protecting the data privacy of user vehicles and effectively assisting the intelligent driving decision-making of intelligent connected vehicles. Summary of the Invention

[0006] In view of this, the purpose of the present invention is to provide a multi-vehicle intelligent driving collaborative training method based on blockchain in the vehicle networking, which is used to solve the problem of large-scale intelligent driving collaborative training in the vehicle networking scenario. User vehicles can collect road data information in real time through on-vehicle sensors, and collaboratively train an intelligent driving model for the vehicle networking through this architecture.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] In the first aspect, according to the vehicle networking data sharing scenario and the requirements for vehicle data privacy protection, the present invention provides a multi-vehicle intelligent driving collaborative training solution based on blockchain in the vehicle networking. The execution process of this solution is as follows:

[0009] S1: The task publisher issues the training task through the central cloud layer;

[0010] S2: The RSU needs to initially screen the CAV nodes participating in the federated learning;

[0011] S3: The CAV obtains the model from the RSU, verifies and aggregates the selected model as the model to be trained;

[0012] S4: The CAV uses the optimal training strategy to train the model obtained in S3 to obtain the latest local model;

[0013] S5: The local model in S4 and the aggregated edge model are uploaded to the blockchain network through the RSU or the BS;

[0014] S6: The RSU evaluates and updates the learning reputation of all CAV nodes and conducts node selection;

[0015] S7: All CAVs repeat S2 - S6 until the model converges, and the model training task is completed.

[0016] In the second aspect, in the embodiment of the present invention, in S1, a data sharing solution based on the end-edge-cloud architecture applicable to the vehicle networking scenario is established. This architecture includes a central cloud layer, an edge service layer, and an intelligent terminal layer. The intelligent terminal layer is composed of CAV terminal nodes deployed in the road network. The edge service layer is composed of the BS and the RSU deployed in the road infrastructure. The central cloud layer is composed of a cloud server and a task publisher (such as an institution or an enterprise) responsible for the training task. The edge service layer adopts a double-layer blockchain architecture. The lower layer is composed of a partition blockchain based on an improved DAG consensus, and the upper layer is composed of a main blockchain based on the PBFT consensus.

[0017] Thirdly, in S2 of the embodiments of the present invention, an initial reputation evaluation strategy for intelligent connected vehicles based on historical information is provided. The initial reputation of a node is evaluated by quantifying its historical performance, and a hierarchical strategy is adopted to establish an initial reputation evaluation mechanism. In order to select initial CAV nodes more evenly and comprehensively, a hierarchical strategy is adopted to establish an initial reputation evaluation mechanism, and the nodes are divided into different reputation levels (high reputation, medium reputation, low reputation), so as to avoid the influence of extreme values on node selection.

[0018] Fourthly, in S3 of the embodiments of the present invention, an asynchronous sharing framework based on an improved DAG is proposed. A DAG-based Tangle network is adopted within each partition to achieve asynchronous data sharing. Each transaction in the DAG contains a model shared by vehicles, and this model is locally trained and updated after aggregating the models in all transactions pointed to by this transaction. That is, when each transaction joins the DAG network, it will point to some transactions to be verified (referred to as Tips), and then verify and aggregate their models for local training. This transaction will become a Tip after connecting to the DAG network.

[0019] The CAV obtains the Tips in the current DAG from the RSU, verifies and aggregates the selected models as the models to be trained. Each Tip contains a model shared and uploaded to the blockchain by other vehicles. The Tips selection algorithm is improved so that the probability of each Tip being selected is positively correlated with the accuracy of the model it contains. That is, the higher the model accuracy, the greater the probability that this Tip will be selected, and it can be verified by other transactions more quickly. In the embodiments of the present invention, the vehicle will randomly select 2 of the current Tips according to the accuracy weight, verify and aggregate their models as the models to be trained.

[0020] Fifthly, in S4 of the embodiments of the present invention, a MARL-based delay optimization algorithm is provided. The common goal of CAVs is to complete the intelligent driving training tasks issued by vehicle service providers. At the same time, each CAV also hopes to complete the high-quality training of its local model as quickly as possible, and prompt its model data to be confirmed and consensus in the blockchain system as soon as possible to obtain corresponding incentives. After local training, the CAV uploads the model to the RSU in its area, and then adds it to the DAG in its area to wait for verification. Therefore, the total delay is the sum of the local training, uploading to the RSU, and block consensus confirmation delays;

[0021] Therefore, for the three parts of the delay:

[0022] 1. Local training delay: The local training delay mainly includes model aggregation delay and model training delay. The model aggregation delay is closely related to the allocation of computing resources, and the model training delay is positively correlated with the selection of training strategies. That is, higher local training accuracy means longer model training delay.

[0023] 2. Model upload communication delay: In the Internet of Vehicles environment, data transmission delay is one of the key factors affecting the efficiency of federated learning. In order to quantify the uplink transmission delay, an orthogonal frequency division multiple access scheme is used for modeling. The uplink transmission delay depends on the size of the local model parameters uploaded by the CAV and the uplink transmission rate.

[0024] 3. Block consensus confirmation delay: The training strategy adopted by CAV will directly affect the blockchain consensus confirmation delay. Generally speaking, higher local training accuracy will lead to longer local training delay, but due to the high quality of the model, the speed of blockchain in consensus confirmation will also increase accordingly, thereby shortening the block consensus confirmation delay. On the contrary, although lower local training accuracy can effectively reduce the single training delay, it may lead to insufficient model parameter updates, thereby affecting the model quality and prolonging the consensus confirmation time. In addition, under the distributed architecture of the DAG blockchain, the training decisions of CAV nodes in the same partition are interdependent. This coupling relationship makes the CAV consensus speed not only depend on its local training decisions, but also affected by the synergistic influence of the training strategies of other CAV nodes.

[0025] Based on the above analysis, it can be determined that the relationship between CAVs is a non-completely cooperative relationship. Under this relationship, the common goal of CAVs is to complete the intelligent driving training tasks issued by the vehicle service provider. At the same time, each CAV also hopes to complete the high-quality training of the local model as quickly as possible, and prompt its model data to be confirmed and agreed upon in the blockchain system as soon as possible to obtain corresponding incentives. Therefore, the present invention transforms the delay minimization problem of CAV collaborative training into a partially observable Markov decision process (POMDP) ​​problem suitable for deep reinforcement learning (DRL) solution. Under the MARL framework, each CAV optimizes its own policy parameters to maximize its objective function.

[0026] After determining the optimal local training strategy, CAV will use the model aggregated in S3 as the initial model for this round of training. According to the optimal sharing strategy obtained by MARL, the vehicle will use the local data set for training to obtain an updated model that meets the accuracy requirements. The new model is packaged into a transaction and sent to the RSU. The transaction header includes the hash values ​​of all transactions selected by it, and the transaction body includes the trained model.

[0027] In the sixth aspect, an embodiment of the present invention provides a model uploading and chaining method in S5. CAV uploads the transaction packaged in S4 to the blockchain network through RSU. Specifically, CAV uploads the transaction to the DAG network of this area through RSU, and RSU completes the block header information of the transaction, including version number, timestamp, random number, hash value, etc. At the same time, since each RSU can directly obtain the latest block data from the DAG blockchain, in the edge aggregation stage, the leading RSU first obtains the local model parameters uploaded by the CAV nodes participating in FL training from the DAG blockchain. Subsequently, the leading RSU performs edge aggregation operations based on these local models to generate edge models. After the leading RSUs of each partition complete the edge aggregation within their partitions, the BS jointly maintains the upper-layer PBFT blockchain for recording and saving edge models. Therefore, in the global aggregation stage, the central server located in the intelligent cloud layer can directly obtain the latest edge models of different partitions from the PBFT blockchain, and aggregate the edge models to obtain the global model.

[0028] In the seventh aspect, the embodiment of the present invention provides a blockchain-based smart connected vehicle reputation evaluation and selection strategy in S6, where the RSU evaluates and updates the learning reputation of all CAV nodes and selects nodes. In the federated learning process, after updating the global model, the RSU will evaluate and update the learning reputation of all CAV nodes. The change in the learning reputation of CAV nodes during the training process is jointly determined by the following three key features: timeliness of training interaction, confirmed site occupancy, and model quality contribution.

[0029] 1. Timely training interaction

[0030] This scheme adopts asynchronous federated learning to perform model training, in which the timeliness of training interaction is used to measure the interaction speed of CAV nodes with RSUs during training, reflecting their enthusiasm for participating in tasks.

[0031] 2. Confirmed site share

[0032] According to the status characteristics of the site, the sites in the DAG blockchain can be divided into initial sites, confirmed sites, sites to be confirmed, terminal sites (Tips), and newly arrived sites. Among them, the confirmed site is a site whose cumulative weight reaches the consensus threshold and is verified by the entire network. Its transaction content is permanently recorded in the blockchain and cannot be tampered with. On the DAG blockchain, the confirmed site occupancy rate of CAV is also an important indicator to measure the contribution of CAV in FL.

[0033] 3. Model quality contribution

[0034] During the model training process, after the edge model is updated, the similarity between the local model uploaded by the CAV to the DAG blockchain and the edge model can be used to reflect its contribution degree in the local area during training. The model quality contribution can be calculated using cosine similarity, which measures the similarity between two vectors by measuring the cosine value of the angle between the two vectors.

[0035] To screen and eliminate malicious nodes and unstable nodes, this solution designs a dynamic threshold screening mechanism. Let the reputation value threshold be defined as the average value of the reputation values of all CAVs. If the reputation value of a CAV is less than the reputation value threshold, then the CAV is marked as a "problem node". If the cumulative marking times of a node exceed the tolerance threshold, its reputation is cleared and it is prohibited from participating in subsequent training tasks.

[0036] In the eighth aspect, the embodiment of the present invention provides a method for terminating a data sharing task in S7. The cloud will request the update status of the current DAG from the BS in real time and analyze the performance of the current model. When it believes that the model has reached the expectation or has converged, it will send task termination information to each partition RSU and BS, and the data sharing task ends.

[0037] The beneficial effects of the present invention are as follows: Aiming at the problems of data privacy protection, sharing efficiency, and system performance bottlenecks faced by machine learning in intelligent driving applications, the present invention proposes an innovative solution. By constructing a cloud-edge-end three-layer network model and combining a double-layer blockchain architecture, a decentralized asynchronous federated learning (FL) framework is realized, enabling connected and autonomous vehicles (CAVs) to complete efficient collaborative training without directly sharing raw data. Based on the blockchain topology, this solution proposes a delay optimization algorithm based on multi-agent reinforcement learning (MARL) for the delay optimization problem in collaborative training, and minimizes the total system delay through a distributed decision-making mechanism. In addition, a multi-dimensional feature dynamic reputation evaluation system is designed. By quantitatively analyzing indicators such as the interaction timeliness (IT) of vehicle nodes in training, the confirmation of site occupancy (CSO), and the model quality contribution (MQC), high-reputation nodes are screened to participate in the training process, significantly improving the performance and convergence stability of the global model. On the premise of ensuring data privacy and security, this solution realizes the efficient sharing and collaborative training of models in the vehicle networking scenario, effectively solves the problem of insufficient scalability of the traditional single-chain blockchain architecture, and provides reliable technical support for the intelligent driving decision-making of CAVs.

[0038] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. Brief Description of the Drawings

[0039] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:

[0040] Figure 1 is a double-layer blockchain system model based on edge-cloud-end;

[0041] Figure 2 is a schematic diagram of collaborative training for multi-vehicle intelligent driving empowered by blockchain;

[0042] Figure 3 is a flowchart for implementing a collaborative training solution for multi-vehicle intelligent driving based on blockchain in the vehicle networking. Detailed Embodiments

[0043] The following uses specific specific examples to illustrate the embodiments of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0044] Among them, the accompanying drawings are only used for exemplary illustration, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the accompanying drawings will be omitted, enlarged, or reduced, and do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.

[0045] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, it is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only used for illustrative purposes and cannot be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.

[0046] Figure 1 Fig. shows a possible structural schematic diagram of the multi-vehicle intelligent driving collaborative training system involved in the embodiments of the present invention. As Figure 1 shown, this network considers three layers of networks: the central cloud layer, the edge service layer, and the intelligent terminal layer.

[0047] Intelligent terminal layer: The intelligent terminal layer is composed of CAV terminal nodes deployed in the road network. Let the CAV set be V = {1, 2,..., n,..., N}, where N represents the total number of vehicles in the network. Each CAV is equipped with limited computing and communication resources. In addition, the CAV is also equipped with a multi-modal sensor system that can collect environmental data in real time and sense the operating state information of the vehicle.

[0048] Edge service layer: The edge service layer is composed of BS and RSU deployed in the road infrastructure. Let the RSU set be R = {1, 2,..., m,..., M}, where M represents the total number of RSUs in the edge layer. BS and RSU, as edge computing nodes, are equipped with high-performance edge servers, which can provide low-latency communication services, distributed storage resources, and edge computing capabilities for CAV. A two-layer blockchain architecture integrating the DAG and PBFT mechanisms is proposed in the edge service layer. The lower layer is composed of a distributed DAG blockchain, which is responsible for processing high-frequency model sharing transactions generated by CAV and achieving asynchronous consensus through a parallel verification mechanism; the upper layer is constructed by a PBFT blockchain to build a global consensus layer to perform final consistency confirmation on cross-regional transactions.

[0049] Central cloud layer: The central cloud layer is composed of a cloud server and a task publisher (such as an institution or enterprise) responsible for training tasks. In this architecture, relying on its excellent computing performance and massive storage resources, the cloud service platform provides a secure and stable collaborative storage solution for the entire system, and is also responsible for performing key task scheduling and management functions.

[0050] 1. Initial reputation evaluation strategy for intelligent connected vehicles based on historical information

[0051] In the dynamic topology environment of the vehicle network, due to a large number of new CAV nodes frequently joining the network and participating in network activities, this poses challenges to the collaborative training of intelligent driving models. If a large number of newly added CAV nodes directly participate in the federated learning process without effective screening, it may affect the stability and reliability of the entire system. Therefore, this solution proposes a CAV initial reputation evaluation strategy based on historical information to evaluate its initial reputation by quantifying the historical performance of nodes. In addition, a hierarchical strategy is adopted to establish an initial reputation evaluation mechanism, which is specifically analyzed as follows:

[0052] In the vehicle network environment, let the initial reputation of CAVn be For intelligent driving model learning tasks with similar task types, there is a certain similarity between models. If a CAV node has recently participated in a similar learning task completely, then a higher weight is assigned; if a CAV has never participated in a similar task, then a normal weight is assigned. Therefore, the historical experience factor σ Hi is introduced to quantify the historical performance of CAVs in similar tasks. The value range of σ Hi is [0, 1], and its assignment rule is as follows: σ Hi = 1 means that the vehicle has recently participated in multiple similar tasks completely and the model performance is good. σ Hi = 0 means that the vehicle has never participated in a similar task. 0 < σ Hi < 1 means that the vehicle has participated in a small number of similar tasks, but has not participated completely or the effect is average. Based on this, the initial reputation R n of can be expressed as:

[0053]

[0054] To select initial CAV nodes more evenly and comprehensively, a hierarchical strategy is adopted to establish an initial reputation evaluation mechanism, which divides nodes into different reputation levels (high reputation, medium reputation, low reputation), so as to avoid the influence of extreme values on node selection. For RSUm, first calculate the reputation mean μ of the active CAV list m applying to join this FL and the standard deviation σ m . Finally, initialize the startup reputation of CAVn:

[0055]

[0056] Among them, γ rep is the reputation coefficient, which is used to quantify the reputation value of nodes after classification; α is the adjustment factor, which is used to control the sensitivity of classification. Specifically, if it is necessary to strictly screen high-reputation nodes, α can be increased, and at this time, more nodes will enter the medium-reputation level. If more CAVs need to be rated as high-reputation nodes, then α can be decreased.

[0057] 2. Solving the Local Optimal Training Strategy Based on MARL

[0058] In the vehicle networking environment, CAVs face a dynamically changing network environment during collaborative training. As the number of CAVs increases, the scale of the delay optimization problem will expand rapidly, and there are complex correlations between optimization variables. These factors make it difficult for traditional optimization methods to obtain the global optimal solution. To address the above challenges, this solution uses the MARL method to calculate the optimal local training strategy for CAVs. However, since the relationship between CAVs is non-fully cooperative, the present invention adopts the multi-agent twin-delayed deep deterministic policy gradient (MATD3) architecture of "centralized training - decentralized decision-making". Specifically, in the training phase, all CAVs optimize the policy network and value function network through centralized training and use global information to improve the model performance; in the execution phase, each CAV deploys the trained policy network and makes approximately optimal decisions independently based on local observation information. Therefore, MATD3 is very suitable for the dynamic game scenario when CAVs in the vehicle network only have partial observation states, so as to minimize the long-term task training delay of the system. The state space, action space, and reward function of the intelligent agent CAVn are introduced in detail below.

[0059] (1) State Space

[0060] State space o n It consists of the incomplete observation information of CAVn. The local observation state at decision time slot t is defined as:

[0061]

[0062] Among them, is the set of the latest local training decisions of some CAVs within the partition where CAVn is located, that is, the set of neighbor nodes within the communication range of CAVn The latest training decision of. is the set of global parameters of the DAG blockchain, including the cumulative weight threshold W t ′, the arrival rate λ of the site t etc.

[0063] (2) Action Space

[0064] Action space Corresponds to the set of training strategies that CAVn can select, and its value range is the continuous interval [0,1]. At decision time slot t, the action selected by CAVn is

[0065] (3) Reward Function

[0066] During the DRL training process, the reward function is used to guide the direction of policy optimization. For CAVn, the goal of its reward function is to minimize the sum of the local training delay and the block consensus confirmation delay, so it can be defined as:

[0067]

[0068] Among them, The local training delay mainly includes the model aggregation delay and the model training delay is the confirmation delay for site n . Through MATD3 training, each CAV can autonomously optimize the decision-making strategy under the condition of limited local information, thereby effectively reducing the overall delay of intelligent driving collaborative training.

[0069] After that, for the latest intelligent driving model for which the CAV performs local model training, the RSU will package the local model, add transaction header information, and then broadcast it to the edge network to be uploaded to the local DAG, waiting for verification of other transactions. As new transactions are continuously uploaded to the chain, the DAG chain grows continuously. After the Tips are verified, they gradually become trusted transactions, and the newly added transactions are supplemented as Tips waiting for connection with subsequent newly added transactions. Therefore, the model accuracy in the Tips in the DAG chain will gradually increase to convergence. Vehicles only need to aggregate the models of some (2 in the embodiments of the present invention) Tips in the current DAG chain for local training, without waiting for other vehicles. Therefore, this method is asynchronous, and all vehicles can participate in training at any time and can go offline at any time.

[0070] 3. CAV Nodes Learn Reputation for Evaluation and Update

[0071] As Figure 2 shown, the CAV uploads the transaction to the DAG network of the local area through the RSU. At the same time, since each RSU can directly obtain the latest block data from the DAG blockchain, in the edge aggregation stage, the leading RSU first obtains the local model parameters uploaded by the CAV nodes participating in the FL training from the DAG blockchain. Subsequently, the leading RSU performs edge aggregation operations based on these local models to generate an edge model. After the leading RSUs in each partition complete the edge aggregation within their partitions, since the BS jointly maintains the upper-layer PBFT blockchain for recording and storing the edge models. Therefore, in the global aggregation stage, the central server located in the intelligent cloud layer can directly obtain the latest edge models of different partitions from the PBFT blockchain and aggregate the edge models to obtain the global model.

[0072] During the federated learning process, after the RSU updates the global model, it will evaluate and update the learning reputation of all CAV nodes. The change in the learning reputation of CAV nodes during the training process is jointly determined by the following three key features: training interaction timeliness, confirmed site occupancy rate, and model quality contribution. The specific analysis is as follows:

[0073] (1) Training interaction timeliness

[0074] In this paper, asynchronous federated learning is adopted for model training, where training interaction timeliness is used to measure the interaction speed between CAV nodes and the RSU during the training process, reflecting their enthusiasm for participating in tasks. Let the RSUm evaluate the learning reputation of CAV nodes in its partition for the qth time, then the training interaction timeliness of CAVn is Its general form is:

[0075]

[0076] where τ′ represents the time slot when CAVn last uploaded its local model to the RSU server; τ represents the current global time slot; λ IT is an adjustment factor (usually a positive number), used to control the decay rate of training interaction timeliness. If a stricter penalty for delay is required, a larger λ can be selected IT ; if the tolerance for delay is higher, a smaller λ can be selected IT .

[0077] (2) Confirmed site occupancy rate

[0078] On the DAG blockchain, the confirmed site occupancy rate of CAV is also an important indicator to measure the contribution of CAV in FL. Let the number of confirmed sites of CAVn be when the RSUm evaluates the learning reputation of CAV nodes in its partition for the qth time, then the confirmed site occupancy rate of CAVn can be defined as:

[0079]

[0080] where β is a trend adjustment coefficient (β≥0), used to control the influence intensity of the historical occupancy rate change trend on the current occupancy rate; is the occupancy rate change trend of CAVn. When q = 0, the occupancy rate of all CAV nodes is initialized to When q = 1, the basic occupancy ratio is directly used. When q≥2, the average change rate in the past Q rounds is calculated to smooth the short-term occupancy rate change fluctuations.

[0081] (3) Model quality contribution

[0082] During the model training process, after the edge model is updated, the similarity between the local model uploaded by the CAV to the DAG blockchain and the edge model can reflect its contribution degree in the local area during training. The model quality contribution can be calculated using cosine similarity, which measures the similarity between two vectors by measuring the cosine value of the angle between the two vectors. For CAVn, assuming that the parameters in the models obtained by its local update and global update are regarded as two M-dimensional vectors, when RSUm evaluates the learning reputation of the CAV node in its partition for the qth time, the model quality contribution of CAVn is:

[0083]

[0084] where and ω edge represent the local model parameters of the vehicle node CAVn and the edge model parameters aggregated by RSU respectively. Finally, by weighted fusion of the three indicators, the comprehensive learning reputation value of the CAV node is calculated.

[0085] In order to screen out malicious nodes and unstable nodes, this scheme designs a dynamic threshold screening mechanism. Let the reputation threshold Reputation value threshold is defined as the average value of all CAV reputation values, that is:

[0086]

[0087] When , mark CAVn as a "problem node". If the cumulative marking times of the node exceed the tolerance threshold, clear its reputation and prohibit it from participating in subsequent training tasks.

[0088] 4. System Process

[0089] Figure 3 The figure shows the execution flow chart of the multi-vehicle intelligent driving collaborative training scheme based on blockchain in the vehicle network. The specific steps are as follows:

[0090] Step 401: Task Release

[0091] The task publisher sends the training task to the BS through the central cloud layer. Subsequently, the RSU obtains the task information from the BS. Then, the RSU retrieves the list of active CAVs within its communication coverage area and sends the task information to these CAVs through the broadcast mechanism.

[0092] Step 402: Initial Node Selection

[0093] To ensure the rapid convergence of the federated learning model training, the RSU needs to screen the CAV nodes participating in the federated learning. Initial node selection is first carried out in the task release stage. When the federated learning task is released, the RSU stores the vehicles that respond within the specified time in the active CAV list in the order of reply. Subsequently, the RSU sorts the CAV list according to the initial reputation evaluation strategy of intelligent connected vehicles based on historical information, and selects nodes with higher reputation values to form the initial node set for participating in the federated learning task.

[0094] Step 403: Model Download

[0095] In the architecture based on the DAG blockchain, the local model parameters uploaded by the CAV are distributed and stored in all RSUs in the area. Therefore, when the CAV is ready to perform a new round of model training, it must first obtain the model parameters contained in the latest Tips from the neighboring RSUs.

[0096] Step 404: Local Training

[0097] In the model training stage, the CAV first aggregates the k model parameters downloaded from the DAG blockchain through the average aggregation method to obtain the aggregated model. Subsequently, the CAV calculates the local optimal training strategy through the local optimal training strategy solving algorithm based on MARL. Finally, the CAV uses the local data for training to obtain the latest local model.

[0098] Step 405: Model Upload

[0099] In the model parameter update stage, the CAV nodes participating in the training encapsulate the locally trained model parameters as site nodes in the DAG blockchain, and then transmit them to the associated RSU through the vehicle-to-infrastructure (V2I) communication protocol.

[0100] Step 406: Edge Aggregation

[0101] Each partition elects a leading RSU through the Raft consensus mechanism to be responsible for coordinating the edge aggregation tasks in this area. Since each RSU can directly obtain the latest block data from the DAG blockchain, in the edge aggregation stage, the leading RSU first obtains the local model parameters uploaded by the CAV nodes participating in the FL training from the DAG blockchain. Subsequently, the leading RSU performs edge aggregation operations based on these local models to generate the latest edge model for the partition.

[0102] Step 407: Global Aggregation

[0103] After the leader RSU of each partition completes the edge aggregation within its partition, since the BS jointly maintains the upper-layer PBFT blockchain for recording and storing the edge models. Therefore, in the global aggregation stage, the central server located in the intelligent cloud layer can directly obtain the latest edge models of P different partitions from the PBFT blockchain and aggregate the edge models to obtain the global model.

[0104] Step 408: Evaluate and update learning reputation

[0105] During the federated learning process, after updating the global model, the RSU will evaluate and update the learning reputation of all CAV nodes. Judge the cumulative number of times the CAV node is marked as a "problem node". If the cumulative number of marks exceeds the tolerance threshold, clear its reputation and prohibit it from participating in subsequent training tasks.

[0106] Step 409: Task termination

[0107] The vehicle service provider observes the model training process. When the effect of the intelligent driving model meets the requirements, the vehicle service provider will send a request to stop training, and the task training is completed. At the same time, according to the model training participation recorded in the RSU, rewards and punishments are given to the participants.

[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A multi-vehicle intelligent driving collaborative training method based on blockchain in an Internet of Vehicles, characterized by: The method comprises the following steps: S1: The task publisher issues the training task through the central cloud layer; S2: RSU needs to perform initial screening of CAV nodes participating in federated learning; S3: CAV obtains models from RSU, verifies and aggregates the selected models as the models to be trained; S4: CAV uses the optimal training strategy to train the model obtained in S3 to obtain the latest local model; S5: upload the local model in S4 and the aggregated edge model to the blockchain network through RSU or BS; S6: RSU evaluates and updates the learning reputation of all CAV nodes and performs node selection; S7: All CAVs repeat S2 to S6 until the model converges, completing the model training task.

2. The multi-vehicle intelligent driving collaborative training method based on blockchain in the Internet of Vehicles according to claim 1 is characterized by: In S1, a data sharing solution based on an end-edge-cloud architecture suitable for the Internet of Vehicles scenario is established, and the architecture includes a central cloud layer, an edge service layer, and an intelligent terminal layer; the edge service layer adopts a two-layer blockchain architecture, the lower layer is composed of a blockchain based on a DAG consensus, and the upper layer is composed of a main blockchain based on a PBFT consensus.

3. The multi-vehicle intelligent driving collaborative training method based on blockchain in the Internet of Vehicles according to claim 1 is characterized by: In S2, the initial reputation of the node is evaluated by quantifying its historical performance, and a grading strategy is used to establish an initial reputation evaluation mechanism; in order to evenly select the initial CAV nodes, a grading strategy is used to establish an initial reputation evaluation mechanism, and the nodes are divided into different reputation levels.

4. The multi-vehicle intelligent driving collaborative training method based on blockchain in the Internet of Vehicles according to claim 1 is characterized by: In the S3, a DAG-based asynchronous sharing framework is provided; a DAG-based Tangle network is used in each partition to realize asynchronous data sharing; each transaction in the DAG contains a vehicle sharing model, which is obtained by aggregating the models in all transactions pointed to by this transaction and then locally training and updating.

5. The multi-vehicle intelligent driving collaborative training method based on blockchain in the Internet of Vehicles according to claim 1 is characterized by: In S4, a MARL-based delay optimization algorithm is adopted; the delay minimization problem of CAV collaborative training is transformed into a POMDP problem suitable for DRL solution; under the MARL framework, each CAV optimizes its own strategy parameters to maximize its objective function; after determining the local optimal training strategy, the CAV is trained to obtain an updated model that meets the accuracy requirements; the new model is packaged into a transaction and sent to the RSU.

6. The multi-vehicle intelligent driving collaborative training method based on blockchain in the Internet of Vehicles according to claim 1 is characterized by: In S5, a model upload and chaining method is provided; CAV uploads the transaction to the local DAG network, and RSU supplements the block header information, including version number, timestamp, random number and hash value, etc.; in the edge aggregation stage, the leading RSU obtains the local model parameters uploaded by the CAV nodes participating in FL training from the DAG blockchain, and performs edge aggregation based on these parameters to generate an edge model; After each partition leader RSU completes edge aggregation, the edge model is recorded in the upper-layer PBFT blockchain maintained by the BS; In the global aggregation stage, the central server of the intelligent cloud layer directly obtains the latest edge models of each partition from the PBFT blockchain and aggregates them to generate a global model.

7. The multi-vehicle intelligent driving collaborative training method based on blockchain in the Internet of Vehicles according to claim 1 is characterized by: In S6, a blockchain-based smart connected vehicle reputation evaluation and selection strategy is provided. The RSU evaluates and updates the learning reputation of all CAV nodes and performs node selection. In the federated learning process, the RSU evaluates and updates the learning reputation of all CAV nodes after updating the global model. The change in the learning reputation of the CAV node during the training process is jointly determined by the following three key features: timeliness of training interaction, confirmed site occupancy, and model quality contribution.

8. The multi-vehicle intelligent driving collaborative training method based on blockchain in the Internet of Vehicles according to claim 1 is characterized by: In S7, a training task termination method is provided; the cloud will request the BS for the update of the current DAG in real time and analyze the performance of the current model; when it believes that the model has reached expectations or has converged, it will send task termination information to each partition RSU and BS, and the training task will end.

Citation Information

Cited By

  • Cross-institution nursing data security sharing platform based on block chain and federal learning

    CN120881086A

  • Blockchain and federated learning based cross-institutional nursing data security sharing platform

    CN120881086B

  • Automobile collaborative driving method and system and medium

    CN120949682A

  • Federal learning method of DAG block chain based on main chain consensus

    CN121328774A

  • A federated learning method of a DAG blockchain based on main chain consensus

    CN121328774B