A Trusted Lane-Changing Method for Autonomous Vehicles Based on Blockchain and Federated Reinforcement Learning

Through a blockchain-based federated learning system and near-end strategy optimization algorithm, combined with the reputation mechanism of edge computing nodes and digital currency rewards, data privacy and data island problems in autonomous driving are solved, and the federated learning training effect of autonomous driving vehicles and the performance of lane-changing tasks are improved.

CN115660077BActive Publication Date: 2025-07-25BEIJING UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211224957.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-09
Publication Date
2025-07-25
Estimated Expiration
2042-10-09

AI Technical Summary

Technical Problem

There are data privacy concerns and data silos in existing autonomous driving technologies, which leads to the poor training of federated learning models and the lack of effective reward mechanisms that make local devices reluctant to participate in training.

Method used

The federated learning system based on blockchain is adopted, combining the near-end strategy optimization algorithm and the reputation mechanism of edge computing nodes, and ensuring data privacy and trustworthy model training is achieved through blockchain technology. The roadside edge computing nodes are used to publish training tasks and provide digital currency rewards. The near-end strategy optimization algorithm is used for local training and model aggregation.

Benefits of technology

It has achieved the improvement of the training reliability and efficiency of the federated learning model under the premise of protecting data privacy, encouraged more autonomous vehicles to participate in federal training, and improved the performance of the lane change task of autonomous vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115660077B_ABST
    Figure CN115660077B_ABST
Patent Text Reader

Abstract

The present invention discloses a trustworthy lane-changing method for autonomous vehicles based on blockchain and federated reinforcement learning. The present invention ensures that autonomous vehicles on the road conduct joint training without disclosing local data. The process of model uploading and downloading is protected based on blockchain. In the embodiment of the present invention, in order to ensure the credibility of the training results of local federated reinforcement learning, witness vehicles are introduced to authenticate the results of participating in federated reinforcement learning training. In addition, a task-based evaluation model for autonomous vehicles and a reputation-based evaluation model for edge computing nodes are proposed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of autonomous driving. Background Art

[0002] Autonomous driving, also known as driverless driving, is a cutting-edge technology that relies on computer and artificial intelligence technologies to complete full, safe, and effective driving without human operation. At present, autonomous driving technology has made great breakthroughs in terms of efficiency. However, the widespread use of data-driven models in the technology has raised users' concerns about personal data privacy. At the same time, the "data island" problem also affects the effect of the autonomous driving model to a certain extent.

[0003] In recent years, federated learning has become a solution to the problem of sensitive data in machine learning. Federated machine learning is a machine learning framework that can effectively help institutions use data and perform machine learning modeling while meeting the requirements of user privacy protection, data security, and government regulations. As a distributed machine learning paradigm, federated learning can effectively solve the data island problem, enabling participating parties to jointly model without sharing data, technically breaking the data island, and achieving collaborative learning. However, in federated learning, it is vulnerable to attacks from malicious clients, resulting in unsatisfactory model training effects. At the same time, federated learning lacks a reward mechanism for vehicles participating in training, leading to local devices being reluctant to participate in federated training. Therefore, how to improve the reliability of model training while ensuring the security of user data is an urgent problem to be solved at present. Summary of the Invention

[0004] The purpose of the present invention is to provide a federated learning system based on blockchain to complete lane-changing tasks, so as to overcome the deficiencies of the prior art. The present invention can effectively solve the data privacy problem while enabling the federated learning training to reach ideal performance indicators, as shown in Figure 1 .

[0005] The present invention includes three parts:

[0006] 1. Federated Learning: Federated learning, also known as collaborative learning, can perform large-scale training on devices that generate data, and these sensitive data are retained by the data owners for local collection and local training. After local training, the central training coordinator obtains the training contributions of each node by acquiring updates to the distributed model without accessing the actual sensitive data. In a typical federated learning scenario, the central server sends model parameters to each node (also known as a client, terminal, or worker). The node trains an initial model for local data and sends the newly trained weights back to the central server, which averages the new model parameters (usually related to the amount of training performed on each node). In this case, the central server or other nodes never directly see the data on any other node.

[0007] 2. Proximal Policy Optimization Algorithm: As the underlying model for performing lane-changing tasks, the proximal policy optimization algorithm is a policy gradient method that uses stochastic gradient descent to optimize the agent objective function. The agent objective function is an approximation of the loss function that limits the ratio between the new policy and the old policy to avoid large losses. The proximal policy optimization algorithm has significant advantages compared to other deep reinforcement learning algorithms. First, importance sampling and generalized advantage estimation help the proximal policy optimization algorithm agent effectively utilize sampled data and improve the network update speed. Second, by performing policy clipping between the new and old policies to limit the update range of the new policy, the proximal policy optimization algorithm exhibits better convergence than other deep reinforcement learning algorithms.

[0008] 3. Blockchain: Blockchain is essentially a decentralized database, referring to a technical solution for collectively maintaining a reliable database in a decentralized and trustless manner. Blockchain technology does not rely on a third party and can achieve the storage, verification, transmission, and exchange of network data through its own distributed nodes. It uses cryptography and distributed consensus algorithms to enable participants to reach consensus on the Internet where trust relationships cannot be established without the intervention of any third-party center, solving the problem of reliable transmission of trust and value at extremely low cost. Additionally, blockchain also has the characteristics of decentralization, openness, independence, security, and anonymity. Description of the Drawings

[0009] Figure 1 is the overall system diagram Detailed Implementation Manner

[0010] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0011] S1: All entities within the road, that is, all autonomous vehicles and roadside edge nodes, perform blockchain network initialization registration in the trusted center to become authorized nodes. The trust center issues a public-private key pair and an electronic wallet account to the authorized nodes.

[0012] Among them, the public and private keys are used for communication encryption between the edge computing node and the autonomous vehicle. The electronic wallet account is used for the management of digital currency. All registered devices in the trusted center need to pay a certain amount of digital currency as a deposit to prevent malicious behavior of the devices. The amount of deposit required for different devices is negatively correlated with their reputation values. In addition, each entity becomes a common node after being authorized in step S1. However, due to the memory limitation of the autonomous vehicle, a lightweight design is adopted for the autonomous vehicle node, that is, the autonomous vehicle only retains the block header of the blockchain.

[0013] S2: The roadside edge computing node publishes a federated learning training task according to road traffic needs to become the server side of the federated learning. Autonomous vehicles within the service range of the roadside edge computing node apply to become clients according to their will. The server establishes a training contract with the distributed clients. To encourage autonomous vehicles to become clients, the roadside edge computing node will provide a digital currency (B) as a reward for the clients participating in the training.

[0014] The federated learning training task published by the roadside edge computing node should include: training task objectives, model size, reputation threshold requirements, the number of verification signatures, and the bonus obtained for completing the task. The roadside edge computing node, as the server, selects clients according to the task reputation ranking of the applied autonomous vehicles. The reputation of autonomous vehicles consists of three indicators: enthusiasm, honesty, and task completion quality. Enthusiasm represents the number of times an autonomous vehicle has served as a client and a witness vehicle since registration. Honesty represents the difference between the number of completed tasks and the number of uncompleted tasks. The task completion quality is the score given by the roadside server for the task completion quality of each completed task or the average reward given by the witness vehicle to the working vehicle model for loading the model with its own computing resources. The calculation method of the reputation of autonomous vehicles is as follows: where p n (t) represents enthusiasm, h n (t) represents maturity, q n(t) represents the task completion quality. After the contract is established, any party's breach of contract will result in the deduction of a portion of the margin as a penalty.

[0015] S3: The roadside edge computing node generates the initial model parameters w initial , which includes randomly initializing the action-value network parameters θ initial and the state-value network parameters as well as the hyperparameter learning rate γ = 0.95. It is sent to the participating clients for local training. After the client configures the local model according to the received model parameters, model training is carried out, and the start time of training is recorded as t initial , the aggregation time is t fed , and the current timestamp is t.

[0016] For the underlying control model for executing the lane-changing task, that is, the model issued by the server and locally trained, the proximal policy optimization algorithm should be adopted. The proximal policy optimization algorithm is a policy-based deep reinforcement learning algorithm using two neural networks. In the proximal policy optimization algorithm, the two neural networks are the action-value network and the state-value network. The algorithm inputs the agent's state S n into the action-value network. After calculating through the action-value network parameters, the action A in the current state is obtained n ; the agent's state S n is input into the state-value network. After calculating through the state-value network parameters, the score of the network for the current state is obtained, that is, by inputting the current "state" of the "agent" into the neural network, corresponding "actions" and "rewards" will ultimately be obtained. Then, according to the "action", the state of the "agent" is updated, and according to the objective function containing "rewards" and "actions", gradient ascent is used to update the weight parameters in the neural network, so as to obtain a "action" judgment that can make the overall reward value larger. The proximal policy optimization algorithm can leverage the PPO module provided by the Stabl e -Baselin e 3 reinforcement learning framework.

[0017] In the lane-changing task, the key model parameters are set as:

[0018] (1) State space: Represents the agent's observation of the surrounding environment. After the local training starts, the values in the state space are used as the input of the algorithm. In this patent, the state space is set as:

[0019] S n = [s ego , s f , s b , s l , s r , s lf , srf , s lb , s rb

[0020] where s ego is the state of the vehicle being trained, s f is the state of the vehicle directly in front of the vehicle being trained, s b is the state of the vehicle directly behind the vehicle being trained, s r is the training state of the vehicle on the right side of the vehicle being trained, s l is the training state of the vehicle on the left side of the vehicle being trained, s lf is the training state of the vehicle in the front left of the vehicle being trained, s rf is the training state of the vehicle in the front right of the vehicle being trained, s lb is the training state of the vehicle in the rear left of the vehicle being trained, s rb is the training state of the vehicle in the rear left of the vehicle being trained. The state of each vehicle includes:

[0021] s i = [(x, y), v Lon , v Lat

[0022] (x, y) represents the longitude and latitude coordinates of vehicle i, v x is the lateral speed of vehicle i, v y is the longitudinal speed of vehicle i.

[0023] (2) Action space: In the proximal policy optimization algorithm model, the action is the policy obtained by the model based on the acquired state. Generally, the output of the action neural network is directly used as the action. The action space includes:

[0024] A n = [-1, 0, 1]

[0025] where -1 means the vehicle being trained changes lanes to the left, 0 means the vehicle being trained keeps the original lane, and 1 means the vehicle being trained changes lanes to the right.

[0026] (3) Reward function: When the action is completed, the environment returns a reward to the agent, and the proximal policy optimization algorithm model is optimized by means of the gradient ascent algorithm. The reward function is calculated as:

[0027] r n = w s r n,s + w c r n,c + w k r n,k

[0028] r n,s is the incentive reward for speed and can be calculated as v max and v​​min is the difference between the maximum and minimum speeds that the vehicle being trained can reach. r n,c is an incentive for safety. When a collision occurs during training r n,c = -1, otherwise r n,c = 0. r n,k is the rule incentive. On a multi-lane highway, we hope to vacate the leftmost lane for overtaking vehicles to use for overtaking. Therefore, the vehicle being trained should occupy the leftmost lane as little as possible except when overtaking. When the vehicle being trained is driving in the right lane r n,k = 0.05, otherwise r n,k = 0. w s is the speed reward weight value, w c is the collision reward weight value, w k is the lane keeping reward weight value. In a multi-lane highway scenario, w s 、w c 、w k can take 1, 0.95, 1. Among them r max = 1.

[0029] For the action value network, a neural network with 3 hidden layers is adopted. The number of nodes in the hidden layer neural network is 256, 512, 256 respectively. The input is the state of the agent, and the output is the action of the agent. The state value network adopts a neural network with 3 hidden layers. The number of nodes in the hidden layer neural network is 256, 512, 256 respectively. The input is the state of the agent, and the output is the value of the agent in this state.

[0030] The PPO algorithm can refer to the paper: [1]Schulman J, Wolski F, Dhariwal P, et al. Proximal Policy Optimization Algorithms[J]. 2017.

[0031] S4: The client invites 5 autonomous vehicles within the communication range to be witness vehicles to verify the training results of the client. To encourage the surrounding vehicles to participate in the authentication, the client distributes a part of the task reward b (b < B) as the authentication reward.

[0032] S5: When the number of training rounds e of the client reaches 5000 times, the client first sends the training results to the witness vehicles for work authentication. After the witness vehicles authenticate, they send the authentication results and signatures to the client. When the training results of the client pass the authentication of all witness vehicles. The client uploads the model with the signatures of each witness vehicle to the server.

[0033] S6: When the aggregation time arrives, the server aggregates the model parameters returned by the clients using the federated averaging aggregation algorithm to obtain the aggregation parameters for this round; when the average reward r n of the aggregation model for this round does not meet the performance condition of the server, i.e., r n < 0.9r max , the aggregation model for this round is used as the initial model for the next round, and the parameters of the next-round initial model are sent to the clients for local training in the next round until the model meets the server performance requirements;

[0034] The method for determining whether the aggregation time has arrived is: t - t initial = t fed , and when the equation holds, the above aggregation algorithm starts to calculate the aggregation model parameters. The model determines whether to accept the aggregation model according to . r fed is the reward obtained by the federated aggregation model in the environment, and r expect = 0.9r max is the performance condition that the server expects to achieve.

[0035] The federated averaging aggregation algorithm is summarized as: The federated averaging aggregation algorithm is summarized as:

[0036]

[0037] where w fed is the model parameter obtained by the aggregation algorithm, including the action-value network parameter θ fed and the network value parameter which can be sent to the clients as the initial model for the next round. N is the number of models collected within the aggregation time interval, and w n are the model parameters uploaded by each client. R n (t) is the reputation of the work vehicle at this time. To reduce the negative impact of clients with poor computing power on the aggregation efficiency, an asynchronous aggregation method is adopted. Each client can upload after completing the local training a certain number of times. Clients that have not completed the local training when the aggregation time arrives will not be able to upload their models. For pre-training needs, a highly simulated road environment is deployed in the edge roadside nodes. Generally, Carla or Sumo can be used as the software for the model pre-training simulation environment of the roadside edge computing nodes.

[0038] S7: In the blockchain network, consensus nodes are selected by voting based on the reputation of the roadside edge computing nodes. Due to the computing and memory limitations of autonomous vehicles, consensus nodes are only selected from the roadside edge computing nodes. The roadside edge computing nodes vote to select consensus nodes, and the top 10 nodes in terms of the number of votes can be selected as consensus nodes. The number of votes held by each node is positively correlated with its reputation. The reputation of the roadside edge computing node j is calculated as:

[0039] R j (t) = (1 + e -F ) -1

[0040] F=F h -F m

[0041] Where F represents the positive contribution of the edge computing node, which is the difference between the node’s honest contribution and malicious contribution. h represents the honest behavior contribution of the roadside edge computing node, that is, the sum of the number of completed federated training and computing resource consumption; F m It represents the malicious behavior contribution of the roadside edge computing node, which is the sum of the number of incomplete federated training and the number of node downtime.

[0042] Each consensus node packs all the contract contents that occur in the network into blocks. The consensus node that first completes the block production broadcasts its consensus proposal to all consensus nodes for certification. After more than half of the consensus nodes verify and agree with the block proposal, the block is written into the blockchain and the new blockchain is broadcast to all nodes.

[0043] The consensus mechanism uses the PoW mechanism, that is, the proof-of-work mechanism, in which the distribution of currency and the determination of bookkeeping rights are carried out according to the workload of the consensus node. The winner of the computing power competition will obtain the corresponding block bookkeeping rights. Therefore, the higher the computing power of the consensus node, the more likely it is to obtain the bookkeeping rights, and computing power requires a huge investment. The reasons for choosing the PoW mechanism are: the solution is simple and easy to implement; nodes can reach a consensus without exchanging additional information; malicious nodes need to invest a huge cost to destroy the system.

[0044] After the consensus process is completed, the roadside edge computing node will read the data on the blockchain and issue rewards to all clients participating in the training after calculation. After the client updates the lightweight blockchain, it uses the block header to cache the data of the latest block and distributes rewards to the witness car based on the contract in the block. Finally, each device recalculates the reputation and writes it to the blockchain at the next consensus.

Claims

1. A trusted lane-changing method for autonomous vehicles based on blockchain and federated learning, characterized in that, include: S1: All entities on the road, i.e. all autonomous vehicles and roadside edge computing nodes, initialize the blockchain network and register as authorized nodes in the trusted center; The trust center issues public and private key pairs and electronic wallet accounts to authorized nodes; S2: The roadside edge computing node publishes federated learning training tasks according to road traffic needs and becomes the federated learning server. The autonomous driving vehicles within the service range of the roadside edge computing node apply to become clients according to their wishes. The server establishes a training contract with the distributed client. To motivate the autonomous vehicle to become a client, the roadside edge computing node will provide a sum of digital currency B as a reward to the client participating in the training. S3: The roadside edge computing node generates the initial model parameters w initial , which includes randomly initializing the action-value network parameters θ initial and the action-value network parameters as well as the hyperparameter learning rate γ = 0.95; Send them to the clients participating in the training for local training. After the clients configure the local model according to the received model parameters, they perform model training. Record the start time of the training as t initial , the aggregation time as t fed , and the current timestamp as t; S4: The client invites five autonomous driving vehicles within the communication range to become witness vehicles to verify the client's training results. In order to encourage surrounding vehicles to participate in the certification, the client allocates a part of the task reward b as the certification reward; where b<B; S5: When the number of training rounds e of the client reaches more than 5000 times, the client first sends the training results to the witness car for work authentication. After authentication, the witness car sends the authentication results and signature to the client. When the client training results are authenticated by all witness cars; The client uploads the model with the signatures of each witness car to the server; S6: When the aggregation time arrives, the server aggregates the model parameters returned by the clients it has received using the federated averaging aggregation algorithm to obtain the aggregation parameters for this round; when the average reward r n of the aggregated model for this round does not meet the performance condition of the server, that is, r n < 0.9r max at this time, the aggregated model for this round is used as the initial model for the next round, and the initial model parameters for the next round are sent to the clients to perform local training for the next round through the clients until the model meets the server performance requirements; r max is the best performance condition of the server. S7: In the blockchain network, consensus nodes are selected by voting among the roadside edge computing nodes; each consensus node packs all contract contents occurring in the network into blocks; the consensus node that first completes block production broadcasts its consensus proposal to all consensus nodes for certification. After more than half of the consensus nodes verify and agree with the block proposal, the block is written into the blockchain and the new blockchain is broadcast to all nodes.

2. For the underlying control model for performing the lane change task, that is, the model issued by the server and locally trained, the proximal policy optimization algorithm should be adopted according to the method described in claim 1; Proximal Policy Optimization algorithm is a policy-based deep reinforcement learning algorithm that uses two neural networks. In the Proximal Policy Optimization algorithm, the two neural networks are the action-value network and the state-value network. The algorithm takes the agent's state S n as the input to the action-value network. After calculating the parameters of the action-value network, the action A in the current state is obtained n ; The agent's state S n is input into the state-value network. After calculating the parameters of the state-value network, the score of the network for the current state is obtained. That is, by inputting the current "state" of the "agent" into the neural network, the corresponding "action" and "reward" will be obtained. Then, according to the "action", the state of the "agent" is updated. According to the objective function containing "reward" and "action", gradient ascent is used to update the weight parameters in the neural network, so as to obtain the "action" judgment that makes the overall reward value larger; The Proximal Policy Optimization algorithm uses the PPO module provided by the Stable-Base1ine3 reinforcement learning framework; In the lane-changing task, the key model parameters are set as: (1) State space: represents the agent’s observation of the surrounding environment. After local training begins, the values in the state space are used as the input of the algorithm. The state space is set to: S n = [s ego , s f , s b , s l , s r , s lf , s rf , s lb , s rb ​ where s ego is the state of the vehicle being trained, s f is the state of the vehicle directly in front of the vehicle being trained, s b is the state of the vehicle directly behind the vehicle being trained, s r is the training state of the vehicle on the right side of the vehicle being trained, s l is the training state of the vehicle on the left side of the vehicle being trained, s lf is the training state of the vehicle in the front left of the vehicle being trained, s rf is the training state of the vehicle in the front right of the vehicle being trained, s lb is the training state of the vehicle in the rear left of the vehicle being trained, s rb is the training state of the vehicle in the rear left of the vehicle being trained; where the state s of the vehicle at each of the above positions i includes: s i = [(x, y), v Lon , v Lat ​ (x, y) represents the longitude and latitude coordinates of vehicle i, and v x is the lateral speed of vehicle i, and v y is the longitudinal speed of vehicle i; (2) Action space: In the proximal policy optimization algorithm model, the action is the strategy derived by the model based on the obtained state. Generally, the output of the action neural network is directly used as the action; the action space includes: A n =[-1,0,1] Among them, -1 means that the trained vehicle changes lanes to the left, 0 means that the trained vehicle stays in the original lane, and 1 means that the trained vehicle changes lanes to the right; (3) Reward function: When the action is completed, the environment returns a reward to the agent, and the proximal policy optimization algorithm model is optimized with the help of the gradient ascent algorithm; the reward function is calculated as: r n = w s r n,s + w c r n,c + w k r n,k r n,s is the incentive reward for speed, v max and v min is the difference between the maximum and minimum speeds that the vehicle under training can reach; r c is the incentive reward for safety. During training, when a collision occurs, r n,c = -1, otherwise r n,c = 0; when the vehicle under training is driving in the right lane, r n,k = 0.05, otherwise r n,k = 0; w s is the speed reward weight value, w c is the collision reward weight value, w k is the lane-keeping reward weight value; in the highway scenario with one-way three-lane or four-lane, w s 、w c 、w k take 1, 0.95, 1; where r max = 1; For the action value network, a neural network with 3 hidden layers is used. The number of nodes in the hidden layer neural network is 256, 512, and 256 respectively. The input is the state of the intelligent agent, and the output is the action of the intelligent agent. For the state value network, a neural network with 3 hidden layers is used. The number of nodes in the hidden layer neural network is 256, 512, and 256 respectively. The input is the state of the intelligent agent, and the output is the value of the intelligent agent in that state.

3. The method according to claim 1, wherein The roadside edge computing node, acting as a server, selects clients based on the task prestige ranking of the applied-for autonomous vehicles. The prestige of an autonomous vehicle consists of three indicators: enthusiasm, honesty, and task completion quality. Enthusiasm represents the number of times an autonomous vehicle has served as a client and a witness vehicle since self-registration. Honesty represents the difference between the number of completed tasks and the number of uncompleted tasks. Task completion quality is the score given by the witness vehicle to the working vehicle model's reward, that is, the average reward of the working vehicle model calculated by the witness vehicle loading the model with its own computing resources. The calculation method for the prestige of autonomous vehicle n is as follows: where p n (t) represents enthusiasm, h n (t) represents maturity, q n (t) represents task completion quality.

4. The method according to claim 1, wherein Each entity becomes a common node after being authorized in step S1; however, due to the memory limitation of the autonomous vehicle, a lightweight design is adopted for the autonomous vehicle node, that is, the autonomous vehicle only retains the block header of the blockchain.

5. The method according to claim 1, wherein The federated average aggregation algorithm can be summarized as follows: Among them, w fed is the model parameter obtained by the aggregation algorithm, including the action-value network parameter θ fed and the network value parameter which is sent to the client as the initial model for the next round; N is the number of models collected within the aggregation time interval, and w n is the model parameter uploaded by each client; R n (t) is the reputation of the work vehicle at this time; in order to reduce the negative impact of clients with poor computing power on the aggregation efficiency, an asynchronous aggregation method is adopted. Each client can upload after completing the number of local training times. Clients that have not completed by the aggregation time will not be able to upload their own models.

6. The method according to claim 1, characterized in that The judgment method for whether the aggregation time is reached in S6 is: t - t initial = t fed , when the formula holds, the above aggregation algorithm starts to calculate the aggregation model parameters; the model calculates whether to accept the aggregation model according to r, where r is the reward obtained by the federated aggregation model in the environment, r expect is the performance condition expected to be achieved by the server, r expect = 0.9r max .

7. The method according to claim 1, characterized in that The method for selecting consensus nodes in S7 is as follows: Due to the computing and memory limitations of autonomous vehicles, the consensus nodes are only selected from roadside edge computing nodes; the roadside edge computing nodes vote to select consensus nodes, and the top 10 nodes in terms of the number of votes can be selected as consensus nodes; the number of votes held by each node is positively correlated with its reputation and prestige, and the reputation and prestige of roadside edge computing node j are calculated as: R j (t) = (1 + e -F ) -1 F = F h -F m Among them, F represents the positive contribution of the edge computing node, which is the difference between the node's honest contribution and malicious behavior contribution; F h represents the honest behavior contribution of the roadside edge computing node, that is, the sum of the number of times of completing federated training and the consumption of computing resources; F m represents the malicious behavior contribution of the roadside edge computing node, that is, the sum of the number of times of incomplete federated training and the number of node outages.

Citation Information

Patent Citations

  • Automatic driving training method and device based on federated learning and medium

    CN112052959A

  • Automatic driving data processing method and system conforming to data privacy security

    CN114861220A