Asynchronous federal near-end strategy optimization reinforcement learning method based on block chain

Through the blockchain-based asynchronous federated proximity strategy optimization reinforcement learning method, combined with dual-gated recurrent unit network and asynchronous federated learning, the problem that autonomous driving systems are difficult to accurately predict the motion trajectory of vehicles and pedestrians in dynamic traffic environments is solved, and efficient trajectory prediction and user privacy protection are achieved.

CN120144948APending Publication Date: 2025-06-13ZHENGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510193218.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-20
Filing Date
2025-02-21
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

It is difficult for autonomous driving systems to accurately predict the movement trajectories of vehicles and pedestrians in dynamic traffic environments, and the privacy protection problems and network security risks brought about by data sharing are difficult to solve.

Method used

The blockchain-based asynchronous federated near-end strategy optimization reinforcement learning (BE-AFRL) method is adopted to extract historical and future trajectory features through a dual-gated recurrent unit network, and combine asynchronous federated learning and curiosity-driven near-end strategy optimization algorithm to achieve real-time updates and optimizations. Blockchain technology ensures data transparency, immutability and security through dynamic grouping.

Benefits of technology

It improves the accuracy and robustness of trajectory prediction of autonomous driving vehicles, ensures user privacy protection, reduces the risk of cyber attacks, and improves system security and data integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144948A_ABST
    Figure CN120144948A_ABST
Patent Text Reader

Abstract

The invention discloses an asynchronous federal near-end strategy optimization reinforcement learning method based on a block chain, and the method is used for the track prediction of an automatic driving vehicle, and employs a double-gating circulation unit network to extract the historical and future track features of a target vehicle. Through asynchronous federated learning, the model can perform joint training among a plurality of participants without directly sharing data, so that real-time updating and optimization are realized, and the adaptability to new data and scenes is enhanced. In addition, a curiosity-driven near-end strategy optimization algorithm is designed, an intelligent agent is stimulated to actively explore a state space, the exploration efficiency is improved, local optimum is avoided, and then the prediction accuracy is improved. And finally, the block chain module adopts a dynamic grouping practical Byzantine fault-tolerant consensus algorithm to enhance the fault-tolerant capability of the block chain network and ensure efficient processing of the trajectory data from the roadside unit. On the whole, the method significantly improves the accuracy and robustness of vehicle trajectory prediction in multiple aspects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent transportation systems, and particularly relates to a blockchain-based asynchronous federated proximal policy optimization reinforcement learning (BE-AFPPO) method for predicting the trajectories of autonomous vehicles. Background Art

[0002] With the rapid development of intelligent transportation systems (ITS), the transportation field is moving towards a more efficient, safe, and convenient direction. As a core technology in ITS, autonomous driving relies on artificial intelligence, sensor technology, and communication networks to achieve autonomous perception, decision-making, and control of vehicles. This technology not only improves road safety, reduces human errors, but also significantly enhances traffic mobility and efficiency, effectively alleviating urban traffic congestion. However, the effectiveness of autonomous driving systems largely depends on their accurate perception and prediction of the surrounding environment, especially the ability to predict the trajectories of other traffic participants. In a dynamic traffic environment, accurately predicting the movement trajectories of vehicles and pedestrians is crucial for ensuring driving safety and optimizing driving routes.

[0003] With the rapid progress of autonomous driving technology, the demand for data by the system has also increased significantly. Sensitive information such as location, speed, and driving status is frequently exchanged between vehicles. Although this data sharing is crucial for improving driving safety and optimizing traffic flow, it also raises serious privacy protection issues. The leakage of personal location information may lead to the infringement of user privacy and even be used for tracking and surveillance. In addition, data is vulnerable to cyberattacks during transmission and storage, facing the risk of being maliciously obtained. With the increasingly strict legal regulations on data privacy, enterprises need to bear compliance risks when processing and sharing data. Therefore, how to effectively protect user privacy while ensuring the performance of autonomous driving systems has become a key challenge that must be addressed in the development of this technology. Summary of the Invention

[0004] To address some of the challenges in predicting the trajectories of autonomous vehicles mentioned above and solve the problem of protecting user privacy while ensuring the performance of autonomous driving systems, this application proposes an innovative blockchain-based asynchronous joint proximal policy optimization reinforcement learning (BE-AFRL) method.

[0005] The blockchain-based asynchronous federated proximal policy optimization reinforcement learning method includes the following steps:

[0006] Step 1: Real-time collect historical trajectory data of the target vehicle and its surrounding traffic participants through sensors to generate corresponding node features; then use statistical methods to remove outliers;

[0007] Step 2: On each participant, locally train the collected historical trajectory data using a dual-gated recurrent unit network;

[0008] Step 3: Adopt the method of asynchronous federated learning to aggregate the updated parameters generated by all participants in local training to generate a new global model;

[0009] Step 4: C-PPO optimizes the policy network to maximize the expected return by restricting the update amplitude of the policy and using the parameters obtained from the global model;

[0010] Step 5: Generate transactions containing updated parameters and hidden states and submit them to the blockchain network to ensure data transparency and immutability; in this process, use the dynamic grouping practical Byzantine fault tolerance consensus algorithm to verify the submitted transactions to ensure the validity and consistency of each transaction;

[0011] Step 6: Distribute the updated global model to all participants, and the participants perform local training according to the received global model to ensure continuous optimization of the model.

[0012] In Step 2, the dual-gated recurrent unit network includes a historical GRU and a future GRU, which are respectively used to extract historical motion features and future trajectory features. The specific calculation formulas are as follows:

[0013] (1) GRU encoding output of historical trajectory:

[0014]

[0015] (2) GRU encoding output of future trajectory:

[0016]

[0017] Among them, represents the GRU encoding output of the historical trajectory T aim of the target vehicle, represents the GRU encoding output of the future trajectory of the target vehicle;

[0018] (3) Combine the final hidden states of the two GRUs together as the final output result. Therefore, the output result of the Bi-GRU network can be expressed as:

[0019]

[0020] In Step 3, in asynchronous federated learning, the clients perform local updates at different time points, and the central server immediately aggregates any received updates; for each client i, the extracted features and updated parameters Δθ i are sent to the federated server;

[0021] (1) The model update of client i is expressed as:

[0022]

[0023] Among them, Δζ i is the model parameter of client i, and T represents the training process;

[0024] (2) The AFL aggregation server performs weighted aggregation on the received local model updates; due to asynchrony, the weight Δζ of each client i is related to its submission time point and participation frequency. The global model parameter update formula is:

[0025]

[0026] Among them, ζ n is the parameter of the global model at time n, w i is the weight of client i, η is the learning rate, and the aggregated global model ζ n+1 is sent to each client for the next round of training;

[0027] (3) The weight w i is calculated by the formula as follows:

[0028]

[0029] Among them, f i represents the participation frequency of client i, t i is the latest submission time point of client i, and α is the adjustment factor; the aggregated global model parameter ζ n+1 is sent to each client for the next round of local training.

[0030] In step 4, after completing the aggregation step of federated learning, a global model integrating multiple local models is obtained. The PPO algorithm is used to adapt the policy network to the complexity and dynamics of the environment using the global model parameters obtained from federated learning:

[0031] (1) First, set the global model parameter ζ n+1 to initialize the parameters of the policy network. Its optimization goal is to maximize the expected return of the policy, expressed as:

[0032]

[0033] Among them, is the advantage function, which describes the advantage of the current action l relative to the old policy. The parameters ψ of the initialized policy network can be expressed as ψ 0 = ζ n+1 ;

[0034] (2) Select the optimal server model parameters by optimizing the policy network, defined as follows:

[0035]

[0036] Among them, ε is the clipping threshold set to 0.1, clip(r n (ψ), 1 - ∈, 1 + ∈) is the probability ratio, which restricts the update amplitude to prevent the policy from changing too much in each iteration. Its calculation method is:

[0037]

[0038] Among them, π ψ (l n ∣q n ) represents the probability that the old policy takes action l n under state q n , while π ψ (l n+1 ∣q n+1 ) represents the probability that the new policy takes action l n+1 under state q n+1 ;

[0039] (3) For the loss of the value function, it is specifically defined as:

[0040]

[0041] Among them, V ψ (q n ) is the value function in the policy network, is the target value function calculated through the advantage function and the reward; the final total loss function is:

[0042]

[0043] (4) Use the optimization algorithm Adam to minimize the loss function to update the policy parameters. After each training iteration, update the policy parameters by calculating the gradient:

[0044]

[0045] Among them, α is the learning rate, ψ n+1 is the updated policy parameter, is the gradient of the PPO policy loss function.

[0046] In step 5 mentioned above, the steps of the dynamic grouping practical Byzantine fault tolerance consensus algorithm are divided into two stages: node grouping and consensus communication:

[0047] First, adopt a grouping method to organize nodes into multiple broadcast domains for facilitating consensus. In the blockchain, the set of nodes is Select G nodes from all nodes as the initial central nodes They are indexed by a positive integer n g (g ∈ [1, G]), and the remaining nodes are denoted as O = {n 1 , n 2 ,..., n O} and are indexed by a positive integer n o (o ∈ [1, N - G]);

[0048] 1) Grouping process: To group all nodes, first calculate the Euclidean distance between all nodes as follows:

[0049]

[0050] Then retain the nearest Euclidean distance from each node n o to the current central node n g . The retained nodes n o will be divided into C g groups, specifically represented as:

[0051]

[0052] 2) Central node optimization: When the sum of the Euclidean distances from a node n g to other nodes n o in the group is the smallest, this node is regarded as the central node of the group. To find the optimal set of central nodes replace the original central node n g with a non - central node n g from C o . When the sum of its distances to n' g is the smallest, the update rule for the central node is as follows:

[0053]

[0054] 3) Stability iteration: Steps 1) and 2) will be repeated until the set of central nodes is consistent with ; if a certain central node n g has a problem, then the central node will be re - selected and regrouped;

[0055] After node grouping, the consortium blockchain nodes are divided into G broadcast domains, and the consensus communication process is divided into four stages: block announcement, intra - group consensus, inter - group consensus, and block synchronization:

[0056] 1) Block announcement: From the set of central nodes The selected primary node n * Pack a section of client transaction requests into a block and encapsulate them as announcement messages:

[0057]

[0058] Then announce it to other aggregation node groups;

[0059] 2) Intra-group consensus: After receiving the announcement message, the aggregation nodes in each group perform PBFT consensus within the group. After the intra-group consensus is completed, a confirmation message is sent:

[0060]

[0061] To the primary node n * , and then enter the inter-group consensus stage;

[0062] 3) Inter-group consensus: The reputation values of the aggregation nodes selected through intra-group consensus are regarded as safe and reliable; when the primary node n * Receives confirmation messages from all aggregation nodes The primary node will send a repeated confirmation message:

[0063]

[0064] After receiving the repeated confirmation message, other aggregation nodes send confirmation messages to the primary node n * again:

[0065]

[0066] Then enter the block synchronization stage;

[0067] 4) Block synchronization: After both the intra-group and inter-group consensus processes are completed, the aggregation node set Sends an execution message:

[0068]

[0069] Notifies the access node set To perform local block writing, and all aggregation nodes Send a response message to the client:

[0070] <Response,v,ts,c,n,n * ,e>

[0071] The consensus process is completed.

[0072] The present invention has made significant improvements in traffic trajectory prediction in multiple key modules. First, the method uses a Bidirectional Gated Recurrent Unit network (Bi-GRU) to extract the historical and future trajectory features of the target vehicle. Through asynchronous federated learning, the model can be jointly trained among multiple participants without directly sharing data, enabling real-time updates and optimizations, and enhancing the adaptability to new data and scenarios. In addition, a Curiosity-driven Proximal Policy Optimization (C-PPO) algorithm is designed to encourage the agent to actively explore the state space, improve the exploration efficiency, avoid local optima, and thus enhance the prediction accuracy. Finally, the blockchain module adopts a Dynamic Group Practical Byzantine Fault Tolerance (DG-PBFT) consensus algorithm to enhance the fault tolerance of the blockchain network and ensure the efficient processing of trajectory data from roadside units.

[0073] In summary, the method combining blockchain technology and asynchronous federated learning can significantly improve the accuracy and robustness of vehicle trajectory prediction in multiple key aspects, laying a solid foundation for the development of future intelligent transportation systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] To more clearly illustrate the technical solutions of the present invention, the accompanying drawings required for the implementation examples or the description of the prior art will be briefly introduced below.

[0075] Figure 1 is the overall principle framework diagram of the present invention;

[0076] Figure 2 is the Bi-GRU network structure described in the present invention;

[0077] Figure 3 is the asynchronous federated learning described in the present invention;

[0078] Figure 4 is the PPO algorithm learning process described in the present invention;

[0079] Figure 5 The communication process of the DG-PBFT consensus algorithm described in the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0080] The following will further describe in detail the specific embodiments of the present invention in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention but are not intended to limit the scope of the present invention.

[0081] Problem description: Assume that the historical state information of the target vehicle is Lobs is the time span of historical observations, and represents the state of the target vehicle at time t, including horizontal and vertical position coordinates speed acceleration and yaw rate To simplify the model, the present invention does not consider information of traffic participants such as pedestrians and traffic facilities, and only considers the historical trajectories of other vehicles within the area around the target vehicle. For neighboring other vehicles, the specific definition of the motion information is as follows:

[0082]

[0083] Among them, represents the set of all other vehicles around the target vehicle, represents the i-th agent, where i ∈ {1, 2,..., N}.

[0084] Therefore, the historical state sequence information of other traffic participants around the target vehicle can be expressed as:

[0085]

[0086] The task of the present invention is to accurately predict the future trajectory of the target vehicle by obtaining the historical trajectory data T of the vehicle within a given time range aim and the historical sequence features S of other vehicles around it other , and using the historical data and dynamic interaction information. The output is the future waypoints within the prediction range t p , which are usually expressed as a sequence of expected positions of the vehicle in the future time period. Specifically, it can be defined as:

[0087]

[0088] Among them, and respectively represent the predicted horizontal coordinate and predicted vertical coordinate at the future time point (n + k), where k ∈ {1, 2,..., t p}}. And there is t p is the prediction range, indicating the number of future time steps or time periods.

[0089] To achieve the above task, as Figure 1 shown, the present invention includes the following steps:

[0090] Step 1: Data collection and preprocessing

[0091] First, the historical trajectory data of the target vehicle and the traffic participants around it are collected in real time through sensors, including information such as position coordinates x and y, speed v, and acceleration a, so as to generate corresponding node features. These features will provide a basis for subsequent model training to help better understand and predict the dynamic behavior of the vehicle, and then statistical methods are used to remove outliers. Suppose there is a set of attribute values First, calculate the mean value μ s and the standard deviation σ s, and then according to a predefined threshold (usually 2 or 3 times the standard deviation) to determine the range of outliers. The specific definition of outliers is as follows:

[0092]

[0093] Then, use the linear interpolation method to fill in the missing values in each attribute. Assume that the missing value in the dataset is at index i, and there are valid values before and after this index i and The following linear interpolation calculation is:

[0094]

[0095] Step 2, Local Training: Based on the Bi-GRU Trajectory Prediction Model

[0096] In trajectory prediction, data-driven methods can effectively extract the historical state information of vehicles and generate trajectories that better conform to the "intentions" of vehicle driving. Widely used models include Recurrent Neural Network (RNN), Long Short-Term Memory Network (LSTM), and Gated Recurrent Unit (GRU). Although these models can all handle time series data, RNN is prone to short-term memory limitations, resulting in information forgetting. While LSTM and GRU solve the "gradient vanishing" problem of RNN through a "gating" mechanism. Among them, GRU has a faster training speed due to fewer gating units and simplified tensor operations. Therefore, GRU is selected in the present invention to extract the historical state features of vehicles, such as Figure 2 As shown, the calculation formula of GRU is as follows:

[0097] ν n = σ(W iγ x t + b iγ + W hγ h (n-1) + b hγ )

[0098]

[0099] ι n = tanh(W in x t + b in + γ(W hn h (n-1) + b hn ))

[0100] Among them, σ represents the Sigmoid function, v t , γ t and ι t are the reset gate, update gate, and cell state respectively. The reset gate vt Determine how much information from the previous cell hidden state ι t needs to be forgotten. Update the gate γ t Determine how much information is passed from the previous hidden layer state to the current hidden layer state ι t These gates control the learning process of neural units formed through a large amount of training data. Additionally, W represents the weight matrix and b represents the bias term.

[0101] To further improve the prediction performance, the present invention adopts a bidirectional GRU (Bi-GRU) model. Bi-GRU combines forward and backward information flows, enabling it to consider both past and future information simultaneously, which is crucial for capturing the driving intention of the vehicle. In the specific implementation, the model processes the historical trajectory and future trajectory data of the target vehicle separately, thereby generating more accurate trajectory prediction results and effectively extracting the motion features of the target vehicle. During the training process, Bi-GRU processes the input data through a gating mechanism, overcoming the problem of gradient disappearance and thus outputting hidden states. These hidden states not only capture the dynamic changes of the vehicle but also provide rich information for the subsequent global model aggregation.

[0102] The input data consists of node features and edge features, where the node features include the historical state time series and future trajectory of the target vehicle. Since these two node features have different time spans, the historical motion features and future motion features are extracted separately. In this model, two GRUs are used to perform forward and backward passes on the traffic flow sequence respectively, thereby generating two candidate hidden states. The specific calculation formulas are as follows:

[0103]

[0104] where, represents the GRU encoded output of the historical trajectory T aim of the target vehicle, represents the future trajectory of the target vehicle and its GRU encoded output.

[0105] Finally, the final states of the two neural networks are combined together as the final output result. Therefore, the output result of the Bi-GRU network can be expressed as:

[0106]

[0107] In this way, it is possible to more comprehensively capture the driving intention of the vehicle and improve the accuracy and reliability of trajectory prediction.

[0108] Step 3: Asynchronous Federated Learning (AFL) aggregates and updates the global model

[0109] Asynchronous Federated Learning (AFL) is a variant of federated learning that allows clients to perform local updates and parameter synchronization at different times. In asynchronous federated learning, the central server no longer waits for all clients to complete local updates but instead aggregates immediately upon receiving updates from any client node. As Figure 3 shown, for each client i, the features it extracts and are sent to the federated server for aggregation, and its model update can be expressed as:

[0110]

[0111] where ζ is the client i's model parameters and T represents the training process.

[0112] The AFL aggregation server performs weighted aggregation on the received local model updates Δζ i . Due to asynchrony, the weight of each client is related to its submission time point and participation frequency. The aggregation formula is:

[0113]

[0114] where, ζ n is the parameter of the global model at time n, w i is the weight of client i, and η is the learning rate. The aggregated global model ζ n+1 is sent to each client for the next round of training. In this way, AFL can effectively utilize the data distributed across different clients, improve the generalization ability and convergence speed of the model, while ensuring data privacy.

[0115] Step 4, Curiosity-Exploration-Based Proximal Policy Optimization Algorithm (C-PPO)

[0116] After completing the aggregation step of federated learning, a global model integrating multiple local models is obtained. This global model will be used as the input for optimizing the policy. Specifically, the PPO algorithm uses the global model parameters obtained from federated learning to adapt the policy network to the complexity and dynamics of the environment. On this basis, the trajectory prediction accuracy and decision-making ability are improved, as Figure 4 shown.

[0117] First, set the global model parameters ζ n+1 , to initialize the parameters ψ of the policy network. It can be expressed as: ψ 0 = ζ n+1 . The key of PPO is to avoid policy failure caused by over-updating by restricting the update amplitude of the policy. The goal is to maximize the expected return of the policy, and the optimization problem can be expressed as follows:

[0118]

[0119] wherein is the advantage function, which describes the advantage of the current action l relative to the old policy.

[0120] Next, the main task of the present invention is to select the best server model parameters by optimizing the policy network. The definitions are as follows:

[0121]

[0122] where clip(r n (ψ), 1 - ∈, 1 + ∈) is a trimming operation on the probability ratio, which restricts the update amplitude. This is to prevent the policy from changing too much in each iteration. r n (ψ) is the probability ratio of the policy, which represents the ratio of the probability of selecting action l between the new policy and the old policy in state q. Its calculation method is as follows:

[0123]

[0124] For the loss of the value function, its specific definition is:

[0125]

[0126] where V ψ (q n ) is the value function in the policy network. In addition, is the target value function, which is usually calculated through the advantage function and the reward.

[0127] Therefore, the final total loss will be obtained. The content is as follows:

[0128]

[0129] The present invention minimizes the loss function by using the optimization algorithm Adam

[0130]

[0131] where ψ n+1 is the updated policy parameter, is the gradient of the PPO policy loss function.

[0132] Step 5, Practical Byzantine Fault Tolerant Consensus Algorithm Based on Dynamic Grouping

[0133] Although these improved algorithms have processed trajectory data, in the federated learning environment, malicious node attacks remain a problem. Introducing blockchain technology and proposing the DG-PBFT consensus algorithm can effectively solve the throughput limitation of traditional blockchain systems and reduce the communication overhead in the consensus process. This method not only improves the security and data integrity of the model under multi-node collaboration but also effectively reduces the risks brought by malicious nodes. Specifically, DG-PBFT can be divided into two stages: node grouping and consensus communication.

[0134] As Figure 5 shown, the DG-PBFT consensus algorithm first uses a grouping method to organize nodes into multiple broadcast domains to facilitate consensus. In the blockchain, assume the node set is Then select G nodes from all nodes as the initial central nodes They are indexed by positive integer n g (g ∈ [1, G]). The remaining nodes are denoted as O = {n 1 , n 2 , …, n O}, indexed by positive integer n o (o ∈ [1, N - G]).

[0135] 1) Grouping process: To group all nodes, first calculate the Euclidean distance between all nodes as follows:

[0136]

[0137] Then retain the nearest Euclidean distance from each node n o to the current central node n g . The retained nodes n o will be divided into C g groups, specifically represented as:

[0138]

[0139] 2) Central node optimization: When the sum of the Euclidean distances from node n g to other nodes n o in the group is the smallest, this node is regarded as the central node of the group. To find the optimal set of central nodes replace the original central node n g with a non-central node n g from C o , and when the sum of its distances to n' g is the smallest, the update rule for the central node is as follows:

[0140]

[0141] 3) Stability Iteration: Steps 1) and 2) will be repeated until the central node set is consistent. If a central node n g has a problem (such as timeout or malicious behavior), then a new central node will be selected and regrouped.

[0142] After the node grouping, the consortium blockchain nodes are divided into G broadcast domains, thus avoiding network congestion caused by broadcast messages. Therefore, the consensus communication process is divided into four stages: block announcement, intra-group consensus, inter-group consensus, and block synchronization:

[0143] 1) Block Announcement: The primary node n selected from the central node set * packages a segment of client transaction requests into a block and encapsulates it as an announcement message:

[0144]

[0145] Then it announces it to other aggregator node groups.

[0146] 2) Intra-group Consensus: After receiving the announcement message, the aggregator nodes in each group perform intra-group PBFT consensus. After the intra-group consensus is completed, a confirmation message is sent:

[0147]

[0148] to the primary node n * , and then enter the inter-group consensus stage.

[0149] 3) Inter-group Consensus: The reputation values of the aggregator nodes selected through intra-group consensus are regarded as secure and reliable. Therefore, inter-group consensus can be completed through log information synchronization. When the primary node n * receives the confirmation messages from all aggregator nodes , the primary node will send a repeated confirmation message:

[0150]

[0151] After other aggregator nodes receive the repeated confirmation message, they send a confirmation message to the primary node n * again:

[0152]

[0153] Then enter the block synchronization stage.

[0154] 4) Block Synchronization: After both the intra-group and inter-group consensus processes are completed, the aggregator node set sends an execution message:

[0155]

[0156] Notification access node set Execute local block writing. All aggregation nodes Send a response message to the client:

[0157] <Response,v,ts,c,n,n * ,e>

[0158] The consensus process is completed.

[0159] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features. And these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.

Claims

1. A blockchain-based asynchronous federated proximal policy optimization reinforcement learning method, characterized by: The method comprises the following steps: Step 1: Use sensors to collect historical trajectory data of the target vehicle and its surrounding traffic participants in real time to generate corresponding node features; then use statistical methods to remove outliers; Step 2: On each participant, a dual-gated recurrent unit network is used to perform local training on the collected historical trajectory data; Step 3: Adopt the asynchronous federated learning method to aggregate the updated parameters generated by all participants in local training to generate a new global model; Step 4: C-PPO optimizes the policy network to maximize the expected return by limiting the update amplitude of the policy and using the parameters obtained from the global model; Step 5: Generate a transaction containing updated parameters and hidden states and submit it to the blockchain network to ensure data transparency and immutability. In this process, a dynamic grouping practical Byzantine fault-tolerant consensus algorithm is used to verify the submitted transactions to ensure the validity and consistency of each transaction. Step 6: Distribute the updated global model to all participants, and the participants perform local training based on the received global model to ensure continuous optimization of the model.

2. According to claim 1, the asynchronous federated proximal strategy optimization reinforcement learning method based on blockchain is characterized in that: In step 2, the dual-gated recurrent unit network includes a historical GRU and a future GRU, which are used to extract historical motion features and future trajectory features respectively. The specific calculation formula is as follows: (1) GRU encoding output of historical trajectory: (2) GRU encoding output of future trajectory: in, Represents the historical trajectory T of the target vehicle aim The GRU encoding output is represents the future trajectory of the target vehicle The GRU encoding output of (3) The final hidden states of the two GRUs are combined as the final output result, so the output result of the Bi-GRU network can be expressed as:

3. According to claim 1, the asynchronous federated proximal strategy optimization reinforcement learning method based on blockchain is characterized in that: In step 3, in asynchronous federated learning, clients make local updates at different time points, and the central server immediately aggregates any updates received; for each client i, its extracted features and updated parameters Δθ i Send to the federation server; (1) The model update of client i is expressed as: where Δζ i are the model parameters of client i, and T represents the training process; (2) The AFL aggregation server performs weighted aggregation on the received local model updates; due to asynchrony, the weight Δζ of each client i Related to its submission time point and participation frequency, the global model parameter update formula is: Among them, n are the parameters of the global model at time n, w i is the weight of client i, η is the learning rate, and the aggregated global model ζ n+1 is sent to each client for the next round of training; (3) Weight w i The calculation is as follows: Among them, f i represents the participation frequency of client i, t i is the latest submission time point of client i, α is the adjustment factor; the aggregated global model parameter ζ n+1 is sent to each client for the next round of local training.

4. According to claim 1, the asynchronous federated proximal strategy optimization reinforcement learning method based on blockchain is characterized in that: In step 4, after completing the aggregation step of federated learning, a global model integrating multiple local models is obtained, and the PPO algorithm is used to use the global model parameters obtained from federated learning to adapt the policy network to the complexity and dynamics of the environment: (1) First, set the global model parameter ζ n+1 To initialize the parameters of the policy network, the optimization goal is to maximize the expected return of the strategy, expressed as: in, is the advantage function, describing the advantage of the current action l over the old strategy. The parameter ψ of the initialization strategy network can be expressed as ψ0 = ζ n+1 ; (2) Select the best server model parameters through the optimization strategy network, which are defined as follows: Where ε is the clipping threshold set to 0.1, clip(r n (ψ), 1-∈, 1+∈) is the probability ratio, which limits the update amplitude to prevent the strategy from changing too much in each iteration. It is calculated as: Among them, π ψ (l n ∣q n ) indicates that the old policy is in state q n Next action is taken n The probability of π ψ (l n+1 ∣q n+1 ) indicates that the new strategy is in state q n+1 Next action is taken n+1 probability; (3) The loss of the value function is specifically defined as: Among them, V ψ (q n ) is the value function in the policy network, is the target value function calculated by the advantage function and the reward; the final total loss function is: (4) Use the optimization algorithm Adam to minimize the loss function To update the policy parameters, after each training iteration, the policy parameters are updated by calculating the gradient: Among them, α is the learning rate, ψ n+1 are the updated policy parameters, is the gradient of the loss function of the PPO strategy.

5. According to claim 1, the asynchronous federated proximal strategy optimization reinforcement learning method based on blockchain is characterized in that: The step 5, the step of dynamically grouping the practical Byzantine fault-tolerant consensus algorithm is divided into two stages: node grouping and consensus communication: First, a grouping method is used to organize nodes into multiple broadcast domains to facilitate consensus. In the blockchain, the node set is Select G nodes from all nodes as the initial central nodes They are positive integers n g (g∈[1,G]) is the index, and the remaining nodes are denoted as O={n1,n2,…,n O }, with a positive integer n o (o∈[1,NG]) is the index; 1) Grouping process: To group all nodes, first calculate the Euclidean distance between all nodes, as follows: Then keep each node n o To the current central node n g The nearest Euclidean distance of the retained node n o will be divided into C g Group, specifically expressed as: 2) Central node optimization: When node n g To other nodes n in the group o When the sum of the Euclidean distances is the smallest, the node is regarded as the central node of the group. To find the best central node set The original central node n g Replace with from C g Non-central node n o , when it reaches n' g When the sum of the distances is the smallest, the update rule of the central node is as follows: 3) Stability iteration: Steps 1) and 2) will be repeated until the central node set and Keep consistent; if a central node n g If a problem occurs, the central node will be reselected and the grouping will be re-done; After the nodes are grouped, the alliance blockchain nodes are divided into G broadcast domains, and the consensus communication process is divided into four stages: block announcement, intra-group consensus, inter-group consensus, and block synchronization: 1) Block announcement: collected from the central node The master node n is selected * Package a client transaction request into a block and encapsulate it into an announcement message: Then announce it to other aggregation node groups; 2) Intra-group consensus: After receiving the announcement message, the aggregation node of each group conducts intra-group PBFT consensus. After the intra-group consensus is completed, a confirmation message is sent: To the master node n * , and then enter the inter-group consensus stage; 3) Inter-group consensus: The reputation value of the aggregation node selected by the intra-group consensus is considered safe and reliable; when the master node n * Received from all aggregation nodes After receiving the confirmation message, the master node will send a duplicate confirmation message: After receiving the duplicate confirmation message, other aggregation nodes send it to the master node n again. * Send confirmation message: Then enter the block synchronization phase; 4) Block synchronization: After the consensus process within and between groups is completed, the aggregation node set Send an execution message: Notify access node set Execute local block writes, all aggregation nodes Send a response message to the client: <Response,v,ts,c,n,n * ,and> The consensus process is complete.