A CAV Decision Method Using Selective Federated Learning

Through the selective federated learning method, combined with deep reinforcement learning and multi-factor-driven model screening strategy, the problems of inconsistent model quality and large communication burden in federated learning of autonomous vehicles are solved, the global model accuracy and efficiency are improved, and efficient autonomous driving decisions are achieved.

CN115660113BActive Publication Date: 2025-07-25XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211310743.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-25
Publication Date
2025-07-25
Estimated Expiration
2042-10-25

AI Technical Summary

Technical Problem

In the existing federated learning method for autonomous driving, inconsistent local model quality among participants leads to low global model accuracy, large communication burden, and the quality of computing resources and communication links affecting efficiency, especially in the case of limited computing resources or poor wireless channels.

Method used

The selective federated learning method is adopted to use a multi-factor-driven model screening strategy, including based on reputation, model convergence and utility screening, and select high-quality local models to upload to the FL server. Combined with deep reinforcement learning, local decision model is designed with relative position and velocity as the state space, acceleration as the action space, and safety and comfort as the reward function.

Benefits of technology

It improves the accuracy and efficiency of the global model, reduces communication overhead and computing time, selectively screening high-quality users accelerates the federated learning training process, and improves the self-learning and decision-making capabilities of autonomous driving vehicles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115660113B_ABST
    Figure CN115660113B_ABST
Patent Text Reader

Abstract

The present invention discloses a CAV decision-making method applying selective federated learning, including: all autonomous vehicles download the globally updated model of the previous iteration from the FL server; all autonomous vehicles use their own driving data to train the local model and update the model parameters; obtain the communication overhead and computing overhead of all autonomous vehicles during the federated learning process; perform multi-factor-driven screening of the local models, and the selected autonomous vehicles upload the updated local models to the FL server, wherein the local model screening includes screening based on reputation, screening based on model convergence, and screening based on maximizing utility; the FL server aggregates the received local models to obtain an improved global model. The present invention proposes a multi-factor-driven model screening strategy before uploading the local models, overcomes the defect of the traditional method that does not distinguish participants, and can select as many high-quality users as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of autonomous driving, and particularly relates to a CAV decision-making method applying selective federated learning. Background Art

[0002] Autonomous vehicles (CAVs) utilize on-vehicle sensors and communication modules to enhance their perception capabilities and achieve their motion planning and automatic control. Machine learning algorithms are widely used to perform autonomous driving tasks, such as pedestrian detection, automatic lane change, collision avoidance, etc. Generally speaking, accurate autonomous driving models need to be built on a large amount of data. However, the scenarios involved in a single CAV are limited, and its perception and computing capabilities cannot cope with complex traffic environments. More importantly, the knowledge obtained from CAVs through model training exists locally, and it is difficult to reuse the knowledge and cooperate with other CAVs. With the concern for user privacy and data security, how to solve the data problem while protecting privacy is a huge challenge faced by artificial intelligence (AI) empowering CAVs. To overcome these challenges, federated learning (FL) can be used as a promising solution to limit the amount of data transmission and accelerate the learning process of CAVs.

[0003] Most existing autonomous driving federated learning methods are that autonomous vehicles participating in federated learning directly upload all the locally trained models to the FL server for aggregation after training. The basic idea is as follows: 1) Model download: All autonomous vehicles download the global model from the FL server for local training; 2) Local training: Autonomous vehicles use local sensing data to independently train their local models; 3) Model upload: All autonomous vehicles upload the trained local models to the FL server; 4) Model aggregation: The received local models are aggregated by the FL server to obtain an improved global model.

[0004] Liang et al. proposed an online federated learning process for real-time knowledge extraction to perform the collision avoidance task of autonomous vehicles. Vehicles locally train a reinforcement learning model and share the local model with the central server, so that vehicles can make corresponding actions according to the knowledge learned by others even when driving in different environments. Fu et al. proposed a distributed machine learning model sharing architecture for autonomous vehicles. Vehicles locally train a neural network model to learn the evolution of positioning errors and correct the positioning errors by sharing the vehicle positioning error evolution model.

[0005] First, for autonomous vehicles, the collected data and trained models are affected by the external environment, performance, and reputation, and there are usually errors to varying degrees. The quality of the local model may seriously affect the accuracy of the global model. Second, when a large number of autonomous vehicles participate in federated learning, the communication burden increases with the increase in the number of participants and the number of iterative rounds. Finally, the computing power and communication link quality affect the efficiency of federated learning. If the computing resources of the participating autonomous vehicles are limited or under poor wireless channel conditions, the model update time will be extended. Summary of the Invention

[0006] To solve the above problems existing in the prior art, the present invention provides a CAV decision-making method applying selective federated learning. The technical problems to be solved by the present invention are realized through the following technical solutions:

[0007] The present invention provides a CAV decision-making method applying selective federated learning, including:

[0008] S1: Model download: All autonomous vehicles download the globally updated model of the previous iteration from the FL server;

[0009] S2: Local training: All autonomous vehicles use their own driving data to train the local model and update the model parameters. The local model uses the relative position and relative speed as the state space, the acceleration as the action space, and the safety and comfort as the reward function;

[0010] S3: Overhead calculation: Obtain the communication overhead and computing overhead of all autonomous vehicles during the federated learning process;

[0011] S4: Model screening and uploading: Perform multi-factor-driven local model screening, and the selected autonomous vehicles upload the updated local model to the FL server. Among them, the local model screening includes screening based on reputation, screening based on model convergence, and screening based on maximizing utility;

[0012] S5: Model aggregation: The FL server aggregates the received local models to obtain an improved global model.

[0013] In an embodiment of the present invention, the state space is expressed as:

[0014] s(t) = {s ov (t), s op (t)}

[0015] where s ov (t) represents the speed v ov (t), position P ov (t), and acceleration a ov (t) of the target vehicle; sop The speed v of other traffic participants is represented by (t). op The position P(t). op And the acceleration a(t). op (t);

[0016] The action space is represented as:

[0017] a(t) = {a(t)}. ov (t)}.

[0018] In an embodiment of the present invention, the reward function includes a reward function regarding safety, a reward function regarding comfort, and a reward function regarding efficiency, where

[0019] The reward function regarding safety is represented as:

[0020] r(t) = -m[v(t) s +α]||{Collsion}, ov (t) 2 +α]||{Collision},

[0021] where v(t) represents the speed of the target vehicle, m and α are weight parameters, and {Collision} represents a value of 1 if a collision occurs, otherwise 0; ov (t) represents the speed of the target vehicle, m and α are weight parameters, {Collision} represents a value of 1 if a collision occurs, otherwise 0;

[0022] The reward function regarding comfort is represented as:

[0023]

[0024] where k and β are weight parameters, and j(t) represents the jerk of the target vehicle, ov (t) represents the jerk of the target vehicle, represents the threshold of acceleration, represents the threshold of jerk, and a(t) represents the acceleration of the target vehicle; ov (t) represents the acceleration of the target vehicle;

[0025] The reward function regarding efficiency is represented as:

[0026]

[0027] where s(t) represents the distance between the target vehicle and other traffic participants, and s represents the threshold for indicating whether the braking distance is appropriate, and δ and l are weight parameters. rel (t) represents the distance between the target vehicle and other traffic participants, and s represents the threshold for indicating whether the braking distance is appropriate, and δ and l are weight parameters. th represents the threshold for indicating whether the braking distance is appropriate, and δ and l are weight parameters.

[0028] In an embodiment of the present invention, S3 includes:

[0029] S31: Obtain the autonomous vehicle v iExecution time for local model training in the r-th iteration:

[0030]

[0031] where, represents the data samples participating in federated learning in the r-th iteration of the autonomous vehicle, f i (r) represents autonomous vehicle v i contributes computing resources for local model training, represents the number of iterations for local model update when the global accuracy is fixed, represents the number of CPU cycles for executing one data sample;

[0032] S32: Obtain the time consumed in the upload process of the autonomous vehicle in the r-th iteration:

[0033]

[0034] where, represents autonomous vehicle v i the size of the local model parameters uploaded in the r-th round, ur i (r) represents autonomous vehicle v i the uplink rate of uploading data in the r-th round;

[0035] S33: Obtain the total time consumption of autonomous vehicle v i in the r-th iteration according to the execution time of local model training and the time consumed in the upload process:

[0036]

[0037] In an embodiment of the present invention, the S4 includes:

[0038] S41: Conduct a reputation assessment on the autonomous vehicle and select multiple local models with high reputation values in each learning iteration;

[0039] S42: Use similarity to screen the selected multiple local models that meet the reputation value requirements based on model convergence, and select local models that meet the model convergence requirements;

[0040] S43: Conduct screening based on maximizing utility on the selection results based on reputation and based on convergence trend, with the constraint:

[0041]

[0042] where, represents autonomous vehicle v iThe execution time of local model training in the r-th round of iteration represents the autonomous vehicle v i The time consumed in the upload process in the r-th round of iteration represents the longest time threshold allowed by the FL server

[0043] In an embodiment of the present invention, the reputation evaluation of the autonomous vehicle specifically includes:

[0044] For the autonomous vehicle v i , the reputation value update in the r-th round is expressed as:

[0045]

[0046] Wherein, is the reputation value of the current autonomous vehicle v i in the previous round, is the comprehensive reputation score obtained by the autonomous vehicle v i after participating in the last FL iteration, represents the reputation score obtained from the completion time reliability θ t , represents the reputation score obtained from the local model quality reliability θ q .

[0047] In an embodiment of the present invention, the S42 includes:

[0048] Obtain the similarity between the updated parameters of the local model and the updated parameters of the global model:

[0049]

[0050] Wherein, sgn() is the sign function, which is used to compare whether the signs of the i-th parameter updated by the local model and the i-th parameter in the global model update are the same. If the signs of the i-th parameter updated locally and the i-th parameter in the global update are the same, then take 1, otherwise take 0;

[0051] For two models with the same similarity to the global model, the cosine similarity algorithm is introduced to perform a secondary screening on the selected models. The parameter cosine similarity between the local model and the global model is expressed as:

[0052]

[0053] Wherein, u i represents the set of parameter updates of the autonomous vehicle v i , u ij represents the j-th parameter of the autonomous vehicle v i . Denotes a set of global model parameter updates, u j Denotes the average value of the j-th parameter of all users participating in the upload;

[0054] A cosine similarity threshold is preset. After obtaining the cosine similarity between the local update parameter and the global update parameter, it is compared with the cosine similarity threshold. The local update parameter greater than the cosine similarity threshold is retained, and the local update parameter with a cosine similarity less than the cosine similarity threshold is removed.

[0055] Another aspect of the present invention provides a storage medium in which a computer program is stored, and the computer program is used to execute the steps of the CAV decision method for applying selective federated learning in any one of the above embodiments.

[0056] Yet another aspect of the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the CAV decision method for applying selective federated learning in any one of the above embodiments are implemented.

[0057] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0058] 1. The CAV decision method applying selective federated learning in the present invention proposes a multi-factor driven model screening strategy before uploading the local model in order to achieve efficient federated learning, overcomes the defect of not differentiating participants in the traditional method, and comprehensively considers the reputation of participants, the quality of the local model, and the time overhead, so as to select as many high-quality users as possible while considering communication resources and time limitations.

[0059] 2. The vehicle local decision model of the present invention comprehensively considers safety, comfort, and traffic efficiency, and is closer to reality.

[0060] 3. When performing screening based on model convergence proposed in the present invention, comprehensive screening is performed through the sign of model update and model similarity, and the similarity between the obtained local model and the global model is higher, accelerating the federated learning training process.

[0061] The present invention will be further described in detail below with reference to the drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] Figure 1 Is a flowchart of a CAV decision method for applying selective federated learning provided by an embodiment of the present invention;

[0063] Figure 2It is a comparison chart of the communication overhead of the selective federated learning method (SFRL) in this embodiment, the existing CMFL method, and the FedAVG method during the FL process. Detailed implementation manners

[0064] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following combines the accompanying drawings and specific implementation manners to detail a CAV decision-making method applying selective federated learning proposed according to the present invention.

[0065] The foregoing and other technical contents, features, and effects of the present invention can be clearly presented in the following detailed description in conjunction with the accompanying drawings. Through the description of the specific implementation manners, a more in-depth and specific understanding of the technical means and effects adopted by the present invention to achieve the predetermined purpose can be obtained. However, the accompanying drawings are only provided for reference and illustration, and are not used to limit the technical solution of the present invention.

[0066] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant is intended to cover non-exclusive inclusion, so that an article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the article or device including the said element.

[0067] Assume that a group of connected autonomous vehicles (CAVs) participating in federated learning (FL) is represented as V = {v1, v2,..., v N}, where N is the total number of connected autonomous vehicles. Different connected autonomous vehicles perform driving tasks and train local models in their respective environments. All connected autonomous vehicles share the same model architecture, that is, the initialization parameter settings are consistent. The local models uploaded by connected autonomous vehicles in each round can be aggregated in the FL server. It should be noted that the local model refers to the model of each connected autonomous vehicle participating in federated learning. After the connected autonomous vehicle participating in federated learning uploads the local model, it is aggregated in the FL server to obtain a global model. In the next round, each connected autonomous vehicle downloads the global model from the FL server and trains a new local model based on its own data, and so on.

[0068] Please refer to Figure 1 , Figure 1The figure is a flowchart of a CAV decision-making method applying selective federated learning provided by an embodiment of the present invention. The CAV decision-making method includes:

[0069] S1: Model download: All autonomous vehicles download the global model from the FL server (such as a base station) for local training.

[0070] S2: Local training: All autonomous vehicles use their own driving data to train the local model and update the model parameters. The local model uses relative position and relative speed as the state space, acceleration as the action space, and safety and comfort as the reward function. The driving data includes speed, acceleration, etc.

[0071] To improve the self-learning and decision-making capabilities of CAVs in the driving environment, this embodiment designs a local autonomous driving decision-making model, that is, the local model, using deep reinforcement learning (DRL). This model uses relative position and relative speed as the state space, acceleration as the action space, and safety and comfort as the reward function. The training process of reinforcement learning is: based on the current state, select an action, and then calculate the reward, continuously trying to maximize the reward.

[0072] 1) State space: A CAV can obtain the driving state information of itself and another traffic participant (such as another vehicle) at time t, and then use relative position, speed, and acceleration as the state space s(t). After receiving the current state s(t) at time t, the DRL agent outputs an operation and obtains a reward value. At this time, the state changes to s(t + 1).

[0073] s(t) = {s ov (t), s op (t)}

[0074] where s ov (t) represents the state of the target vehicle OV, including speed v ov (t), position P ov (t), and acceleration a ov (t); s op (t) represents the state of other traffic participants, including speed v op (t), position P op (t), and acceleration a op (t).

[0075] 2) Action space: In a dangerous situation, a CAV (including the target vehicle OV) avoids collisions by automatic braking. Therefore, the action space a(t) includes the acceleration a ov (t) for longitudinal control of OV.

[0076] a(t) = {a ov (t)}

[0077] where a ov (t) ∈ [-8, 0] represents different braking levels, a ov (t) = 0 means the OV is traveling at a constant speed, and a ov (t) = -8 means hard braking.

[0078] 3) Reward function: The reward function is used to guide the adjustment direction of the local model network parameters so that the OV can execute according to the required actions. According to the collision avoidance goal, the design of the reward function should follow the following principles: (1) Safety, that is, the OV can brake appropriately to eliminate risks. (2) Comfort, that is, the OV can effectively control the acceleration or acceleration change during braking. (3) Efficiency, that is, the braking distance of the OV cannot be too conservative.

[0079] Regarding safety, if the OV fails to avoid an accident in the current event, a penalty is required. The severity of the collision can be reflected by the current speed of the OV. Therefore, the reward function r s (t) regarding safety can be expressed as:

[0080] r s (t) = -m[v ov (t) 2 +α]||{Collsion},

[0081] where v ov (t) represents the speed of the target vehicle OV, and m and α are weight parameters. The term {Collsion} means the value is 1 if a collision occurs, otherwise 0.

[0082] Regarding comfort, if the acceleration or jerk change of the OV is too large during braking, it will bring a large inertial impact to the passengers. Therefore, it is necessary to penalize the uneven braking process. The reward function r c (t) regarding comfort can be expressed as:

[0083]

[0084] where k and β are weight parameters. j ov (t) represents the jerk of the OV, represents the threshold of acceleration, represents the threshold of jerk, and a ov (t) represents the acceleration of the target vehicle OV.

[0085] Regarding efficiency, if the OV brakes too early, the braking distance will be too conservative, which will affect traffic efficiency. In this embodiment, the reward function r e(t) can be expressed as:

[0086]

[0087] where s rel (t) represents the distance between OV and other traffic participants, and s th represents the threshold used to indicate whether the braking distance is appropriate, and δ and l are weight parameters.

[0088] The final reward function is the weighted sum of three reward functions, that is, the final reward function can be expressed as:

[0089] r(t) = r s (t) + r c (t) + r e (t).

[0090] S3: Overhead calculation: Obtain the communication overhead and computing overhead of all autonomous vehicles during the federated learning process;

[0091] Considering that each autonomous vehicle participating in FL will obtain the same initialized local model and then use local data to train this local model. Therefore, it can be considered that for each autonomous vehicle, in each iterative training, the size of the uploaded and downloaded models is the same, denoted as Z. Downloading the global model and uploading the local model require the autonomous vehicle to communicate with the FL server.

[0092] Download process: In the r-th round of iterative training, all N autonomous vehicles will download the latest global model, so the communication overhead of the download process is where Z (r) represents the size of the global model downloaded in the r-th round of iterative training.

[0093] Upload process: In the r-th round of iterative training, assume that there are n r (n r ≤ N) autonomous vehicles upload their local models to the FL server, that is, in the r-th round, the number of autonomous vehicles that need to upload the updated local model to the FL server is n r , then the communication overhead of the upload process is Z (r) represents the size of the local model uploaded in the r-th round of iterative training. As described above, in each iterative training, the size of the uploaded local model and the downloaded global model is the same.

[0094] Therefore, the total communication cost C comm after R rounds of iterative training is:

[0095]

[0096] Among them, R represents the total number of rounds of federated learning.

[0097] It can be seen that for the download process, since all autonomous vehicles need to participate, the communication overhead is actually fixed. Therefore, the key to reducing the communication overhead lies in the upload process. That is to say, on the premise that the global model converges, it is necessary to be able to reduce the number n of autonomous vehicles uploading in each iteration of the model r .

[0098] In one iteration, both communication (downloading and uploading) and local computing take time. Due to the powerful computing power of the FL server, the time consumed by the model aggregation process can be ignored. Therefore, for any autonomous vehicle v i , the time consumption in the r-th round of iterative training can be expressed as:

[0099]

[0100] Among them, and respectively represent the time consumed in the r-th round of download process, local model training process, and upload process. In synchronous FL, the communication and computing of each autonomous vehicle are parallel. Therefore, the time cost of each iteration should be determined by the autonomous vehicle with the longest total time (the slowest autonomous vehicle). It can be inferred that the time cost of each iteration can be expressed as:

[0101]

[0102] Among them, is the synchronization delay required by the FL server, and the synchronization delay is a preset maximum time threshold allowed by a server. Therefore, the total time cost T after R rounds cost is:

[0103]

[0104] Next, the communication time cost in the local model upload, local model training, and global model download processes will be analyzed in detail.

[0105] S31: Obtain the execution time of autonomous vehicle v i for local model training in the r-th round of iteration.

[0106] Local model training time cost: Assume that autonomous vehicle v with a local dataset i uses data samples to participate in federated learning, where represents the data samples of autonomous vehicle v i in the r-th round of iteration. Autonomous vehicle vi Contribute computing resources for local model training, i.e., f i (r) , denoted as the local processing capacity, i.e., the frequency of the CPU of the autonomous vehicle. Therefore, for autonomous vehicle v i , the execution time of one iteration in local model training is:

[0107]

[0108] where denotes the number of CPU cycles to execute one data sample.

[0109] Generally speaking, the better the data quality, the fewer the number of iterations in local model training and global model training. Therefore, autonomous vehicle v i The computing time in one round of training iteration can be expressed as:

[0110]

[0111] where denotes the number of iterations for local model update when the global accuracy is fixed, ε i denotes autonomous vehicle v i 's local data quality. The better the data quality, i.e., the larger the value of ε i .

[0112] S32: Obtain the time consumed by the autonomous vehicle in the uploading process during the r-th iteration.

[0113] Communication time cost: For the uploading process of the local model, the time consumed by autonomous vehicle v i in the uploading process during the r-th iteration can be expressed as:

[0114]

[0115] where denotes the size of the local model parameters uploaded by autonomous vehicle v i in the r-th round. Since the structure of the model is fixed, it can be assumed that the size of the model is constant for any autonomous vehicle throughout the FL process, i.e., Considering the orthogonal frequency division multiple access scheme, ur i (r) The uplink rate is expressed as:

[0116]

[0117] where B, P i , h i , α, δ i$N_0$ respectively represent the bandwidth, transmission power, channel gain, path loss exponent, received interference power, and noise power spectral density of the base station. $d$ i represents the distance between the autonomous vehicle $v$ i and the base station.

[0118] During the download process, the base station can have a higher transmission power and a larger bandwidth. Therefore, compared with the uplink transmission delay, the delay of downloading the global model is generally considered negligible. Therefore, for the autonomous vehicle $v$ i the total time consumption in the $r$-th round of iteration can be modified to:

[0119]

[0120] S4: Model screening and uploading: Conduct multi-factor-driven local model screening, and the selected autonomous vehicles will upload the updated local models to the FL server. Among them, local model screening includes reputation-based screening, model convergence-based screening, and maximum utility-based screening.

[0121] During each round of iteration, how to select the participants for model uploading according to the reputation of the autonomous vehicle and the model quality within the time limit, including reputation-based screening, model convergence-based screening, and maximum utility-based screening.

[0122] Specifically, step S4 of this embodiment includes:

[0123] S41: Conduct reputation evaluation on the autonomous vehicles and select multiple local models with high reputation values in each learning iteration.

[0124] Reputation-based screening: Generally speaking, the higher the reputation value of an autonomous vehicle, the greater the possibility of providing high-quality data, so the local models trained thereby have higher accuracy in the FL process. Therefore, the embodiments of the present invention first conduct reputation evaluation on the autonomous vehicles and propose a reputation-based selection strategy. In each learning iteration (such as the $r$-th round), select multiple local models with higher reputation values, where

[0125] Specifically, there should be a threshold standard for participant selection. For autonomous vehicles participating in FL for the first time, the initialized reputation value is set as the current reputation value threshold of the system This reputation value threshold is the manually set initial reputation value. In addition, the reputation value range of the autonomous vehicle is $\eta\in[0,\xi]$, where $\xi$ represents the upper bound of the reputation value. If the reputation value of the autonomous vehicle becomes negative after update, the reputation value of this vehicle is considered 0. For the autonomous vehicle $v$ i, the reputation value update in the r-th round can be expressed as:

[0126]

[0127] Among them, is the reputation value of the current autonomous driving vehicle v i in the previous round, is the comprehensive reputation score obtained after the autonomous driving vehicle v i participates in the last FL iteration, and φ() represents the reputation value update function. Specifically, the reputation scoring requirements are jointly determined by the completion time reliability θ t and the local model quality reliability θ q . Therefore, the embodiments of the present invention comprehensively consider objective factors and adopt a hierarchical strategy to establish reputation evaluation rules.

[0128] Here, only the autonomous driving vehicles participating in the local model upload in the r-th round need to be considered. For non-participating autonomous driving vehicles, the reputation value remains unchanged. For the autonomous driving vehicle v i , the completion time reliability in the (r - 1)-th round can be expressed as:

[0129]

[0130] Among them, and represent the average completion time and the completion time of the autonomous driving vehicle v i . The average completion time represents the average completion time of all autonomous driving vehicles participating in the (r - 1)-th round of federated learning, indicating that the greater the time reliability value, the higher the efficiency of the autonomous driving vehicle v i . The reputation score obtained from the time reliability can be graded as follows:

[0131]

[0132] Among them, ∈ is the weight parameter.

[0133] Next, the similarity between the local model and the global model is expressed as the local model quality reliability The similarity between the local model i of the autonomous driving vehicle v and the global model g (r-1) aggregated in the (r - 1)-th round can be measured by the following formula, using the cosine similarity algorithm:

[0134]

[0135] Among them, The closer to 1, the higher the similarity (higher reliability of the local model quality). Similar to It can also be divided into three levels:

[0136]

[0137] For convenience, the same weight is set for the completion time reliability and the local model quality reliability. Combining the completion time reliability θ t and the local model quality reliability θ q , The calculation rule of

[0138] Therefore, for the autonomous vehicle v i , the update rule of the reputation can be expressed as

[0139]

[0140] S42: Use the similarity to screen multiple local models that meet the reputation value requirements based on model convergence, and select local models that meet the model convergence requirements.

[0141] Screening based on model convergence:

[0142] First, compare whether the signs of the updated parameters of the local model are consistent with the updated parameters in the global model, count the number of parameters with consistent signs, and normalize the result. The obtained ratio shows how many parameters of the two models have consistent signs. Compare this ratio with a pre-set threshold. If the ratio is less than the threshold, it proves that the update of this local model has a very low correlation with the global model update, and this update is deleted; the local model updates with a ratio greater than the threshold are considered models with better quality, and the updates of these models are uploaded to the central server for global model update. The global model update and the local model update can be regarded as two M-dimensional vectors:

[0143] The update of the M model parameters of user (i.e., autonomous vehicle) 1 is expressed as: u1 = <u 11 , u 12 ,..., u 1M >, where u1 represents the set of parameter updates of user 1, with a total of M model parameters, and u 11 represents the first parameter of user 1, and so on.

[0144] The global update of the M model parameters is expressed as: Among them, It represents the set of global model parameter updates. There are a total of M model parameters. u1 represents the average value of the first parameter of all users participating in the upload, and so on.

[0145] The similarity between the updated parameters of the local model and the updated parameters of the global model is expressed as:

[0146]

[0147] Among them, sgn() is the sign function, which is used to compare whether the signs of the i-th parameter updated by the local model and the i-th parameter in the global model update are the same. If the signs of the i-th parameter updated locally and the i-th parameter in the global update are the same, then take 1, otherwise take 0. In this way, the number of parameters with the same sign among the N parameters is obtained, and the ratio result obtained after normalization is the similarity between the local update and the global update.

[0148] The positive and negative signs of the parameters in the update indicate the update direction of the model. The higher the proportion of parameters with the same sign, the higher the consistency between the local model and the global model. However, for two models with the same ratio of parameters with the same sign as the global model, the above method cannot screen out the better one. Considering that even if the number of components with the same sign in two vectors and a third vector is equal, the degree of direction consistency between the two vectors and the third vector is not the same, the cosine similarity algorithm is introduced to perform a secondary screening on the models screened by the above method.

[0149] Cosine similarity is a method to judge the similarity between two vectors by measuring the cosine value of the included angle between the two vectors. The cosine value of the included angle between the two vectors determines whether the two vectors point in roughly the same direction. The cosine similarity between the two M-dimensional vectors of the local update and the global update can be obtained through the following formula:

[0150]

[0151] Preset a cosine similarity threshold V. After obtaining the cosine similarity between the local update parameters and the global update parameters, compare it with the cosine similarity threshold. Those greater than the cosine similarity threshold are considered to have a higher degree of correlation, and those with a cosine similarity less than the cosine similarity threshold are considered to have insufficient correlation, and then this local update is not uploaded.

[0152] S43: Perform screening based on maximizing utility on the selection results based on reputation and the selection results based on the convergence trend.

[0153] Screening based on maximizing utility:

[0154] For the results of reputation-based selection and convergence-trend-based selection, the goal in this selection step is to enable the FL server to aggregate as many local updates as possible within the time threshold for each round. That is, select as many autonomous vehicles ( where N mu ≤N ct ) has the following constraints:

[0155]

[0156] where, v i ∈N ct , represents the time cost of the model aggregation process in the r-th round.

[0157] Since the embodiments of the present invention consider the synchronous FL process, and the time costs of the download process and the model aggregation process can be ignored. Therefore, the constraint can be expressed as:

[0158]

[0159] where v i ∈N ct , represents the execution time of the local model training of autonomous vehicle v i in the r-th round of iteration, represents the time consumed by the upload process of autonomous vehicle v i in the r-th round of iteration, represents the longest time threshold allowed by the FL server.

[0160] S5: Model aggregation: The FL server aggregates the received local models to obtain an improved global model, and the aggregation process is implemented by FedAvg.

[0161] The effect of the CAV decision method applying selective federated learning in the embodiments of the present invention can be further illustrated by the following simulation experiments.

[0162] (1) Simulation conditions

[0163] The embodiments of the present invention use Python for simulation.

[0164] The methods compared in the experiment are as follows:

[0165] One is an autonomous driving federated learning method based on traditional model-free screening, denoted as FedAVG in the experiment, and the reference is: H.B. McMahan, E. Moore, D. Ramage, S. Hampson, and B.A.Y. Arcas, "Communication-efficient learning of deep networks from decentralized data," in Proc. 20th Int. Conf. Artif. Intell. Statist., 2017, pp. 1273-1282.

[0166] The other is an autonomous driving federated learning method based on reducing communication overhead, denoted as CMFL in the experiment, and the reference is: L. Wang, W. Wang, and B. Li, "CMFL: Mitigating Communication Overhead for Federated Learning," in Proc. 39th Int. Conf. Distributed Computing Systems (ICDCS), 2019, pp. 954-964.

[0167] (2) Simulation content

[0168] According to the embodiments of the present invention, calculate the communication overhead under different model accuracies and compare it with the overheads of the existing FedAVG method and CMFL method above. The results are as Figure 2 shown, where Figure 2 the abscissa represents the model accuracy set in advance, and the ordinate represents the proportion of communication overhead.

[0169] This experiment compares the communication overheads of the selective federated learning method (SFRL) proposed in this embodiment, as well as the existing CMFL method and FedAVG method, during the FL process. Figure 2It shows the communication overhead ratios of SFRL, CMFL, and FedAVG when reaching different accuracies. Generally speaking, due to the users' choices, the number of CAVs participating in the FL process decreases, and the accelerated convergence speed leads to a reduction in the number of communication rounds. Therefore, at the same accuracy, the communication overheads of the SFRL method and the existing CMFL method proposed in the embodiments of the present invention are less than that of FedAVG. As the accuracy increases, the proportion of the reduced communication overhead becomes more obvious. Specifically, when the accuracy rates reach 50%, 70%, and 85%, the ratios of the communication overhead of the proposed SFRL method to that of the FedAVG method are 90.2%, 38%, and 19.98% respectively. Correspondingly, the ratios of the communication overhead of the CMFL method to that of the FedAVG method are 92.7%, 57.6%, and 41.9% respectively. It can be seen that the CAV decision method applying selective federated learning in the embodiments of the present invention can further reduce the communication overhead while achieving the same accuracy.

[0170] In the embodiments of the present invention, for the CAV decision method applying selective federated learning, in order to achieve efficient federated learning, a multi-factor-driven model screening strategy is proposed before uploading the local model, which overcomes the defect of the traditional method not differentiating participants, and comprehensively considers the reputation of participants, the quality of local models, and time overhead, so as to select as many high-quality users as possible while considering communication resources and time limitations.

[0171] Another embodiment of the present invention provides a storage medium in which a computer program is stored, and the computer program is used to execute the steps of the CAV decision method applying selective federated learning in the above embodiments. Another aspect of the present invention provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps of the CAV decision method applying selective federated learning as described in the above embodiments are implemented. Specifically, the integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above software functional module is stored in a storage medium and includes several instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs and other various media that can store program codes.

[0172] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. A CAV decision-making method applying selective federated learning, characterized in that Including: S1: Model Download: All autonomous vehicles download the global model updated in the previous iteration from the FL server. S2: Local Training: All autonomous vehicles use their own driving data to train the local model and update the model parameters. The local model takes relative position and relative speed as the state space, acceleration as the action space, and safety and comfort as the reward function. S3: Overhead Calculation: Obtain the communication overhead and computing overhead of all autonomous vehicles during the federated learning process. S4: Model Screening and Upload: Conduct multi-factor-driven local model screening, and the selected autonomous vehicles upload the updated local model to the FL server. Among them, local model screening includes reputation-based screening, model convergence-based screening, and maximization-of-utility-based screening. S5: Model Aggregation: The FL server aggregates the received local models to obtain an improved global model. The state space is represented as: Among them, represents the speed of the target vehicle , position and acceleration ; represents the speed of other traffic participants , position and acceleration ; The action space is represented as: ; The reward function includes a reward function for safety, a reward function for comfort, and a reward function for efficiency. Among them, The reward function for safety is represented as: Among them, represents the speed of the target vehicle, m and α are weight parameters, { Collsion } means that if a collision occurs, the value is 1, otherwise it is 0; The reward function for comfort is represented as: Among them, k and β are weight parameters, represents the jerk of the target vehicle, represents the threshold of acceleration, represents the threshold of jerk, represents the acceleration of the target vehicle; The reward function for efficiency is represented as: Among them, represents the distance between the target vehicle and other traffic participants, represents a threshold for indicating whether the braking distance is appropriate, and δ and l are weighting parameters.

2. The CAV decision-making method applying selective federated learning according to claim 1, wherein, S3 includes: S31: Obtain an autonomous vehicle In the r execution time of local model training in the nth iteration: Among them, represents the data samples participating in federated learning in the r th round of iteration of the autonomous vehicle, represents that the autonomous vehicle contributes computing resources for local model training, represents the number of iterations of local model update when the global accuracy is fixed, represents the number of CPU cycles for executing one data sample; S32: Obtain the time consumed by the uploading process of the autonomous vehicle in the r round iteration: Among them, represents the size of the local model parameters uploaded by the autonomous vehicle in the round, and represents the uplink rate of the data uploaded by the autonomous vehicle in the round. S33: Obtain the total time consumption of the autonomous vehicle in the during the round of iteration: 。 3. The CAV decision-making method applying selective federated learning according to claim 1, characterized in that, S4 includes: S41: Conduct a reputation assessment on autonomous vehicles and select multiple local models with high reputation values in each learning iteration. S42: Use similarity to conduct model convergence-based screening on the selected multiple local models that meet the reputation value requirements, and select local models that meet the model convergence requirements. S43: Conduct maximization-of-utility-based screening on the results of reputation-based selection and convergence-trend-based selection, with the constraint that: Among them, represents the execution time of the local model training of the autonomous driving vehicle in the r round of iteration, represents the time consumed by the autonomous driving vehicle in the r upload process in the round of iteration, represents the longest time threshold allowed by the FL server.

4. The CAV decision method applying selective federated learning according to claim 3, characterized in that Conduct a reputation assessment on autonomous vehicles, specifically including: For autonomous vehicles , the credit value update of the r wheel is expressed as: Among them, is the credit value of the current autonomous vehicle in the previous round, is the comprehensive credit score obtained after the autonomous vehicle participates in the last FL iteration, represents the credit score obtained from the reliability of the completion time and represents the credit score obtained from the reliability of the local model quality .

5. The CAV decision-making method applying selective federated learning according to claim 4, wherein S42 includes: Obtain the similarity between the updated parameters of the local model and the updated parameters of the global model: Among them, sgn() is the sign function, which is used to compare whether the sign of the i th parameter of the local model update is the same as the sign of the i th parameter in the global model update. If the sign of the i th parameter in the local update is the same as the sign of the th parameter in the global update, take 1, otherwise take 0; For two models with the same similarity to the global model, introduce the cosine similarity algorithm to conduct secondary screening on the selected models. The cosine similarity between the parameters of the local model and the global model is represented as: Among them, represents the set of parameter updates of the autonomous driving vehicle, represents the th j parameter of the autonomous driving vehicle, represents the set of global model parameter updates, represents the average value of the j th parameter of all participating uploading users; Preset a cosine similarity threshold. After obtaining the cosine similarity between the local updated parameters and the global updated parameters, compare it with the cosine similarity threshold. The local updated parameters greater than the cosine similarity threshold are retained, and the local updated parameters with a cosine similarity less than the cosine similarity threshold are removed.

6. A storage medium, characterized in that, The storage medium stores a computer program, and the computer program is used to execute the steps of the CAV decision-making method for applying selective federated learning according to any one of claims 1 to 5.

7. An electronic device, characterized in that, Including a memory and a processor. When the processor calls the computer program stored in the memory, it implements the steps of the CAV decision-making method for applying selective federated learning according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Federated deep reinforcement learning-based intelligent decision-making implementation method for automatic driving group vehicle

    CN112348201A

  • Aggregation server selection method for decentralized federated learning

    CN115081002A