An asynchronous federation optimization method to defend against Byzantine attacks

By configuring trusted data sets and asynchronous federal aggregation mechanism in on-board edge computing, the problems of privacy and security and Byzantine attacks during vehicle data upload are solved, and the accuracy and security of the global model are improved.

CN116542342BActive Publication Date: 2025-05-16北京极乘科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310553063.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-16
Publication Date
2025-05-16
Estimated Expiration
2043-05-16

AI Technical Summary

Technical Problem

In the prior art, vehicles have privacy and security problems when uploading local data to roadside units, and some vehicles with large training time will lead to large global aggregation time and may be attacked by Byzantine, affecting the accuracy of the global model.

Method used

Using an asynchronous federated optimization method that can defend against Byzantine attacks, the trusted data set DRSU is configured to the roadside unit, and select the vehicles required for asynchronous federated aggregation, and after the local training of the vehicle is completed, the trusted data set of the roadside unit is used for model comparison to ensure that the vehicle is not subject to malicious attacks before the global model is updated.

Benefits of technology

Effectively filter out vehicles that are maliciously attacked to avoid the impact of global model accuracy, improve the accuracy of global models at roadside units, and prevent malicious tampering of vehicle data by Byzantine attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116542342B_ABST
    Figure CN116542342B_ABST
Patent Text Reader

Abstract

The present invention relates to an asynchronous federated optimization method capable of defending against Byzantine attacks, which includes: configuring a trusted dataset D RSU to roadside units; selecting vehicles required for asynchronous federated aggregation; the selected vehicles download the global model from the roadside units, and the roadside units copy the global model; the selected vehicles use local data to train the downloaded global model to obtain vehicle local models and vehicle loss values L wk ; the roadside units use the trusted dataset D RSU to train the copied global model to obtain roadside local models and roadside loss values L RSU ; the selected vehicles upload the vehicle local models and vehicle loss values L wk to the roadside units; when L wk ≤β R ·L RSU , the vehicle local models and the global model are federated and aggregated to obtain an updated global model; where β R is a preset parameter. The present invention can effectively screen out vehicles under malicious attacks, thereby avoiding the influence on the accuracy of the global model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of vehicle-mounted network, and in particular to an asynchronous federation optimization method capable of defending against Byzantine attacks. Background Art

[0002] In traditional vehicle networks, vehicles send the required computing tasks to the cloud for processing. However, this often results in a large delay. This is not applicable in high-speed moving vehicle scenarios. So vehicle-mounted edge computing was born. In vehicle-mounted edge computing, roadside units with certain computing capabilities can be used as edge terminals to collect and process vehicle data.

[0003] However, when vehicles upload local data to roadside units, privacy and security issues arise, which hinders users from uploading data. Therefore, federated learning came into being. Federated learning allows vehicles to train local models using local data locally, and upload local models instead of original data to roadside units, thereby greatly protecting user privacy. However, vehicles with long training time will result in a long global aggregation time.

[0004] In asynchronous federated learning, each time a roadside unit receives a local model, it performs a global aggregation to update the global model, thereby effectively reducing the aggregation delay. However, since each vehicle may be affected by Byzantine attacks during its own training process, it may maliciously tamper with the data and labels in the data set carried by the vehicle itself, thereby affecting the accuracy of the vehicle's local model, further affecting the update of the global model, and reducing the accuracy of the global model. Summary of the invention

[0005] The technical problem to be solved by the present invention is to overcome the deficiencies in the prior art and provide an asynchronous federated optimization method that can defend against Byzantine attacks, which can effectively screen out vehicles that have been maliciously attacked, thereby avoiding affecting the accuracy of the global model.

[0006] According to the technical solution provided by the present invention, the asynchronous federation optimization method capable of defending against Byzantine attacks includes:

[0007] Configure the trusted dataset D RSU to the roadside unit; select the vehicle required for asynchronous federated aggregation;

[0008] The selected vehicle downloads the global model from the roadside unit, and the roadside unit copies the global model;

[0009] The selected vehicle uses local data to train the downloaded global model to obtain the vehicle local model and vehicle loss value The vehicle local model and the vehicle loss value Upload to the roadside unit; the roadside unit uses the trusted data set D RSU Train the copied global model to obtain the roadside local model and roadside loss value L RSU ;

[0010] When the vehicle local model loss value And the roadside loss value L RSU satisfy , the vehicle local model is aggregated with the global model to obtain an updated global model; otherwise, the vehicle local model is discarded and the step of asynchronous federation aggregation of the required vehicle is returned; wherein, β R are preset parameters.

[0011] In one embodiment of the present invention, a trained global model is obtained after multiple updates. During the training of the global model, multiple selections of vehicles required for asynchronous federated aggregation include:

[0012] Constructing a DDPG model, wherein the DDPG model includes a system reward function;

[0013] Get system status;

[0014] The DDPG model selects actions based on the system state;

[0015] Select the vehicle required for asynchronous federated aggregation based on the selected action;

[0016] The DDPG model is based on the vehicle loss value And the system reward function outputs the reward;

[0017] Return to the step of obtaining the system state until the global model training is completed;

[0018] The system status, actions and rewards form historical data. During the vehicle selection process, the DDPG model is trained based on the historical data.

[0019] In one embodiment of the present invention, the system reward function is:

[0020]

[0021] Where r(t) is the system reward at time slot t, ω 1 and ω 2 is a non-negative weight factor, a di (t) is the system action in time slot t, λ i (t), i∈[1, K] represents the probability of selecting vehicle i, Loss(t) is the vehicle loss value at time slot t, is the delay caused by local training of vehicle i, is the transmission delay of vehicle i uploading the local model at time slot t, a(t) is the system action at time slot t, and s(t) is the system state at time slot t.

[0022] In one embodiment of the present invention, the training delay Determined according to the following formula:

[0023]

[0024] in, is the delay caused by local training of vehicle i, C 0 The number of CPU cycles required to train a data set, μ i is the computing resource of vehicle i, measured by the CPU cycle frequency. Each vehicle i (1≤i≤K) carries a different amount of data D i .

[0025] In one embodiment of the present invention, the transmission delay Determined according to the following formula:

[0026]

[0027]

[0028] d i (t)=||P i (t)-P r ||

[0029] in, is the transmission delay of vehicle i uploading the local model in time slot t, |w| is the size of the local model obtained by local training for each vehicle, tr i (t) is the transmission rate of vehicle i in time slot t, B is the transmission bandwidth, p 0 is the transmission power of each vehicle, a fixed value, h i (t) is the channel gain at time slot t, α is the path loss exponent, σ 2 is the noise power, the position P of vehicle i at time slot t i (t) is set to (d ix (t), d y , 0), where d ix (t) and d y are the distances of vehicle i from the antenna of the roadside unit along the x-axis and y-axis at time slot t, respectively. y is a fixed value, d ix (t) = d i0 +vt,d i0 is the coordinate of the initial position of vehicle i along the x-axis, v is the vehicle speed, t is the time slot, and the antenna height of the roadside unit is set to H r, then the antenna position of the roadside unit is expressed as P r =(0,0,H r ).

[0030] In one embodiment of the present invention, after obtaining the vehicle local model and before uploading the vehicle local model to the roadside unit, the vehicle local model is weight optimized to obtain a weight-optimized local model, taking into account the hysteresis effect of training delay and transmission delay on the local model trained by the vehicle.

[0031] In one embodiment of the present invention, the weight includes a training weight and a transmission weight, and the training weight is:

[0032]

[0033] Among them, β 1,k is the training weight, m 1 ∈(0,1) is a parameter that makes β 1,k As the local training latency increases, it decreases. For vehicle V k Local computing latency;

[0034] The transmission weight is:

[0035]

[0036] Among them, β 2,k (t) is the transmission weight, m 2 ∈(0,1) is a parameter that makes β 2,k (t) decreases as the transmission delay increases, For vehicle V k transmission delay.

[0037] In one embodiment of the present invention, according to the formula w kw =w k *β 1,k *β 2,k , and obtain the vehicle local model after weight optimization, where W k is the local model of the vehicle, W kw is the vehicle local model after weight optimization, β 1,k is the training weight, β 2,k (t) is the transmission weight.

[0038] In one embodiment of the present invention, federation aggregation is performed according to the following formula:

[0039] w new =βw old +(1-β)w kw

[0040] Among them, w old is the current global model at the roadside unit, W new is the updated global model, w kw is the vehicle local model after weight optimization, and β∈(0, 1) is the aggregation ratio.

[0041] In one embodiment of the present invention, based on the system reward at time slot t, the expected long-term discounted reward of the system can be expressed as:

[0042]

[0043] Among them, γ∈(0,1) is the discount factor, N is the total number of time slots, μ is the system strategy, and J(μ) is the expected long-term discounted reward of the system.

[0044] The above technical solution of the present invention has the following advantages compared with the prior art:

[0045] The roadside unit has a clean and reliable dataset, that is, a dataset that will not be maliciously attacked or contaminated, called a trusted dataset D. RSU . When the roadside unit first sends the global model to each vehicle for local training, the roadside unit also uses its own data set to train the roadside local model. When the vehicle local training is completed and the vehicle local model is uploaded, the roadside unit will compare the uploaded vehicle local model with its own trained roadside local model. If the vehicle loss value uploaded by the vehicle is The roadside loss value L of the roadside unit itself training RSU satisfy It is considered that the vehicle has not been attacked maliciously and can participate in the update of the global model. This method can prevent Byzantine attacks from maliciously tampering with the data and labels in the data set carried by the vehicle itself, and improve the accuracy of the global model at the roadside unit. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to make the contents of the present invention more clearly understood, the present invention is further described in detail below based on specific embodiments of the present invention in conjunction with the accompanying drawings.

[0047] Figure 1 is a flow chart of the asynchronous federated optimization method of the present invention;

[0048] Figure 2 This is the accuracy comparison between the proposed scheme and the Byzantine-robust scheme under the Class flip attack method;

[0049] Figure 3 Comparison of the accuracy of the proposed scheme and the Byzantine-robust scheme under the Data flip attack method;

[0050] Figure 4 Comparison of the loss between the proposed scheme and the Byzantine-robust scheme under the Class flip attack method;

[0051] Figure 5 Comparison of the loss between the proposed scheme and the Byzantine-robust scheme under the Data flip attack method;

[0052] Figure 6 The test error rate comparison between the proposed scheme and the Byzantine-robust scheme under the Class flip attack method;

[0053] Figure 7 The test error rate comparison between the proposed scheme and the Byzantine-robust scheme under the Data flip attack method;

[0054] Figure 8 For the testing phase, the loss of our scheme is compared with that of traditional asynchronous federated learning and traditional federated learning in the presence of bad nodes;

[0055] Fig. 9 For the testing phase, under the same node selection conditions, the proposed scheme is compared with the loss of traditional asynchronous federated learning without local weight processing and traditional federated learning. DETAILED DESCRIPTION

[0056] The present invention is further described below in conjunction with the accompanying drawings and specific embodiments so that those skilled in the art can better understand the present invention and implement it, but the embodiments are not intended to limit the present invention.

[0057] Reference Figure 1 As shown, in order to prevent Byzantine attacks from maliciously tampering with data and labels in the data set carried by the vehicle itself and improve the accuracy of the global model at the roadside unit, the present invention includes:

[0058] Configure the trusted dataset D RSU to the roadside unit; select the vehicle required for asynchronous federated aggregation;

[0059] The selected vehicle downloads the global model from the roadside unit, and the roadside unit copies the global model;

[0060] The selected vehicle uses local data to train the downloaded global model to obtain the vehicle local model and vehicle loss value The vehicle local model and the vehicle loss value Upload to the roadside unit; the roadside unit uses the trusted data set D RSU Train the copied global model to obtain the roadside local model and roadside loss value L RSU ;

[0061] When the vehicle local model loss value And the roadside loss value L RSU satisfy , the vehicle local model is aggregated with the global model to obtain an updated global model; otherwise, the vehicle local model is discarded and the step of asynchronous federation aggregation of the required vehicle is returned; wherein, β R are preset parameters.

[0062] Specifically, the roadside unit is equipped with a clean and reliable dataset, that is, a dataset that will not be maliciously attacked or contaminated, called a trusted dataset D RSU . When the roadside unit first sends the global model to each vehicle for local training, the roadside unit also uses its own data set to train the roadside local model. When the vehicle local training is completed and the vehicle local model is uploaded, the roadside unit will compare the uploaded vehicle local model with its own trained roadside local model. If the vehicle loss value L uploaded by the vehicle is wk The roadside loss value L of the roadside unit itself training RSU satisfy It is considered that the vehicle has not been attacked maliciously and can participate in the update of the global model. This method can prevent Byzantine attacks from maliciously tampering with the data and labels in the data set carried by the vehicle itself, and improve the accuracy of the global model at the roadside unit.

[0063] The specific process is as follows: First, the roadside unit initializes the global model as w 0 , the entire training is done by E pi The selected K DDPG The car first downloads the global model. Then it performs local training. DDPG Vehicle in vehicle V Dk ,k∈[1,K DDPG ] as an example. Vehicle V Dk First download the global model, then perform l rounds of local iterations to calculate the vehicle local model. Then calculate w k The loss value At the same time, the vehicle local model w is calculated based on the updated vehicle usage weights. Dkw Then the vehicle loss value and the vehicle local model w Dkw Upload to the roadside unit. The roadside unit calculates the roadside local model w based on the current global model and its own data set RSU And the roadside loss value L RSU , if satisfied If the global model is updated, the global model is updated. Otherwise, no update is performed and the local model and vehicle loss value of the next vehicle training are uploaded. piAfter one round, the roadside unit stops updating the global model and obtains the final global model.

[0064] The detailed algorithm pseudo code is shown in Algorithm 1.

[0065]

[0066]

[0067] Through the above experiments, the method of the present invention has the following conclusions:

[0068] 1. If Figure 2 and Figure 3 As shown in the figure, under class flip attack or data flip attack, the asynchronous federated optimization method of the present invention has higher accuracy in the global model than the existing Byzantine-robust scheme. The Byzantine-robust scheme is referenced in "Huang S, Zhou Y, Wang T, et al. Byzantine-Resilient Federated Machine Learning via Over-the-Air Computation[C]. 2021 IEEE International Conference on Communications Workshops (ICC Workshops), Montreal, QC, Canada, 2021: 1-6.".

[0069] 2. If Figure 4 and Figure 5 As shown, under a Class flip attack or a Data flip attack, the asynchronous federation optimization method of the present invention has lower loss than the existing Byzantine-robustness scheme.

[0070] 3. If Figure 6 and Figure 7 As shown, under the Class flip attack or Data flip attack, the asynchronous federation optimization method of the present invention has a lower test error rate than the existing Byzantine-robustness scheme.

[0071] Furthermore, in order to select the following vehicles with good performance when selecting vehicles, remove the bad nodes that may exist in the vehicles, and obtain the trained global model after multiple updates, during the training of the global model, the vehicles required for asynchronous federated aggregation are selected multiple times, including:

[0072] Constructing a DDPG model, wherein the DDPG model includes a system reward function;

[0073] Get system status;

[0074] The DDPG model selects actions based on the system state;

[0075] Select the vehicle required for asynchronous federated aggregation based on the selected action;

[0076] The DDPG model is based on the vehicle loss value And the system reward function outputs the reward;

[0077] Return to the step of obtaining the system state until the global model training is completed;

[0078] The system status, actions and rewards form historical data. During the vehicle selection process, the DDPG model is trained based on the historical data.

[0079] Specifically, a deep reinforcement learning algorithm is used to select vehicles participating in the training based on the vehicle's own transmission rate, the size of available computing resources, and the vehicle's location. The selected vehicles then use asynchronous federation technology to train the vehicle's local model, which is then uploaded to the roadside unit to ultimately obtain a more accurate global model.

[0080] Since the mobility of a vehicle can be reflected by its position change, the training time and upload time of the vehicle's local model are related to the vehicle's own time-varying available computing resources and the current channel conditions. Therefore, the system state s(t) at time slot t is defined as:

[0081] s(t)=(Tr(t),μ(t),d x (t), a(t-1))

[0082] Where s(t) is the system state at time slot t, Tr(t) represents the set of transmission rates of all vehicles at time slot t, μ(t) is the set of available computing resources of all vehicles at time slot t, and d x (t) is the set of position coordinates of all vehicles along the x-axis at time slot t, and a(t-1) is the system action at time slot t-1.

[0083] Since the purpose of the present invention is to select a better vehicle for asynchronous federated learning training according to the current state, the system action a(t) in time slot t is defined as:

[0084] a(t)=(λ 1 (t), λ 2 (t),…,λ K (t))

[0085] Where a(t) is the system action at time slot t, λ i (t), i∈[1,K] represents the probability of selecting vehicle i, let λ 1 (0) = λ2 (0) = ... = λ K (0)=1.

[0086] The present invention aims to select vehicles with better performance for asynchronous federated training to obtain a more accurate global model at the roadside unit, while considering the delay and the accuracy of the global model. Therefore, the system reward r(t) at time slot t is defined as:

[0087]

[0088] Where r(t) is the system reward at time slot t, ω 1 and ω 2 is a non-negative weight factor, a di (t) is the system action in time slot t, λ i (t), i∈[1, K] represents the probability of selecting vehicle i, Loss(t) is the loss value calculated in asynchronous federated training, is the delay caused by local training of vehicle i, is the transmission delay of vehicle i uploading the local model in time slot t.

[0089] Then the expected long-term discounted reward of the system can be expressed as:

[0090]

[0091] Among them, γ∈(0,1) is the discount factor, N is the total number of time slots, μ is the system strategy, and J(μ) is the expected long-term discounted reward of the system.

[0092] To select a specific vehicle, let set a d (t)=(a d1 (t), a d2 (t),…,a dK (t)), i (t) is normalized and λ is set i (t) values ​​greater than or equal to 0.5 correspond to a di (t) is recorded as 1, otherwise it is 0. The final set a d (t) is composed of 0 and 1, 1 means selecting a vehicle, and 0 means not selecting a vehicle.

[0093] The selected vehicle uses local data for local training to obtain the corresponding local model, including the following steps:

[0094] S1: At time slot t, vehicle V k Download the global model w from the roadside unit t-1 , where, at time slot 1, the global model at the roadside unit is initialized to w using a convolutional neural network 0 ;

[0095] S2: Vehicle V k The local data is trained based on the convolutional neural network, and the local training consists of l rounds. In the mth (m∈[1, l]) round of local training, the vehicle V k First, the label probability of each local data a, i.e., y a Input to the local model w k,m The convolutional neural network is then used to obtain the predicted probability of the convolutional neural network for each data label. The cross entropy loss function is used to calculate w k,m The loss value is calculated as follows:

[0096]

[0097] S3: Update the local model using the stochastic gradient descent algorithm. The formula is as follows:

[0098]

[0099] in, f k (w k,m ) is the gradient, η is the learning rate;

[0100] S4: Vehicle V k Use the updated local model to perform m+1 rounds of local training. When the number of local training rounds reaches l, the local training stops and the vehicle obtains the updated local model W. k .

[0101] Furthermore, when the vehicle is performing local training, training delay and transmission delay will occur. The training delay is:

[0102]

[0103] in, is the delay caused by local training of vehicle i, C 0 The number of CPU cycles required to train a data set, μ i is the computing resource of vehicle i, measured by the CPU cycle frequency. Each vehicle i (1≤i≤K) carries a different amount of data D i ;

[0104] The transmission delay is:

[0105]

[0106]

[0107] d i (t)=||P i(t)-P r ||

[0108] in, is the transmission delay of vehicle i uploading the local model in time slot t, |w| is the size of the local model obtained by local training for each vehicle, tr i (t) is the transmission rate of vehicle i in time slot t, B is the transmission bandwidth, p 0 is the transmission power of each vehicle, a fixed value, h i (t) is the channel gain at time slot t, α is the path loss exponent, σ 2 is the noise power, the position P of vehicle i at time slot t i (t) is set to (d ix (t), d y , 0), where d ix (t) and d y are the distances of vehicle i from the antenna of the roadside unit along the x-axis and y-axis at time slot t, respectively. y is a fixed value, d ix (t) = d i0 +vt,d i0 is the coordinate of the initial position of vehicle i along the x-axis, v is the vehicle speed, t is the time slot, and the antenna height of the roadside unit is set to H r , then the antenna position of the roadside unit is expressed as P r =(0,0,H r ).

[0109] Among them, the autoregressive model is used to construct h i (t) and h i (t-1), namely:

[0110]

[0111] Among them, ρ i is the normalized channel correlation coefficient between consecutive time slots, e(t) is the error vector that follows a complex Gaussian distribution and is related to h i (t) related, according to Jack's fading spectrum, Among them J 0 (·) is a zero-order Bessel function of the first kind and is the Doppler frequency of vehicle i Λ is the wavelength, θ is the moving direction, that is, x 0 =(1,0,0) and the uplink communication direction is P r -P i (t), so

[0112] Unlike traditional asynchronous federated learning, the present invention takes into account the hysteresis effect of training delay and transmission delay on the local model trained by the vehicle. Specifically, since there will be a certain delay in the local training of the vehicle and uploading of the vehicle local model to the roadside unit, when a vehicle is in the process of local training and uploading to the roadside unit, it is possible that the roadside unit has received the vehicle local model uploaded by other vehicles and updated the global model. In this case, the vehicle local model trained by this vehicle has a certain hysteresis. Therefore, the present invention performs certain weight processing on the vehicle local model of vehicle Vk, that is, setting training weights and transmission weights. The specific calculation method is as follows:

[0113] The vehicle local model is weight optimized, and the weight includes training weight and transmission weight. The training weight is:

[0114]

[0115] Among them, β 1,k is the training weight, m 1 ∈(0,1) is a parameter that makes β 1,k As the local training latency increases, it decreases. For vehicle V k Local computing latency;

[0116] The transmission weight is:

[0117]

[0118] Among them, β 2,k (t) is the transmission weight, m 2 ∈(0,1) is a parameter that makes β 2,k (t) decreases as the transmission delay increases, For vehicle V k transmission delay;

[0119] According to the formula w kw =w k *β 1,k *β 2,k , get the vehicle local model after weight optimization;

[0120] Among them, W k is the local model of the vehicle, w kw is the vehicle local model after weight optimization, β 1,k is the training weight, β 2,k (t) is the transmission weight.

[0121] Furthermore, the trained vehicle asynchronously uploads the weight-optimized vehicle local model to the roadside unit for asynchronous federated aggregation. After multiple rounds of repeated training, the roadside unit finally obtains the global model, which specifically includes:

[0122] When the vehicle V k After uploading the weight-optimized vehicle local model to the roadside unit, the roadside unit performs a global aggregation, and the formula is as follows:

[0123] W new =βw old +(1-β)w kw

[0124] Among them, w old is the current global model at the roadside unit, W new , is the updated global model, w kw , is the vehicle local model after weight optimization, β∈(0,1) is the aggregation ratio;

[0125] At the beginning of each time slot, when the roadside unit receives the first uploaded local model, w old =w t-1 , when the roadside unit receives the local models of all selected vehicles and gets the updated K 1 The global model w t After that, the global model update of this time slot ends.

[0126] At the same time, the average loss Loss(t) of the vehicles participating in the training can be obtained, which can be expressed as:

[0127]

[0128] Among them, f k (w k ) is the local model w k The loss value.

[0129] In order to further illustrate the principles and beneficial effects of the present invention, a description is given below in conjunction with specific experiments.

[0130] The present invention aims to find an optimal strategy μ * to maximize the expected long-term discounted reward of the system.

[0131] The overall algorithm specifically adopted in the present invention includes two parts, an algorithm of a training phase based on a DAFL (Data-Free Learning) framework and an algorithm of a testing phase based on the DAFL framework.

[0132] The algorithm steps of the training phase based on the DAFL framework are shown in Table 1.

[0133] Table 1

[0134]

[0135] The present invention uses the DDPG algorithm to optimize the asynchronous federation method, wherein the DDPG algorithm is based on the actor-critic network architecture. The actor network is used to perform policy improvement, and the critic network is used to perform policy evaluation. Specifically, the actor network is used to approximate the policy μ, and its approximate policy is represented as μ δ The actor network is based on the policy μ δ And observe the state and output actions.

[0136] The present invention improves and evaluates the strategy through iteration to finally obtain the optimal strategy. In order to ensure the stability of the algorithm, the DDPG algorithm also uses a target network composed of a target actor network and a target critic network, and its architecture is the same as that of the actor network and the critic network.

[0137] Set δ as the actor network parameter, ξ as the critic network parameter, δ * is the optimized actor network parameter, ξ * is the optimized critic network parameter, δ 1 is the target actor network parameter, ξ 1 is the target critic network parameter. τ is the update parameter of the target network, Δ t is the noise of action exploration in time slot t. I is the mini-batch size. Next, the algorithm of the training phase will be introduced in detail.

[0138] First, randomly initialize δ and ξ, and at the same time replace δ in the target network with 1 and 1 Initialize to δ and ξ respectively. At the same time, the experience playback buffer R b Initialize.

[0139] Next, the algorithm will execute E max In the first round, the positions of all vehicles, channel status, and available computing resources of the vehicles are reset. And set λ 1 (0) = λ 2 (0) = ... = λ K (0) = 1, then in the first time slot, the system can obtain the initial state s(1) = (Tr(1), μ(1), d x(1), a(0)). At the same time, CNN (Convolutional Neural Networks) is used to initialize the global model w at the roadside unit. 0 .

[0140] After that, the algorithm will be executed continuously from time slot 1 to the maximum number of time slots N. In the first time slot, the actor network gets the output μ according to the state δ (s|δ), where a random noise Δ is added to the action t , so the system gets action a(1) = μ δ (s(1)|δ)+Δ t . Then calculate a according to the action d (1), determine the vehicle selected for this time slot. The selected vehicle performs asynchronous federated training, that is, the vehicle trains a local model based on local data, and then asynchronously uploads it to the roadside unit to update the global model, and then calculates the loss value Loss(1). At the same time, the local training delay and transmission delay of the vehicle are calculated, so that the system reward under time slot 1 can be obtained. Then, the vehicle position is updated, the channel conditions and the vehicle's own available computing resources are recalculated, and the vehicle's transmission rate is updated so that the system can observe the next state s(2). The tuple (s(1), a(1), r(1), s(2)) is then stored in R b middle.

[0141] When R b When the number of tuples in is less than or equal to I, the system directly inputs the next state into the actor network and proceeds to the next iteration.

[0142] When R b When the number of tuples in is greater than I, the parameters δ, ξ, δ in the actor network, critic network, and target network are 1 and 1 Start updating to maximize J(μ δ ). The parameters δ of the actor network are moving towards J(μ δ ) is the gradient direction of Update. Will follow the strategy μ under s(t) and a(t) δ The action value function is set to Q μδ (s(t), a(t)), its expression is:

[0143]

[0144] It represents the long-term expected discounted reward of the system in time slot t.

[0145] Solution By solving Q uδThe gradient of (s(t), a(t)) Instead, the critic network uses the parameter ξ to μδ (s(t), a(t)) is approximately Q ξ (s(t), a(t)).

[0146] Next, we will introduce the parameters δ, ξ, δ under time slot t. 1 and 1 The update method. When R b When the number of tuples in R is greater than I, the system starts from R b Randomly select I tuples from to form a mini-batch. Let (s x , a x , r x , s′ x ), x∈[1, 2, ..., I] is the xth tuple in the mini-batch. Then the system first converts s′ x Input the target actor network to get the output action Then s′ x and a′ x Input the target critic network and get the output action value function The target value can then be calculated as:

[0147]

[0148] Then, according to s x and a x , the critic network will have an output Q ξ (s x , a x ), then the loss of tuple x can be calculated as:

[0149] L x =[y x -Q ξ (s x , a x )] 2

[0150] When all tuples are input into the critic network and the target network, the loss function is obtained:

[0151]

[0152] The critic network is constructed by Use the gradient descent method to minimize the loss function L(ξ) and update the parameter ξ.

[0153] Similarly, actor networks are implemented by Use the gradient ascent method to maximize J(μ δ ) to update the parameter δ. The action value function is calculated by approximating the critic network, and the formula is as follows:

[0154]

[0155] Where Q ξ The input is

[0156] At the end of time slot t, update the parameters of the target network, and the update formula is:

[0157] ξ 1 ←τξ+(1-τ)ξ 1

[0158] δ 1 ←τδ+(1-τ)δ 1

[0159] Where τ is a constant and satisfies τ<<1.

[0160] Finally, the system inputs s′ into the actor network and starts the iterative calculation of the next time slot. When the time slot t reaches the maximum value N, the round ends. Then the system reinitializes the state value s(1) = (Tr(1), μ(1), d x (1), a(0)), and proceed to the next round of training. When the number of rounds reaches the maximum value E max The training ends when , and the parameters of the optimized actor network, critic network, target actor network and target critic network are obtained, namely δ * , * , and

[0161] The test phase simulates the critic network, target actor network, and target critic network in the training phase. And uses the optimal parameter δ * The optimal strategy.

[0162] The algorithm steps of the testing phase based on the DAFL framework are shown in Table 2.

[0163] Table 2

[0164] <![CDATA[1. For each episode 1 ≤ epi ≤ E′ max Execute:]]> 2. Reset the simulation parameters of the system model and initialize the global model at the roadside unit 3. Get the initial state s(1) 4. For each time slot 1≤t≤N, execute: <![CDATA[5. Generate the action a = μ δ (s|δ)]]> <![CDATA[6. Calculate a d , and determine the selected vehicle]]> 7. The selected vehicle undergoes weight-based AFL update training 8. Obtain reward r and next state s′ from the current system

[0165] The present invention sets time slots according to the vehicle's own transmission rate, available computing resource size and vehicle position, selects vehicles participating in training, and removes possible bad nodes in the vehicle; the selected vehicle uses local data for local training to obtain a corresponding local model, and when the vehicle performs local model training, the hysteresis effect caused by the training delay and transmission delay on the local model trained by the vehicle is considered, and the weight of the local model is optimized, thereby improving the accuracy of the global model at the roadside unit; the trained vehicle asynchronously uploads the weight-optimized local model to the roadside unit for asynchronous federated aggregation, and through multiple rounds of repeated training, the roadside unit finally obtains the global model. The vehicle of the present invention adopts asynchronous federated training, and the roadside unit aggregates the global model once each time it receives a local model uploaded from the vehicle, which can update the global model at the roadside unit faster without waiting for the upload of other vehicles. The method of the present invention is simple to calculate, and the system model is reasonable. Simulation experiments verify that the method can obtain a higher global model accuracy in a vehicle environment.

[0166] Obviously, the above embodiments are merely examples for clear explanation and are not intended to limit the implementation methods. For those skilled in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to list all the implementation methods here. The obvious changes or modifications derived from these are still within the protection scope of the invention.

Claims

1. An asynchronous federation optimization method capable of defending against Byzantine attacks, characterized in that: include: Configuring a trusted dataset to the roadside unit; select the vehicle required for asynchronous federated aggregation; The selected vehicle downloads the global model from the roadside unit, and the roadside unit copies the global model; The selected vehicle uses local data to train the downloaded global model to obtain the vehicle local model and vehicle loss value , and the vehicle local model and vehicle loss value Upload to a roadside unit; the roadside unit uses the trusted data set Train the copied global model to obtain the roadside local model and roadside loss value ; When the vehicle local model loss value and roadside loss value satisfy When the local vehicle model is federated with the global model to obtain an updated global model, otherwise, the local vehicle model is discarded and the step of asynchronous federation aggregation of the required vehicle is returned; wherein, are preset parameters.

2. The asynchronous federation optimization method capable of defending against Byzantine attacks according to claim 1, characterized in that: After multiple updates, the trained global model is obtained. During the training of the global model, the vehicles required for asynchronous federated aggregation are selected multiple times, including: Constructing a DDPG model, wherein the DDPG model includes a system reward function; Get system status; The DDPG model selects actions based on the system state; Select the vehicle required for asynchronous federated aggregation based on the selected action; The DDPG model is based on the vehicle loss value And the system reward function outputs the reward; Return to the step of obtaining the system state until the global model training is completed; The system status, actions and rewards form historical data. During the vehicle selection process, the DDPG model is trained based on the historical data.

3. The asynchronous federation optimization method capable of defending against Byzantine attacks according to claim 2, characterized in that: The system reward function is: , in, is the system reward for time slot t, and is a non-negative weight factor, is the system action in time slot t, represents the probability of selecting vehicle i, is the vehicle loss value at time slot t, is the delay caused by local training of vehicle i, is the transmission delay of vehicle i uploading the local model in time slot t.

4. The asynchronous federation optimization method capable of defending against Byzantine attacks according to claim 3, characterized in that: Based on the system reward at time slot t, the expected long-term discounted reward of the system can be expressed as: , in, is the discount factor, is the total number of time slots, For the system strategy, is the expected long-term discounted reward for the system.

5. The asynchronous federation optimization method capable of defending against Byzantine attacks according to claim 1, characterized in that: After obtaining the vehicle local model, before uploading the vehicle local model to the roadside unit, the vehicle local model is weight optimized to obtain the weight-optimized local model, taking into account the hysteresis effect of training delay and transmission delay on the local model trained by the vehicle.

6. The asynchronous federation optimization method capable of defending against Byzantine attacks according to claim 5, characterized in that: The training delay Determined according to the following formula: , in, is the delay caused by local training of vehicle i, The number of CPU cycles required to train a data set, is the computing resource of vehicle i, measured by CPU cycle frequency, and each vehicle They carry different amounts of data .

7. The asynchronous federation optimization method capable of defending against Byzantine attacks according to claim 5, characterized in that: The transmission delay Determined according to the following formula: , in, For vehicles In time slot The transmission delay of uploading the local model, The size of the local model trained locally for each vehicle, for Time slot vehicle The transmission rate is, B is the transmission bandwidth, is the transmission power of each vehicle, is a fixed value, for The channel gain of the time slot, is the path loss exponent, is the noise power, vehicle In time slot Location Set to ,in and are the distances of vehicle i from the antenna of the roadside unit along the x-axis and y-axis at time slot t, respectively. is a fixed value, , is the coordinate of the initial position of vehicle i along the x-axis, v is the vehicle speed, and the antenna height of the roadside unit is set to , then the antenna position of the roadside unit is expressed as .

8. The asynchronous federation optimization method capable of defending against Byzantine attacks according to claim 5, characterized in that: The weights include training weights and transmission weights, and the training weights are: , in, is the training weight, As a parameter, it makes As the local training latency increases, it decreases. For vehicles Local computing latency; The transmission weight is: , in, is the transmission weight, As a parameter, it makes As the transmission delay increases, it decreases. For vehicles transmission delay.

9. The asynchronous federation optimization method capable of defending against Byzantine attacks according to claim 5, characterized in that: According to the formula , and obtain the vehicle local model after weight optimization, where is the local model of the vehicle, is the vehicle local model after weight optimization, is the training weight, is the transmission weight.

10. The asynchronous federation optimization method capable of defending against Byzantine attacks according to claim 1, characterized in that: Federated aggregation is performed according to the following formula: , in, is the current global model at the roadside unit, is the updated global model, is the vehicle local model after weight optimization, is the polymerization ratio.

Citation Information

Patent Citations

  • Federated learning toilet vehicle attack defense method based on block chain

    CN112714106A

  • Asynchronous federated optimization method for selecting vehicles based on DDPG algorithm

    CN116055489A