Internet of vehicles data dynamic protection method and system in federated learning process
By building a vehicle dynamic position set and a digital model of the Internet of Vehicles, combining differential privacy and graph neural networks, the privacy leakage and model training difficulties of data sharing in the Internet of Vehicles are solved, efficient data protection and model accuracy are achieved, malicious attacks are prevented, and the security and reliability of the Internet of Vehicles system are ensured.
Patent Information
- Application Number
- CN202510483204.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-01
AI Technical Summary
There are risks of privacy leakage during data sharing in the Internet of Vehicles and difficulty in training models under complex topology structures. The existing security protection mechanism is difficult to deal with complex attacks, affecting data value mining and model accuracy.
A vehicle dynamic position set and vehicle network digital model is built, differential privacy protection and graph neural network are adopted, and a hyperellipsoid equation and reputation algorithm are combined to achieve dual-layer protection of dynamic anonymity and model parameters, and screen trusted vehicles to participate in federated learning.
It realizes effective protection of vehicle user privacy during data sharing, improves model accuracy and adaptability, prevents malicious attacks, and ensures the security and reliability of Internet of Vehicle data sharing.
Smart Images

Figure CN120416804A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle data security, and in particular, to a method and system for dynamically protecting vehicle networking data during the federated learning process. Background Art
[0002] There are a large number of data interaction requirements among vehicles, roadside units (RSUs), and cloud servers in the vehicle networking to support applications such as intelligent traffic management and autonomous driving assistance. However, the data generated by vehicles contains a large amount of sensitive information, such as vehicle location, driving trajectory, and user identity. Traditional data sharing methods are difficult to effectively protect user privacy while ensuring efficient data sharing, and the risk of privacy leakage seriously hinders the full exploration of the value of vehicle networking data.
[0003] In addition, the vehicle networking has a complex topological structure, and there are differences in communication links and data distributions among different vehicles and devices. When performing data-driven model training, centralized training faces problems such as high data transmission costs and high privacy exposure risks; in distributed training, how to coordinate the training processes of different nodes and ensure the accuracy and convergence of the model are problems that need to be solved urgently. The vehicle networking faces various security threats. Malicious vehicle users may deliberately upload incorrect model parameters to interfere with the federated learning process, and attackers may also reverse-engineer user privacy information by analyzing model parameters. Existing security protection mechanisms are difficult to comprehensively handle these complex attack means and cannot guarantee the security and reliability of vehicle networking data sharing. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for dynamically protecting vehicle networking data during the federated learning process, aiming to solve at least one problem in the background art.
[0005] In a first aspect, the present invention provides a method for dynamically protecting vehicle networking data during the federated learning process, the method comprising:
[0006] Constructing a vehicle dynamic position set, the vehicle dynamic position set at least including the position coordinates, driving direction, and speed of the vehicle at different times;
[0007] Building a vehicle networking digital model, the vehicle networking digital model including a plurality of node models, each node model corresponding to a vehicle, and training the node models under the same vehicle according to the vehicle dynamic position set to obtain the trained node models;
[0008] After performing differential privacy protection on the first model parameters of the trained node model, upload them to the corresponding roadside unit, so that the roadside unit aggregates the first model parameters once to obtain the first aggregated parameter, and perform differential privacy protection on the second model parameters under all roadside units and upload them to the server, so that the server performs secondary aggregation on the second model parameters under all roadside units to obtain the second aggregated parameter, and construct an aggregated global model according to the second aggregated parameter;
[0009] Intercept a set of position points within a preset time period covering the current moment from the dynamic position set, and calculate the geometric center of the set of position points, so as to construct a super-ellipsoid equation according to the geometric center of gravity, and obtain a number of virtual positions corresponding to the true position at the current moment according to the super-ellipsoid equation, and combine the true position and the number of virtual positions into a target information set;
[0010] Screen out target vehicles with a reputation higher than the first preset threshold and not identified as an attack according to the pre-trained attack recognition model, and upload the target information set of the target vehicles to perform federated learning on the aggregated global model.
[0011] In a second aspect, the present invention provides a vehicle networking data dynamic protection system during the federated learning process. The system includes:
[0012] A position set construction module, configured to construct a vehicle dynamic position set, where the vehicle dynamic position set at least includes the position coordinates, driving direction, and speed of the vehicle at different times;
[0013] A digital model building module, configured to build a vehicle networking digital model, where the vehicle networking digital model includes a plurality of node models, each node model corresponds to a vehicle, and train the node models under the same vehicle according to the vehicle dynamic position set to obtain trained node models;
[0014] A model parameter upload module, configured to perform differential privacy protection on the first model parameters of the trained node model and upload them to the corresponding roadside unit, so that the roadside unit aggregates the first model parameters once to obtain the first aggregated parameter, and perform differential privacy protection on the second model parameters under all roadside units and upload them to the server, so that the server performs secondary aggregation on the second model parameters under all roadside units to obtain the second aggregated parameter, and construct an aggregated global model according to the second aggregated parameter;
[0015] The super-ellipsoid equation construction module is used to intercept the position point set within a preset time period covering the current moment from the dynamic position set, calculate the geometric center of the position point set, construct a super-ellipsoid equation based on the geometric center of gravity, and obtain several virtual positions corresponding to the true position at the current moment according to the super-ellipsoid equation, and combine the true position and several virtual positions into a target information set;
[0016] The federated learning module is used to screen out target vehicles with a reputation higher than the first preset threshold and not identified as an attack according to the pre-trained attack recognition model, and upload the target information set of the target vehicle to perform federated learning on the aggregated global model.
[0017] In a third aspect, the present invention provides a storage medium that stores one or more programs, and when the program is executed by a processor, it implements the vehicle networking data dynamic protection method in the above-mentioned federated learning process.
[0018] In a fourth aspect, the present invention provides an electronic device, which includes a memory and a processor, wherein:
[0019] The memory is used to store computer programs;
[0020] When the processor is used to execute the computer program stored on the memory, it implements the vehicle networking data dynamic protection method in the above-mentioned federated learning process.
[0021] Compared with the prior art, the present invention has the following advantages:
[0022] 1. By constructing a vehicle networking digital model, digital modeling of physical entities and the establishment of a vehicle dynamic position set are realized, providing a comprehensive and accurate data basis for data sharing. Based on the two-layer structure data sharing method of federated learning, combined with the data and features provided by the digital model, while realizing efficient data sharing, dynamic anonymous privacy protection and differential privacy technology are used to protect the privacy of vehicle users from two levels of position information and model parameters, effectively balancing the relationship between data sharing and privacy protection.
[0023] 2. In the construction of the vehicle networking digital model, the use of graph neural networks (GNNs) and spatio-temporal attention mechanisms can effectively learn the complex relationships between nodes in the vehicle networking, accurately capture spatio-temporal change features, and improve the response ability of the model to the dynamic changes in the vehicle networking. The aggregation optimization in the federated learning process and the model protection based on differential privacy consider the importance of model parameters and the spatio-temporal features of different regions, improving the accuracy and adaptability of the global model, making it more in line with the complex and changeable actual environment of the vehicle networking.
[0024] 3. By constructing a vehicle user reputation algorithm and determining reputation evaluation indicators in combination with the vehicle location privacy protection situation and the stability of model parameters, the credibility of vehicle users can be effectively evaluated. Based on the machine learning-based attack recognition algorithm and the adaptive vehicle user selection strategy, potential attack behaviors can be identified, trustworthy vehicle users can be selected to participate in federated learning, vehicle users can be encouraged to maintain good behaviors, malicious attacks on the vehicle networking system can be prevented, and the security and reliability of vehicle networking data sharing can be comprehensively improved. Brief Description of the Drawings
[0025] Figure 1 It is a flowchart of a method for dynamically protecting vehicle networking data during the federated learning process proposed in an embodiment of the present invention;
[0026] Figure 2 It is a schematic structural diagram of a system for dynamically protecting vehicle networking data during the federated learning process proposed in an embodiment of the present invention.
[0027] The following specific embodiments will further illustrate the present invention in conjunction with the above-mentioned drawings. Specific Embodiments
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention. Unless otherwise defined, the technical terms or scientific terms used herein shall have the ordinary meanings understood by those of ordinary skill in the art in the field to which the present invention belongs. The words such as "including" used herein mean that the elements or items appearing before this word cover the elements or items listed after this word and their equivalents, without excluding other elements or items.
[0029] As Figure 1 shown, an embodiment of the present invention proposes a method for dynamically protecting vehicle networking data during the federated learning process. The method includes steps S101 to S105, where:
[0030] Step S101: Construct a vehicle dynamic location set, where the vehicle dynamic location set at least includes the location coordinates, driving direction, and speed of the vehicle at different times;
[0031] It should be noted that in the digital modeling of physical entities, various physical entities in the vehicle network, such as vehicles, roadside units (RSUs), and cloud servers, all need to undergo in-depth digital modeling. In the construction of the vehicle digital model, multi-sensor fusion technology is introduced to comprehensively perceive the operating state of the vehicle. Suppose the vehicle is equipped with an acceleration sensor, a gyroscope sensor, and a GPS positioning sensor, etc. The extended Kalman filter algorithm (EKF) is used to fuse and process the multi-source sensor data. Taking the vehicle's position, speed, and acceleration as state variables, the state equation is constructed as follows:
[0032] X k = F(X k-1 , u k-1 ) + Q k-1 ;
[0033] where X k is the state vector at time k, including the vehicle position x k , speed v k , and acceleration a k , that is, X k = [x k , v k , a k T , F(X k-1 , u k-1 ) is the state transition function, considering the vehicle's dynamic characteristics and control input u k-1 , Q k-1 is the process noise, and the measurement equation can be expressed as:
[0034] Z k = H(X k ) + R k ;
[0035] where Z k is the sensor measurement value vector, H(X k ) is the measurement function, R k is the measurement noise that maps the state variable to the measurement space. By continuously iterating and updating the state estimation through the EKF algorithm, the operating state of the vehicle at different times can be accurately obtained, providing a reliable data basis for subsequent vehicle network applications.
[0036] Establish a vehicle dynamic position set. Construct the vehicle dynamic position set DLD, which not only includes the position coordinates of the vehicle at different times but also incorporates dynamic information such as the driving direction and speed of the vehicle. Suppose the vehicle is at time series t1, …, t i , …, t n times, and the position coordinates at time t i are (x i , y i ), and the driving direction is Ai with a speed of v i the dynamic location set DLD can be expressed as:
[0037] DLD = {(x i , y i , A i , v i , t i ) | i = 1, 2,..., n}.
[0038] Step S102: Build a vehicle networking digital model. The vehicle networking digital model includes multiple node models, each node model corresponding to a vehicle, and train the node models under the same vehicle according to the vehicle dynamic location set to obtain the trained node models;
[0039] It should be noted that in constructing the vehicle networking digital model, when building the vehicle networking digital model, the graph neural network (GNN) is used to model the complex topological structure in the vehicle networking. Vehicles, RSUs, cloud servers, etc. in the vehicle networking can be regarded as nodes in the graph, and the communication links between them are edges. For each node a i , its feature vector h i contains node information such as the position, speed, and driving direction of the vehicle, the coverage range and signal strength of the RSU, and the processing capacity and storage capacity of the cloud server. Through graph convolution operations, the updated feature of node a i is calculated according to the following formula
[0040]
[0041] where N(i) is the set of neighbor nodes of node a i , d i and d j are the degrees of node a i and a j respectively, W is a learnable weight matrix, b is a bias term, and σ is an activation function.
[0042] By stacking multiple graph convolutional networks, the complex relationships between nodes in the vehicle networking can be effectively learned, and the accurate perception and analysis of the overall state of the vehicle networking can be realized. At the same time, using the spatio-temporal attention mechanism, the information at different times and different spatial positions is weighted and fused to highlight key information and improve the response ability of the model to the dynamic changes of the vehicle networking. For example, using the spatio-temporal attention mechanism, the node information at different times and different spatial positions is weighted and fused, where the attention weight in the time dimension and the attention weight in the spatial dimension are expressed as follows:
[0043]
[0044] Among them, exp(f t (t i )) performs an exponential transformation on the fractional function value corresponding to time t i . is to sum up the exponentially transformed fractional values corresponding to all times. exp(f s d(x i , y i )) maps the original score to the positive space. ∑ j∈N(i) exp(f s (x j , y j )) is to sum up the exponential scores of all neighbors j ∈ N(i) of node a i . By weighted fusion of time and space information, the spatio-temporal change characteristics in the vehicle network can be captured more accurately, providing strong support for the efficient operation of the vehicle network.
[0045] In addition, in some embodiments, regarding the training of the local vehicle node model, based on the completion of the construction of the vehicle network digital model, the data sharing is realized by using the federated learning mechanism. The vehicle, as a local node, conducts local model training based on the data collected by sensors and processed by multi-sensor fusion during the construction of the vehicle network digital twin model. Taking the neural network model in deep learning as an example, assume that the vehicle uses a multi-layer perceptron (MLP) for model training, and its loss function adopts the cross-entropy loss function. For an MLP with L hidden layers, the output a l of the l-th layer is calculated through forward propagation:
[0046] a l = σ(W l a l-1 + b l );
[0047] Among them, W l is the weight matrix of the l-th layer, b l is the bias vector, and a l-1 is the output of the (l - 1)-th layer;
[0048] Let the training data set of the vehicle be D i , which contains N samples. For the sample (x n , y n ) ∈ D i , the cross-entropy loss function is:
[0049]
[0050] Among them, C is the number of categories, y nkis the one - hot encoded value of the true label of sample n in the k - th class, is the predicted value of the model output in the k - th class;
[0051] Update the model parameters of the node model according to the following formula:
[0052]
[0053] where, are the first model parameters of the node model corresponding to vehicle i at time t + 1 and time t respectively, η is the learning rate, is the loss function L i with respect to gradient.
[0054] Step S103: After performing differential privacy protection on the first model parameters of the trained node model, upload them to the corresponding roadside unit, so that the roadside unit aggregates the first model parameters once to obtain the first aggregated parameter, and after performing differential privacy protection on the second model parameters under all roadside units, upload them to the server, so that the server aggregates the second model parameters under all roadside units twice to obtain the second aggregated parameter, and construct an aggregated global model according to the second aggregated parameter;
[0055] It should be noted that the aggregated global model is the general term for the node model after updating the second aggregated parameter.
[0056] After the vehicle completes local training, it uploads the trained model parameters to the nearby roadside unit (RSU). The RSU collects the model parameters from multiple vehicles and performs federated average aggregation.
[0057] To ensure privacy during data upload, differential privacy protection is adopted when uploading to the roadside unit or cloud server. Specifically, the first model parameters are clipped, and noise is added to the clipped first model parameters to obtain the first model parameters with added noise, and the first model parameters with added noise are uploaded to the roadside unit;
[0058] The roadside unit restores the first model parameters with added noise to obtain the first model parameters, and performs the first aggregation according to the following formula:
[0059]
[0060] where, θ RSU is the first aggregated parameter, w ij is the adjustment factor, n i is the data volume. If vehicle i trains the local model (node) based on 1000 trajectory data (speed, position, etc.) collected by its sensors, then n iis one thousand, θ i is the first model parameter corresponding to vehicle i, M is the number of vehicles participating in the training within the RSU coverage, λ is the adjustment parameter, sim ij is the cosine similarity. Introducing a weight adjustment factor based on cosine similarity is to further optimize the aggregation process and consider the difference degree of model parameters, θ j is the first model parameter of the node model of vehicle j;
[0061] It should be noted that the differential privacy protection method of the first model parameter is exactly the same as that of the first aggregation parameter. Therefore, it will not be repeated in this embodiment. Then, the cloud server restores the first aggregation parameter after adding noise, and then performs secondary aggregation according to the following formula:
[0062]
[0063] Among them, K is the number of participating RSUs, θ RSU,k is the first aggregation parameter aggregated by the kth RSU, N k is the total data volume within the cloud server coverage, θ global is the second aggregation parameter, w k is the weight.
[0064] In addition, in some embodiments, in the data upload and aggregation from the RSU to the cloud server, the RSU uploads the aggregated model parameters to the cloud server. The cloud server also uses the federated average aggregation method to re-aggregate the model parameters from multiple RSUs. At the same time, the cloud server can use the spatio-temporal attention mechanism in the vehicle network digital model constructed in the construction of the vehicle network digital twin model to weight the model parameters uploaded by different RSUs. According to the spatio-temporal characteristics such as traffic flow and vehicle density in the area where each RSU is located, the weight w k .
[0065] For example, by calculating the cosine similarity between the spatio-temporal feature vector s k and a preset key feature vector s key to determine the weight:
[0066]
[0067] In addition, in some embodiments, the following formula is used to perform clipping on:
[0068]
[0069] Among them, E′ i is the first model parameter or the first aggregation parameter after clipping, E i is the first model parameter or the first aggregation parameter, and T0 is the self-set clipping threshold;
[0070] Add noise according to the following formula:
[0071]
[0072] where E″ i is the first model parameter or the first aggregation parameter after adding noise, ε is the privacy budget, which controls the intensity of the added noise, controls the intensity of the added noise, represents the noise value sampled from the Laplace distribution with a mean of 0 and a scale parameter of , S(E′ i ) is the sensitivity of E′ i , E i (D) is the first model parameter or the first aggregation parameter of the node model trained from the dataset D, E i (D′) is the first model parameter or the first aggregation parameter of the node model trained from the dataset D′. The dataset D and the dataset D′ are two datasets of adjacent nodes. In federated learning, it can be understood as a slight change in the vehicle local dataset.
[0073] It should also be noted that in the process of protecting the entire vehicle network model in the privacy budget allocation and adjustment, it is crucial to reasonably allocate the privacy budget. Considering that the dynamic k-anonymity privacy protection in the dynamic anonymous privacy protection has protected the data privacy to a certain extent, based on this, in the model protection based on differential privacy, the privacy budget can be dynamically adjusted according to the intensity of the vehicle location privacy protection and the importance of the model parameters. The privacy budget is specifically calculated according to the following formula:
[0074]
[0075] where ε total is the self-set total privacy budget, α is the adjustment coefficient, which is used to balance the weights of the location privacy protection and the model parameter privacy protection, I k is the intensity of the vehicle location privacy protection, which can be measured by factors such as the k value in the dynamic k-anonymity and the size of the privacy area, I m is the importance of the model parameters, which can be measured by the sensitivity of the parameters to the model performance. In addition, the privacy budget ε RSU when the RSU uploads to the cloud serverSimilarly, in a similar manner, the overall situation of vehicle location privacy protection within the RSU coverage area and the importance of RSU aggregation model parameters can be combined for dynamic adjustment. Through this differential privacy-based protection mechanism, in coordination with the dynamic k-anonymous privacy protection of dynamic anonymous privacy protection, the privacy security of model parameters in the data sharing process in the vehicle network is further enhanced, preventing attackers from reverse-inferring user privacy information through model parameters, while reasonably allocating privacy budgets to ensure an optimal privacy-utility balance in different privacy protection links.
[0076] In summary, through this data sharing method with a two-layer structure based on federated learning, combined with the data and features provided by the digital twin model constructed in the digital twin model of the vehicle network, it is possible to not only achieve efficient data sharing in the vehicle network, but also protect the privacy of vehicle users to a certain extent, while using spatio-temporal features to optimize the aggregation process and improve the accuracy and adaptability of the global model.
[0077] Step S104: Intercept a set of location points within a preset time period covering the current moment from the dynamic location set, and calculate the geometric center of the set of location points, so as to construct a hyper-ellipsoid equation based on the geometric center of gravity, and obtain a number of virtual locations corresponding to the real location at the current moment according to the hyper-ellipsoid equation, and combine the real location and the number of virtual locations into a target information set;
[0078] In this step, regarding the precise construction of the privacy area based on the dynamic location set, in the data sharing process of the two-layer structure of federated learning in the data sharing of the two-layer structure of federated learning, the security of vehicle location information is crucial during local training of the vehicle and uploading data to the RSU. With the help of the vehicle dynamic location set DLD generated by the vehicle network digital twin model, privacy protection operations are carried out. First, assume that the actual location of the current vehicle is P0(x0, y0), and intercept the set of location points from DLD:
[0079] S = {(x i , y i , t i )|i = 0, 1,..., m};
[0080] To accurately determine the geometric center of gravity of the set of location points, the weighted average method is used, fully considering the importance of location points at different times. Introduce a time decay factor, and specifically calculate the geometric center of gravity of the set of location points according to the following formula:
[0081]
[0082] where α i is the introduced time decay factor, 0 < α i < 1, x g , y grespectively represent the horizontal and vertical coordinates of the geometric centroid G. A privacy region is constructed based on the dispersion degree of the position point set, and the Mahalanobis Distance is introduced to more accurately measure the dispersion characteristics of the point set. Calculate the covariance matrix of the position point set:
[0083]
[0084] Among them, is the covariance matrix of the position point set, and are the means of the vector set X in the j-th and k-th dimensions respectively, x ij is the value of x i in the j-th dimension, and x ik is the value of x i in the k-th dimension.
[0085] Taking the geometric centroid G as the center, construct a hyper-ellipsoid as the privacy region. The equation of the hyper-ellipsoid is:
[0086]
[0087] Among them, T represents the transpose, c is a constant, is the inverse matrix of the covariance matrix of the position point set. By adjusting the value of c, the size of the hyper-ellipsoid can be controlled to meet the requirements of different privacy protection intensities.
[0088] In addition, within the hyper-ellipsoid privacy region constructed above, 2k - 2 candidate virtual positions are generated using the Quasi-Monte Carlo method. Compared with the traditional Monte Carlo method, the points generated by the Quasi-Monte Carlo method are more evenly distributed, which can improve the effectiveness of virtual positions. In the parametric equation of the hyper-ellipsoid, the angular values are determined by sampling on a low-discrepancy sequence (such as the Halton sequence) to generate the candidate virtual position P j (x j , y j ).
[0089] To select k - 1 virtual positions from the 2k - 2 candidate virtual positions, a screening algorithm based on Bayesian inference of model parameters is designed. In the two-layer structure data sharing based on federated learning, the RSU aggregates the model parameters θ RSU of multiple vehicles, which follows a certain probability distribution p(θ RSU ).
[0090] For each candidate virtual position P j , construct a virtual vehicle model and assume that it completes local training based on the simulated data around this position to obtain the virtual model parameter θ vj . Then calculate θ vj when given θ RSUThe posterior probability is:
[0091]
[0092] Among them, p(θ vj |θ RSU ) is the likelihood function, and the posterior probability threshold T is set p , if p(θ RSU |θ vj )≥T p , then the candidate virtual position P j Select the final set of virtual positions V until the set V contains k-1 virtual positions, p(θ RSU |θ vj ) is θ vj Time θ RSU The posterior probability, posterior probability p(θ RSU ) is θ RSU The historical distribution of is based on the aggregated results of past federated learning.
[0093] The real location P0 and the selected k-1 virtual locations are combined into a target information set F = {P0} ∪ V containing k locations, achieving a dynamic k-anonymity effect. When providing information related to the vehicle's location to the outside, only the target information set F is provided, effectively preventing attackers from accurately inferring the vehicle's real location from the set, and effectively protecting the vehicle's location privacy. As the vehicle continues to drive and time passes, the dynamic location set DLD is continuously updated, and the privacy area construction, virtual location generation and screening processes are also dynamically adjusted. Moreover, this process is deeply integrated with the federated learning process in the two-layer structure data sharing based on federated learning. With the help of information such as the probability distribution of model parameters generated in the federated learning process, the privacy protection strategy is optimized to ensure that the vehicle's location privacy is always effectively protected in the complex and changing Internet of Vehicles environment.
[0094] Step S105: Target vehicles having a credibility higher than a first preset threshold and not identified as attacks are screened out according to the pre-trained attack recognition model, and the target information set of the target vehicles is uploaded to perform federated learning on the aggregated global model.
[0095] After completing the model protection operation based on differential privacy, in order to further ensure the security and efficiency of vehicle network data sharing, it is necessary to build a vehicle user reputation evaluation system. Combining the vehicle location privacy protection in dynamic anonymous privacy protection and the stability of model parameters in differential privacy model protection, the reputation evaluation index is determined. Let the performance of vehicle i in dynamic k anonymous privacy protection be P i,k, which can be measured by the degree of compliance with the standard process in aspects such as privacy area construction and virtual location screening, and the value range is [0,1]. In the model protection based on differential privacy, after the model parameters are trimmed and noise is added, the change rate of the model performance is ΔM i By comparing the accuracy and recall of the model on the validation set before and after adding noise, the reputation of vehicle i is defined as R i for:
[0096] R i =βP i,k (1-β)·(1-|ΔM i |);
[0097] In addition, machine learning algorithms are used to identify attack behaviors in the Internet of Vehicles, such as malicious vehicle users deliberately uploading incorrect model parameters to interfere with the federated learning process. The support vector machine (SVM) algorithm is used to build an attack identification model, which takes the model parameter features uploaded by the vehicle and its privacy protection and model protection related features in dynamic anonymous privacy protection and differential privacy-based model protection as input. Let the feature vector of the first model parameter uploaded by vehicle i be B i,m , its characteristic vector in privacy protection and model protection is B i,p , then the input feature vector For linear SVM, the decision function is:
[0098] f(B i )=sgn(w T B i +b);
[0099] Among them, f(B i ) is the decision function, sgn is the sign function, this function maps the input value to +1, 0 or -1, which is used to determine which category the input sample belongs to. In SVM, sgn(w T B i +b) determines whether the input sample is classified as positive (i.e., normal behavior when the output is 1). When the output is 0, the sample is on the decision boundary and needs further processing or is considered uncertain and discarded. Or negative (i.e., attack behavior when the output is -1). w is the weight vector and b is the bias term. By training on a training set consisting of known normal and attack samples, the optimization problem is solved:
[0100]
[0101] Among them, d i is the sample label, if d i =1 indicates a normal sample, if d i =-1 represents the attack sample, C is the penalty parameter, ξ iis a slack variable. The SVM model obtained through training can classify the behaviors of vehicle users and identify potential attack behaviors.
[0102] When sharing data in federated learning, an adaptive selection is made according to the credibility of vehicle users and the attack recognition results. Specifically, the trained linear SVM is used to identify the vehicles that have not been attacked, and the target vehicles with a credibility higher than the first preset threshold are selected from the vehicles that have not been attacked:
[0103] U selected ={i∈U|R i ≥R th and f(B i )=1};
[0104] Among them, R th is the first preset threshold, U is the set of vehicles participating in federated learning, and U selected is the set containing all target vehicles.
[0105] In addition, in order to encourage vehicle users to maintain good behaviors, certain rewards are given to vehicle users who participate in federated learning and contribute positively, such as increasing their resource allocation priority in subsequent vehicle networking services. Through this adaptive vehicle user selection and attack recognition mechanism, it is connected with the model protection based on differential privacy, further improving the security and reliability of vehicle networking data sharing, ensuring that only trusted vehicle users participate in model training, and preventing malicious attacks from damaging the vehicle networking system.
[0106] In summary, the present invention has the following advantages:
[0107] 1. By constructing a vehicle networking digital model, the digital modeling of physical entities and the establishment of a set of vehicle dynamic positions are realized, providing a comprehensive and accurate data basis for data sharing. Based on the two-layer structure data sharing method of federated learning, combined with the data and features provided by the digital model, while realizing efficient data sharing, the dynamic anonymous privacy protection and differential privacy technologies are used to protect the privacy of vehicle users from two levels of location information and model parameters, effectively balancing the relationship between data sharing and privacy protection.
[0108] 2. In the construction of the vehicle networking digital model, the graph neural network (GNN) and the spatio-temporal attention mechanism are used, which can effectively learn the complex relationships between various nodes in the vehicle networking, accurately capture the spatio-temporal change features, and improve the response ability of the model to the dynamic changes of the vehicle networking. The aggregation optimization in the federated learning process and the model protection based on differential privacy consider the importance of model parameters and the spatio-temporal features of different regions, improving the accuracy and adaptability of the global model, making it more in line with the complex and changeable actual environment of the vehicle networking.
[0109] 3. By constructing a vehicle user reputation algorithm and determining reputation evaluation indicators in combination with the vehicle location privacy protection situation and the stability of model parameters, the credibility of vehicle users can be effectively evaluated. Based on the machine learning-based attack recognition algorithm and the adaptive vehicle user selection strategy, potential attack behaviors can be identified, trusted vehicle users can be selected to participate in federated learning, vehicle users can be encouraged to maintain good behaviors, malicious attacks on the vehicle networking system can be prevented, and the security and reliability of vehicle networking data sharing can be comprehensively improved.
[0110] As Figure 2 shown, an embodiment of the present invention further provides a vehicle networking data dynamic protection system during the federated learning process. The system includes:
[0111] A location set construction module 10, configured to construct a vehicle dynamic location set, where the vehicle dynamic location set at least includes the position coordinates, driving direction, and speed of the vehicle at different times;
[0112] A digital model building module 20, configured to build a vehicle networking digital model. The vehicle networking digital model includes multiple node models, each node model corresponding to a vehicle, and training the node models under the same vehicle according to the vehicle dynamic location set to obtain the trained node models;
[0113] A model parameter uploading module 30, configured to upload the first model parameters of the trained node models to the corresponding roadside units after differential privacy protection, so that the roadside units perform a first aggregation on the first model parameters to obtain a first aggregation parameter, and upload the second model parameters under all roadside units to the server after differential privacy protection, so that the server performs a second aggregation on the second model parameters under all roadside units to obtain a second aggregation parameter, and constructs an aggregated global model according to the second aggregation parameter;
[0114] A hyperellipsoid equation construction module 40, configured to intercept a position point set within a preset time period covering the current moment from the dynamic location set, calculate the geometric center of the position point set, construct a hyperellipsoid equation according to the geometric center of gravity, and obtain a plurality of virtual positions corresponding to the real position at the current moment according to the hyperellipsoid equation, and combine the real position and the plurality of virtual positions into a target information set;
[0115] A federated learning module 50, configured to screen out target vehicles with a reputation higher than a first preset threshold and not identified as an attack according to a pre-trained attack recognition model, and upload the target information sets of the target vehicles to perform federated learning on the aggregated global model.
[0116] On the other hand, the present invention also provides a storage medium, on which one or more programs are stored, and when the program is executed by a processor, it implements the method for dynamically protecting vehicle networking data in the above-mentioned federated learning process.
[0117] On the other hand, the present invention also provides an electronic device, including a memory and a processor, where the memory is used to store a computer program, and the processor is used to execute the computer program stored on the memory to implement the method for dynamically protecting vehicle networking data in the above-mentioned federated learning process.
[0118] Those skilled in the art can understand that the logic and / or steps represented in the flowchart or described in other ways herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch instructions from the instruction execution system, apparatus, or device and execute the instructions), or used in combination with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.
[0119] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, then editing, interpreting, or processing it in other suitable ways as necessary, and then storing it in a computer memory.
[0120] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known techniques in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.
[0121] Although the embodiments of the present invention have been described in detail above, it will be apparent to those skilled in the art that various modifications and variations can be made to these embodiments. However, it should be understood that such modifications and variations are all within the scope and spirit of the present invention as described in the claims. Moreover, the present invention described herein may have other embodiments and can be implemented or realized in various ways.
Claims
1. A method for dynamically protecting vehicle network data in the process of federated learning, characterized in that The method comprises: Constructing a vehicle dynamic position set, wherein the vehicle dynamic position set includes at least the position coordinates, driving direction, and speed of the vehicle at different times; Building a digital model of the Internet of Vehicles, the digital model comprising a plurality of node models, each node model corresponding to a vehicle, and training the node models under the same vehicle according to the vehicle dynamic position set to obtain a trained node model; Perform differential privacy protection on the first model parameters of the trained node model and upload them to the corresponding roadside unit, so that the first model parameters of the roadside unit are aggregated once to obtain a first aggregated parameter. Perform differential privacy protection on the second model parameters of all roadside units and upload them to the server, so that the server performs secondary aggregation on the second model parameters of all roadside units to obtain a second aggregated parameter. An aggregated global model is constructed based on the second aggregated parameter. A set of position points within a preset time period covering the current moment is intercepted from the dynamic position set, and a geometric center of gravity of the position point set is calculated to construct a hyperellipsoid equation based on the geometric center of gravity, and a plurality of virtual positions corresponding to the real position at the current moment are obtained based on the hyperellipsoid equation, and the real position and the plurality of virtual positions are combined into a target information set; Target vehicles whose credibility is higher than a first preset threshold and are not identified as attacks are screened out according to the pre-trained attack recognition model, and target information sets of the target vehicles are uploaded to perform federated learning on the aggregated global model.
2. The method for dynamically protecting vehicle networking data in the federated learning process according to claim 1, wherein The step of constructing a vehicle dynamic position set, wherein the vehicle dynamic position set includes at least the position coordinates, driving direction, and speed of the vehicle at different times, comprises: Vehicle information is measured according to the following formula: Z k = H(X k ) + R k ; Among them, Z k is the sensor measurement value vector, H(X k ) is the measurement function, and R k is the measurement noise that maps the state variable to the measurement space; The equation of state is constructed according to the following formula: X k = F(X k-1 , u k-1 ) + Q k-1 ; Among them, X k is the state vector at time k, including the vehicle position x k , speed v k , acceleration a k , that is, X k = [x k , v k , a k T , F(X k-1 , u k-1 ) is the state transition function, and Q k-1 is the process noise; Let the vehicle be at time series \(t_1,\ldots,t\) i ,\ldots,t n At time \(t\) i , the position coordinates are \((x i , y i ), the driving direction is \(A i \), and the speed is \(v i \). Then the dynamic position set DLD can be expressed as: DLD={(x i ,y i ,A i ,v i ,t i )|i=1,2,…,n}。 3. The method for dynamically protecting vehicle networking data in the federated learning process according to claim 2, wherein, The steps of building a digital model of the Internet of Vehicles, wherein the digital model includes a plurality of node models, and each node model corresponds to a vehicle, include: Model the vehicle network using a graph neural network, with vehicles, RSUs, and cloud servers in the vehicle network as nodes, and the communication links between nodes as edges. For each node a i , its feature vector h i contains node information such as the position, speed, and driving direction of the vehicle; Through the graph convolution operation, calculate the updated feature of node a according to the following formula i where N(i) is the set of neighbor nodes of node a i and d i and d j are the degrees of node a i and a j respectively, W is a learnable weight matrix, b is a bias term, and σ is an activation function; Using a spatio-temporal attention mechanism, the node information at different times and different spatial positions is weighted and fused. Among them, the attention weights in the time dimension and the attention weights in the spatial dimension are expressed as follows: Among them, exp(f t (t i )) performs an exponential transformation on the fractional function value corresponding to time t i . is to sum up the exponentially transformed fractional values corresponding to all times. exp(f s (x i , y i )) maps the original score to the positive space. is to sum up the exponential scores of all neighbors j ∈ N(i) of node a i .
4. The method for dynamically protecting vehicle networking data in the federated learning process according to claim 3, wherein The step of training the node model under the same vehicle according to the vehicle dynamic position set to obtain the trained node model includes: In the training of the node model, for an MLP with L hidden layers, the output a of the l-th layer l is calculated through forward propagation: a l = σ(W l a l-1 + b l ); Among them, W l is the weight matrix of the l-th layer, b l is the bias vector, and a l-1 is the output of the (l - 1)-th layer; Assume that the vehicle training data set is D i , contains N samples, for the sample (x n ,y n )∈D i , the cross entropy loss function is: where C is the number of classes, and y nk is the one-hot encoded value of the true label of sample n in the k-th class, is the predicted value of the model output in the k-th class; Update the model parameters of the node model according to the following formula: Among them, are the first model parameters of the node model corresponding to vehicle i at time t + 1 and time t respectively, η is the learning rate, is the loss function L i with respect to gradient.
5. The method for dynamically protecting vehicle networking data in the federated learning process according to claim 4, wherein The steps of performing differential privacy protection on the first model parameters of the trained node model and uploading them to the corresponding roadside unit so that the first model parameters of the roadside unit are aggregated once to obtain a first aggregated parameter, performing differential privacy protection on the second model parameters of all roadside units and uploading them to the server so that the server performs secondary aggregation on the second model parameters of all roadside units to obtain a second aggregated parameter, and constructing an aggregated global model based on the second aggregated parameter include: trimming the first model parameters, adding noise to the trimmed first model parameters to obtain first model parameters after adding noise, and uploading the first model parameters after adding noise to the roadside unit; The roadside unit restores the first model parameters after adding noise to obtain the first model parameters, and performs a first aggregation according to the following formula: Among them, θ RSU is the first aggregation parameter, w ij is the adjustment factor, n i is the data volume, θ i is the first model parameter corresponding to vehicle i, M is the number of vehicles participating in training within the RSU coverage range, λ is the adjustment parameter, sim ij is the cosine similarity, θ j is the first model parameter of the node model of vehicle j; The differential privacy protection method for the first model parameter is exactly the same as the first aggregation parameter. The cloud server restores the first aggregation parameter after adding noise, and then performs secondary aggregation according to the following formula: Among them, K is the number of participating RSUs, θ RSU,k is the first aggregation parameter aggregated by the k-th RSU, N k is the total amount of data within the coverage of the cloud server, θ global is the second aggregation parameter, w k is the weight.
6. The method for dynamically protecting vehicle networking data in the federated learning process according to claim 5, wherein The steps of differential privacy protection include: Clip according to the following formula: Among them, E′ i is the first model parameter or first aggregation parameter after pruning, E i is the first model parameter or the first aggregation parameter, T0 is the self-set clipping threshold; Add noise according to the following formula: Among them, E″ i is the first model parameter or the first aggregation parameter after adding noise, ε is the privacy budget, which controls the intensity of adding noise, indicating the noise value sampled from the Laplace distribution with a mean of 0 and a scale parameter of , and S(E′ i ) is the sensitivity of E′ i , E i (D) is the first model parameter or the first aggregation parameter of the node model trained from the dataset D, and E i (D′) is the first model parameter or the first aggregation parameter of the node model trained from the dataset D′. The dataset D and the dataset D′ are two datasets of adjacent nodes; Calculate the privacy budget according to the following formula: Among them, ε total is the total privacy budget set by oneself, α is the adjustment coefficient, I k is the vehicle location privacy protection strength, I m is the importance of model parameters.
7. The method for dynamically protecting vehicle networking data in the federated learning process according to claim 6, wherein The step of intercepting a set of position points within a preset time period covering the current moment from the dynamic position set and calculating the geometric center of the set of position points to construct a hyperellipsoid equation based on the geometric center of gravity includes: Let the actual position of the current vehicle be P0(x0, y0), and intercept the set of position points from the DLD: S = {(x i , y i , t i ) | i = 0, 1, …, m}; Calculate the geometric center of gravity of the set of position points according to the following formula: Among them, α i is the introduced time decay factor, 0 < α i < 1, x g , y g respectively represent the abscissa and ordinate of the geometric centroid G; Let vector X i =[x i ,y i ] T , and all vectors X i Summarizing the vector set X, the covariance matrix of the position point set is: Among them, is the covariance matrix of the set of position points, and are the means of the vector set X in the j-th and k-th dimensions respectively, and x ij is the value of x i in the j-th dimension, and x ik is the value of x i in the k-th dimension; Construct a hyperellipsoid equation according to the following formula: where T represents transpose and c is a constant, is the inverse matrix of the covariance matrix of the set of position points.
8. The method for dynamically protecting vehicle networking data in the federated learning process according to claim 7, wherein The step of obtaining a plurality of virtual positions corresponding to the true position at the current moment according to the hyperellipsoid equation and combining the true position and the plurality of virtual positions into a target information set includes: In the parametric equation of the super-ellipsoid, generate the candidate virtual position P j (x j , y j ). For each candidate virtual position P j , construct a virtual vehicle model and assume that it completes local training based on the simulation data around this position to obtain the virtual model parameter θ vj , and calculate the posterior probability of θ vj when given θ RSU as: where p(θ vj |θ RSU ) is the likelihood function. Set the posterior probability threshold T p . If p(θ RSU |θ vj ) ≥ T p , then select the candidate virtual position P j into the final virtual position set V until the set V contains k - 1 virtual positions. p(θ RSU |θ vj ) is the posterior probability of θ vj when θ RSU . The posterior probability p(θ RSU ) is the historical distribution of θ RSU , which is statistically obtained based on the aggregation results of past federated learning; Combine the true position P0 with the selected k - 1 virtual positions to form a target information set F = {P0} ∪ V containing k positions.
9. The method for dynamically protecting vehicle networking data in the federated learning process according to claim 8, wherein, The step of screening out target vehicles with a credibility higher than a first preset threshold and not identified as an attack according to the pre-trained attack recognition model and uploading the target information set of the target vehicles to perform federated learning on the aggregated global model includes: Assume that the performance of vehicle i in dynamic k-anonymity privacy protection is P i,k , the value range is [0,1], defining the reputation R of vehicle i i for: R i = βP i,k + (1 - β)·(1 - |ΔM i |); Among them, β is a regulation factor used to balance the weights of location privacy protection performance and model performance stability in reputation evaluation, and its value range is [0, 1], ΔM i is the change rate of the node model performance; the higher the reputation, the better the vehicle user performs in terms of privacy protection and model contribution. Let the eigenvector of the first model parameter uploaded by vehicle i be B i,m , and its eigenvector in privacy protection and model protection is B i,p , then the input eigenvector For a linear SVM, its decision function is: f(B i ) = sgn(w T B i + b); where f(B i ) is the decision function, sgn is the sign function, w is the weight vector, and b is the bias term. By training on a training set composed of known normal and attack samples, the optimization problem is solved: Among them, d i is the sample label. If d i = 1, it indicates a normal sample. If d i = -1, it indicates an attack sample. C is the penalty parameter, and ξ i is the slack variable; Identify the vehicles not under attack according to the trained linear SVM, and screen out target vehicles with a credibility higher than the first preset threshold from the vehicles not under attack: U selected = {i ∈ U | R i ≥ R th and f(B i ) = 1}; Among them, R th is the first preset threshold, U is the set of vehicles participating in federated learning, and U selected is the set containing all target vehicles.
10. A dynamic protection system for Internet of Vehicles data in a federated learning process, characterized in that: The system includes: A position set construction module for constructing a vehicle dynamic position set, where the vehicle dynamic position set at least includes the position coordinates, driving direction, and speed of the vehicle at different times; A digital model building module for building a vehicle networking digital model, where the vehicle networking digital model includes multiple node models, each node model corresponding to a vehicle, and training the node models under the same vehicle according to the vehicle dynamic position set to obtain trained node models; A model parameter uploading module for uploading the first model parameters of the trained node models to the corresponding roadside units after differential privacy protection, so that the roadside units perform primary aggregation on the first model parameters to obtain a first aggregation parameter, and uploading the second model parameters under all roadside units to the server after differential privacy protection, so that the server performs secondary aggregation on the second model parameters under all roadside units to obtain a second aggregation parameter, and constructing an aggregated global model according to the second aggregation parameter; A hyperellipsoid equation construction module for intercepting a set of position points within a preset time period covering the current moment from the dynamic position set, calculating the geometric center of the set of position points, constructing a hyperellipsoid equation based on the geometric center of gravity, obtaining a plurality of virtual positions corresponding to the true position at the current moment according to the hyperellipsoid equation, and combining the true position and the plurality of virtual positions into a target information set; The federated learning module is used to screen out target vehicles whose credibility is higher than a first preset threshold and which have not been identified as attacks based on the pre-trained attack recognition model, and upload the target information set of the target vehicles to perform federated learning on the aggregated global model.
Citation Information
Cited By
Federal learning training method suitable for dynamic Internet of Vehicles based on vehicle state perception
CN121745225A
Federal learning training method based on vehicle state perception suitable for dynamic vehicle networking
CN121745225B