Federal learning data aggregation method for predicting traffic flow, medium and equipment
Through nuclear ridge regression and homomorphic encryption technology, the weight allocation and screening of Internet of Vehicle data is optimized, which solves the problem of insufficient traffic flow prediction accuracy caused by data heterogeneity in the Internet of Vehicles environment, and achieves higher prediction accuracy and data security.
Patent Information
- Application Number
- CN202510501906.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-07-22
AI Technical Summary
The problem of data heterogeneity in the Internet of Vehicles environment leads to insufficient model generalization capabilities and prediction accuracy in traffic flow prediction in traditional data aggregation methods.
The vehicle data is weighted by nuclear ridge regression technology, and data aggregation is carried out in the ciphertext state through homomorphic encryption technology, combining the vehicle's credibility and data quality evaluation, and weight allocation and data screening are optimized.
It improves the accuracy of traffic flow prediction and generalization capabilities of the model, reduces the amount of data aggregation, and ensures the privacy and security of the data.
Smart Images

Figure CN120356331A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of vehicle networking and federated learning, and particularly relates to a federated learning data aggregation method, medium and device for predicting traffic flow. Background Art
[0002] With the rapid development of vehicle networking technology, vehicles, as mobile data collection terminals, can obtain a large amount of driving data, such as road conditions, vehicle speeds, driving behaviors, etc. Federated learning provides a privacy protection solution for the collaborative utilization of data in vehicle networking, enabling vehicles to jointly train a global model to serve various vehicle networking applications, such as traffic flow prediction, etc., without directly sharing the original data.
[0003] However, the data in the vehicle networking environment has characteristics such as high-dimensionality, sparsity, uneven distribution, and noise interference. Traditional data aggregation methods (FedAvg algorithm) have problems of data heterogeneity, resulting in deficiencies in the generalization ability and prediction accuracy of the model. As a powerful machine learning technology, Kernel Ridge Regression (RNN) can effectively process complex discrete data and alleviate the overfitting problem. Introducing it into the federated learning data aggregation process for traffic flow prediction is expected to improve the prediction accuracy of the model and reduce the impact of data heterogeneity in federated learning, so as to better meet the complex and changeable application requirements of vehicle networking. Summary of the Invention
[0004] A federated learning data aggregation method, device and storage medium for predicting traffic flow proposed by the present invention can at least solve one of the technical problems in the background art.
[0005] To achieve the above object, the present invention adopts the following technical solutions:
[0006] A federated learning data aggregation method for predicting traffic flow includes the following steps:
[0007] S1. Establish a connection between the cloud server side and the vehicle side to establish a corresponding system;
[0008] S2. Collect vehicle data through the system for local training and data preprocessing;
[0009] S3. The vehicle side encrypts the preprocessed data through homomorphic encryption technology and uploads it to the cloud server side;
[0010] S4. Perform weight allocation on the encrypted and uploaded data based on kernel ridge regression technology;
[0011] S5. Aggregate the encrypted update amount based on the adjusted weights;
[0012] S6. Decrypt the aggregated encrypted update amount to complete the update of the pre-trained model.
[0013] Further, the method for preprocessing the local training set data of the vehicle data in step S2 of the present invention includes:
[0014] During the driving process of the vehicle, various driving data are collected through on-vehicle sensors, such as vehicle speed v, acceleration a, longitude and latitude position information pos, and the number of surrounding vehicles n near , and road condition information road;
[0015] Among them, the road condition information includes: road surface flatness, slope;
[0016] Preprocess the data collected by the on-vehicle sensors:
[0017] Data cleaning: Remove obviously abnormal data points; Set a reasonable vehicle speed range according to the legal maximum vehicle speed of the road type, and exclude data points when the vehicle speed suddenly changes; Set a reasonable vehicle acceleration range according to the vehicle performance, and exclude data points with abnormal acceleration
[0018] Data normalization: Normalize data in different ranges to the same interval; Use linear normalization, and the formula is: Normalize the vehicle speed to the interval [0, 1], where v is the original vehicle speed of a single individual, v min and v max are respectively the minimum and maximum vehicle speeds in this collection of data, and at the same time, normalize the acceleration and the number of surrounding vehicles;
[0019] Based on the preprocessed data, use the local machine learning model for training, and the model form is:
[0020]
[0021] Input features x1, w2, …, x n : Multi-source data of vehicle speed, acceleration, position, and the number of surrounding vehicles;
[0022] Output predicted value Traffic flow predicted value;
[0023] Parameters w0, w1, …, w n : Model weights, indicating the contribution of each feature to the traffic flow;
[0024] Calculate the update amount Δw of the model parameters, and use the stochastic gradient descent method SGD for training. In each iteration, for a training sample (x i , y i ), calculate the gradient:
[0025]
[0026] Among them, x i is the feature vector of the i-th training sample, and y i is its corresponding true label. is the gradient of the loss function with respect to the parameter w, and w T is the weight vector of the data x;
[0027] Update the model parameters according to the learning rate η:
[0028] After several iterations, the updated amount of the model parameters for this training is obtained as Δw = w - w0;
[0029] Among them, w0 is the initialized global model parameter.
[0030] Furthermore, the method for weight allocation of the encrypted uploaded data based on the kernel ridge regression technology in step S4 of the present invention includes:
[0031] S410. Evaluate the credibility and data quality of each vehicle and allocate initial weights;
[0032] The influencing factors for evaluation include: the historical data contribution of the vehicle, the current network connection stability, the integrity and accuracy of the data;
[0033] Allocate an initial weight to each vehicle The specific implementation of the allocation of the initial weight adopts the linear weighting method as follows:
[0034]
[0035] Among them, α1, α2, α3, α4 are weight coefficients, satisfying α1 + α2 + α3 + α4 = 1, and f time is the number of successful historical authentications of the vehicle with the RSU within a period of time, c hist,i is the historical data contribution score of vehicle i, s net,i is the network connection stability score, and q data,i is the data quality score. These scores are quantified according to the above evaluation factors;
[0036] S420. Use the kernel ridge regression technology to further adjust the weights of the vehicle update amount;
[0037] Construct a Gaussian kernel function
[0038] Among them, x i and x j are the encrypted update amounts of vehicle i and vehicle j;
[0039] Calculate the similarity matrix S. For each pair of vehicles i and j, where i calculates the similarity
[0040] Calculate the kernel matrix K = [S ij n×n ;
[0041] Then, according to the formula of kernel ridge regression, calculate (K + λI) -1 , where I is the identity matrix and λ is the regularization parameter. Then, calculate a = (K + λI) -1 y, where y is a column vector, is the initial weight; perform matrix multiplication operation on the ciphertext to obtain the adjusted weight vector a.
[0042] Furthermore, the aggregation process of the encrypted update amount in step S5 of the present invention includes:
[0043] Secure data aggregation: The server adopts a secure aggregation protocol to aggregate the encrypted update amounts based on the reallocated weights;
[0044] The aggregation process is carried out in the ciphertext state. Utilizing the characteristics of homomorphic encryption, calculate the aggregated update amount by means of weighted average; Let the encrypted update amount of vehicle i be Δw i , for the encrypted update amount Enc(Δw i ) and the adjusted weight a i , the number of vehicles participating in federated learning is n, calculate the aggregated encrypted update amount Enc(Δw agg ), and the calculation formula is as follows:
[0045]
[0046] Furthermore, step S6 of the present invention decrypts the aggregated encrypted update amount to complete the pre-trained model update method, including:
[0047] The server uses the homomorphic encryption decryption algorithm to decrypt the aggregated encrypted update amount Enc(Δw agg ) to obtain Δw agg ; Apply it to the global model to complete the update operation of the model: w g = w0 + Δw agg , where g ∈ N. When g = 1, that is, w1 is the global model parameter after the first round of update. The server feeds back the updated global model parameter w1 to the vehicles participating in federated learning;
[0048] After the vehicle receives the updated global model parameters fed back by the server, it repeats the operations from step S2 to step S6, updates the global model parameters, and feeds the updated parameters back to the vehicle. This process repeats until the global model reaches the predetermined convergence condition and the loss function value is less than the threshold L th . When the model meets the above conditions, the model weights are screened to obtain the lowest weight a that can participate in the aggregation i > δ, the data of this vehicle will be uploaded to the server and participate in the aggregation, and the updated amount Δw after aggregation agg The calculation formula is:
[0049]
[0050] where m is the number of vehicles that meet the weight conditions selected in the final round, n is the total number of vehicles participating in federated learning, and the qualified vehicle parameters are used to predict traffic flow problems
[0051] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to execute the steps of the above method
[0052] On yet another aspect, the present invention also discloses a computer device including a memory and a processor, where the memory stores a computer program, which, when executed by the processor, causes the processor to execute the steps of the above method
[0053] From the above technical solutions, it can be seen that the present invention has the following advantages compared with the traditional ones
[0054] 1. Improvement in the accuracy of prediction problems: By optimizing the weights of vehicle update amounts through kernel ridge regression technology, it is possible to better explore the potential patterns and relationships in vehicle networking data. For example, if the update amounts of vehicle i and vehicle j are highly similar, it indicates that they are more consistent in the model update direction. Then, when allocating weights, their weights can be appropriately increased to make the vehicle update amounts that contribute more similarly and reliably to the global model update play a more important role in the prediction process
[0055] 2. Optimization of data aggregation: Since the server screens the weights of all vehicles in the final training round, eliminating some vehicles with too low weights, it reduces the total amount of data that needs to be aggregated and ensures a certain similarity of the remaining data, making it easier to perform data aggregation on the remaining data Description of the Drawings
[0056] Figure 1 It is a flowchart of a federated learning data aggregation method for predicting traffic flow Detailed Embodiments
[0057] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention.
[0058] As Figure 1 shown, the federated learning data aggregation method for predicting traffic flow in this embodiment is executed by a computer device through the following steps:
[0059] A federated learning data aggregation method for predicting traffic flow mainly includes the following steps:
[0060] S1. Establish a connection between the cloud server side and the vehicle side to establish a corresponding system;
[0061] S2. Collect vehicle data through the system for local training and data preprocessing;
[0062] S3. The vehicle side encrypts the preprocessed data through homomorphic encryption technology and uploads it to the cloud server side;
[0063] S4. Perform weight assignment on the encrypted and uploaded data based on kernel ridge regression technology;
[0064] S5. Aggregate the encrypted update amount based on the reallocated weights;
[0065] S6. Decrypt the aggregated encrypted update amount to complete the update of the pre-trained model.
[0066] The following will explain each step in detail:
[0067] S1. Establish a connection between the cloud server side and the vehicle side to establish a corresponding system;
[0068] Cloud server side:
[0069] Initialize the global model parameters w0, w1, …, w n , using the methods of random initialization and loading pre-trained model parameters; at the same time, set the relevant parameters of kernel ridge regression, such as the bandwidth parameter σ of the Gaussian kernel function, the regularization parameter γ, etc. The initial values of these parameters are set according to the method of cross-validation and are dynamically adjusted according to the model performance in the subsequent process. Establish a communication connection with the vehicle and prepare to receive the data uploaded by the vehicle.
[0070] Vehicle side:
[0071] After the vehicle starts, establish a communication connection with the server and obtain the initial global model parameters w0, w1, …, w n ;
[0072] Initialize the local model training environment and prepare to collect driving data for local training.
[0073] S2. Collect vehicle data through the system for local training and data preprocessing;
[0074] Perform local vehicle training and data preprocessing:
[0075] During vehicle driving, collect various driving data through on-vehicle sensors, such as vehicle speed v, acceleration a, position information pos (latitude and longitude), the number of surrounding vehicles n near , road condition information road, etc.;
[0076] Among them, the road condition information includes: road surface flatness, slope, etc.
[0077] Preprocess the data collected by on-vehicle sensors:
[0078] Data cleaning: Remove significantly abnormal data points;
[0079] For example, the vehicle speed suddenly becomes an unreasonably large or small value (which can be judged by setting a reasonable vehicle speed range, such as 0 ≤ v ≤ v max , v max is the legal maximum vehicle speed for this road type), the acceleration exceeds the normal range (such as -a max ≤ a ≤ a max , a max set according to vehicle performance), and other abnormal data.
[0080] Data normalization: Normalize data in different ranges to the same interval;
[0081] For example, for vehicle speed, a linear normalization method can be used to normalize it to the [0, 1] interval. The formula is: where v is the original vehicle speed of a single individual, v min and v max are the minimum and maximum vehicle speeds in this data collection respectively. Similarly, corresponding normalization processing is also performed on other data such as acceleration and the number of surrounding vehicles.
[0082] Based on the preprocessed data, use a local machine learning model (such as a linear regression model, neural network model, etc.) for training. Here, we use a linear regression model as an example. The model form is:
[0083]
[0084] Input features x1, x2, …, x n : Multi-source data such as vehicle speed, acceleration, position, and the number of surrounding vehicles.
[0085] Output predicted value Traffic flow predicted value.
[0086] Parameters w0, w1, …, w n : Model weights, representing the contributions of each feature to traffic flow.
[0087] Calculate the update amount Δw of the model parameters. Use the Stochastic Gradient Descent (SGD) method for training. In each iteration, for a training sample (x i , y i ), calculate the gradient:
[0088]
[0089] where x i is the feature vector of the i-th training sample, and y i is its corresponding true label. is the gradient of the loss function with respect to the parameter w, and w T is the weight vector of the data x. Here, the mean squared error loss function is used;
[0090] Update the model parameters according to the learning rate η:
[0091] After several iterations, obtain the update amount of the model parameters for this training, Δw = w - w0;
[0092] where w0 is the initialized global model parameter.
[0093] S3. The vehicle encrypts the preprocessed data using homomorphic encryption technology and uploads it to the cloud server;
[0094] The vehicle encrypts the calculated update amount Δw of the model parameters using homomorphic encryption technology,
[0095] Encrypt each element of the update amount of the model parameters to obtain the encrypted update amount Enc(Δw).
[0096] The vehicle uploads the encrypted update amount Enc(Δw), as well as its own identification information (such as vehicle ID), metadata such as the timestamp of data collection, to the cloud server through the vehicle networking communication module.
[0097] The vehicle uses homomorphic encryption technology to encrypt the calculated model update amount, ensuring the privacy of data during transmission and subsequent aggregation. Homomorphic encryption allows specific mathematical operations to be directly performed on ciphertexts, enabling the server to aggregate the encrypted update amounts without decryption, effectively preventing the leakage of vehicle data privacy. If other encryption methods are used, it will affect the results. Taking the use of asymmetric encryption (such as RSA encryption) as an example, it does not have homomorphic properties and problems will occur during data aggregation and weight assignment because RSA encryption does not support distance calculation on ciphertexts and cannot directly calculate the similarity matrix and kernel matrix according to the above formula. If weights are to be calculated, Enc(Δw i ) needs to be decrypted to obtain Δw i , and then the similarity matrix and kernel matrix are calculated based on the plaintext to adjust the weights. However, doing so will also expose the vehicle data privacy and change the original calculation logic, which may lead to inaccurate weight assignment, thereby affecting the prediction accuracy and generalization ability of the model. The vehicle uploads the encrypted update amount and metadata such as its own identification information to the server.
[0098] S4. Perform weight assignment on the encrypted uploaded data based on kernel ridge regression technology
[0099] After the server receives the encrypted update amounts and metadata uploaded by each vehicle, the weight assignment process is as follows:
[0100] S410. Evaluate the credibility and data quality of each vehicle and assign initial weights;
[0101] The evaluation factors include but are not limited to the number of historical verification passes between the vehicle and the RSU within a period of time, the historical data contribution of the vehicle (for example, the contribution degree of the data provided during past participation in federated learning to the performance improvement of the global model, which can be measured by recording the changes in performance metrics before and after each model update), the current network connection stability (judged according to indicators such as the packet loss rate and latency of the data uploaded by the vehicle. For example, if the packet loss rate is lower than p th (threshold) and the latency is less than t th (threshold), the network connection is considered stable), the integrity and accuracy of the data (by checking whether the uploaded data is complete and comparing it with the data collected by other vehicles in the same time period and the same area to judge whether the data is reasonable and accurate), etc.
[0102] Based on these factors, an initial weight is assigned to each vehicle The initial weight is assigned using the linear weighting method. The specific implementation method is as follows:
[0103]
[0104] Among them, α1, α2, α3, and α4 are weight coefficients, satisfying α1 + α2 + α3 + α4 = 1, and f time is the number of successful historical authentications between the vehicle and the RSU within a period of time, and c hist,i is the historical data contribution score of vehicle i, and s net,i is the network connection stability score, and q data,i is the data quality score. These scores can be quantified according to the above evaluation factors (for example, the network connection stability score can be obtained by linearly mapping based on the packet loss rate and latency).
[0105] S420. Further adjust the weights of the vehicle update amounts using kernel ridge regression technology;
[0106] Construct a Gaussian kernel function
[0107] where x i and x j are the encrypted update amounts of vehicle i and vehicle j (in subsequent calculations, the server can perform corresponding operations on the ciphertext of homomorphic encryption to protect data privacy).
[0108] Calculate the similarity matrix S. For each pair of vehicles i and j (i = 1, 2,..., n; j = 1, 2,..., n), calculate the similarity
[0109] Calculate the kernel matrix K = [S ii n×n .
[0110] Then, according to the formula of kernel ridge regression, calculate (K + λI) -1 , where I is the identity matrix and λ is the regularization parameter. Then, calculate a = (K + λI) -1 y, where y is a column vector, (initial weight). Perform matrix multiplication operations on the ciphertext to obtain the adjusted weight vector a.
[0111] After the server receives the encrypted update amounts uploaded by each vehicle, it first evaluates the credibility and data quality of the vehicles. The evaluation factors can include the number of historical verification passes between the vehicle and the RSU in this area, the historical data contribution situation, the current network connection stability, the integrity and accuracy of the data, etc. According to the evaluation results, an initial weight is assigned to each vehicle. Then, kernel ridge regression technology is used to further adjust the weights of the vehicle update amounts. Specifically, the vehicle update amounts are regarded as sample data in the kernel ridge regression model. By constructing a suitable kernel function (such as a Gaussian kernel function), the similarity between the vehicle update amounts is calculated, and the weights are adjusted based on this, so that vehicles with high similarity to other high-quality update amounts obtain higher weights, thus more reasonably reflecting the vehicle situation.
[0112] S5. Aggregate the encrypted update amount based on the reallocated weights;
[0113] The process of aggregating the encrypted update amount includes:
[0114] Secure data aggregation: The server uses a secure aggregation protocol to aggregate the encrypted update amount based on the reallocated weights;
[0115] The aggregation process is carried out in the ciphertext state. Utilizing the characteristics of homomorphic encryption, the aggregated update amount is calculated by means of weighted average. Assume the encrypted update amount of vehicle i is Δw i , for the encrypted update amount Enc(Δw i ) and the adjusted weight a i , and the number of vehicles participating in federated learning is n, calculate the aggregated encrypted update amount Enc(Δw agg ), and the calculation formula is as follows:
[0116]
[0117] S6. Decrypt the aggregated encrypted update amount to complete the update of the pre-trained model;
[0118] Global model update and feedback (Round 1)
[0119] The server decrypts the aggregated encrypted update amount Enc(Δw agg ) (using the corresponding homomorphic encryption decryption algorithm) to obtain Δw agg . Apply it to the global model to complete the model update operation: w g = w0 + Δw agg , where g ∈ N. When g = 1, that is, w1 is the global model parameters after the first round of update. The server feeds back the updated global model parameters w1 to the vehicles participating in federated learning.
[0120] After the vehicle receives the updated global model parameters fed back by the server, it repeats the operations in steps two to six to update the global model parameters and feeds back the updated parameters to the vehicle. This process repeats until the global model reaches the predetermined convergence condition (for example, the loss function value is less than a certain threshold L th . When the model meets the above conditions, the model weights are screened, that is, when a i > δ (the minimum weight that can participate in aggregation), the data of this vehicle will be uploaded to the server and participate in the aggregation. Then the aggregated update amount Δw agg The calculation formula is:
[0121]
[0122] Where m is the number of vehicles that meet the weight conditions in the final round of screening, and n is the total number of vehicles participating in federated learning. The qualified vehicle parameters are used to predict traffic flow problems.
[0123] In summary, the method of the present invention aims to improve the accuracy of prediction problems. By optimizing the weights of vehicle update amounts through kernel ridge regression technology, the potential patterns and relationships in Internet of Vehicles data can be better mined. For example, if the update amounts of vehicle i and vehicle j are very similar, it means that they are relatively consistent in the direction of model update. Then, when allocating weights, it is possible to consider appropriately increasing their weights, so that the update amounts of vehicles that contribute more similarly and reliably to the global model update play a more important role in the prediction process.
[0124] Regarding the data aggregation optimization problem: In the final training round, the server screens the weights of all vehicles, eliminates some vehicles with too low weights, reduces the total amount of data that needs to be aggregated and ensures a certain similarity of the remaining data, making it easier to aggregate the remaining data.
[0125] On the other hand, the present invention further discloses a computer-readable storage medium storing a computer program, wherein when the computer program is executed by a processor, the processor executes the steps of the above method.
[0126] On the other hand, the present invention further discloses a computer device, including a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the above method.
[0127] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any of the federated learning data aggregation methods for predicting traffic flow in the above embodiments.
[0128] It is understandable that the system, device and storage medium provided in the embodiments of the present invention correspond to the method provided in the embodiments of the present invention, and the explanation, examples and beneficial effects of the relevant contents can refer to the corresponding parts in the above methods.
[0129] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can access, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0130] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0131] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.
[0132] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A federated learning data aggregation method for predicting traffic flow, characterized in that, It includes the following steps: S1. Set up a connection between the cloud server side and the vehicle side to establish a corresponding system; S2. Collect vehicle data through the system for local training and data preprocessing; S3. The vehicle side encrypts the preprocessed data through homomorphic encryption technology and uploads it to the cloud server side; S4. Allocate weights to the encrypted and uploaded data based on the kernel ridge regression technology; S5. Aggregate the encrypted update amount based on the adjusted weights; S6. Decrypt the aggregated encrypted update amount to complete the update of the pre-trained model.
2. The federated learning data aggregation method for predicting traffic flow according to claim 1, wherein, The method for preprocessing the local training set data of the vehicle data in step S2 includes: During the driving process of the vehicle, various driving data are collected through in-vehicle sensors, such as vehicle speed v, acceleration a, longitude and latitude position information pos, and the number of surrounding vehicles n near , and road condition information road; Among them, the road condition information includes: road surface flatness, slope; Preprocess the data collected by the vehicle-mounted sensor: Data cleaning: Remove obviously abnormal data points; Set a reasonable vehicle speed range according to the legal maximum vehicle speed of the road type, and exclude data points when the vehicle speed suddenly changes; Set a reasonable vehicle acceleration range according to the vehicle performance, and exclude data points with abnormal acceleration. Data normalization: Normalize data in different ranges to the same interval; use linear normalization, and the formula is: Normalize the vehicle speed to the interval [0, 1], where v is the original vehicle speed of a single individual, v min and v max are respectively the minimum and maximum vehicle speeds in the data collected this time, and at the same time, normalize the acceleration and the number of surrounding vehicles; Based on the preprocessed data, use the local machine learning model for training. The model form is: Input features x1, x2, …, x n : Multi-source data of vehicle speed, acceleration, position, and the number of surrounding vehicles; Output predicted value Traffic flow predicted value; Parameters w0, w1, …, w n : Model weights, representing the contributions of each feature to traffic flow; Calculate the update amount Δw of the model parameters and use the Stochastic Gradient Descent (SGD) method for training. In each iteration, for a training sample (x i , y i ), calculate the gradient: Among them, x i is the feature vector of the i-th training sample, and y i is its corresponding true label. is the gradient of the loss function with respect to the parameter w, and w T is the weight vector of the data x; Update the model parameters according to the learning rate η: After several iterations, the model parameter update amount Δw = w - w0 of this training is obtained; Among them, w0 is the initialized global model parameter.
3. The federated learning data aggregation method for predicting traffic flow according to claim 1, wherein The method for allocating weights to the encrypted and uploaded data based on the kernel ridge regression technology in step S4 includes: S410. Evaluate the credibility and data quality of each vehicle and allocate initial weights; The influencing factors for evaluation include: the contribution of the vehicle's historical data, the current network connection stability, the integrity and accuracy of the data; Assign an initial weight to each vehicle The initial weight is assigned by the linear weighting method, and the specific implementation method is as follows: where α1, α2, α3, and α4 are weight coefficients satisfying α1 + α2 + α3 + α4 = 1, and f time is the number of successful historical authentications between the vehicle and the RSU within a period of time, and c hist,i is the contribution score of the historical data of vehicle i, and s net,i is the network connection stability score, and q data,i is the data quality score, and is quantified according to the above evaluation factors; S420. Use the kernel ridge regression technology to further adjust the weights of the vehicle's update amount; Construct Gaussian kernel function where x i and x j are the encrypted update amounts of vehicle i and vehicle j; Calculate the similarity matrix S. For each pair of vehicles i and j, where i calculates the similarity Calculate the kernel matrix K = [S ij n×n ; Then, according to the formula of kernel ridge regression, calculate (K + λI) -1 , where I is the identity matrix and λ is the regularization parameter. Then, calculate a = (K + λI) -1 y, where y is a column vector, is the initial weight; perform matrix multiplication on the ciphertext to obtain the adjusted weight vector a.
4. The federated learning data aggregation method for predicting traffic flow according to claim 1, wherein The process of aggregating the encrypted update amount in step S5 includes: Secure data aggregation: The server adopts a secure aggregation protocol to aggregate the encrypted update amount based on the reallocated weights; The aggregation process is carried out in the ciphertext state. Utilizing the characteristics of homomorphic encryption, the aggregated update amount is calculated by means of weighted average. Let the encrypted update amount of vehicle i be Δw i , for the encrypted update amount Enc(Δw i ) and the adjusted weight a i , the number of vehicles participating in federated learning is n, and the encrypted update amount Enc(Δw agg ) after aggregation is calculated. The calculation formula is as follows:
5. The federated learning data aggregation method for predicting traffic flow according to claim 4, wherein: The method for decrypting the aggregated encrypted update amount in step S6 to complete the update of the pre-trained model includes: The server uses the homomorphic encryption and decryption algorithm to decrypt the aggregated encrypted update amount Enc(Δw agg ), and obtains Δw agg ; Apply it to the global model to complete the model update operation: w g = w0 + Δw agg , where g ∈ N. When g = 1, that is, w1 is the global model parameter after the first round of update. The server feeds back the updated global model parameter w1 to the vehicles participating in federated learning; After the vehicle receives the updated global model parameters fed back by the server, it repeats the operations from step S2 to step S6, updates the global model parameters, and feeds the updated parameters back to the vehicle. This process repeats until the global model reaches the predetermined convergence condition and the loss function value is less than the threshold L th . After the model meets the above conditions, the model weights are screened, that is, the lowest weight a participating in aggregation i > δ, the data of this vehicle will be uploaded to the server and participate in aggregation, then the updated amount Δw after aggregation agg The calculation formula is: Among them, m is the number of vehicles that meet the weight conditions selected in the final round, n is the total number of vehicles participating in federated learning, and the qualified vehicle parameters are used to predict traffic flow problems.
6. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the processor is caused to execute the method according to any one of claims 1 to 5.
7. A computer device, comprising a memory and a processor, characterized in that, The memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the method according to any one of claims 1 to 5.
Citation Information
Cited By
Traffic flow prediction model training method and system for intelligent networked vehicles
CN121686803A
A traffic flow prediction model training method and system for intelligent networked vehicles
CN121686803B
Federal learning training method suitable for dynamic Internet of Vehicles based on vehicle state perception
CN121745225A
Federal learning training method based on vehicle state perception suitable for dynamic vehicle networking
CN121745225B