Traffic congestion prediction method, system, device and medium based on swarm learning

By calculating robustness values ​​and adaptive weight adjustments in bee colony learning, and combining localized differential privacy and blockchain technology, the privacy and security challenges of bee colony learning in traffic data processing are solved, enabling more efficient and secure training and application of traffic prediction models.

CN117173882BActive Publication Date: 2026-05-08BEIJING INFORMATION SCI & TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING INFORMATION SCI & TECH UNIV
Filing Date
2023-08-21
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Bee colony learning presents challenges in traffic data processing and prediction, including privacy protection, data heterogeneity, scalability, and security. In particular, in dynamic environments, node joining and leaving can lead to network instability, and sharing model parameters poses security risks.

Method used

By employing robustness value calculation, local model noise addition, blockchain registration, and adaptive weight adjustment, localized differential privacy technology is used to protect model parameters. Combined with blockchain technology, network trustworthiness and the legitimacy of participating nodes are ensured, and the model aggregation weights are dynamically adjusted to improve security and accuracy.

Benefits of technology

It improves the privacy and security of traffic prediction models, prevents poisoning attacks, ensures privacy protection and security during model training, and enhances the scalability and accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117173882B_ABST
    Figure CN117173882B_ABST
Patent Text Reader

Abstract

The application discloses a traffic congestion prediction method, system, device and medium based on swarm learning, and relates to the technical field of swarm learning. The method comprises the following steps: determining the robustness value of a local model of each participant node; under the current aggregation round, obtaining the local model training evaluation weight of each participant node under the current aggregation round according to the local data credibility, the robustness value of each participant node under the last aggregation round, and the machine learning evaluation index under the current aggregation round; adopting the local differential privacy technology to add noise to the local model of the current aggregation round to obtain the noisy local model under the current aggregation round; obtaining the local model under the next aggregation round according to the local model of the current aggregation round, the local model training evaluation weight, and the noisy local model; and determining the local model under the next aggregation round as a traffic prediction model. The application can improve the privacy and security in the process of data processing and traffic prediction model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bee colony learning technology, and in particular to a method, system, device and medium for predicting traffic congestion based on bee colony learning. Background Technology

[0002] Modern vehicles are equipped with onboard sensors that collect various traffic data in real time, such as traffic flow, vehicle speed, and road conditions. This data can be used to train traffic prediction models for subsequent traffic congestion prediction. However, traditional centralized machine learning methods face several challenges when handling large-scale data. First, centralized machine learning requires collecting all raw data to a central server for training and model merging, which can involve significant data transfer and storage requirements, and may also raise privacy and security concerns. Furthermore, some data owners may be unwilling to share their sensitive data, limiting the feasibility of centralized machine learning methods. Second, centralized machine learning methods cannot adapt to data variations and heterogeneity in distributed environments. In a distributed environment, the distribution, quantity, and quality of data from different participating nodes may differ, making it difficult for centralized models to adapt to different data characteristics and changes.

[0003] To overcome these problems, bee swarm learning emerged. Bee swarm learning is a decentralized machine learning framework that leverages edge computing and peer-to-peer networking technologies to achieve secure model merging and parameter sharing while protecting data privacy. It uses blockchain technology to ensure network trustworthiness and provides a higher level of data security and protection, enabling participants to collaborate securely on machine learning tasks. It allows for model training and updating of data distributed across different locations while protecting data privacy. In bee swarm learning, participating nodes (such as devices, sensors, or edge nodes) train the model locally, using local data for learning. Then, through peer-to-peer networks and edge computing technologies, participating nodes share model parameters, achieving model merging and aggregation. Compared to centralized machine learning methods, bee swarm learning has the following advantages:

[0004] 1. Data Privacy Protection: Bee swarm learning avoids centralized sharing of raw data through localized learning and parameter sharing. Participating nodes only need to share some model parameters, without sharing the original data, thus protecting data privacy and security.

[0005] 2. Highly efficient communication and computation: Bee swarm learning reduces data transmission and computational overhead through peer-to-peer networks and edge computing. Participating nodes train the model locally, sharing only model parameters, thereby reducing communication volume and computational load.

[0006] 3. Adaptive Model Aggregation: Bee colony learning allows participating nodes to adaptively aggregate models based on their weights.

[0007] However, the data measured by the aforementioned sensors is directly related to the privacy of drivers and vehicles, so it is very sensitive. Therefore, when training traffic models, it is necessary to ensure the privacy of this data, protect the privacy of drivers and vehicles, and achieve better traffic management and driving experience.

[0008] However, bee colony learning has the following drawbacks:

[0009] 1. Scalability: Bee colony learning requires intensive communication and collaboration among participating nodes. As the number of participating nodes increases, the complexity and management difficulty of the system also increase. This may pose a challenge to the scalability of the system, especially in dynamic environments where the joining and leaving of participating nodes may lead to network instability.

[0010] 2. Data heterogeneity: In bee colony learning, the data distribution, scale, and quality of different participating nodes may differ, which may lead to model bias or inaccuracy.

[0011] 3. Security and privacy risks: Bee colony learning involves the sharing and communication of model parameters among participating nodes, which may introduce security and privacy risks. Shared parameters that are not properly protected may be attacked by malicious parties, leading to the risk of information leakage or model tampering.

[0012] These drawbacks can compromise privacy and security when using existing bee colony learning for data processing and traffic prediction model training. Summary of the Invention

[0013] The purpose of this invention is to provide a traffic congestion prediction method, system, device, and medium based on bee colony learning, which can improve privacy and security during data processing and traffic prediction model training.

[0014] To achieve the above objectives, the present invention provides the following solution:

[0015] A traffic congestion prediction method based on bee colony learning includes:

[0016] Using vehicles as participating nodes in a swarm network, the robustness value of the local model of each participating node is determined based on the size of the traffic dataset of each participating node in the swarm network; the size of the traffic dataset is the total amount of data in the traffic dataset.

[0017] In the current aggregation round, the local data credibility of each participating node in the previous aggregation round is obtained based on the local data credibility of each participating node and the robustness value of the local model of each participating node.

[0018] Based on the machine learning evaluation metrics under the current aggregation round, the credibility of the local model of each participating node under the current aggregation round is obtained; the machine learning evaluation metrics include: true positives, false positives, and false negatives;

[0019] Based on the local data credibility of each participating node in the previous aggregation round, the credibility of the local model of each participating node in the previous aggregation round, and the loss function of each participating node in the traffic dataset corresponding to each participating node in the previous aggregation round, the training evaluation weights of the local model of each participating node in the current aggregation round are obtained.

[0020] The localized differential privacy technique is used to add noise to the model parameters of the local model at the current aggregation round number to obtain the noisy local model at the current aggregation round number;

[0021] Based on the local model of the current aggregation round, the local model training and evaluation weights of each participating node under the current aggregation round, and the noisy local model under the current aggregation round, the local model for the next aggregation round is obtained.

[0022] Determine whether the current aggregation round number has reached the set aggregation round number, and obtain the first determination result;

[0023] If the first judgment result is yes, then the local model under the next aggregation round is determined to be the traffic prediction model, and the traffic prediction model is used to predict traffic congestion.

[0024] If the first judgment result is negative, then update the aggregation round number and proceed to the next aggregation.

[0025] Optionally, based on the size of the traffic dataset of each participating node in the bee colony network, the robustness value of the local model of each participating node is determined, specifically including:

[0026] For any participating node N k According to the formula C(N) k ) = log m S Nk Calculate participant node N k The robustness value of the local model, where C(N) k ) represents the participating node N k The robustness value of the local model, where m represents the size of the largest traffic dataset with the total amount of data from all participating nodes, and S represents the robustness value of the local model. Nk Indicates the participating node N k The size of the traffic dataset.

[0027] Optionally, the local data credibility of each participating node in the current aggregation round is obtained based on the local data credibility of each participating node in the previous aggregation round and the robustness value of the local model of each participating node, specifically including:

[0028] For any participating node N k According to the formula Calculate the participating node N under the i-th aggregation round number. k The local data credibility, where Ci(Nk) represents the number of participating nodes N in the i-th aggregation round. k The credibility of local data, C(N) k ) represents the participating node N k The robustness value of the local model, C i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k The local data reliability, where n represents the total number of participating nodes in the bee colony network, C(N) j ) represents the robustness value of the local model of the participating node Nj.

[0029] Optionally, based on the machine learning evaluation metrics under the current aggregation round, the credibility of the local model of each participating node under the current aggregation round is obtained, specifically including:

[0030] For any participating node N k According to the formula Calculate the participating node N under the i-th aggregation round number k The credibility of the local model, where A i (N k ) represents the participating node N in the i-th aggregation round. k The reliability of the local model, β represents the first adjustment parameter, TP i Let FP represent the true instance under the i-th aggregation round number. i FN represents a false positive in the i-th aggregation round. i This represents a false counterexample in the i-th aggregation round.

[0031] Optionally, based on the local data credibility of each participating node in the previous aggregation round, the local model credibility of each participating node in the previous aggregation round, and the loss function of each participating node in the traffic dataset corresponding to each participating node in the previous aggregation round, the local model training evaluation weights of each participating node in the current aggregation round are obtained, specifically including:

[0032] For any participating node N k According to the formula

[0033] Calculate the participating node N under the i-th aggregation round number k The local model training evaluation weights, where W i (N k ) represents the participating node N in the i-th aggregation round. k Local model training evaluation weights, C i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k The credibility of local data, C i-1 (N j ) represents the participating node N in the (i-1)th aggregation round. j The credibility of local data, A i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k The credibility of the local model, A i-1 (N j ) represents the participating node N in the (i-1)th aggregation round. j The reliability of the local model, where n represents the total number of participating nodes in the bee colony network, α represents the second adjustment parameter, exp() represents the exponential function with base e, and L i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k At participant node N k The corresponding loss function W on the traffic dataset i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k Local model training evaluation weights.

[0034] Optionally, based on the local model of the current aggregation round, the local model training and evaluation weights of each participating node in the current aggregation round, and the noisy local model in the current aggregation round, the local model for the next aggregation round is obtained, specifically including:

[0035] According to the formula Calculate the local model for the next aggregation round, where M(i+1) represents the local model for the (i+1)th aggregation round, M(i) represents the local model for the ith aggregation round, n represents the total number of participating nodes in the bee colony network, and Wi(N) represents the local model for the next aggregation round. k ) represents the participating node N in the i-th aggregation round. k Local model training evaluation weights, ∠ represents the noisy local model under the i-th aggregation round, and · represents the multiplication operation.

[0036] A traffic congestion prediction system based on bee colony learning includes:

[0037] The robustness value calculation module uses vehicles as participating nodes in a swarm network and determines the robustness value of the local model of each participating node based on the size of the traffic dataset of each participating node in the swarm network; the size of the traffic dataset is the total amount of data in the traffic dataset.

[0038] The local data credibility calculation module is used to obtain the local data credibility of each participating node in the current aggregation round based on the local data credibility of each participating node in the previous aggregation round and the robustness value of the local model of each participating node.

[0039] The local model credibility calculation module is used to obtain the credibility of the local models of each participating node under the current aggregation round based on the machine learning evaluation metrics under the current aggregation round; the machine learning evaluation metrics include: true positives, false positives, and false negatives;

[0040] The local model training evaluation weight calculation module is used to obtain the local model training evaluation weight of each participating node in the current aggregation round based on the local data credibility of each participating node in the previous aggregation round, the local model credibility of each participating node in the previous aggregation round, and the loss function of each participating node in the traffic dataset corresponding to each participating node in the previous aggregation round.

[0041] The localized differential privacy module is used to add noise to the model parameters of the local model at the current aggregation round number using localized differential privacy technology to obtain a noisy local model at the current aggregation round number.

[0042] The local model update module is used to obtain the local model for the next aggregation round based on the local model of the current aggregation round, the local model training evaluation weights of each participating node under the current aggregation round, and the noisy local model under the current aggregation round.

[0043] The judgment module is used to determine whether the current number of aggregation rounds has reached the set number of aggregation rounds, and obtain the first judgment result;

[0044] The traffic prediction model determination module is used to determine the local model under the next aggregation round as the traffic prediction model if the first judgment result is yes. The traffic prediction model is used to predict traffic congestion.

[0045] The iteration module is used to update the aggregation round number and proceed to the next aggregation if the first judgment result is negative.

[0046] Optionally, the robustness value calculation module specifically includes:

[0047] Robustness value calculation unit, used for any participant node N k According to the formula C(N) k ) = log m S Nk Calculate participant node N k The robustness value of the local model, where C(N) k ) represents the participating node N k The robustness value of the local model, where m represents the size of the largest traffic dataset with the total amount of data from all participating nodes, and S represents the robustness value of the local model. Nk Indicates the participating node N k The size of the traffic dataset.

[0048] An electronic device, comprising:

[0049] A memory and a processor, the memory for storing a computer program, the processor for running the computer program to cause the electronic device to perform the traffic congestion prediction method based on bee colony learning as described above.

[0050] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned traffic congestion prediction method based on bee colony learning.

[0051] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0052] Using vehicles as participating nodes in a swarm network, the robustness value of each participating node's local model is determined based on the size of its traffic dataset. At the current aggregation round, the local data reliability of each participating node is obtained based on the local data reliability and the robustness value of its local model at the previous aggregation round. The reliability of each participating node's local model at the current aggregation round is obtained based on the machine learning evaluation metric. The training and evaluation weights of each participating node's local model at the current aggregation round are obtained based on the local data reliability, the local model reliability, and the loss function of each participating node in the traffic dataset corresponding to its respective participating node at the previous aggregation round. Localized differential mapping is then employed. Privacy technology adds noise to the model parameters of the local model at the current aggregation round to obtain a noisy local model at the current aggregation round. Based on the local model at the current aggregation round, the local model training and evaluation weights of each participating node at the current aggregation round, and the noisy local model at the current aggregation round, the local model for the next aggregation round is obtained. It is determined whether the current aggregation round has reached the set aggregation round number, and a first judgment result is obtained. If the first judgment result is yes, the local model for the next aggregation round is determined to be a traffic prediction model, which is used to predict traffic congestion. If the first judgment result is no, the aggregation round number is updated, and the next aggregation begins. This invention uses localized differential privacy technology to add noise to the model parameters of the local model at the current aggregation round to protect the privacy of intermediate parameters, which can improve privacy and security during data processing and traffic prediction model training. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0054] Figure 1 A flowchart illustrating the traffic congestion prediction method based on bee colony learning provided in an embodiment of the present invention;

[0055] Figure 2 The flowchart illustrates the execution of the traffic congestion prediction method based on bee colony learning provided in this embodiment of the invention. Detailed Implementation

[0056] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0057] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0058] The traffic congestion prediction method based on bee colony learning provided by this invention (1) requires each participating node to register through a blockchain smart contract in the initial stage of the system. This process ensures the legitimacy of each participating node's identity. Each participating node in bee colony learning is clearly defined, and only legally authorized participating nodes can execute transactions. Each participating node records information from the smart contract, such as acquiring the model and immediately executing local training after meeting the synchronization conditions. After registering with the blockchain smart contract, participating nodes do not need to upload local data; they only need to upload a local data summary (data type, data format, data size, etc.) to the blockchain record. Subsequently, the system distributes a unique identifier ID to each participating node.

[0059] (2) During the transaction request processing phase, participating nodes send transaction requests to the system, such as requesting the latest global model for a new round of local training. Transaction requests are recorded on the blockchain. Transaction information initiated by each participating node is audited, packaged into blocks, and broadcast to the entire network for verification. After verification, the information is added to the blockchain system.

[0060] (3) Swarm Learning Phase. The Swarm API facilitates parameter exchange between nodes and merges and updates the model before the next training round. After the final round of parameter sharing and aggregation, the framework checks if the stopping criterion has been met. If it has, training stops; otherwise, it resumes. During model updates, localized differential privacy is used to add noise to intermediate parameters to ensure their privacy. The system adaptively aggregates based on the local model training quality and reliability, using aggregation weights.

[0061] like Figure 1 As shown, this embodiment of the invention provides a traffic congestion prediction method based on bee colony learning, including:

[0062] Step 101: Using vehicles as participating nodes in the swarm network, determine the robustness value of the local model for each participating node based on the size of its traffic dataset. The traffic dataset refers to traffic data collected by vehicle sensors, including data from different vehicles or traffic equipment, such as vehicle speed, acceleration, location information, traffic light status, and traffic flow. Size refers to the amount of data. This data will be used for local training of each vehicle during the swarm learning process. In SL, the larger the dataset during training, the better the model's robustness. Therefore, if each vehicle's local dataset is rich and diverse enough, the swarm learning algorithm will have more opportunities to learn more comprehensive and accurate traffic patterns and features, thereby improving the model's robustness. Therefore, during global model aggregation in the swarm network SN, models trained on larger datasets can be assigned higher weights.

[0063] Step 102: Under the current aggregation round, based on the local data credibility of each participating node under the previous aggregation round and the robustness value of the local model of each participating node, obtain the local data credibility of each participating node under the current aggregation round.

[0064] Step 103: Based on the machine learning evaluation metrics under the current aggregation round, obtain the credibility of the local model of each participating node under the current aggregation round; the machine learning evaluation metrics include: true positives, false positives, and false negatives.

[0065] Step 104: Based on the local data credibility of each participating node in the previous aggregation round, the local model credibility of each participating node in the previous aggregation round, and the loss function of each participating node in the traffic dataset corresponding to each participating node in the previous aggregation round, obtain the local model training evaluation weights of each participating node in the current aggregation round.

[0066] Step 105: Use localized differential privacy technology to add noise to the model parameters of the local model at the current aggregation round to obtain the noisy local model at the current aggregation round.

[0067] Step 106: Based on the local model of the current aggregation round, the local model training evaluation weights of each participating node in the current aggregation round, and the noisy local model in the current aggregation round, obtain the local model for the next aggregation round.

[0068] Step 107: Determine whether the current number of aggregation rounds has reached the set number of aggregation rounds, and obtain the first determination result.

[0069] Step 108: If the first judgment result is yes, then determine the local model under the next aggregation round as the traffic prediction model, which is used to predict traffic congestion.

[0070] Step 109: If the first judgment result is negative, update the aggregation round number and proceed to the next aggregation.

[0071] In practical applications, traffic prediction models can be used to:

[0072] 1. Traffic congestion prediction: Traffic prediction models can predict traffic flow and congestion in different areas of a city based on traffic data collected by vehicle sensors. Such predictions can help drivers avoid congested road sections, improve traffic efficiency, and reduce commuting time.

[0073] 2. Route optimization: Based on the prediction results of traffic prediction models, intelligent transportation systems can provide drivers with more optimized driving routes. These routes will take into account traffic flow, congestion, and estimated travel time, helping drivers choose the fastest route.

[0074] 3. Fuel efficiency optimization: Traffic prediction models can analyze driving behavior and vehicle conditions to provide optimization suggestions for fuel consumption, helping car owners save fuel costs and reduce environmental impact.

[0075] 4. Urban planning and traffic management: By analyzing data from traffic forecasting models, urban planners and traffic management departments can better understand traffic demand and mobility, providing decision support for urban planning and traffic policy formulation.

[0076] In practical applications, before determining the robustness value of the local model of each participating node based on the size of the traffic dataset of each participating node in the bee colony network, the following steps are also included:

[0077] Each participating node is registered with a smart contract to obtain its identifier. Subsequently, the identifier is used to determine whether each participating node can execute the model training steps.

[0078] In practical applications, the robustness value of the local model of each participating node is determined based on the size of the traffic dataset of each participating node in the bee colony network, specifically including:

[0079] For any participating node N k According to the formula C(N) k ) = log m S Nk (1) Calculate the participating node N k The robustness value of the local model, where C(N) k ) represents the participating node N kThe robustness value of the local model, where m represents the size of the largest traffic dataset with the total amount of data from all participating nodes, and S represents the robustness value of the local model. Nk Indicates the participating node N k The size of the traffic dataset.

[0080] In practical applications, the local data credibility of each participating node in the current aggregation round is obtained based on the local data credibility of each participating node in the previous aggregation round and the robustness value of the local model of each participating node. Specifically, this includes:

[0081] For any participating node N k According to the formula Calculate the participating node N under the i-th aggregation round number. k The credibility of local data, of which C i (N k ) represents the participating node N in the i-th aggregation round. k The credibility of local data, C(N) k ) represents the participating node N k The robustness value of the local model, C i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k The local data reliability, where n represents the total number of participating nodes in the bee colony network, C(N) j ) represents the participating node N j Robustness value of the local model.

[0082] In practical applications, the credibility of the local models of each participating node under the current aggregation round is obtained based on the machine learning evaluation metrics, specifically including:

[0083] For any participating node N k According to the formula Calculate the participating node N under the i-th aggregation round number k The credibility of the local model, where A i (N k ) represents the participating node N in the i-th aggregation round. k The reliability of the local model, β represents the first adjustment parameter, TP i Let FP represent the true instance under the i-th aggregation round number. i FN represents a false positive in the i-th aggregation round. iTP, FP, and FN represent false negatives in the i-th aggregation round. TP, FP, and FN are directly obtained from the machine learning of the participating nodes in each round and are machine learning metrics. True Positive (TP): The number of positive samples (actually positive) correctly predicted as positive by the model. False Positive (FP): The number of negative samples (actually negative) incorrectly predicted as positive by the model. False Negative (FN): The number of positive samples (actually positive) incorrectly predicted as negative by the model.

[0084] In practical applications, the local model training evaluation weights of each participating node in the current aggregation round are obtained based on the local data reliability of each participating node in the previous aggregation round, the local model reliability of each participating node in the previous aggregation round, and the loss function of each participating node in the traffic dataset corresponding to each participating node in the previous aggregation round. Specifically, this includes:

[0085] For any participating node N k According to the formula Calculate the participating node N under the i-th aggregation round number k The local model training evaluation weights, where W i (N k ) represents the participating node N in the i-th aggregation round. k Local model training evaluation weights, C i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k The credibility of local data, C i-1 (N j ) represents the participating node N in the (i-1)th aggregation round. j The credibility of local data, A i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k The credibility of the local model, A i-1 (N j ) represents the participating node N in the (i-1)th aggregation round. j The reliability of the local model, where n represents the total number of participating nodes in the bee colony network, α represents the second adjustment parameter, exp() represents the exponential function with base e, and L i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k At participant node N k The corresponding loss function W on the traffic dataset i-1 (N k) represents the participating node N in the (i-1)th aggregation round. k Local model training evaluation weights, This is the normalization coefficient, used to normalize the weights and ensure that the sum of the weights of all participating nodes is 1.

[0086] In practical applications, based on the local model of the current aggregation round, the local model training and evaluation weights of each participating node under the current aggregation round, and the noisy local model under the current aggregation round, the local model for the next aggregation round is obtained, specifically including:

[0087] According to the formula Calculate the local model for the next aggregation round, where M(i+1) represents the local model for the (i+1)th aggregation round, M(i) represents the local model for the ith aggregation round, n represents the total number of participating nodes in the bee colony network, and W... i (N k ) represents the participating node N in the i-th aggregation round. k Local model training evaluation weights, ∠ represents the noisy local model under the i-th aggregation round, and · represents the multiplication operation.

[0088] In practical applications, localized differential privacy techniques are used to add noise to the model parameters of the local model at the current aggregation round, resulting in a noisy local model at the current aggregation round. Specifically, localized differential privacy techniques are applied according to the formula... intermediate parameter m during model update k (t) Laplace noise is added to ensure the privacy of intermediate parameters, where m k (i) represents the model parameters of the local model under the i-th aggregation round. Here are the model parameters of the local model after adding noise in the (i+1)th aggregation round, where Lap() is the Laplacian noise, ε is the total privacy budget (preset), and Δf is the total privacy budget. LS For local sensitivity, α i This is the learning rate (step size), which can be adjusted automatically. L i (m k ) is the participating node N k The loss function used to train the model on its dataset is a loss function that measures the degree of inconsistency between the model's predictions and the true values. The loss function is optimized by updating the parameters in the negative direction of the objective function's gradient. k The parameters representing model training, ΔL i (m k ) is the participating node N k The increment of the loss function for model training on its dataset.

[0089] The present invention has the following technical effects:

[0090] 1. The total privacy budget cost added to the entire model can be broken down into the sum of privacy budget additions in each iteration. Each model iteration satisfies ε-differential privacy. Let the total number of iterations be T and the total privacy budget be ε, then the privacy budget added in each iteration is ε. i The privacy budget is ε / T, because differential privacy has sequential composability, and the privacy budget for each round is ε. i Since the sum is less than or equal to ε, adding noise to the model satisfies ε-differential privacy. In bee colony learning, when participating nodes train the model and share parameters, each participating node adds a certain amount of Laplace noise to its local model parameters. The purpose of this is to confuse and obscure the model parameters, making it impossible for malicious parties to accurately infer the individual data of the participating nodes. Due to the randomness of Laplace noise, attackers cannot accurately deduce the individual data of the participating nodes, thus increasing the difficulty of privacy leakage. Therefore, adding Laplace noise to intermediate parameters provides a certain degree of privacy security for the intermediate parameters.

[0091] 2. To ensure the security and trustworthiness of swarm learning, this invention employs blockchain technology. Blockchain provides a decentralized consensus mechanism, ensuring the trustworthiness of interactions between participating nodes and preventing potential fraud. This invention, through blockchain smart contract registration, ensures the legitimacy of participating node identities and the trustworthiness of the network. It verifies the identities of participating nodes using blockchain technology and ensures that only legally authorized nodes can execute transactions, enhancing system security and resistance to poisoning attacks (a poisoning attack is a security threat against machine learning models that aims to manipulate or tamper with training data, causing the model to produce incorrect results during inference or prediction, or to be controlled by an attacker). This prevents malicious parties from disrupting the results of collaborative learning through malicious data or manipulation of model parameters. Therefore, this invention overcomes the limitations of traditional methods and provides a more efficient, secure, and privacy-preserving collaborative solution for machine learning tasks.

[0092] 3. This invention can ensure the accuracy and reliability of the model training process, protect the privacy of intermediate parameters, ensure the legitimacy of the identities of participating nodes, and prevent the risk of poisoning attacks.

[0093] 4. This invention enables participating nodes to adaptively aggregate models based on their weights, dynamically adjust weights according to the credibility of participating nodes and model quality, and adaptively aggregate models based on their weights to better reflect the contributions and data characteristics of different participating nodes.

[0094] 5. Differential privacy is a privacy protection technique that aims to protect individual privacy while allowing limited, statistically significant analysis and inference of data. The core idea of ​​differential privacy is to introduce a suitable amount of noise into the data, making it impossible to accurately infer the information of sensitive individuals. By adding noise during data publishing or querying, differential privacy provides a mathematical guarantee that even if an attacker possesses background knowledge beyond the individual data and powerful computational capabilities, they cannot accurately infer the privacy information of a specific individual. The Laplace distribution is a special form of exponential distribution, with as the location parameter and as the scale parameter. Its probability distribution function is symmetric about and reaches its maximum value. The Laplace mechanism adds noise to the query results, making the originally deterministic query results conform to a Laplace distribution. If an attacker cannot distinguish the probability distribution of the query results of two adjacent nodes, the purpose of differential privacy is achieved. This invention employs localized differential privacy technology to protect the privacy of intermediate parameters. By adding Laplace noise, the privacy of intermediate parameters is protected, and the privacy of individual data is protected when sharing model parameters, preventing malicious parties from inferring sensitive information through parameters.

[0095] like Figure 2 As shown, the specific execution flow of the above method includes:

[0096] Step 1: Participating node users register smart contracts, and the system distributes unique identifiers (IDs) to participating node users.

[0097] Step 2: Participating nodes send training requests to the system.

[0098] Step 3: After the system approves the training request, the participating nodes train the model.

[0099] Step 4: Calculate the local data credibility C for the i-th training round according to formula (2). i (N k ).

[0100] Step 5: Calculate the model credibility A in the i-th training round according to formula (3). i (N k ).

[0101] Step 6: Calculate the training evaluation weight W of the participating nodes in the i-th round of training according to formula (4). i (N k ).

[0102] Step 7: According to formula (6), add Laplacian noise (Lap()) to the intermediate parameters, where ε is the privacy budget. The smaller the privacy budget, the higher the degree of privacy protection. This value can be defined by the user. Perform adaptive model aggregation according to formula (5) and update model M(i) to obtain M(i+1).

[0103] Step 8: If the model reaches the stopping condition or the number of aggregation rounds reaches the maximum, output model M(i+1); otherwise, repeat steps 3 to 8 to continue the next round of model training.

[0104] This invention also provides an embodiment to demonstrate that the above method can resist poisoning attacks and satisfy ε-differential privacy:

[0105] (1) Can resist poisoning attacks

[0106] Assume participant node N k Poisoning attacks, where labels are altered or data is tampered with, can significantly increase the model's loss. We will use ΔL. i This represents the increment of the model loss. When the number of participating nodes N... k When the virus is poisoned in the i-th round, let it be denoted as N. k The weight update formula for the next round i+1 can be rewritten as:

[0107]

[0108] If N k The formula for updating weights that are not attacked remains unchanged:

[0109]

[0110] Now, we will compare N k 'and N k The weight ratio in round i+1:

[0111]

[0112] Simplified as follows:

[0113]

[0114] Based on the properties of exponential functions, and given that the data reliability and model reliability in round i are determined by the previous round, we have:

[0115]

[0116] Right now:

[0117]

[0118] Because the increment ΔL of the model lossi >0, according to the properties of exponential functions, exp(-α·ΔL) i Since ) < 1, we can conclude that:

[0119]

[0120] In other words, the participating node N that was poisoned was... k In the next round of aggregation weights compared to N that were not attacked k This will reduce the impact of poisoning attacks on the final model. Therefore, the adaptive bee colony learning model aggregation algorithm can minimize the influence of poisoned attack participants on the final model. Each round of weight updates is adjusted based on data quality and reliability, thereby resisting poisoning attacks and ensuring the accuracy and reliability of the model.

[0121] (2) Satisfies ε-differential privacy

[0122] The local differential privacy method satisfies ε-differential privacy and is denoted as Algorithm L. In this invention, it is denoted as Algorithm A.

[0123] Let S be the set of all possible output ranges of algorithm L, and t be any one of the output ranges of A. Given two adjacent datasets D and D', which differ from each other by at most one data point, i.e., |DΔD'|≤1, we have:

[0124]

[0125] L() is a local differential privacy method, where the probability Pr{*} is controlled by its internal algorithm, reflecting the risk of privacy disclosure. A() is an adaptive bee colony learning model aggregation algorithm with Laplace noise. The privacy budget parameter, ε, represents the degree of privacy protection; therefore, the adaptive bee colony learning model aggregation algorithm with Laplace noise satisfies ε-differential privacy.

[0126] In view of the above method, this embodiment of the invention provides a traffic congestion prediction system based on bee colony learning, including:

[0127] The robustness value calculation module takes vehicles as participating nodes in the swarm network and determines the robustness value of the local model of each participating node based on the size of the traffic dataset of each participating node in the swarm network.

[0128] The local data credibility calculation module is used to obtain the local data credibility of each participating node in the current aggregation round based on the local data credibility of each participating node in the previous aggregation round and the robustness value of the local model of each participating node.

[0129] The local model credibility calculation module is used to obtain the credibility of the local model of each participating node under the current aggregation round based on the machine learning evaluation index under the current aggregation round; the machine learning evaluation index includes: true positives, false positives and false negatives.

[0130] The local model training evaluation weight calculation module is used to obtain the local model training evaluation weight of each participating node in the current aggregation round based on the local data credibility of each participating node in the previous aggregation round, the local model credibility of each participating node in the previous aggregation round, and the loss function of each participating node in the traffic dataset corresponding to each participating node in the previous aggregation round.

[0131] The localized differential privacy module is used to add noise to the model parameters of the local model at the current aggregation round using localized differential privacy technology to obtain a noisy local model at the current aggregation round.

[0132] The local model update module is used to obtain the local model for the next aggregation round based on the local model of the current aggregation round, the local model training evaluation weights of each participating node in the current aggregation round, and the noisy local model in the current aggregation round.

[0133] The judgment module is used to determine whether the current number of aggregation rounds has reached the set number of aggregation rounds, and obtain the first judgment result.

[0134] The traffic prediction model determination module is used to determine the local model under the next aggregation round as the traffic prediction model if the first judgment result is yes. The traffic prediction model is used to predict traffic congestion.

[0135] The iteration module is used to update the aggregation round number and proceed to the next aggregation if the first judgment result is negative.

[0136] In practical applications, the robustness value calculation module specifically includes:

[0137] Robustness value calculation unit, used for any participant node N k According to the formula C(N) k ) = log m S Nk Calculate participant node N k The robustness value of the local model, where C(N) k ) represents the participating node N k The robustness value of the local model, where m represents the size of the largest traffic dataset with the total amount of data from all participating nodes, and S represents the robustness value of the local model. Nk Indicates the participating node N k The size of the traffic dataset.

[0138] This invention also provides an electronic device, comprising:

[0139] A memory and a processor, the memory for storing a computer program, the processor for running the computer program to cause the electronic device to perform the traffic congestion prediction method based on bee colony learning as described above.

[0140] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned traffic congestion prediction method based on bee colony learning.

[0141] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.

[0142] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A traffic congestion prediction method based on bee colony learning, characterized in that, include: Using vehicles as participating nodes in a swarm network, the robustness value of the local model of each participating node is determined based on the size of the traffic dataset of each participating node in the swarm network; the size of the traffic dataset is the total amount of data in the traffic dataset. In the current aggregation round, the local data credibility of each participating node in the previous aggregation round is obtained based on the local data credibility of each participating node and the robustness value of the local model of each participating node. Based on the machine learning evaluation metrics under the current aggregation round, the credibility of the local model of each participating node under the current aggregation round is obtained; The machine learning evaluation metrics include: true positives, false positives, and false negatives; Based on the local data credibility of each participating node in the previous aggregation round, the credibility of the local model of each participating node in the previous aggregation round, and the loss function of each participating node in the traffic dataset corresponding to each participating node in the previous aggregation round, the training evaluation weights of the local model of each participating node in the current aggregation round are obtained. The localized differential privacy technique is used to add noise to the model parameters of the local model at the current aggregation round number to obtain the noisy local model at the current aggregation round number; Based on the local model of the current aggregation round, the local model training and evaluation weights of each participating node under the current aggregation round, and the noisy local model under the current aggregation round, the local model for the next aggregation round is obtained. Determine whether the current aggregation round number has reached the set aggregation round number, and obtain the first determination result; If the first judgment result is yes, then the local model under the next aggregation round is determined to be the traffic prediction model, and the traffic prediction model is used to predict traffic congestion. If the first judgment result is negative, then update the aggregation round number and proceed to the next aggregation.

2. The traffic congestion prediction method based on bee colony learning according to claim 1, characterized in that, Based on the size of the traffic dataset of each participating node in the bee colony network, the robustness value of the local model of each participating node is determined, specifically including: For any participating node N k According to the formula C(N) k ) = log m S Nk Calculate participant node N k The robustness value of the local model, where C(N) k ) represents the participating node N k The robustness value of the local model, where m represents the size of the largest traffic dataset with the total amount of data from all participating nodes, and S represents the robustness value of the local model. Nk Indicates the participating node N k The size of the traffic dataset.

3. The traffic congestion prediction method based on bee colony learning according to claim 1, characterized in that, Based on the local data credibility of each participating node in the previous aggregation round and the robustness value of the local model of each participating node, the local data credibility of each participating node in the current aggregation round is obtained, specifically including: For any participating node N k According to the formula Calculate the participating node N under the i-th aggregation round number. k The credibility of local data, of which C i (N k ) represents the participating node N in the i-th aggregation round. k The credibility of local data, C(N) k ) represents the participating node N k The robustness value of the local model, C i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k The local data reliability, where n represents the total number of participating nodes in the bee colony network, C(N) j ) represents the participating node N j Robustness value of the local model.

4. The traffic congestion prediction method based on bee colony learning according to claim 1, characterized in that, Based on the machine learning evaluation metrics under the current aggregation round, the credibility of the local model of each participating node under the current aggregation round is obtained, specifically including: For any participating node N k According to the formula Calculate the participating node N under the i-th aggregation round number k The credibility of the local model, where A i (N k ) represents the participating node N in the i-th aggregation round. k The reliability of the local model, β represents the first adjustment parameter, TP i Let FP represent the true instance under the i-th aggregation round number. i FN represents a false positive in the i-th aggregation round. i This represents a false counterexample in the i-th aggregation round.

5. The traffic congestion prediction method based on bee colony learning according to claim 1, characterized in that, Based on the local data reliability of each participating node in the previous aggregation round, the local model reliability of each participating node in the previous aggregation round, and the loss function of each participating node in the traffic dataset corresponding to each participating node in the previous aggregation round, the local model training evaluation weights of each participating node in the current aggregation round are obtained, specifically including: For any participating node N k According to the formula Calculate the participating node N under the i-th aggregation round number k The local model training evaluation weights, where W i (N k ) represents the participating node N in the i-th aggregation round. k Local model training evaluation weights, C i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k The credibility of local data, C i-1 (N j ) represents the participating node N in the (i-1)th aggregation round. j The credibility of local data, A i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k The credibility of the local model, A i-1 (N j ) represents the participating node N in the (i-1)th aggregation round. j The reliability of the local model, where n represents the total number of participating nodes in the bee colony network, α represents the second adjustment parameter, exp() represents the exponential function with base e, and L i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k At participant node N k The corresponding loss function W on the traffic dataset i-1 (N k ) represents the participating node N in the (i-1)th aggregation round. k Local model training evaluation weights.

6. The traffic congestion prediction method based on bee colony learning according to claim 1, characterized in that, Based on the local model of the current aggregation round, the local model training and evaluation weights of each participating node in the current aggregation round, and the noisy local model of the current aggregation round, the local model for the next aggregation round is obtained, specifically including: According to the formula Calculate the local model for the next aggregation round, where M(i+1) represents the local model for the (i+1)th aggregation round, M(i) represents the local model for the ith aggregation round, n represents the total number of participating nodes in the bee colony network, and W... i (N k ) represents the participating node N in the i-th aggregation round. k Local model training evaluation weights, ∠ represents the noisy local model under the i-th aggregation round, and · represents the multiplication operation.

7. A traffic congestion prediction system based on bee colony learning, characterized in that, include: The robustness value calculation module uses vehicles as participating nodes in a swarm network and determines the robustness value of the local model of each participating node based on the size of the traffic dataset of each participating node in the swarm network; the size of the traffic dataset is the total amount of data in the traffic dataset. The local data credibility calculation module is used to obtain the local data credibility of each participating node in the current aggregation round based on the local data credibility of each participating node in the previous aggregation round and the robustness value of the local model of each participating node. The local model credibility calculation module is used to obtain the credibility of the local model of each participating node under the current aggregation round based on the machine learning evaluation index under the current aggregation round. The machine learning evaluation metrics include: true positives, false positives, and false negatives; The local model training evaluation weight calculation module is used to obtain the local model training evaluation weight of each participating node in the current aggregation round based on the local data credibility of each participating node in the previous aggregation round, the local model credibility of each participating node in the previous aggregation round, and the loss function of each participating node in the traffic dataset corresponding to each participating node in the previous aggregation round. The localized differential privacy module is used to add noise to the model parameters of the local model at the current aggregation round number using localized differential privacy technology to obtain a noisy local model at the current aggregation round number. The local model update module is used to obtain the local model for the next aggregation round based on the local model of the current aggregation round, the local model training evaluation weights of each participating node under the current aggregation round, and the noisy local model under the current aggregation round. The judgment module is used to determine whether the current number of aggregation rounds has reached the set number of aggregation rounds, and obtain the first judgment result; The traffic prediction model determination module is used to determine the local model under the next aggregation round as the traffic prediction model if the first judgment result is yes. The traffic prediction model is used to predict traffic congestion. The iteration module is used to update the aggregation round number and proceed to the next aggregation if the first judgment result is negative.

8. The traffic congestion prediction system based on bee colony learning according to claim 7, characterized in that, The robustness value calculation module specifically includes: Robustness value calculation unit, used for any participant node N k According to the formula C(N) k ) = log m S Nk Calculate participant node N k The robustness value of the local model, where C(N) k ) represents the participating node N k The robustness value of the local model, where m represents the size of the largest traffic dataset with the total amount of data from all participating nodes, and S represents the robustness value of the local model. Nk Indicates the participating node N k The size of the traffic dataset.

9. An electronic device, characterized in that, include: A memory and a processor, the memory for storing a computer program, the processor for running the computer program to cause the electronic device to perform the traffic congestion prediction method based on bee colony learning according to any one of claims 1 to 6.

10. A computer-readable storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the traffic congestion prediction method based on bee colony learning as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Methods and systems for intelligent collection and analysis of vehicle data

    US20190025813A1

  • Signal randomization method and device of communication apparatus

    WO2022025321A1