Urban traffic detection system and method based on bidirectional memory federated learning, medium and electronic equipment
By introducing a two-way memory mechanism into the federated learning system, the personalized model optimization of each traffic node is solved, the problem of model homogeneity in the existing system is significantly improved, the detection accuracy and response speed are enhanced, and the data privacy protection capabilities are enhanced.
Patent Information
- Application Number
- CN202411968825.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-12-30
AI Technical Summary
The existing federated learning system has the problem of model homogeneity in urban road state and abnormal detection, and cannot be personalized to adjust according to the dynamic heterogeneous characteristics of each traffic node, resulting in insufficient detection accuracy and response speed.
Using federated learning technology based on two-way memory, the data acquisition module, global prediction head update module, forward inheritance module and comparison learning module are used to realize personalized model optimization of each traffic node. On the premise of ensuring data privacy, the system establishes an efficient knowledge migration path between nodes and servers through a two-way memory mechanism to optimize the adaptability of node traffic model.
It significantly improves the accuracy and real-timeness of each node in abnormal event detection, reduces the burden on the central server, improves the system's monitoring effect in complex traffic scenarios, and enhances data privacy protection capabilities.
Smart Images

Figure CN119989111A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of federated learning, urban road status and anomaly detection, and in particular to an urban road status and anomaly detection system, method, medium and electronic equipment based on bidirectional memory federated learning. Background Art
[0002] The rapid development of smart transportation and smart cities has made how to efficiently and accurately monitor the status of urban roads, especially the rapid detection of abnormal events such as traffic accidents, one of the core requirements for the optimization of current smart transportation systems. Road status and anomaly detection systems based on federated learning have gradually emerged and made significant progress in privacy protection. However, such systems still have obvious deficiencies in anomaly detection accuracy. The core problem lies in the homogeneity of the models of various traffic nodes, the lack of the ability to make personalized adjustments based on the dynamic heterogeneous characteristics of the nodes, and the inability to meet the needs of each node for adaptive adjustment of dynamic heterogeneous characteristics.
[0003] For example, at high-traffic intersections, the system needs to improve the sensitivity of accident detection through an adaptive adjustment mechanism for traffic density in order to more quickly identify traffic accidents. On some low-traffic sections, especially at night or at intersections with poor lighting, the system needs to enhance the dynamic perception sensitivity of low-speed vehicles based on an adaptive adjustment mechanism for lighting intensity. In a complex and changing urban traffic environment, homogeneous models are insufficient in detection accuracy and response speed, making it difficult to achieve a rapid response to abnormal events.
[0004] Therefore, for this field, how to provide a detection technology that not only needs to be able to adapt to the global environment, but also should make corresponding adjustments according to the dynamic heterogeneous characteristics of each traffic node, so that it can be personalized for the dynamic heterogeneous characteristics such as traffic density characteristics and lighting intensity characteristics of different nodes, so as to improve the accuracy and real-time performance of the system in abnormal event detection, has become an urgent problem to be solved. Summary of the invention
[0005] In order to effectively solve the problem of homogeneity of traffic models of nodes in the prior art, the present invention provides a system, method, medium and electronic device for urban road status and anomaly detection based on two-way memory federated learning. By adopting the technical solution of the present invention, each traffic node can realize personalized model optimization according to its own traffic characteristics and environmental conditions, thereby significantly improving the accuracy and real-time performance of different nodes in abnormal event detection.
[0006] In addition, the present invention uses federated learning technology to achieve distributed model training while ensuring the privacy of public data, avoiding the risk of privacy leakage when all public data is concentrated on the server. At the same time, through the two-way memory mechanism, the present invention establishes a more efficient knowledge transfer path between nodes and servers, making node model updates faster and detection more accurate. Furthermore, the two-way memory strategy designed by the present invention further optimizes the adaptability of the node traffic model to dynamic traffic data, which can significantly reduce the burden on the central server and improve the monitoring effect of the system in high-frequency and complex traffic scenarios.
[0007] The present invention adopts the following technical scheme:
[0008] An urban traffic detection system based on bidirectional memory federated learning includes the following modules:
[0009] Data collection module: This module is deployed at various nodes of urban roads. It is responsible for collecting data from equipment on each road section in real time, converting the raw data into prototype representation vectors through the extractor, and then uploading these vectors to the city computer room server.
[0010] Global prediction head update module: This module is located in the city computer room server and is responsible for receiving the mean prototype representation vectors uploaded from each road node. Based on these uploaded prototype representation vectors, the module optimizes and updates the global prediction head and generates a new global prediction head model to provide global guidance for subsequent node training.
[0011] Forward inheritance module: This module is deployed at each node of the urban road, and it maintains a hypernetwork model internally to store and transmit information from the historical training process. The hypernetwork model stacks all the local prediction heads generated by the road node in the previous training rounds, thereby integrating the historical learning experience of the node. These stacked historical information provides the basis for generating the intermediate prediction heads of the current round, so that the model can make full use of past knowledge during the current training, improve learning efficiency and prediction accuracy. Through the forward memory mechanism, the model can dynamically absorb and utilize historical experience, thereby avoiding dependence on outdated information and enhancing the timing and continuity of the model. In the road traffic status and anomaly detection tasks, the forward memory mechanism effectively improves the generalization ability of the model in different environments and significantly improves the accuracy and robustness of the prediction.
[0012] Contrastive learning module: This module is deployed at each node of the urban road, aiming to achieve knowledge transfer between the global prediction head and the intermediate prediction head, and work together with the forward inheritance module to jointly realize the bidirectional memory mechanism. Specifically, through the contrastive learning method, this module calculates the contrast loss between the global prediction head and the intermediate prediction head generated by the forward inheritance module, and constructs a joint loss function in combination with the supervised learning loss. If the joint loss is less than the threshold, the training is completed, and a personalized anomaly detection model for local dynamic heterogeneous features is obtained. The local real-time road data is input into the model to obtain the local real-time road detection status. During the training process, the global prediction head represents the model's understanding and generalization ability of global data, while the intermediate prediction head carries the knowledge accumulated in the historical training of the road node. Through contrastive learning, the model promotes the consistency of the two in the feature space, ensuring that the two can share valuable information in similar tasks or environments. This contrastive optimization process not only strengthens knowledge transfer, but also improves the personalized adaptability of the model through the optimization of the loss function, so that the local model can better meet the needs of dynamic heterogeneous features.
[0013] The combination of the contrastive learning module and the forward inheritance module constitutes a bidirectional memory mechanism of the system. In the forward inheritance module, historical information is stacked into the intermediate prediction head through the hypernetwork, while in the contrastive learning module, the comparison between the global prediction head and the intermediate prediction head further deepens the transfer and update of knowledge. The two work together to enable the local prediction head of each road node to better integrate the global perspective and historical experience, significantly improving the accuracy and robustness of the model in road state and anomaly detection tasks.
[0014] Preferably, the data acquisition module is used to extract the prototype representation vector, specifically as follows:
[0015] In the detection system, N road nodes are represented by a set C = {C1, C2, C3, … ,C N} represents that the local data set of each road node is a set D = {D1, D2, D3, … ,D N}, where N represents a natural number greater than zero, C i represents the i-th road node, D i Indicates the corresponding local dataset.
[0016] For any road node C i For example, Figure 2 As shown in step ①, its local dataset D i Extracted into prototype representation vector R by local extractor i Specifically, node C i A sample image of (category y) will be extracted as the corresponding prototype representation vector Where y is the image label category, m is the index of the image sample in category y, 0≤m<|M|, |M| indicates that category y is in the road node C i These prototype representation vectors condense the feature information of each type of sample, which is extracted from the local dataset by the local extractor, and finally forms a prototype representation vector that can represent the category characteristics.
[0017] Next, for category y, calculate its mean prototype representation vector The mean prototype representation vector can be regarded as the node C i The overall feature representation of category y:
[0018]
[0019] Finally, if Figure 2 As shown in step ②, node C i The mean prototype representation vector set of all categories will be calculated and reported, denoted as Where Y is node C i Local dataset D i The set of all categories in D i The mean prototype representation vector of all categories in is used for subsequent training of the server-side global prediction head.
[0020] Preferably, the global prediction head update module is used for training the server global prediction head, specifically as follows:
[0021] The server collects the mean prototype representation vector set uploaded by each road node Afterwards, these mean prototype representation vectors are combined into a small training sample to update the global prediction head to learn the global knowledge that runs through all nodes. The following A1-A3 is the specific process:
[0022] A1. Collect the mean prototype representation vector:
[0023] Mean prototype representation vector is the category y at road node C i The overall feature representation on the local data can comprehensively summarize the category y in C i Dynamic heterogeneous features in local datasets. For characteristics such as traffic density and lighting intensity in road nodes, the mean prototype representation vector can condense C iThe core features of the local data set show the distribution patterns of these features in space and time. Each node effectively transmits dynamic heterogeneous features to the server by reporting the mean prototype representation vector. On the one hand, this transmission method reduces the communication overhead caused by uploading the original data; on the other hand, it provides key support for the server to integrate the characteristics of different nodes and train the global prediction head.
[0024] A2. Construct small training samples:
[0025] The mean prototype representation vector reported by the server at the receiving node Finally, these vectors and their corresponding class labels y are integrated into a small set of training samples. Each mean prototype representation vector condenses the core characteristics of a specific category of data on a road node, such as the peak trend of traffic density or a specific distribution pattern under lighting conditions. This small training sample not only retains the global heterogeneous information across nodes, but also significantly reduces the communication cost. At the same time, the diverse dynamic heterogeneous features in the set provide a rich source of information for the optimization of the global prediction head, enabling it to capture the complex and changing traffic patterns between different nodes.
[0026] The construction of this small training sample reflects the collaborative expression of dynamic heterogeneous features, which is both lightweight and diverse. In the training process of the global prediction head, the small training sample set plays a bridging role, providing the server model with distributed knowledge across nodes, so as to improve the generalization ability and adaptability of the global prediction head in complex traffic scenarios.
[0027] A3. Training of global prediction head:
[0028] Finally, if Figure 2 As shown in step ③, the server uses the constructed small training sample set to optimize the global prediction head. The global prediction head aims to extract common features across road nodes to improve the model's generalization and adaptability to diverse traffic scenarios. During the training process, the server uses the mean prototype representation vector The corresponding relationship between the class label y and the global prediction head is iteratively optimized using the cross entropy loss function.
[0029] The server inputs a small training sample into the global prediction head, which predicts the category probability distribution through the global prediction head, compares it with the actual category label, calculates the error and back-propagates the updated parameters. In this process, the global prediction head gradually captures the consistency of feature distribution across nodes. For example, at nodes with dense traffic, feature distribution may tend to show high dynamic changes, while at nodes with lower traffic, it shows a stable pattern. By learning these dynamic heterogeneous features, the global prediction head can build a global model that integrates the features of various road nodes.
[0030] Preferably, the forward inheritance module is used to generate the intermediate prediction head, as follows:
[0031] like Figure 2 As shown in step ④, the intermediate prediction head The generation of is realized through the hypernetwork structure in the forward inheritance module. The main goal of the forward inheritance module is to use the historical model parameters to generate the intermediate prediction head of the current round, thereby effectively integrating historical information and enhancing the adaptability to dynamic heterogeneous traffic data.
[0032] In order to clearly explain the process of generating the intermediate prediction head, it is divided into the following three stages B1-B3:
[0033] B1. Hypernetwork input stage:
[0034] The input of the hypernetwork structure includes the local prediction head of the node in the previous round And earlier historical prediction head parameters:
[0035]
[0036] Where HN represents the hypernetwork function, and the stacked historical parameter vector It is its only input source. These historical parameters carry the knowledge of the previous rounds. The hypernetwork generates the intermediate prediction head of the current round by analyzing its feature distribution. Stacking history parameters not only improves the depth of forward memory, but also enhances the ability to capture long-term dynamic features.
[0037] In each round, the local prediction head parameters will be stored and added to the history, updated to:
[0038]
[0039] in, It is node C i The stacked set of historical parameters at the tth round. The hypernetwork takes the stacked set as input and dynamically adjusts the weights to generate new intermediate prediction head parameters.
[0040] B2, intermediate prediction head generation stage:
[0041] Getting the historical parameter stack After that, the hypernetwork generates an intermediate prediction head through the following steps
[0042] The hypernetwork first stacks the historical parameters Encode and extract high-order features through a multi-layer perceptron (MLP):
[0043]
[0044] Among them, Z i It is the extracted high-order feature representation, which contains the global and local features of the historical parameters.
[0045] Then the hypernetwork is based on the high-order features Z i Dynamically assign importance weights to historical parameters:
[0046] ω j =Softamx(WZ i +b),j∈{0,1,…,t-1}
[0047] Among them, ω j is the weight assigned to the historical parameters of the jth round, W and b are the learnable parameters of the hypernetwork. The Softmax function ensures that the weights are normalized so that ∑ω j =1.
[0048] Finally, the hypernetwork uses dynamic weights to perform a weighted summation of historical parameters to generate an intermediate prediction head:
[0049]
[0050] This stage ensures that the generated intermediate prediction head can integrate historical information and adapt to current data characteristics based on dynamic weight allocation.
[0051] B3, forward memory and dynamic adjustment stage:
[0052] In order to enhance the model's adaptability to dynamic features, the forward inheritance module further introduces a smoothing adjustment mechanism to balance the influence of historical and current parameters when generating intermediate prediction heads.
[0053]
[0054] Among them, α is a smoothing coefficient, which is used to control the influence of historical parameters on the current round. Usually, α can be determined by linear attenuation or a dynamic adjustment strategy based on data characteristics.
[0055] The dynamic adjustment characteristics of the forward memory mechanism enable the model to flexibly adapt to data heterogeneity in time and space. For example, in the case of a surge in peak traffic or a stable mode in off-peak periods, the importance distribution of historical parameters will adjust with data changes, thereby improving the personalized adaptability of the prediction head to different traffic nodes.
[0056] Preferably, the contrast learning module is used to update the local prediction head, as follows:
[0057] like Figure 2As shown in step ⑤, the update of the local prediction head of the traffic node is achieved through the contrastive learning module based on the previous steps. The core goal of the contrastive learning module is to transfer global knowledge to the local prediction head while retaining the characteristics of local data, thereby improving the prediction performance and generalization ability of the model. The following is a detailed description of the process of updating the local prediction head in two steps (C1-C2):
[0058] C1. Positive and negative sample selection
[0059] In contrastive learning, the selection of positive and negative samples directly determines the effect of knowledge transfer. In order to effectively realize the transfer and integration of global prediction head knowledge, the system of the present invention focuses on the following three key points in the design of positive and negative samples: globality, historicality and difference. The specific selection method is as follows:
[0060] The positive sample is selected as the global prediction head parameter θ of this round t ,This design is based on the following two considerations. First, the global prediction head θ t It is trained by the mean prototype representation vector uploaded by all road nodes, which contains the dynamic heterogeneous features and shared knowledge of different nodes, helping the local prediction head to calibrate and absorb global knowledge. Secondly, by using the global prediction head as a positive sample, the local prediction head can integrate global shared information while training in a personalized way, thereby improving the generalization ability of the model in data heterogeneity scenarios.
[0061] Negative samples select the local prediction head parameters of the road node itself in the previous round This design emphasizes history and difference. Local prediction head As the final result of the last round of the local model, the characteristics and historical information of the local data are completely preserved. By introducing negative samples, the model can be prevented from over-relying on historical characteristics during optimization and avoid falling into the local optimal solution. At the same time, the difference between negative samples and positive samples can provide a clear optimization direction for the contrast loss, that is, to increase the distance between the current prediction head and the historical prediction head, prevent the model from overfitting to historical data, and enhance the ability to absorb global knowledge.
[0062] C2, contrastive learning loss and local prediction head update
[0063] The basic framework of contrastive learning is as follows Figure 3 As shown in , the key is to define a reasonable loss function to quantify the similarities and differences between model parameters. + ,I a ) and maximize the objective function δ(I - ,I a ) idea, contrast loss function l con The specific definitions are as follows:
[0064]
[0065] Where: sim(·,·) represents the similarity function, usually cosine similarity, and τ is the temperature coefficient, which is used to adjust the sensitivity of the similarity.
[0066] The optimization goal of the loss function is to maximize the similarity between the local prediction head and the global prediction head, while minimizing the similarity with the historical prediction head, thereby achieving effective knowledge transfer.
[0067] During the optimization of the local prediction head, in addition to the contrastive learning loss e con , it is also necessary to consider the supervised learning loss e of local data sup , the total loss function is:
[0068] e=μ*e con +(1-μ)*l sup
[0069] Among them, μ is the balance coefficient, which is used to adjust the weights of contrastive learning and supervised learning in the total loss.
[0070] After calculating the total loss, the local prediction head parameters are updated through the back-propagation algorithm:
[0071]
[0072] Among them, η is the learning rate, which controls the parameter update step size. By balancing the contrast learning loss and the supervised learning loss, the road node can learn global knowledge and local characteristics at the same time, thereby optimizing the personalized performance and generalization ability of the model.
[0073] The present invention also discloses a method for urban traffic detection based on bidirectional memory federated learning, which is used to execute the above system, and comprises the following steps:
[0074] S1. Collecting raw road data in real time, converting the raw data into prototype representation vectors, and then uploading the prototype representation vectors to the server;
[0075] S2, the server collects the uploaded mean prototype representation vector, uses the global prediction head for training, optimizes and updates, and generates a new global prediction head;
[0076] S3. The super network stacks all local prediction heads generated by the road nodes in the previous training rounds to generate the intermediate prediction head of the current round. In this step, each terminal uses the super network and the forward inheritance mechanism to stack the historical local prediction heads to generate the intermediate prediction head.
[0077] S4. Using contrastive learning, calculate the contrast loss between the global prediction head and the intermediate prediction head generated by the forward inheritance module, and construct a joint loss in combination with the supervised learning loss. If the joint loss is greater than or equal to the threshold, generate a local prediction head, and stack the local prediction head to the super network, and return to step S1; if the joint loss is less than the threshold, the training is completed, and a personalized anomaly detection model for local dynamic heterogeneous features is obtained. The local real-time road data is input into the model to obtain the local real-time road detection status. In this step, the terminal device transfers knowledge between the intermediate prediction head and the global prediction head through contrastive learning, optimizes and generates a local prediction head, and stacks the local prediction head to the super network to realize a two-way memory mechanism.
[0078] The present invention also discloses a storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the system or method.
[0079] The present invention also discloses an electronic device, comprising:
[0080] processor;
[0081] The memory is used to store a program. When the program is called and executed by the processor, the processor executes the above-mentioned system or method.
[0082] The urban traffic detection system and method based on bidirectional memory federated learning provided by the present invention have the following significant advantages compared with the existing centralized system:
[0083] 1. Data privacy protection: The present invention adopts a distributed training architecture to ensure that the traffic data of each intersection is always processed locally, avoiding the privacy leakage problems that may occur during data transmission and storage from the source, and providing strong protection for user data security.
[0084] 2. Improve detection accuracy: Through the introduction of bidirectional memory technology, the present invention realizes the organic combination of local model and global model in optimization training, which not only inherits the generalization ability of the global model, but also makes personalized adjustments for the dynamic heterogeneous characteristics of each intersection, thereby significantly improving the detection accuracy of road status and abnormal events. Figure 4 As shown, the present invention achieves an accuracy of 0.62 in only 21 rounds of training, which is achievable by the traditional federated average algorithm after 100 rounds of training. The convergence speed is increased by 4.76 times, which significantly enhances the learning efficiency and detection capability of the system.
[0085] 3. Improve real-time performance: This invention achieves efficient processing and rapid response to massive real-time traffic data on the basis of protecting data privacy by separating local training from global model updates. This architecture ensures that the system can quickly adapt to dynamic traffic changes and output high-quality status monitoring and anomaly detection results in a timely manner, significantly enhancing the real-time performance and adaptability of the system.
[0086] 4. System adaptability and scalability: The present invention can flexibly select the appropriate extractor model according to the computing power configuration of different intersections, and perform adaptive optimization and adjustment for the dynamic heterogeneous characteristics of each intersection (such as traffic density, lighting intensity, etc.). This personalized adaptability to dynamic heterogeneous characteristics enables the system to operate efficiently in a complex and changeable urban traffic environment, not only meeting diverse needs, but also having good deployment and expansion capabilities, further improving the robustness and adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] Figure 1 This is a processing flow chart of a road state and anomaly detection system based on bidirectional memory federated learning in a preferred embodiment of the present invention;
[0088] Figure 2 This is a schematic diagram of the overall process of a bidirectional memory federated learning method according to a preferred embodiment of the present invention;
[0089] Figure 3 Schematic diagram of updating the local prediction head loss calculation for the contrastive learning module;
[0090] Figure 4 This is an experimental comparison chart between the existing federated averaging method and the method proposed in the present invention. DETAILED DESCRIPTION
[0091] The preferred embodiments of the present invention are described in detail below.
[0092] The present invention solves the problem of homogeneity of node models in existing systems through a bidirectional memory federated learning framework, fully utilizes node computing resources, and realizes efficient integration of local knowledge and global knowledge in a dynamic heterogeneous feature environment.
[0093] In order to better understand the system of the present invention, the bidirectional memory federated learning method mentioned in the present invention is first introduced in detail, and its specific training process is as follows:
[0094]
[0095]
[0096] in, Representative node C i The intermediate prediction head generated in the tth round, HN represents the hypernetwork structure, lsup represents the loss of the supervision task, e con represents the loss of the contrastive learning task and η represents the learning rate.
[0097] The present invention is described in more detail below in conjunction with the accompanying drawings. The embodiments given are only used to explain the present invention but not to limit the scope of the present invention.
[0098] First, the workflow of the system of the present invention is briefly summarized. The system of the present invention solves the problem of homogeneity of road node models by constructing a framework based on a two-way memory federated learning method. The system of the present invention is composed of multiple modules, including a data acquisition module, a forward inheritance module, a contrastive learning module, and a global prediction head update module. In each round of training, the data acquisition module is responsible for extracting the local road data of each node and generating a prototype representation vector, while calculating the mean prototype representation vector of each category and reporting it to the server. The server uses the uploaded mean prototype representation vector to train and update the global prediction head through the global prediction head update module. At the same time, the node-side forward inheritance module generates an intermediate prediction head through a super network structure, and the prediction head enters the contrastive learning module for optimization and generates a local prediction head. Finally, the local prediction head is stacked into the super network structure to complete a round of training. Among them, the contrastive learning module promotes the knowledge transfer between the global prediction head and the intermediate prediction head through the contrastive learning method, thereby optimizing the local prediction head, and finally generating a personalized anomaly detection model for local dynamic heterogeneous features, and inputs the local real-time road data into the model to obtain the local real-time road detection status.
[0099] According to a preferred embodiment of the present invention, Figure 1 As shown in Figure 1, the steps of updating the local prediction head in each round of bidirectional memory federated learning in the system go through the following modules: data acquisition module, global prediction head update module, forward inheritance module, and contrastive learning module. Figure 1 The workflow is described in detail from four steps to implement each round of the two-way memory federated learning method, and an urban traffic detection system based on two-way memory federated learning in this embodiment is described.
[0100] 1. Extract prototype representation vector (data acquisition module):
[0101] In the detection system, N road nodes are represented by a set C = {C1, C2, C3, …, C N} represents that the local data set of each road node is a set D = {D1, D2, D3, … ,D N}, where N represents a natural number greater than zero, C i represents the i-th road node, D i Indicates the corresponding local dataset.
[0102] For any road node C i For example, Figure 2 As shown in step ①, its local dataset D i Extracted into prototype representation vector R by local extractor i Specifically, node C i A sample image of (category y) will be extracted as the corresponding prototype representation vector Where y is the image label category, m is the index of the image sample in category y, 0≤m<|M|, |M| indicates that category y is in the road node C i These prototype representation vectors condense the feature information of each type of sample, which is extracted from the local dataset by the local extractor, and finally forms a prototype representation vector that can represent the category characteristics.
[0103] Next, for category y, calculate its mean prototype representation vector The mean prototype representation vector can be regarded as the node C i The overall feature representation of category y:
[0104]
[0105] Finally, if Figure 2 As shown in step ②, node C i The mean prototype representation vector set of all categories will be calculated and reported, denoted as Where Y is the node C i Local dataset D i The set of all categories in D i The mean prototype representation vector of all categories in is used for subsequent training of the server-side global prediction head.
[0106] 2. Training of server global prediction head (global prediction head update module):
[0107] The server collects the mean prototype representation vector set uploaded by each road node Afterwards, these mean prototype representation vectors are combined into a small training sample to update the global prediction head to learn the global knowledge that runs through all nodes. The following A1-A3 is the specific process:
[0108] A1. Collect the mean prototype representation vector:
[0109] Mean prototype representation vector is the category y at road node C i The overall feature representation on the local data can comprehensively summarize the category y in C iDynamic heterogeneous features in local datasets. For characteristics such as traffic density and lighting intensity in road nodes, the mean prototype representation vector can condense C i The core features of the local data set show the distribution patterns of these features in space and time. Each node effectively transmits dynamic heterogeneous features to the server by reporting the mean prototype representation vector. On the one hand, this transmission method reduces the communication overhead caused by uploading the original data; on the other hand, it provides key support for the server to integrate the characteristics of different nodes and train the global prediction head.
[0110] A2. Construct small training samples:
[0111] The mean prototype representation vector reported by the server at the receiving node Finally, these vectors and their corresponding class labels y are integrated into a small set of training samples. Each mean prototype representation vector condenses the core characteristics of a specific category of data on a road node, such as the peak trend of traffic density or a specific distribution pattern under lighting conditions. This small training sample not only retains the global heterogeneous information across nodes, but also significantly reduces the communication cost. At the same time, the diverse dynamic heterogeneous features in the set provide a rich source of information for the optimization of the global prediction head, enabling it to capture the complex and changing traffic patterns between different nodes.
[0112] The construction of this small training sample reflects the collaborative expression of dynamic heterogeneous features, which is both lightweight and diverse. In the training process of the global prediction head, the small training sample set plays a bridging role, providing the server model with distributed knowledge across nodes, so as to improve the generalization ability and adaptability of the global prediction head in complex traffic scenarios.
[0113] A3. Training of global prediction head:
[0114] Finally, if Figure 2 As shown in step ③, the server uses the constructed small training sample set to optimize the global prediction head. The global prediction head aims to extract common features across road nodes to improve the model's generalization and adaptability to diverse traffic scenarios. During the training process, the server uses the mean prototype representation vector The corresponding relationship between the class label y and the global prediction head is iteratively optimized using the cross entropy loss function.
[0115] The server inputs a small training sample into the global prediction head, which predicts the category probability distribution through the global prediction head, compares it with the actual category label, calculates the error and back-propagates the updated parameters. In this process, the global prediction head gradually captures the consistency of feature distribution across nodes. For example, at nodes with dense traffic, feature distribution may tend to show high dynamic changes, while at nodes with lower traffic, it shows a stable pattern. By learning these dynamic heterogeneous features, the global prediction head can build a global model that integrates the features of various road nodes.
[0116] 3. Generation of intermediate prediction head (forward inheritance module):
[0117] like Figure 2 As shown in step ④, the intermediate prediction head The generation of is realized through the hypernetwork structure in the forward inheritance module. The main goal of the forward inheritance module is to use the historical model parameters to generate the intermediate prediction head of the current round, thereby effectively integrating historical information and enhancing the adaptability to dynamic heterogeneous traffic data.
[0118] In order to clearly explain the process of generating the intermediate prediction head, it is divided into the following three stages B1-B3:
[0119] B1. Hypernetwork input stage:
[0120] The input of the hypernetwork structure includes the local prediction head of the node in the previous round And earlier historical prediction head parameters:
[0121]
[0122] Where HN represents the hypernetwork function, and the stacked historical parameter vector It is its only input source. These historical parameters carry the knowledge of the previous rounds. The hypernetwork generates the intermediate prediction head of the current round by analyzing its feature distribution. Stacking history parameters not only improves the depth of forward memory, but also enhances the ability to capture long-term dynamic features.
[0123] In each round, the local prediction head parameters will be stored and added to the history, updated to:
[0124]
[0125] in, It is node C i The stacked set of historical parameters at the tth round. The hypernetwork takes the stacked set as input and dynamically adjusts the weights to generate new intermediate prediction head parameters.
[0126] B2, intermediate prediction head generation stage:
[0127] Getting the historical parameter stack After that, the hypernetwork generates an intermediate prediction head through the following steps
[0128] The hypernetwork first stacks the historical parameters Encode and extract high-order features through a multi-layer perceptron (MLP):
[0129]
[0130] Among them, Z i It is the extracted high-order feature representation, which contains the global and local features of the historical parameters.
[0131] Then the hypernetwork is based on the high-order features Z i Dynamically assign importance weights to historical parameters:
[0132] ω j =Softmax(WZ i +b),j∈{0,1,…,t-1}
[0133] Among them, ω j is the weight assigned to the historical parameters of the jth round, W and b are the learnable parameters of the hypernetwork. The Softmax function ensures that the weights are normalized so that ∑ω j =1.
[0134] Finally, the hypernetwork uses dynamic weights to perform a weighted summation of historical parameters to generate an intermediate prediction head:
[0135]
[0136] This stage ensures that the generated intermediate prediction head can integrate historical information and adapt to current data characteristics based on dynamic weight allocation.
[0137] B3, forward memory and dynamic adjustment stage:
[0138] In order to enhance the model's adaptability to dynamic features, the forward inheritance module further introduces a smoothing adjustment mechanism to balance the influence of historical and current parameters when generating intermediate prediction heads.
[0139]
[0140] Among them, α is a smoothing coefficient, which is used to control the influence of historical parameters on the current round. Usually, α can be determined by linear attenuation or a dynamic adjustment strategy based on data characteristics.
[0141] The dynamic adjustment characteristics of the forward memory mechanism enable the model to flexibly adapt to data heterogeneity in time and space. For example, in the case of a surge in peak traffic or a stable mode in off-peak periods, the importance distribution of historical parameters will adjust with data changes, thereby improving the personalized adaptability of the prediction head to different traffic nodes.
[0142] 4. Update of local prediction head (contrastive learning module):
[0143] like Figure 2 As shown in step ⑤, the update of the local prediction head of the traffic node is achieved through the contrastive learning module based on the previous steps. The core goal of the contrastive learning module is to transfer global knowledge to the local prediction head while retaining the characteristics of local data, thereby improving the prediction performance and generalization ability of the model. The following is a detailed description of the process of updating the local prediction head in two steps (C1-C2):
[0144] C1. Positive and negative sample selection
[0145] In contrastive learning, the selection of positive and negative samples directly determines the effect of knowledge transfer. In order to effectively realize the transfer and integration of global prediction head knowledge, the system of the present invention focuses on the following three key points in the design of positive and negative samples: globality, historicality and difference. The specific selection method is as follows:
[0146] The positive sample is selected as the global prediction head parameter θ of this round t ,This design is based on the following two considerations. First, the global prediction head θ t It is trained by the mean prototype representation vector uploaded by all road nodes, which contains the dynamic heterogeneous features and shared knowledge of different nodes, helping the local prediction head to calibrate and absorb global knowledge. Secondly, by using the global prediction head as a positive sample, the local prediction head can integrate global shared information while training in a personalized way, thereby improving the generalization ability of the model in data heterogeneity scenarios.
[0147] Negative samples select the local prediction head parameters of the road node itself in the previous round This design emphasizes history and difference. Local prediction head As the final result of the last round of the local model, the characteristics and historical information of the local data are completely preserved. By introducing negative samples, the model can be prevented from over-relying on historical characteristics during optimization and avoid falling into the local optimal solution. At the same time, the difference between negative samples and positive samples can provide a clear optimization direction for the contrast loss, that is, to increase the distance between the current prediction head and the historical prediction head, prevent the model from overfitting to historical data, and enhance the ability to absorb global knowledge.
[0148] C2, contrastive learning loss and local prediction head update
[0149] The basic framework of contrastive learning is as follows Figure 3 As shown in , the key is to define a reasonable loss function to quantify the similarities and differences between model parameters. + ,I a ) and maximize the objective function δ(I - ,I a ) idea, contrast loss function l con The specific definitions are as follows:
[0150]
[0151] Where: sim(·,·) represents the similarity function, usually cosine similarity, and τ is the temperature coefficient, which is used to adjust the sensitivity of the similarity.
[0152] The optimization goal of the loss function is to maximize the similarity between the local prediction head and the global prediction head, while minimizing the similarity with the historical prediction head, thereby achieving effective knowledge transfer.
[0153] During the optimization of the local prediction head, in addition to the contrastive learning loss l con , it is also necessary to consider the supervised learning loss l of local data sup , the total loss function is:
[0154] l=μ*l con +(1-μ)*l sup
[0155] Among them, μ is the balance coefficient, which is used to adjust the weights of contrastive learning and supervised learning in the total loss.
[0156] After calculating the total loss, the local prediction head parameters are updated through the back-propagation algorithm:
[0157]
[0158] Among them, η is the learning rate, which controls the parameter update step size. By balancing the contrast learning loss and the supervised learning loss, the road node can learn global knowledge and local characteristics at the same time, thereby optimizing the personalized performance and generalization ability of the model.
[0159] Through the calculation of the above total loss, if the total loss is less than the threshold, the training is completed, and a personalized anomaly detection model for local dynamic heterogeneous features is obtained. The local real-time road data is input into the model to obtain the local real-time road detection status.
[0160] In summary, this embodiment proposes a bidirectional memory federated learning method and applies it to the urban road status and anomaly detection system. By dividing the model structure into two modules, the extractor and the prediction head, the computing resources of each node are fully utilized and the performance overhead is reduced. The design of the extractor allows each road node to select an adaptive model according to its data characteristics, thereby solving the problem of inconsistent model scales of different nodes in personalized federated learning and simplifying the integration and convergence process of the model.
[0161] In terms of algorithm implementation, the generation of local prediction heads of nodes combines local historical data with global migration knowledge to realize a bidirectional memory mechanism. The forward inheritance module generates intermediate prediction heads through a hypernetwork (HN), and improves the adaptability to dynamic heterogeneous traffic data by stacking and dynamically weighting the parameters of historical prediction heads. In the global contrast learning process, the node uses the global prediction head of the previous round as a positive sample and the local prediction head of the previous round as a negative sample. By optimizing the contrast loss function, the accuracy and generalization ability of the local prediction head are further improved, ensuring that the model can effectively absorb global knowledge while retaining local characteristics.
[0162] The present invention integrates local and global information through a bidirectional memory mechanism, effectively improving the performance of local personalized models of each node and achieving higher applicability and accuracy. The separation design of the extractor and the prediction head and the efficient comparative learning mechanism minimize the communication overhead while giving full play to the advantages of computing resources, so that the present invention can maintain high performance even in a resource-constrained environment.
[0163] This embodiment also compares the experimental results with the prior art. Figure 4 As shown, the present invention achieves an accuracy of 0.62 in just 21 rounds of training, which is achievable by the traditional federated average algorithm after 100 rounds of training. The convergence speed is increased by 4.76 times, which significantly enhances the learning efficiency and detection capability of the system.
[0164] The preferred embodiment of the present invention also discloses a method for urban traffic detection based on bidirectional memory federated learning, which is used to execute the above system, and includes the following steps:
[0165] S1. Collecting raw road data in real time, converting the raw data into prototype representation vectors, and then uploading the prototype representation vectors to the server;
[0166] S2, the server collects the uploaded mean prototype representation vector, uses the global prediction head for training, optimizes and updates, and generates a new global prediction head;
[0167] S3. The super network stacks all local prediction heads generated by the road nodes in the previous training rounds to generate the intermediate prediction head of the current round. In this step, each terminal uses the super network and the forward inheritance mechanism to stack the historical local prediction heads to generate the intermediate prediction head.
[0168] S4. Using contrastive learning, calculate the contrast loss between the global prediction head and the intermediate prediction head generated by the forward inheritance module, and construct a joint loss in combination with the supervised learning loss. If the joint loss is greater than or equal to the threshold, generate a local prediction head, and stack the local prediction head to the super network, and return to step S1; if the joint loss is less than the threshold, the training is completed, and a personalized anomaly detection model for local dynamic heterogeneous features is obtained. The local real-time road data is input into the model to obtain the local real-time road detection status. In this step, the terminal device transfers knowledge between the intermediate prediction head and the global prediction head through contrastive learning, optimizes and generates a local prediction head, and stacks the local prediction head to the super network to realize a two-way memory mechanism.
[0169] For other contents of this embodiment, please refer to the above system embodiment.
[0170] A preferred embodiment of the present invention further discloses a storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the above system or method.
[0171] A preferred embodiment of the present invention further discloses an electronic device, comprising:
[0172] processor;
[0173] The memory is used to store a program. When the program is called and executed by the processor, the processor executes the above-mentioned system or method.
[0174] The above contents are only preferred embodiments of the present invention. After understanding the basic concept of the present invention, ordinary technicians can make various changes and improvements, and these changes and improvements should be considered as the protection scope of the present invention.
Claims
1. The urban traffic detection system based on bidirectional memory federated learning is characterized by: Includes the following modules: Data collection module: deployed on nodes of urban roads, responsible for collecting raw road data in real time, converting the raw data into prototype representation vectors, and then uploading the vectors to the server; Global prediction head update module: deployed on the server, responsible for receiving the prototype representation vector uploaded from the road node, and optimizing and updating the global prediction head based on the uploaded prototype representation vector to generate a new global prediction head; Forward inheritance module: nodes deployed on urban roads maintain a hypernetwork inside to store and transmit information from historical training processes; the hypernetwork stacks all local prediction heads generated by road nodes in previous training rounds to generate intermediate prediction heads for the current round; Contrastive learning module: It is deployed at nodes on urban roads. Through contrastive learning, it calculates the contrast loss between the global prediction head and the intermediate prediction head generated by the forward inheritance module, and constructs a joint loss in combination with the supervised learning loss. If the joint loss is less than the threshold, the training is completed and a personalized anomaly detection model for local dynamic heterogeneous features is obtained. The local real-time road data is input into the model to obtain the local real-time road detection status.
2. The urban traffic detection system based on bidirectional memory federated learning as claimed in claim 1, characterized in that: The data collection module is as follows: Assume that N road nodes are a set C = {C1, C2, C3, ..., C N } represents that the local data set of each road node is a set D = {D1, D2, D3, ..., D N }, where N represents a natural number greater than zero, C i represents the i-th road node, D i Indicates the corresponding local dataset; For any road node C i For example, its local dataset D i Extracted into prototype representation vector R by local extractor i Specifically, node C i A sample image of Will be extracted as the corresponding prototype representation vector Among them, y is the image label category, m is the index of the image sample in category y, 0≤m<|M|, |M| means category y is in road node C i The number of pictures; For category y, calculate its mean prototype representation vector The mean prototype representation vector is regarded as the node C i The overall feature representation of category y: Node C i The mean prototype representation vector set of all categories will be calculated and reported, denoted as Where Y is the node C i Local dataset D i The collection of all categories in .
3. The urban traffic detection system based on bidirectional memory federated learning as claimed in claim 2, characterized in that: In the global prediction head update module, the server collects the mean prototype representation vector set uploaded by the road nodes After that, the mean prototype representation vector is combined into a training sample to update the global prediction head. The specific process is as follows: A1. Collect the mean prototype representation vector: Mean prototype representation vector is the category y at road node C i The overall feature representation on the local data can comprehensively summarize the category y in C i Dynamic heterogeneous features in local datasets, mean prototype representation vector can condense C i Core features of local datasets; Each node effectively transmits dynamic heterogeneous features to the server by reporting the mean prototype representation vector; A2. Construct training samples: The mean prototype representation vector reported by the server at the receiving node After that, the vectors and their corresponding category labels y are integrated into a training sample set; A3. Training of global prediction head: The server uses the constructed training sample set to optimize the global prediction head. During the training process, the server uses the mean prototype representation vector The corresponding relationship between the class label y and the global prediction head is iteratively optimized using the cross entropy loss function.
4. The urban traffic detection system based on bidirectional memory federated learning as claimed in claim 3 is characterized in that the forward In the inheritance module, the process of generating the intermediate prediction head is as follows: B1. Hypernetwork input stage: The input of the hypernetwork includes the local prediction head of the node in the previous round And the historical forecast header parameters: Where HN represents the hypernetwork function, and the stacked historical parameter vector It is its only input source. These historical parameters carry the knowledge of the previous rounds. The hypernetwork generates the intermediate prediction head of the current round by analyzing its feature distribution. In each round, the local prediction head parameters will be stored and added to the history, updated to: in, It is node C i The stacked set of historical parameters at the tth round; the hypernetwork takes the stacked set as input and dynamically adjusts the weights to generate new intermediate prediction head parameters; B2, intermediate prediction head generation stage: Getting the historical parameter stack After that, the hypernetwork generates an intermediate prediction head through the following steps: Historical parameters of the hypernetwork pair stack Encode and extract high-order features through a multi-layer perceptron: Among them, Z i It is the extracted high-order feature representation, which contains the global and local features of the historical parameters; The hypernetwork is based on the high-order features Z i Dynamically assign importance weights to historical parameters: oh j =Softmax(WZ i +b), j∈{0, 1,..., t-1} Among them, ω j is the weight assigned to the historical parameters of the jth round, W and b are the learnable parameters of the hypernetwork; the Softmax function ensures that the weights are normalized so that ∑ω j =1; The hypernetwork uses dynamic weights to perform weighted summation of historical parameters to generate an intermediate prediction head: B3, forward memory and dynamic adjustment stage: The forward inheritance module introduces a smoothing adjustment mechanism to balance the influence of historical and current parameters when generating intermediate prediction heads; Among them, α is the smoothing coefficient, which is used to control the influence of historical parameters in the current round.
5. The urban traffic detection system based on bidirectional memory federated learning as claimed in claim 4, characterized in that: In the contrastive learning module, the process of updating the local prediction head is divided into two steps: C1. Positive and negative sample selection The positive sample is selected as the global prediction head parameter θ of this round t ; Negative samples select the local prediction head parameters of the road node itself in the previous round C2, contrastive learning loss and local prediction head update Contrastive loss function l con The specific definitions are as follows: Where: sim(·,·) represents the similarity function, τ is the temperature coefficient; During the optimization of the local prediction head, in addition to the contrastive learning loss l con , it is also necessary to consider the supervised learning loss l of local data sup , the total loss function is: l=μ*l con +(1-μ)*l sup Among them, μ is the balance coefficient; After calculating the total loss, the local prediction head parameters are updated through the back-propagation algorithm: Where η is the learning rate.
6. A method for urban traffic detection based on bidirectional memory federated learning, used to execute the system as claimed in any one of claims 1 to 5, characterized in that: The steps include: S1. Collecting raw road data in real time, converting the raw data into prototype representation vectors, and then uploading the prototype representation vectors to the server; S2, the server collects the uploaded mean prototype representation vector, uses the global prediction head for training, optimizes and updates, and generates a new global prediction head; S3, the super network stacks all local prediction heads generated by the road nodes in the previous training rounds to generate the intermediate prediction head of the current round; S4. Calculate the contrast loss between the global prediction head and the intermediate prediction head generated by the forward inheritance module by using contrastive learning, and construct a joint loss in combination with the supervised learning loss. If the joint loss is greater than or equal to the threshold, generate a local prediction head, and stack the local prediction head to the super network, and return to step S1. If the joint loss is less than the threshold, the training is completed, and a personalized anomaly detection model for local dynamic heterogeneous features is obtained. Input the local real-time road data into the model to obtain the local real-time road detection status.
7. A storage medium, characterized in that: The computer program product stores computer instructions, wherein the computer instructions are used to cause a computer to execute the system according to any one of claims 1 to 5 or the method according to claim 6.
8. An electronic device, characterized in that: include: processor; A memory for storing a program, wherein when the program is called and executed by a processor, the processor executes the system according to any one of claims 1 to 5 or the method according to claim 6.
Citation Information
Patent Citations
Internet of vehicles intrusion detection model construction method based on federated learning
CN118784305A
Distributed traffic flow prediction method based on personalized federal learning
CN118840871A
Wireless service traffic prediction method based on weighted federated learning
WO2021169577A1
Method and apparatus for training traffic flow prediction model, electronic device, and storage medium
WO2022116424A1
Cited By
Cross-domain data security early warning method and system based on privacy calculation
CN120180451A
A Cross-Domain Data Security Early Warning Method and System Based on Privacy Computing
CN120180451B