Traffic sign image recognition method based on personalized federal learning

By adopting a personalized federated learning method in traffic identification image recognition, using hypernetwork structure to aggregate client models and update layer granularity coefficients, the problems caused by data privacy and non-independent homogeneous data are solved, and the accuracy and security of the model are improved.

CN120125955APending Publication Date: 2025-06-10BEIJING GUOTENG INNOVATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311675820.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In traffic sign image recognition, traditional centralized machine learning algorithms have data privacy and security issues, and non-independent and homogeneously distributed client data has negative impacts on model aggregation and update.

Method used

Using a method based on personalized federated learning, a weight matrix is ​​generated to aggregate client models by introducing a hypernetwork structure on the server, and the layer granularity coefficients are updated on the client to optimize model updates.

Benefits of technology

It effectively protects the data privacy and security of vehicle clients, reduces the negative impact of non-independent and homogeneous data on model aggregation and update, and improves the accuracy of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention provides a traffic sign image recognition method based on personalized federal learning, is applied to the field of intelligent Internet of Things, and provides an improvement scheme for the influence of common non-independent identically distributed data on each vehicle client while realizing basic traffic sign image recognition. A super network mechanism is introduced into a server side, a unique super network is formed for each vehicle client side in the process, and hierarchical learning of a model and a similar data model is facilitated from all the vehicle client sides. And meanwhile, when the server model is loaded on each vehicle client, judging a layer granularity coefficient when the client aggregates the global model according to a difference value between the local model and the global model, and updating a corresponding aggregation weight according to the layer granularity coefficient. The system can be used in the fields of intelligent Internet of Things traffic sign recognition, automatic navigation traffic sign recognition systems and the like.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This patent provides a traffic sign image recognition method based on personalized federated learning. Specifically, it involves protecting the privacy and security of vehicle client data, ensuring that data does not leave the local for training, reducing the negative impact of non-independent and identically distributed data on model aggregation and update, and introducing a hypernetwork structure on the server to find a client model that better conforms to the data sample distribution for aggregation operations when aggregating models on the server, thereby improving the model accuracy. Background Art

[0002] With the development of mobile devices and Internet of Things technologies, traffic sign image recognition has become an important application scenario. However, traditional centralized machine learning algorithms require all data to be centralized for training, but this approach involves data privacy and security issues. To address these issues, federated learning has been proposed and widely applied to traffic sign image recognition.

[0003] Federated learning is a decentralized machine learning framework that allows multiple participants to cooperate in training a model without sharing data. Each participant maintains its local dataset and then uses the local dataset to train the local model. In each round of federated learning, the participants send the parameter updates of the local model to the central server, and the server aggregates these parameters and returns the updated model parameters to each participant. The participants use the updated model parameters to continue training the local model and then send the updated model parameters to the central server again. This process iterates until the model reaches a convergence state.

[0004] The main reason why federated learning is applicable to traffic sign image recognition is that it can solve data privacy and security problems. Traditional centralized machine learning algorithms require all data to be centralized for training, which involves data privacy and security issues. With federated learning, each participant can train the model on its local device without sharing data, thus ensuring data privacy and security. In addition, since federated learning does not require all data to be centralized for training, it can reduce the use of network bandwidth and improve the efficiency of the algorithm.

[0005] However, the application of federated learning in traffic sign image recognition still needs to face some challenges. First, since the local datasets of each participant may be different, some special algorithms need to be adopted to solve the problem of data deviation. For example, the weighted aggregation algorithm in federated learning can be used to solve this problem. Second, due to the differences in the number of participants and device performance, it may lead to problems such as too long training time and performance degradation. To address this problem, some optimization algorithms can be used to accelerate training and optimize the performance of the model. Summary of the Invention

[0006] This patent proposes a traffic sign image recognition method based on personalized federated learning for the above core problems, aiming to reduce the impact of non-independent and identically distributed data of each vehicle client on model training and aggregation while protecting the privacy and security of vehicle client data. The supernetwork on the server is responsible for integrating the models uploaded by each client, and performing comparison operations based on the models uploaded by each client, and then backpropagating a weight matrix for aggregating all client models. This weight matrix can adopt the model data of clients with similar data distributions from each client and give relatively larger weights to improve the accuracy of the client model. On the client side, according to the difference between the models before and after the update, the layer granularity coefficient is updated to determine the weight ratio of each layer when the client model is updated. This reduces the negative effects brought by non-independent and identically distributed data.

[0007] The overall structure of the traffic sign image recognition based on personalized federated learning described in this patent is as Figure 1 shown, and it mainly consists of two modules, namely: (1) The server side, including the model matrix θ generated according to all client models, and the supernetwork model responsible for generating the weight matrix α. The model matrix θ will be updated each time a client uploads a model, and the supernetwork model will also be updated at the same time. After the supernetwork model generates the weight matrix α according to the client model, a new client model will be generated for each client by fitting the model matrix θ and the weight matrix α; (2) The client side, including the layer granularity coefficient module and the local update module. The layer granularity coefficient module is generated based on the comparison of the effects before and after the update of the client local model data. After receiving the personalized client model from the server, it aggregates to the local model according to the layer granularity coefficient. The aggregated model is used for local update operations, and the updated model is sent to the server for cyclic operations.

[0008] The traffic sign image recognition process of the personalized federated learning described in this patent is as Figure 2 shown, and is described in detail as follows:

[0009] Step 1: The client performs normalization preprocessing operations on the local data and conducts a certain number of batches of pre-training locally. After pre-training, it is uploaded to the server.

[0010] Step 2: The server collects the pre-training model parameters uploaded by all clients and generates a model matrix θ to be updated, where the model matrix θ includes the unique model structures of all clients.

[0011] Step 3: The server inputs the client models uploaded by each client into the hypernetwork. The hypernetwork model generates a weight matrix α that matches each client's unique model parameters. Then, the server uses the model matrix θ and each weight matrix α to generate a unique personalized client model for each client and distribute it.

[0012] Step 4: After each client receives the personalized client model distributed by the server, it updates the layer granularity coefficient and aggregates the personalized client model onto the local model according to the layer granularity coefficient. After the local model aggregation is completed, a certain number of batches of SGD optimization are performed based on the local data to update the local model, and it is decided whether to continue participating in the training. If so, it returns to Step 1.

[0013] In the traffic sign image recognition method based on personalized federated learning described in this patent, the architecture mode of federated learning is used to ensure that the data of vehicle clients does not leave the local for training, thereby ensuring the security of data privacy. In addition, the hypernetwork structure is used on the server, enabling each client to focus on learning the data of models with a data distribution similar to its own from all models, thus reducing the negative impact of non-independent and identically distributed data on model updates. Moreover, on the client side, layer granularity screening is performed on the model distributed by the server. On the basis of giving priority to ensuring the efficiency of the local model, the model parameters from other clients are selectively learned, once again avoiding the negative impact of non-independent and identically distributed data on the model effect. Brief Description of the Drawings

[0014] To more clearly illustrate the content of the present patent and the technical solutions in the embodiments, the accompanying drawings used will be briefly introduced below. The accompanying drawings in the following description are only some overall architectures and embodiments of the present patent. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0015] Figure 1 It is the architecture diagram of personalized federated learning traffic sign image recognition provided by this patent;

[0016] Figure 2 It is the flow chart of personalized federated learning traffic sign image recognition provided by this patent;

[0017] Figure 3 It is the global model update diagram based on the hypernetwork provided by this patent;

[0018] Figure 4 It is the vehicle client model aggregation example diagram provided by this patent. Detailed Description of the Embodiments

[0019] To make the objectives, technical solutions, and advantages of the embodiments of this patent clearer, the following will clearly and completely describe the technical solutions in the embodiments of this patent with reference to the accompanying drawings in the embodiments of this patent. Obviously, the described embodiments are some, rather than all, of the embodiments of this patent. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this patent without creative efforts fall within the scope of protection of this patent.

[0020] In the embodiments of this patent, the data privacy and security of each vehicle client and the situation of data being non-independent and identically distributed are mainly considered. First, each client needs to preprocess the original data. For each participant, it is necessary to preprocess its local traffic sign images, including operations such as image enhancement, image cropping, and image normalization. These preprocessing operations can improve the quality and accuracy of the images, thereby improving the training effect of the model.

[0021] Next, design a traffic sign image recognition model and train it on local data. The training of local data can use traditional deep learning methods, such as Convolutional Neural Networks (CNNs), or newer methods, such as meta-learning, transfer learning, etc. In this example, the LENET-5 convolutional neural network is used. The basic structure of LeNet-5 includes 7 network layers (excluding the input layer), including 2 convolutional layers, 2 downsampling layers (pooling layers), 2 fully connected layers, and an output layer. 1. The input layer receives a handwritten digit image of size 32*32, which includes grayscale values (0 - 255). 2. The convolutional layer C1 includes 6 convolutional kernels, each with a size of 5*5, a stride of 1, and a padding of 0. Therefore, each convolutional kernel generates a feature map of size 28*28 (the number of output channels is 6). 3. The subsampling layer S2 uses max-pooling operation, with each window size of 2*2 and a stride of 2. Therefore, each pooling operation selects the maximum value from 4 adjacent feature maps, generating a feature map of size 14*14 (the number of output channels is 6). This can reduce the size of the feature map, improve computational efficiency, and maintain a certain invariance to slight position changes. 4. The convolutional layer C3 includes 16 convolutional kernels, each with a size of 5*5, a stride of 1, and a padding of 0. Therefore, each convolutional kernel generates a feature map of size 10*10 (the number of output channels is 16). 5. The subsampling layer S4 uses max-pooling operation, with each window size of 2*2 and a stride of 2. Therefore, each pooling operation selects the maximum value from 4 adjacent feature maps, generating a feature map of size 5*5 (the number of output channels is 16). 6. The fully connected layer C5 flattens each feature map of size 5*5 into a vector of length 400 and connects it through a fully connected layer with 120 neurons. 7. The fully connected layer F6 connects 120 neurons to 84 neurons. 8. The output layer consists of 10 neurons, each corresponding to a digit from 0 - 9 and outputs the final classification result. We set the local batch size to 32, the local number of iterations to 10, the global number of iterations to 50, the number of clients to 100, the client sampling ratio to 0.2, select the SGD algorithm as the optimizer, and set the learning rate to 0.005.First, perform 10 batches of local training locally. After the training is completed, save the trained LENET5 model parameters and upload them to the server.

[0022] After the server receives the LENET5 model parameters uploaded by all clients, it generates a LENET5 model parameter matrix θ. This model parameter matrix θ contains all the client models and is hierarchical by layer. Specifically, each row of this matrix is a set of parameters for a layer in the LENET5 model, and each column is the LENET5 model parameters of a certain client. At the same time, the server will use the hypernetwork model to generate a unique weight matrix α for each client participating in the training. Regarding the hypernetwork, we define the hypernetwork structure: use a hypernetwork structure with multiple branches, and each branch corresponds to a layer in the LENET5 convolutional neural network. Each branch contains several candidate convolutional kernels and fully connected weight matrices, where each candidate weight matrix is generated by a set of hyperparameters, including the convolutional kernel size, the number of convolutional kernels, the number of fully connected neurons, etc. Hypernetwork input: The tensor input to the hypernetwork is the parameters of each layer of the trained LENET5 convolutional neural network model. Hypernetwork output: The output is a tensor that contains the weight matrices corresponding to each layer of the LENET5 convolutional neural network. Training the hypernetwork: For each input, we calculate the output of the LENET5 convolutional neural network and compare it with the initial value. According to the loss calculated by the loss function, we use the backpropagation algorithm to update the weight matrices corresponding to each branch in the hypernetwork. After generating the weight matrix α, the server side will perform a fitting operation according to the weight matrix α of each client and the model parameter matrix θ on the server to generate the corresponding personalized client model and distribute it.

[0023] After the client receives the personalized model, it first updates the layer granularity coefficient according to the local model. Specifically, the update of the layer granularity coefficient includes the following steps:

[0024] Local model update: Each participant trains the model locally using local data and calculates the difference between the local model and the global model.

[0025] Calculation of the gradients of layer weights and biases: For each layer, calculate the difference between the local model and the global model and calculate the corresponding gradients of the weights and biases.

[0026] Calculation of layer granularity coefficient: Calculate the layer granularity coefficient for each layer according to the weighted average of the gradients of each layer and the global gradient.

[0027] Update of layer weights and biases: Update the weights and biases according to the layer granularity coefficient of each layer.

[0028] After the layer granularity coefficient is updated locally on the client side, the personalized model and the local model are aggregated according to the layer granularity coefficient, and the new model is loaded into the client for 10 batches of local updates. After the update, the model is transmitted to the server again until the desired effect is achieved.

[0029] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this patent and are not intended to limit them. Although the present patent has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of this patent.

Claims

1. A method for traffic sign image recognition based on personalized federated learning, characterized in that: By distributing model training across multiple devices to protect personal data privacy and reduce the need for data transmission.

2. The method according to claim 1, characterized in that, During the image recognition process, the model is trained locally on the client side, loads the global model according to the layer granularity coefficient, and aggregates the model on the server using the hypernetwork model.

3. The layer granularity coefficient according to claim 2, characterized in that, On the client side, the weight coefficient of each layer of the client data model is divided by comparing the parameter differences before and after the local model is updated, which is called the layer granularity coefficient. When introducing the global model, the global model parameters are loaded according to the layer granularity coefficient.

4. The hypernetwork model according to claim 2, characterized in that, Save the models uploaded from each client, and then generate a unique weight matrix to aggregate all client models according to the comparison differences between each client's model and all models, and update the weight matrix in each round.