Super-network personalized federal learning method for garbage classification
Through the hyper-network personalized federated learning method, the aggregation weights of the client sharing layer are dynamically generated, which solves the problem of insufficient model adaptability caused by data heterogeneity in traditional federated learning, and achieves higher garbage classification accuracy and stability.
Patent Information
- Application Number
- CN202510344071.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-23
- Publication Date
- 2025-07-04
AI Technical Summary
Traditional garbage classification methods have problems with high labor costs and low efficiency. In the face of data heterogeneity, the simple average aggregation method cannot fully utilize the advantages of personalized models, resulting in insufficient adaptability of the model to certain devices.
The hypernetwork personalized federated learning method is adopted, and the client model is decoupled to the personalization layer and the sharing layer, and the hypernetwork is used to generate dynamic aggregation weights, flexibly adjust its contribution to the global model according to the client's local data distribution and training effect, and optimize the aggregation strategy with the backpropagation algorithm.
It improves the generalization ability and accuracy of the global model, ensures the optimization of the personalized layer, can better adapt to data heterogeneity, and enhances the adaptability and stability of the model.
Smart Images

Figure CN120258093A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of garbage classification, and specifically to a hypernetwork personalized federated learning method for garbage classification. Background Art
[0002] With the acceleration of the urbanization process and the improvement of living standards, the output of urban domestic waste continues to grow. Garbage classification has become an important link in environmental protection and resource recycling. However, traditional garbage classification methods have problems such as high labor costs and low efficiency. Using machine learning methods can improve the efficiency of automatic image recognition for classification, but traditional centralized learning faces the high cost of massive data transmission and processing.
[0003] To solve the above problems, the architecture based on federated learning has become an ideal choice for processing large-scale data. However, the garbage classification scenario has high complexity and diversity, and there are significant differences in garbage types, distributions, and characteristics in different environments. The global classification model trained by traditional federated learning frameworks is difficult to adapt to diverse application scenarios, affecting the generalization ability and classification accuracy of the model.
[0004] In traditional federated learning frameworks, the problem of data heterogeneity (such as uneven data distribution, different characteristics, etc.) may lead to poor model performance. Especially in personalized applications, FedPer is a solution to the personalized problem in federated learning. Its main goal is to improve the learning effect of each participant by making personalized adjustments on the basis of the global model. The core idea of the FedPer framework is that each device (or client) can maintain a locally personalized model, and at the same time, on the basis of the globally shared model, perform regular model fusion.
[0005] However, there are certain defects in the aggregation method of FedPer. Especially when facing data heterogeneity, the aggregation method may not be intelligent enough, resulting in the inability to fully utilize the advantages of personalized models. The aggregation method of Fedper is usually simple average aggregation. However, some devices have much more data than other devices, and the simple average method does not consider the importance difference of model updates. When the model updates uploaded by devices are weighted and averaged, devices with a large amount of data may dominate the direction of the final model, resulting in insufficient adaptability of the model to some devices with a small amount of data. Even if the data volumes of two devices are the same, the quality of model updates may be very different. If the updates of some devices are poor due to noise or training problems, the model updates of these devices may interfere with the optimization of the global model. Summary of the Invention
[0006] To solve the technical problems mentioned in the current background art, the present invention proposes a hypernetwork personalized federated learning method for garbage classification.
[0007] To this end, the technical solution adopted by the present invention is as follows:
[0008] A hypernetwork personalized federated learning method for garbage classification, the architecture of the federated learning consists of a server, multiple clients and a hypernetwork. Each client stores a client model, and the client model consists of a personalized layer and a shared layer. The server stores a global shared layer, and the shared layer in the client is obtained by training and updating the global shared layer. The method includes:
[0009] S1. Initialize the parameters of the personalized layer, shared layer and hypernetwork of all clients; the server distributes the global shared layer to each client, and each client performs local training on the personalized layer and the shared layer, updates the parameters of the personalized layer, and calculates the parameters of the shared layer;
[0010] S2. Generate the aggregation weights of the shared layers of each client through the hypernetwork model, upload all the client shared layer parameters to the server, and the server performs weighted aggregation on all the client shared layer parameters according to the aggregation weights to obtain the global shared layer parameters;
[0011] S3. Based on the backpropagation algorithm, the hypernetwork adjusts the strategy for generating the aggregation weights by optimizing the global loss function; according to the adjustment of the strategy of the aggregation weights, repeat the processes of the local training and aggregation.
[0012] Furthermore, each client performs local training on the personalized layer and the shared layer based on local data, and updates the parameters of the personalized layer.
[0013] The local training objective of each client is to minimize the local loss function. The calculation formula of the local loss of client i is:
[0014]
[0015] where, f si (x) is the output of the shared layer of client i; f pi (x) is the output of the personalized layer of client i; L local is the local training loss function; y is the true label; is the parameter of the shared layer of client i, is the parameter of the personalized layer of client i; L i represents the local loss of client i.
[0016] Furthermore, the hypernetwork model is jointly composed of a multi-layer neural network and a fully connected layer. The hypernetwork model generates the aggregation weights through the following formula, and the formula is:
[0017] w i = σ(FCj(clienti training effect,client i local data))
[0018] Among them, FC j represents the jth layer of the hypernetwork model; w i is the shared layer aggregation weight of client i, client i The training effect represents the training accuracy of client i. i local data represents the distribution of local data of client i in all client datasets, and σ represents the activation function.
[0019] Furthermore, the global shared layer parameters are obtained through the weighted aggregation, and the formula is:
[0020]
[0021] Where N is the number of clients; is the global shared layer parameter.
[0022] Further, after calculating the aggregation weight, the hypernetwork model parameters are optimized by the global loss function, which is obtained by weighted average of the local losses of each client, and the formula is as follows:
[0023]
[0024] Among them, L i is the local loss of the ith client; w i is the aggregate weight of the i-th client; L global is the global model loss.
[0025] Furthermore, the strategy of the aggregation weight is adjusted by the back propagation algorithm, and the adjustment formula is:
[0026]
[0027] Where θ is the learning rate; is the global loss function with respect to the aggregate weight w i The gradient of the local training and aggregation process is repeated according to the adjustment of the aggregation weight strategy.
[0028] Compared with the prior art, the advantages of the present invention are:
[0029] 1. The present invention overcomes the limitation of the simple average aggregation method in the traditional FedPer method by introducing the hypernetwork to dynamically generate the aggregation weight of the client sharing layer;
[0030] 2. The algorithm of the present invention can flexibly adjust the contribution of each client in the global model according to the local data distribution and training effect of the client, avoiding performance loss;
[0031] 3. The aggregation strategy based on the hypernetwork enables the model to more finely adapt to data heterogeneity, improving the generalization ability of the global model, while ensuring the optimization of the personalized layer, thereby enhancing the accuracy and adaptability of the global model. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0033] Figure 1 is the local training and aggregation flowchart of the present invention;
[0034] Figure 2 is the aggregation weight generation flowchart of the present invention;
[0035] Figure 3 is the performance diagram of each algorithm for the TrashNet dataset with independent and identically distributed data;
[0036] Figure 4 is the performance of each algorithm under non-independent and identically distributed data Figure 1 ;
[0037] Figure 5 is the performance of each algorithm under non-independent and identically distributed data Figure 2 ;
[0038] Figure 6 is the distribution diagram of the self-built garbage dataset. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To achieve the above objectives, the present invention is implemented through the following technical solutions. The present invention provides a hypernetwork personalized federated learning method for garbage classification, and the method includes, as Figure 1 and 2 shown:
[0040] S1. The client model is decoupled into a personalized layer and a shared layer, and the parameters of the personalized layer, the shared layer, and the hypernetwork are initialized; the server outputs the global shared layer parameters to each client, and generates the aggregation weights of the shared layer of each client through the hypernetwork model.
[0041] In this embodiment, the client model uses the ResNet34 model, which has 34 layers, combines the traditional convolutional neural network and the residual module, and improves the training effect of the deep network.
[0042] The client model better addresses the issue of data heterogeneity by decoupling into a personalized layer and a shared layer. The personalized layer is trained locally to ensure that each client can be optimized according to its specific data. The shared layer learns global model knowledge by aggregating updates from different clients.
[0043] Each client independently trains and updates the parameters of the personalized layer based on local data, and the training process of the personalized layer does not perform cross-client aggregation. The personalized layer is designed to enable the local training layer to be adjusted according to the local data characteristics of each client, thereby improving the performance of the local training layer. The personalized layer includes the subsequent parts outside the shared layer, especially the residual blocks containing advanced feature extraction and the final classification layer. These parts are optimized on local data to ensure that the model can adapt to the specific data distribution of each client. The training objective of the personalized layer for each client is to minimize its local loss function, thereby improving the accuracy and adaptability of local tasks. The formula is:
[0044]
[0045] where, f si (x) is the output of the shared layer of client i; f pi (x) is the output of the personalized layer of client i; L local is the local training loss function; y is the true label; are the parameters of the shared layer of client i, are the parameters of the personalized layer of client i; L i represents the local loss of client i.
[0046] The goal of the shared layer is to extract the common features of all clients. The common features are low-level features related to the task. The shared layer only includes the earlier convolutional layers and the first few residual blocks in the network. These parts will learn similar features in most clients. The shared layer is updated by aggregating among all clients. The shared layer of each client is trained locally and calculates the corresponding shared layer parameters
[0047] S2. Generate the aggregation weights of the shared layer for each client through the hypernetwork model. The shared layer parameters are weighted and aggregated according to the aggregation weights to obtain the global shared layer parameters. The shared layer parameters are uploaded to the server, and the server performs weighted aggregation on the shared layer parameters according to the aggregation weights.
[0048] Utilize the generation ability of the hypernetwork model to dynamically generate the aggregation weights w i, so that the contribution of each client in the global model can be flexibly adjusted according to its local data distribution and training situation. The hypernetwork optimizes the aggregation strategy by learning the contribution of each client and its influence in the global model, so as to automatically adjust its weight in the global model according to the local data characteristics and training effect of the client. In this way, the hypernetwork can make fine-grained adjustments for data heterogeneity and client performance differences, avoiding the performance loss caused by traditional simple average aggregation methods,
[0049] The hypernetwork model is composed of multiple layers of neural networks combined with fully connected layers. During the training process, the parameters of the hypernetwork model are optimized so that it can learn an optimal weight generation strategy. The formula is as follows:
[0050] w i = σ(FC j (client i training effect, client i local data))
[0051] Among them, FC j represents the j-th layer of the hypernetwork model; w i is the aggregation weight of the shared layer of the i-th client; taking the training effect of each client and the local data distribution characteristics of each client as input, learning the sharing of each client in the global model and the relationship between clients, so as to obtain the aggregation weight of each client,
[0052] The shared layer parameters of all clients are weighted and aggregated according to the aggregation weight w i generated by the hypernetwork to obtain the global shared layer parameters The specific formula is as follows:
[0053]
[0054] Among them, N is the number of clients.
[0055] S3. Based on the backpropagation algorithm, the hypernetwork adjusts the strategy for generating the aggregation weight by optimizing the global loss function. According to the adjustment of the aggregation weight strategy, the processes of the local training and aggregation are repeated,
[0056] The global loss function is obtained by weighted averaging the local losses of each client. The formula is as follows:
[0057]
[0058] Among them, L i is the local loss of the i-th client; w i is the aggregation weight generated by the hypernetwork; L global is the global model loss;
[0059] After each round of training, the strategy for generating the aggregation weights is adjusted through the backpropagation algorithm, and the update process is as follows:
[0060]
[0061] where θ is the learning rate; is the gradient of the loss function with respect to the aggregation weight w i ; According to the adjustment of the strategy of the aggregation weight, the processes of local training and aggregation are repeated.
[0062] In summary, the advantage of the present invention is that by introducing a hypernetwork to dynamically generate the aggregation weights of the client sharing layer, the limitations of the simple average aggregation method in the traditional FedPer method are overcome. Different from the traditional simple average method, the algorithm of the present invention can flexibly adjust the contribution of each client in the global model according to the local data distribution and training effect of the client, avoiding performance loss. This aggregation strategy based on the hypernetwork enables the model to more finely adapt to data heterogeneity, improves the generalization ability of the global model, and at the same time ensures the optimization of the personalized layer, thereby enhancing the accuracy and adaptability of the global model.
[0063] The experimental verification results based on this invention are as follows:
[0064] Table 1 Best performance of each algorithm on the TrashNet dataset with independent and identically distributed data
[0065]
[0066] Figure 3 Table 1 and Table 1 show the performance of each algorithm in the case of independent and identically distributed data. In this environment, the data distributions of each client are basically the same. Therefore, due to its simple global aggregation method, the traditional FedAvg algorithm can make good use of the data advantages of each client and finally obtain strong performance. However, the FedPer based on model decoupling and the pFedHN algorithm proposed in this paper have relatively poor performance in this environment. The method based on model decoupling mainly adapts to the data distributions of different clients by generating personalized model parameters for each client. However, when the data is independent and identically distributed, the data distributions of all clients are almost the same, resulting in this personalized strategy not playing its advantages. Due to the similarity of the data distributions, the personalized updates between models are interfered by too many parameter updates and fail to make full use of the global shared information, resulting in its performance being inferior to simple aggregation methods such as FedAvg.
[0067] Table 2 Best performance of each algorithm under non-independent and identically distributed data - Dirichlet(0.5)
[0068]
[0069] Table 3 Best Performance of Each Algorithm under Non - independent and Identically Distributed - Dirichlet(0.3)
[0070]
[0071] Figure 4 、 Figure 5 And Tables 2 and 3 show the performance of each algorithm when data heterogeneity is increasing. As the data heterogeneity increases, the performance of the FedAvg algorithm begins to decline significantly. In the initial stage with relatively low data heterogeneity, FedAvg can handle global model aggregation well. However, as the difference in data distribution increases, FedAvg cannot effectively adapt to the heterogeneity among clients, resulting in large fluctuations in its training accuracy and ultimately affecting the stability and convergence speed of the model. The FedALA algorithm shows strong advantages in the case of relatively low data heterogeneity and can balance the local models and the global model of each client by adaptively adjusting parameters, thus obtaining better training results. However, as the data heterogeneity further deepens, the performance of FedALA begins to decline. In contrast, the algorithm proposed in this paper shows significant advantages in the environment of increasing data heterogeneity. As the data heterogeneity gradually deepens, this algorithm can better adapt to the data differences among clients and can more effectively cope with the data differences among clients, showing stronger adaptability and finally achieving better performance in the strong heterogeneous data environment.
[0072] According to the experimental results, FedAvg performs excellently when the data heterogeneity is weak, but its performance significantly declines as the heterogeneity increases. FedALA performs well under medium heterogeneity but gradually fails in a strong heterogeneous environment. FedPer performs poorly under data independent and identically distributed or weak data heterogeneity and only shows advantages under strong data heterogeneity. In contrast, the algorithm proposed in this paper shows strong adaptability in various data heterogeneous scenarios and can maintain good performance whether in a weak or strong data heterogeneous environment. Especially in a strong data heterogeneous environment, its performance is significantly better than other algorithms.
[0073] Table 4 Best Performance of Each Algorithm on the Self - built Dataset
[0074]
[0075] Figure 6Table 4 shows the performance of each algorithm in the real garbage classification scenario. Due to the different environments of the three clients, affected by factors such as lighting and acquisition angles, the data has strong heterogeneity. In this case, the traditional FedAvg algorithm performs poorly, with large fluctuations in training accuracy, unable to adapt to the significant differences between clients, resulting in unstable and low performance. When facing a data environment with strong heterogeneity, the FedALA algorithm fails to effectively utilize its advantage of personalized adjustment. Eventually, its performance is similar to that of the FedAvg algorithm and also has large fluctuations.
[0076] The FedPer method performs strongly in this environment with strong heterogeneity. Its personalized model can better adapt to the unique data characteristics of each client. Especially in the case of large data differences, it shows higher accuracy and stability. The method proposed in this paper performs optimally among all algorithms, capable of maintaining higher accuracy and lower fluctuations in a strong heterogeneous data environment, demonstrating stronger adaptability and robustness.
[0077] The FedAMP method also shows certain advantages in a strongly data heterogeneous environment. However, when dealing with extreme data heterogeneity, it still faces certain challenges. Its weighting strategy may be affected by data noise and extreme data differences, resulting in an unstable training process.
[0078] In summary, pFedHN demonstrates significant advantages in coping with data heterogeneity and improving model performance, and is suitable for application in real garbage data scenarios. Through experimental verification, when facing a complex and heterogeneous data environment, pFedHN can not only effectively improve the accuracy of the model, but also shows strong stability and robustness. Its personalized and adaptive characteristics enable it to better adapt to the differences between clients, ensuring excellent performance in complex scenarios such as garbage classification.
[0079] The above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.
Claims
1. A hypernetwork personalized federated learning method for garbage classification, the architecture of the federated learning consists of a server, multiple clients and a hypernetwork. Each client stores a client model, and the client model consists of a personalized layer and a shared layer. The server stores a global shared layer, and the shared layer in the client is obtained by training and updating the global shared layer. It is characterized in that, The method includes: S1. Initialize the personalized layer, shared layer, and hypernetwork parameters of all clients; the server distributes the global shared layer to each client, and each client performs local training on the personalized layer and the shared layer, updates the parameters of the personalized layer, and calculates the parameters of the shared layer; S2. Generate the aggregation weights of the shared layers of all clients through the hypernetwork model, upload the shared layer parameters of all clients to the server, and the server performs weighted aggregation on the shared layer parameters of all clients according to the aggregation weights to obtain the global shared layer parameters; S3. Based on the backpropagation algorithm, the hypernetwork adjusts the strategy for generating the aggregation weights by optimizing the global loss function; according to the adjustment of the strategy of the aggregation weights, repeat the processes of the local training and aggregation.
2. The hypernetwork personalized federated learning method for garbage classification according to claim 1, wherein, Each client performs local training on the personalized layer and the shared layer based on local data and updates the parameters of the personalized layer. The local training objective of each client is to minimize the local loss function, and the calculation formula for the local loss of client i is: Among them, f si (x) is the output of the shared layer of client i; f pi (x) is the output of the personalized layer of client i; L local is the local training loss function; y is the true label; is the parameter of the shared layer of client i, is the parameter of the personalized layer of client i; L i represents the local loss of client i.
3. A hypernetwork personalized federated learning method for waste sorting according to claim 2, characterized in that, The hypernetwork model is jointly composed of multiple layers of neural networks and fully connected layers. The hypernetwork model generates the aggregation weights through the following formula: w i = σ(FC j (client i training effect, client i (local data)) Among them, FC j represents the j-th layer of the hypernetwork model; w i is the shared layer aggregation weight of client i, client i trainingeffect represents the training accuracy of client i, client i local data represents the distribution characteristics of the local data of client i among the local data of all clients, and σ represents the activation function.
4. The hypernetwork personalized federated learning method for garbage classification according to claim 3, wherein, The global shared layer parameters are obtained through the weighted aggregation, and the formula is: where N is the number of clients; are global shared layer parameters; Distribute the global shared layer parameters to each client, and replace the shared layer parameters of all clients with the global shared layer.
5. The hypernetwork personalized federated learning method for garbage classification according to claim 4, wherein, After calculating the aggregation weights, optimize the hypernetwork model parameters through the global loss function. The global loss function is obtained by weighted averaging the local losses of each client, and the formula is as follows: Among them, L i is the local loss of the i-th client; w i is the aggregated weight of the i-th client; L global is the global model loss.
6. The hypernetwork personalized federated learning method for garbage classification according to claim 5, wherein The strategy of the aggregation weights is adjusted through the backpropagation algorithm, and the adjustment formula is: where θ is the learning rate; is the gradient of the global loss function with respect to the aggregation weight w i ; is the adjusted aggregation weight.
Citation Information
Cited By
Personalized federal learning method and system for data heterogeneous and resource constrained environment
CN120893526A
A personalized federated learning method and system for data heterogeneous and resource constrained environments
CN120893526B
Financial fraud detection model training method, system and device, medium and equipment
CN121052831A
Financial fraud detection model training method, system, device, medium and equipment
CN121052831B
Kitchen garbage classification and tracking method and device, medium and product
CN122067233A