An efficient hierarchical soft clustering federated learning method, device, and medium
By using the fuzzy C-mean clustering algorithm in the Internet of Vehicles, the problems of low privacy protection and clustering efficiency in the Internet of Vehicles are solved, and efficient generation of personalized models and data privacy protection are achieved.
Patent Information
- Application Number
- CN202510655163.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2045-05-21
AI Technical Summary
In the scenario of Internet of Vehicles, it is difficult for the existing technology to achieve efficient and flexible client clustering under the premise of protecting privacy. Traditional hard clustering methods may lose key information or misclassify, and traditional clustering algorithms cannot effectively handle the fuzzy relationship between the automobile end and the cluster.
The fuzzy C mean clustering algorithm is used to hierarchically cluster the BN layer parameters of the neural network, and the scaling_factor parameter of the BN layer is used as the clustering basis. A personalized model is generated through the fuzzy C mean clustering algorithm, and combined with the sparse local objective function, hierarchical clustering is achieved.
Without adding additional communication overhead, effective clustering on the automotive side is realized, data privacy is protected, and the personalization and generalization capabilities of the model are improved, enhancing the robustness and efficiency of clustering.
Smart Images

Figure CN120181143B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of federated learning, and more particularly to an efficient hierarchical soft clustering federated learning method, device, and medium. Background Art
[0002] In recent years, the rapid development of large models has profoundly changed daily life. Their intelligent decision-making capabilities rely on training with massive amounts of data. However, if the distributed data generated by IoT devices is centrally transmitted to a central node, there are concerns about privacy leakage and high communication costs. Federated Learning (FL) has emerged to address this issue. By training models through decentralized collaboration, FL eliminates the need to share raw data, protecting privacy and reducing transmission costs.
[0003] In the IoV scenario, the federated learning framework consists of a server and multiple vehicles. During each round of training, the server sends a global model to each vehicle. The latter updates the model using local data, uploading only the parameters to the server for aggregation to form a new global model. The commonly used aggregation algorithm, FedAvg, weights the average model parameters based on the amount of data from each client. Clients with larger data volumes have a greater impact on the global model. However, IoV data often exhibits non-independent and identically distributed (non-IID) characteristics. For example, the types of images collected by different vehicles vary significantly, resulting in low efficiency and insufficient accuracy in global model training.
[0004] To address data heterogeneity, cluster federated learning divides clients into multiple clusters and trains a personalized model for each cluster. However, practical applications present two major challenges: first, direct data clustering is impossible while maintaining privacy; second, the fuzzy boundaries between clusters mean that traditional hard clustering methods may miss key information or misclassify clients. For example, some vehicles may belong to multiple clusters simultaneously. Forcing a single cluster to belong to a specific client significantly reduces model effectiveness.
[0005] Therefore, achieving efficient and flexible clustering of IoV clients while protecting privacy and building a personalized federated learning framework is key to improving model performance and user experience. This requires exploring new clustering mechanisms while also taking into account the dynamic adjustment capabilities of clusters to address the complex and ever-changing needs of real-world scenarios.
[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this application, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention
[0007] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical elements or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.
[0008] The embodiments of the present disclosure provide an efficient hierarchical soft clustering federated learning method, device, and medium to solve the technical problems of clustering federated learning and clustering training in the Internet of Vehicles object recognition scenario.
[0009] In some embodiments, the method comprises:
[0010] The cloud server initializes a global model consisting of L model blocks, where each model block includes a convolutional layer or a fully connected layer and its corresponding BN layer;
[0011] The global model is sent to each vehicle, and each vehicle trains the model based on local data, with the local objective function being:
[0012] ,
[0013] in, Indicates the car side The local objective function on Indicates the car side The amount of data on It is The cross entropy loss of sample data is It's the car side Previous The one-norm regularization loss of the scaling_factor parameter of the BN layer of the model block, is a super parameter, Indicates the car side Personalized models on
[0014] The cloud server receives the model parameters uploaded by each car end and obtains the scaling_factor parameters of each model block of the locally trained models of all cars end;
[0015] Based on all the scaling_factor parameters of each model block, the fuzzy C-means clustering algorithm is used to perform hierarchical clustering on each model block to generate a personalized model for each car end;
[0016] The personalized model is sent to the corresponding car end for application.
[0017] Preferably, the model blocks are hierarchically clustered using the fuzzy C-means clustering algorithm according to the scaling_factor parameter of each model block to generate a personalized model for each car end, specifically in the following manner:
[0018] All scaling_factor parameters of each model block are divided into fuzzy clusters and obtain the membership matrix;
[0019] According to the division Fuzzy clusters and membership matrices are aggregated into each model block. Cluster model blocks;
[0020] According to each model block The cluster model blocks and membership matrix generate L personalized model blocks for each car end;
[0021] The car side combines L personalized model blocks into a complete personalized model.
[0022] Preferably, the fuzzy C-means clustering algorithm uses cosine distance to measure the similarity between the scaling_factor parameter and the cluster center.
[0023] Preferably, the objective function of the fuzzy C-means clustering algorithm is:
[0024] ,
[0025] in, represents the clustering optimization objective, is the fuzzy factor; It is The membership matrix of the model blocks; Indicates that for Model Nugget Client Cluster The membership degree satisfies and ; Indicates the number of car terminals; It is The centers of the clusters of the model blocks, Indicates the The first cluster centers; For the car Previous The scaling_factor parameter of the BN layer of the model block and the cluster center The cosine distance.
[0026] Preferably, the objective function of the fuzzy C-means clustering algorithm is minimized in the following manner:
[0027] Normalize the scaling_factor parameter, that is, convert it into a unit vector , For the car Previous A unit vector of scaling_factor parameters for each model block;
[0028] Random initialization cluster centers and calculate each unit vector cosine distance from each cluster center;
[0029] Update the membership and update the membership matrix based on the membership:
[0030] ,
[0031] Update the cluster centers and recalculate each The cosine distance from each cluster center until the cluster centers converge.
[0032] Preferably, the parameters of the cluster model block are obtained by aggregating the convolutional layer or fully connected layer parameters of all model blocks in the corresponding cluster according to the membership matrix, specifically in the following manner:
[0033] ,
[0034] ,
[0035] or
[0036] ,
[0037] in, For the car Previous The scaling_factor parameter of the BN layer of the model block, Indicates the car side Previous The convolutional layer parameters in the model block, Indicates the car side Previous The fully connected layer parameters in each model block; Representation model nugget The corresponding The scaling_factor parameter of the BN layer in the cluster model block; Representation model nugget The corresponding The convolutional layer parameters in each cluster model block; Representation model nugget The corresponding The connection layer parameters in the cluster model nugget.
[0038] Preferably, the parameters of the personalized model block are based on the The cluster model blocks and membership matrix are generated as follows:
[0039] ,
[0040] ,
[0041] or
[0042] ,
[0043] in, Indicates the car side Previous The scaling_factor parameter of the BN layer in the personalized model block, Indicates the car side Previous The convolutional layer parameters in the personalized model blocks, Indicates the car side Previous The fully connected layer parameters in each personalized model block.
[0044] Preferably, the fuzzy factor .
[0045] In some embodiments, the apparatus includes a processor and a memory storing program instructions, and the processor is configured to execute the efficient hierarchical soft clustering federated learning method when running the program instructions.
[0046] In some embodiments, the storage medium stores a computer program thereon, which, when executed by a processor, implements the efficient hierarchical soft clustering federated learning method as described above.
[0047] The embodiments of the present disclosure provide an efficient hierarchical soft clustering federated learning method, device, and medium, which can achieve the following technical effects:
[0048] The present invention does not increase additional communication overhead. Under the premise of protecting data privacy, it only uses model parameters to achieve effective vehicle-side clustering with almost no additional communication overhead.
[0049] Using the parameters of the BN layer as the basis for clustering achieves efficient clustering while protecting data privacy. Furthermore, applying a one-norm sparsification to the BN layer during local vehicle training enhances the diversity of different vehicle models, further facilitating clustering.
[0050] Since the parameters of the sparse BN layer are used, the cosine distance, which is more robust to high-dimensional sparse data, is used during clustering to achieve better clustering results.
[0051] Fuzzy clustering is used to effectively address the situation where fuzzy cluster relationships exist between vehicles and clusters. The fuzzy C-means algorithm is used to determine the membership relationships between vehicles and clusters. A fuzzy factor m ≥ 1 is set to control the degree of fuzziness in the membership. When m = 1, fuzzy C-means degenerates into a hard k-means clustering algorithm. When m approaches positive infinity, the membership tends to the average. A cluster model is derived based on the membership aggregation parameters, and a personalized model for each vehicle is derived based on the cluster model and the membership.
[0052] Given that different layers in a neural network often have distinct structures and functions, it's inappropriate to use a unified clustering relationship for the entire model. This paper, combined with commonly used neural network structures, proposes a hierarchical clustering federated learning algorithm. This algorithm uses model blocks as units for clustering, cluster model generation, and personalized model generation, enhancing the generalization capabilities of personalized models. Notably, because the clustering between layers is relatively independent, it can be effectively parallelized without extending training time.
[0053] The above general description and the following description are exemplary and explanatory only and are not intended to limit the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] One or more embodiments are exemplarily described by corresponding drawings. These exemplary descriptions and drawings do not limit the embodiments. Elements with the same reference numerals in the drawings are shown as similar elements. The drawings do not constitute a scale limitation. In addition,
[0055] Figure 1 It is a schematic flow chart of the method of the present invention;
[0056] Figure 2 This is a schematic diagram of the cloud server process provided by an embodiment of the present disclosure;
[0057] Figure 3 This is a schematic diagram of the vehicle-side process provided by an embodiment of the present disclosure;
[0058] Figure 4 Schematic diagram of the similarity matrix of the car side generated when the last BatchNorm layer is used as the clustering basis in the embodiment of the present disclosure;
[0059] Figure 5 Schematic diagram of the similarity matrix of the car side generated when the penultimate BatchNorm layer is used as the clustering basis in the embodiment of the present disclosure;
[0060] Figure 6 Schematic diagram of the similarity matrix of the car side generated when the first BatchNorm layer is used as the clustering basis in the embodiment of the present disclosure;
[0061] Figure 7 It is a schematic diagram of the device structure provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0062] In order to be able to understand the features and technical content of the embodiments of the present disclosure in more detail, the implementation of the embodiments of the present disclosure is described in detail below in conjunction with the accompanying drawings. The accompanying drawings are for reference only and are not used to limit the embodiments of the present disclosure. In the following technical description, for the sake of convenience of explanation, a full understanding of the disclosed embodiments is provided through multiple details. However, one or more embodiments can still be implemented without these details. In other cases, to simplify the drawings, well-known structures and devices can be simplified for display.
[0063] In the description and claims of the embodiments of the present disclosure, as well as in the accompanying drawings, the terms "first," "second," and the like are used to distinguish similar items and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate to describe the embodiments of the present disclosure herein. In addition, the terms "including," "having," and any variations thereof are intended to cover non-exclusive inclusions.
[0064] Unless otherwise stated, the term "plurality" means two or more.
[0065] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B means: A or B.
[0066] The term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0067] The term "correspondence" may refer to an association relationship or a binding relationship. The correspondence between A and B means that there is an association relationship or a binding relationship between A and B.
[0068] Example 1
[0069] Currently, in order to achieve clustering that protects data privacy, the intuitive idea is to directly use the local model uploaded by the car end during the federated learning process. Existing methods use the local model uploaded by each car end, and calculate the distance between local models to achieve clustering. However, this method has two problems. One is the problem of computational efficiency. When the number of model parameters is large, the overhead of calculating the distance between models is relatively large; the other is the problem of clustering effect. The clustering result obtained by calculating the distance between the parameters of the entire model is not necessarily good. In addition, in the clustering algorithm, the commonly used hard clustering algorithm that forces each car end to be assigned to a single cluster cannot effectively handle the situation where there is an ambiguous cluster relationship between the car end and the cluster, that is, a car end is at the boundary of multiple clusters and may belong to multiple clusters at the same time. In order to solve the above problems, the present invention proposes an efficient hierarchical soft clustering federated learning method, ignoring the pooling layer and the nonlinear layer because they usually have no trainable parameters. Specifically, Figure 1 Shown, including:
[0070] S1: The cloud server initializes a global model consisting of L model blocks, where each model block includes a convolutional layer or a fully connected layer and its corresponding BN layer.
[0071] As a refinement of the above embodiment, a neural network is composed of several layers with different functions, such as convolutional layers, pooling layers, fully connected layers, or normalization layers. Considering that the BatchNorm (BN) layer, which is widely used in neural networks, has a small number of parameters and is representative, the present invention selects the parameters of the BN layer as a representation of local data and adds a norm sparsification to the BN layer parameters in the local objective function, so that the BN parameters can better reflect the differences in local data distribution.
[0072] The BN layer is usually located after the fully connected layer and the convolutional layer. Compared with the fully connected layer and the convolutional layer, the batch normalization layer has fewer training parameters. The reduction of parameters will improve the efficiency of distance calculation. Specifically, taking the BN layer after the convolutional layer as an example, for a (bias is not calculated here) convolutional layer with trainable parameters, and the subsequent BN is only (Bias is not calculated here) trainable parameters (called scaling_factor), the reduction of parameters greatly reduces the complexity of distance calculation (where is the number of convolution kernels, is the number of channels of the input feature map, is the size of the convolution kernel). Furthermore, the scaling_factor parameter of the BN layer is closely tied to the training data and can therefore serve as a good representation of local data. Since the BN layer itself is part of the model parameters and needs to be uploaded to the server for aggregation during training, using the BN layer parameters as the basis for clustering theoretically does not incur additional communication overhead.
[0073] Modern deep neural networks are typically composed of several network layers, which often have different structures and functions and therefore need to be treated differently. That is, each layer of the network needs to be clustered independently in a hierarchical manner. The BN layer is usually located after the fully connected layer and the convolutional layer and is closely related to them. Compared with the fully connected layer and the convolutional layer, it has fewer training parameters. Therefore, the present invention treats the fully connected layer or convolutional layer and the subsequent BN layer as a whole and clusters them according to the scaling_factor parameter of the BN layer.
[0074] S2: Send the global model to each vehicle end, and each vehicle end trains the model based on local data, and its local objective function is:
[0075] ,
[0076] in, Indicates the car side The local objective function on Indicates the car side The amount of data on It is The cross entropy loss of sample data is It's the car side Previous The one-norm regularization loss of the scaling_factor parameter of the BN layer of the model block, Is a hyperparameter used to weigh the cross entropy loss and the one-norm regularization loss. Indicates the car side Personalized model on.
[0077] The above modification of the local objective function is to make the scaling_factor parameter of each model block have a certain sparsity, that is, to make the differences in scaling_factor parameters trained on different data more significant, which is more conducive to subsequent clustering based on the scaling_factor parameter.
[0078] S3: The cloud server receives the model parameters uploaded by each vehicle and obtains the scaling_factor parameter of each model block of the locally trained models of all vehicles.
[0079] S4: Based on all the scaling_factor parameters of each model block, the fuzzy C-means clustering algorithm is used to perform hierarchical clustering on each model block to generate a personalized model for each car end.
[0080] As a refinement of the above embodiment, for the case where there is a fuzzy cluster relationship between the car end and the cluster, the present invention adopts a fuzzy clustering method, using the fuzzy C-means clustering algorithm (fuzzyc-means algorithm). The present invention performs clustering based on the scaling factor parameter of the BN layer. Specifically, the scaling factor parameter of the BN layer is a vector. The fuzzyc-means algorithm divides the scaling factor parameter of each car end into C clusters and obtains the car end. Cluster Membership ,According to the membership aggregation parameters, a cluster model is obtained and the cluster model is used to obtain the personalized model of each car end.
[0081] Specifically, this embodiment adopts the idea of hierarchical clustering. For each model block , according to the scaling_factor parameter of the BN layer, the A car-side model block Divide into Clusters, here we use the fuzzy C-means clustering algorithm, and use cosine similarity to measure the similarity between two vectors. A car-side model block The membership of each cluster is calculated and the first cluster is obtained according to the membership aggregation. Model blocks On this basis, we get the first cluster model according to the membership degree. Model blocks In this way, a personalized model is obtained for each car end in a hierarchical manner.
[0082] The following is Taking the model block as an example, the process of generating a personalized model is described in detail:
[0083] Indicates the car side Previous The scaling_factor parameter of the BN layer of the model block is used to cluster the Car side Divide into fuzzy clusters, so that the objective function is minimized:
[0084] ,
[0085] in, Represents the clustering optimization objective, fuzzy factor , is a hyperparameter that determines the degree of blur, It is The membership matrix of the model blocks, Indicates that for Model Nugget Client Cluster The membership degree satisfies and ; Indicates the number of car terminals. It is The centers of the clusters of the model blocks, Indicates the The first Cluster centers, For the car Previous The scaling_factor parameter of the BN layer of the model block and the cluster center The cosine distance is used here instead of the Euclidean distance because is a high-dimensional sparse vector.
[0086] In order to minimize the objective function, the following optimization method is used:
[0087] (1) Yes Normalize it, that is, convert it into a unit vector:
[0088] ,
[0089] in, For the car Previous A unit vector of scaling_factor parameters for each model block.
[0090] (2) Random initialization Cluster Center ,in All are related to Unit vectors of the same dimension; compute each Cosine distance from each cluster center:
[0091] ,
[0092] in, Represents the dot product of vectors.
[0093] (3) Update the membership as follows, and update the membership matrix based on the membership:
[0094] ,
[0095] (4) Update the cluster center as follows and recalculate each Cosine distance from each cluster center:
[0096] ,
[0097] Iterate the above (3)-(4) until the cluster centers converge.
[0098] After the above steps, the membership matrix is obtained , that is, each The membership relationship with each cluster center. Aggregate the model block according to the membership matrix of Cluster model blocks, it is worth noting that It is not standardized.
[0099] ,
[0100] ,
[0101] or
[0102] ,
[0103] in, For the car Previous The scaling_factor parameter of the BN layer of the model block, Indicates the car side Previous The convolutional layer parameters in the model block, Indicates the car side Previous The fully connected layer parameters in each model block; Representation model nugget The corresponding The scaling_factor parameter of the BN layer in the cluster model block; Representation model nugget The corresponding The convolutional layer parameters in each cluster model block; Representation model nugget The corresponding The connection layer parameters in the cluster model nugget.
[0104] Finally, according to the model block of The cluster model blocks and membership matrix generate personalized model blocks for each car end, namely:
[0105] The parameters of the personalized model block are based on the The cluster model blocks and membership matrix are generated as follows:
[0106] ,
[0107] ,
[0108] or
[0109] ,
[0110] in, Indicates the car side Previous The scaling_factor parameter of the BN layer in the personalized model block, Indicates the car side Previous The convolutional layer parameters in the personalized model blocks, Indicates the car side Previous The fully connected layer parameters in each personalized model block.
[0111] After the above process is done for each model block, the car side of personalized model blocks are combined into a complete personalized model .
[0112] S5: Send the personalized model to the corresponding car end for application.
[0113] If the training reaches the preset number of rounds or the model reaches the preset accuracy, the training is stopped; otherwise, S2 is continued.
[0114] The present invention verifies the effectiveness of the algorithm on the public dataset cifar10; the classic network architecture VGG16 is used to help the car side identify obstacles, pedestrians and other vehicles ahead. In the simulation experiment, in order to simulate the data heterogeneity between different clusters, a data set with 10 categories is divided into three clusters, where each cluster contains 3-4 categories of data. The data of each cluster is then randomly divided into 10 car sides to simulate that the data between different car sides can complement each other and learn jointly. The scaling_factor of the last BN layer, the second to last BN layer and the first BN layer of VGG16 are used as the basis for clustering, and the following is obtained: Figure 4-Figure 6 The three similarity matrices shown.
[0115] From the experimental results:
[0116] The similarity matrix using the last BN layer as the clustering basis is clearly divided into three clusters, which shows that applying a norm regularization to the scaling_factor parameter of the BN layer and clustering based on the scaling_factor parameter can well capture the correlation between the car-side data.
[0117] The similarity matrix using the first BN layer as the clustering basis shows that the cluster relationship reflected by the parameters of this layer is fuzzy, so it is necessary to use a fuzzy clustering algorithm based on membership.
[0118] The clustering results obtained by using the scaling_factor of the last BN layer, the second to last BN layer, and the first BN layer as the basis for clustering are very different. In order to maximize the knowledge of each car-side model, it is necessary to use a hierarchical clustering strategy.
[0119] Example 2
[0120] Combine Figure 7 As shown, an embodiment of the present disclosure provides an efficient hierarchical soft clustering federated learning apparatus 300, comprising a processor 304 and a memory 301. Optionally, the apparatus may further comprise a communication interface 302 and a bus 303. The processor 304, the communication interface 302, and the memory 301 may communicate with each other via the bus 303. The communication interface 302 may be used for information transmission. The processor 304 may invoke logic instructions in the memory 301 to execute the efficient hierarchical soft clustering federated learning method of the above embodiment.
[0121] In addition, the logic instructions in the memory 301 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.
[0122] Memory 301, as a computer-readable storage medium, can be used to store software programs and computer-executable programs, such as program instructions / modules corresponding to the methods in the embodiments of the present disclosure. Processor 304 executes the program instructions / modules stored in memory 301 to perform functional applications and data processing, thereby implementing the efficient hierarchical soft clustering federated learning method in the above-mentioned embodiments.
[0123] The memory 301 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data generated based on the use of the terminal device. Furthermore, the memory 301 may include high-speed random access memory and non-volatile memory.
[0124] An embodiment of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the above-mentioned efficient hierarchical soft clustering federated learning method.
[0125] The aforementioned computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.
Claims
1. An efficient hierarchical soft clustering federated learning method, characterized by: The following steps are involved: The cloud server initializes a global model consisting of L model blocks, where each model block includes a convolutional layer or a fully connected layer and its corresponding BN layer; The global model is sent to each vehicle, and each vehicle trains the model based on local data, with the local objective function being: , in, Indicates the car side The local objective function on Indicates the amount of data on the car side, It is The cross entropy loss of sample data is It's the car side Previous The one-norm regularization loss of the scaling_factor parameter of the BN layer of the model block, is a super parameter, Indicates the car side Personalized models on The cloud server receives the model parameters uploaded by each car end and obtains the scaling_factor parameters of each model block of the locally trained models of all cars end; Based on all the scaling_factor parameters of each model block, the fuzzy C-means clustering algorithm is used to perform hierarchical clustering on each model block to generate a personalized model for each car end; Sending the personalized model to the corresponding car end for application; The fuzzy C-means clustering algorithm is used to hierarchically cluster the model blocks according to the scaling_factor parameter of each model block to generate a personalized model for each car end. The specific method is as follows: All scaling_factor parameters of each model block are divided into fuzzy clusters and obtain the membership matrix; According to the division Fuzzy clusters and membership matrices are aggregated into each model block. Cluster model blocks; According to each model block The cluster model blocks and membership matrix generate L personalized model blocks for each car end; The car side combines L personalized model blocks into a complete personalized model.
2. The efficient hierarchical soft clustering federated learning method according to claim 1, characterized in that: The fuzzy C-means clustering algorithm uses cosine distance to measure the similarity between the scaling_factor parameter and the cluster center.
3. The efficient hierarchical soft clustering federated learning method according to claim 2, characterized in that: The objective function of the fuzzy C-means clustering algorithm is: , in, represents the clustering optimization objective, is the fuzzy factor; It is The membership matrix of the model blocks; Indicates that for Model Nugget Client Cluster The membership degree satisfies and ; Indicates the number of car terminals; It is The centers of the clusters of the model blocks, Indicates the The first cluster centers; For the car Previous The scaling_factor parameter of the BN layer of the model block and the cluster center The cosine distance.
4. The efficient hierarchical soft clustering federated learning method according to claim 3, characterized in that: The specific method of minimizing the objective function of the fuzzy C-means clustering algorithm is as follows: Normalize the scaling_factor parameter, that is, convert it into a unit vector , For the car Previous A unit vector of scaling_factor parameters for each model block; Random initialization cluster centers and calculate each unit vector cosine distance from each cluster center; Update the membership and update the membership matrix based on the membership: , Update the cluster centers and recalculate each The cosine distance from each cluster center until the cluster centers converge.
5. The efficient hierarchical soft clustering federated learning method according to claim 3, characterized in that: The parameters of the cluster model block are obtained by aggregating the convolutional layer or fully connected layer parameters of all model blocks in the corresponding cluster according to the membership matrix. The specific method is as follows: , , or , in, For the car Previous The scaling_factor parameter of the BN layer of the model block, Indicates the car side Previous The convolutional layer parameters in the model block, Indicates the car side Previous The fully connected layer parameters in each model block; Representation model nugget The corresponding The scaling_factor parameter of the BN layer in the cluster model block; Representation model nugget The corresponding The convolutional layer parameters in each cluster model block; Representation model nugget The corresponding The connection layer parameters in the cluster model nugget.
6. The efficient hierarchical soft clustering federated learning method according to claim 5, characterized in that: The parameters of the personalized model block are based on the The cluster model blocks and membership matrix are generated as follows: , , or , in, Indicates the car side Previous The scaling_factor parameter of the BN layer in the personalized model block, Indicates the car side Previous The convolutional layer parameters in the personalized model blocks, Indicates the car side Previous The fully connected layer parameters in each personalized model block.
7. The efficient hierarchical soft clustering federated learning method according to claim 3, characterized in that: Fuzzy Factor .
8. An efficient hierarchical soft clustering federated learning device, comprising a processor and a memory storing program instructions, characterized in that: The processor is configured to execute the efficient hierarchical soft clustering federated learning method according to any one of claims 1 to 7 when running the program instructions.
9. A computer-readable storage medium, characterized in that A computer program is stored thereon, which, when executed by a processor, implements the efficient hierarchical soft clustering federated learning method as described in any one of claims 1 to 7 above.
Citation Information
Patent Citations
Two-stage clustering algorithm based on difference evolution and fuzzy C-means
CN104881688A
Federal learning multi-granularity grouping fine tuning method with high communication efficiency
CN119109943A