A clustering federation method and device based on data distribution differences
By calculating the label probability and feature vector differences of the data set, iterative clustering is performed to generate client clusters, which solves the model training bias problem caused by data distribution differences in Non-IID scenarios and improves the accuracy of training results and clustering effect.
Patent Information
- Application Number
- CN202310149034.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-02-22
AI Technical Summary
In the Non-IID scenario, the client dataset has problems such as insufficient sample number, missing features of some samples, and large differences in the same feature values between two samples, which makes the global model unable to adapt well to each client dataset.
By calculating the label probability distribution and eigenvector differences of the original data set, and using relative entropy and covariance matrix to calculate the data distribution differences between clients, clustering operations are performed to generate client clusters. Iterative clustering is performed in each round of training, and a new cluster center model is generated using the weighted average method of training accuracy.
It improves the accuracy of client clustering and the accuracy of training results, balances the data feature extraction model performance of each client, avoids the deviation of clustering results, and optimizes the client clustering process.
Smart Images

Figure CN116910587B_ABST
Abstract
Description
Technical Field
[0001] The technical field to which the present invention relates is the field of data distribution technology, and in particular to a clustering federation method and device based on data distribution differences. Background Art
[0002] Federated learning is a machine learning framework that allows local clients to participate in the training of joint statistical models on large-scale distributed data while protecting the privacy of their data. The most common setup for federated learning is to use a central entity (the server) to aggregate local learning models from many clients to learn a global optimal model, with the goal of achieving global model performance that approximates that of the centralized training model. Federated learning has been shown to work well when data is independent and identically distributed (IID). The most commonly used federated averaging algorithm (FedAvg) uses the ratio of the number of samples from a single client to the number of samples from all clients as weights, and aggregates the global model on the server by taking a weighted average of the model parameters. In real-world scenarios, data is distributed across a large number of clients and may not be independent and identically distributed (Non-IID), which is not conducive to training an optimal global model.
[0003] In existing work, researchers have proposed methods to improve model training problems in Non-IID scenarios from two perspectives: local model updates and server-side model aggregation, and have achieved certain results. However, in Non-IID scenarios, the data sets of different clients show large differences, and the global model cannot be well adapted to the training of each client's data set. The proposed clustering federation method effectively alleviates this problem. The client uploads the local training model to the server, and the server calculates the differences between the local training models, performs client clustering, and assigns clients corresponding to models with large differences to different clusters, thereby achieving the purpose of improving the training accuracy of the cluster center model.
[0004] Existing federated clustering algorithms often operate at the model level, primarily using differences in model parameters as the basis for client clustering. In non-IID scenarios, client datasets can suffer from issues such as insufficient sample size, missing features in some samples, and significant differences in the value of the same feature between two samples. These issues can lead to a certain degree of deviation in model parameters during training, adversely affecting the clustering process. Therefore, it is necessary to consider the impact of raw data on clustering results. Summary of the Invention
[0005] The purpose of this section is to summarize some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this section and the abstract and title of this application to avoid obscuring the purpose of this section, the abstract and the title of the invention, and such simplifications or omissions should not be used to limit the scope of the present invention.
[0006] In view of the above-mentioned problems, the present invention is proposed.
[0007] Therefore, the technical problem solved by the present invention is that the client data set has problems such as insufficient number of samples, missing features of some samples, and large difference in the same feature value between two samples.
[0008] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0009] In a first aspect, an embodiment of the present invention provides a clustering federation method based on data distribution differences, including:
[0010] Calculate the label probability distribution and distribution difference of the original data set based on the label category and the number of samples corresponding to the label;
[0011] Using the feature extraction network, extract the data feature vector related to the original data set;
[0012] Based on the data feature vector, calculating the difference in data feature distribution between client data sets;
[0013] Based on the difference between the label probability distribution of the data set and the data feature distribution, a clustering operation is performed to obtain client clusters, and the client clusters are iteratively clustered.
[0014] As a preferred solution of clustering federation method based on data distribution differences,
[0015] The calculation of the label probability distribution of the original data set includes: for a certain client p, the statistical data set D p The total number of samples M and the number of label types n included; count the number of samples M included in each label category separately i (i=1,2,……,n), with the ratio M i / M is the probability distribution value of the i-th category label.
[0016] As a preferred solution of clustering federation method based on data distribution differences,
[0017] The calculation of the label probability distribution difference of the original data set includes: defining the data set D p 、D q The difference in label probability distribution between is calculated as:
[0018]
[0019] where p(i|D p ) represents the dataset D p In the matrix, the number of samples corresponding to label i accounts for the proportion of the total number of samples, n represents the number of categories of labels in the dataset; s1 is a component of the matrix S1 representing the difference in label probability distribution.
[0020] As a preferred solution of clustering federation method based on data distribution differences,
[0021] The method of extracting the data feature vector related to the original data set includes: selecting a feature extraction network, each client p receives the original data set D p As the network input, the network training is carried out; after the network training is completed, the 2048-dimensional feature vector v is output. i , complete feature extraction.
[0022] As a preferred solution of clustering federation method based on data distribution differences,
[0023] The clustering operation includes: obtaining clients (K1, K2...K m ) After obtaining the label probability distribution of the dataset, the label probability distribution differences between different clients are calculated according to the relative entropy formula; after obtaining the data feature distribution of each client dataset, the data feature distribution differences between different clients are calculated according to the covariance matrix formula; the two parts of the difference are integrated to synthesize the distance matrix S, which is an m*m matrix, where s ij represents the data distribution difference between client i and client j. S is added as an input parameter to the clustering operation process to obtain client clusters (C1, C2...C n ).
[0024] As a preferred solution of clustering federation method based on data distribution differences,
[0025] The iterative clustering includes: in each round of training, after the client completes the data feature extraction operation, the server aggregates all feature extraction models, and then returns the aggregated global feature extraction model to the client for use in the next round of data feature extraction operation.
[0026] As a preferred solution of clustering federation method based on data distribution differences,
[0027] The server performs model aggregation on all feature extraction models, including: the aggregation method adopts a weighted average method based on training accuracy, and the specific steps include: the client cluster sends the cluster center model to all clients in the cluster, the clients perform local training, and send the local model and training accuracy obtained after each round of training to the cluster; the cluster records the training accuracy of each client in each round, and uses the average of the training accuracy of the last three rounds as the aggregation weight of the local model, and generates a new cluster center model by weighted average.
[0028] In a second aspect, an embodiment of the present invention provides a system, characterized by including:
[0029] The label probability distribution difference calculation module is used to calculate the label probability distribution and distribution difference of the original data set based on the label category and the number of samples corresponding to the label of the original data set;
[0030] A data feature distribution difference calculation module is used to extract data feature vectors related to the original data set using a feature extraction network; based on the data feature vectors, calculate the differences in data feature distribution between client data sets;
[0031] The clustering module is used to perform clustering operations based on the difference between the label probability distribution of the data set and the data feature distribution, obtain client clusters, and iteratively cluster the client clusters.
[0032] In a third aspect, an embodiment of the present invention provides a computing device, including:
[0033] memory and processor;
[0034] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the one or more programs are executed by the one or more processors, the one or more processors implement the clustering federation method based on data distribution differences as described in any embodiment of the present invention.
[0035] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the clustering federation method based on data distribution differences.
[0036] Beneficial effects of the present invention: The present invention mainly considers the impact of differences in data distribution on client clustering results in federated learning, and provides a clustering federation method based on data distribution differences. Starting from the two directions of data set label probability distribution and data internal feature distribution, the present invention introduces the calculation of differences between different client data distributions, combines existing clustering methods, generates more accurate client clusters, and obtains better training results. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. Those skilled in the art can also derive other drawings based on these drawings without inventive effort. Among them:
[0038] Figure 1 This is an overall framework diagram of the clustering federation method based on data distribution differences described in the first embodiment of the present invention;
[0039] Figure 2 1. The figure shows experimental results based on the MNIST dataset and the DDoS attack traffic dataset in a simulation experiment of the clustering federation method based on data distribution differences described in the second embodiment of the present invention, and experimental results compared with commonly used methods;
[0040] Figure 3 This is a diagram of the ablation experiment results in a simulation experiment of the clustering federation method based on data distribution differences described in the second embodiment of the present invention. DETAILED DESCRIPTION
[0041] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0042] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from those described herein. Those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0043] Secondly, the term "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in various places throughout this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive of other embodiments.
[0044] The present invention is described in detail with reference to schematic diagrams. For ease of illustration, cross-sectional views of device structures may be partially enlarged and not to scale when describing embodiments of the present invention. Furthermore, the schematic diagrams are merely illustrative and should not limit the scope of the present invention. Furthermore, in actual production, the three-dimensional dimensions of length, width, and depth should be included.
[0045] In the description of the present invention, it should be noted that the terms "upper, lower, inner, and outer" and other references to orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first, second, or third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0046] In this disclosure, unless otherwise specified or limited, the terms "mounted," "connected," and "connected" should be interpreted broadly. For example, they may refer to fixed, removable, or integral connections. They may also refer to mechanical, electrical, or direct connections, indirect connections through an intermediary, or internal communication between two components. Those skilled in the art will understand the specific meanings of these terms in this disclosure.
[0047] Example 1
[0048] Reference Figure 1 , which is the first embodiment of the present invention, provides a clustering federation method based on data distribution differences, including:
[0049] S1: Calculate the label probability distribution and distribution difference of the original data set based on the label category and the number of samples corresponding to the label;
[0050] Specifically, the calculation of the label probability distribution of the original data set includes: for a certain client p, the statistical data set D p The total number of samples M and the number of label types n included; count the number of samples M included in each label category separately i (i=1,2,……,n), with the ratio M i / M is the probability distribution value of the i-th category label.
[0051] The calculation of the label probability distribution difference of the original data set includes: defining the data set D p 、D q The difference in label probability distribution between is calculated as:
[0052]
[0053] where p(i|D p ) represents the dataset D p In the matrix, the number of samples corresponding to label i accounts for the proportion of the total number of samples, n represents the number of categories of labels in the dataset; s1 is a component of the matrix S1 representing the difference in label probability distribution.
[0054] S2: Using a feature extraction network, extract data feature vectors related to the original data set; based on the data feature vectors, calculate the differences in data feature distribution between client data sets;
[0055] Specifically, the extraction of data feature vectors related to the original data set includes: selecting a feature extraction network, each client p takes the original data set D p As the network input, the network training is carried out; after the network training is completed, the 2048-dimensional feature vector v is output. i , complete feature extraction.
[0056] S3: performing a clustering operation based on the difference between the label probability distribution of the data set and the data feature distribution to obtain client clusters, and iteratively clustering the client clusters.
[0057] Specifically, the clustering operation includes: obtaining clients (K1, K2...K m ) After obtaining the label probability distribution of the dataset, the label probability distribution differences between different clients are calculated according to the relative entropy formula; after obtaining the data feature distribution of each client dataset, the data feature distribution differences between different clients are calculated according to the covariance matrix formula; the two parts of the difference are integrated to synthesize the distance matrix S, which is an m*m matrix, where s ij represents the data distribution difference between client i and client j. S is added as an input parameter to the clustering operation process to obtain client clusters (C1, C2...C n ).
[0058] The iterative clustering includes: in each round of training, after the client completes the data feature extraction operation, the server aggregates all feature extraction models, and then returns the aggregated global feature extraction model to the client for use in the next round of data feature extraction operation.
[0059] The server performs model aggregation on all feature extraction models, including: the aggregation method adopts a weighted average method based on training accuracy, and the specific steps include: the client cluster sends the cluster center model to all clients in the cluster, the clients perform local training, and send the local model and training accuracy obtained after each round of training to the cluster; the cluster records the training accuracy of each client in each round, and uses the average of the training accuracy of the last three rounds as the aggregation weight of the local model, and generates a new cluster center model by weighted average.
[0060] It should be noted that the purpose of iterative clustering is to ensure the accuracy and timeliness of the clustering results. The client iterative clustering method can balance the performance of the feature extraction model of each client data to avoid huge deviations in the clustering results due to poor feature extraction effects of individual clients, thereby achieving continuous optimization of the client clustering process.
[0061] Example 2
[0062] Reference Figure 2-3 , is an embodiment of the present invention, which provides a clustering federation method based on data distribution differences. In order to verify the beneficial effects of the present invention, scientific demonstration is carried out through simulation experiments.
[0063] S1: Calculation of label probability distribution differences
[0064] Step 1 is used to calculate the difference between the probability distributions of different client data labels. Assume that D p 、D q It is a private dataset from two clients. If the number of samples in a certain category label differs greatly, that is, the label probability values differ greatly, then it is considered that there are large differences in the data distribution of the two datasets; conversely, if the difference in their corresponding label probability values is small, then it is considered that the data distribution of the two datasets is similar to a certain extent. The label category of a dataset is the most intuitive feature of a dataset. Considering that relative entropy can be used to measure the difference between two probability distributions, the relative entropy formula is selected as the calculation formula for the difference in the probability distribution of dataset labels. Assume that P(x) and Q(x) are two probability distributions on the random variable x. In the case of discrete random variables, the definition of relative entropy is as follows:
[0065]
[0066] Define dataset D p 、D q The calculation method of the difference in label probability distribution between :
[0067]
[0068] where p(i|D p ) represents the dataset Dp Where, the number of samples corresponding to label i accounts for the proportion of the total number of samples, and n represents the number of categories of labels in the dataset. s1 is a component of matrix S1, representing the difference in label probability distribution between the two clients.
[0069] S2: Calculation of data feature distribution differences
[0070] Step 1 measures the difference in label probabilities between the two datasets and does not reflect the correlation between the data within the datasets. In step 2, a calculation method based on the difference in internal data features is defined to measure the difference in internal features between the two datasets.
[0071] The original dataset is fed into the feature extraction network for training, and the data feature vector of the local dataset is extracted before the network output layer. Since the extracted data features are multi-dimensional distributions, the covariance matrix can measure the correlation between multiple dimensions. Therefore, the mean and covariance matrix are used to calculate the difference between the two high-dimensional distributions:
[0072]
[0073] where u Dp Represented by the dataset D p The mean of the extracted eigenvectors, ∑ Dp represents the covariance matrix calculated from the eigenvectors, and Tr represents the matrix trace. s2 is a component of matrix S2, representing the difference in the data feature distribution between the two clients. Because the data feature vectors extracted by the feature extraction network, rather than the original data, are transmitted between the clustering center and the client during the calculation process, data privacy and security are guaranteed.
[0074] S3: Iterative clustering process
[0075] At the beginning of each training round, each client extracts feature vectors from its local dataset using a feature extraction network. The local feature extraction network models are then aggregated on the server to form a global feature extraction model. The server then returns this model to each client, which serves as the feature extraction model for the next round of training. Simultaneously, the distance matrix representing data distribution differences dynamically changes.
[0076] To ensure the accuracy and timeliness of clustering results, a client-side iterative clustering approach is employed. During each round of training, after the client completes data feature extraction, the server aggregates all feature extraction models and then returns the aggregated global feature extraction model to the client for the next round of data feature extraction. This client-side iterative clustering approach balances the performance of each client's data feature extraction model, preventing significant deviations in clustering results due to poor feature extraction performance on individual clients, thereby ensuring continuous optimization of the client-side clustering process.
[0077] After completing client clustering, the cluster aggregates the trained models of the internal clients to generate an initial cluster center model. The cluster distributes the cluster center model to all clients within the cluster. The clients perform local training and send the local models and training accuracy obtained after each round of training to the cluster. The cluster records the training accuracy of each client for each round and uses the average of the training accuracy of the last three rounds as the aggregation weight for the local models. A new cluster center model is generated by weighted average. After each round of training, all cluster center models are aggregated on the server for model aggregation to generate a global training model for testing. Training stops when the specified number of training rounds is reached or the model converges.
[0078] The following table shows the impact of factors such as the distance calculation formula on the experimental results in the simulation experiment under the DDoS attack traffic dataset:
[0079]
[0080] It should be noted that in order to measure the impact of different distance calculation formulas on the experimental results, the distance calculation formulas used are: covariance matrix formula (cov.), cosine similarity distance formula (cos.), Manhattan distance formula (L1), and Euclidean distance formula (L2).
[0081] In summary, the clustering federation method proposed in this paper, based on data distribution differences, utilizes both the label and feature information of the original data, compared to existing clustering federation methods based on model differences. This method achieves higher accuracy during client-side clustering. Furthermore, in tests on selected datasets, the method designed by this invention achieved higher experimental accuracy than existing methods, validating the effectiveness of the method.
Claims
1. A clustering federation method based on data distribution differences, characterized in that: include: Calculate the label probability distribution and distribution difference of the original data set based on the label category and the number of samples corresponding to the label; Using the feature extraction network, extract the data feature vector related to the original data set; Based on the data feature vector, calculating the difference in data feature distribution between client data sets; Performing a clustering operation based on the difference between the label probability distribution of the data set and the data feature distribution to obtain client clusters, and iteratively clustering the client clusters; The clustering operation includes: obtaining clients (K1, K2...K m ) After obtaining the label probability distribution of the dataset, the label probability distribution differences between different clients are calculated according to the relative entropy formula; after obtaining the data feature distribution of each client dataset, the data feature distribution differences between different clients are calculated according to the covariance matrix formula; the two parts of the difference are integrated to synthesize the distance matrix S, which is an m*m matrix, where s ij represents the data distribution difference between client i and client j. S is added as an input parameter to the clustering operation process to obtain client clusters (C1, C2...C n ).
2. The clustering federation method based on data distribution differences according to claim 1, characterized in that: The calculation of the label probability distribution of the original data set includes: for a certain client p, the statistical data set D p The total number of samples M and the number of label types n included; count the number of samples M included in each label category separately i (i=1,2,……,n), with the ratio M i / M is the probability distribution value of the i-th category label.
3. The clustering federation method based on data distribution differences according to claim 2, characterized in that: The calculation of the label probability distribution difference of the original data set includes: defining the data set D p 、D q The difference in label probability distribution between is calculated as: where p(i|D p ) represents the dataset D p In the matrix, the number of samples corresponding to label i accounts for the proportion of the total number of samples, n represents the number of categories of labels in the dataset; s1 is a component of the matrix S1 representing the difference in label probability distribution.
4. The clustering federation method based on data distribution differences according to claim 3, characterized in that: The method of extracting the data feature vector related to the original data set includes: selecting a feature extraction network, each client p receives the original data set D p As the network input, the network training is carried out; after the network training is completed, the 2048-dimensional feature vector v is output. i , complete feature extraction.
5. The clustering federation method based on data distribution differences according to claim 4, characterized in that: The iterative clustering includes: in each round of training, after the client completes the data feature extraction operation, the server aggregates all feature extraction models, and then returns the aggregated global feature extraction model to the client for use in the next round of data feature extraction operation.
6. The clustering federation method based on data distribution differences according to claim 5, characterized in that: The server performs model aggregation on all feature extraction models, including: the aggregation method adopts a weighted average method based on training accuracy, and the specific steps include: the client cluster sends the cluster center model to all clients in the cluster, the clients perform local training, and send the local model and training accuracy obtained after each round of training to the cluster; the cluster records the training accuracy of each client in each round, and uses the average of the training accuracy of the last three rounds as the aggregation weight of the local model, and generates a new cluster center model by weighted average.
7. A clustering federation system based on data distribution differences, characterized in that: include: The label probability distribution difference calculation module is used to calculate the label probability distribution and distribution difference of the original data set based on the label category and the number of samples corresponding to the label of the original data set; A data feature distribution difference calculation module is used to extract data feature vectors related to the original data set using a feature extraction network; based on the data feature vectors, calculate the differences in data feature distribution between client data sets; A clustering module, configured to perform clustering operations based on the difference between the label probability distribution of the data set and the data feature distribution, obtain client clusters, and iteratively cluster the client clusters; The clustering operation includes: obtaining clients (K1, K2...K m ) After obtaining the label probability distribution of the dataset, the label probability distribution differences between different clients are calculated according to the relative entropy formula; after obtaining the data feature distribution of each client dataset, the data feature distribution differences between different clients are calculated according to the covariance matrix formula; the two parts of the difference are integrated to synthesize the distance matrix S, which is an m*m matrix, where s ij represents the data distribution difference between client i and client j. S is added as an input parameter to the clustering operation process to obtain client clusters (C1, C2...C n ).
8. A computing device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the clustering federation method based on data distribution differences described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the clustering federation method based on data distribution differences as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Training method for improving federal learning model performance based on two-stage clustering, and storage equipment
CN113313266A
Active learning customer selection method and device under clustering federated learning framework
CN114723074A