Security verification method, terminal and medium for cluster federated learning clustering process
Through cluster entry attack methods, the attacker disguises the clustering process of entering the cluster federated learning, verifies the security of the clustering process, solves the problem that the security of the clustering algorithm is not fully studied in the existing technology, and achieves efficient cluster entry and inference success rates.
Patent Information
- Application Number
- CN202510176178.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-18
AI Technical Summary
During cluster federated learning clustering, the security of clustering algorithm has not been thoroughly studied. Malicious clients can lead to performance degradation of cluster models and data leakage through cluster attacks.
By designing a cluster attack method, the attacker eavesdrops on some parameters of the victim model, builds an approximate victim model, and uses auxiliary data sets to filter heterogeneous data, trains victim similar models, and disguise them into the victim cluster to verify the security of the clustering process.
Achieve a higher success rate of cluster entry and inference success rate, revealing the vulnerability of existing cluster federated learning defenses, helping to develop safer clustering and parameter aggregation methods.
Smart Images

Figure CN119646811B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning technology, and in particular to a security verification method for a cluster federated learning clustering process, as well as a computer terminal and a computer-readable storage medium applying the method. Background Art
[0002] Federated learning (FL) is a distributed learning paradigm that facilitates multiple clients to collaboratively train shared models without the need for private data exchange. This approach ensures efficient communication, significantly reduces overhead, and maintains the privacy of client data. Therefore, federated learning is widely used in edge intelligence with a large number of smart terminal devices. However, in real-world scenarios, data on different clients are not independent and identically distributed (IID). Although early federated learning studies proposed methods to deal with data heterogeneity, training global models on non-IID data isomorphism will inevitably lead to performance degradation. To address this issue, cluster federated learning (CFL) was proposed as a paradigm to mitigate the impact of data heterogeneity and improve the performance of global models.
[0003] In the CFL scenario, clients with similar data distribution are aggregated into clusters, and different models are trained for each cluster. This approach ensures data homogeneity within each cluster, thereby improving the performance of the cluster model. Although CFL can improve model performance in the presence of data heterogeneity, the security of the clustering algorithm used in the CFL process has not been thoroughly studied. If a malicious client has entered a specific cluster, it will inevitably pose a greater threat to other clients within the cluster, which may include targeted operations such as using GAN (Generative Adversarial Network) to directly leak private data or embed backdoors in the cluster model. This operation emphasizes the secondary attack that CFL may suffer from malicious clustering. It is necessary to assume that it has entered the cluster for subsequent verification. Therefore, this attack method has limitations, which limits the effective security verification of the clustering algorithm used in the CFL process. Summary of the invention
[0004] In order to solve the technical problems existing in the prior art, the present invention provides a security verification method, terminal and medium for the cluster federated learning clustering process. By designing a unique intra-cluster attack method, it makes it easier for malicious clients to illegally enter a specific cluster, providing more effective security verification for different clustering processes.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] The present invention discloses a security verification method for a cluster federated learning clustering process, comprising: S1-S7.
[0007] S1. Perform cluster federated learning tasks according to a specific clustering scheme. The architecture of this task consists of a parameter server and multiple benign clients.
[0008] S2. During the execution of the task, a malicious client, namely an attacker, is deployed in the architecture; the attacker has an auxiliary dataset D a , D a contains data similar to that of a designated benign client, the victim.
[0009] S3. The attacker eavesdrops on some of the victim’s model parameters W v ,Will W v Cluster model parameters sent by the parameter server W c Weighted to build an approximate victim model M v .
[0010] S4. Auxiliary Dataset D a The samples in are sequentially input into the approximate victim model M v In the example, a shadow dataset is constructed based on the output results. D s .
[0011] S5. Attacker builds k shadow model, the shadow dataset D s Divide into k Group-to-group shadow training set D train and shadow test set D test , and then use the shadow training set and shadow test set to train the corresponding shadow model; each group of shadow training set D train and shadow test set D test The samples in are input into the corresponding shadow model to obtain the prediction metric to generate a prediction set D pred .
[0012] S6. Using the prediction set D pred Train a meta-classifier to classify the shadow dataset D s Input into the victim model, obtain the predicted output information and input it into the meta-classifier for classification to obtain the screened data set Df .
[0013] S7. Use of Dataset D f Update the attacker model and upload it to the parameter server to determine whether the attacker successfully enters the victim's cluster to verify the security of the clustering process.
[0014] As a further improvement of the present invention, in step S1, the execution of the cluster federated learning task includes: S11~S13.
[0015] S11. The parameter server sends the initialized model and hyperparameters to each benign client;
[0016] S12. Each benign client uses the initialization model, hyperparameters, and its own private data to train client parameters, and then uploads the trained client parameters to the parameter server;
[0017] S13. The parameter server determines whether the current round meets the clustering conditions. If so, it obtains clustering according to a specific clustering scheme; otherwise, it performs ordinary federated learning. If the clustering conditions cannot be met for consecutive preset rounds, the clustering phase ends, and the parameter server continues to perform ordinary federated learning until the task is completed.
[0018] As a further improvement of the present invention, the specific clustering scheme is composed of one or more of four clustering algorithms: k-means, agglomerative clustering, spectral clustering and DBSCAN.
[0019] As a further improvement of the present invention, in step S13, the parameter server calculates the maximum norm of the parameters uploaded by each client in the cluster 1 and mean norm 2 ,like 1 and 2 If both exceed the predefined threshold, the current round is judged to meet the clustering condition, otherwise it does not meet the clustering condition.
[0020] As a further improvement of the present invention, in step S5, the attacker constructs multiple shadow models based on the model structure information, hyperparameter settings, and knowledge about the training algorithm sent by the parameter server;
[0021] The training set of each shadow model is either an intersecting subset or a non-intersecting subset of the shadow training set;
[0022] The prediction set D predIt includes a prediction training set and a prediction test set; wherein the prediction training set corresponds to the output of the shadow training set, and the prediction test set corresponds to the output of the shadow test set.
[0023] As a further improvement of the present invention, in step S6, the meta-classifier can learn the difference representation between the shadow training set and the shadow test set, thereby generalizing to the difference representation of the training set and the test set in the cluster federated learning scenario.
[0024] As a further improvement of the present invention, in step S2, the auxiliary data set D a The malware contains data that is 20%-50% similar to the victim.
[0025] As a further improvement of the present invention, in step S4, after the auxiliary data set D a The samples in are sequentially input into the approximate victim model M v After that, by calculating the cross entropy loss, samples whose cross entropy loss is less than the loss threshold are extracted to form the shadow dataset. D s .
[0026] The present invention also discloses a computer terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the security verification method for the cluster federated learning clustering process as described above are implemented.
[0027] The present invention also discloses a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the security verification method for the cluster federated learning clustering process as described above are implemented.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] The present invention provides a clustering attack method for cluster federated learning clustering process. The attacker uses part of the victim model parameters that are eavesdropped on to build an approximate victim model, and then removes the heterogeneous data of the attacker's auxiliary data set in combination with the trained inference model. Finally, the victim similar model is trained using the remaining auxiliary data set and uploaded to the server to mislead its clustering process, thereby having a high clustering success rate and inference success rate, and having a significant impact on the clustering process of cluster federated learning. This damage causes the server to produce erroneous clustering results. When cluster federated learning is applied, this attack method can be used to perform security verification of the clustering process, revealing the vulnerability of the existing cluster federated learning defense, and helping to develop a safer cluster federated learning clustering and parameter aggregation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a flowchart of a security verification method for a cluster federated learning clustering process in Example 1 of the present invention.
[0031] Figure 2 This is a diagram of a cluster attack model of a malicious client in Example 1 of the present invention.
[0032] Figure 3 This is a model diagram of the training phase of the data set filtering algorithm in Example 1 of the present invention.
[0033] Figure 4 This is a model diagram of the filtering stage of the data set filtering algorithm in Example 1 of the present invention.
[0034] Figure 5 This is a schematic diagram of the process of training and clustering clients using an agglomerative clustering algorithm in Example 1 of the present invention.
[0035] Figure 6 This is a schematic diagram of the process of training and clustering clients using the k-means algorithm in Example 1 of the present invention.
[0036] Figure 7 Schematic diagram of the process of training and clustering clients using the spectral clustering algorithm in Embodiment 1 of the present invention.
[0037] Figure 8 This is a schematic diagram of the process of training and clustering clients using the DBSCAN clustering algorithm in Example 1 of the present invention.
[0038] Fig. 9 Schematic diagram of the influence of the monitoring parameter ratio α on various clustering algorithms under the EMNIST data set in Example 1 of the present invention.
[0039] Fig.10 Schematic diagram of the influence of the monitoring parameter ratio α on various clustering algorithms under the FMNIST data set in Example 1 of the present invention.
[0040] Fig.11 Schematic diagram of the influence of similarity β on various clustering algorithms under the EMNIST data set in Example 1 of the present invention.
[0041] Fig.12 Schematic diagram of the influence of similarity β on various clustering algorithms under the FMNIST data set in Example 1 of the present invention.
[0042] Fig.13 This is an architecture diagram of a computer terminal according to Embodiment 2 of the present invention. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] Example 1
[0045] See also Figure 1 , this embodiment provides a security verification method for cluster federated learning clustering process, including: S1~S7.
[0046] S1. Obtain the architecture of traditional cluster federated learning and perform cluster federated learning tasks according to a specific clustering scheme.
[0047] See also Figure 2 ,The architecture consists of a parameter server and multiple benign clients. In this system there is a malicious client i.e. the attacker, who wants to attack a benign client i.e. the victim, will perform multiple intra-cluster attacks in order to enter the same cluster as the victim.
[0048] In step S1, the execution of the cluster federated learning task includes: S11~S13.
[0049] S11. The parameter server sends the initialized model and hyperparameters to each benign client;
[0050] S12. Each benign client uses the initialization model, hyperparameters, and its own private data to train client parameters, and then uploads the trained client parameters to the parameter server;
[0051] S13. The parameter server determines whether the current round meets the clustering conditions. If so, it obtains clustering according to a specific clustering scheme; otherwise, it performs ordinary federated learning. If the clustering conditions cannot be met for consecutive preset rounds, the clustering phase ends, and the parameter server continues to perform ordinary federated learning until the task is completed.
[0052] In step S13, the parameter server calculates the maximum norm of the parameters uploaded by each client in the cluster 1 and mean norm 2 ,like 1 and 2 If both exceed the predefined threshold, the current round is judged to meet the clustering condition, otherwise it does not meet the clustering condition.
[0053] In this embodiment, the specific clustering scheme is composed of one or more of four clustering algorithms: k-means, agglomerative clustering, spectral clustering and DBSCAN. It should be noted that the number of clustering rounds is dynamic and cannot be determined in advance, but if the client data is relatively fixed, the number of clustering rounds and the results are quite fixed. Therefore, the number of clustering rounds and the specific classification can be directly assumed, or the characteristics of the clustering algorithm can be combined, such as gradually clustering in the early hierarchical clustering to improve the clustering effect, and the clients in each cluster are relatively similar in the later stage, so k-means fast clustering is adopted.
[0054] S2. During the cluster federated learning task, a malicious client, namely an attacker, is deployed in the architecture; the attacker has an auxiliary dataset D a , D a contains data similar to that of a designated benign client, the victim.
[0055] In step S2, the auxiliary data set D a The data contains data that is 20%-50% similar to the victim, which can be generated through similar distribution or through model synthesis.
[0056] S3. The attacker eavesdrops on some of the victim’s model parameters according to the preset eavesdropping parameter ratio α W v ,Will W v Cluster model parameters sent by the parameter server W c Weighted to build an approximate victim model M v .
[0057] S4. Auxiliary Dataset D a The samples in are sequentially input into the approximate victim model M v In the example, a shadow dataset is constructed based on the output results. D s .
[0058] In step S4, the auxiliary data set D a The samples in are sequentially input into the approximate victim model M v After that, by calculating the cross entropy loss, samples whose cross entropy loss is less than the loss threshold are extracted to form the shadow dataset. D s .
[0059] In the present invention, the attacker first constructs a local model similar to the victim by using the partial parameters of the eavesdropped victim and the received global or cluster model parameters. Subsequently, the attacker uses the constructed approximate victim model and its own auxiliary data set to filter out data similar to the victim's local data. Finally, the attacker uses the filtered data similar to the victim to participate in the clustering process, thereby increasing the success rate of joining the victim's cluster.
[0060] Assume that the victim model is more consistent with homogeneous data, so the loss of homogeneous data is often smaller than that of heterogeneous data. D s The proportion must be higher than the initial auxiliary dataset D a However, not all homogeneous data losses are smaller than heterogeneous data losses, so the following dataset filtering algorithm is urgently needed. D s Homogeneous data is retained while more heterogeneous data is efficiently extracted.
[0061] The dataset filtering algorithm is mainly divided into the training phase and the filtering phase. The training phase is used to train the classifier, while the filtering phase uses the classifier to extract the shadow dataset. D s Filter out new datasets D f .
[0062] S5. Attacker builds k A shadow model is used to approximate the victim's model behavior, and the shadow dataset is D s Divide into k Group-to-group shadow training set D train and shadow test set D test , and then use the shadow training set and shadow test set to train the corresponding shadow model; each group of shadow training set D train and shadow test set D test The samples in are input into the corresponding shadow model to obtain the prediction metric to generate a prediction set D pred .
[0063] See also Figure 3 , Figure 3 The training phase of dataset filtering is described. First, based on the structural information and hyperparameter settings of the model published by the server, as well as the knowledge about the training algorithm, multiple shadow models are built to approximate the victim's model behavior.
[0064] In step S5, the attacker constructs multiple shadow models based on the model structure information, hyperparameter settings, and knowledge about the training algorithm sent by the parameter server.
[0065] The training set of each shadow model is an intersecting subset or a non-intersecting subset of the shadow training set; it should be noted that the shadow training set is a complete set, which is the one mentioned above. D s , in order to train multiple shadow models, we need to D s Suppose there are five shadow models. D s It can be divided into five shadow subsets evenly, or divided into five shadow subsets that overlap with each other.
[0066] The prediction set D pred It includes a prediction training set and a prediction test set; wherein the prediction training set corresponds to the output of the shadow training set, and the prediction test set corresponds to the output of the shadow test set.
[0067] By taking the training set of each shadow model D train and test set D test Samples in x Input into the corresponding shadow model to obtain prediction metrics, marked as "In" and "Out" respectively. Here, "In" means that the data sample is used to train the shadow model, and "Out" means the opposite. This embodiment processes the shadow dataset in this way. D s All the data of to obtain a new data set, namely the prediction set D pred .
[0068] S6. Using the prediction set D pred Train a meta-classifier to classify the shadow dataset D s Input into the victim model, obtain the predicted output information and input it into the meta-classifier for classification to obtain the screened data set D f .
[0069] It should be noted that the output information is the shadow dataset. D sAfter inputting into the victim model, the output result is a 62-dimensional vector, corresponding to the probability of predicting the 62 categories. The original classifier can be divided into two types, one is the training set and the other is the non-training set. The training set is the data used to train the victim model, which is the victim's own data, and the non-training set is not the data used for the victim. D s The data classified as training set by the meta-classifier is used to construct the dataset D f .
[0070] In step S6, the meta-classifier can learn the difference representation between the shadow training set and the shadow test set, thereby generalizing to the difference representation of the training set and the test set in the cluster federated learning scenario.
[0071] Figure 4 Depicts the filtering stage of dataset filtering. At this point, a classifier has been obtained and the initially filtered shadow dataset is then fed into the victim model to obtain the predicted output information. This information is then fed into the classifier to classify it. Data that is classified as “In” is retained as the filtered set. D f .
[0072] S7. Use of Dataset D f Update the attacker model and upload it to the parameter server to determine whether the attacker successfully enters the victim's cluster to verify the security of the clustering process.
[0073] Attacker uses dataset D f To train the model, the attacker uploads the parameters before the next attack to participate in the cluster federated learning process. The cluster federated learning process receives the parameters uploaded by the attacker and mistakenly assumes that the attacker and the victim are similar, and mistakenly classifies the attacker and the victim into the same group, that is, successfully clustered. At this time, there is a hidden danger in verifying the security of the clustering process.
[0074] This embodiment also applies the above method to carry out simulation and experiment.
[0075] A Experimental setup
[0076] EMNIST is an image classification dataset containing 814,255 images with 62 categories. The dataset is constructed by dividing the data in the extended MNIST into 3550 parts based on the source of digital / character images.
[0077] The FMNIST (Fashion-MNIST) dataset contains grayscale images of 10 categories. The training dataset contains 6000 samples per category, and the test dataset contains 1000 samples per category. The image is a 28×28 pixel matrix.
[0078] CIFAR-10 is a small dataset for recognizing common objects. It contains RGB color images of 10 categories. The size of each image is 32×32, and there are 6000 images for each category.
[0079] In the experiment of this embodiment, the following settings are made for all data sets. First, the data set is divided into two mutually exclusive subsets: client data set D c and attacker dataset D a A total of 20 benign clients were established C i ( i ∈0 to 19), evenly distributed among the 4 clusters. Each client is assigned a random subset of data D c i ( i ∈0 to 19), the data in the same cluster is independent and identically distributed, while the data between different clusters are not independent and identically distributed. In addition, malicious clients are introduced C 20 , which has an auxiliary dataset D c 20 , whose distribution is D c The overall distribution is similar.
[0080] The clustering process and attack results are shown in Figures 5 to 8 As shown in the figure, the training and clustering of clients using four different clustering algorithms are shown. The entire training process produces one to three clusters, and the clustering results are shown at the top of the figure. The colored lines represent the changes in similarity between different clients and the victim; the similarity of clients that are not in the same cluster as the victim is not shown in the visualization.
[0081] B. Performance of Intra-Cluster Attacks
[0082] This embodiment conducts 100 rounds of experiments on three different data sets for each of the four clustering algorithms (k-means, agglomerative clustering, spectral clustering, and DBSCAN). For agglomerative clustering, k-means, and spectral clustering, when the clustering conditions are met, binary clustering is achieved. For DBSCAN, this embodiment sets the parameter eps to 0.5 and the minimum number of samples to 5. In the experiments of this embodiment, the proportion of parameters obtained by eavesdropping is set to 50%, and the similarity between the victim data set and the attacker auxiliary data set is also set to 50%. The attack of this embodiment is performed in the early rounds of each clustering.
[0083] The evaluation criteria of this embodiment focuses on the success rate of the clustering algorithm and the accuracy of the attack model. The success rate refers to the number of times the attacker and the victim are placed in the same cluster after executing the algorithm. The accuracy rate measures how similar the data filtered by the algorithm is to the victim data. The higher the accuracy rate, the more similar the data is to the victim data, and the easier it is for the server to mistakenly group the attacker and the victim together, thereby achieving the goal of the algorithm.
[0084] Figures 5 to 8 The training process of cluster federated learning under four clustering algorithms is shown. The black vertical line indicates when the server executes the clustering algorithm on the client in that round. The top of the picture shows the change of client clustering in the three clusters. Each bracket represents a cluster, and the number in the bracket represents the client that is finally grouped into that cluster. The colored lines in the figure represent the change of similarity between different clients and victims in the same cluster. Clients that do not belong to the same cluster no longer appear in the figure subsequently. It can be seen from the figure that from the first intra-cluster attack to the end of the clustering stage, the gradient similarity of the attacker and the victim remains high. The results show that after the intra-cluster attack, the parameters uploaded by the attacker are more similar to the parameters of the victim than to the parameters of the isomorphic client. Therefore, the server is more likely to group the attacker and the victim into the same cluster.
[0085] Table 1 below shows the successful clustering rates of the attack method of this embodiment under various clustering algorithms. It is worth noting that on the FMNIST dataset, the attack success rate of the agglomerative clustering algorithm is as high as 90%, and exceeds 83% in other scenarios. This experiment assumes that half of the attacker's dataset is isomorphic to the victim's dataset, resulting in a baseline success rate of 50%.
[0086] Table 1: Intra-cluster attack success rate
[0087] ;
[0088] The core of the cluster attack algorithm is to identify the identity of the data, which is similar to the inference attack. Therefore, this embodiment compares the cluster attack algorithm with the existing SIF algorithm and MIA algorithm, and performs a total of three rounds of attacks during the clustering process. The attack accuracy is shown in Table 2. The results show that after the third attack, the attack accuracy of the present invention can reach 80%, and after the first attack, it can reach 70%. This is because as the clustering stage proceeds, the distribution of client data in each cluster becomes more similar. Therefore, the cluster model performs better and is more effective in fitting similar data, which makes it easier for the attack model of the present invention to determine member identities, resulting in higher accuracy.
[0089] Table 2: Intra-cluster attack accuracy
[0090] ;
[0091] C. Hyperparameter settings
[0092] Next, this example observes the robustness of the clustering algorithm by adjusting some hyperparameter settings.
[0093] (1) Monitoring parameter ratio α: In the real world, it is difficult for an attacker to monitor all parameters uploaded by the client during the communication process. Therefore, it is assumed that the attacker can only intercept a certain ratio of the uploaded parameters, represented by α%. This embodiment experiments with cluster attacks with a monitoring ratio α between 20% and 50%. The experimental results are shown in Figure 2. Fig. 9 , Fig.10 As shown. It can be observed that as α% increases, the attack success rate of the clustering algorithm increases. The reason behind this trend is that the higher the monitoring ratio, the more accurate the interception model. This leads to more detailed information leaked in the model output, making it easier for the attack model of the present invention to distinguish members. Therefore, this improves the similarity of the attacker's data set, thereby increasing the success rate of the attack.
[0094] (2) Similarity of auxiliary data set β: As the demands of various entities for data privacy and security increase, it becomes more difficult for attackers to collect a large amount of data similar to the victim. Therefore, this embodiment studies the impact on the clustering algorithm by changing the amount of data in the attacker's data set that is isomorphic to the victim's data set. Fig.11 , Fig.12 The performance impact of clustering algorithms when the proportion of homogeneous data increases from 20% to 50% is shown. The results show that the larger the proportion of homogeneous data, the more obvious the attack effect. Even with only 20% homogeneous data, it is easier to induce malicious clustering results.
[0095] In summary, this embodiment provides a means of intra-cluster attack against the clustering process of cluster federated learning, and uses four clustering algorithms to evaluate its effectiveness on three data sets, and compares it with the existing SIF method and MIA method. Experimental results show that the intra-cluster attack method of the present invention achieves high clustering success rate and inference success rate, and has a significant impact on the clustering process of cluster federated learning. This damage causes the server to produce incorrect clustering results. This embodiment further explores the factors that affect the effectiveness of intra-cluster attacks, especially the number of parameters eavesdropped by the attacker and the influence of the similarity between the attacker's auxiliary data set and the victim data set. The results show that the more the number of eavesdropped parameters, the higher the data similarity, and the more successful the attack. The present invention helps to develop safer cluster federated learning clustering and parameter aggregation methods.
[0096] Example 2
[0097] This embodiment provides a computer terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the security verification method for the cluster federated learning clustering process as described in Example 1 are implemented.
[0098] like Fig.13 As shown, the computer terminal provided in this embodiment includes: at least one processor 101, and a memory 102 connected to the at least one processor 101. The specific connection medium between the processor 101 and the memory 102 is not limited in this embodiment. Fig.13 In the example, the processor 101 and the memory 102 are connected via the bus 100. The bus 100 is connected to the memory 102 via the bus 100. Fig.13 The connections between other components are shown by thick lines, which are only schematic and not limiting. The bus 100 can be divided into an address bus, a data bus, a control bus, etc. For convenience of representation, Fig.13 In the figure, only one thick line is used, but it does not mean that there is only one bus or one type of bus. Alternatively, the processor 101 can also be called a controller, and there is no limitation on the name.
[0099] In this embodiment, the memory 102 stores instructions that can be executed by at least one processor 101 , and the at least one processor 101 can perform the aforementioned method by executing the instructions stored in the memory 102 .
[0100] Among them, the processor 101 is the control center of the device, and can use various interfaces and lines to connect the various parts of the entire control device. By running or executing instructions stored in the memory 102 and calling data stored in the memory 102, the various functions of the device and process data, the device can be monitored as a whole.
[0101] In one possible design, the processor 101 may include one or more processing units, and the processor 101 may integrate an application processor and a modem processor, wherein the application processor mainly processes an operating system, a user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the modem processor may not be integrated into the processor 101. In some embodiments, the processor 101 and the memory 102 may be implemented on the same chip, and in some embodiments, they may also be implemented separately on separate chips.
[0102] The processor 101 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component, and may implement or execute the methods, steps, and logic block diagrams disclosed in this embodiment. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the security verification method disclosed in Embodiment 1 may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor 101.
[0103] The memory 102 is a non-volatile computer-readable storage medium that can be used to store non-volatile software programs, non-volatile computer executable programs and modules. The memory 102 may include at least one type of storage medium, such as a flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (Random Access Memory, RAM), a static random access memory (Static Random Access Memory, SRAM), a programmable read-only memory (Programmable Read Only Memory, PROM), a read-only memory (Read Only Memory, ROM), an electrically erasable programmable read-only memory (Electrically Erasable Programmable Read-Only Memory, EEPROM), a magnetic memory, a disk, an optical disk, etc. The memory 102 is any other medium that can be used to carry or store a desired program code in the form of an instruction or data structure and can be accessed by a computer, but is not limited thereto. The memory 102 in this embodiment can also be a circuit or any other device that can realize a storage function, for storing program instructions and / or data.
[0104] By programming the processor 101, the code corresponding to the security verification method described in the above embodiment can be fixed into the chip, so that the chip can execute the code when running. Figure 1The steps of the security verification method of the embodiment shown are as follows: How to design and program the processor 101 is a technique known to those skilled in the art and will not be described in detail here.
[0105] Example 3
[0106] This embodiment provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the security verification method for the cluster federated learning clustering process as described in Example 1 are implemented.
[0107] The computer-readable storage medium may include flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the storage medium may be an internal storage unit of a computer device, such as a hard disk or memory of the computer device. In other embodiments, the storage medium may also be an external storage device of a computer device, such as a plug-in hard disk equipped on the computer device, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. Of course, the storage medium may also include both an internal storage unit of a computer device and an external storage device thereof. In this embodiment, the memory is generally used to store an operating system and various application software installed on the computer device, etc. In addition, the memory may also be used to temporarily store various types of data that have been output or are to be output.
[0108] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A security verification method for cluster federated learning clustering process, characterized in that: include: S1. Perform cluster federated learning tasks according to a specific clustering scheme. The architecture of this task consists of a parameter server and multiple benign clients. S2. During the execution of the task, a malicious client, namely an attacker, is deployed in the architecture; The attacker has an auxiliary dataset D a , D a contains data similar to that of a designated benign client, the victim; S3. The attacker eavesdrops on some of the victim’s model parameters W v ,Will W v Cluster model parameters sent by the parameter server W c Weighted to build an approximate victim model M v ; S4. Auxiliary Dataset D a The samples in are sequentially input into the approximate victim model M v In the example, a shadow dataset is constructed based on the output results. D s ; S5. Attacker builds k shadow model, the shadow dataset D s Divide into k Group-to-group shadow training set D train and shadow test set D test , and then use the shadow training set and shadow test set to train the corresponding shadow model; each group of shadow training set D train and shadow test set D test The samples in are input into the corresponding shadow model to obtain the prediction metric to generate a prediction set D pred ; S6. Using the prediction set D pred Train a meta-classifier to classify the shadow dataset D s Input into the victim model, obtain the predicted output information and input it into the meta-classifier for classification to obtain the screened data set D f ; S7. Use of Dataset D f Update the attacker model and upload it to the parameter server to determine whether the attacker successfully enters the victim's cluster to verify the security of the clustering process.
2. The security verification method for cluster federated learning clustering process according to claim 1 is characterized in that: In step S1, the execution of cluster federated learning tasks includes: S11. The parameter server sends the initialized model and hyperparameters to each benign client; S12. Each benign client uses the initialization model, hyperparameters, and its own private data to train client parameters, and then uploads the trained client parameters to the parameter server; S13. The parameter server determines whether the current round meets the clustering conditions. If so, it obtains clustering according to a specific clustering scheme; otherwise, it performs ordinary federated learning. If the clustering conditions cannot be met for consecutive preset rounds, the clustering phase ends, and the parameter server continues to perform ordinary federated learning until the task is completed.
3. The security verification method for cluster federated learning clustering process according to claim 2 is characterized in that: The specific clustering scheme is composed of one or more of four clustering algorithms: k-means, agglomerative clustering, spectral clustering and DBSCAN.
4. The security verification method for cluster federated learning clustering process according to claim 2 is characterized in that: In step S13, the parameter server calculates the maximum norm of the parameters uploaded by each client in the cluster 1 and mean norm 2. If 1 and 2 exceed the predefined threshold, then the current round is judged to meet the clustering condition, otherwise it does not meet the clustering condition.
5. The security verification method for cluster federated learning clustering process according to claim 1 is characterized in that: In step S5, the attacker constructs multiple shadow models based on the model structure information, hyperparameter settings, and knowledge about the training algorithm sent by the parameter server; The training set of each shadow model is either an intersecting subset or a non-intersecting subset of the shadow training set; The prediction set D pred It includes a prediction training set and a prediction test set; wherein the prediction training set corresponds to the output of the shadow training set, and the prediction test set corresponds to the output of the shadow test set.
6. The security verification method for cluster federated learning clustering process according to claim 1 is characterized in that: In step S6, the meta-classifier can learn the difference representation between the shadow training set and the shadow test set, thereby generalizing to the difference representation of the training set and the test set in the cluster federated learning scenario.
7. The security verification method for cluster federated learning clustering process according to claim 1 is characterized in that: In step S2, the auxiliary data set D a The malware contains data that is 20%-50% similar to the victim.
8. The security verification method for cluster federated learning clustering process according to claim 1 is characterized in that: In step S4, the auxiliary data set D a The samples in are sequentially input into the approximate victim model M v After that, by calculating the cross entropy loss, samples whose cross entropy loss is less than the loss threshold are extracted to form the shadow dataset. D s .
9. A computer terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the security verification method for the cluster federated learning clustering process as described in any one of claims 1 to 8.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by the processor, the steps of the security verification method for the cluster federated learning clustering process as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Federated learning protocol interaction security verification method and apparatus, and electronic equipment
CN114021188A
Clustering effect verification method in cluster federated learning, terminal and storage medium
CN117150255A