A Federated Learning Method for Spam Detection Model Based on Random Strategy

By performing similarity evaluation and clustering on client nodes, randomly selecting nodes to participate in training, and combining singular value decomposition and weight correction, the problems of server computing and network bandwidth limitations in federated learning are solved, and the efficiency and accuracy of model training are improved.

CN120151315BActive Publication Date: 2025-09-05BEIJING TOPSEC NETWORK SECURITY TECH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510404346.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-01
Publication Date
2025-09-05
Estimated Expiration
2045-04-01

AI Technical Summary

Technical Problem

In the federated learning training scenario for large-scale users, the computing processing performance and network bandwidth of the email security monitoring center server are limited, resulting in processing delays and communication delays. At the same time, the global model update gradient variance is large, which reduces the parameter convergence speed of model training.

Method used

By performing similarity evaluation and clustering on the local spam training samples of client nodes, randomly selecting nodes to participate in training, and combining singular value decomposition and weight correction, the global model parameters are generated to optimize resource utilization and gradient variance.

Benefits of technology

It reduces the processing delay and communication delay of the email security monitoring center server, improves the convergence speed of model parameters, reduces the model training time, and improves the accuracy and efficiency of model training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151315B_ABST
    Figure CN120151315B_ABST
Patent Text Reader

Abstract

The present application provides a federated learning method for a spam detection model based on a random strategy, which includes: evaluating the similarity of the feature matrices of samples in all nodes, clustering nodes based on the similarity, and randomly determining the nodes that can participate in training in each node based on a preset upper limit on the number of nodes to select clustering clusters; receiving and aggregating the local model parameters fed back by each node to obtain the global model parameters, and then determining whether to continue the cycle or output the final spam detection model based on the training termination condition. It can be seen that this method can reduce the consumption of computing, storage and network communication resources by limiting the number of participating nodes, thereby solving the problems of processing and communication delays. At the same time, the selection of classification clusters can also reduce the gradient variance of each round of global model updates, improve the parameter convergence speed of model training, and reduce the time required for model training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the fields of spam detection and federated learning technology, and more specifically, to a federated learning method, electronic device, readable storage medium, and computer program product for a spam detection model based on a random strategy. Background Art

[0002] Currently, in the federated learning training scenario of spam detection models for large-scale users to protect the privacy of user email data, due to the large number of model parameters, when all user nodes simultaneously participate in each round of federated learning iterative training, the computing processing performance and network bandwidth size of the email security monitoring center server will impose certain limitations on the efficiency of federated learning.

[0003] A common solution is to randomly select a small number of nodes to participate in each round of iterative training of federated learning, so as to reduce the computing, storage and network communication resource consumption of the email security monitoring center server, and alleviate the processing delay and communication delay problems caused by the limited computing storage performance and network bandwidth of the email security monitoring center server.

[0004] However, in practice, it is found that although the above solution can solve the above problems to a certain extent, it will lead to a large variance of the global model update gradient in each round, thereby reducing the parameter convergence speed of model training and increasing the time required for model training. Summary of the Invention

[0005] The purpose of this application is to provide a federated learning method for a spam detection model based on a random strategy, which is used to solve the problems of huge computing, storage and network communication resource consumption of the email security monitoring center server in the federated learning scenario, processing delay and communication delay caused by the limited computing processing performance and network bandwidth of the email security monitoring center server, and at the same time solve the problem that the global model update gradient variance in each round of training is large, thereby reducing the parameter convergence speed of model training and increasing the time required for model training.

[0006] In a first aspect, the present application provides a federated learning method for a spam detection model based on a random strategy, the method being applied to an email security monitoring center server, the method comprising:

[0007] Performing a similarity evaluation on a feature matrix of local spam training samples of all client nodes to obtain multiple similarity evaluation results; wherein the feature matrix is ​​generated based on email header features and email content features, the email header features including at least one of the sender's email domain name type, the number of recipients, and the attachment type; and the email content features including at least one of the email subject and keywords;

[0008] Cluster all client nodes based on multiple similarity evaluation results to obtain multiple node clusters;

[0009] Calculate the number of nodes that can participate in training in each node cluster based on the upper limit of the number of client nodes to be selected;

[0010] Randomly selecting client nodes for each node cluster based on the number of nodes to obtain multiple node selection clusters;

[0011] Send the latest global model parameters to all client nodes in the selected clusters, so that each client node can train its own local model parameters after the local model is updated.

[0012] Based on multiple nodes, the local model parameters uploaded by all client nodes in the cluster are selected and aggregated to generate global model parameters;

[0013] When it is determined based on the global model parameters and the number of training rounds that the spam detection model meets a preset training termination condition, the global model parameters are output as final parameters to obtain a trained spam detection model.

[0014] In the above implementation process, this method can perform similarity assessments on local spam training samples of client nodes, paving the way for subsequent reasonable node classification and reducing the variance of global model update gradients. Furthermore, client nodes are clustered based on the similarity assessment results, allowing similar data nodes to aggregate, which can reduce the variance of global model update gradients and facilitate parameter convergence. The method then calculates the number of nodes that can participate in training in each cluster based on the upper limit of the planned number of client nodes selected, rationally controlling the training scale and balancing resource utilization and training efficiency. Nodes are then randomly selected within each cluster, achieving a balance between randomness and scale control. Finally, the method can achieve local training and parameter upload through federated learning, thereby protecting user data privacy. Furthermore, the parameter aggregation method can correct the aggregation weights of each node cluster, reducing the estimation error of the global model update gradient, improving the convergence speed of model training, and reducing the time required for model training.

[0015] Furthermore, the similarity evaluation is performed on the feature matrices of the local spam training samples of all client nodes to obtain multiple similarity evaluation results, including:

[0016] Receiving the number of non-zero singular values ​​uploaded by the client node and determining a minimum number; wherein the client node obtains a left singular matrix, a singular value matrix, and a right singular matrix when performing singular value decomposition on a feature matrix of a local spam training sample, and the number of non-zero singular values ​​is obtained by the client node based on statistics of the singular value matrix;

[0017] Sending the minimum number value to all client nodes, so that the client nodes select multiple non-zero singular values ​​with the largest singular values ​​in the singular value matrix based on the minimum number value, and generate a pruned singular value pruned matrix and a left singular pruned matrix; wherein the number of the multiple non-zero singular values ​​is the same as the minimum number value;

[0018] Receive the left singular clipping matrix uploaded by the client node, and calculate the angle between the column vectors of any two left singular clipping matrices;

[0019] Selecting a plurality of minimum angles satisfying an orthogonality condition from all calculated angles; wherein the number of the plurality of minimum angles is the same as the minimum value of the number;

[0020] Based on multiple minimum angles, the similarity between the feature matrices of local spam training samples of any two client nodes is calculated to obtain a similarity evaluation result.

[0021] In the above implementation process, this method can accurately calculate the similarity of local spam training samples of client nodes through singular value decomposition, clipping and angle calculation, thereby laying the foundation for subsequent reasonable clustering, node selection, reducing computing and network resource consumption, improving federated learning efficiency and reducing model training time.

[0022] Furthermore, all client nodes are clustered based on multiple similarity evaluation results to obtain multiple node clusters, including:

[0023] Converting the multiple similarity evaluation results into multiple similarity distances;

[0024] Based on the multiple similarity distances, K-means algorithm and silhouette coefficient method, the client nodes are clustered to obtain multiple node clusters.

[0025] In the above implementation process, this method can cluster client nodes more accurately and reasonably, thereby improving the efficiency of federated learning training and reducing model training time.

[0026] Furthermore, the method further comprises:

[0027] Obtaining computing processing performance parameters, network bandwidth parameters, and number of neural network model parameters of the email security monitoring center server;

[0028] Based on the computing processing performance parameter, the network bandwidth parameter and the number of neural network model parameters, an upper limit on the number of client nodes planned to be selected is determined.

[0029] In the above implementation process, this method can determine the upper limit of the planned number of client nodes by obtaining the computing processing performance, network bandwidth and number of neural network model parameters of the email security monitoring center server, which can effectively avoid processing and communication delays caused by resource limitations, thereby ensuring the efficiency of federated learning training.

[0030] Furthermore, the calculation of the number of nodes that can participate in training in each node cluster based on the upper limit of the number of client nodes selected in the plan includes:

[0031] Based on the preset constraint calculation formula, the upper limit of the number of client nodes planned to be selected, the number of clusters in multiple node clusters, the number of nodes in each node cluster and the total number of client nodes, the number of nodes that can participate in training in each node cluster is calculated.

[0032] In the above implementation process, this method can calculate the number of nodes that can participate in training for each cluster based on the preset constraint calculation formula, taking into account the upper limit of the number of nodes, the number of clusters, the number of nodes in each cluster and the total number of nodes, thereby achieving a reasonable allocation of resources, which is beneficial to improving the efficiency of federated learning training and ensuring the speed of model convergence.

[0033] Furthermore, the method of selecting local model parameters uploaded by all client nodes in the cluster based on multiple nodes and aggregating them to generate global model parameters includes:

[0034] Calculate the first total sample size of local spam training samples in each node cluster;

[0035] Calculate the second total sample size of local spam training samples in the cluster selected by each node;

[0036] Calculating a weight correction factor for each node cluster based on the first total sample size and the second total sample size;

[0037] Based on the local sample size of the local spam training sample of the client node, weighted calculation is performed on the local model parameters uploaded by the client node to obtain the local model weighted parameters;

[0038] Summing all local model weighted parameters in the target node selected cluster to obtain a local model weighted parameter sum; wherein the target node selected cluster is any one of the multiple node selected clusters;

[0039] Calculate the product of the weight correction factor of the target node cluster and the sum of the weighted parameters of the local model to obtain the cluster model parameters of the target node cluster;

[0040] The global model parameters are obtained by summing up all cluster model parameters and dividing them by the sum of all first total sample sizes.

[0041] In the above implementation process, this method can fully consider the differences in sample distribution and improve the accuracy of global model parameter updates and the efficiency of federated learning training through weight correction.

[0042] A second aspect of the present application provides a spam detection method, the spam detection method comprising:

[0043] Get the email to be tested;

[0044] The email to be detected is input into a spam detection model so that the spam detection model outputs a detection result of whether the email to be detected is spam; wherein the spam detection model is trained by the federated learning method of the spam detection model based on random strategy according to any one of the first aspects of this application.

[0045] A third aspect of the present application provides a federated learning device for a spam detection model based on a random strategy, wherein the federated learning device for the spam detection model based on a random strategy is a mail security monitoring center server, and the federated learning device for the spam detection model based on a random strategy includes:

[0046] A similarity evaluation unit is configured to perform similarity evaluation on a feature matrix of local spam training samples of all client nodes to obtain a plurality of similarity evaluation results; wherein the feature matrix is ​​generated based on email header features and email content features, the email header features including at least one of the sender's email domain name type, the number of recipients, and the type of attachments, and the email content features including at least one of the email subject and keywords;

[0047] A clustering unit, configured to cluster all client nodes based on multiple similarity evaluation results to obtain multiple node clusters;

[0048] A calculation unit, configured to calculate the number of nodes in each node cluster that can participate in training based on the upper limit of the number of client nodes to be selected;

[0049] a random selection unit, configured to randomly select client nodes from each node cluster based on the number of nodes, to obtain a plurality of node selection clusters;

[0050] The sending unit is used to send the latest global model parameters to all client nodes in the cluster selected by multiple nodes, so that each client node can train and obtain the updated local model parameters of the local model;

[0051] Aggregation unit, used to select local model parameters uploaded by all client nodes in the cluster based on multiple nodes, and aggregate them to generate global model parameters;

[0052] The output unit is configured to output the global model parameters as final parameters to obtain a trained spam detection model when it is determined based on the global model parameters and the number of training rounds that the spam detection model meets a preset training termination condition.

[0053] Furthermore, the similarity evaluation unit includes:

[0054] a receiving subunit, configured to receive the number of non-zero singular values ​​uploaded by the client node and determine a minimum number; wherein the client node obtains a left singular matrix, a singular value matrix, and a right singular matrix when performing singular value decomposition on the feature matrix of the local spam training sample, and the number of non-zero singular values ​​is obtained by the client node based on the statistics of the singular value matrix;

[0055] a sending subunit, configured to send the minimum number value to all client nodes, so that the client nodes select multiple non-zero singular values ​​with the largest singular values ​​in the singular value matrix based on the minimum number value, and generate a pruned singular value pruned matrix and a left singular pruned matrix; wherein the number of the multiple non-zero singular values ​​is the same as the minimum number value;

[0056] The receiving subunit is further configured to receive a left singular clipping matrix uploaded by a client node, and calculate an angle between column vectors of any two left singular clipping matrices based on the two columns;

[0057] A selection subunit is used to select a plurality of minimum angles that meet the orthogonality condition from all calculated angles; wherein the number of the plurality of minimum angles is the same as the minimum value of the number;

[0058] The calculation subunit is used to calculate the similarity between the feature matrices of local spam training samples of any two client nodes based on multiple minimum angles to obtain a similarity evaluation result.

[0059] Furthermore, the clustering unit includes:

[0060] a conversion subunit, configured to convert the plurality of similarity evaluation results into a plurality of similarity distances;

[0061] The clustering subunit is used to cluster the client nodes based on the multiple similarity distances, K-means algorithm and silhouette coefficient method to obtain multiple node clusters.

[0062] Furthermore, the federated learning device of the random strategy-based spam detection model further includes:

[0063] An acquisition unit, configured to acquire computing processing performance parameters, network bandwidth parameters, and number of neural network model parameters of the email security monitoring center server;

[0064] A determination unit is used to determine an upper limit on the number of client nodes planned to be selected based on the computing processing performance parameter, the network bandwidth parameter and the number of neural network model parameters.

[0065] Furthermore, the calculation unit is specifically used to calculate based on a preset constraint calculation formula, the upper limit of the number of client nodes planned to be selected, the number of clusters of multiple node clusters, the number of nodes in each node cluster and the total number of client nodes, to obtain the number of nodes that can participate in training in each node cluster.

[0066] Furthermore, the aggregation unit is specifically used to calculate a first total sample size of local spam training samples in each node cluster; calculate a second total sample size of local spam training samples in each node selected cluster; calculate a weight correction factor for each node cluster based on the first total sample size and the second total sample size; perform weighted calculation on the local model parameters uploaded by the client node based on the local sample size of the local spam training samples of the client node to obtain local model weighted parameters; sum all local model weighted parameters in the target node selected cluster to obtain a sum of local model weighted parameters; wherein the target node selected cluster is any one of multiple node selected clusters; calculate the product of the weight correction factor of the target node cluster and the sum of the local model weighted parameters to obtain the cluster cluster model parameters of the target node cluster; sum all cluster cluster model parameters and divide them by the sum of all first total sample sizes to obtain global model parameters.

[0067] In a fourth aspect, the present application provides an electronic device, comprising a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the federated learning method of the spam detection model based on a random strategy as described in any one of the first aspects of the present application.

[0068] In a fifth aspect, the present application provides a computer-readable storage medium, wherein the computer program instructions are stored in the computer-readable storage medium. When the computer program instructions are read and executed by a processor, the federated learning method of the spam detection model based on random strategy described in any one of the first aspects of the present application is executed.

[0069] A fifth aspect of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it executes the federated learning method of the spam detection model based on random strategy described in any one of the first aspects of the present application.

[0070] In a sixth aspect, the present application provides a computer program product, comprising a computer program. When the computer program is executed by a processor, the computer program executes the federated learning method of the spam detection model based on random strategy described in any one of the first aspects of the present application.

[0071] The beneficial effect of this application is that the method of randomly selecting client nodes to participate in the training of the federated learning model can reduce the computing processing performance of the email security monitoring center server in the federated learning scenario and the processing delay and communication delay caused by the limited network bandwidth.

[0072] Randomly selecting client nodes according to clusters based on client node clustering can reduce the variance of global model parameter updates in each round, thereby improving the convergence speed of model parameters in federated learning and reducing the model training time of federated learning.

[0073] When aggregating to generate global model parameters, the aggregation weights of the local model parameters of each client node are corrected, which can reduce the deviation between the updated values ​​of the global model parameters when randomly selecting client nodes to participate in model training and the updated values ​​of the global model parameters when all client nodes participate in model training, thereby ensuring that randomly selecting client nodes will not significantly change the convergence direction of the model, and improving the correctness and accuracy of the model training results. BRIEF DESCRIPTION OF THE DRAWINGS

[0074] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0075] Figure 1 A flowchart of a federated learning method for a spam detection model based on a random strategy provided in an embodiment of the present application;

[0076] Figure 2 A flowchart of another federated learning method for a spam detection model based on a random strategy provided in an embodiment of the present application;

[0077] Figure 3 A schematic diagram of the technical implementation process of a federated learning method for a spam detection model based on a random strategy provided in an embodiment of the present application;

[0078] Figure 4 A flowchart of a spam detection method provided in an embodiment of the present application;

[0079] Figure 5A schematic diagram of the structure of a federated learning device for a spam detection model based on a random strategy provided in an embodiment of the present application. DETAILED DESCRIPTION

[0080] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.

[0081] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0082] Example 1

[0083] Please see Figure 1 , Figure 1 The present embodiment provides a flow chart of a federated learning method for a spam detection model based on a random strategy. The federated learning method for a spam detection model based on a random strategy is applied to an email security monitoring center server and includes:

[0084] S101. Perform similarity evaluation on the feature matrices of local spam training samples of all client nodes to obtain multiple similarity evaluation results; wherein the feature matrix is ​​generated based on email header features and email content features, the email header features include at least one of the sender's email domain name type, the number of recipients, and the attachment type, and the email content features include at least one of the email subject and keywords.

[0085] In this embodiment, each client node can extract information such as the sender's email domain name type (such as corporate email domain name, personal email domain name, government institution email domain name, etc.), the number of recipients, and the attachment type (executable file, document, compressed package, image and video, no attachment, etc.) from the email header of the local spam training sample, and use it as the email header feature.

[0086] At the same time, each client node can also extract information related to the email content, such as email subject and keywords, from the email body of the local spam training sample and use it as the email content feature.

[0087] On this basis, the method can generate a feature matrix of local spam training samples of each client node according to the above-mentioned email header features and email content features.

[0088] S102: Cluster all client nodes based on multiple similarity evaluation results to obtain multiple node clusters.

[0089] S103: Calculate the number of nodes that can participate in training in each node cluster based on the upper limit of the number of client nodes that are planned to be selected.

[0090] S104: Randomly select client nodes for each node cluster based on the number of nodes to obtain multiple node selection clusters.

[0091] S105 , issuing the latest global model parameters to all client nodes in the selected cluster of multiple nodes, so that each client node can train and obtain the updated local model parameters of the local model.

[0092] S106 , selecting local model parameters uploaded by all client nodes in the cluster based on multiple nodes, and aggregating them to generate global model parameters.

[0093] S107: When it is determined based on the global model parameters and the number of training rounds that the spam detection model meets the preset training termination condition, the global model parameters are output as final parameters to obtain a trained spam detection model.

[0094] In this embodiment, the execution subject of the method may be a computing device such as a computer or a server, and this is not limited in this embodiment.

[0095] It can be seen that the implementation of the federated learning method of the random strategy-based spam detection model described in this embodiment can solve the processing delay and communication delay problems caused by the limited computing performance and network bandwidth of the email security monitoring center server in the federated learning scenario, and at the same time solve the problem that the global model update gradient variance in each round of training is large, thereby reducing the parameter convergence speed of model training and increasing the model training time.

[0096] Example 2

[0097] Please see Figure 2 , Figure 2 The present embodiment provides a flow chart of a federated learning method for a spam detection model based on a random strategy. The federated learning method for a spam detection model based on a random strategy is applied to an email security monitoring center server and includes:

[0098] S201. Receive the number of non-zero singular values ​​uploaded by the client node and determine a minimum number; wherein, when the client node performs singular value decomposition on the feature matrix of the local spam training sample, it obtains a left singular matrix, a singular value matrix, and a right singular matrix, and the number of non-zero singular values ​​is obtained by the client node based on statistics of the singular value matrix.

[0099] For example, the total number of client nodes (client hosts) with spam detection models is Nc = 1000000. The feature matrix of all local spam training samples of the i-th client host , X i Each component contains a feature column vector of a spam training sample. The dimension of the feature column vector of each spam training sample is M=500 (each column vector represents the header feature, content feature or part of the above features of a spam email). i Represents the total number of local spam training samples of the i-th client host. Where 1≤i≤N c , i, N i Is a positive integer.

[0100] At this time, the i-th client host pair feature matrix X i Perform singular value decomposition, that is , It's X i The left singular matrix of It's X i The singular value matrix of It's X i The client host will ∑ i The above singular values ​​are arranged from large to small, and the number of non-zero singular values ​​(ie, the number of non-zero singular values ​​uploaded by the client node in step S201) is counted and uploaded to the email security monitoring center server S.

[0101] S202. Send the minimum value to all client nodes, so that the client nodes select multiple non-zero singular values ​​with the largest singular values ​​in the singular value matrix based on the minimum value, and generate a pruned singular value pruned matrix and a left singular pruned matrix; wherein the number of the multiple non-zero singular values ​​is the same as the minimum value.

[0102] For example, the email security monitoring center server S receives N c The number of non-zero singular values ​​uploaded by the client hosts is counted, the minimum value among the above numbers is calculated, and the minimum value is sent to all client hosts. This minimum value is denoted as P, where P is a positive integer; in this example, P = 100.

[0103] S203: Receive the left singular clipping matrix uploaded by the client node, and calculate the angle between the column vectors of any two left singular clipping matrices based on the two columns.

[0104] For example, after the i-th client host receives the value P sent by the email security monitoring center server S, it selects ∑ i The largest P non-zero singular values ​​in generate the pruned singular value matrix , select U iThe column vectors corresponding to the P non-zero singular values ​​in the above generate the pruned left singular matrix Each client host will have its own Upload to the email security monitoring center server S.

[0105] Here, the column vector of a matrix refers to the vector composed of all the values ​​in a column of the matrix, and the column vector corresponding to the singular value in the left singular matrix refers to the column vector whose column number in the left singular matrix is ​​equal to the row number of the singular value in the singular value matrix.

[0106] Furthermore, the email security monitoring center server S receives N c For any two pruned left singular matrices, the angle between any two column vectors belonging to the two matrices is calculated according to the following formula, where

[0107] .

[0108] S204. Select multiple minimum angles that meet the orthogonality condition from all calculated angles; wherein the number of the multiple minimum angles is the same as the minimum value.

[0109] For example, the method continues to select from all the above angles the angle that satisfies the following orthogonality condition: The column vector and The P smallest angles between the column vectors 、 … , 1≤j≤N c , j≠i, j is a positive integer.

[0110] Orthogonality condition: When p≥2, Must be with ,…, are all orthogonal, Must be with ,…, All orthogonal.

[0111] in, express is a matrix Column vector of , express is a matrix Column vector of , Represents a vector With vector The inner product of Represents the modulus of vector x, 1≤p≤P, where p is a positive integer.

[0112] S205 : Based on the multiple minimum angles, calculate the similarity between the feature matrices of the local spam training samples of any two client nodes to obtain a similarity evaluation result.

[0113] For example, the email security monitoring center server S calculates the similarity of spam training data of any two client hosts according to the following formula:

[0114] ;

[0115] Among them, The client host and The similarity of spam training data of client hosts is recorded as S(i, j).

[0116] S206: Convert the multiple similarity evaluation results into multiple similarity distances.

[0117] S207 , clustering the client nodes based on multiple similarity distances, the K-means algorithm, and the silhouette coefficient method to obtain multiple node clusters.

[0118] In this embodiment, the method can traverse different numbers of cluster centers based on multiple similarity distances, K-means algorithm and silhouette coefficient method, and determine the optimal number of cluster centers when the clustering result is the best. In this case, the optimal clustering result is the multiple node clusters finally obtained in step S207.

[0119] In this embodiment, the method may also re-cluster the client nodes based on multiple similarity distances, the K-means algorithm and the optimal number of cluster centers, thereby obtaining multiple node clusters corresponding to the optimal clustering results.

[0120] For example, the email security monitoring center server S uses PS (i, j) to measure the distance between the i-th client host and the j-th client host, and uses the K-means algorithm to calculate the distance between N c client hosts to cluster (number of cluster centers Determined by the silhouette coefficient method), the final clustering is = 10 sets of client hosts .in, Is a positive integer.

[0121] S208. Obtain computing processing performance parameters, network bandwidth parameters, and number of neural network model parameters of the email security monitoring center server.

[0122] S209. Based on the computing processing performance parameters, the network bandwidth parameters and the number of neural network model parameters, determine the upper limit of the number of client nodes to be selected.

[0123] For example, the email security monitoring center server S initializes the parameters of the neural network model (such as Bert, LSTM, RNN, etc.) used, and the initialization parameters are w 0 The email security monitoring center server S determines the upper limit N of the number of client hosts to be randomly selected for model training based on its own computing performance (such as the system's million instructions per second MIPS), network bandwidth (such as 500Mbps bandwidth), and the number of parameters of the above-mentioned neural network model (such as the number of Bert-Base parameters is 110 million). T =100000.

[0124] S210. Calculate the number of nodes that can participate in training in each node cluster based on a preset constraint calculation formula, the upper limit of the number of client nodes to be selected, the number of clusters in multiple node clusters, the number of nodes in each node cluster, and the total number of client nodes.

[0125] For example, the email security monitoring center server S can solve and determine whether the constraint conditions are met. The maximum value of the variable x * , and calculate and N s =98500.

[0126] in, Indicates the A set of client nodes, Representing a collection The cardinality, Is a temporary variable, x * is the maximum value of x, Indicates the total number of client nodes actually randomly selected to participate in model training, Represents a collection The number of client nodes actually randomly selected to participate in model training, Express The operation of rounding up, , , , Z + Represents the set of positive integers.

[0127] S211 : Randomly select client nodes for each node cluster based on the number of nodes to obtain multiple node selection clusters.

[0128] For example, in the rth round of model training, the email security monitoring center server S Randomly select client hosts, the above The serial numbers of the client hosts constitute a set The email security monitoring center server S selects N s client hosts participate in this round of model training.

[0129] S212: Send the latest global model parameters to all client nodes in the selected cluster of multiple nodes, so that each client node can train and obtain the updated local model parameters of the local model.

[0130] For example, the email security monitoring center server S will set the parameter w r-1 Sent to the client host selected in this round. Among them, 1≤r≤R end , R end is the training round threshold constant, r, R end Is a positive integer.

[0131] Then, in the rth round of model training, if the i-th client host is a client host randomly selected by the email security monitoring center server S in this round, the client host will receive the w r-1 The client host updates the local model parameter value to w r-1 , the gradient descent algorithm is used to train and update the parameters of the local model on the local spam training sample set. The updated local model parameters are The client host will Upload to the email security monitoring center server S.

[0132] S213 , selecting local model parameters uploaded by all client nodes in the cluster based on multiple nodes, and aggregating them to generate global model parameters.

[0133] As an optional implementation, the local model parameters uploaded by all client nodes in the cluster are selected based on multiple nodes, and the global model parameters are aggregated and generated, including:

[0134] Calculate the first total sample size of local spam training samples in each node cluster;

[0135] Calculate the second total sample size of local spam training samples in the cluster selected by each node;

[0136] Calculating a weight correction factor for each node cluster based on the first total sample size and the second total sample size;

[0137] Based on the local sample size of the local spam training sample of the client node, weighted calculation is performed on the local model parameters uploaded by the client node to obtain the local model weighted parameters;

[0138] Summing all local model weighted parameters in the target node selected cluster to obtain a local model weighted parameter sum; wherein the target node selected cluster is any one of the multiple node selected clusters;

[0139] Calculate the product of the weight correction factor of the target node cluster and the sum of the weighted parameters of the local model to obtain the cluster model parameters of the target node cluster;

[0140] The global model parameters are obtained by summing up all cluster model parameters and dividing them by the sum of all first total sample sizes.

[0141] For example, in the rth round of model training, the email security monitoring center server S receives N s The updated parameters of the local models uploaded by the client host are aggregated to generate the global model parameters w r .

[0142] ;

[0143] in, Representing a collection The sum of the number of samples in the local training sample sets of all client nodes, is the number of samples in the local training sample set of the i-th client node.

[0144] S214. When it is determined based on the global model parameters and the number of training rounds that the spam detection model meets the preset training termination condition, the global model parameters are output as final parameters to obtain a trained spam detection model.

[0145] For example, if w r Convergence or r ≥ R end =200, the email security monitoring center server S terminates all model training processes and sets w r As the final parameter output of the global model; otherwise, let r=r+1, the email security monitoring center server S starts the next round of model training cycle, and repeats the above steps S211~S213.

[0146] Please see Figure 3 , Figure 3 The following is a technical implementation diagram of a federated learning method for a spam detection model based on a random strategy. Figure 3 The process shown is Figure 1 and Figure 2 The processes in have corresponding correspondence.

[0147] In this embodiment, the execution subject of the method may be a computing device such as a computer or a server, and this is not limited in this embodiment.

[0148] It can be seen that the implementation of the federated learning method of the random strategy-based spam detection model described in this embodiment can solve the processing delay and communication delay problems caused by the limited computing performance and network bandwidth of the email security monitoring center server in the federated learning scenario, and at the same time solve the problem that the global model update gradient variance in each round of training is large, thereby reducing the parameter convergence speed of model training and increasing the time required for model training.

[0149] Example 3

[0150] Please see Figure 4 , Figure 4 This is a flow chart of a spam detection method provided in this embodiment. The method includes:

[0151] S301: Obtain the email to be detected.

[0152] S302: Input the email to be detected into a spam detection model, so that the spam detection model outputs a detection result of whether the email to be detected is spam.

[0153] In this embodiment, the spam detection model is trained by the federated learning method of the spam detection model based on the random strategy in Example 1 or Example 2 of the present application.

[0154] In this embodiment, the execution subject of the method may be a computing device such as a computer or a server, and this is not limited in this embodiment.

[0155] It can be seen that the implementation of the federated learning method of the random strategy-based spam detection model described in this embodiment can solve the processing delay and communication delay problems caused by the limited computing performance and network bandwidth of the email security monitoring center server in the federated learning scenario, and at the same time solve the problem that the global model update gradient variance in each round of training is large, thereby reducing the parameter convergence speed of model training and increasing the time required for model training.

[0156] Example 4

[0157] Please see Figure 5 , Figure 5 This is a structural diagram of a federated learning device for a spam detection model based on a random strategy provided in this embodiment. The federated learning device for a spam detection model based on a random strategy may be a mail security monitoring center server. Figure 5 As shown, the federated learning device of the spam detection model based on random strategy includes:

[0158] Similarity evaluation unit 410 is configured to perform similarity evaluation on a feature matrix of local spam training samples of all client nodes to obtain multiple similarity evaluation results; wherein the feature matrix is ​​generated based on email header features and email content features, wherein the email header features include at least one of the sender's email domain name type, the number of recipients, and the type of attachments; and the email content features include at least one of the email subject and keywords.

[0159] A clustering unit 420 is configured to cluster all client nodes based on multiple similarity evaluation results to obtain multiple node clusters;

[0160] A calculation unit 430 is configured to calculate the number of nodes in each node cluster that can participate in training based on the upper limit of the number of client nodes to be selected;

[0161] A random selection unit 440 is configured to randomly select client nodes from each node cluster based on the number of nodes to obtain a plurality of node selection clusters;

[0162] The sending unit 450 is used to send the latest global model parameters to all client nodes in the cluster selected by the multiple nodes, so that each client node can train and obtain the updated local model parameters of the local model;

[0163] Aggregation unit 460, configured to select local model parameters uploaded by all client nodes in the cluster based on multiple nodes, and aggregate them to generate global model parameters;

[0164] The output unit 470 is configured to output the global model parameters as final parameters to obtain a trained spam detection model when it is determined based on the global model parameters and the number of training rounds that the spam detection model meets a preset training termination condition.

[0165] As an optional implementation, the similarity evaluation unit 410 includes:

[0166] The receiving subunit 411 is configured to receive the number of non-zero singular values ​​uploaded by the client node and determine a minimum number. When the client node performs singular value decomposition on the feature matrix of the local spam training sample, it obtains a left singular matrix, a singular value matrix, and a right singular matrix. The number of non-zero singular values ​​is obtained by the client node based on statistics of the singular value matrix.

[0167] The sending sub-unit 412 is configured to send the minimum value to all client nodes, so that the client nodes select multiple non-zero singular values ​​with the largest singular value in the singular value matrix based on the minimum value, and generate a pruned singular value pruned matrix and a left singular pruned matrix; wherein the number of the multiple non-zero singular values ​​is the same as the minimum value;

[0168] The receiving subunit 411 is further configured to receive a left singular clipping matrix uploaded by a client node, and calculate an angle between column vectors of any two left singular clipping matrices based on the two columns;

[0169] The selection subunit 413 is used to select a plurality of minimum angles that meet the orthogonality condition from all calculated angles; wherein the number of the plurality of minimum angles is the same as the minimum value;

[0170] The calculation subunit 414 is configured to calculate the similarity between the feature matrices of the local spam training samples of any two client nodes based on the multiple minimum angles to obtain a similarity evaluation result.

[0171] As an optional implementation, the clustering unit 420 includes:

[0172] A conversion subunit 421 is configured to convert the multiple similarity evaluation results into multiple similarity distances;

[0173] The clustering subunit 422 is configured to cluster the client nodes based on multiple similarity distances, the K-means algorithm, and the silhouette coefficient method to obtain multiple node clusters.

[0174] As an optional implementation, the federated learning device of the spam detection model based on the random strategy further includes:

[0175] An acquisition unit 480 is configured to acquire computing processing performance parameters, network bandwidth parameters, and number of neural network model parameters of the email security monitoring center server;

[0176] The determination unit 490 is used to determine the upper limit of the number of client nodes to be selected based on the computing processing performance parameters, the network bandwidth parameters and the number of neural network model parameters.

[0177] As an optional implementation, the calculation unit 430 is specifically used to calculate based on a preset constraint calculation formula, the upper limit of the number of client nodes planned to be selected, the number of clusters of multiple node clusters, the number of nodes in each node cluster and the total number of client nodes, to obtain the number of nodes that can participate in training in each node cluster.

[0178] As an optional implementation, the aggregation unit 460 is specifically used to calculate a first total sample size of local spam training samples in each node cluster; calculate a second total sample size of local spam training samples in each node selected cluster; calculate a weight correction factor for each node cluster based on the first total sample size and the second total sample size; perform weighted calculation on the local model parameters uploaded by the client node based on the local sample size of the local spam training samples of the client node to obtain local model weighted parameters; sum all local model weighted parameters in the target node selected cluster to obtain a sum of local model weighted parameters; wherein the target node selected cluster is any one of multiple node selected clusters; calculate the product of the weight correction factor of the target node cluster and the sum of the local model weighted parameters to obtain the cluster cluster model parameters of the target node cluster; sum all cluster cluster model parameters and divide them by the sum of all first total sample sizes to obtain global model parameters.

[0179] In this embodiment, the explanation of the federated learning device of the spam detection model based on the random strategy can refer to the description in Example 1 or Example 2, and will not be repeated in this embodiment.

[0180] It can be seen that the federated learning device that implements the random strategy-based spam detection model described in this embodiment can solve the processing delay and communication delay problems caused by the limited computing processing performance and network bandwidth of the email security monitoring center server in the federated learning scenario, and at the same time solve the problem that the global model update gradient variance in each round of training is large, thereby reducing the parameter convergence speed of model training and increasing the time required for model training.

[0181] An embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the federated learning method of the spam detection model based on the random strategy in Example 1 or Example 2 of the present application.

[0182] An embodiment of the present application provides a computer-readable storage medium storing computer program instructions. When the computer program instructions are read and executed by a processor, the federated learning method of the spam detection model based on random strategy in Example 1 or Example 2 of the present application is executed.

[0183] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it executes the federated learning method of the spam detection model based on random strategy in Example 1 or Example 2 of the present application.

[0184] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0185] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0186] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks or optical disks.

[0187] The foregoing is merely an embodiment of the present application and is not intended to limit the scope of protection of the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.

[0188] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

[0189] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.

Claims

1. A federated learning method for spam detection model based on random strategy, characterized by: The method is applied to a mail security monitoring center server, and the method includes: Performing a similarity evaluation on a feature matrix of local spam training samples of all client nodes to obtain multiple similarity evaluation results; wherein the feature matrix is ​​generated based on email header features and email content features, the email header features including at least one of the sender's email domain name type, the number of recipients, and the attachment type; and the email content features including at least one of the email subject and keywords; Cluster all client nodes based on multiple similarity evaluation results to obtain multiple node clusters; Calculate the number of nodes that can participate in training in each node cluster based on the upper limit of the number of client nodes to be selected; Randomly selecting client nodes for each node cluster based on the number of nodes to obtain multiple node selection clusters; Send the latest global model parameters to all client nodes in the selected clusters, so that each client node can train its own local model parameters after the local model is updated. Based on multiple nodes, the local model parameters uploaded by all client nodes in the cluster are selected and aggregated to generate global model parameters; When it is determined based on the global model parameters and the number of training rounds that the spam detection model meets a preset training termination condition, the global model parameters are output as final parameters to obtain a trained spam detection model; The similarity evaluation is performed on the feature matrices of the local spam training samples of all client nodes to obtain multiple similarity evaluation results, including: Receiving the number of non-zero singular values ​​uploaded by the client node and determining a minimum number; wherein the client node obtains a left singular matrix, a singular value matrix, and a right singular matrix when performing singular value decomposition on a feature matrix of a local spam training sample, and the number of non-zero singular values ​​is obtained by the client node based on statistics of the singular value matrix; Sending the minimum number value to all client nodes, so that the client nodes select multiple non-zero singular values ​​with the largest singular values ​​in the singular value matrix based on the minimum number value, and generate a pruned singular value pruned matrix and a left singular pruned matrix; wherein the number of the multiple non-zero singular values ​​is the same as the minimum number value; Receive the left singular clipping matrix uploaded by the client node, and calculate the angle between the column vectors of any two left singular clipping matrices; Selecting a plurality of minimum angles satisfying an orthogonality condition from all calculated angles; wherein the number of the plurality of minimum angles is the same as the minimum value of the number; Based on multiple minimum angles, the similarity between the feature matrices of local spam training samples of any two client nodes is calculated to obtain a similarity evaluation result.

2. The federated learning method for the spam detection model based on random strategy according to claim 1, characterized in that: The method clusters all client nodes based on multiple similarity evaluation results to obtain multiple node clusters, including: Converting the multiple similarity evaluation results into multiple similarity distances; Based on the multiple similarity distances, K-means algorithm and silhouette coefficient method, the client nodes are clustered to obtain multiple node clusters.

3. The federated learning method for the spam detection model based on random strategy according to claim 1, characterized in that: The method further comprises: Obtaining computing processing performance parameters, network bandwidth parameters, and number of neural network model parameters of the email security monitoring center server; Based on the computing processing performance parameter, the network bandwidth parameter and the number of neural network model parameters, an upper limit on the number of client nodes planned to be selected is determined.

4. The federated learning method for a spam detection model based on a random strategy according to claim 1, characterized in that: The calculation of the number of nodes that can participate in training in each node cluster based on the upper limit of the number of client nodes selected in the plan includes: Based on the preset constraint calculation formula, the upper limit of the number of client nodes planned to be selected, the number of clusters in multiple node clusters, the number of nodes in each node cluster and the total number of client nodes, the number of nodes that can participate in training in each node cluster is calculated.

5. The federated learning method for a spam detection model based on a random strategy according to claim 1, characterized in that: The method of selecting local model parameters uploaded by all client nodes in the cluster based on multiple nodes and aggregating them to generate global model parameters includes: Calculate the first total sample size of local spam training samples in each node cluster; Calculate the second total sample size of local spam training samples in the cluster selected by each node; Calculating a weight correction factor for each node cluster based on the first total sample size and the second total sample size; Based on the local sample size of the local spam training sample of the client node, weighted calculation is performed on the local model parameters uploaded by the client node to obtain the local model weighted parameters; Summing all local model weighted parameters in the target node selected cluster to obtain a local model weighted parameter sum; wherein the target node selected cluster is any one of the multiple node selected clusters; Calculate the product of the weight correction factor of the target node cluster and the sum of the weighted parameters of the local model to obtain the cluster model parameters of the target node cluster; The global model parameters are obtained by summing up all cluster model parameters and dividing them by the sum of all first total sample sizes.

6. A spam detection method, characterized in that: The method comprises: Get the email to be tested; The email to be detected is input into a spam detection model so that the spam detection model outputs a detection result of whether the email to be detected is spam; wherein the spam detection model is trained by the federated learning method of the spam detection model based on the random strategy according to any one of claims 1 to 5.

7. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the federated learning method of the spam detection model based on a random strategy according to any one of claims 1 to 5.

8. A readable storage medium, characterized in that: The readable storage medium stores computer program instructions, and when the computer program instructions are read and executed by a processor, the federated learning method of the spam detection model based on random strategy according to any one of claims 1 to 5 is executed.

9. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the computer program executes the federated learning method of the spam detection model based on random strategy according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Cleaning method and device for training data set and server

    CN114398350A

  • Client selection federal learning method based on DBSCAN clustering

    CN114819069A