Federal learning method of junk mail detection model based on model splitting cooperation
By performing feature matrix similarity evaluation and clustering on client nodes, the spam detection model is split into a subnet, which solves the problem of resource consumption and gradient variance in federated learning, improves model training efficiency and resource utilization, reduces delay, and achieves faster model training speed.
Patent Information
- Application Number
- CN202510404344.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
In the federated learning scenario of spam detection model for large-scale users, the computing storage and network communication resources of the email security monitoring center server and user nodes consume huge amounts, resulting in processing delay and communication delay problems. At the same time, the global model update gradient variance is large, which reduces the convergence speed of model training parameters.
By evaluating and clustering the local spam training samples of client nodes, splitting the spam detection model into multiple subnets, randomly selecting subsets of nodes for training, and judging the termination conditions based on global model parameters and training rounds, reducing resource consumption and gradient variance, and improving training efficiency.
It reduces the computing storage and network communication resource consumption of the mail security monitoring center server and user nodes, reduces processing delays and communication delays, improves resource utilization efficiency, shortens model training time, and improves the parameter convergence speed of model training.
Smart Images

Figure CN120336896A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of spam detection and federated learning. Specifically, it relates to a federated learning method, an electronic device, a readable storage medium, and a computer program product for a spam detection model based on model splitting and cooperation. Background Art
[0002] Currently, in the scenario of federated learning training of a spam detection model for protecting the privacy of user email data for a large number of users, due to the large number of model parameters, when all user nodes participate in each round of iterative training process of federated learning, the computing and storage performance and network bandwidth of the email security monitoring center server and user nodes will impose certain limitations on the efficiency of federated learning.
[0003] A common solution is that in each round of iterative training process of federated learning, a small number of nodes are randomly selected to participate in the training process of this round, so as to reduce the consumption of computing, storage, and network communication resources of the email security monitoring center server, and alleviate the processing delay and communication delay problems caused by the limited computing and storage performance and network bandwidth of the email security monitoring center server.
[0004] However, it is found in practice that although the above solution can solve the above problems faced by the email security monitoring center server to a certain extent, it cannot reduce the consumption of computing, storage, and network communication resources of user nodes, alleviate the processing delay and communication delay problems caused by the limited computing and storage performance and network bandwidth of user nodes, and will also lead to a large variance in the global model update gradient in each round, thereby reducing the parameter convergence speed of model training and increasing the time required for model training. Summary of the Invention
[0005] The purpose of the present application is to provide a federated learning method for a spam detection model based on model splitting and cooperation, which is used to solve the problems such as huge consumption of computing, storage, and network communication resources of the email security monitoring center server and user nodes in the federated learning scenario, processing delay and communication delay caused by the limited computing and storage performance and network bandwidth of the email security monitoring center server and user nodes, and at the same time solve the problem of large variance in the global model update gradient in each round during the training process, thereby reducing the parameter convergence speed of model training and increasing the time required for model training.
[0006] The first aspect of the present application provides a federated learning method for a spam detection model based on model splitting and cooperation. The method is applied to an email security monitoring center server, and the method includes:
[0007] Perform similarity evaluation on the feature matrices of local spam training samples for all client nodes to obtain multiple similarity evaluation results; wherein, the feature matrices are generated based on email header features and email content features, and the email header features include at least one of the sender email domain name type, the number of recipients, and the attachment type, and the email content features include at least one of the email subject and keywords;
[0008] Cluster all client nodes based on multiple similarity evaluation results to obtain multiple node clustering clusters;
[0009] Split the spam detection model into sub-networks to obtain multiple model sub-networks;
[0010] Randomly select multiple node subsets from each node clustering cluster; wherein, the multiple node subsets have a one-to-one training relationship with the multiple model sub-networks;
[0011] Send the latest global model parameters to all client nodes so that each client node can train its own sub-network parameters and average prediction loss values after updating the local model;
[0012] Aggregate the sub-network parameters uploaded by all client nodes to generate global model parameters;
[0013] When it is determined that the spam detection model meets the preset training termination condition based on the global model parameters and the number of training rounds, output the global model parameters as the final parameters to obtain a trained spam detection model.
[0014] In the above implementation process, this method can provide a data basis for subsequent node clustering through similarity evaluation, thus facilitating the classification of nodes with similar data characteristics into one category. This method can also enable nodes within the same cluster to cooperate in training during model training through the clustering method, thereby improving the training efficiency and reducing the problem of large variance in the global model update gradient caused by excessive data differences. This method can also split the spam detection model into multiple sub-networks, facilitating different subsets of nodes to train different sub-networks, thereby minimizing the constraints of computing performance and network bandwidth on federated learning. This method can also train different sub-networks based on similar training samples, thus facilitating subsequent intra-cluster division of labor and inter-cluster cooperation, and further improving the model training efficiency. This method can also perform corresponding dynamic adjustments of division of labor and cooperation based on the updated sub-network parameters and average prediction loss values, thereby contributing to the continuous learning and optimization of the global model, improving the utilization efficiency of resources, and reducing the time required for model training. This method can also aggregate the training results of each node to enable the global model to learn more extensive data characteristics and reduce the impact that partial data deviation can cause to the global model. Finally, appropriate training termination conditions can ensure the effectiveness and rationality of model training, thereby avoiding over-training or under-training and ensuring that the obtained spam detection model has good performance.
[0015] Further, the similarity evaluation of the feature matrices of the local spam training samples of all client nodes is performed to obtain multiple similarity evaluation results, including:
[0016] Receiving the number of non-zero singular values uploaded by the client node and determining the minimum value of the number; wherein, when the client node performs singular value decomposition on the feature matrix of the local spam training sample, a left singular matrix, a singular value matrix, and a right singular matrix are obtained, and the number of non-zero singular values is statistically obtained by the client node based on the singular value matrix;
[0017] Sending the minimum value of the number to all client nodes, so that the client nodes select multiple non-zero singular values with the largest singular values in the singular value matrix based on the minimum value of the number, and generate a clipped singular value clipped matrix and a left singular clipped matrix; wherein, the number of multiple non-zero singular values is the same as the minimum value of the number;
[0018] Receiving the left singular clipped matrix uploaded by the client node and calculating the angle between the column vectors of any two left singular clipped matrices;
[0019] Selecting multiple minimum angles that meet the orthogonality condition from all the calculated angles; wherein, the number of multiple minimum angles is the same as the minimum value of the number;
[0020] Calculate the similarity between the feature matrices of the local spam training samples of any two client nodes based on multiple minimum angles to obtain a similarity evaluation result.
[0021] In the above implementation process, this method can evaluate the similarity of the local spam training samples of client nodes through singular value decomposition and vector angle calculation, thereby providing an accurate basis for subsequent node clustering and targeted model training, and further helping to reduce resource consumption and improve model training efficiency.
[0022] Further, clustering all client nodes based on the multiple similarity evaluation results to obtain multiple node clustering clusters includes:
[0023] Convert the multiple similarity evaluation results into multiple similarity distances;
[0024] Cluster the client nodes based on the multiple similarity distances, the K-means algorithm, and the silhouette coefficient method to obtain multiple node clustering clusters.
[0025] In the above implementation process, this method can determine the optimal number of clustering centers based on the sample data similarity, the K-means algorithm, and the silhouette coefficient method, thereby realizing the reasonable clustering of client nodes, and further facilitating the improvement of the pertinence and efficiency of model training in federated learning and reducing the time required for model training.
[0026] Further, the spam detection model has a neural network;
[0027] Each model sub-network is composed of some neurons of the neural network;
[0028] The neurons of all model sub-networks together constitute all the neurons of the neural network.
[0029] In the above implementation process, this provides a structural basis for parallel training and targeted optimization of the model.
[0030] Further, when the current training round is 1, randomly selecting multiple node subsets from each node clustering cluster includes:
[0031] When the number of nodes in the node clustering cluster is divisible by the number of multiple model sub-networks, evenly divide the nodes in the node clustering cluster based on the number of multiple model sub-networks to obtain multiple node subsets; or
[0032] When the number of nodes in the node clustering cluster is not divisible by the number of multiple model sub-networks, divide the nodes in the node clustering cluster based on the number of multiple model sub-networks by rounding up or down to obtain multiple node subsets;
[0033] Among them, the number of nodes in a single node clustering cluster is the same as the total number of nodes in the corresponding multiple node subsets.
[0034] In the above implementation process, this method can reasonably divide node clustering clusters into node subsets in the case of different numbers of nodes, ensuring that each model sub-network has appropriate nodes participating in training and no nodes are omitted.
[0035] Further, when the current training round is not 1, randomly selecting multiple node subsets from each node clustering cluster includes:
[0036] Based on the average prediction loss value, calculate the overall average prediction loss value corresponding to the model target sub-network; where the model target sub-network is any one of the multiple model sub-networks;
[0037] Based on the overall average prediction loss value, calculate the division ratio corresponding to each model target sub-network;
[0038] Using the method of rounding up or rounding down, divide the nodes in the node clustering cluster based on all division ratios to obtain multiple node subsets;
[0039] Among them, the number of nodes in a single node clustering cluster is the same as the total number of nodes in the corresponding multiple node subsets.
[0040] In the above implementation process, this method can dynamically divide node clustering clusters based on the average prediction loss value, so as to allocate a more appropriate number of nodes to different sub-networks, thereby improving the utilization efficiency of resources and the training efficiency of federated learning.
[0041] The second aspect of this application provides a spam detection method, and the method includes:
[0042] Obtain the email to be detected;
[0043] Input the email to be detected into the spam detection model, so that the spam detection model outputs the detection result of whether the email to be detected is spam; where the spam detection model is trained according to the federated learning method of the spam detection model based on model splitting and cooperation described in any item of the first aspect of this application.
[0044] The third aspect of this application provides a federated learning device for a spam detection model based on model splitting and cooperation. The federated learning device for the spam detection model based on model splitting and cooperation is a mail security monitoring center server, and the federated learning device for the spam detection model based on model splitting and cooperation includes:
[0045] A similarity evaluation unit for evaluating the similarity of the feature matrices of the local spam training samples of all client nodes to obtain multiple similarity evaluation results; wherein, the feature matrix is generated based on the email header features and the email content features, and the email header features include at least one of the sender email domain name type, the number of recipients, and the attachment type, and the email content features include at least one of the email subject and keywords;
[0046] A clustering unit for clustering all client nodes based on multiple similarity evaluation results to obtain multiple node clustering clusters;
[0047] A model splitting unit for splitting the spam detection model into sub-networks to obtain multiple model sub-networks;
[0048] A selection unit for randomly selecting multiple node subsets from each node clustering cluster; wherein, the multiple node subsets have a one-to-one training relationship with the multiple model sub-networks;
[0049] A distribution unit for distributing the latest global model parameters to all client nodes so that each client node can train to obtain the sub-network parameters and the average prediction loss value after updating the local model;
[0050] An aggregation unit for aggregating and generating global model parameters based on the sub-network parameters uploaded by all client nodes;
[0051] An output unit for, when it is determined that the spam detection model meets the preset training termination condition based on the global model parameters and the number of training rounds, outputting the global model parameters as the final parameters to obtain a trained spam detection model.
[0052] Further, the similarity evaluation unit includes:
[0053] A receiving sub-unit for receiving the number of non-zero singular values uploaded by the client node and determining the minimum value; wherein, when the client node performs singular value decomposition on the feature matrix of the local spam training sample, a left singular matrix, a singular value matrix, and a right singular matrix are obtained, and the number of non-zero singular values is statistically obtained by the client node based on the singular value matrix;
[0054] A distribution sub-unit for distributing the minimum value to all client nodes so that the client node selects multiple non-zero singular values with the largest singular values in the singular value matrix based on the minimum value and generates a clipped singular value clipped matrix and a left singular clipped matrix; wherein, the number of the multiple non-zero singular values is the same as the minimum value;
[0055] The receiving subunit is further configured to receive the left singular pruning matrix uploaded by the client node, and calculate the angle between the column vectors of any two left singular pruning matrices;
[0056] The selection subunit is configured to select multiple minimum angles that meet the orthogonality condition from all the calculated angles; wherein, the number of the multiple minimum angles is the same as the minimum value of the quantity;
[0057] The first calculation subunit is configured to calculate the similarity between the feature matrices of the local spam training samples of any two client nodes based on the multiple minimum angles, and obtain a similarity evaluation result.
[0058] Further, the clustering unit includes:
[0059] The conversion subunit is configured to convert the multiple similarity evaluation results into multiple similarity distances;
[0060] The clustering subunit is configured to cluster the client nodes based on the multiple similarity distances, the K-means algorithm, and the silhouette coefficient method, and obtain multiple node clustering clusters.
[0061] Further, the spam detection model has a neural network;
[0062] Each model subnet is composed of some neurons of the neural network;
[0063] The neurons of all model subnets together constitute all the neurons of the neural network.
[0064] Further, when the current training round is 1, the selection unit is specifically configured to evenly divide the nodes in the node clustering cluster based on the number of multiple model subnets when the number of nodes in the node clustering cluster is divisible by the number of multiple model subnets, to obtain multiple node subsets; or
[0065] When the number of nodes in the node clustering cluster is not divisible by the number of multiple model subnets, divide the nodes in the node clustering cluster based on the number of multiple model subnets by rounding up or rounding down, to obtain multiple node subsets;
[0066] Wherein, the number of nodes in a single node clustering cluster is the same as the total number of nodes in the corresponding multiple node subsets.
[0067] Further, when the current training round is not 1, the selection unit includes:
[0068] The second calculation subunit is configured to calculate the overall average prediction loss value corresponding to the model target subnet based on the average prediction loss value; wherein, the model target subnet is any one of the multiple model subnets;
[0069] The second calculation subunit is further configured to calculate a division ratio corresponding to each model target sub-network based on the overall average prediction loss value;
[0070] The division subunit is configured to divide the nodes in the node clustering cluster based on all the division ratios in a way of rounding up or rounding down to obtain a plurality of node subsets;
[0071] Wherein, the number of nodes in a single node clustering cluster is the same as the total number of nodes in the corresponding plurality of node subsets.
[0072] In a fourth aspect of the present application, an electronic device is provided. The electronic device includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the federated learning method of the spam detection model based on model splitting and cooperation according to any one of the first aspects of the present application.
[0073] In a fifth aspect of the present application, a computer-readable storage medium is provided. Computer program instructions are stored in the readable storage medium. When the computer program instructions are read and run by a processor, the federated learning method of the spam detection model based on model splitting and cooperation according to any one of the first aspects of the present application is executed.
[0074] In a sixth aspect of the present application, a computer program product is provided. The computer program product includes a computer program. When the computer program is run by a processor, the federated learning method of the spam detection model based on model splitting and cooperation according to any one of the first aspects of the present application is executed.
[0075] The beneficial effects of the present application are as follows: The method of randomly selecting client nodes according to the clustering cluster based on client node clustering can reduce the variance of the global model update gradient in each round compared with the traditional method of randomly selecting some client nodes to train all the parameters of the model, thereby improving the parameter convergence speed of model training, reducing the time required for model training, reducing resource consumption, and improving resource utilization efficiency.
[0076] Through the method of model splitting and cooperation, each client node can only train part of the parameters of the model, thereby reducing the local model training workload of the client node, reducing the number of parameters uploaded by the client node in each round, reducing the local model training time and communication time of the client node, improving the overall training speed of the federated learning model, reducing the processing delay and communication delay of the spam security monitoring center server, and further reducing the requirements of the federated learning technology for the computing and communication resources of the client node, promoting the popularization and application of the federated learning technology in scenarios where the computing and communication resources of client nodes such as Internet of Things devices and embedded devices are limited.
[0077] When randomly selecting the participating client nodes for different sub-networks, the number of client nodes can be dynamically allocated according to the average prediction loss value of each sub-network, so as to improve the utilization efficiency of client node resources, increase the convergence speed of the global model, and reduce the overall training time of the federated learning model. BRIEF DESCRIPTION OF THE DRAWINGS
[0078] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0079] Figure 1 Schematic flowchart of a federated learning method for a spam detection model based on model splitting and collaboration provided by an embodiment of the present application;
[0080] Figure 2 Schematic flowchart of another federated learning method for a spam detection model based on model splitting and collaboration provided by an embodiment of the present application;
[0081] Figure 3 Schematic diagram of the technical implementation process of a federated learning method for a spam detection model based on model splitting and collaboration provided by an embodiment of the present application;
[0082] Figure 4 Schematic flowchart of a spam detection method provided by an embodiment of the present application;
[0083] Figure 5 Schematic diagram of the structure of a federated learning device for a spam detection model based on model splitting and collaboration provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0084] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application.
[0085] It should be noted that similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings. At the same time, in the description of the present application, the terms "first", "second", etc. are only used for differential description and cannot be understood as indicating or implying relative importance.
[0086] Embodiment 1
[0087] Please refer to Figure 1 , Figure 1Schematic flow chart of a federated learning method for a spam detection model based on model splitting and collaboration. Among them, the federated learning method for the spam detection model based on model splitting and collaboration is applied to the mail security monitoring center server, and the method includes:
[0088] S101. Evaluate the similarity of the feature matrices of the local spam training samples of all client nodes to obtain multiple similarity evaluation results; among them, the feature matrix is generated based on the mail header features and mail content features, and the mail header features at least include one of the sender email domain name type, the number of recipients, and the attachment type, and the mail content features at least include one of the mail subject and keywords.
[0089] In this embodiment, each client node can extract information such as the sender email domain name type (such as enterprise email domain name, personal email domain name, government institution email domain name, etc.), the number of recipients, and the attachment type (executable file, document, compressed package, image / video, no attachment, etc.) from the mail header of the local spam training sample, and use it as the mail header feature.
[0090] At the same time, each client node can also extract information related to the mail content such as the mail subject and keywords from the mail body of the local spam training sample, and use it as the mail content feature.
[0091] On this basis, the method can generate the feature matrix of the local spam training sample of each client node according to the above mail header features and mail content features.
[0092] S102. Cluster all client nodes based on multiple similarity evaluation results to obtain multiple node clusters.
[0093] S103. Split the spam detection model into multiple model sub-networks.
[0094] In this embodiment, the spam detection model has a neural network;
[0095] Each model sub-network is composed of some neurons of the neural network;
[0096] The neurons of all model sub-networks together constitute all the neurons of the neural network.
[0097] S104. Randomly select multiple node subsets from each node cluster; among them, the multiple node subsets have a one-to-one training relationship with the multiple model sub-networks.
[0098] S105. Send the latest global model parameters to all client nodes so that each client node can train the sub-network parameters and the average prediction loss value after the local model is updated respectively.
[0099] S106. Aggregate and generate global model parameters based on the sub-network parameters uploaded by all client nodes.
[0100] S107. When it is determined that the spam detection model meets the preset training termination condition based on the global model parameters and the number of training rounds, output the global model parameters as the final parameters to obtain the trained spam detection model.
[0101] In this embodiment, the execution subject of this method can be a computing device such as a computer or a server, and no limitation is made in this embodiment.
[0102] It can be seen that implementing the federated learning method of the spam detection model based on model splitting and collaboration described in this embodiment can solve the processing delay problem and communication delay problem caused by the limited computing and processing performance and network bandwidth of the mail security monitoring center server and user nodes in the federated learning scenario, reduce the consumption of computing storage and network communication resources of the mail security monitoring center server and user nodes, improve the resource utilization efficiency of user nodes, and at the same time solve the problem that the global model update gradient variance in each round of the training process is relatively large, thereby reducing the parameter convergence speed of model training and increasing the time required for model training.
[0103] Embodiment 2
[0104] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a federated learning method for a spam detection model based on model splitting and collaboration provided in this embodiment. Among them, the federated learning method for the spam detection model based on model splitting and collaboration is applied to the mail security monitoring center server, and this method includes:
[0105] S201. Receive the number of non-zero singular values uploaded by the client node and determine the minimum value; wherein, when the client node performs singular value decomposition on the feature matrix of the local spam training sample, a left singular matrix, a singular value matrix, and a right singular matrix are obtained, and the number of non-zero singular values is statistically obtained by the client node based on the singular value matrix.
[0106] For example, the total number of client nodes (client hosts) with a spam detection model is set to N c = 1000000. The feature matrix of all local spam training samples of the i-th client host X iEach component included represents a feature column vector of a spam training sample, and the dimension of the feature column vector of each spam training sample is M = 500 (each column vector represents the header features, content features or some of the above features of a spam email), N i represents the total number of local spam training samples of the i-th client host. Among them, 1 ≤ i ≤ N c , i, N i are positive integers.
[0107] At this time, the i-th client host performs singular value decomposition on the feature matrix X i , that is U i ∈R M×M is the left singular matrix of X i , is the singular value matrix of X i , is the right singular matrix of X i . This client host arranges the singular values above ∑ i in descending order, counts the number of non-zero singular values (that is, the number of non-zero singular values uploaded by the client node in step S201), and uploads this number to the email security monitoring center server S.
[0108] S202. Send the minimum value of the quantity to all client nodes, so that the client nodes select multiple non-zero singular values with the largest singular values in the singular value matrix based on the minimum value of the quantity, and generate a trimmed singular value trimmed matrix and a left singular trimmed matrix; among them, the number of multiple non-zero singular values is the same as the minimum value of the quantity.
[0109] For example, the email security monitoring center server S receives the number of non-zero singular values uploaded by N c client hosts, counts the minimum value of the above numbers, and sends the minimum value to all client hosts. Denote this minimum value as P, where P is a positive integer; in this example, P = 100.
[0110] S203. Receive the left singular trimmed matrix uploaded by the client node, and calculate the angle between the column vectors of any two left singular trimmed matrices.
[0111] For example, after the i-th client host receives the value P sent by the email security monitoring center server S, it selects the largest P non-zero singular values in ∑ i to generate a trimmed singular value matrix ∑′ i ∈R P×P , selects the column vectors corresponding to the above P non-zero singular values in U i to generate a trimmed left singular matrix U′ i ∈RM×P Each client host uploads its respective U′ i to the email security monitoring center server S.
[0112] Among them, the column vector of a matrix refers to the vector composed of all the numerical values in a certain column of the matrix. The column vector in the left singular matrix corresponding to a singular value refers to the column vector in the left singular matrix whose column number is equal to the row number of this singular value in the singular value matrix.
[0113] Furthermore, the email security monitoring center server S receives N c clipped left singular matrices. For any two clipped left singular matrices, calculate the angle between any two column vectors belonging to the two matrices according to the following formula, where
[0114]
[0115] S204. Select multiple minimum angles that meet the orthogonality condition from all the calculated angles; among them, the number of multiple minimum angles is the same as the minimum quantity.
[0116] For example, the method continues to select the P minimum angles between the column vectors of U′ i and the column vectors of U′ i that meet the following orthogonality condition from all the above angles 1 ≤ j ≤ N c , j ≠ i, and j is a positive integer.
[0117] Orthogonality condition: When p ≥ 2, must be orthogonal to both, must be orthogonal to both.
[0118] Among them, represents is the column vector of matrix U′ i , represents is the column vector of matrix U′ i , represents the inner product of vector and vector , |x| represents the modulus of vector x, 1 ≤ p ≤ P, and p is a positive integer.
[0119] S205. Calculate the similarity between the feature matrices of the local spam training samples of any two client nodes based on multiple minimum angles, and obtain the similarity evaluation result.
[0120] For example, the email security monitoring center server S calculates the spam training data similarity between any two client hosts according to the following formula:
[0121]
[0122] Among them, the spam training data similarity between the i-th client host and the j-th client host is denoted as S(i, j).
[0123] S206. Convert multiple similarity evaluation results into multiple similarity distances.
[0124] S207. Cluster the client nodes based on multiple similarity distances, the K-means algorithm, and the silhouette coefficient method to obtain multiple node clustering clusters.
[0125] In this embodiment, the method can traverse different numbers of clustering centers based on multiple similarity distances, the K-means algorithm, and the silhouette coefficient method, and determine the optimal number of clustering centers when the clustering result performs best. At this time, the optimal clustering result is the multiple node clustering clusters finally obtained in step S207.
[0126] In this embodiment, the method can also re-cluster the client nodes based on multiple similarity distances, the K-means algorithm, and the optimal number of clustering centers, so as to obtain multiple node clustering clusters corresponding to the optimal clustering result.
[0127] For example, the email security monitoring center server S measures the distance between the i-th client host and the j-th client host using P-S(i, j), and uses the K-means algorithm to cluster N c client hosts (the number of clustering centers N a is determined by the silhouette coefficient method), and finally clusters to obtain N a = 10 sets composed of client hosts Among them, N a is a positive integer.
[0128] S208. Split the spam detection model into multiple model sub-networks.
[0129] For example, the email security monitoring center server S initializes the parameters of the adopted neural network model (such as Bert, LSTM, RNN, etc.), and the initialized parameters are w 0 . The email security monitoring center server S splits the neural network model into a total of N l = 2 sub-networks M1 and M2 for the feature extraction layer and the classification layer (when N l is greater than 2, the N l sub-networks are )。The mail security monitoring center server S distributes the architecture and sub-network splitting results of the neural network model to all client hosts, and all client hosts use the architecture of the neural network model as the architecture of the local model. Among them, N l is a positive integer.
[0130] In this embodiment, after step S209, the step of randomly selecting multiple node subsets from each node clustering cluster is performed. Among them, this step has different specific processes based on whether the current training round is 1. Specifically, when the current training round is 1, the following optional implementation manner is performed; when the current training round is not 1, step S210 is performed.
[0131] As an optional implementation manner, when the current training round is 1, randomly selecting multiple node subsets from each node clustering cluster includes:
[0132] When the number of nodes in the node clustering cluster is divisible by the number of multiple model sub-networks, the nodes in the node clustering cluster are evenly divided based on the number of multiple model sub-networks to obtain multiple node subsets; or
[0133] When the number of nodes in the node clustering cluster is not divisible by the number of multiple model sub-networks, the nodes in the node clustering cluster are divided based on the number of multiple model sub-networks by using the ceiling or floor method to obtain multiple node subsets;
[0134] Among them, the number of nodes in a single node clustering cluster is the same as the total number of nodes in the corresponding multiple node subsets.
[0135] For example, in the r-th round of the model training cycle, the mail security monitoring center server S randomly divides each set A a randomly into N l = 2 subsets and The client hosts in only responsible for the parameter training of the sub-network M l . Among them, 1 ≤ a ≤ N a , 1 ≤ l ≤ N l , 1 ≤ r ≤ R end , R end is the training round threshold constant, and a, l, r, R end are positive integers.
[0136] When r = 1 and the cardinality of the set A a can be divided by N l = 2, the subset division of the set A a adopts the equal division strategy, that is, the cardinality of each subset is the cardinality of the set A a divided by N lThe obtained quotient;
[0137] When r = 1 and the cardinality of set A a cannot be divided evenly by N l = 2, the subset partitioning of set A a adopts an approximate equal - division strategy, that is, round up or round down the result of dividing the cardinality of set A a by N l , but it must satisfy that the intersection of any two subsets is empty and the union of all subsets is A a .
[0138] S209. When the current training round is not 1, calculate the overall average prediction loss value corresponding to the model target sub - network based on the average prediction loss value; where the model target sub - network is any one of multiple model sub - networks.
[0139] S210. Calculate the partitioning ratio corresponding to each model target sub - network based on the overall average prediction loss value.
[0140] S211. Adopt the method of rounding up or rounding down, and partition the nodes in the node clustering cluster based on all partitioning ratios to obtain multiple node subsets.
[0141] In this embodiment, the number of nodes in a single node clustering cluster is the same as the total number of nodes in the corresponding multiple node subsets.
[0142] For example, when r ≥ 2, the subset partitioning of set A a adopts a non - equal - division strategy, calculate the average prediction loss value of the l - th sub - network and the ratio P of the cardinality of the partitioned subset to the cardinality of A a . l .
[0143]
[0144] Among them, if P l ×|A a | is an integer, then randomly select P a ×|A l | client hosts from A a to form a subset
[0145] If P l ×|A a | is not an integer, then round up or round down P l ×|A a |, and randomly select P a ×|A l ×|Aa The integer - valued subset composed of the client hosts However, it must be satisfied that the intersection of any two subsets is empty and the union of all subsets is A a . Wherein, |A a represents the set A a of the cardinality.
[0146] S212. Send the latest global model parameters to all client nodes so that each client node can train the sub - network parameters and the average prediction loss value after the local model is updated respectively.
[0147] For example, in the r - th round of the model training cycle, the mail security monitoring center server S sends the model parameter w r-1 and the sub - network serial number l i to the i - th client host. l i represents the serial number of the sub - network responsible for training by the i - th client node after the subsets are divided. Where 1 ≤ l i ≤ N l , l i is a positive integer.
[0148] Furthermore, in the r - th round of the model training cycle, the i - th client host receives the model parameter w r-1 and the sub - network serial number l i sent by the mail security monitoring center server S. The client host updates the local model parameter value to w r-1 , fixes the values of all parameters of the local model except the l i - th sub - network, and uses the gradient descent algorithm to train and update the parameters of the l i - th sub - network of the local model on the local spam training sample set. The parameters of the l i - th sub - network of the updated local model are When the parameters of the l i - th sub - network are updated to and the parameters of other sub - networks are the aforementioned fixed values, calculate the average prediction loss value of the local model of this client host on the local spam training sample set This client host will and upload to the mail security monitoring center server S.
[0149] Among them, the average prediction loss value refers to the arithmetic mean of the prediction loss values of the local model for all local training samples. The prediction loss value refers to the error between the predicted value and the true value of the model for a given sample under a given error calculation method (such as mean square error, cross - entropy error).
[0150] S213. Aggregate and generate global model parameters based on the sub-network parameters uploaded by all client nodes.
[0151] As an alternative implementation method, aggregating and generating global model parameters based on the sub-network parameters uploaded by all client nodes includes:
[0152] Calculate the first total sample size of the local spam training samples in each node cluster;
[0153] Calculate the second total sample size of the local spam training samples in each node subset;
[0154] Calculate the weight correction factor of each node cluster based on the first total sample size and the second total sample size;
[0155] Perform weighted calculation on the sub-network parameters uploaded by the client nodes based on the local sample size of the local spam training samples of the client nodes to obtain the weighted local model parameters;
[0156] Sum up all the weighted local model parameters in the target node subset to obtain the sum of the weighted local model parameters; where the target node subset is a part of the target node cluster;
[0157] Calculate the product of the weight correction factor of the target node cluster and the sum of the weighted local model parameters to obtain the cluster model parameters of the target node cluster;
[0158] Sum up all the cluster model parameters and divide by the sum value of all the first total sample sizes to obtain the sub-network model parameters;
[0159] Link all the sub-network model parameters together in a certain order to obtain the global model parameters.
[0160] For example, in the r-th round of the model training cycle, the mail security monitoring center server S receives the sub-network parameters after the local model is updated uploaded by each client host and Aggregate and generate the global model parameter w r . Among them,
[0161]
[0162] Among them, represents the set A a The sum of the sample numbers of the local training sample sets (local spam training sample sets) of all client nodes, |D i | is the sample number of the local training sample set of the i-th client node, represents that the i-th client node belongs to the set w ris N l vectors linked together in a certain order to form a new vector, vector w r has a length of N l vectors the sum of the lengths.
[0163] Implementing this implementation method can correct the aggregation weights of the local model parameters of each client node when aggregating and generating global model parameters, thereby reducing the deviation between the updated value of the global model parameters under the condition of randomly selecting client nodes to participate in model training and the updated value of the global model parameters when all client nodes participate in model training. Furthermore, it can ensure that randomly selecting client nodes will not significantly change the convergence direction of the model and improve the correctness and accuracy of the model training results.
[0164] S214. When it is determined that the spam detection model meets the preset training termination condition based on the global model parameters and the number of training rounds, output the global model parameters as the final parameters to obtain the trained spam detection model.
[0165] For example, if w r converges or r ≥ R end = 200, then the mail security monitoring center server S terminates the entire model training process and outputs w r as the final parameters of the global model; otherwise, set r = r + 1, and the mail security monitoring center server S starts the next round of model training cycle, and repeats the steps of S209 - S213 in the above loop.
[0166] Please refer to Figure 3 , Figure 3 which shows a schematic diagram of the technical implementation process of a federated learning method for a spam detection model based on model splitting and collaboration. Among them, Figure 3 the shown process has a corresponding relationship with Figure 1 and Figure 2 the processes in.
[0167] In this embodiment, the execution subject of this method can be a computing device such as a computer or a server, and no limitation is made in this embodiment.
[0168] It can be seen that implementing the federated learning method of the spam detection model based on model splitting and collaboration described in this embodiment can solve the processing delay problem and communication delay problem caused by the limited computing and processing performance and network bandwidth of the email security monitoring center server and user nodes in the federated learning scenario, reduce the consumption of computing storage and network communication resources of the email security monitoring center server and user nodes, improve the resource utilization efficiency of user nodes, and at the same time solve the problem that the global model update gradient variance in each round of the training process is relatively large, thereby reducing the parameter convergence speed of model training and increasing the time required for model training.
[0169] Embodiment 3
[0170] Please refer to Figure 4 , Figure 4 which is a schematic flowchart of a spam detection method provided in this embodiment. Among them, the method includes:
[0171] S301. Obtain the email to be detected.
[0172] S302. Input the email to be detected into the spam detection model so that the spam detection model outputs a detection result indicating whether the email to be detected is spam.
[0173] In this embodiment, the spam detection model is trained according to the federated learning method of the spam detection model based on model splitting and collaboration in Embodiment 1 or Embodiment 2 of this application.
[0174] In this embodiment, the execution subject of the method can be a computing device such as a computer or a server, and no limitation is made in this embodiment.
[0175] It can be seen that implementing the federated learning method of the spam detection model based on model splitting and collaboration described in this embodiment can solve the processing delay problem and communication delay problem caused by the limited computing and processing performance and network bandwidth of the email security monitoring center server and user nodes in the federated learning scenario, reduce the consumption of computing storage and network communication resources of the email security monitoring center server and user nodes, improve the resource utilization efficiency of user nodes, and at the same time solve the problem that the global model update gradient variance in each round of the training process is relatively large, thereby reducing the parameter convergence speed of model training and increasing the time required for model training.
[0176] Embodiment 4
[0177] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of a federated learning device of a spam detection model based on model splitting and collaboration. Among them, the federated learning device of the spam detection model based on model splitting and collaboration can be an email security monitoring center server. AsFigure 5 As shown in Figure 5 , the federated learning device of the spam detection model based on model splitting and collaboration includes:
[0178] A similarity evaluation unit 410, configured to perform similarity evaluation on the feature matrices of the local spam training samples of all client nodes to obtain multiple similarity evaluation results; wherein, the feature matrix is generated based on the email header features and the email content features, and the email header features include at least one of the sender email domain name type, the number of recipients, and the attachment type, and the email content features include at least one of the email subject and keywords;
[0179] A clustering unit 420, configured to cluster all client nodes based on the multiple similarity evaluation results to obtain multiple node clustering clusters;
[0180] A model splitting unit 430, configured to split the spam detection model into sub-networks to obtain multiple model sub-networks;
[0181] A selection unit 440, configured to randomly select multiple node subsets from each node clustering cluster; wherein, the multiple node subsets have a one-to-one training relationship with the multiple model sub-networks;
[0182] A distribution unit 450, configured to distribute the latest global model parameters to all client nodes, so that each client node independently trains to obtain the sub-network parameters and the average prediction loss value after the local model is updated;
[0183] An aggregation unit 460, configured to aggregate and generate global model parameters based on the sub-network parameters uploaded by all client nodes;
[0184] An output unit 470, configured to, when it is determined that the spam detection model meets the preset training termination condition based on the global model parameters and the number of training rounds, output the global model parameters as the final parameters to obtain the trained spam detection model.
[0185] As an optional implementation manner, the similarity evaluation unit 410 includes:
[0186] A receiving subunit 411, configured to receive the number of non-zero singular values uploaded by the client node and determine the minimum value of the number; wherein, when the client node performs singular value decomposition on the feature matrix of the local spam training sample, a left singular matrix, a singular value matrix, and a right singular matrix are obtained, and the number of non-zero singular values is statistically obtained by the client node based on the singular value matrix;
[0187] A sending subunit 412, configured to send the minimum quantity to all client nodes, so that the client nodes select multiple non-zero singular values with the largest singular values in the singular value matrix based on the minimum quantity, and generate a clipped singular value matrix and a left singular clipped matrix; wherein, the number of the multiple non-zero singular values is the same as the minimum quantity.
[0188] A receiving subunit 411 is further configured to receive the left singular clipped matrix uploaded by the client node, and calculate the angle between the column vectors of any two left singular clipped matrices.
[0189] A selecting subunit 413, configured to select multiple minimum angles that meet the orthogonality condition from all the calculated angles; wherein, the number of the multiple minimum angles is the same as the minimum quantity.
[0190] A first calculating subunit 414, configured to calculate the similarity between the feature matrices of the local spam training samples of any two client nodes based on the multiple minimum angles, and obtain a similarity evaluation result.
[0191] As an optional implementation manner, the clustering unit 420 includes:
[0192] A conversion subunit 421, configured to convert the multiple similarity evaluation results into multiple similarity distances.
[0193] A clustering subunit 422, configured to cluster the client nodes based on the multiple similarity distances, the K-means algorithm, and the silhouette coefficient method, and obtain multiple node clustering clusters.
[0194] In this embodiment, the spam detection model has a neural network.
[0195] Each model sub-network is composed of some neurons of the neural network.
[0196] The neurons of all the model sub-networks together constitute all the neurons of the neural network.
[0197] As an optional implementation manner, when the current training round is 1, the selecting unit 440 is specifically configured to, when the number of nodes in the node clustering cluster can be divided evenly by the number of multiple model sub-networks, evenly divide the nodes in the node clustering cluster based on the number of multiple model sub-networks to obtain multiple node subsets; or
[0198] when the number of nodes in the node clustering cluster cannot be divided evenly by the number of multiple model sub-networks, divide the nodes in the node clustering cluster based on the number of multiple model sub-networks in a way of rounding up or rounding down to obtain multiple node subsets.
[0199] Among them, the number of nodes in a single node clustering cluster is the same as the total number of nodes in the corresponding multiple node subsets.
[0200] As an alternative implementation, when the current training round is not 1, the selection unit 440 includes:
[0201] A second calculation subunit 441, configured to calculate an overall average prediction loss value corresponding to the model target sub-network based on the average prediction loss value; wherein, the model target sub-network is any one of multiple model sub-networks;
[0202] The second calculation subunit 441 is further configured to calculate a division ratio corresponding to each model target sub-network based on the overall average prediction loss value;
[0203] A division subunit 442, configured to divide the nodes in the node clustering cluster based on all the division ratios in a ceiling or floor manner to obtain multiple node subsets;
[0204] Among them, the number of nodes in a single node clustering cluster is the same as the total number of nodes in the corresponding multiple node subsets.
[0205] In this embodiment, the explanation of the federated learning device for the spam detection model based on model splitting and cooperation can refer to the description in Embodiment 1 or Embodiment 2, and thus will not be elaborated herein.
[0206] It can be seen that implementing the federated learning device for the spam detection model based on model splitting and cooperation described in this embodiment can solve the processing delay problem and communication delay problem caused by the limited computing and processing performance and network bandwidth of the mail security monitoring center server and user nodes in the federated learning scenario, reduce the consumption of computing storage and network communication resources of the mail security monitoring center server and user nodes, improve the resource utilization efficiency of user nodes, and at the same time solve the problem that the global model update gradient variance in each round of the training process is relatively large, thereby reducing the parameter convergence speed of model training and increasing the time required for model training.
[0207] An embodiment of the present application provides an electronic device, including a memory and a processor, where the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to execute the federated learning method for the spam detection model based on model splitting and cooperation in Embodiment 1 or Embodiment 2 of the present application.
[0208] An embodiment of the present application provides a computer-readable storage medium, which stores computer program instructions, and when the computer program instructions are read and run by a processor, the federated learning method for the spam detection model based on model splitting and cooperation in Embodiment 1 or Embodiment 2 of the present application is executed.
[0209] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is run by a processor, it executes the federated learning method of the spam detection model based on model splitting and collaboration in Embodiment 1 or Embodiment 2 of the present application.
[0210] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0211] In addition, in each embodiment of the present application, the functional modules can be integrated together to form an independent part, or each module can exist separately, or two or more modules can be integrated to form an independent part.
[0212] If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.
[0213] The above are only embodiments of the present application and are not intended to limit the protection scope of the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application. It should be noted that similar reference numerals and letters indicate similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0214] As described above, this is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or replacements within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
[0215] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element.
Claims
1. A federated learning method for a spam detection model based on model splitting and collaboration, characterized in that The method is applied to the server of the email security monitoring center, and the method includes: Performing similarity evaluation on the feature matrices of the local spam training samples of all client nodes to obtain multiple similarity evaluation results; wherein, the feature matrix is generated based on the email header features and the email content features, and the email header features include at least one of the sender email domain name type, the number of recipients, and the attachment type, and the email content features include at least one of the email subject and keywords; Clustering all client nodes based on the multiple similarity evaluation results to obtain multiple node clustering clusters; Splitting the spam detection model into multiple model sub-networks; Randomly selecting multiple node subsets from each node clustering cluster; wherein, there is a one-to-one training relationship between the multiple node subsets and the multiple model sub-networks; Sending the latest global model parameters to all client nodes so that each client node can train to obtain the sub-network parameters and the average prediction loss value after updating the local model; Aggregating the sub-network parameters uploaded by all client nodes to generate global model parameters; When it is determined that the spam detection model meets the preset training termination condition based on the global model parameters and the number of training rounds, outputting the global model parameters as the final parameters to obtain the trained spam detection model.
2. The federated learning method for spam detection model based on model splitting and collaboration according to claim 1, characterized in that The performing similarity evaluation on the feature matrices of the local spam training samples of all client nodes to obtain multiple similarity evaluation results includes: Receiving the number of non-zero singular values uploaded by the client node and determining the minimum value; wherein, when the client node performs singular value decomposition on the feature matrix of the local spam training sample, a left singular matrix, a singular value matrix, and a right singular matrix are obtained, and the number of non-zero singular values is statistically obtained by the client node based on the singular value matrix; Sending the minimum value to all client nodes so that the client node selects multiple non-zero singular values with the largest singular values in the singular value matrix based on the minimum value and generates a trimmed singular value trimmed matrix and a left singular trimmed matrix; wherein, the number of the multiple non-zero singular values is the same as the minimum value; Receiving the left singular trimmed matrix uploaded by the client node and calculating the angle between the column vectors of any two left singular trimmed matrices; Selecting multiple minimum angles that meet the orthogonal condition from all the calculated angles; wherein, the number of the multiple minimum angles is the same as the minimum value; Calculating the similarity between the feature matrices of the local spam training samples of any two client nodes based on the multiple minimum angles to obtain the similarity evaluation result.
3. The federated learning method for spam detection model based on model splitting and collaboration as claimed in claim 1, wherein The clustering all client nodes based on the multiple similarity evaluation results to obtain multiple node clustering clusters includes: Converting the multiple similarity evaluation results into multiple similarity distances; Clustering the client nodes based on the multiple similarity distances, the K-means algorithm, and the silhouette coefficient method to obtain multiple node clustering clusters.
4. The federated learning method for spam detection model based on model splitting and collaboration according to claim 1, characterized in that The spam detection model has a neural network; Each model sub-network is composed of some neurons of the neural network; The neurons of all model sub-networks together constitute all neurons of the neural network.
5. The federated learning method for spam detection model based on model splitting and collaboration according to claim 1, characterized in that When the current training round is 1, randomly select multiple node subsets from each node cluster, including: When the number of nodes in the node cluster can be evenly divided by the number of multiple model sub-networks, evenly divide the nodes in the node cluster based on the number of multiple model sub-networks to obtain multiple node subsets; or When the number of nodes in the node cluster cannot be evenly divided by the number of multiple model sub-networks, divide the nodes in the node cluster based on the number of multiple model sub-networks by rounding up or rounding down to obtain multiple node subsets; Among them, the number of nodes in a single node cluster is the same as the total number of nodes in the corresponding multiple node subsets.
6. The federated learning method for spam detection model based on model splitting and collaboration according to claim 1, characterized in that When the current training round is not 1, randomly select multiple node subsets from each node cluster, including: Based on the average prediction loss value, calculate the overall average prediction loss value corresponding to the model target sub-network; wherein, the model target sub-network is any one of the multiple model sub-networks; Based on the overall average prediction loss value, calculate the division ratio corresponding to each model target sub-network; Divide the nodes in the node cluster based on all division ratios by rounding up or rounding down to obtain multiple node subsets; Among them, the number of nodes in a single node cluster is the same as the total number of nodes in the corresponding multiple node subsets.
7. A spam detection method, characterized in that, The method includes: Obtain the email to be detected; Input the email to be detected into the spam detection model so that the spam detection model outputs the detection result of whether the email to be detected is spam; wherein, the spam detection model is trained according to the federated learning method of the spam detection model based on model splitting and cooperation described in any one of claims 1 to 6.
8. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory is used to store a computer program, and the processor runs the computer program so that the electronic device executes the federated learning method of the spam detection model based on model splitting and cooperation described in any one of claims 1 to 6.
9. A readable storage medium, characterized in that, Computer program instructions are stored in the readable storage medium. When the computer program instructions are read and run by a processor, the federated learning method of the spam detection model based on model splitting and cooperation described in any one of claims 1 to 6 is executed.
10. A computer program product, characterized in that, The computer program product includes a computer program. When the computer program is run by a processor, the federated learning method of the spam detection model based on model splitting and cooperation described in any one of claims 1 to 6 is executed.