A Distributed ADMM Spam Classification Method Based on MPI
By adopting the MPI-based distributed ADMM spam classification method in the big data environment, the text data is divided into multiple sub-problems and calculated in parallel, solving the problem of insufficient spam classification efficiency and accuracy in the big data environment, and achieving fast and efficient spam classification.
Patent Information
- Application Number
- CN202111477718.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-12-06
AI Technical Summary
In the big data environment, it is difficult for the existing technology to efficiently classify spam, especially when processing large amounts of data, the efficiency and accuracy of existing methods are insufficient.
The distributed ADMM spam classification method based on MPI is used to vectorize and segment the text data into multiple subproblems. Through the improved stochastic gradient descent algorithm, the distributed ADMM framework is optimized alternately until the global optimal consensus is reached.
It improves the efficiency and accuracy of spam classification, can quickly solve large-scale optimization and learning problems, and is suitable for spam classification tasks in big data scenarios.
Smart Images

Figure CN114154581B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of distributed machine learning, and in particular to a distributed ADMM spam classification method based on MPI. Background Art
[0002] The classification problem is a very important and universal problem faced by humans. It is a supervised learning problem of identifying which class a new instance belongs to based on a known training set. Correctly classifying things helps people understand the world and makes the chaotic real world organized. For example, automatic text classification is to automatically classify a large number of natural language texts according to certain topic categories, which is a very important problem in natural language processing; text classification is mainly applied to tasks such as information retrieval, machine translation, automatic summarization, information filtering, and email classification.
[0003] The Alternating Direction Method of Multipliers (ADMM) was first proposed by Stephen Boyd et al. in 2010. As a computational framework for solving optimization problems, it is applicable to solving distributed convex optimization problems. The ADMM algorithm provides the possibility for the efficient distributed solution of constrained optimization problems in machine learning. The original ADMM algorithm has been widely applied in fields such as statistical machine learning, data mining, and computer vision. As a powerful tool that can effectively coordinate the optimization of sub-global model variables among several nodes, ADMM plays a crucial role in distributed optimization and statistical learning and has received great attention from research scholars. Since its development to date, ADMM has been widely applied to fields such as machine learning, data mining, and signal processing.
[0004] MPI (Massage Passing Interface) is a cross-language communication protocol. Based on a message-passing programming model, it is used for communication between different processes in a single-machine mode and for communication between different machines in a cluster mode. It can be installed and run on Linux, Windows, and Mac OS. There are several versions such as MPICH and OpenMPI available, and it can be developed using multiple programming languages such as C++ and Python. The mpi4py library implemented by Python can be well combined with the numpy library to ensure the program running speed while achieving efficient development. MPI can implement various communication methods such as point-to-point communication, one-to-many broadcast communication, and many-to-one reduction. MPI can complete both blocking communication and non-blocking communication. Currently, it is widely applied in the field of distributed computing in academia. Summary of the Invention
[0005] The object of the present invention is to propose a distributed ADMM spam classification method based on MPI, which is used to solve the spam classification problem in the big data environment, and to propose a suitable distributed framework based on ADMM for this problem, so as to improve performance and reduce time.
[0006] The present invention divides the problem to be studied into several sub-problems that can be computed in parallel. Each sub-problem is solved by an improved stochastic gradient descent algorithm, and the distributed ADMM framework is used to alternately optimize and gradually reach the global optimal consensus.
[0007] The technical solution for implementing the present invention: A distributed ADMM spam classification method based on MPI includes the following steps:
[0008] Step 1: Vectorize the text data into a dataset in digital format;
[0009] Step 2: Divide the dataset into a training set and a test set, perform oversampling on the training set, and then divide it into several parts and save them on several slave nodes respectively;
[0010] Step 3: MPI executes the code on all nodes in parallel, and the slave nodes update the local model in parallel;
[0011] Step 4: The master node aggregates the local models of the slave nodes through the MPI reduction function;
[0012] Step 5: The master node updates the global model and distributes the global model to each slave node by using the MPI broadcast function;
[0013] Step 6: Alternately update the models of the slave nodes and the master node in a loop until the termination condition is met;
[0014] Step 7: Save the global model of the master node as the classifier model;
[0015] Step 8: Use the trained classifier model to classify the test set and output the classification result.
[0016] Further, in step 1 of the present invention, NLP (Natural Language Processing) technology, such as the Word2vec algorithm, is used to vectorize the text data into a dataset in digital format. The processed dataset can be expressed as where n is the number of samples, x is the sample data vector of d dimensions, and y is the sample label. The present invention uses the L2 loss support vector machine (SVM) with L2 regularization as the main linear classification model. The objective function of this method can be expressed as:
[0017]
[0018] where \(C > 0\) is a hyperparameter used to control the ratio between the regularization term and the loss term, preventing overfitting. \(w\) is the parameter vector of the model, and \(w\in\mathbb{R}\). d ; L2 regularization means that the regularization term \(\|w\|^2\) in the above objective function 2 is squared, and L2 loss means that the loss term is \(\max(1 - y i w T x i , 0)\) 2 is squared.
[0019] Since spam belongs to the minority class in the dataset, the SMOTE algorithm is used to oversample the training set to solve the imbalanced data classification problem, making the number of positive and negative samples in the training set equal. Then it is divided into several parts and saved on several slave nodes respectively. The basic idea of the SMOTE algorithm is to analyze the minority class samples and artificially synthesize new samples according to the minority class samples and add them to the dataset.
[0020] Furthermore, the dataset is divided into a training set and a test set in a ratio of 4:1 and divided into several parts and saved on several slave nodes respectively. Suppose the data is stored on \(m\) nodes (\(D_1, D_2, \ldots, D\) m ), then equation (9) is rewritten as:
[0021]
[0022] s.t. \(w\) j - \(z = 0, j = 1, \ldots, m\)
[0023] where \(\rho\) is a hyperparameter, \(w\) j is the local model parameter of the \(j\)-th node, \(z\) is the global variable updated on the master node, and \(z\in\mathbb{R}\) d .
[0024] Rewrite the original function (10) into the augmented Lagrangian form, that is, the dual problem of the original problem:
[0025]
[0026] where \(\theta\) j is the dual variable of the \(j\)-th node.
[0027] Furthermore, MPI executes the code on all nodes in parallel, which is completed through the mpiexec command. First, the global model variable \(z\) is randomly initialized on the master node, and the local model variable \(w\) j and its dual variable \(\theta\) j are randomly initialized on each slave node. For convenience, they are generally initialized to all zeros. Then the local model variables are updated in parallel on the slave nodes. According to the update rules of the ADMM algorithm, \(w\), \(z\), and \(\theta\) are iteratively updated according to the following formulas:
[0028]
[0029]
[0030]
[0031] where k is the number of iterations. Since the Lagrangian function L(w, z, θ) is decomposable with respect to w j the present invention solves problem (12) in parallel on each slave node to update the local model variable w j :
[0032]
[0033] Furthermore, the master node aggregates the local models of the slave nodes through the MPI reduction function. First, the present invention creates a communication sub-comm of MPI through the following code:
[0034] comm = MPI.COMM_WORLD
[0035] Then, the reduction function is implemented using the communication sub, which is expressed as follows:
[0036] comm.Reduce(sendbuf, recvbuf, Op, root)
[0037] where sendbuf represents the content sent by the slave node, and here specifically the local model variable and its dual variable recvbuf represents the variable used by the master node for reception, Op represents the specific reduction operation function. Here, the present invention uses the MPI.SUM function, indicating that the values of the slave nodes are added and then passed to the master node, and root represents the root node number. Here, the present invention passes in the master node number 0.
[0038] Furthermore, the master node collects and integrates the local model variables and their dual variables of each slave node, updates the global variables, and sends the updated values to each slave node. (13) can be written as a closed-form solution to update the global variable z:
[0039]
[0040] Then, the broadcast function is implemented using the communication sub, which is expressed as follows:
[0041] comm.Bcast(buf, root)
[0042] where buf represents the data to be broadcast. Here, the present invention passes in the global model variable z of the master node k+1, root represents the serial number of the root node. Here, the main node serial number 0 is passed into the present invention.
[0043] Further, the present invention makes θ j = ρu j , and the update formulas for w j , z, u j can be obtained:
[0044]
[0045]
[0046]
[0047] where u j replaces θ j to become the dual variable of the present invention and is iteratively updated on each slave node.
[0048] Further, the present invention uses the variance-reduced gradient descent algorithm to specifically solve the optimization problem in (9) on the slave nodes. The SVM loss function part is taken out separately and defined as:
[0049] f i (w) = max(1 - y i w T x i , 0) 2 (12)
[0050] In the k-th iteration, first calculate the average gradient of the samples
[0051]
[0052] where n j is the number of training set samples of the j-th slave node. Then let start the inner iteration T - 1 times. The one-step update formula for the variable in the t-th iteration is:
[0053]
[0054] where η is the learning rate. After completing T - 1 inner iterations, take the mean as the update value for the (k + 1)-th iteration:
[0055]
[0056] Further, the present invention updates the local model variable w j on the slave nodes with (15), and then the master node collects the local model variables w j and the dual variable u j, the master node updates the global variable z with (10), then the master node distributes the updated global model variable z to each slave node, and then the slave node updates the dual variable u with (11). j , continuously loop and execute the above steps until the termination condition is met, that is, the residual of the variable updated in two iterations is small enough, which indicates that the algorithm converges and stops iterating. For convenience, the present invention sets the maximum number of iterations K.
[0057] After steps 1 to 6, the model training stage of the present invention ends. In step 7, the global model variable z of the master node is saved as the classifier model.
[0058] Furthermore, step 8 is the testing stage. The trained classifier model is used to classify the test set. Let (x i , y i ) be the i-th sample data in the test set, x i ∈R d , y i ∈{-1, +1}, then the predicted value is:[[]]
[0059]
[0060] where z is the global variable output by the algorithm in the last iteration, and the superscript T of z T means transpose, and sign() is the sign function, that is:[[]]
[0061]
[0062] Since more attention is paid to the minority class in the imbalance problem, that is, whether the spam is classified correctly, so when the predicted value , this sample is determined to be a positive sample, that is, spam, otherwise it is determined to be a negative sample, that is, normal mail.
[0063] Compared with the prior art, the present invention has the following remarkable advantages: Most random ADMM methods can only achieve a convergence speed slower than O(1 / T) in general convex problems, where T is the number of iterations and O is a notation to describe the speed of the algorithm. The present invention can obtain the same convergence speed O(1 / T) as batch ADMM in general convex problems, and can be used for large-scale optimization and learning problems to quickly solve the spam classification problem. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] Figure 1 FIG. is a flowchart of a distributed ADMM spam classification method based on MPI provided by an embodiment of the present invention.
[0065] Figure 2 FIG. is a graph showing the change of experimental accuracy, recall rate, and F1 score with the model training time provided by an embodiment of the present invention. Detailed implementation mode
[0066] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0067] The present invention discloses a distributed ADMM spam classification method based on MPI. The present invention explores the distributed ADMM framework, divides the distributed classification problem into some small problems, and adopts a master-slave distributed architecture. These small problems can be solved in parallel by dispersing resources. The present invention includes the following steps: vectorizing text data into a dataset in digital format; dividing the dataset into a training set and a test set, oversampling the training set, and then dividing it into several parts and storing them on several slave nodes respectively; MPI parallelly executing the code on all nodes, and the slave nodes parallelly updating the local models; the master node summarizing the local models of the slave nodes through the MPI reduction function; the master node updating the global model and distributing the global model to each slave node by using the MPI broadcast function; alternately updating the models of the slave nodes and the master node in a loop until the termination condition is met; saving the global model of the master node as the classifier model; using the trained classifier model to classify the test set and outputting the classification result. The present invention uses the ADMM parallel method to solve the optimization problem under the MPI distributed framework, making the model training faster, suitable for spam classification tasks in big data scenarios, and effectively improving the efficiency and accuracy of classification.
[0068] As Figure 1 shown, the present invention provides a distributed ADMM spam classification method based on MPI, including the following steps:
[0069] Step 1: Vectorize text data into a dataset in digital format;
[0070] Step 2: Divide the dataset into a training set and a test set, oversample the training set, and then divide it into several parts and store them on several slave nodes respectively;
[0071] Step 3: MPI parallelly execute the code on all nodes, and the slave nodes parallelly update the local models;
[0072] Step 4: The master node summarizes the local models of the slave nodes through the MPI reduction function;
[0073] Step 5: The master node updates the global model and distributes the global model to each slave node by using the MPI broadcast function;
[0074] Step 6: Alternately update the models of the slave nodes and the master node in a loop until the termination condition is met;
[0075] Step 7: Save the global model of the master node as the classifier model;
[0076] Step 8: Classify the test set using the trained classifier model and output the classification results.
[0077] To verify the present invention, the following experiments were conducted:
[0078] The present invention uses the unigram version of the public dataset webspam from the LibSVM website, which can be obtained from https: / / www.csie.ntu.edu.tw / ~cjlin / libsvmtools / datasets / binary.html to obtain this dataset. This dataset has been preprocessed, that is, as in Step 1, the text data is vectorized into a digital format dataset. After processing, it contains 350,000 samples, including 137,811 positive samples and 212,189 negative samples. Each sample contains 254 features, and each sample corresponds to a label value of +1 or -1. In the dataset preprocessing stage, the SMOTE library function of the third-party library imblearn is used to perform oversampling on the training set to make the number of positive and negative samples equivalent.
[0079] The experimental parameters are set as K = 100, T = n j , C = 10 -5 , ρ = 1, η = 4, m = 24. The present invention uses 24 linux servers, all of which are installed with MPICH and mpi4py libraries and developed using the python programming language.
[0080] Then, as in Step 2, the dataset is split into a training set and a test set in a ratio of 4:1 and split into several parts and saved on several slave nodes respectively. The above algorithm pseudocode is written as python code, and the file is named ADMM-MPI.py. Several copies of it are saved in the same directory of 24 servers. The following command is executed in the command line to execute in parallel:
[0081] mpiexec -n 24 python ADMM-MPI.py
[0082] Initialize the global model variable z on the master node 0 , and initialize the local model variables on each slave node and their dual variables Iteratively update each variable according to the pseudocode process of the algorithm, complete Steps 3 to 6. Finally, the present invention obtains the output z K As the classifier model parameters, on the test set, according to the formula:
[0083]
[0084] Predict and classify each test sample, and the predicted value If it is determined that the sample is a positive sample at this time, that is, it is spam, otherwise it is determined as a negative sample, that is, a normal email, and steps 7 and 8 are completed.
[0085] For the imbalanced data classification problem, the present invention uses accuracy, recall, and F1-score to measure the experimental effect. The definitions of the above evaluation indicators are given below: TP (True Positive) represents the number of positive classes predicted as positive classes, FN (False Negative) represents the number of positive classes predicted as negative classes, FP (False Positive) represents the number of negative classes predicted as positive classes, and TN (True Negative) represents the number of negative classes predicted as negative classes. Then the accuracy is defined as (TP + TN) / (TP + TN + FN + FP), which represents the proportion of the total number of correctly predicted cases in the total number of predictions. The recall is defined as TP / (TP + FN), which represents the proportion of positive examples in the sample that are correctly predicted. The precision is defined as TP / (TP + FP), which represents how many of the samples predicted as positive are truly positive samples. The F1-score is defined as 2 × precision × recall / (precision + recall), which represents the harmonic mean of precision and recall.
[0086] During the execution of the algorithm, continuously calculate the accuracy, recall, and F1-score of the classification model on the test set, as Figure 2 , the present invention obtains a curve graph of the experimental accuracy, recall, and F1-score changing with the model training time. It can be found that the invention proposed by the present invention has a relatively fast convergence speed, that is, it can reach stability quickly, and obtains an accuracy of 92.92%, a recall of 91.57%, and an F1-score of 90.44%, completing the spam classification task.
[0087] The pseudo-code summary of the specific algorithm process is as follows:
[0088]
[0089]
[0090] The present invention provides a distributed ADMM spam classification method based on MPI. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using existing technologies.
Claims
1. A distributed ADMM spam classification method based on MPI, characterized in that It includes the following steps: Step 1: Vectorize the text data into a dataset in digital format; Step 2: Split the dataset into a training set and a test set, perform oversampling on the training set, and then split it into several parts and save them on several slave nodes respectively; Step 3: Use MPI to execute the code on all nodes in parallel, and the slave nodes update the local model in parallel; Step 4: The master node aggregates the local models of the slave nodes through the MPI reduction function; Step 5: The master node updates the global model and distributes the global model to each slave node using the MPI broadcast function; Step 6: Alternately update the models of the slave nodes and the master node in a loop until the termination condition is met; Step 7: Save the global model of the master node as the classifier model; Step 8: Use the trained classifier model to classify the test set and output the classification results; Step 1: Use NLP technology to vectorize the text data into a dataset in digital format; the processed dataset is represented as x i ∈R d ,y i ∈{-1,+1}, where n is the number of samples, x i is the i-th d-dimensional sample data vector, y i is the i-th sample label, R d represents the d-dimensional real number set, and i ranges from 1 to n; use the L2-loss support vector machine SVM with L2 regularization as the linear classification model, and the objective function is expressed as: where \(C > 0\) is a hyperparameter used to control the ratio between the regularization term and the loss term, \(w\) is the variable of the classification model, and \(w\in\mathbb{R}\) d ; Use the SMOTE algorithm to perform oversampling on the training set so that the number of positive and negative samples in the training set is equal, and then split it into several parts and save them on several slave nodes respectively; Step 2 divides the dataset into a training set and a test set in a ratio of 4:1, and divides them into several parts and stores them on several slave nodes respectively. At the same time, copy the code files to several slave nodes. Suppose the data is stored on m nodes (D1, D2, …, D m ), Equation (1) is rewritten as: where ρ is a hyperparameter, w j is the local model variable of the j-th slave node, z is the global model variable updated on the master node, and z ∈ R d ; Rewrite Equation (2) into the augmented Lagrangian form to obtain Equation (3), that is: where θ j is the model dual variable of the j-th slave node; In Step 3, use MPI to execute the code on all nodes in parallel, which is completed through the mpiexec command; Randomly initialize the global model variable z on the master node, and randomly initialize the local model variable w on each slave node j and its dual variable θ j , initialized to all zeros; the slave nodes update the local model variables in parallel, obtained by the ADMM algorithm update rule, and w, z, θ are iteratively updated according to the following formula: where k is the number of iterations, and the Lagrangian function L(w, z, θ) is decomposable with respect to w j Solve equation (4) in parallel on each slave node to update the local model variable w j :
2. The distributed ADMM spam classification method based on MPI according to claim 1, wherein In Step 4, the master node aggregates the local models of the slave nodes through the MPI reduction function, and creates a communication sub-comm of MPI through the following code: comm = MPI.COMM_WORLD, where COMM_WORLD is a built-in object of MPI, and then use the communication sub to implement the reduction function, which is expressed as follows: comm.Reduce(sendbuf, recvbuf, Op, root), where Reduce is the reduction function, sendbuf represents the content sent by slave nodes, and the local model variables are passed in and their dual variables recvbuf represents the variable used by the master node for receiving, Op represents the specific reduction operation function. The MPI.SUM function is used, indicating that the values of the slave nodes are added and then passed to the master node. root represents the root node number, and the master node number 0 is passed in.
3. A distributed ADMM spam classification method based on MPI according to claim 2, characterized in that, In Step 5, the master node updates the global model and distributes the global model to each slave node using the MPI broadcast function; Formula (5) is written as a closed-form solution to update the global variable z: Then use the communication sub to implement the broadcast function, which is expressed as follows: comm.Bcast(buf,root) Among them, Bcast is the broadcast function, buf represents the data to be broadcast, and the global model variable z of the master node is passed in k+1 , root represents the root node number, and the master node number 0 is passed in.
4. The distributed ADMM spam classification method based on MPI according to claim 3, wherein, Let θ j = ρu j to obtain the update equations for j w, z, u j : where u j replaces θ j to become the dual variable and is iteratively updated at each slave node.
5. A distributed ADMM spam classification method based on MPI according to claim 4, characterized in that Use the variance-reduced gradient descent algorithm to optimize Formula (9) on the slave nodes, separately take out the SVM loss function part, and define: f i (w) = max(1 - y i w T x i , 0) 2 (12) In the k-th iteration, first calculate the average gradient of the samples where n j is the number of training set samples of the j-th slave node; then let Start the inner iteration for T - 1 times. The one-step update formula for the variable in the t-th iteration is as follows: where η is the learning rate, and take the mean value after T-1 inner iterations as the update value of the (k+1)-th iteration:
6. The distributed ADMM spam classification method based on MPI according to claim 5, characterized in that Update the local model variable w at the slave node using formula (15) j , then the master node collects the local model variables w j and the dual variable u j . The master node updates the global variable z using formula (10), then the master node distributes the updated global model variable z to each slave node, and then the slave node updates the dual variable u using formula (11) j . Continuously loop through the above steps until the termination condition is met, that is, the residual of the variable updated in two iterations is small enough, indicating that the algorithm converges, stop the iteration, and set the maximum number of iterations K; In Step 6, loop through Steps 3 to 5 to alternately update the model variables of the slave nodes and the master node until the termination condition is met, that is, the residual of the variable update in two iterations is small enough to converge, stop the iteration, and the model training stage ends. In Step 7, save the global model variable z of the master node as the classifier model.
7. A distributed ADMM spam classification method based on MPI according to claim 6, characterized in that Step 8 is the testing stage. The trained classifier model is used to classify the test set. Let (x i , y i ) be the i-th sample data in the test set, x i ∈ R d , y i ∈ {-1, +1}, then the predicted value is: where z is the global variable z output by the last iteration of the algorithm K , z T with superscript T denoting transpose, and sign() being the sign function, i.e.: Predicted value When the condition is met, the sample is determined to be a positive sample, that is, spam, otherwise it is determined to be a negative sample, that is, normal mail.
Citation Information
Patent Citations
ADMM-based unbalanced big data distributed classification method
CN113627485A