Mail classification method based on cascade clustering high-dimensional multi-modal feature selection
Through the combination of cascade clustering and multi-sub population framework, multiple equivalent feature subsets are found, which solves the problem of degradation of email classification accuracy under high-dimensional features in the existing technology, and achieves more accurate and flexible email classification.
Patent Information
- Application Number
- CN202510193233.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-23
AI Technical Summary
Existing email classification technology is difficult to find multiple subsets of equivalent feature under high-dimensional features, resulting in a decrease in classification accuracy.
Using a high-dimensional multimodal feature selection method based on cascade clustering, the population is initialized through Latin hypercube sampling, the guidance vector is updated, environmental selection and subpopulation division are performed, and multiple equivalent feature subsets are found.
It realizes finding multiple subsets of equivalent feature in high-dimensional feature cases, improving the accuracy of email classification and decision makers' choice space.
Smart Images

Figure CN120030430A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of big data analysis, and in particular to an email classification method based on cascade clustering high-dimensional multimodal feature selection. Background Art
[0002] Email classification, as a text classification task, contains a large number of features. It is expected that a relatively small number of features will achieve a higher classification accuracy. However, these two goals often conflict in practice, so email classification is a multi-objective optimization problem. Among the many features, there are many redundant or unimportant features, and it is necessary to select the most critical feature combination to achieve accurate email classification. In addition, email classification tasks often have multimodal characteristics, that is, the number of selected features and classification accuracy are the same, but the feature combinations are not the same.
[0003] Feature selection is an important preprocessing technique in machine learning and data mining. Its main purpose is to select the most representative features from high-dimensional data to improve the performance and interpretability of the model. High-dimensional data usually contains a large number of redundant or irrelevant features, which may lead to problems such as low computational efficiency and waste of resources. At present, many scholars have developed many intelligent optimization algorithms to solve combinatorial optimization problems such as feature selection.
[0004] The same number of selected features but different corresponding feature subsets often achieve similar or identical effects, so feature selection has multimodal characteristics. Different feature subsets often have huge differences in the cost consumed in practice, so it is necessary to find more feature subsets for decision makers to choose from and provide multiple alternatives so that decision makers can choose the most appropriate solution based on actual conditions. In addition, some feature selection problems contain a large number of features and have high-dimensional characteristics. It is necessary to develop a multimodal multi-objective optimization algorithm that can cope with such high-dimensional feature selection problems and find multiple equivalent feature subsets.
[0005] A single feature combination may cause the method to be overly dependent on a specific vocabulary combination, and the spam identification accuracy may drop significantly after the feature combination is detected.
[0006] Therefore, how to find a combination of multiple features to achieve more effective email classification has become a technical problem that needs to be solved urgently. Summary of the invention
[0007] The purpose of the present invention is to solve the defect in the existing email classification technology that it is impossible to find multiple equivalent feature subsets to achieve accurate email classification in the case of high-dimensional features, and to provide an email classification method based on cascade clustering high-dimensional multimodal feature selection to solve the above problem.
[0008] In order to achieve the above object, the technical solution of the present invention is as follows:
[0009] A method for email classification based on cascade clustering high-dimensional multimodal feature selection, comprising the following steps:
[0010] 11) Preprocessing of email data and construction of data set;
[0011] 12) Initialize the population using Latin hypercube sampling method according to the number of features in the email data;
[0012] 13) Update the guidance vector of each subpopulation;
[0013] 14) Complete the evaluation of the solution;
[0014] 15) Perform environmental selection on each subpopulation;
[0015] 16) Division of multiple subpopulations;
[0016] 17) Adjustment of sub-populations;
[0017] 18) Recording of equivalent feature subsets and classification of emails: Based on the features selected in the equivalent feature subsets, the corresponding words in the email text are obtained, and these word combinations are used to classify the emails to determine whether they are spam.
[0018] The preprocessing of the email data and the construction of the data set are as follows: obtaining the email data and preprocessing it, the preprocessing includes removing noise and word segmentation; converting the text data in the email into a numerical feature matrix and a label vector, using the processed email data as a data set for email classification and dividing it into a training set and a test set; including the following steps:
[0019] 21) Perform preprocessing on the acquired email data, including noise removal and word segmentation. In the noise removal operation, regular expressions are used to remove HTML / XML tags, numbers and dates, extra spaces and special characters, and then stop words are removed. For English characters, the upper and lower cases are unified, spelling errors are corrected, and stems are extracted. In the word segmentation process, the Chinese sentence is segmented into word sequences using a word segmentation tool, and the English words are segmented by spaces and punctuation. A vocabulary is established based on the processed data.
[0020] 22) Convert the text data in the email into a numerical feature matrix and a label vector, calculate the word frequency-inverse document frequency of each word in each email, combine them to construct a numerical feature matrix and add labels as a data set, and randomly divide the data set into 80% training set and 20% validation set.
[0021] The calculation of term frequency-inverse document frequency TF-IDF(t,d) is as follows:
[0022] TF-IDF(t,d)=TF(t,d)×IDF(t)
[0023] Where TF(t,d) represents the term frequency, that is, the frequency of word t appearing in email d, and IDF(t) represents the inverse document frequency, that is, the logarithm of the reciprocal number of emails containing word t, which is calculated as follows:
[0024]
[0025] Among them, Q represents the total number of emails, df t is the number of emails containing word t.
[0026] The method of initializing the population using the Latin hypercube sampling method according to the number of features in the email data includes the following steps:
[0027] 31) According to the number of words in the vocabulary of the email classification dataset, determine the number of features to be processed D, and set the dimension of the decision variable to D. Assume that each solution consists of two parts: a real vector and a binary vector. Each dimension of the real vector is a real number, and each dimension of the binary vector has only two cases: 0 or 1. If it is 1, it means that the feature represented by the dimension is selected. The binary vector is used for fast dimensionality reduction.
[0028] Suppose the solution x is constructed as follows:
[0029]
[0030] Among them, bin represents the binary vector of solution x, dec represents the real vector of solution x, D represents the dimension of decision variable, bin represents the binary vector of solution x, dec represents the real vector of solution x, D represents the dimension of decision variable, D is the Dth element of the binary vector, dec D is the D-th element of the real vector;
[0031] 32) The initial population obtains N binary vectors and real vectors of solutions through Latin hypercube sampling. The dimension of each solution vector is D. The Latin hypercube sampling is as follows:
[0032] X i,j =(rand(i)+π(j)) / N,
[0033] Among them, i represents the index of the sample, ranging from 1 to N, N represents the size of the population, j represents the index of the dimension, ranging from 1 to D, D represents the decision variable dimension, X i,j The value of the i-th sample in the j-th dimension, rand(i) is a random number uniformly distributed in the interval [0,1), and π(j) is a random permutation of the interval number set {1,2,…,N};
[0034] 33) Since the result of Latin hypercube sampling is a real number, the binary vector of the solution is discretized. If the dimension is not less than 0.5, it is set to 1, otherwise it is set to 0, as shown in the formula below:
[0035]
[0036] H(x) is a step function. If x is less than 0, H(x)=0, otherwise it is 1. H(x-0.5) means that if x is less than 0.5, H(x)=0, otherwise it is 1.
[0037] The updating of each subpopulation guidance vector is as follows: updating the subpopulation guidance vector according to the population history information and the sparse distribution of the current non-dominated solution; including the following steps:
[0038] 41) Assume that the guidance vector shows the evolution direction of each subpopulation. The guidance vector of each subpopulation is composed of the historical guidance vector information of the population and the sparse distribution information of the current non-dominated solution. The formula is as follows:
[0039]
[0040] Among them, lv i represents the historical guidance vector of the i-th subpopulation, |R| represents the number of non-dominated solutions in the current subpopulation, and V represents the sparse distribution of non-dominated solutions in the current subpopulation;
[0041] 42) Set the composition of V as follows:
[0042]
[0043] Among them, v i represents the i-th element of V, and and Respectively represent the i-th element of the binary vector of solution X and its neighbor Y.
[0044] The evaluation of the completed solution is as follows: discretizing each solution, obtaining a feature subset and classification accuracy, and recording the words in the email text corresponding to the selected features in the feature subset; including the following steps:
[0045] 51) Acquisition of feature subsets:
[0046] Discretize the real vector of each solution. If the dimension is not less than 0.5, set it to 1, indicating that the feature is selected. Otherwise, set it to 0, indicating that the feature is not selected. Obtain the feature subset of each solution and record the text words in the email data represented by each selected feature in the feature subset.
[0047] 52) Complete the solution evaluation:
[0048] The selected feature subset is sent to the KNN classifier, and the distance between all solutions in the test set and all solutions in the training set is calculated. The three solutions with the closest distance are found to vote to determine the category of each solution in the test set, and the accuracy is calculated according to the actual category of each solution to obtain the classification accuracy, and the classification error rate is calculated based on this. The number of features and the classification error rate are selected from the feature subset as the two goals to be optimized in the target space to complete the evaluation of the solution. As a minimization problem, the smaller the number of features and the classification error rate, the better;
[0049] 53) For each sub-population, non-dominated sorting is performed, and the parent generation is selected using the binary bidding method and the child generation solution is generated through crossover mutation. The real number vector of the child generation solution is also discretized, and the parent generation and child generation solutions are merged.
[0050] The performing of environment selection for each subpopulation is: calculating the distance between each solution and the guidance vector, and retaining the solution with a small distance from the guidance vector in the environment selection; comprising the following steps:
[0051] 61) Set a zero vector r of length D. For each subpopulation, count the dimensions of its guidance vector greater than 0.5, and set these dimensions to 1 in r. The formula is as follows:
[0052]
[0053] Among them, lv i represents the historical guidance vector of the ith subpopulation, r i is the i-th element of vector r;
[0054] 62) Calculate the relationship between each solution and vector r in the subpopulation i The calculation of Hamming distance is as follows:
[0055]
[0056] Among them, x and y represent two different non-dominated solutions, and and Respectively represent the i-th element of the two solution binary vectors;
[0057] 63) As an environmental selection operation for each subpopulation, select i The number of solutions with the smallest distance is N / K.
[0058] Where N represents the population size and K represents the number of subpopulations;
[0059] The distance is used as the third target value of the objective function. The smaller the distance is, the closer it is to the guidance vector. Non-dominated sorting is performed and the crowding distance is calculated. N / K solutions are selected according to the non-dominated level and crowding distance.
[0060] The division of the multiple subpopulations is as follows: selecting a cluster center and distributing the remaining solutions according to the distance from the cluster center to form multiple subpopulations; including the following steps:
[0061] 71) Perform cluster center selection operations on each subpopulation in turn, calculate the mean of the Hamming distances from any non-dominated solution of each subpopulation to the remaining non-dominated solutions, sort the distances, and select half of the solutions with larger distances as cluster centers. The distance calculation of each solution is as follows:
[0062]
[0063] In each solution, x i With x j They represent two different non-dominated solutions from the same subpopulation, and |R| represents the number of non-dominated solutions in the subpopulation;
[0064] 72) Calculate the Hamming distance between the remaining non-cluster center solutions and each cluster center, and assign them to the cluster with the nearest cluster center;
[0065] 73) All solutions are allocated, and multiple sub-populations are finally formed.
[0066] The subpopulation adjustment is as follows: calculating the similarity between all subpopulations, merging similar subpopulations, updating the number of subpopulations, and adjusting the size of subpopulations; including the following steps:
[0067] 81) Calculate the similarity between any two sub-populations and merge all sub-populations with similarity greater than 0.5;
[0068] The calculation of the similarity between populations is achieved by calculating representative non-dominated solutions. The calculation of the similarity sim(x, y) between any two solutions is as follows:
[0069]
[0070] in, and Respectively represent the i-th element of any two solutions X and Y binary vectors, and D represents the decision variable dimension;
[0071] 82) Similarity between populations Select a non-dominated solution from each population as a representative and calculate the similarity between the two solutions. The similarity between sub-populations is calculated as follows:
[0072] sim(subP a ,subP b ) = sim(x,y)
[0073] Among them, X and Y represent representative non-dominated solutions from two different sub-populations, subP a With subP b They represent two different subpopulations;
[0074] 83) Count the number of sub-populations after similarity detection and merging, and update the number of sub-populations K;
[0075] At this time, the size of each subpopulation is inconsistent, so keep the size of each subpopulation at N / K. For each subpopulation, assume that the current number of solutions is T. If T is greater than N / K, then the number of solutions with low non-dominated level and large crowding distance, TN / K, is deleted through non-dominated sorting and crowding distance, so that the number of solutions remains at N / K; if T is less than N / K, then the number of solutions, N / KT, is regenerated and added to the subpopulation, so that the number of solutions remains at N / K.
[0076] The recording of the equivalent feature subset and the classification of the mails include the following steps:
[0077] 91) Determine whether the maximum number of iterations has been reached. If so, merge all subpopulations and output equivalent feature subsets, and end the iteration; if not, repeat steps 13)-17);
[0078] 92) Merge all solutions in the subpopulations, perform non-dominated sorting, and extract all non-dominated solutions;
[0079] 93) Recording of equivalent feature subsets and email classification: Extract the feature subsets of each non-dominated solution, remove duplicate feature subsets, and record the classification accuracy corresponding to feature subsets with different number of features; record all equivalent feature subsets, that is, the number of features selected is the same as the classification accuracy, but the features are not the same;
[0080] 94) According to the selected features in the equivalent feature subset, the words corresponding to the email text information are obtained, and the selected different text word combinations are used for email classification to distinguish whether it is spam.
[0081] A computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the email classification method based on cascade clustering high-dimensional multimodal feature selection is implemented.
[0082] Beneficial Effects
[0083] Compared with the prior art, the email classification method based on cascade clustering high-dimensional multimodal feature selection of the present invention can find multiple equivalent feature subsets, provide more selection space for decision makers, and realize accurate classification of emails. Specifically, it has the following advantages:
[0084] 1. The present invention uses a cascade clustering method and a multi-subpopulation framework to find and retain multiple equivalent feature subsets with the same number of features and the same accuracy, thereby ensuring the diversity of the final solution. Through the similarity detection mechanism, each subpopulation retains a different feature subset, and ultimately more equivalent feature subsets can be obtained, providing decision makers with multiple options for distinguishing whether it is spam.
[0085] 2. The present invention uses a guidance vector to represent the convergence direction of each subpopulation and integrates it into the environment selection so that each subpopulation accelerates convergence in different directions. The final feature subset is of better quality, that is, the number of selected features is smaller and the target effect is better. The solution is represented by double coding and discretized so that non-sparse positions can be quickly located, that is, which features are determined to be the key features that affect the accuracy of email classification, and fast convergence can be achieved in the face of high-dimensional data with a large number of features. BRIEF DESCRIPTION OF THE DRAWINGS
[0086] Figure 1 It is a method sequence diagram of the present invention. DETAILED DESCRIPTION
[0087] In order to have a further understanding and recognition of the structural features and the effects achieved by the present invention, a preferred embodiment and accompanying drawings are used for detailed description as follows:
[0088] like Figure 1 As shown, the email classification method based on cascade clustering high-dimensional multimodal feature selection described in the present invention can find multiple equivalent feature subsets for email classification, providing decision makers with multiple optional feature subset solutions, effectively solving the problem that traditional algorithms can only search for a fixed feature subset and it is difficult to find a feature subset with a small number of features and high classification accuracy when facing high-dimensional features with a large number of features. The method described in the present invention is difficult to deal with the increasingly updated and changing spam emails, and ensures a high accuracy and robustness, and includes the following steps:
[0089] The first step is to preprocess the email data and construct the data set: obtain the email data and preprocess it, which includes noise removal and word segmentation; convert the text data in the email into a numerical feature matrix and label vector, use the processed email data as the data set for email classification and divide it into a training set and a test set.
[0090] (1) Preprocess the obtained email data, including noise removal and word segmentation. In the noise removal operation, regular expressions are used to remove HTML / XML tags, numbers and dates, extra spaces and special characters, and then stop words are removed. For English characters, the upper and lower cases are unified, spelling errors are corrected, and stems are extracted. In the word segmentation process, the Chinese sentence is segmented into word sequences using a word segmentation tool, and the English words are segmented by spaces and punctuation. A vocabulary is established based on the processed data.
[0091] (2) Convert the text data in the email into a numerical feature matrix and a label vector, calculate the word frequency-inverse document frequency of each word in each email, combine them to construct a numerical feature matrix and add labels as a data set, and randomly divide the data set into 80% training set and 20% validation set.
[0092] The calculation of term frequency-inverse document frequency TF-IDF(t,d) is as follows:
[0093] TF-IDF(t,d)=TF(t,d)×IDF(t)
[0094] Where TF(t,d) represents the term frequency, that is, the frequency of word t appearing in email d, and IDF(t) represents the inverse document frequency, that is, the logarithm of the reciprocal number of emails containing word t, which is calculated as follows:
[0095]
[0096] Among them, Q represents the total number of emails, df t is the number of emails containing word t.
[0097] In the second step, the population is initialized using the Latin hypercube sampling method according to the number of features in the email data.
[0098] (1) According to the number of words in the vocabulary of the email classification dataset, determine the number of features to be processed D, and set the dimension of the decision variable to D. Assume that each solution consists of two parts: a real vector and a binary vector. Each dimension of the real vector is a real number, while each dimension of the binary vector has only two cases: 0 or 1. If it is 1, it means that the feature represented by the dimension is selected. The binary vector is used for fast dimensionality reduction.
[0099] Suppose the solution x is constructed as follows:
[0100] x = bin × dec
[0101] =(bin 1 ×dec 1 ,...,bin D ×dec D )
[0102] Among them, bin represents the binary vector of solution x, dec represents the real vector of solution x, D represents the dimension of decision variable, bin represents the binary vector of solution x, dec represents the real vector of solution x, D represents the dimension of decision variable, D is the Dth element of the binary vector, dec D is the Dth element of the real vector.
[0103] (2) The initial population obtains N binary vectors and real vectors of solutions through Latin hypercube sampling. The dimension of each solution vector is D. The Latin hypercube sampling is as follows:
[0104] X i,j =(rand(i)+π(j)) / N,
[0105] Among them, i represents the index of the sample, ranging from 1 to N, N represents the size of the population, j represents the index of the dimension, ranging from 1 to D, D represents the decision variable dimension, X i,j The value of the i-th sample in the j-th dimension, rand(i) is a random number uniformly distributed in the interval [0,1), and π(j) is a random permutation of the interval number set {1,2,…,N}.
[0106] (3) Since the result of Latin hypercube sampling is a real number, the binary vector of the solution is discretized. If the dimension is not less than 0.5, it is set to 1, otherwise it is set to 0, as shown in the following formula:
[0107]
[0108] H(x) is a step function. If x is less than 0, H(x)=0, otherwise it is 1. H(x-0.5) means that if x is less than 0.5, H(x)=0, otherwise it is 1.
[0109] The third step is to update the guidance vector of each sub-population: based on the historical information of the population and the sparse distribution of the current non-dominated solutions, update the guidance vector of the sub-population.
[0110] (1) Assume that the guidance vector represents the evolutionary direction of each subpopulation. The guidance vector of each subpopulation is composed of the historical guidance vector information of the population and the sparse distribution information of the current non-dominated solution. The formula is as follows:
[0111]
[0112] Among them, lv i represents the historical guidance vector of the i-th subpopulation, |R| represents the number of non-dominated solutions in the current subpopulation, and V represents the sparse distribution of non-dominated solutions in the current subpopulation.
[0113] (2) Set the composition of V as follows:
[0114]
[0115] Among them, v i represents the i-th element of V, and and Respectively represent the i-th element of the binary vector of solution X and its neighbor Y.
[0116] Step 4: Complete the evaluation of the solution: Since the above method is a continuous optimization process and cannot directly handle discrete feature selection problems such as email classification, it is impossible to compare the quality of the solution. It is necessary to discretize each solution to determine which features (words in the email data) are selected. According to the discretization results, obtain the feature subset and calculate the classification accuracy, record the words in the email text corresponding to the selected features in the feature subset, and use this feature combination for subsequent email classification.
[0117] (1) Acquisition of feature subsets:
[0118] The real number vector of each solution is discretized. If the dimension is not less than 0.5, it is set to 1, indicating that the feature is selected. Otherwise, it is set to 0, indicating that the feature is not selected. The feature subset of each solution is obtained and the text words in the email data represented by each selected feature in the feature subset are recorded.
[0119] (2) Complete the solution evaluation:
[0120] The selected feature subset is sent to the KNN classifier, and the distance between all solutions in the test set and all solutions in the training set is calculated. The three solutions with the closest distance are found to vote to determine the category of each solution in the test set, and the accuracy is calculated according to the actual category of each solution to obtain the classification accuracy, and the classification error rate is calculated based on this. The number of features and the classification error rate are selected from the feature subset as the two goals to be optimized in the target space to complete the evaluation of the solution. As a minimization problem, the smaller the number of features and the classification error rate, the better.
[0121] (3) For each sub-population, non-dominated sorting is performed, and the parent generation is selected using the binary bidding method. The child generation solution is generated through crossover mutation. The real number vector of the child generation solution is also discretized, and the parent generation and child generation solutions are merged.
[0122] The fifth step is to perform environmental selection on each subpopulation: calculate the distance between each solution and the guidance vector, and retain the solutions with a small distance to the guidance vector in the environmental selection.
[0123] (1) Set a zero vector r of length D. For each subpopulation, count the dimensions of its guidance vector greater than 0.5 and set these dimensions to 1. The formula is as follows:
[0124]
[0125] Among them, lv i represents the historical guidance vector of the i-th subpopulation, r i is the i-th element of vector r.
[0126] (2) Calculate each solution and vector r in the subpopulation i The calculation of Hamming distance is as follows:
[0127]
[0128] Among them, x and y represent two different non-dominated solutions, and and Respectively represent the i-th element of the two solution binary vectors.
[0129] (3) As an environmental selection operation for each subpopulation, select the species with r i The number of solutions with the smallest distance is N / K.
[0130] Where N represents the population size and K represents the number of subpopulations;
[0131] The distance is used as the third target value of the objective function. The smaller the distance is, the closer it is to the guidance vector. Non-dominated sorting is performed and the crowding distance is calculated. N / K solutions are selected according to the non-dominated level and crowding distance.
[0132] Step 6: Division of multiple sub-populations: Select the cluster center and distribute the remaining solutions according to the distance from the cluster center to form multiple sub-populations.
[0133] (1) Perform cluster center selection operations on each subpopulation in turn, calculate the mean of the Hamming distances from any non-dominated solution in each subpopulation to the remaining non-dominated solutions, sort the distances, and select half of the solutions with larger distances as cluster centers. The distance calculation of each solution is as follows:
[0134]
[0135] In each solution, x i With x j They represent two different non-dominated solutions from the same subpopulation, and |R| represents the number of non-dominated solutions in the subpopulation.
[0136] (2) Calculate the Hamming distance between the remaining non-cluster center solutions and each cluster center, and assign them to the cluster with the nearest cluster center.
[0137] (3) All solutions are allocated, and finally multiple sub-populations are formed.
[0138] The seventh step is to adjust the sub-populations: calculate the similarity between all sub-populations, merge similar sub-populations, update the number of sub-populations, and adjust the sub-population size.
[0139] (1) Calculate the similarity between any two sub-populations and merge all sub-populations with similarity greater than 0.5;
[0140] The calculation of the similarity between populations is achieved by calculating representative non-dominated solutions. The calculation of the similarity sim(x, y) between any two solutions is as follows:
[0141]
[0142] in, and They represent the i-th element of any two binary vectors of solutions X and Y, respectively, and D represents the dimension of the decision variable.
[0143] (2) Similarity between populations Select a non-dominated solution from each population as a representative and calculate the similarity between the two solutions. The similarity between sub-populations is calculated as follows:
[0144] sim(subP a ,subP b )=sim(x,y)
[0145] Among them, X and Y represent representative non-dominated solutions from two different sub-populations, subP a With subP b They represent two different subpopulations.
[0146] (3) Count the number of subpopulations after similarity detection and merging, and update the number of subpopulations K;
[0147] At this time, the size of each subpopulation is inconsistent, so the size of each subpopulation is kept at N / K. For each subpopulation, assume that the number of its current solutions is T. If T is greater than N / K, then the number of solutions with low non-dominated levels and large crowding distances (TN / K) is deleted through non-dominated sorting and crowding distance, so that the number of solutions remains at N / K. If T is less than N / K, then the number of solutions (N / KT) is regenerated and added to the subpopulation, so that the number of solutions remains at N / K.
[0148] Step 8. Recording of equivalent feature subsets and classification of emails: Based on the features selected in the equivalent feature subsets, the corresponding words in the email text are obtained, and these word combinations are used to classify the emails to determine whether they are spam.
[0149] (1) Determine whether the maximum number of iterations has been reached. If so, merge all subpopulations and output equivalent feature subsets, and end the iteration. If not, repeat steps 3 to 7.
[0150] (2) Merge the solutions in all subpopulations, perform non-dominated sorting, and extract all non-dominated solutions.
[0151] (3) Recording of equivalent feature subsets and email classification: Extract the feature subsets of each non-dominated solution, remove duplicate feature subsets, and record the classification accuracy corresponding to feature subsets with different numbers of features; record all equivalent feature subsets, that is, the number of features selected is the same as the classification accuracy, but the features are not the same;
[0152] (4) The words corresponding to the email text information are obtained according to the selected features in the equivalent feature subset, and the selected different text word combinations are used for email classification to determine whether it is spam.
[0153] Here, a computer-readable storage medium is also provided, on which a computer program is stored. When the computer program is executed by a processor, an email classification method based on cascade clustering high-dimensional multimodal feature selection can be implemented.
[0154] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. The present invention may be subject to various changes and improvements without departing from the spirit and scope of the present invention. These changes and improvements fall within the scope of the present invention. The scope of protection claimed by the present invention is defined by the attached claims and their equivalents.
Claims
1. A mail classification method based on cascade clustering high-dimensional multimodal feature selection, characterized in that: The following steps are involved: 11) Preprocessing of email data and construction of data set; 12) Initialize the population using Latin hypercube sampling method according to the number of features in the email data; 13) Update the guidance vector of each subpopulation; 14) Complete the evaluation of the solution; 15) Perform environmental selection on each subpopulation; 16) Division of multiple subpopulations; 17) Adjustment of sub-populations; 18) Recording of equivalent feature subsets and classification of emails: Based on the features selected in the equivalent feature subsets, the corresponding words in the email text are obtained, and these word combinations are used to classify the emails to determine whether they are spam.
2. The email classification method based on cascade clustering high-dimensional multimodal feature selection according to claim 1 is characterized in that: The preprocessing of the mail data and the construction of the data set are as follows: obtaining the mail data and preprocessing it, the preprocessing including noise removal and word segmentation; The text data in the email is converted into a numerical feature matrix and a label vector, and the processed email data is used as a dataset for email classification and divided into a training set and a test set; the following steps are included: 21) Perform preprocessing on the acquired email data, including noise removal and word segmentation. In the noise removal operation, regular expressions are used to remove HTML / XML tags, numbers and dates, extra spaces and special characters, and then stop words are removed. For English characters, the upper and lower cases are unified, spelling errors are corrected, and stems are extracted. In the word segmentation process, the Chinese sentence is segmented into word sequences using a word segmentation tool, and the English words are segmented by spaces and punctuation. A vocabulary is established based on the processed data. 22) Convert the text data in the email into a numerical feature matrix and a label vector, calculate the word frequency-inverse document frequency of each word in each email, combine them to construct a numerical feature matrix and add labels as a data set, and randomly divide the data set into 80% training set and 20% validation set. The calculation of term frequency-inverse document frequency TF-IDF(t,d) is as follows: TF-IDF(t,d)=TF(t,d)×IDF(t) Where TF(t,d) represents the term frequency, that is, the frequency of word t appearing in email d, and IDF(t) represents the inverse document frequency, that is, the logarithm of the reciprocal number of emails containing word t, which is calculated as follows: Among them, Q represents the total number of emails, df t is the number of emails containing word t.
3. The email classification method based on cascade clustering high-dimensional multimodal feature selection according to claim 1 is characterized in that: The method of initializing the population using the Latin hypercube sampling method according to the number of features in the email data includes the following steps: 31) According to the number of words in the vocabulary of the email classification dataset, determine the number of features to be processed D, and set the dimension of the decision variable to D. Assume that each solution consists of two parts: a real vector and a binary vector. Each dimension of the real vector is a real number, and each dimension of the binary vector has only two cases: 0 or 1. If it is 1, it means that the feature represented by the dimension is selected. The binary vector is used for fast dimensionality reduction. Suppose the solution x is constructed as follows: x = bin × dec =(bin1×dec1,...,bin D ×dec D ), Among them, bin represents the binary vector of solution x, dec represents the real vector of solution x, D represents the dimension of decision variable, bin represents the binary vector of solution x, dec represents the real vector of solution x, D represents the dimension of decision variable, D is the Dth element of the binary vector, dec D is the D-th element of the real vector; 32) The initial population obtains N binary vectors and real vectors of solutions through Latin hypercube sampling. The dimension of each solution vector is D. The Latin hypercube sampling is as follows: X i,j =(rand(i)+π(j)) / N, Among them, i represents the index of the sample, ranging from 1 to N, N represents the size of the population, j represents the index of the dimension, ranging from 1 to D, D represents the decision variable dimension, X i,j The value of the i-th sample in the j-th dimension, rand(i) is a random number uniformly distributed in the interval [0,1), and π(j) is a random permutation of the interval number set {1,2,…,N}; 33) Since the result of Latin hypercube sampling is a real number, the binary vector of the solution is discretized. If the dimension is not less than 0.5, it is set to 1, otherwise it is set to 0, as shown in the formula below: H(x) is a step function. If x is less than 0, H(x)=0, otherwise it is 1. H(x-0.5) means that if x is less than 0.5, H(x)=0, otherwise it is 1.
4. The email classification method based on cascade clustering high-dimensional multimodal feature selection according to claim 1 is characterized in that: The updating of each subpopulation guidance vector is as follows: updating the subpopulation guidance vector according to the population history information and the sparse distribution of the current non-dominated solution; including the following steps: 41) Assume that the guidance vector shows the evolution direction of each subpopulation. The guidance vector of each subpopulation is composed of the historical guidance vector information of the population and the sparse distribution information of the current non-dominated solution. The formula is as follows: Among them, lv i represents the historical guidance vector of the i-th subpopulation, |R| represents the number of non-dominated solutions in the current subpopulation, and V represents the sparse distribution of non-dominated solutions in the current subpopulation; 42) Set the composition of V as follows: Among them, v i represents the i-th element of V, and bin i x With bin i y Represent the i-th element of the binary vector of solution X and its neighbor Y respectively.
5. The email classification method based on cascade clustering high-dimensional multimodal feature selection according to claim 1 is characterized in that: The evaluation of the completed solution is as follows: discretizing each solution, obtaining a feature subset and classification accuracy, and recording the words in the email text corresponding to the selected features in the feature subset; The following steps are involved: 51) Acquisition of feature subsets: Discretize the real vector of each solution. If the dimension is not less than 0.5, set it to 1, indicating that the feature is selected. Otherwise, set it to 0, indicating that the feature is not selected. Obtain the feature subset of each solution and record the text words in the email data represented by each selected feature in the feature subset. 52) Complete the solution evaluation: The selected feature subset is sent to the KNN classifier, and the distance between all solutions in the test set and all solutions in the training set is calculated. The three solutions with the closest distance are found to vote to determine the category of each solution in the test set, and the accuracy is calculated according to the actual category of each solution to obtain the classification accuracy, and the classification error rate is calculated based on this. The number of features and the classification error rate are selected from the feature subset as the two goals to be optimized in the target space to complete the evaluation of the solution. As a minimization problem, the smaller the number of features and the classification error rate, the better; 53) For each sub-population, non-dominated sorting is performed, and the parent generation is selected using the binary bidding method and the child generation solution is generated through crossover mutation. The real number vector of the child generation solution is also discretized, and the parent generation and child generation solutions are merged.
6. The email classification method based on cascade clustering high-dimensional multimodal feature selection according to claim 1 is characterized in that: The performing of environment selection for each subpopulation is: calculating the distance between each solution and the guidance vector, and retaining the solution with a small distance from the guidance vector in the environment selection; comprising the following steps: 61) Set a zero vector r of length D. For each subpopulation, count the dimensions of its guidance vector greater than 0.5, and set these dimensions to 1 in r. The formula is as follows: Among them, lv i represents the historical guidance vector of the i-th subpopulation, r i is the i-th element of vector r; 62) Calculate the relationship between each solution and vector r in the subpopulation i The calculation of Hamming distance is as follows: Among them, x and y represent two different non-dominated solutions, and and Respectively represent the i-th element of the two solution binary vectors; 63) As an environmental selection operation for each subpopulation, select i The number of solutions with the smallest distance is N / K. Where N represents the population size and K represents the number of subpopulations; The distance is used as the third target value of the objective function. The smaller the distance is, the closer it is to the guidance vector. Non-dominated sorting is performed and the crowding distance is calculated. N / K solutions are selected according to the non-dominated level and crowding distance.
7. The email classification method based on cascade clustering high-dimensional multimodal feature selection according to claim 1 is characterized in that: The division of the multiple subpopulations is as follows: selecting a cluster center and distributing the remaining solutions according to the distance from the cluster center to form multiple subpopulations; including the following steps: 71) Perform cluster center selection operations on each subpopulation in turn, calculate the mean of the Hamming distances from any non-dominated solution of each subpopulation to the remaining non-dominated solutions, sort the distances, and select half of the solutions with larger distances as cluster centers. The distance calculation of each solution is as follows: In each solution, x i With x j They represent two different non-dominated solutions from the same subpopulation, and |R| represents the number of non-dominated solutions in the subpopulation; 72) Calculate the Hamming distance between the remaining non-cluster center solutions and each cluster center, and assign them to the cluster with the nearest cluster center; 73) All solutions are allocated, and multiple sub-populations are finally formed.
8. The email classification method based on cascade clustering high-dimensional multimodal feature selection according to claim 1 is characterized in that: The subpopulation adjustment is as follows: calculating the similarity between all subpopulations, merging similar subpopulations, updating the number of subpopulations, and adjusting the size of subpopulations; including the following steps: 81) Calculate the similarity between any two sub-populations and merge all sub-populations with similarity greater than 0.5; The calculation of the similarity between populations is achieved by calculating representative non-dominated solutions. The calculation of the similarity sim(x, y) between any two solutions is as follows: in, and Respectively represent the i-th element of any two solutions X and Y binary vectors, and D represents the decision variable dimension; 82) Similarity between populations Select a non-dominated solution from each population as a representative and calculate the similarity between the two solutions. The similarity between sub-populations is calculated as follows: sim(subP a ,subP b )=sim(x,y) Among them, X and Y represent representative non-dominated solutions from two different sub-populations, subP a With subP b They represent two different subpopulations; 83) Count the number of sub-populations after similarity detection and merging, and update the number of sub-populations K; At this time, the size of each subpopulation is inconsistent, so keep the size of each subpopulation at N / K. For each subpopulation, assume that the current number of solutions is T. If T is greater than N / K, then the number of solutions with low non-dominated level and large crowding distance, TN / K, is deleted through non-dominated sorting and crowding distance, so that the number of solutions remains at N / K; if T is less than N / K, then the number of solutions, N / KT, is regenerated and added to the subpopulation, so that the number of solutions remains at N / K.
9. The email classification method based on cascade clustering high-dimensional multimodal feature selection according to claim 1 is characterized in that: The recording of the equivalent feature subset and the classification of the mails include the following steps: 91) Determine whether the maximum number of iterations has been reached. If so, merge all subpopulations and output equivalent feature subsets, and end the iteration; if not, repeat steps 13)-17); 92) Merge all solutions in the subpopulations, perform non-dominated sorting, and extract all non-dominated solutions; 93) Recording of equivalent feature subsets and email classification: Extract the feature subsets of each non-dominated solution, remove duplicate feature subsets, and record the classification accuracy corresponding to feature subsets with different number of features; record all equivalent feature subsets, that is, the number of features selected is the same as the classification accuracy, but the features are not the same; 94) According to the selected features in the equivalent feature subset, the words corresponding to the email text information are obtained, and the selected different text word combinations are used for email classification to distinguish whether it is spam.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by the processor, the email classification method based on cascade clustering high-dimensional multimodal feature selection according to any one of claims 1 to 9 is implemented.
Citation Information
Patent Citations
Text feature extraction method, system and device based on feature coding
CN109977227A
Text sentiment classification method based on genetic algorithm
CN117668225A