Improved GADT model assisted spam detection method based on genetic algorithm

Through the improved GADT model based on genetic algorithms, combined with text preprocessing, feature dimensionality reduction and decision tree optimization, the high-dimensional sparsity and semantic redundancy problems in spam detection are solved, achieving higher detection accuracy and efficiency, and adapting to multi-language and multi-scenario email data classification.

CN120450665AInactive Publication Date: 2025-08-08XIAMEN UNIV MALAYSIA BRANCH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510544911.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-08-08
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing spam detection technology faces the problems of high-dimensional spamness, semantic redundancy, empirical selection defects of decision tree pruning parameters, and the separation of feature dimensionality reduction and model optimization, which makes it difficult for the detection system to balance accuracy, robustness and computational efficiency.

Method used

The GADT model improved based on genetic algorithm is adopted to optimize the decision tree pruning parameters through text normalization, stop word deletion, stemming extraction, TF-IDF weighting, PCA dimensionality reduction and genetic algorithms, and build a GADT hybrid classifier to dynamically balance the complexity and generalization capabilities of the model.

Benefits of technology

It significantly improves the accuracy, robustness and computing efficiency of spam detection, can effectively deal with high-dimensional sparse characteristics and dynamic spam patterns, and has good scalability and application value.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450665A_ABST
    Figure CN120450665A_ABST
Patent Text Reader

Abstract

The invention discloses an improved GADT model auxiliary junk mail detection method based on a genetic algorithm, and relates to the technical field of junk mail detection and classification, and the method comprises the following steps: S10, carrying out structured preprocessing on an input text, including text standardization, stop word deletion and stem extraction; according to the method, feature space redundancy and noise are effectively reduced, key semantic information is reserved, and the data scale is compressed by performing structured preprocessing on the mail text and combining TF-IDF feature coding and PCA dimension reduction; a decision tree pruning parameter confidence factor is adaptively optimized by using a genetic algorithm, the complexity and generalization ability of the decision tree are dynamically balanced, and the classification accuracy and the model robustness are remarkably improved; the feature dimension reduction and model optimization cooperate to reduce the training reasoning complexity and improve the detection real-time performance, and the method is significantly superior to the prior art in accuracy, robustness and calculation efficiency, and has good expansibility and application and popularization value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of spam detection and classification, and in particular to a GADT model-assisted spam detection method based on an improved genetic algorithm. Background Art

[0002] With the widespread adoption of internet communication technologies, email systems have become a core means of information exchange. However, the prevalence of spam poses a serious threat to network resource allocation, information security, and user experience. Spam often contains malicious code, phishing links, or false information. Its spread not only consumes server storage and bandwidth resources but can also lead to financial losses or privacy breaches for users. Therefore, building an efficient spam detection system has become a key technical requirement in the field of information security.

[0003] Current mainstream detection methods include rule-based blacklist / whitelist mechanisms, content-based filtering techniques, and machine learning classification algorithms. Among these, machine learning methods (such as Naive Bayes, support vector machines, and decision tree models) have become a hot topic of research and application due to their data-driven adaptive capabilities. However, in actual engineering applications, these methods face the following specific technical bottlenecks:

[0004] (1) High-dimensional sparsity and semantic redundancy issues in text feature processing:

[0005] In content-based detection scenarios, email text needs to be converted into a numerical feature matrix through word segmentation and vectorization. Although traditional TF-IDF (term frequency-inverse document frequency) encoding can quantify the importance of terms, the feature space generated after direct processing has significant high-dimensional sparsity. Taking English emails as an example, the vocabulary size after word segmentation can usually reach tens of thousands of dimensions, while the proportion of non-zero features of a single email is often less than 1%. High-dimensional sparse features lead to a surge in computational complexity during model training (for example, the time and space complexity of matrix operations are O(n) / ... 2 ) growth), and it is easy to cause the "curse of dimensionality", which makes the classifier overfit in small sample scenarios.

[0006] In addition, there are technical gaps in the existing preprocessing process for handling semantic redundant features:

[0007] Incomplete filtering of short words and stop words: Without N-chars filtering, meaningless words of length 1-2 (such as "a" and "an") remain in the feature space. These words can account for 15%-20% of the total, but their contribution to classification approaches zero.

[0008] Lack of stem normalization: When stem extraction is not performed, different grammatical forms of the same root word (such as "connect", "connected", and "connecting") are regarded as independent features, resulting in an increase of about 20%-30% in feature dimension redundancy and fragmented semantic information.

[0009] (2) Defects in the empirical selection of pruning parameters for decision tree models:

[0010] Decision tree classifiers (such as the J-48 algorithm) control tree structure complexity through a pruning parameter called the "Confidence Factor," which determines the statistical significance threshold for subtree pruning. Existing technologies typically use manually preset confidence factors (e.g., a default value of 0.25) or simple grid searches (search step size ≥ 0.1), lacking a data-driven global optimization mechanism. Specific technical pain points are as follows:

[0011] Fixed parameters cannot adapt to dynamic feature distributions: Spam features are time-sensitive (for example, the high-frequency words in promotional spam vary with the seasons). When the confidence factor is fixed, it is difficult for the decision tree to dynamically adjust the degree of pruning. For example, in scenarios where new spam uses low-frequency words to evade detection, a too small confidence factor (such as 0.1) will retain a large number of irrelevant branches, causing the model to memorize training set noise. A too large confidence factor (such as 0.9) will lead to excessive pruning and miss samples with new feature patterns. Real-world measurements show that the accuracy variance across batches of datasets can exceed 15%.

[0012] Local optimal solution problem: Simple grid search is limited by the preset parameter range (usually [0.05, 0.95]) and step size, and cannot perform refined search in the continuous parameter space. It is easy to fall into local optimality, resulting in a non-optimal decision tree structure (such as the presence of redundant branches or the pruning of key discriminant nodes).

[0013] (3) Technical separation between feature dimensionality reduction and model optimization:

[0014] In existing technologies, feature dimensionality reduction and classifier parameter optimization are usually implemented independently, lacking a collaborative optimization mechanism:

[0015] Semantic information loss due to PCA dimensionality reduction: Traditional principal component analysis projects features solely based on data variance contribution, failing to consider the business needs of spam detection. This can result in the filtering of key semantic features. For example, after PCA processing that retains 95% of the cumulative variance, some low-frequency but highly discriminative lexical features are discarded due to their low variance contribution, resulting in a 5%-8% decrease in classification accuracy.

[0016] The contradiction between computational efficiency and accuracy: When high-dimensional features are directly input into the classifier, the training time increases exponentially with the dimension (for example, the training time complexity of SVM is O(n 3)), while the existing dimensionality reduction methods compress the dimensions, they do not simultaneously optimize the classifier parameters, resulting in the risk of underfitting the model in low-dimensional space.

[0017] The core flaws of existing technologies are: the text preprocessing process lacks refined processing of spam semantic features (such as ineffective filtering of short words and failure to perform stem normalization), decision tree pruning parameters rely on manual experience and the optimization mechanism is inefficient, and feature dimensionality reduction and classifier parameter optimization are independent of each other and lack coordination. These problems are superimposed on each other, making it difficult for the detection system to achieve a balance between accuracy, robustness, and computational efficiency when faced with high-dimensional sparse features and dynamically changing spam patterns. Therefore, there is an urgent need for a new detection method that can systematically solve the above technical bottlenecks to meet the high-performance requirements for spam detection in practical applications.

[0018] In view of this, a GADT model-assisted spam detection method based on an improved genetic algorithm is provided to overcome the above problems. Summary of the Invention

[0019] The purpose of the present invention is to provide a GADT model-assisted spam detection method based on an improved genetic algorithm to solve the problems raised in the above background technology.

[0020] To solve the above technical problems, the present invention provides a GADT model-assisted spam detection method based on an improved genetic algorithm, comprising the following steps:

[0021] S10: Perform structural preprocessing on the input text, including text normalization, stop word removal, and stemming;

[0022] S20: Use the TF-IDF method to perform feature representation and weighted calculation on the preprocessed text to generate a feature matrix;

[0023] S30: reducing the dimension of the feature matrix by principal component analysis, retaining the principal components of the first 95% of the cumulative variance information, and forming a low-dimensional feature matrix;

[0024] S40: Use genetic algorithm to adaptively optimize the confidence factor of decision tree pruning parameters and construct GADT hybrid classifier;

[0025] S50: Input the low-dimensional feature matrix after dimensionality reduction into the GADT hybrid classifier for spam detection, and output the classification result.

[0026] Furthermore, in step S10, text normalization specifically includes:

[0027] Remove punctuation from the text and convert it to lowercase. Use the N-chars filter to remove words that contain fewer than N characters.

[0028] Furthermore, in step S10, stem extraction uses the PorterStemmer algorithm to remove prefixes and suffixes of words to restore the root form.

[0029] Furthermore, in step S20, the TF-IDF weight calculation formula is:

[0030]

[0031] Among them, TF(t,d) represents the number of occurrences of term t in document d, |D| is the total number of documents in the set, and DF(t) is the number of documents containing term t.

[0032] Furthermore, in step S40, the process of optimizing the decision tree parameters using the genetic algorithm includes:

[0033] The classification accuracy of the training set is used as the fitness function to initialize the population size and the maximum evolutionary generation, and randomly generate an initial confidence factor population in the range of [0.05, 0.95].

[0034] Individuals were screened using a roulette wheel selection mechanism, offspring were generated using the BLX-α crossover method, and confidence factor values of some individuals were perturbed using mutation probabilities;

[0035] The population is iteratively updated until the maximum evolutionary generation is reached or the fitness converges, and the optimal confidence factor is obtained to configure the J-48 decision tree classifier.

[0036] Furthermore, the calculation formula of the BLX-α crossover method is:

[0037] offspring=(1-α)·parent1+α·parent2;

[0038] where parent 1 and parent 2 They represent two parent chromosome individuals respectively, α is the crossover coefficient, and its value range is 0.1 to 0.9.

[0039] Furthermore, in step S40, K-Fold cross validation is used to evaluate the performance of the GADT hybrid classifier. The data set is divided into K folds, with one fold used as the validation set and the rest as the training set each time. The average of the K evaluation results is finally taken.

[0040] Furthermore, in step S50, the email text to be detected is subjected to the same standardization preprocessing, TF-IDF weighting and PCA dimension reduction processing as in the training phase, and then input into the GADT hybrid classifier for classification.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] This invention significantly improves spam detection performance through a systematic solution of "feature preprocessing-dimensionality reduction-model optimization". The specific advantages are as follows:

[0043] 1. Multi-stage feature processing improves data quality:

[0044] Structural preprocessing: Through text standardization (punctuation removal, lowercase unification, N-chars filtering), stop word removal, and PorterStemmer stemming, low-frequency noise and grammatical redundancy are removed, and semantically similar words are normalized (for example, words in the "connect" family are unified into roots). This reduces data sparsity while focusing on key semantic features, providing concise and efficient input for subsequent analysis.

[0045] TF-IDF weighting and PCA dimensionality reduction: TF-IDF is used to suppress high-frequency, low-discrimination words, highlight discriminative keywords, and generate a sparse feature matrix. PCA is further used to extract the principal components of the first 95% of the cumulative variance, compressing high-dimensional features into a low-dimensional dense space. This method filters out redundant dimensions while retaining core information, making the feature space more suitable for the classification model and significantly reducing computational complexity and the risk of overfitting.

[0046] 2. Genetic algorithm driven decision tree optimization:

[0047] Adaptive parameter tuning: For the "confidence factor," a key parameter in decision tree pruning, a genetic algorithm is used with classification accuracy as the fitness function. Through roulette wheel selection, BLX-α crossover, and mutation operations, a global search for the optimal value is performed within the range of [0.05, 0.95]. This dynamically balances the complexity and generalization ability of the decision tree, avoiding the blindness of traditional manual parameter tuning and enabling the model to achieve the optimal configuration of pruning levels on different datasets.

[0048] Hybrid model synergy: The optimized confidence factor is configured in the J-48 decision tree to build a GADT hybrid classifier. K-Fold cross-validation is combined with stability assessment to ensure the robustness of the model in complex feature scenarios, fundamentally solving the core problem of overfitting or underfitting of traditional decision trees.

[0049] 3. Comprehensive performance breakthrough and application value:

[0050] Improved detection efficiency: Multi-dimensional technological innovations work together to address the high-dimensional sparsity and feature variability of spam. Compared with traditional methods (such as Naive Bayes and unoptimized decision trees), classification accuracy and anti-interference capabilities are significantly improved, especially in complex email scenarios with a large number of invalid features.

[0051] Efficiency and Scalability: Feature dimensionality reduction and parameter optimization jointly reduce training and inference time, supporting real-time detection of large-scale email data. The method framework is universal and can adapt to email data preprocessing and classification in multiple languages and scenarios, providing a scalable and efficient solution for the information security field.

[0052] In summary, the present invention achieves triple breakthroughs in accuracy, efficiency, and robustness, providing an optimal solution for spam detection that combines theoretical breakthroughs with engineering value. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 This is a flow chart of a method for detecting spam using a GADT hybrid classifier in a GADT model assisted spam detection method based on an improved genetic algorithm according to the present invention;

[0054] Figure 2 This is a flow chart of a method for structured preprocessing of input text in a GADT model assisted spam detection method based on an improved genetic algorithm according to the present invention;

[0055] Figure 3 The present invention provides a flow chart of a method for optimizing decision tree model parameters using a genetic algorithm (GA) in a GADT model assisted spam detection method based on an improved genetic algorithm. DETAILED DESCRIPTION

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0057] See also Figure 1-Figure 3 , the present invention provides a technical solution:

[0058] See Figure 1-Figure 3 As shown, an embodiment of a GADT model-assisted spam detection method based on an improved genetic algorithm is shown:

[0059] S10: Perform structured preprocessing on the input text:

[0060] The representation of text data greatly affects the accuracy of the results. In an embodiment, the text analysis problem needs to be converted into a representation suitable for the applied method.

[0061] The method of structured preprocessing of input text is as follows (see Figure 2 ):

[0062] S101: Normalize input text:

[0063] In this embodiment, the email text is obtained and the input text is standardized.

[0064] The specific steps are as follows: remove all punctuation marks from the text data and convert the text data into lowercase to make it consistent in format and convenient for subsequent analysis; use N-chars character filtering to delete words containing fewer than N characters, remove meaningless or low-frequency parts that have little contribution to semantics, and reduce the sparsity of the data.

[0065] Standardization converts text into a unified format, simplifies subsequent processing, and improves the model's attention to text semantics.

[0066] S102: Delete stop words in the text:

[0067] Stop words are frequently used words in a language but contain limited semantic information. They include conjunctions, prepositions, and pronouns, such as "about," "above," "after," "until," and "again." They are often used to complete sentence structures and connect expressions. In this embodiment, the input text is scanned using a predefined stop word list. The text sentence sequence is traversed and the relevant stop words in the stop word list are deleted. This reduces noise interference, improves the model's ability to focus on key content, and reduces data dimensionality, improving model training efficiency.

[0068] S103: Extract word stems based on different grammatical forms:

[0069] According to the different grammatical forms of words, they are simplified to their root forms, and the basic forms of the words are retained, so that words with the same or similar semantics are grouped together. In this embodiment, the PorterStemmer algorithm is applied to traverse the input text, remove the suffixes or prefixes of the words, and restore them to fixed root forms. For example, words such as connection, connective, connected, connecting, and disconnect can be lemmatized into the word connect and used directly for subsequent analysis. Through stem extraction, the complexity of data dimensions and feature space can be effectively reduced, providing more concise input data for downstream tasks and improving the efficiency of semantic analysis.

[0070] S20: Introduce TF-IDF for feature representation and weighted calculation:

[0071] The representation of text data greatly affects the accuracy of the results. In the embodiment, the term frequency-inverse document frequency (TF-IDF) method is introduced to perform feature representation and weighted calculation on the text, and the text is converted into numerical features that can be processed by the classifier. The specific steps are: based on the preprocessed email text, the frequency of occurrence of each word in the current email is counted to obtain the term frequency (TF) of the word to measure the local importance of the word in the current email. Combined with the occurrence of each word in the entire training set email collection, the inverse document frequency (IDF) of the word is obtained to evaluate the rarity of the word in the global scope. Given a term t, a document d and a document set D, the TF-IDF weight of each word is calculated by the following formula:

[0072]

[0073] Here, TF(t,d) represents the number of occurrences of term t in document d, |D| is the total number of documents in the set, and DF(t) is the number of documents containing term t. TF-IDF weighting can suppress common but weakly discriminative high-frequency words in a text, highlighting discriminative keywords for classification, and generating a sparse and efficient feature matrix, providing optimized input for subsequent feature dimensionality reduction and classification model training.

[0074] S30: Use PCA to complete feature matrix dimensionality reduction:

[0075] In order to reduce the dimension of the feature space and reduce redundant information, the principal component analysis (PCA) method is used to reduce the dimension of the TF-IDF feature matrix. In the embodiment, the covariance analysis is performed on the TF-IDF weighted email text feature matrix to extract the orthogonal principal components of the first 95% of the cumulative variance information; the original high-dimensional sparse features are projected into a low-dimensional space to form a new low-dimensional feature representation Z = XV k , as the input of the subsequent GADT hybrid model.

[0076] Through linear transformation, PCA compresses high-dimensional and sparse TF-IDF text features into a low-dimensional dense matrix, effectively reducing the data size while retaining the main variance information in the data, filtering out irrelevant noise features, and effectively reducing the computational complexity and overfitting risk of subsequent classification models.

[0077] S40: Construct GADT hybrid classifier:

[0078] In this embodiment, a genetic algorithm (GA) is used to adaptively optimize the confidence factor (ConfidenceFactor) of the pruning parameter in the decision tree, thereby constructing a GADT hybrid classifier to obtain better email classification performance.

[0079] The specific method of optimizing the decision tree model parameters by genetic algorithm GA is as follows (see Figure 3 ):

[0080] In this embodiment, before building the decision tree model, the pruning degree is adaptively optimized by a genetic algorithm (GA) to find the confidence factor that best suits the current data set, thereby enhancing the robustness and generalization ability of the classifier.

[0081] The classification accuracy of the hybrid classifier on the training set is used as the fitness function:

[0082] f(x i )=Accuracy(x i );

[0083] The fitness function is used to measure the quality of each individual chromosome. The higher the fitness value, the higher the quality of the corresponding chromosome and the better the classification performance of the confidence factor.

[0084] Initialize the GA algorithm parameters. In this embodiment, it includes setting the maximum evolutionary generation gen max , the population size N, and the initial population is randomly generated using real number encoding within the feasible range. Each chromosome is represented by a real value in the range [0.05, 0.95], corresponding to the confidence factor of the decision tree (Confidence Factor) example:

[0085] In each generation of evolution, the classification performance of each individual is evaluated using the fitness function and the classification performance is calculated according to f(x i ) value continuously updates the historical optimal fitness of the individual chromosome and the corresponding parameter p i , as well as the global optimal fitness of the entire population and the corresponding parameter g. Use the RouletteWheelSelection mechanism according to the fitness f(x i )

[0086] The chromosome individuals in the population are selected, and the selection probability is calculated according to the following formula:

[0087]

[0088] Among them, f(x i ) is the fitness of the i-th individual, and N is the population size. Individuals with higher fitness have a greater probability of selection, thus retaining more high-quality solutions and ensuring that good characteristics are inherited in the population.

[0089] The selected chromosome individuals are crossovered in a random pairing manner, and the BLX-α crossover method is used to generate new offspring individuals, which are updated according to the following formula:

[0090] offspring=(1-α)·parent1+α·parent2;

[0091] where parent 1 and parent 2 where α represents two parent chromosome individuals, and α is the crossover coefficient, which in this embodiment is set in the range of 0.1 to 0.9. BLX-α crossover generates offspring by linearly interpolating between the two parent individuals at a certain ratio, which helps expand the search space, promotes global exploration capabilities, and prevents the population from converging to a local optimal solution early on.

[0092] To further increase population diversity and avoid falling into local optima, the implementation introduces a mutation operation. With a certain mutation probability, the confidence factor values of some individuals are randomly perturbed, causing chromosomes to shift slightly near their current positions. This destabilizes the current search state and improves the algorithm's ability to escape local extremes.

[0093] The new chromosome set generated by crossover or mutation is combined with some of the parent individuals with higher fitness to form a new population. This completes the population update and the new population enters the next evolutionary cycle, continuing with fitness evaluation, selection, crossover, and mutation. This evolutionary process is repeated until the maximum number of iterations is reached or the population fitness converges. If the global optimal fitness does not significantly improve over several consecutive generations, the optimal confidence factor parameter is obtained.

[0094] The model is trained using the indicators selected by the GA algorithm. In this embodiment, the optimal confidence factor x output by the genetic algorithm is * , used to configure the J-48 decision tree classifier and complete the construction of the GADT hybrid classifier. K-Fold cross-validation is used to evaluate the performance of the GADT hybrid classification model. For example, the dataset is divided into K folds of equal size. In each iteration, one fold is selected as the validation set, and the remaining K-1 folds are used as the training set. This step is repeated until each fold has been used as a validation set. The final K evaluation results are averaged to obtain the performance indicators of the final model, which comprehensively measures the model's generalization performance and stability.

[0095] S50: Email detection using GADT hybrid classifier:

[0096] The GADT hybrid classification method, based on a genetic algorithm-optimized decision tree, employs a trained classification model to detect spam based on email text features. In this example, the email text to be tested is imported and subjected to standardization preprocessing, TF-IDF weighting, and PCA dimensionality reduction to extract its text features. This feature data is then fed into a decision tree classifier configured with genetic algorithm-optimized parameters. Based on the decision rules generated by the GADT hybrid classifier, the resulting email classification is output, effectively distinguishing between spam and legitimate emails.

[0097] Summarize:

[0098] The present invention effectively reduces the redundancy and noise in the feature space and improves the quality of feature representation through systematic text preprocessing and TF-IDF feature encoding. It also introduces principal component analysis (PCA) to perform feature dimensionality reduction, reducing the dimension of the feature space while retaining key information and improving the adaptability of text features. A genetic algorithm is used to adaptively optimize the confidence factor of the decision tree and dynamically control the pruning process, significantly improving classification accuracy and model generalization capabilities. At the same time, feature dimensionality reduction and model optimization jointly reduce the computational complexity of the training and inference processes, improving the real-time and stability of spam detection. Compared with existing technologies, the present invention achieves significant improvements in accuracy, robustness, and computational efficiency, and has good scalability and application promotion value.

Claims

1. A GADT model-assisted spam detection method based on an improved genetic algorithm, characterized in that: The following steps are involved: S10: Perform structural preprocessing on the input text, including text normalization, stop word removal, and stemming; S20: Using TF- The IDF method performs feature representation and weighted calculation on the preprocessed text to generate a feature matrix; S30: reducing the dimension of the feature matrix by principal component analysis, retaining the principal components of the first 95% of the cumulative variance information, and forming a low-dimensional feature matrix; S40: Use genetic algorithm to adaptively optimize the confidence factor of decision tree pruning parameters and construct GADT hybrid classifier; S50: Input the low-dimensional feature matrix after dimensionality reduction into the GADT hybrid classifier for spam detection, and output the classification result.

2. The GADT model-assisted spam detection method based on an improved genetic algorithm according to claim 1, characterized in that: In step S10, text normalization specifically includes: Remove punctuation from the text and convert it to lowercase. Use the N-chars filter to remove words that contain fewer than N characters.

3. The GADT model-assisted spam detection method based on an improved genetic algorithm according to claim 1, characterized in that: In step S10, stem extraction uses the PorterStemmer algorithm to remove the prefix and suffix of the word to restore the root form.

4. The GADT model-assisted spam detection method based on an improved genetic algorithm according to claim 1, wherein: In step S20, the TF-IDF weight calculation formula is: Among them, TF(t,d) represents the number of occurrences of term t in document d, |D| is the total number of documents in the set, and DF(t) is the number of documents containing term t.

5. The GADT model-assisted spam detection method based on an improved genetic algorithm according to claim 1, wherein: In step S40, the process of optimizing the decision tree parameters using the genetic algorithm includes: The classification accuracy of the training set is used as the fitness function to initialize the population size and the maximum evolutionary generation, and randomly generate an initial confidence factor population in the range of [0.05, 0.95]. Individuals were screened using a roulette wheel selection mechanism, and BLX- The α-crossover method generates offspring and perturbs the confidence factor values of some individuals with the mutation probability; The population is iteratively updated until the maximum evolutionary generation is reached or the fitness converges, and the optimal confidence factor is obtained to configure the J-48 decision tree classifier.

6. The GADT model-assisted spam detection method based on an improved genetic algorithm according to claim 5, characterized in that: The calculation formula of the BLX-α crossover method is: offspring=(1-α)·parent1+α·parent2; where parent 1 and parent 2 Represents two parent chromosome individuals, α is the cross coefficient, ranging from 0.1 to 0.

9.

7. The GADT model-assisted spam detection method based on an improved genetic algorithm according to claim 1, characterized in that: In step S40, K- Fold cross-validation is used to evaluate the performance of the GADT hybrid classifier. The dataset is divided into K folds, with one fold used as the validation set and the rest as the training set. The average of the K evaluation results is finally taken.

8. As claimed in claim 1- The GADT model-assisted spam detection method based on an improved genetic algorithm according to any one of the preceding claims, characterized in that: In step S50 , the email text to be detected is subjected to the same standardization preprocessing, TF-IDF weighting and PCA dimension reduction processing as in the training phase, and then input into the GADT hybrid classifier for classification.