Phishing website detection method and system, terminal and storage medium

By combining CNN, Bi-LSTM, sparse attention mechanism and Stacking-XGBoost classifier, the problem of low detection accuracy in the existing technology is solved, and efficient and accurate detection of phishing websites is achieved.

CN120498727APending Publication Date: 2025-08-15NANCHANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510551752.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing phishing website detection methods have low accuracy and cannot effectively cope with the rapid changes in phishing websites.

Method used

By performing feature extraction of website sample data, feature extraction and context information processing using CNN and Bi-LSTM, combining the sparse attention mechanism and Stacking-XGBoost classifier, a meta-feature matrix is ​​constructed and dynamic attention weighted is performed, and finally a meta-learner is used for phishing website detection.

Benefits of technology

It significantly improves the accuracy of phishing website detection, enhances the model's ability to capture phishing website features, reduces interference with irrelevant features, and improves computing efficiency and detection stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120498727A_ABST
    Figure CN120498727A_ABST
Patent Text Reader

Abstract

The invention provides a phishing website detection method and system, a terminal and a storage medium, and the method comprises the steps: carrying out the feature extraction of website sample data, and obtaining sample local features; performing context information extraction on the sample local features to obtain sample context features, and performing sparse attention mechanism calculation on the sample context features to obtain attention sample features; constructing a meta-feature matrix according to the attention sample features and a base learner; performing dynamic attention weighting on the meta-feature matrix to obtain a meta-weighted matrix, and inputting the meta-weighted matrix into a meta-learner for training until the meta-learner converges; and inputting to-be-detected website data into the converged meta-learner for phishing detection to obtain a phishing website detection result. According to the embodiment of the invention, the dynamic attention weighting is carried out on the meta-feature matrix, so that the extraction of key features in the meta-feature matrix by the meta-learner is improved, the accuracy of the meta-learner is improved, and the accuracy of phishing website detection is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to a phishing website detection method, system, terminal and storage medium. Background Art

[0002] With the widespread use of the internet, network security issues are becoming more frequent, and the issue of network information and communication security is receiving increasing attention. Phishing websites are a serious threat to network security, not only causing personal information leakage and financial loss, but also potentially threatening national security and stability. Therefore, to improve network security and protect sensitive information and network services from attacks, phishing website detection is receiving increasing attention.

[0003] In the existing phishing website detection process, a blacklist is generally used for phishing website detection. However, due to the rapid changes in phishing websites, the accuracy of phishing website detection is low. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a phishing website detection method, system, terminal and storage medium to solve the problem of low phishing website detection accuracy in the prior art.

[0005] The embodiment of the present invention is implemented as follows: a method for detecting phishing websites, the method comprising: Obtaining website sample data, and performing feature extraction on the website sample data to obtain sample local features; Extracting context information from the sample local features to obtain sample context features, and performing sparse attention mechanism calculation on the sample context features to obtain attention sample features; Training a base learner according to the attention sample features to obtain sample prediction values, and constructing a meta-feature matrix according to the sample prediction values; Performing dynamic attention weighting on the meta-feature matrix to obtain a meta-weight matrix, and inputting the meta-weight matrix into a meta-learner for training until the meta-learner converges; The website data to be detected is input into the meta-learner after convergence to perform phishing detection to obtain phishing website detection results.

[0006] Preferably, the sample context features are calculated using a sparse attention mechanism to obtain attention sample features, including: Performing a linear transformation on the sample context feature to obtain a sample linear feature, and performing a nonlinear activation on the sample linear feature to obtain a nonlinear feature; Calculating an attention score of the nonlinear feature according to the learnable attention vector, and filtering the sample linear features according to the attention score to obtain a sample filtering feature; The attention score corresponding to the sample screening feature is normalized to obtain an attention weight, and the sample context feature is weightedly fused according to the attention weight to obtain the attention sample feature.

[0007] Preferably, training a base learner according to the attention sample features to obtain sample prediction values, and constructing a meta-feature matrix according to the sample prediction values, including: Dividing sample subsets according to the attention sample features, and constructing a sample validation set based on the sample subsets; Inputting the sample verification set into the base learner to perform sample prediction to obtain a sample prediction value, and performing vector splicing on the sample prediction value according to the sample index in the website sample data to obtain a sample prediction vector; The sample prediction vectors corresponding to different base learners are horizontally spliced by column to obtain the meta-feature matrix.

[0008] Preferably, dynamic attention weighting is performed on the meta-feature matrix to obtain a meta-weighted matrix, including: Calculating a classification performance index of the base learner according to the sample prediction value, and calculating a model weight of the base learner according to the classification performance index; Dynamic attention weighting is performed on the meta-feature matrix according to the model weight to obtain the meta-weighted matrix.

[0009] Preferably, the formula used to calculate the model weight of the base learner according to the classification performance index includes: Where τ represents the temperature coefficient, represents the model weight, Indicates the k The classification performance index of the base learner on the sample validation set, Indicates the m The classification performance index of the base learner, M represents the total number of base learners; The formula used to dynamically weight the meta-feature matrix according to the model weight includes: in, a k(j) Indicates the j The model weights of the base learners, Indicates thei The confidence weight of the sample, P i j Indicates the j The sample prediction value of the attention sample feature.

[0010] Preferably, the objective function of the meta-learner includes: Objective function in, n represents the number of samples, Indicates the i The true labels of samples, Indicates that the meta-learner is for i The predicted value of the sample, represents the loss function of the meta-learner, represents the regularization term.

[0011] Preferably, feature extraction is performed on the website sample data to obtain local features of the sample, including: Performing character splitting on the website sample data to obtain sample split characters, and performing word embedding processing on the sample split characters to obtain word embedding vectors; The word embedding vector is convolved to obtain a convolution feature, and the convolution feature is pooled to obtain the sample local feature.

[0012] Another object of an embodiment of the present invention is to provide a phishing website detection system, the system comprising: A feature extraction module is used to obtain website sample data and perform feature extraction on the website sample data to obtain sample local features; An attention mechanism module is used to extract context information from the local features of the sample to obtain sample context features, and perform sparse attention mechanism calculation on the sample context features to obtain attention sample features; A matrix construction module is used to train a base learner according to the attention sample features to obtain sample prediction values, and to construct a meta-feature matrix according to the sample prediction values; a training module, configured to perform dynamic attention weighting on the meta-feature matrix to obtain a meta-weight matrix, and input the meta-weight matrix into a meta-learner for training until the meta-learner converges; The phishing detection module is used to input the website data to be detected into the meta-learner after convergence to perform phishing detection and obtain phishing website detection results.

[0013] The embodiments of the present invention can effectively extract local features of samples in the website sample data by performing feature extraction on the website sample data, and can effectively capture complex phishing website features in the website sample data by performing context information extraction on the sample local features, so as to improve the detection accuracy of the meta-learner after training, thereby improving the accuracy of phishing website detection. By performing sparse attention mechanism calculation on the sample context features, the interference of irrelevant features is effectively reduced, and the computational efficiency of the base learner and the meta-learner is improved. By constructing a meta-feature matrix, the prediction results of multiple base learners can be effectively converted into high-order features for further fusion and optimization of the meta-learner. By performing dynamic attention weighting on the meta-feature matrix, a meta-weighted matrix is obtained to improve the meta-learner's extraction of key features in the meta-feature matrix, thereby improving the accuracy of the meta-learner after training, and thereby improving the accuracy of phishing website detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a flow chart of a phishing website detection method provided by the first embodiment of the present invention; Figure 2 1 is a schematic diagram of feature extraction based on deep learning of sparse attention mechanism provided by the first embodiment of the present invention; Figure 3 is a schematic diagram of a phishing website detection model provided by the first embodiment of the present invention; Figure 4 is a structural diagram of a phishing website detection system provided by a second embodiment of the present invention; Figure 5 It is a structural diagram of a terminal device provided by the third embodiment of the present invention. DETAILED DESCRIPTION

[0015] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0016] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.

[0017] Example 1 This embodiment provides an XGBoost phishing detection model based on the Stacking framework. The core concept of the phishing detection model is to use the Stacking framework to leverage the diversity of multiple basic learners, such as decision trees, random forests, LightGBM, and CatBoost, to integrate their prediction results, and use XGBoost as a meta-learner to further improve the performance of the phishing detection model.

[0018] See also Figures 1 to 3 , is a flow chart of a phishing website detection method provided by a first embodiment of the present invention. The phishing website detection method can be applied to any device or system. The phishing website detection method includes the following steps: Step S10, obtaining website sample data, and performing feature extraction on the website sample data to obtain sample local features; The website sample data includes Uniform Resource Locator (URL) information, and features of the website sample data can be extracted based on a CNN network.

[0019] Optionally, feature extraction is performed on the website sample data to obtain local features of the sample, including: Performing character splitting on the website sample data to obtain sample split characters, and performing word embedding processing on the sample split characters to obtain word embedding vectors; The URL is broken down into character-level sequences. For example, the URL "http: / / example.com" is broken down into ['h', 't', 't', 'p', ':', ' / ', ' / ', 'e', 'x', 'a', 'm', 'p', 'l', 'e', '.', 'c', 'o', 'm']. Each character is mapped to a dense vector, typically using word embeddings. Word embeddings represent each character as a low-dimensional, dense vector, helping to capture the semantic relationships between characters. This allows subsequent models to more effectively understand potential malicious patterns in URLs. Embeddings are learned during model training and are continuously optimized as training progresses. Convolving the word embedding vector to obtain convolution features, and performing pooling processing on the convolution features to obtain the sample local features; The character sequence processed with word embeddings is input into a convolutional neural network (CNN). The CNN uses convolutional layers to extract local features, such as common malicious substrings or character combinations in URLs. The convolution operation slides a convolution kernel over the input character sequence, scanning the word-embedded character sequence to extract local character combinations, which are then input into the CNN. The convolution layer extracts different features using multiple convolution kernels, each sized to extract features for substrings of varying lengths in the URL. The pooling layer processes the resulting feature map, reducing its dimensionality and preserving important information. Ultimately, the CNN outputs multiple feature maps (local features of the samples) to help capture important local patterns in the URL.

[0020] Specifically, the convolution operation of CNN can be expressed by the following formula: in, W j is the weight of the convolution kernel, u i + j −1 is the i + j −1 character, y i is the local feature of the sample, b is the bias term. CNN can identify characteristic patterns in URLs through convolution operations, which helps to distinguish phishing websites from legitimate websites.

[0021] Step S20, extracting context information from the sample local features to obtain sample context features, and performing sparse attention mechanism calculation on the sample context features to obtain attention sample features; Among them, a bidirectional long short-term memory network (Bi-LSTM) can be used to extract contextual information from local features of samples. In this step, although CNN can extract local features in URLs, it is not good at capturing long-distance dependencies in character sequences. Therefore, a bidirectional long short-term memory network (Bi-LSTM) is introduced. Bi-LSTM can process URL character sequences from two directions simultaneously, extracting contextual information from the forward and backward directions respectively. By considering the front and back information of the character sequence simultaneously, Bi-LSTM enables the model to more comprehensively understand the structure of the URL and potential attack characteristics. The output of Bi-LSTM is expressed by the following formula: h i =Bi-LSTM( u i ) in, h i It is i Bi-LSTM output of characters, u i It is i Bi-LSTM can effectively capture long-distance dependencies in URLs.

[0022] Preferably, this step also introduces a sparse attention mechanism to weight the features extracted by the CNN and Bi-LSTM. By limiting the computational scope and focusing only on the most relevant parts of the input data, the sparse attention mechanism effectively reduces interference from irrelevant features, lowers computational complexity, and improves model efficiency. The sparse attention mechanism automatically selects the most discriminative features, thereby improving the expressiveness of features and optimizing the model's detection performance.

[0023] Optionally, a sparse attention mechanism is performed on the sample context features to obtain attention sample features, including: Performing a linear transformation on the sample context feature to obtain a sample linear feature, and performing a nonlinear activation on the sample linear feature to obtain a nonlinear feature; Among them, the set of input sample context features {f1,f2,...,f n}, preliminarily encodes the sample context features through linear transformation and nonlinear activation operations. For each sample context feature, a linear mapping is performed using a learnable parameter matrix and bias term, and the nonlinear expression capability is enhanced through the hyperbolic tangent function tanh(⋅) to obtain a nonlinear feature; Calculating an attention score of the nonlinear feature according to the learnable attention vector, and filtering the sample linear features according to the attention score to obtain a sample filtering feature; The attention score of each nonlinear feature is calculated through a learnable attention vector. The attention score quantifies the importance of the feature to the classification task. In this step, only the first preset number of features with the highest attention scores or features exceeding the preset score threshold are retained to obtain the sample screening features. The preset number and preset score threshold can be set according to needs. Normalizing the attention score corresponding to the sample screening feature to obtain an attention weight, and weightedly fusing the sample context feature according to the attention weight to obtain the attention sample feature; The attention scores corresponding to the sample screening features are normalized using the Softmax function to generate attention weights, which are then weighted and fused with the sample context features. By dynamically focusing on key features, such as unusual character combinations in URLs, the model significantly improves its ability to detect dynamically rendered pages and adversarial attacks, reduces its computational complexity, and is suitable for processing long sequences of data.

[0024] In this step, the sparse attention mechanism automatically assigns a weight to each feature, allowing the model to focus on the most critical features while ignoring redundant or insignificant ones. By calculating the attention weights for each feature, the sparse attention mechanism allows the model to "focus" on the most important parts, thereby improving the model's classification ability. The features of the attended samples are then input into the fully connected layer for full connection processing and output.

[0025] Step S30, training a base learner according to the attention sample features to obtain sample prediction values, and constructing a meta-feature matrix according to the sample prediction values; Decision trees, random forests, LightGBM, and CatBoost were selected as base learners. Each base learner was trained independently, and prediction results were generated through 5-fold cross-validation. During the feature processing stage, the decision tree recursively selects the optimal splitting feature based on information gain or the Gini coefficient, constructs classification rules by discretizing continuous features such as URL length, and processing missing values. Its advantage lies in the intuitive assessment of feature importance. Random forests introduce feature random subset sampling and bagging strategies based on decision trees. The training set for each tree is generated through sampling with replacement, and candidate subsets are randomly selected from all features, effectively enhancing model diversity and reducing variance. Ultimately, importance is assessed by the average split contribution of features in the forest. LightGBM uses gradient unilateral sampling and mutually exclusive feature bundling technology to retain high gradient absolute value samples to accelerate training, merge sparse and mutually exclusive features to reduce dimensionality, and directly support categorical feature processing, avoiding the computational overhead of one-hot encoding. CatBoost processes categorical features through ordered target encoding, uses mean encoding of the target variable to prevent information leakage, and employs a symmetric tree structure to force nodes in each layer to split according to the same features. This, combined with automatically generated high-order feature combinations, enhances the model's expressiveness. After independently training the four base learners on the training set, predictions are generated through 5-fold cross-validation. This involves partitioning the training set into five subsets at each iteration, using one subset as the validation set and the remaining four subsets for training. This ensures that predictions are independent of a single data partition, thereby reducing the risk of overfitting.

[0026] Optionally, training a base learner according to the attention sample features to obtain sample prediction values, and constructing a meta-feature matrix according to the sample prediction values, including: Dividing sample subsets according to the attention sample features, and constructing a sample validation set based on the sample subsets; The cross-validation strategy divides the attention sample features into five sample subsets as the training set. Each sample subset is used as the validation set in turn, and the remaining four sample subsets are used to train the base learner to ensure that the prediction results of the base learner do not depend on a single data partition, thereby reducing the risk of overfitting. Inputting the sample verification set into the base learner to perform sample prediction to obtain a sample prediction value, and performing vector splicing on the sample prediction value according to the sample index in the website sample data to obtain a sample prediction vector; All validation set prediction results are concatenated into a complete probability vector based on the original sample indexes to obtain the sample prediction vector. For binary classification tasks, each base learner outputs the probability value of each sample belonging to the positive class, forming a vector of dimension N × 1 (N is the total number of training samples).

[0027] The sample prediction vectors corresponding to different base learners are horizontally spliced by column to obtain the meta-feature matrix; Among them, the sample prediction vectors corresponding to all base learners are spliced horizontally by column to form a meta-feature matrix. Each row of the meta-feature matrix corresponds to the multi-model joint representation of a sample, and each column corresponds to the prediction output of a single base learner.

[0028] In this embodiment, in the feature extraction stage, CNN+Bi-LSTM generates a set of highly comprehensive feature vectors after weighting by the sparse attention mechanism. In order to further enhance the classification effect, the stacking method is used to combine the prediction results of multiple basic learners into a meta-feature matrix.

[0029] The Stacking-XGBoost classifier is used to classify features. Stacking-XGBoost further improves classification performance by integrating the outputs of multiple base learners, enhancing the model's ability to detect phishing websites. The XGBoost algorithm itself has strong classification capabilities and resistance to overfitting, making it suitable for processing large and complex datasets.

[0030] The stacking approach combines the strengths of different models, enabling the meta-learner to effectively correct errors in base learners, thereby improving overall classification accuracy. In the stacking ensemble framework, the meta-feature matrix is constructed through a systematic cross-validation strategy. Its core goal is to transform the predictions of multiple base learners into high-level features for further fusion and optimization by the meta-learner.

[0031] Step S40, performing dynamic attention weighting on the meta-feature matrix to obtain a meta-weighted matrix, and inputting the meta-weighted matrix into a meta-learner for training until the meta-learner converges; Among them, after the base learners (decision tree, random forest, LightGBM, CatBoost) generate prediction probabilities through five-fold cross-validation, a dynamic attention weighting strategy is introduced to optimize the meta-feature expression from two dimensions: model performance and sample characteristics.

[0032] Optionally, dynamic attention weighting is performed on the meta-feature matrix to obtain a meta-weighted matrix, including: Calculating a classification performance index of the base learner according to the sample prediction value, and calculating a model weight of the base learner according to the classification performance index; Dynamic attention weighting is performed on the meta-feature matrix according to the model weight to obtain the meta-weighted matrix.

[0033] Furthermore, the formula used to calculate the model weight of the base learner according to the classification performance index includes: Among them, τ represents the temperature coefficient, which is used to smooth the weight distribution and ensure that the contribution of the high-stability base learner is significantly improved. represents the model weight, Indicates the k The classification performance index of the base learner on the sample validation set, Indicates the m The classification performance index of the base learner, M represents the total number of base learners; The formula used to dynamically weight the meta-feature matrix according to the model weight includes: in, a k(j) Indicates the j The model weights of the base learners, Indicates the i The confidence weight of the sample, P i j Indicates the j The sample prediction value of the attention sample feature is calculated. By adjusting the weights, the meta-learner can focus on efficient models and high-confidence samples, improving feature discriminability.

[0034] XGBoost's objective function consists of two parts: a loss function and a regularization term. The loss function measures the difference between the model's predictions and the true labels, while the regularization term controls the model's complexity and prevents overfitting. XGBoost's goal is to optimize the model by minimizing this objective function.

[0035] Preferably, the objective function of the meta-learner includes: Objective function in, n represents the number of samples, Indicates the i The true labels of samples, Indicates that the meta-learner is for i The predicted value of the sample, represents the loss function of the meta-learner, represents the regularization term.

[0036] The meta-learner uses the XGBoost algorithm, whose input is the meta-weight matrix P and output is the final classification result. XGBoost gradually corrects the prediction error of the base learner through the gradient boosting tree and introduces a regularization term to control the model complexity.

[0037] Step S50, inputting the website data to be detected into the meta-learner after convergence to perform phishing detection and obtain phishing website detection results; Among them, the data of the website to be detected is input into the CNN+Bi-LSTM based on the sparse attention mechanism for feature extraction to obtain the features to be detected, and the features to be detected are input into the converged meta-learner for classification to obtain the phishing website detection results.

[0038] The phishing website detection model (SCB-SXGBoost model) combines the CNN+Bi-LSTM feature extraction method with a sparse attention mechanism and the Stacking-XGBoost classifier. The goal of SCB-SXGBoost is to improve the accuracy and robustness of phishing website detection tasks through effective feature extraction and model fusion. Figure 3 The framework consists of data preprocessing, feature extraction, a sparse attention mechanism, and a Stacking-XGBoost classifier. The data preprocessing module cleans and formats the input URL data for subsequent processing. The feature extraction module extracts deep features using convolutional neural networks (CNNs) and bidirectional long short-term memory (Bi-LSTM) networks. The sparse attention mechanism weights the extracted features, selectively focusing on the most relevant ones and reducing interference from irrelevant ones. Finally, the meta-learner optimizes classification performance through a multi-model fusion strategy, improving the accuracy and stability of phishing website detection.

[0039] In this example, by combining the feature extraction capabilities of CNN and Bi-LSTM, along with the weighted function of the sparse attention mechanism, we can effectively capture complex phishing website characteristics in URLs, significantly improving detection accuracy. Faced with diverse phishing attack methods, the model can more accurately distinguish phishing websites from legitimate ones, reducing false positives and missed detections.

[0040] The introduction of the sparse attention mechanism addresses the computational burden of traditional deep learning methods when processing high-dimensional features. By focusing only on the most relevant features in the input data, the sparse attention mechanism effectively reduces the interference of irrelevant features and improves the computational efficiency of the model.

[0041] The Stacking-XGBoost classifier integration strategy effectively combines the advantages of multiple base learners to further optimize classification performance. XGBoost demonstrates high efficiency and strong anti-overfitting capabilities when processing large-scale datasets, ensuring the robustness of the model in complex environments. This embodiment effectively extracts local features from website sample data by performing feature extraction. By extracting contextual information from these local features, it effectively captures complex phishing website features within the sample data, thereby improving the detection accuracy of the trained meta-learner and, consequently, the accuracy of phishing website detection. By applying a sparse attention mechanism to sample contextual features, it effectively reduces interference from irrelevant features, improving the computational efficiency of both the base learner and the meta-learner. By constructing a meta-feature matrix, it effectively transforms the prediction results of multiple base learners into high-order features for further fusion and optimization by the meta-learner. By applying dynamic attention weighting to the meta-feature matrix, it generates a meta-weighted matrix to enhance the meta-learner's ability to extract key features from the meta-feature matrix, improving the accuracy of the trained meta-learner and, consequently, the accuracy of phishing website detection. This embodiment is adaptable to evolving phishing website attack methods and exhibits good generalization capabilities. The model maintains high recognition capabilities across various phishing attack scenarios, offering enhanced adaptability.

[0042] Example 2 See also Figure 4 , is a schematic diagram of the structure of a phishing website detection system 100 provided in a second embodiment of the present invention, comprising: The feature extraction module 10 is used to obtain website sample data and perform feature extraction on the website sample data to obtain sample local features.

[0043] Optionally, the feature extraction module 10 is further configured to: perform character segmentation on the website sample data to obtain sample segmentation characters, and perform word embedding processing on the sample segmentation characters to obtain a word embedding vector; The word embedding vector is convolved to obtain a convolution feature, and the convolution feature is pooled to obtain the sample local feature.

[0044] The attention mechanism module 11 is used to extract context information of the sample local features to obtain sample context features, and perform sparse attention mechanism calculation on the sample context features to obtain attention sample features.

[0045] Optionally, the attention mechanism module 11 is further configured to: perform a linear transformation on the sample context feature to obtain a sample linear feature, and perform a nonlinear activation on the sample linear feature to obtain a nonlinear feature; Calculating an attention score of the nonlinear feature according to the learnable attention vector, and filtering the sample linear features according to the attention score to obtain a sample filtering feature; The attention score corresponding to the sample screening feature is normalized to obtain an attention weight, and the sample context feature is weightedly fused according to the attention weight to obtain the attention sample feature.

[0046] The matrix construction module 12 is used to train the base learner according to the attention sample features to obtain sample prediction values, and construct a meta-feature matrix according to the sample prediction values.

[0047] Optionally, the matrix construction module 12 is further configured to: divide sample subsets according to the attention sample features, and construct a sample verification set based on the sample subsets; Inputting the sample verification set into the base learner to perform sample prediction to obtain a sample prediction value, and performing vector splicing on the sample prediction value according to the sample index in the website sample data to obtain a sample prediction vector; The sample prediction vectors corresponding to different base learners are horizontally spliced by column to obtain the meta-feature matrix.

[0048] The training module 13 is configured to perform dynamic attention weighting on the meta-feature matrix to obtain a meta-weighted matrix, and input the meta-weighted matrix into a meta-learner for training until the meta-learner converges.

[0049] Optionally, the training module 13 is further configured to: calculate a classification performance index of the base learner according to the sample prediction value, and calculate a model weight of the base learner according to the classification performance index; Dynamic attention weighting is performed on the meta-feature matrix according to the model weight to obtain the meta-weighted matrix.

[0050] Furthermore, the formula used to calculate the model weight of the base learner according to the classification performance index includes: Where τ represents the temperature coefficient, represents the model weight, Indicates the k The classification performance index of the base learner on the sample validation set, Indicates the m The classification performance index of the base learner, M represents the total number of base learners; The formula used to dynamically weight the meta-feature matrix according to the model weight includes: in, a k(j) Indicates the jThe model weights of the base learners, Indicates the i The confidence weight of the sample, P i j Indicates the j The sample prediction value of the attention sample feature.

[0051] Preferably, the objective function of the meta-learner includes: Objective function in, n represents the number of samples, Indicates the i The true labels of samples, Indicates that the meta-learner is for i The predicted value of the sample, represents the loss function of the meta-learner, represents the regularization term.

[0052] The phishing detection module 14 is configured to input the website data to be detected into the meta-learner after convergence to perform phishing detection and obtain phishing website detection results.

[0053] In this embodiment, by performing feature extraction on website sample data, local features of samples in the website sample data can be effectively extracted, and by performing contextual information extraction on the local features of samples, complex phishing website features in the website sample data can be effectively captured to improve the detection accuracy of the meta-learner after training, thereby improving the accuracy of phishing website detection. By performing sparse attention mechanism calculation on sample context features, the interference of irrelevant features is effectively reduced, and the computational efficiency of the base learner and the meta-learner is improved. By constructing a meta-feature matrix, the prediction results of multiple base learners can be effectively converted into high-order features for further fusion and optimization of the meta-learner. By performing dynamic attention weighting on the meta-feature matrix, a meta-weighted matrix is obtained to improve the meta-learner's extraction of key features in the meta-feature matrix, thereby improving the accuracy of the meta-learner after training, thereby improving the accuracy of phishing website detection.

[0054] Example 3 Figure 5 This is a block diagram of a terminal device 2 provided in the third embodiment of the present application. Figure 5 As shown, the terminal device 2 of this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program for the phishing website detection method. When the processor 20 executes the computer program 22, the steps of each embodiment of the phishing website detection method described above are implemented.

[0055] Exemplarily, the computer program 22 may be divided into one or more modules, which are stored in the memory 21 and executed by the processor 20 to implement the present application. The one or more modules may be a series of computer program instruction segments capable of performing specific functions, which are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, a processor 20 and a memory 21.

[0056] The processor 20 may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0057] The memory 21 can be an internal storage unit of the terminal device 2, such as a hard drive or memory of the terminal device 2. The memory 21 can also be an external storage device of the terminal device 2, such as a plug-in hard drive, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. equipped with the terminal device 2. Furthermore, the memory 21 can include both an internal storage unit of the terminal device 2 and an external storage device. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 can also be used to temporarily store data that has been output or is about to be output.

[0058] In addition, the functional modules in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0059] If the integrated module is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be either non-volatile or volatile. Based on this understanding, the present application can implement all or part of the process steps in the above-mentioned method embodiments by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. Computer-readable storage media can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content contained in computer-readable storage media can be appropriately increased or decreased according to the requirements of legislation and patent practices in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practices, computer-readable storage media do not include electrical carrier signals and telecommunications signals.

[0060] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for detecting phishing websites, characterized in that: The method comprises: Obtaining website sample data, and performing feature extraction on the website sample data to obtain sample local features; Extracting context information from the sample local features to obtain sample context features, and performing sparse attention mechanism calculation on the sample context features to obtain attention sample features; Training a base learner according to the attention sample features to obtain sample prediction values, and constructing a meta-feature matrix according to the sample prediction values; Performing dynamic attention weighting on the meta-feature matrix to obtain a meta-weight matrix, and inputting the meta-weight matrix into a meta-learner for training until the meta-learner converges; The website data to be detected is input into the meta-learner after convergence to perform phishing detection to obtain phishing website detection results.

2. The phishing website detection method according to claim 1, wherein: The sample context features are calculated using a sparse attention mechanism to obtain attention sample features, including: Performing a linear transformation on the sample context feature to obtain a sample linear feature, and performing a nonlinear activation on the sample linear feature to obtain a nonlinear feature; Calculating an attention score of the nonlinear feature according to the learnable attention vector, and filtering the sample linear features according to the attention score to obtain a sample filtering feature; The attention score corresponding to the sample screening feature is normalized to obtain an attention weight, and the sample context feature is weightedly fused according to the attention weight to obtain the attention sample feature.

3. The phishing website detection method according to claim 1, wherein: The base learner is trained according to the attention sample features to obtain sample prediction values, and a meta-feature matrix is constructed according to the sample prediction values, including: Dividing sample subsets according to the attention sample features, and constructing a sample validation set based on the sample subsets; Inputting the sample verification set into the base learner to perform sample prediction to obtain a sample prediction value, and performing vector splicing on the sample prediction value according to the sample index in the website sample data to obtain a sample prediction vector; The sample prediction vectors corresponding to different base learners are horizontally spliced by column to obtain the meta-feature matrix.

4. The phishing website detection method according to claim 3, wherein: Dynamic attention weighting is performed on the meta-feature matrix to obtain a meta-weighted matrix, including: Calculating a classification performance index of the base learner according to the sample prediction value, and calculating a model weight of the base learner according to the classification performance index; Dynamic attention weighting is performed on the meta-feature matrix according to the model weight to obtain the meta-weighted matrix.

5. The phishing website detection method according to claim 4, wherein: The formula used to calculate the model weight of the base learner according to the classification performance index includes: Where τ represents the temperature coefficient, represents the model weight, Indicates the k The classification performance index of the base learner on the sample validation set, Indicates the m The classification performance index of the base learner, M represents the total number of base learners; The formula used to dynamically weight the meta-feature matrix according to the model weight includes: in, a k(j) Indicates the j The model weights of the base learners, Indicates the i The confidence weight of the sample, P i j Indicates the j The sample prediction value of the attention sample feature.

6. The phishing website detection method according to claim 1, wherein: The objective function of the meta-learner includes: Objective function in, n represents the number of samples, Indicates the i The true labels of samples, Indicates that the meta-learner is for i The predicted value of the sample, represents the loss function of the meta-learner, represents the regularization term.

7. The phishing website detection method according to claim 1, wherein: Feature extraction is performed on the website sample data to obtain local features of the sample, including: Performing character splitting on the website sample data to obtain sample split characters, and performing word embedding processing on the sample split characters to obtain word embedding vectors; The word embedding vector is convolved to obtain a convolution feature, and the convolution feature is pooled to obtain the sample local feature.

8. A phishing website detection system, characterized in that: The system comprises: A feature extraction module is used to obtain website sample data and perform feature extraction on the website sample data to obtain sample local features; An attention mechanism module is used to extract context information from the local features of the sample to obtain sample context features, and perform sparse attention mechanism calculation on the sample context features to obtain attention sample features; A matrix construction module is used to train a base learner according to the attention sample features to obtain sample prediction values, and to construct a meta-feature matrix according to the sample prediction values; a training module, configured to perform dynamic attention weighting on the meta-feature matrix to obtain a meta-weight matrix, and input the meta-weight matrix into a meta-learner for training until the meta-learner converges; The phishing detection module is used to input the website data to be detected into the meta-learner after convergence to perform phishing detection and obtain phishing website detection results.

9. A terminal device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.