Efficient cascade multi-stage adaptive threshold phishing detection method

Through the efficient cascading multi-stage adaptive threshold phishing detection method, the multi-dimensional data and multi-modal detection model are used to solve the problems of low detection accuracy and high time overhead in the existing technology, and efficient and accurate phishing website detection is achieved.

CN120034399AActive Publication Date: 2025-05-23SOUTHWEST PETROLEUM UNIV

Patent Information

Application Number
CN202510506019.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
2045-04-22

AI Technical Summary

Technical Problem

Existing phishing website detection methods rely on limited information from single mode or dual mode, and cannot fully utilize the correlation and interactivity between multiple modes, resulting in low detection accuracy and high time overhead.

Method used

The efficient cascaded multi-stage adaptive threshold phishing detection method is adopted. By collecting multi-dimensional data (URL features, visual elements, HTML source code), pre-processing and standardized feature extraction, a single-modal detection model is built and adversarial defense training is carried out, and the multi-modal detection model is integrated and optimized, and the weight is dynamically adjusted to improve detection efficiency and accuracy.

Benefits of technology

The optimization of computing resources is achieved, reducing the computing overhead by 70%, and the average inference delay is reduced to within 150ms. The efficiency is 2 times higher than that of the traditional multimodal model, and the false alarm rate is stable below 2%, while enhancing the robustness and scalability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120034399A_ABST
    Figure CN120034399A_ABST
Patent Text Reader

Abstract

The invention discloses an efficient cascade multi-stage self-adaptive threshold fishing detection method, and provides a cascade multi-stage self-adaptive threshold fishing detection framework aiming at the problems of insufficient anti-robustness, insufficient multi-mode collaboration, unbalanced detection efficiency and precision and the like in the existing fishing detection technology. In the first stage, millisecond-level primary screening is achieved through a lightweight single-mode detection model, and 80% or above of low-level threats are intercepted; and in the second stage, a multi-modal fusion model with high anti-interference capability is adopted, the anti-robustness is improved by combining frequency domain feature enhancement and a dynamic noise injection technology, deep correlation features are mined through a cross-modal attention mechanism, and accurate recognition of complex phishing attacks is realized. And meanwhile, an online learning module is integrated, so that the system continuously adapts to a novel attack mode, and finally, the high-precision, low-delay and strong-anti-interference detection capability is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security, and in particular to an efficient cascade multi-stage adaptive threshold phishing detection method. Background Art

[0002] As the digitalization process accelerates, phishing attacks have become a major threat to global cybersecurity. According to the 2023 Global Cybersecurity Assessment Report, the proportion of phishing attacks in various types of cybercrime continues to rise, and has become one of the most destructive forms of attack in the digital economy era.

[0003] In case of phishing attacks, the attacker will carefully design a phishing website that is very similar to the target website, or a website with false information. Once the victim visits the website and believes the content of the website, the attacker may obtain the victim's sensitive information, such as account number, password, etc. Payment transactions, financial securities, etc. can directly cause property losses to network users. The production of phishing websites does not require much technical content, but instead uses people's psychological weaknesses to deceive. Therefore, the number of phishing websites is growing. This type of attack not only uses technical loopholes, but also digs deeper into human weaknesses, forming a dual challenge of technology and network security awareness. Therefore, it is the unshirkable responsibility of all sectors of the Internet to combat phishing websites.

[0004] In the early days, there were three main types of phishing website detection: (1) Based on the blacklist and whitelist mechanism: It mainly relies on collecting relevant information about known phishing websites, such as URLs, and listing them in the blacklist database. When a user visits a URL, it is matched with the URL on the blacklist and whitelist to identify potential phishing websites. (2) Heuristic-based detection method: Through artificial rule construction and feature extraction including URL features and text content features, feature vectors are constructed based on some common features and behaviors of phishing to train the classifier model. (3) Detection model based on visual similarity: Compare the images of phishing web pages with those of regular web pages. If the matching degree is higher than the threshold, the web page is considered to be a phishing web page.

[0005] At present, most research focuses on improving feature representation, model structure and algorithm to improve detection accuracy. Secondly, most existing methods rely on limited information of a single modality or dual modality for detection, which can only capture certain abnormal features of the website and do not have strong generalization capabilities. In addition, a one-size-fits-all approach is adopted for the detection results, and only one model is input for detection. However, using only a single modality for detection cannot fully utilize the correlation and interactivity between multiple modalities. Although the multimodal fusion method makes up for the shortcoming that a single modality only focuses on one feature of the website, the time overhead of detection also increases. Summary of the invention

[0006] In view of the above-mentioned deficiencies in the prior art, the present invention provides an efficient cascaded multi-stage adaptive threshold fishing detection method.

[0007] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: An efficient cascade multi-stage adaptive threshold fishing detection method, comprising the following steps: S1. Collect and annotate data from phishing websites and legitimate websites, and construct a multi-dimensional dataset including URL features, visual elements, and HTML source code; S2, preprocessing the constructed multidimensional data set to obtain multimodal standardized features; S3. Build a unimodal detection model for each standardized feature and perform adversarial defense training; S4. Integrate and optimize the detection models of each modality to obtain the optimized multimodal decision fusion result.

[0008] Furthermore, the S1 specifically includes the following steps: S11. Synchronously crawl the URL features of phishing and legitimate websites, including protocol type, domain name composition, path level, file type, and parameter fields; S12. Capture key visual elements of web pages through Selenium automation, and identify brand logos, payment-related icons, and misleading images; S13. Parse the HTML source code, extract the dynamic behavior of hyperlinks, user privacy collection forms and text content features, and build structured metadata based on the source code length.

[0009] Furthermore, the S2 specifically includes the following steps: S21, performing noise filtering, word segmentation, and filling truncation on the URL features and HTML text in the multidimensional data set in turn, generating a semantic matrix using One-Hot encoding and word embedding, and extracting semantic features of the URL features and HTML text using a text feature extraction algorithm; S22. Perform high-frequency filtering and wavelet energy analysis on the visual elements in the multidimensional data set to extract the semantic features of the spatial dimension of the original visual elements and the frequency domain features of the visual elements after high-frequency filtering, and use the channel attention mechanism to dynamically weighted fuse the obtained semantic features and frequency domain features of the spatial dimension.

[0010] Furthermore, the specific method of performing high-frequency filtering and wavelet energy analysis on the visual elements in the multi-dimensional data set in S22 is:

[0011]

[0012] In the formula, is the Gaussian filter kernel, is the distance from the frequency point (u,v) to the center of the spectrum, is the cut-off frequency; is the wavelet energy characteristic, For the HH subband at position , MN is the size of the HH subband after wavelet decomposition.

[0013] Furthermore, the specific method of obtaining the semantic features of the spatial dimension and the frequency domain features by dynamically weighted fusion using the channel attention mechanism in S23 is as follows:

[0014] In the formula, The semantic features and frequency domain features are weighted and fused through the channel attention mechanism; is the semantic feature of spatial dimension; is the frequency domain feature; is the weight.

[0015] Furthermore, the S3 specifically includes the following steps: S31. Establish a URL feature detection model, use a 64-channel convolution kernel with a height of 2 to extract the local sequence pattern of URL features, use a bidirectional LSTM to capture temporal context information, perform batch normalization through a fully connected layer, and perform logistic regression to output the final result; S32. Establish an HTML content detection model, semantically encode the text in HTML into a fixed-length token sequence, input the lightweight pre-trained model, extract the semantic vector of the [CLS] tag as a global feature, and input the vector into a logistic regression classifier, and output the classification probability through linear transformation and Sigmoid function; S33. Build a visual element detection model, use EfficientNet-B3 transfer learning to extract normal sample features on regular websites crawled using selemium, compare the Logo area located by YOLOv5 with the visual elements in the established normal sample library in real time, and determine whether the web page is legitimate by comparing the domain name of the web page with the domain name of the brand associated with the logo.

[0016] Furthermore, the specific method of comparing the Logo area located by YOLOv5 with the normal sample library in real time in S33 is to calculate the L2 distance between the new sample and the sample library in real time, and directly determine it as abnormal if it exceeds the dynamic threshold. The specific calculation method is:

[0017] In the formula, The i-th test vector of the area to be detected, The i-th feature vector of the area to be detected.

[0018] Furthermore, the S4 specifically includes the following steps: S41, constructing a feature matrix using the confidence seat features of the multiple single-modal detection models established in S3, and training the constructed feature matrix to obtain the contribution weight of each modality; S42. According to the confidence distribution of historical samples, use a sliding window to count the confidence mean and standard deviation of the most recent K samples, set a threshold based on the confidence distribution of the samples, dynamically adjust the weights based on the performance of each modality, and perform weighted voting to obtain the final confidence result.

[0019] Furthermore, the specific method of training the constructed feature matrix in S41 to obtain the contribution weight of each mode is:

[0020]

[0021] In the formula, is the feature matrix, They are the confidence of the URL feature, HTML, and visual element monitoring models, respectively; is the Sigmoid function, is the modal label of the ith modality, is the i-th feature matrix, is the regression parameter, T is the matrix transpose, is a constant.

[0022] The present invention has the following beneficial effects: 1. The cascade architecture optimizes computing resources. In the first stage, a lightweight single-modal model is used to intercept high-confidence phishing samples in real time, and only low-confidence samples (the second-stage multimodal deep model is enabled. After testing, it can reduce computing overhead by 70%, and the average inference delay is reduced to within 150ms, which is 2 times more efficient than the traditional multimodal model. It solves the problem that the traditional stage only focuses on one feature, and solves the problem of high time overhead of the feature-level multimodal fusion method that focuses on multiple features.

[0023] 2. Dynamic weight fusion enhances the robustness of the model. In the second stage, an adaptive weight allocation mechanism based on real-time confidence is introduced to dynamically adjust the contribution of URL, HTML and visual modality. The confidence mean μ and standard deviation σ are statistically calculated based on the sliding window through an adaptive threshold, and the judgment threshold is dynamically adjusted to keep the false alarm rate stable below 2%. 3. Modular design improves expansion flexibility. By decoupling the single-modal detection model and the fusion decision layer, new feature modalities (such as JS behavior analysis) can be quickly integrated, reducing engineering adaptation costs by 60%.

[0024] 4. The website visual feature module has been added with anti-adversarial attacks, and the identification of new attack features can effectively deal with new attacks such as local tampering and adversarial samples. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 It is a schematic diagram of the efficient cascade multi-stage adaptive threshold fishing detection method of the present invention. DETAILED DESCRIPTION

[0026] The specific implementation modes of the present invention are described below so that those skilled in the art can understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation modes. For those of ordinary skill in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the attached claims, these changes are obvious, and all inventions and creations utilizing the concept of the present invention are protected.

[0027] An efficient cascade multi-stage adaptive threshold phishing detection method, such as Figure 1 As shown, the following steps are included: S1. Collect and annotate data from phishing websites and legitimate websites, and construct a multi-dimensional dataset including URLs, visual elements, and HTML source codes; In this embodiment, the following steps are specifically included: S11. Synchronously crawl the URL features of phishing and legitimate websites, including protocol type, domain name composition, path level, file type, and parameter fields; S12. Capture key visual elements of web pages through Selenium automation, and identify brand logos, payment-related icons, and misleading images; S13. Parse the HTML source code, extract the dynamic behavior of hyperlinks, user privacy collection forms and text content features, and build structured metadata based on the source code length, including: the number of hyperlinks, the number of empty links, the ratio of external links, the login form for collecting and submitting personal data, the length of the HTML source code, and the HTML text content.

[0028] S2, preprocessing the constructed multidimensional data set to obtain multimodal standardized features; In this embodiment, the following steps are specifically included: S21, performing noise filtering, word segmentation, and filling truncation on the URLs and HTML texts in the multidimensional data set in turn, generating a semantic matrix using One-Hot encoding and word embedding, and extracting semantic features of the URLs and HTML texts using a text feature extraction algorithm; The collected text data is first filtered for illegal characters, word segmentation, stop words removal, and related words replacement to remove irrelevant noise and retain the core text content. Then, the total length of characters and keywords in each URI is determined to be 300 based on the URL dataset and sensitive vocabulary. If L exceeds 300, the extra characters are truncated at the end of the URL; if L is less than 300, the characters are replaced at the end. <pad>The mark is used as an additional word. If unknown characters appear in the URL, they are marked with unknown characters. <unk>Representation. Then, One-Hot encoding is used to digitize the text data. According to the mapping table, characters and sensitive words are assigned unique codes to construct the encoding matrix of the URL. Then, the matrix U′ is converted into a 300×64 two-dimensional dense matrix X containing semantic information through the word embedding layer as the input of the convolution layer.

[0029]

[0030]

[0031] Where: u' is the encoding of the character or word in the URL, X i is a 64-dimensional column vector.

[0032] S22. Perform high-frequency filtering and wavelet energy analysis on the visual elements in the multidimensional data set to extract the semantic features of the spatial dimension of the original visual elements and the frequency domain features of the visual elements after high-frequency filtering, and use the channel attention mechanism to dynamically weighted fuse the obtained semantic features and frequency domain features of the spatial dimension.

[0033] Next is the feature extraction stage, which uses the text feature extraction algorithm TF-IDF algorithm to convert text data into a vector form that can be processed by computers. TF-IDF, as a commonly used weight in information retrieval, text mining and other fields, can capture the relationship and semantic information between words and put them into the model for training.

[0034]

[0035]

[0036] in: Expressing words In the article The number of times it appears in Indicates article The sum of the number of occurrences of all words; Represents the total number of articles in the corpus. Indicates that it contains words The number of articles (i.e. ), if the term does not appear in any other article in the corpus, the value is zero.

[0037] In order to analyze the interference resistance of web page images, this patent proposes a DSF-Net defense method that combines frequency domain analysis and a dual-stream network. First, the Gaussian filter kernel and wavelet energy characteristics are calculated:

[0038] in is the distance from the frequency point (u,v) to the center of the spectrum, is the cut-off frequency (controls the filter range)

[0039] in For the HH subband at position The coefficient of , MN is the size of the HH subband after wavelet decomposition.

[0040] After calculating the Gaussian filter kernel and wavelet energy features, dual-stream feature fusion is performed, and the original image is input into the lightweight MobileNetV3 network to extract the semantic features of its extracted spatial dimension to obtain the output feature map Then, the high-frequency filtered image is input into the MobileNetV3 network to capture the high-frequency disturbance features and edge information to obtain the feature map , then we will get , Dynamically weighted fusion via channel attention mechanism: .

[0041] S3. Build a unimodal detection model for each standardized feature and perform adversarial defense training; In this embodiment, the following steps are specifically included: S31. Establish a URL feature detection model, use a 64-channel convolution kernel with a height of 2 to extract the local sequence pattern of URL features, use a bidirectional LSTM to capture temporal context information, perform batch normalization through a fully connected layer, and perform logistic regression to output the final result; The CNN-BiLSTM model combines a convolutional neural network (CNN) and a bidirectional long short-term memory network (BiLSTM). It extracts local spatial features (such as textures and temporal fragments) through CNN, uses BiLSTM to capture temporal context information bidirectionally, enhances long-distance dependency modeling, and introduces Dropout and regularization techniques to prevent overfitting for the final result output.

[0042] The processed URL vector matrix in S2 is input into the convolutional neural network, and local features are automatically extracted from the feature matrix through the convolution kernel. The height h of the convolution kernel is set to 2, the width is 64, which is consistent with the dimension of the character vector, the number of convolution kernels is 200, and the convolution kernel sliding step is set to 1. For a certain convolution kernel f, the URL embedding matrix obtained at the i-th sliding window is set to X;

[0043] Where: x i It is the vector representation of characters or sensitive words.

[0044] The data is then batch normalized through a fully connected layer to perform logistic regression and output the final result.

[0045]

[0046]

[0047]

[0048] in is the Sigmoid function: .

[0049] S32. Establish an HTML content detection model, semantically encode the text in HTML into a fixed-length token sequence, input the lightweight pre-trained model, extract the semantic vector of the [CLS] tag as a global feature, and input the vector into a logistic regression classifier, and output the classification probability through linear transformation and Sigmoid function; After word segmentation, the text in HTML is semantically encoded and converted into a fixed-length token sequence, which is input into the lightweight pre-trained DistilBERT model to extract the semantic vector of the [CLS] tag as the global feature. The vector is then input into the logistic regression (LR) classifier, and the classification probability is output through linear transformation and Sigmoid function.

[0050]

[0051]

[0052] Optimize using the binary cross entropy loss function:

[0053] Where: Y i is the true label, P i is the predicted label.

[0054] S33. Build a visual element detection model, use EfficientNet-B3 transfer learning to extract visual features, compare the Logo area located by YOLOv5 with the normal sample library in real time, and determine whether the web page is legitimate by comparing the domain name of the web page with the domain name of the brand associated with the logo.

[0055] EfficientNet-B3 is used for transfer learning. Its composite scaling strategy achieves efficient feature extraction for anomaly detection by balancing the network depth, width and resolution. After extracting the feature vector of the target area, a normal sample feature library is established. When a user opens a web page, the detection system first uses yolov5 to locate the key areas in the image (such as brand logos, input forms) in real time. After obtaining the located key areas, the part of the image is extracted, and the L2 distance between the new sample and the sample library is calculated in real time. If it exceeds the dynamic threshold, it is directly judged as abnormal. If the abnormality of the image is not identified, the detector is used to determine the brand to which the logo belongs. The system then determines whether the web page is legal by comparing the domain name of the web page with the domain name of the brand associated with the logo.

[0056]

[0057] Where: V test The feature vector of the area to be detected, V ref The feature vector of the area to be detected.

[0058] S4. Integrate and optimize the detection models of each modality to obtain the optimized multimodal decision fusion result.

[0059] In this embodiment, the following steps are specifically included: The specific steps include: S41, constructing a feature matrix using the confidence seat features of the multiple single-modal detection models established in S3, and training the constructed feature matrix to obtain the contribution weight of each modality; The confidence output of each modality is used as a feature to construct a feature matrix X, train a logistic regression model as a meta-classifier, and learn the contribution weight of each modality through maximum likelihood estimation. The predicted probability of each modality on the validation set is collected, and a feature matrix is ​​constructed and put into the model training to obtain the contribution weight of each modality.

[0060]

[0061]

[0062] In the formula, is the feature matrix, They are the confidence of the URL feature, HTML, and visual element monitoring models, respectively; is the Sigmoid function, is the modal label of the ith modality, is the i-th feature matrix, is the regression parameter, T is the matrix transpose, is a constant.

[0063] S42. Use a sliding window to statistically maintain the confidence mean and standard deviation of the most recent K samples according to the confidence distribution of historical samples, set a threshold according to the confidence distribution of the samples, dynamically adjust the weights according to the performance of each modality, and obtain the final result of the confidence through weighted voting.

[0064] Considering the dynamic changes in the real-time confidence of each modality due to the adjustment of the sample size, the present patent invented a multi-modal model with adaptive dynamic weight adjustment. First, use the Softmax function to p m normalize the modality confidence, and then perform dynamic fusion on it to obtain the result p final , use a sliding window to statistically maintain the most recent K samples p final mean and standard deviation according to the confidence distribution of historical samples. According to the empirical learning model, dynamically set a threshold

[0065]

[0066]

[0067]

[0068]

[0069]

[0070] Where: Where: α is the temperature coefficient, which controls the steepness of the weight distribution. α ->0, the weights tend to be evenly distributed, α ->∞, only the modality with the highest confidence dominates. is the weight, λ is the sensitivity coefficient (e.g., λ =2 corresponds to the 95% confidence interval).

[0071] After obtaining the real-time confidence of each modality, dynamically adjust the weights according to the performance of each modality, and then perform weighted voting to obtain the final result.

[0072] .

[0073] The above model is used to conduct a comprehensive evaluation of predictions on real websites. Through a large number of experimental designs of anti-phishing detection solutions at home and abroad, the performance of the cascade multi-stage framework is judged based on evaluation indicators. The main performance indicators are summarized as follows: True Positive Rate (TPR) and False Positive Rate (FPR) as decisive reference points for distinguishing good and bad performance; Precision and Recall reflect the ability to identify phishing web pages. The F1 value takes into account both precision and accuracy, and is a weighted average of the two, which can comprehensively evaluate the performance of the detection model. The specific calculation method is as follows:

[0074]

[0075]

[0076]

[0077]

[0078]

[0079] Among them: TP represents the number of predicted phishing web pages that are actually phishing web pages; FP represents the number of predicted phishing web pages that are actually legal web pages; TN represents the number of predicted legal web pages that are actually legal web pages; FN represents the number of predicted legal web pages that are actually phishing web pages.

[0080] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0081] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0083] The present invention uses specific embodiments to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.

[0084] Those skilled in the art will appreciate that the embodiments described herein are intended to help readers understand the principles of the present invention, and should be understood that the protection scope of the present invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific variations and combinations that do not deviate from the essence of the present invention based on the technical revelations disclosed by the present invention, and these variations and combinations are still within the protection scope of the present invention.< / unk> < / pad>

Claims

1. An efficient cascade multi-stage adaptive threshold fishing detection method, characterized in that: The steps include: S1. Collect and annotate data from phishing websites and legitimate websites, and construct a multi-dimensional dataset including URL features, visual elements, and HTML source code; S2, preprocessing the constructed multidimensional data set to obtain multimodal standardized features; S3. Build a unimodal detection model for each standardized feature and perform adversarial defense training; S4. Integrate and optimize the detection models of each modality to obtain the optimized multimodal decision fusion result.

2. According to claim 1, an efficient cascade multi-stage adaptive threshold fishing detection method is characterized in that: The S1 specifically includes the following steps: S11. Synchronously crawl the URL features of phishing and legitimate websites, including protocol type, domain name composition, path level, file type, and parameter fields; S12. Capture key visual elements of web pages through Selenium automation, and identify brand logos, payment-related icons, and misleading images; S13. Parse the HTML source code, extract the dynamic behavior of hyperlinks, user privacy collection forms and text content features, and build structured metadata based on the source code length.

3. According to claim 1, an efficient cascade multi-stage adaptive threshold fishing detection method is characterized in that: The S2 specifically includes the following steps: S21, performing noise filtering, word segmentation, and filling truncation on the URL features and HTML text in the multidimensional data set in turn, generating a semantic matrix using One-Hot encoding and word embedding, and extracting semantic features of the URL features and HTML text using a text feature extraction algorithm; S22. Perform high-frequency filtering and wavelet energy analysis on the visual elements in the multidimensional data set to extract the semantic features of the spatial dimension of the original visual elements and the frequency domain features of the visual elements after high-frequency filtering, and use the channel attention mechanism to dynamically weighted fuse the obtained semantic features and frequency domain features of the spatial dimension.

4. The efficient cascade multi-stage adaptive threshold fishing detection method according to claim 3, characterized in that: The specific method of performing high-frequency filtering and wavelet energy analysis on the visual elements in the multi-dimensional data set in S22 is: In the formula, is the Gaussian filter kernel, is the distance from the frequency point (u,v) to the center of the spectrum, is the cut-off frequency; is the wavelet energy characteristic, For the HH subband at position , MN is the size of the HH subband after wavelet decomposition.

5. The efficient cascade multi-stage adaptive threshold fishing detection method according to claim 3, characterized in that: The specific method of obtaining the semantic features of spatial latitude and frequency domain features by dynamically weighted fusion using the channel attention mechanism in S23 is as follows: In the formula, The semantic features and frequency domain features are weighted and fused through the channel attention mechanism; is the semantic feature of spatial dimension; is the frequency domain feature; is the weight.

6. The efficient cascade multi-stage adaptive threshold fishing detection method according to claim 1, characterized in that: The S3 specifically includes the following steps: S31. Establish a URL feature detection model, use a 64-channel convolution kernel with a height of 2 to extract the local sequence pattern of URL features, use a bidirectional LSTM to capture temporal context information, perform batch normalization through a fully connected layer, and perform logistic regression to output the final result; S32. Establish an HTML content detection model, semantically encode the text in HTML into a fixed-length token sequence, input the lightweight pre-trained model, extract the semantic vector of the [CLS] tag as a global feature, and input the vector into a logistic regression classifier, and output the classification probability through linear transformation and Sigmoid function; S33. Build a visual element detection model, use EfficientNet-B3 transfer learning to extract normal sample features on regular websites crawled using selemium, compare the Logo area located by YOLOv5 with the visual elements in the established normal sample library in real time, and determine whether the web page is legitimate by comparing the domain name of the web page with the domain name of the brand associated with the logo.

7. The efficient cascade multi-stage adaptive threshold fishing detection method according to claim 6, characterized in that: The specific method of comparing the Logo area located by YOLOv5 with the normal sample library in real time in S33 is to calculate the L2 distance between the new sample and the sample library in real time. If it exceeds the dynamic threshold, it is directly determined to be abnormal. The specific calculation method is: In the formula, The i-th test vector of the area to be detected, The i-th feature vector of the area to be detected.

8. The efficient cascade multi-stage adaptive threshold fishing detection method according to claim 1, characterized in that: The S4 specifically includes the following steps: S41, constructing a feature matrix using the confidence seat features of the multiple single-modal detection models established in S3, and training the constructed feature matrix to obtain the contribution weight of each modality; S42. According to the confidence distribution of historical samples, use a sliding window to count the confidence mean and standard deviation of the most recent K samples, set a threshold based on the confidence distribution of the samples, dynamically adjust the weights based on the performance of each modality, and perform weighted voting to obtain the final confidence result.

9. The efficient cascade multi-stage adaptive threshold fishing detection method according to claim 8, characterized in that: The specific method of training the constructed feature matrix in S41 to obtain the contribution weight of each mode is: In the formula, is the feature matrix, They are the confidence of the URL feature, HTML, and visual element monitoring models, respectively; is the Sigmoid function, is the modal label of the ith modality, is the i-th feature matrix, is the regression parameter, T is the matrix transpose, is a constant.

Citation Information

Patent Citations

  • Detection method for phishing site

    CN102571768A

  • Phishing website detection method based on multi-feature fusion

    CN108777674A

  • Network intrusion detection method based on machine learning integration model

    CN112769752A

  • Multi-modal fusion feature-based phishing detection method and system

    CN115051817A

  • Generalized phishing website detection method and system oriented to combination of URL (Uniform Resource Locator) and label

    CN115622739A

Cited By

  • Phishing website detection method based on multi-model cascade and genetic algorithm optimization

    CN121530761A

  • A phishing website detection method based on multi-model cascade and genetic algorithm optimization

    CN121530761B

  • High-precision Ethereum phishing account detection method based on high-order topology

    CN122179130A