Phishing website detection method, device and equipment

By combining the website source code and screenshot images as information sources, the multimodal analysis model is trained to generate a phishing website detection model, which solves the problem of low detection accuracy in the existing technology, and achieves higher phishing website detection accuracy and network security.

CN120223393APending Publication Date: 2025-06-27CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510369922.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing phishing website detection technology relies on a single information source, resulting in poor detection of phishing websites using screenshots as backgrounds, and the specific content and code of the web page cannot be extracted, resulting in low detection accuracy.

Method used

By combining the website source code and the website screenshot image as information sources, the multimodal analysis model is trained to obtain the phishing website detection model. The model generates the first and second comprehensive eigenvectors by training the target website and suspicious website data in the data set, and determines the probability value of the phishing website based on these eigenvectors, the weight matrix of the fully connected layer, the bias term and the activation function, and finally optimizes the model weight through the cross entropy loss function and the consistency loss value.

Benefits of technology

It improves the accuracy of detection of phishing websites, improves the network security of users, and solves the problems of poor detection results and low accuracy in the existing technology.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223393A_ABST
    Figure CN120223393A_ABST
Patent Text Reader

Abstract

The invention discloses a phishing website detection method, device and equipment, and the method comprises the steps: inputting target website data into a multi-modal analysis model to obtain a first comprehensive feature vector and a first target consistency loss value, and inputting suspicious website data into the multi-modal analysis model to obtain a second comprehensive feature vector and a second target consistency loss value; determining a probability value of the suspicious website based on the first comprehensive feature vector, the second comprehensive feature vector, the full connection layer and the activation function; and obtaining an optimized weight matrix and an optimized bias item based on a gradient descent algorithm and the total loss function value to realize training of a multi-modal analysis model and obtain a phishing website detection model, thereby inputting the to-be-detected website data into the phishing website detection model, and outputting a category corresponding to the to-be-detected website data. According to the method, the website source code and the website screenshot image are combined to serve as the information source, and the accuracy of phishing website detection is improved through training of the multi-modal analysis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and particularly to a phishing website detection method, device and equipment. Background Art

[0002] Phishing websites usually use deceptive emails or links to induce users to visit fake websites disguised as legitimate websites to steal users' sensitive personal information, such as login credentials, passwords, bank card information, credit card information, etc. Since phishing websites are created by attackers and are very similar to legitimate websites, users are misled into thinking they are legitimate websites and provide personal sensitive information, resulting in economic losses and privacy leaks.

[0003] Existing phishing website detection technologies often rely on a single information source. For example, only using the website source code as the information source, this method has poor detection effect for phishing websites using screenshots as the background, and the method of only using website screenshot images as the information source generally has a general detection effect because the specific content and code of the web page cannot be extracted. Therefore, how to combine the website source code and website screenshot images as information sources to improve the accuracy of phishing website detection and thus enhance users' network security has become a technical problem to be solved urgently. Summary of the Invention

[0004] The present invention provides a phishing website detection method, device and equipment to solve the technical problems of poor detection effect and low accuracy rate in the existing technology for phishing website detection.

[0005] In a first aspect, the present application provides a phishing website detection method, the method comprising:

[0006] Obtain data of the website to be detected;

[0007] Input the data of the website to be detected into a phishing website detection model, and output a category corresponding to the data of the website to be detected, wherein the category is used to represent whether the website to be detected is a phishing website;

[0008] Wherein, the phishing website detection model is trained by the following method:

[0009] Input the target website data in the training data set into a multi-modal analysis model to obtain a first comprehensive feature vector and a first target consistency loss value, and input the suspicious website data in the training data set into the multi-modal analysis model to obtain a second comprehensive feature vector and a second target consistency loss value;

[0010] Determine a probability value based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully connected layer, the bias term of the fully connected layer and the activation function;

[0011] After calculating the cross - entropy loss function value based on the probability value and the cross - entropy loss function, calculate the total loss function value based on the cross - entropy loss function value, the first target consistency loss value, and the second target consistency loss value;

[0012] Based on the gradient descent algorithm and the total loss function value, calculate the optimized weight matrix and the optimized bias term to implement the training of the multi - modal analysis model and obtain the phishing website detection model.

[0013] In a possible implementation manner, inputting the target website data in the training dataset into the multi - modal analysis model to obtain the first comprehensive feature vector includes:

[0014] Taking the image sequence corresponding to the target website data as the first input feature and inputting it into the first multi - modal sub - model to obtain the first output feature, where the first multi - modal sub - model includes multiple first target layers and multiple second target layers, the output feature of the first target layer is used as the input feature of the second target layer, and the output feature of the second target layer is used as the input feature of the next first target layer;

[0015] Taking the text sequence corresponding to the target website data as the second input feature and inputting it into the second multi - modal sub - model to obtain the second output feature, where the second multi - modal sub - model includes multiple third target layers and multiple fourth target layers, the output feature of the third target layer is used as the input feature of the fourth target layer, and the output feature of the fourth target layer is used as the input feature of the next third target layer;

[0016] After constructing the first target feature and the second target feature based on the first output feature, and constructing the third target feature based on the second output feature, determine the target comprehensive feature based on the first target feature, the second target feature, and the third target feature;

[0017] Based on the target comprehensive feature and the average pooling method, obtain the first comprehensive feature vector.

[0018] In a possible implementation manner, constructing the first target feature and the second target feature based on the first output feature, and constructing the third target feature based on the second output feature includes:

[0019] Taking the product of the first output feature and a preset first weight matrix as the first target feature;

[0020] Taking the product of the first output feature and a preset second weight matrix as the second target feature;

[0021] Take the product of the second output feature and a preset third weight matrix as the third target feature.

[0022] In a possible implementation manner, determining the target comprehensive feature based on the first target feature, the second target feature, and the third target feature includes:

[0023] After calculating the first product of the transpose of the first target feature and the third target feature, calculate the first quotient value obtained by dividing the first product by a first preset value;

[0024] Calculate the first activation function value of the first quotient value based on the activation function;

[0025] Based on a fully connected layer, calculate the second product obtained by multiplying the first activation function value and the second target feature, and take the second product as the target comprehensive feature.

[0026] In a possible implementation manner, inputting the target website data in the training dataset into the multi-modal analysis model to obtain the first target consistency loss value includes:

[0027] Determine the first consistency loss value based on the intermediate layer output feature of the first multi-modal sub-model and the second output feature;

[0028] Determine the second consistency loss value based on the first output feature and the second output feature;

[0029] Determine the third consistency loss value based on the target comprehensive feature, the first output feature, and the second output feature;

[0030] Take the sum value of the first consistency loss value, the second consistency loss value, and the third consistency loss value as the first target consistency loss value.

[0031] In a possible implementation manner, determining the third consistency loss value based on the target comprehensive feature, the first output feature, and the second output feature includes:

[0032] For each target website data, calculate the first squared value of the difference between the target comprehensive feature vector and the first output feature, and calculate the second squared value of the difference between the target comprehensive feature vector and the second output feature;

[0033] After calculating the first sum value of the first squared value and the second squared value, calculate the second sum value obtained by adding the first sum values of each target website data;

[0034] After calculating a second quotient value obtained by dividing the second sum value by the number of eigenvalue, use a third quotient value obtained by dividing the second quotient value by a second preset value as the third consistency loss value.

[0035] In a possible implementation, determining the probability value based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully connected layer, the bias term of the fully connected layer, and the activation function includes:

[0036] Concatenate the first comprehensive feature vector and the second comprehensive feature vector to obtain a concatenated comprehensive feature vector;

[0037] After calculating a third product obtained by multiplying the weight matrix by the concatenated comprehensive feature vector, calculate a third sum value obtained by adding the third product and the bias term;

[0038] After calculating a second activation function value of the third sum value based on the activation function, use the second activation function value as the probability value.

[0039] In a possible implementation, calculating the cross entropy loss function value based on the probability value and the cross entropy loss function includes:

[0040] Calculate a first difference obtained by subtracting a third preset value from 1, a second difference obtained by subtracting the probability value from 1, and a first derivative value of the probability value;

[0041] Calculate a second derivative value of the second difference;

[0042] After calculating a fourth product obtained by multiplying the third preset value by the first derivative value and a fifth product obtained by multiplying the first difference by the second derivative value, calculate a fourth sum value obtained by adding the fourth product and the fifth product;

[0043] Use the negative value of the fourth sum value as the cross entropy loss function value.

[0044] In a second aspect, an embodiment of the present application further provides a phishing website detection device, where the device includes:

[0045] A category determination module, configured to obtain data of a website to be detected; input the data of the website to be detected into a phishing website detection model, and output a category corresponding to the data of the website to be detected, where the category is used to represent whether the website to be detected is a phishing website;

[0046] A model training module, configured to input target website data in a training dataset into a multi-modal analysis model to obtain a first comprehensive feature vector and a first target consistency loss value, and input suspicious website data in the training dataset into the multi-modal analysis model to obtain a second comprehensive feature vector and a second target consistency loss value; determine a probability value based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully-connected layer, the bias term of the fully-connected layer, and the activation function; calculate a cross-entropy loss function value based on the probability value and the cross-entropy loss function, and then calculate a total loss function value based on the cross-entropy loss function value, the first target consistency loss value, and the second target consistency loss value; calculate an optimized weight matrix and an optimized bias term based on the gradient descent algorithm and the total loss function value to implement the training of the multi-modal analysis model and obtain the phishing website detection model.

[0047] In a third aspect, an embodiment of the present application further provides a phishing website detection device, including at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of the first aspect.

[0048] The beneficial effects of the present invention are as follows:

[0049] An embodiment of the present application provides a phishing website detection method, device, and equipment. The present application trains a multi-modal analysis model to obtain a phishing website detection model. When obtaining data of a website to be detected, the data of the website to be detected is input into the phishing website detection model, and a category corresponding to the data of the website to be detected is output, where the category is used to characterize whether the website to be detected is a phishing website. Specifically, the multi-modal analysis model is trained by the following method: First, the target website data in the training data set is input into the multi-modal analysis model to obtain a first comprehensive feature vector and a first target consistency loss value, and the suspicious website data in the training data set is input into the multi-modal analysis model to obtain a second comprehensive feature vector and a second target consistency loss value. Second, based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully connected layer, the bias term of the fully connected layer, and the activation function, the probability value that the suspicious website is a phishing website is determined. Finally, after calculating the cross-entropy loss function value based on the probability value and the cross-entropy loss function, the total loss function value is calculated based on the cross-entropy loss function value, the first target consistency loss value, and the second target consistency loss value. The optimized weight matrix and the optimized bias term are calculated based on the gradient descent algorithm and the total loss function value to implement the training of the multi-modal analysis model and obtain the phishing website detection model. The present application combines the website source code and the website screenshot image as information sources, and improves the accuracy of phishing website detection through the training of the multi-modal analysis model, thereby enhancing the network security of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.

[0051] Figure 1 It is a schematic flowchart of a phishing website detection method provided by an embodiment of the present application;

[0052] Figure 2 It is a schematic flowchart of a phishing website detection method provided by an embodiment of the present application;

[0053] Figure 3 It is a schematic flowchart of another phishing website detection method provided by an embodiment of the present application;

[0054] Figure 4 It is a schematic flowchart of another phishing website detection method provided by an embodiment of the present application;

[0055] Figure 5 It is a schematic flowchart of another phishing website detection method provided by an embodiment of the present application;

[0056] Figure 6 Schematic flow diagram of another phishing website detection method provided by an embodiment of the present application;

[0057] Figure 7 Schematic structural diagram of a training method for a phishing website detection model provided by an embodiment of the present application;

[0058] Figure 8 Schematic structural diagram of another training method for a phishing website detection model provided by an embodiment of the present application;

[0059] Figure 9 Schematic diagram of a phishing website detection device provided by an embodiment of the present application;

[0060] Figure 10 Schematic diagram of a phishing website detection device provided by an embodiment of the present application. Detailed implementation manners

[0061] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0062] It should be noted that the terms "including" and "having" and their variations involved in the documents of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0063] The terms "first" and "second" in the text are only used for descriptive purposes and cannot be construed as explicitly or implicitly indicating relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0064] The term "exemplary" used hereinafter means "serving as an example, an embodiment or an illustration". Any embodiment described as "exemplary" does not have to be construed as superior to or better than other embodiments.

[0065] Phishing websites usually use deceptive emails or links to induce users to visit fake websites disguised as legitimate ones to steal users' sensitive personal information, such as login credentials, passwords, bank card information, credit card information, etc. Since phishing websites are created by attackers and are very similar to legitimate websites, users may mistake them for legitimate websites and provide personal sensitive information, resulting in financial losses and privacy leaks. Therefore, phishing website detection technology is particularly important.

[0066] However, existing phishing website detection technologies often rely on a single information source. For example, only using the website source code as the information source, this method has poor detection effect in dealing with phishing websites with screenshots as the background. And the method of only using website screenshot images as the information source generally has a mediocre detection effect because it cannot extract the specific content and code of the web page. Therefore, how to combine the website source code and website screenshot images as information sources to improve the accuracy of phishing website detection and thus enhance users' network security has become a technical problem to be solved urgently.

[0067] To solve this problem, the embodiments of the present application provide a phishing website detection method, device and equipment. By combining the website source code and website screenshot images as information sources and training a multimodal analysis model, the accuracy of phishing website detection is improved, thus enhancing users' network security.

[0068] For ease of understanding, the following explanations are made for the technical terms involved in the embodiments of the present application:

[0069] (1) Multimodal Analysis Model (MAM): It refers to an information model that can simultaneously process and fuse data of different modalities. Common modalities include text, image, video, audio, and sensor data. By combining information of different modalities, the model can obtain more comprehensive and accurate knowledge, improving performance and application effects.

[0070] (2) Self-Attention Mechanism (SAM): The self-attention mechanism generates a weighted representation for each element by calculating the similarity between each element in the input sequence and other elements. That is to say, the self-attention mechanism is a technique for capturing global dependencies by calculating the relationships between elements in a sequence and is widely used in the fields of natural language processing and computer vision.

[0071] (3) Feedforward Neural Network (FNN) is the most basic artificial neural network structure, where data flows unidirectionally in the network in a hierarchical order. It consists of an input layer, one or more hidden layers, and an output layer. Neurons in each layer are connected to neurons in the next layer through weights. The main task of the feedforward network is to extract features from the input data and generate an output through an activation function.

[0072] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.

[0073] Figure 1 It is a schematic flowchart of a phishing website detection method provided by an embodiment of the present application. As Figure 1 shown, a phishing website detection method provided by an embodiment of the present application specifically includes the following processes

[0074] S101. Obtain the data of the website to be detected;

[0075] S102. Input the data of the website to be detected into the phishing website detection model, and output the category corresponding to the data of the website to be detected.

[0076] Among them, the category is used to represent whether the website to be detected is a phishing website.

[0077] In the embodiments of the present application, through the training of the multimodal analysis model, a phishing website detection model is obtained. When obtaining the data of the website to be detected; inputting the data of the website to be detected into the phishing website detection model, and outputting the category corresponding to the data of the website to be detected, where the category is used to represent whether the website to be detected is a phishing website. Specifically, in the phishing website detection model, by determining the probability value that the data of the website to be detected is a phishing website of the target website, and then determining whether the website to be detected is a phishing website. Exemplarily, a preset probability value is set. When the probability value that the website to be detected determined by the phishing website detection model is a phishing website of the target website is greater than or equal to the preset probability value, it is determined that the website to be detected is a phishing website; when the probability value that the website to be detected determined by the phishing website detection model is a phishing website of the target website is less than the preset probability value, it is determined that the website to be detected is not a phishing website.

[0078] Specifically, the phishing website detection model is trained through the following method. As Figure 2 , it is a schematic flowchart of a training method for a phishing website detection model provided by an embodiment of the present application, and specifically includes the following steps:

[0079] S201. Input the target website data in the training dataset into the multi-modal analysis model to obtain the first comprehensive feature vector and the first target consistency loss value, and input the suspicious website data in the training dataset into the multi-modal analysis model to obtain the second comprehensive feature vector and the second target consistency loss value;

[0080] It should be noted that before training the multi-modal analysis model, it is necessary to construct a training dataset including the target website and the phishing website of the target website, where the target website represents the website imitated by the phishing website.

[0081] In a specific embodiment, if the data in the training dataset obtained during the data collection process is in the URL format, it is necessary to use a crawler program, such as the requests library and BeautifulSoup library in Python, to access the URL, obtain the source code of the web page, and take a screenshot of the web page. Specifically, convert the obtained source code into text format and save the obtained web page screenshot in image format.

[0082] Exemplarily, it is also necessary to integrate the data in the training dataset into a preset format. In this application, the preset format can be {<origin i , sample i , y i >}. When y i = 1, sample i represents a phishing website, and origin i represents the target website imitated by the phishing website; when y i = 0, both sample i and origin i represent non-phishing websites, or sample i represents a phishing website, but origin i is not the target website corresponding to the phishing website.

[0083] For example, {<origin1, sample1, 1>} represents that sample1 is a phishing website and origin1 is the target website imitated by the phishing website sample1; {<origin2, sample2, 0>} can represent that both sample2 and origin2 are non-phishing websites, or that sample2 is a phishing website, but origin2 is not the target website corresponding to the phishing website sample2.

[0084] In a possible implementation, in order to combine the website source code and the website screenshot image as information sources to comprehensively obtain the data content of the web page, this application extracts features from the image sequence corresponding to the website data and the text sequence corresponding to the website data, and obtains the target comprehensive features by fusing the image features and the text features. As Figure 3 shown, it is a schematic flowchart of another method for training a phishing website detection model provided by an embodiment of this application. The methods for obtaining the first comprehensive feature vector and the second comprehensive feature vector are similar. Here, taking the example of inputting the target website data in the training dataset into the multi-modal analysis model to obtain the first comprehensive feature vector for detailed description.

[0085] S301. Use the image sequence corresponding to the target website data as the first input feature and input it into the first multi-modal sub-model to obtain the first output feature;

[0086] Among them, the first multi-modal sub-model includes multiple first target layers and multiple second target layers. The output feature of the first target layer is used as the input feature of the second target layer, and the output feature of the second target layer is used as the input feature of the next first target layer.

[0087] Specifically, the first multi-modal sub-model can be a vision transformer (ViT) model. As an example, before using the image sequence corresponding to the target website data in the training dataset as the first input feature and inputting it into the first multi-modal sub-model, the image data of the website screenshot of the obtained target website data is cut to obtain multiple image patches.

[0088] Exemplarily, for example, the image data of the website screenshot of the target website data is an RGB image of 224×224 pixels. This application cuts this image into small images of 16×16, that is, this application divides this image into 196 small pieces. Since the vision transformer model cannot directly process 2D data, therefore, by flattening the small images into 1D vectors, the image sequence corresponding to the target website data is determined so as to use this image sequence as the first input feature and input it into the first multi-modal sub-model. It should be noted that 2D data is represented as data in a two-dimensional space and usually exists in the form of a matrix, 1D data is represented as data in a one-dimensional space and usually exists in the form of a vector, and an RGB image is represented as a color image, and each pixel is composed of numerical values of three color channels: red, green, and blue. In addition, since each small image has 16×16 pixels and each pixel has 3 color channels, the size of each small image is 16×16×3 = 768, that is, the dimension of each small image is 768.

[0089] In the embodiment of the present application, the vision transformer model consists of 12 layers of self-attention mechanism (Multi-head Attention) and feed-forward neural network (FFN). Therefore, the input of the h-th layer is set as where h represents the number of layers. First, at the h-th layer, the query (Q), key (K), and value (V) are calculated through the self-attention mechanism, that is:

[0090]

[0091] where W1, W2, and W3 are pre-set weight matrices.

[0092] Secondly, the dot product of Q1 and K1 is calculated, that is, the product of the transpose of K1 and Q1 is calculated, and the activation function softmax is used to obtain the attention weights between each image patch, that is where represents the square root of the dimension of the input vector.

[0093] Finally, the attention weights are weighted and summed with V1, and the output features of the h-th layer are obtained based on the feed-forward neural network

[0094] It should be noted that the output features of the first target layer are used as the input features of the second target layer, and the output features of the second target layer are used as the input features of the next first target layer. Among them, the first target layer can be the h-th layer, the second target layer can be the h + 1-th layer, and the next first target layer can be the h + 2-th layer. That is, the output features of the first target layer h-th layer are used as the input features of the h + 1-th layer, and the output features of the h + 1-th layer are used as the input features of the h + 2-th layer.

[0095] Exemplarily, taking an image sequence with k = 3 as the first input feature as an example, for instance, there are 4 image patches, which are x0 = [1, 2, 3, 4], x1 = [5, 6, 7, 8], x2 = [9, 10, 11, 12], and x3 = [13, 14, 15, 16] respectively. At the first layer, the input image sequence as the first input feature is Through the self-attention mechanism, the query (Q), key (K), and value (V) are calculated to obtain:

[0096]

[0097] After that, Furthermore, the output features of the first layer are obtained

[0098] Taking the output features of the first layer as the input features of the second layer Until the output features of the 12th layer are obtained.

[0099] S302. Use the text sequence corresponding to the target website data as the second input feature and input it into the second multi-modal sub-model to obtain second output features;

[0100] Among them, the second multi-modal sub-model includes multiple third target layers and multiple fourth target layers. The output features of the third target layer serve as the input features of the fourth target layer, and the output features of the fourth target layer serve as the input features of the next third target layer.

[0101] The second multi-modal sub-model can be a BERT (Bidirectional Encoder Representations from Transformers) model. As an example, before using the text sequence corresponding to the target website data in the training dataset as the second input feature and inputting it into the second multi-modal sub-model, after converting the source code of the obtained target website data into text format, convert the data in this text format into a text sequence.

[0102] Exemplarily, perform word segmentation on the obtained text format data to split it into multiple small unit texts. For example, when the input text is: "Hello, this is a test.", after word segmentation, it is obtained: ["hello", ",", "this", "is", "a", "test", "."]. Then, based on a preset look-up table, convert each word segment into a unique digital ID. For example, convert ["hello", ",", "this", "is", "a", "test", "."] into [101, 102, 103, 104, 105, 106, 107], thereby obtaining the second input feature.

[0103] In the embodiments of the present application, the BERT model is composed of 6 layers of self-attention mechanism (Multi-head Attention) and feed-forward neural network (FFN). Therefore, the input of the hth layer is set as First, where h represents the layer number. At the hth layer, calculate the query (Q), key (K), and value (V) through the self-attention mechanism, that is:

[0104]

[0105] Among them, W1, W2, and W3 are pre-set weight matrices.

[0106] Secondly, calculate the dot product of Q2 and K2, that is, calculate the product of the transpose of K2 and Q2, and use the activation function softmax to obtain the attention weights between each text sequence value.

[0107] That is Among them is characterized as the square root of the dimension of the input vector.

[0108] Finally, calculate the attention weights perform weighted summation with V2, and obtain the output features of the h-th layer based on the feed-forward neural network

[0109] It should be noted that the output features of the third target layer serve as the input features of the fourth target layer, and the output features of the fourth target layer serve as the input features of the next third target layer. Among them, the third target layer can be the h-th layer, the fourth target layer can be the (h + 1)-th layer, and the next third target layer can be the (h + 2)-th layer. That is, the output features of the third target layer, the h-th layer, serve as the input features of the fourth target layer, the (h + 1)-th layer, and the output features of the fourth target layer, the (h + 1)-th layer, serve as the input features of the next third target layer, the (h + 2)-th layer.

[0110] Exemplarily, taking the text sequence with k = 3 as the second input feature and inputting it into the second multi-modal sub-model as an example. For instance, there are 4 text sequence features, namely s0 = [1, 2, 3, 4], s1 = [5, 6, 7, 8], s2 = [9, 10, 11, 12], and s3 = [13, 14, 15, 16]. At the first layer, the input image sequence as the first input feature is calculate the query (Q), key (K), and value (V) through the self-attention mechanism to obtain

[0111]

[0112] After that, calculate and then obtain the output features of the first layer

[0113] Take the output features of the first layer as the input features of the second layer until the output features of the sixth layer are obtained.

[0114] S303. After constructing the first target feature and the second target feature based on the first output feature, and constructing the third target feature based on the second output feature, determine the target comprehensive feature based on the first target feature, the second target feature, and the third target feature;

[0115] In a possible implementation manner, constructing the first target feature and the second target feature based on the first output feature, and constructing the third target feature based on the second output feature specifically include the following steps:

[0116] Take the product of the first output feature and a preset first weight matrix as the first target feature; take the product of the first output feature and a preset second weight matrix as the second target feature; take the product of the second output feature and a preset third weight matrix as the third target feature.

[0117] As an example, the first output feature is determined by the above steps as The second output feature is The first weight matrix W1, the second weight matrix W2, and the third weight matrix W3 are pre-set weight matrices.

[0118] Therefore, the first target feature is calculated as The second target feature

[0119] In a possible implementation manner, as Figure 4 shown, it is a schematic flowchart of another method for training a phishing website detection model provided by an embodiment of the present application. Specifically, it includes the following steps:

[0120] S401. After calculating the first product of the transpose of the first target feature and the third target feature, calculate the first quotient value obtained by dividing the first product by a preset value;

[0121] As an example, the first product = Q3K3 T , where the preset value represents the square root of the dimension of the input vector.

[0122] S402. Calculate the first activation function value of the first quotient value based on an activation function;

[0123] S403. Based on a fully connected layer, calculate the second product obtained by multiplying the first activation function value by the second target feature, and take the second product as the target comprehensive feature.

[0124] As an example, The target comprehensive feature is

[0125] S304. Obtain the first comprehensive feature vector based on the target comprehensive feature and an average pooling method.

[0126] Based on the target comprehensive feature [e0,...e k and the average pooling method, obtain the first comprehensive feature vector For example, the target comprehensive feature is [e0, e1, e2, e3], and the first comprehensive feature vector determined based on the average pooling method

[0127] It should be noted that the method of inputting the suspicious website data in the training dataset into the multi-modal analysis model to obtain the second comprehensive feature vector is similar to the principle of the method used to obtain the first comprehensive feature vector above, and will not be elaborated here.

[0128] In a possible implementation, in order to make the features extracted from different modalities, such as website source code and website screenshot images, have consistency in the shared semantic space, this application processes the features through a hierarchical consistency model, so as to minimize the feature differences between the website source code and website screenshot images at different levels. As Figure 5 shown, it is a schematic flowchart of another method for training a phishing website detection model provided by an embodiment of this application. The method for obtaining the first target consistency loss value and the method for obtaining the second target consistency loss value are similar. Here, taking the input of the target website data in the training dataset into the multi-modal analysis model to obtain the first target consistency loss value as an example for illustration.

[0129] S501. Determine the first consistency loss value based on the intermediate layer output features and the second output features of the first multi-modal sub-model;

[0130] As an example, taking the first multi-modal sub-model as having 12 layers and the second multi-modal sub-model as having 6 layers, the intermediate layer of the first multi-modal sub-model, that is, the output features of the 6th layer are The second output features of the second multi-modal sub-model are Specifically, the calculated first consistency loss value is where n represents that there are n output samples, and i represents taking the i-th sample.

[0131] Exemplarily, taking the intermediate layer output features of the first multi-modal sub-model as as an example, taking the second output features of the second multi-modal sub-model as an example, where the value of k is 4, and defining the value of the number of samples n as 3, the calculated consistency loss value is

[0132] S502. Determine the second consistency loss value based on the first output features and the second output features;

[0133] As an example, taking the first multi-modal sub-model as having 12 layers and the second multi-modal sub-model as having 6 layers, the first output features of the first multi-modal sub-model The second output features of the second multi-modal sub-model are Specifically, the calculated first consistency loss value is where n represents that there are n output samples, and i represents taking the i-th sample.

[0134] Exemplarily, taking the first output feature of the first multi-modal sub-model as an example, and taking the second output feature of the second multi-modal sub-model as an example, where the value of n is 3, the calculated consistency loss value is

[0135] S503. Determine a third consistency loss value based on the target comprehensive feature, the first output feature, and the second output feature;

[0136] As an example, taking the first multi-modal sub-model as 12 layers and the second multi-modal sub-model as 6 layers as an example, the target comprehensive feature is [e0,...e k , the first output feature of the first multi-modal sub-model The second output feature of the second multi-modal sub-model is

[0137] In a specific embodiment, as Figure 6 shown, it is a schematic flowchart of another training method of the phishing website detection model provided by the embodiment of the present application, and the specific steps are as follows:

[0138] S601. For each target website data, calculate the first square value of the difference between the target comprehensive feature vector and the first output feature, and calculate the second square value of the difference between the target comprehensive feature vector and the second output feature;

[0139] As an example, taking the first multi-modal sub-model as 12 layers and the second multi-modal sub-model as 6 layers as an example, the target comprehensive feature is [e0,...e k , the first output feature of the first multi-modal sub-model The second output feature of the second multi-modal sub-model is Then the first square value is The second square value is

[0140] S602. After calculating the first sum value of the first square value and the second square value, calculate the second sum value of the addition of the first sum values of each target website data;

[0141] As an example, where n represents the number of feature values in the output feature.

[0142] S603. After calculating the second quotient value obtained by dividing the second sum value by the number of feature values, use the third quotient value obtained by dividing the second quotient value by the second preset value as the third consistency loss value.

[0143] As an example, the second quotient value is The second preset value is 2, that is, the third consistency loss value is

[0144] Exemplarily, it is defined that the number of samples n = 3, and the i-th target comprehensive feature is e i , and the first output feature of the first multi-modal sub-model is Taking the second output feature of the second multi-modal sub-model as an example, the first squared value is The second squared value is The second sum value is Since n = 3, therefore, the third consistency loss value

[0145] S504. Use the sum value of the first consistency loss value, the second consistency loss value, and the third consistency loss value as the first target consistency loss value.

[0146] Exemplarily, the first target consistency loss value is characterized as loss consistency_origin , and the second target consistency loss value is characterized as loss consistency_sample . The first target consistency loss value loss consistency_sample = loss1 + loss2 + loss3.

[0147] It should be noted that the method of inputting the suspicious website data in the training dataset into the multi-modal analysis model to obtain the second target consistency loss value is similar to the principle of the method used to obtain the first target consistency loss value above, and will not be elaborated here.

[0148] S202. Determine the probability value based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully connected layer, the bias term of the fully connected layer, and the activation function;

[0149] In a specific embodiment, the first comprehensive feature vector and the second comprehensive feature vector are concatenated to obtain a concatenated comprehensive feature vector; after calculating the third product of the weight matrix and the concatenated comprehensive feature vector, calculate the third sum value of adding the third product and the bias term; after calculating the second activation function value of the third sum value based on the activation function, use the second activation function value as the probability value that the suspicious website is a phishing website.

[0150] Exemplarily, taking the first comprehensive feature vector obtained after inputting the target website data into the multi-modal analysis model as e origin , and the second comprehensive feature vector obtained after inputting the suspicious website data into the multi-modal analysis model as e sample , and the weight matrix is W and the bias term is b as an example, the concatenated comprehensive feature vector is [e origin , e sample. The third product is W × [e origin , e sample , and the third sum value is W × [e origin , e sample + b. The probability value that the suspicious website calculated therefrom is a phishing website

[0151] S203. After calculating the cross - entropy loss function value based on the probability value and the cross - entropy loss function, calculate the total loss function value based on the cross - entropy loss function value, the first target consistency loss value, and the second target consistency loss value;

[0152] In a specific embodiment, calculate the first difference between 1 and the third preset value, the second difference between 1 and the probability value, and the first derivative value of the probability value; calculate the second derivative value of the second difference; calculate the fourth product of the third preset value multiplied by the first derivative value, and the fifth product of the first difference multiplied by the second derivative value, and then calculate the fourth sum value of the addition of the fourth product and the fifth product; take the negative value of the fourth sum value as the cross - entropy loss function value.

[0153] Exemplarily, taking the probability value as and the third preset value as y, the first difference is (1 - y), the second difference is and the first derivative value is The second derivative value is The fourth product is The fifth product is The cross - entropy loss function value calculated therefrom

[0154] In a specific embodiment, the total loss function value Loss is the sum of the cross - entropy loss function value loss ce , the first target consistency loss value loss consistency_origin and the second target consistency loss value loss consistency_sample .

[0155] Specifically, the calculation formula of the total loss function value Loss is:

[0156] Loss = loss ce + loss consistency_origin + loss consistency_sample .

[0157] S204. Calculate the optimized weight matrix and optimized bias term based on the gradient descent algorithm and the total loss function value to realize the training of the multi - modal analysis model and obtain the phishing website detection model.

[0158] In a specific embodiment, at the initial stage of model training, the weight matrix and bias term are set to randomly initialized values. For example, a standard normal distribution (mean = 0, variance = 1) can be used to initialize the weights, and constants such as 0, 1, or 2 can be used to initialize the bias term. During the training process of the multi-modal analysis model, the optimized weight matrix and optimized bias term are calculated based on the gradient descent algorithm and the total loss function value, thereby improving the accuracy of the multi-modal analysis model in detecting phishing websites.

[0159] Furthermore, for a better introduction of the embodiments of the present application, an example is given based on Figure 7 and Figure 8 the training process of the phishing website detection model shown. It should be noted that the following examples in the embodiments of the present application are only for illustration and do not constitute a limitation to the embodiments of the present application:

[0160] Exemplarily, as Figure 7 shown, it is a schematic structural diagram of a training method for a phishing website detection model provided by an embodiment of the present application. The image sequence corresponding to the target website data is used as the first input feature and input into the first multi-modal sub-model to obtain the first output feature. The first multi-modal sub-model is composed of a 12-layer self-attention mechanism (Multi-head Attention) and a feed-forward neural network (FFN). The text sequence corresponding to the target website data is used as the second input feature and input into the second multi-modal sub-model to obtain the second output feature. The second multi-modal sub-model is composed of a 6-layer self-attention mechanism (Multi-head Attention) and a feed-forward neural network (FFN). After constructing the first target feature and the second target feature based on the first output feature, and constructing the third target feature based on the second output feature, the target comprehensive feature is determined based on the first target feature, the second target feature, and the third target feature. Among them, the target comprehensive feature is determined through a feed-forward neural network and a cross-attention mechanism, and the first comprehensive feature vector e is obtained based on the target comprehensive feature and the average pooling method.

[0161] As Figure 7 shown, the first consistency loss value CL1 is determined based on the intermediate layer output feature and the second output feature of the first multi-modal sub-model, the second consistency loss value CL2 is determined based on the first output feature and the second output feature, and the third consistency loss value CL3 is determined based on the target comprehensive feature, the first output feature, and the second output feature.

[0162] Exemplarily, as Figure 8As shown in the figure, it is a schematic structural diagram of another method for training a phishing website detection model provided by an embodiment of the present application. Input the target website data in the training dataset into the multi-modal analysis model to obtain the first target consistency loss value, and input the suspicious website data in the training dataset into the multi-modal analysis model to obtain the second target consistency loss value; calculate the total loss function value based on the cross-entropy loss function value, the first target consistency loss value, and the second target consistency loss value; calculate the optimized weight matrix and optimized bias term based on the gradient descent algorithm and the total loss function value to implement the training of the multi-modal analysis model and obtain the phishing website detection model.

[0163] An embodiment of the present application provides a phishing website detection method. This method trains a multi-modal analysis model to obtain a phishing website detection model. When obtaining the website data to be detected; input the website data to be detected into the phishing website detection model, and output the category corresponding to the website data to be detected, where the category is used to characterize whether the website to be detected is a phishing website; specifically, the multi-modal analysis model is trained by the following method: First, input the target website data in the training dataset into the multi-modal analysis model to obtain the first comprehensive feature vector and the first target consistency loss value, and input the suspicious website data in the training dataset into the multi-modal analysis model to obtain the second comprehensive feature vector and the second target consistency loss value; Second, based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully connected layer, the bias term of the fully connected layer, and the activation function, determine the probability value that the suspicious website is a phishing website; Finally, after calculating the cross-entropy loss function value based on the probability value and the cross-entropy loss function, calculate the total loss function value based on the cross-entropy loss function value, the first target consistency loss value, and the second target consistency loss value; calculate the optimized weight matrix and optimized bias term based on the gradient descent algorithm and the total loss function value to implement the training of the multi-modal analysis model and obtain the phishing website detection model. The present application combines the website source code and the website screenshot image as the information source, and improves the accuracy of phishing website detection through the training of the multi-modal analysis model, thereby enhancing the network security of users.

[0164] Based on the same inventive concept, an embodiment of the present application also provides a phishing website detection device. The principle of this phishing website detection device is similar to the above-mentioned phishing website detection method, and the repeated parts will not be elaborated. As Figure 9 shown, it is a schematic diagram of a phishing website detection device provided by an embodiment of the present application, including:

[0165] A category determination module 901, configured to obtain the website data to be detected; input the website data to be detected into the phishing website detection model, and output the category corresponding to the website data to be detected, where the category is used to characterize whether the website to be detected is a phishing website;

[0166] The model training module 902 is configured to input the target website data in the training dataset into the multi-modal analysis model to obtain a first comprehensive feature vector and a first target consistency loss value, and input the suspicious website data in the training dataset into the multi-modal analysis model to obtain a second comprehensive feature vector and a second target consistency loss value; determine a probability value based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully connected layer, the bias term of the fully connected layer, and the activation function; calculate a cross-entropy loss function value based on the probability value and the cross-entropy loss function, and then calculate a total loss function value based on the cross-entropy loss function value, the first target consistency loss value, and the second target consistency loss value; calculate an optimized weight matrix and an optimized bias term based on the gradient descent algorithm and the total loss function value to implement the training of the multi-modal analysis model and obtain the phishing website detection model.

[0167] An embodiment of the present application provides a phishing website detection method and device. The present application trains a multi-modal analysis model to obtain a phishing website detection model. When obtaining data of a website to be detected, the data of the website to be detected is input into the phishing website detection model, and a category corresponding to the data of the website to be detected is output, where the category is used to represent whether the website to be detected is a phishing website. Specifically, the multi-modal analysis model is trained by the following method: First, input the target website data in the training dataset into the multi-modal analysis model to obtain a first comprehensive feature vector and a first target consistency loss value, and input the suspicious website data in the training dataset into the multi-modal analysis model to obtain a second comprehensive feature vector and a second target consistency loss value; Second, determine the probability value that the suspicious website is a phishing website based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully connected layer, the bias term of the fully connected layer, and the activation function; Finally, calculate a cross-entropy loss function value based on the probability value and the cross-entropy loss function, and then calculate a total loss function value based on the cross-entropy loss function value, the first target consistency loss value, and the second target consistency loss value; calculate an optimized weight matrix and an optimized bias term based on the gradient descent algorithm and the total loss function value to implement the training of the multi-modal analysis model and obtain the phishing website detection model. The present application combines the website source code and the website screenshot image as information sources, and improves the accuracy of phishing website detection by training the multi-modal analysis model, thereby enhancing the network security of users.

[0168] In a possible implementation manner, the above model training module 902 is specifically configured to:

[0169] Use the image sequence corresponding to the target website data as the first input feature and input it into the first multi-modal sub-model to obtain a first output feature. Among them, the first multi-modal sub-model includes multiple first target layers and multiple second target layers. The output feature of the first target layer is used as the input feature of the second target layer, and the output feature of the second target layer is used as the input feature of the next first target layer;

[0170] Use the text sequence corresponding to the target website data as the second input feature and input it into the second multi-modal sub-model to obtain a second output feature. Among them, the second multi-modal sub-model includes multiple third target layers and multiple fourth target layers. The output feature of the third target layer is used as the input feature of the fourth target layer, and the output feature of the fourth target layer is used as the input feature of the next third target layer;

[0171] After constructing a first target feature and a second target feature based on the first output feature, and constructing a third target feature based on the second output feature, determine a target comprehensive feature based on the first target feature, the second target feature, and the third target feature;

[0172] Obtain the first comprehensive feature vector based on the target comprehensive feature and the average pooling method.

[0173] In a possible implementation manner, the above model training module 902 is specifically used for:

[0174] Use the product of the first output feature and a preset first weight matrix as the first target feature;

[0175] Use the product of the first output feature and a preset second weight matrix as the second target feature;

[0176] Use the product of the second output feature and a preset third weight matrix as the third target feature.

[0177] In a possible implementation manner, the above model training module 902 is specifically used for:

[0178] After calculating a first product obtained by multiplying the transpose of the first target feature by the third target feature, calculate a first quotient value obtained by dividing the first product by a first preset value;

[0179] Calculate a first activation function value of the first quotient value based on the activation function;

[0180] Based on a fully connected layer, calculate a second product obtained by multiplying the first activation function value by the second target feature, and use the second product as the target comprehensive feature.

[0181] In a possible implementation, the above-mentioned model training module 902 is specifically configured to:

[0182] Determine a first consistency loss value based on the intermediate layer output features of the first multi-modal sub-model and the second output feature;

[0183] Determine a second consistency loss value based on the first output feature and the second output feature;

[0184] Determine a third consistency loss value based on the target comprehensive feature, the first output feature, and the second output feature;

[0185] Take the sum of the first consistency loss value, the second consistency loss value, and the third consistency loss value as the first target consistency loss value.

[0186] In a possible implementation, the above-mentioned model training module 902 is specifically configured to:

[0187] For each target website data, calculate the first squared value of the difference between the target comprehensive feature vector and the first output feature, and calculate the second squared value of the difference between the target comprehensive feature vector and the second output feature;

[0188] After calculating the first sum value of the first squared value and the second squared value, calculate the second sum value obtained by adding the first sum values of each target website data;

[0189] After calculating the second quotient value obtained by dividing the second sum value by the number of feature values, take the third quotient value obtained by dividing the second quotient value by a second preset value as the third consistency loss value.

[0190] In a possible implementation, the above-mentioned model training module 902 is specifically configured to:

[0191] Concatenate the first comprehensive feature vector and the second comprehensive feature vector to obtain a concatenated comprehensive feature vector;

[0192] Calculate the third product of the weight matrix and the concatenated comprehensive feature vector, and then calculate the third sum value obtained by adding the third product and the bias term;

[0193] Based on an activation function, calculate the second activation function value of the third sum value, and take the second activation function value as the probability value.

[0194] In a possible implementation, the above-mentioned model training module 902 is specifically configured to:

[0195] Calculate a first difference by subtracting a third preset value from 1, calculate a second difference by subtracting the probability value from 1, and calculate a first derivative value of the probability value;

[0196] Calculate a second derivative value of the second difference;

[0197] After calculating a fourth product by multiplying the third preset value by the first derivative value and calculating a fifth product by multiplying the first difference by the second derivative value, calculate a fourth sum value by adding the fourth product and the fifth product;

[0198] Use the negative value of the fourth sum value as the cross-entropy loss function value.

[0199] Based on the same inventive concept, an embodiment of the present application further provides a phishing website detection device. The principle of this phishing website detection device is similar to that of the above phishing website detection method, and the repeated parts will not be elaborated here.

[0200] As Figure 10 shown, it is a schematic diagram of a phishing website detection device provided by an embodiment of the present application. The device includes at least one processor 1001; and a memory 1002 communicatively connected to the at least one processor; in the embodiment of the present application, the memory 1002 stores instructions executable by the at least one processor 1001, and the instructions are executed by the at least one processor 1001 to enable the at least one processor 1001 to execute the phishing website detection method in the above embodiment.

[0201] The processor 1001 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 1001 is used to implement the phishing website detection method in the above embodiment when calling a computer program stored in the memory 1002.

[0202] In a possible design, the memory 1002 may be a volatile memory, such as a random-access memory (RAM); the memory 1002 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD), or the memory 1002 is any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1002 may be a combination of the above memories.

[0203] In the embodiments of the present application, the specific connection medium between the above-mentioned memory 1002 and the processor 1001 is not limited. In the embodiments of the present application, Figure 10 it is shown that the memory 1002 and the processor 1001 are connected through a bus 1003. The bus 1003 is represented by a thick line in Figure 10 the figure. The connection manners between other components are only for illustrative purposes and are not restrictive. The bus 1003 can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, Figure 10 only one thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0204] The embodiments of the present application provide a phishing website detection method, device and equipment. In the present application, a phishing website detection model is obtained by training a multimodal analysis model. When obtaining data of a website to be detected, the data of the website to be detected is input into the phishing website detection model, and a category corresponding to the data of the website to be detected is output, where the category is used to represent whether the website to be detected is a phishing website. Specifically, the multimodal analysis model is trained by the following method: First, the target website data in the training data set is input into the multimodal analysis model to obtain a first comprehensive feature vector and a first target consistency loss value, and the suspicious website data in the training data set is input into the multimodal analysis model to obtain a second comprehensive feature vector and a second target consistency loss value. Secondly, based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully connected layer, the bias term of the fully connected layer and the activation function, the probability value that the suspicious website is a phishing website is determined. Finally, after calculating the cross-entropy loss function value based on the probability value and the cross-entropy loss function, the total loss function value is calculated based on the cross-entropy loss function value, the first target consistency loss value and the second target consistency loss value. The optimized weight matrix and the optimized bias term are calculated based on the gradient descent algorithm and the total loss function value to implement the training of the multimodal analysis model and obtain the phishing website detection model. In the present application, the website source code and the website screenshot image are combined as information sources. By training the multimodal analysis model, the accuracy of phishing website detection is improved, thereby enhancing the network security of users.

[0205] The present application is described above with reference to the block diagrams and / or flowcharts showing methods, apparatuses (systems) and / or computer program products according to embodiments of the present application. It should be understood that one block of the block diagrams and / or flowcharts and combinations of blocks in the block diagrams and / or flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, and / or other programmable data processing devices to generate a machine, so that the instructions executed via the computer processor and / or other programmable data processing devices create a method for implementing the functions / actions specified in the blocks of the block diagrams and / or flowcharts.

[0206] Accordingly, the present application can also be implemented by hardware and / or software (including firmware, resident software, microcode, etc.). Further, the present application can take the form of a computer program product on a computer-usable or computer-readable storage medium, having computer-usable or computer-readable program code embodied in the medium for use by or in connection with an instruction execution system. In the context of the present application, a computer-usable or computer-readable medium can be any medium that can contain, store, communicate, transmit, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0207] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.

Claims

1. A method for detecting phishing websites, characterized in that: The method comprises: Get the data of the website to be tested; Inputting the website data to be detected into a phishing website detection model, and outputting a category corresponding to the website data to be detected, wherein the category is used to characterize whether the website to be detected is a phishing website; The phishing website detection model is trained by the following method: Inputting the target website data in the training data set into the multimodal analysis model to obtain a first comprehensive feature vector and a first target consistency loss value, and inputting the suspicious website data in the training data set into the multimodal analysis model to obtain a second comprehensive feature vector and a second target consistency loss value; Determining a probability value based on the first comprehensive feature vector, the second comprehensive feature vector, a weight matrix of a fully connected layer, a bias term of the fully connected layer, and an activation function; After a cross entropy loss function value is calculated based on the probability value and the cross entropy loss function, a total loss function value is calculated based on the cross entropy loss function value, the first target consistency loss value, and the second target consistency loss value; Based on the gradient descent algorithm and the total loss function value, an optimized weight matrix and an optimized bias term are calculated to train the multimodal analysis model and obtain the phishing website detection model.

2. The method according to claim 1, characterized in that The target website data in the training data set is input into the multimodal analysis model to obtain a first comprehensive feature vector, including: Inputting the image sequence corresponding to the target website data as the first input feature into the first multimodal sub-model to obtain the first output feature, wherein the first multimodal sub-model includes a plurality of first target layers and a plurality of second target layers, the output feature of the first target layer is used as the input feature of the second target layer, and the output feature of the second target layer is used as the input feature of the next first target layer; Inputting the text sequence corresponding to the target website data as a second input feature into the second multimodal sub-model to obtain a second output feature, wherein the second multimodal sub-model includes a plurality of third target layers and a plurality of fourth target layers, the output feature of the third target layer is used as the input feature of the fourth target layer, and the output feature of the fourth target layer is used as the input feature of the next third target layer; After constructing a first target feature and a second target feature based on the first output feature, and constructing a third target feature based on the second output feature, determining a target comprehensive feature based on the first target feature, the second target feature, and the third target feature; The first comprehensive feature vector is obtained based on the target comprehensive features and the average pooling method.

3. The method according to claim 2, characterized in that The step of constructing a first target feature and a second target feature based on the first output feature, and constructing a third target feature based on the second output feature, comprises: Taking the product of the first output feature and a preset first weight matrix as the first target feature; The product of the first output feature and a preset second weight matrix is ​​used as the second target feature; The product of the second output feature and a preset third weight matrix is ​​used as the third target feature.

4. The method according to claim 2, characterized in that The determining of a target comprehensive feature based on the first target feature, the second target feature and the third target feature comprises: After calculating a first product of the transpose of the first target feature and the third target feature, calculating a first quotient of the first product divided by a first preset value; Calculate a first activation function value of the first quotient value based on the activation function; Based on the fully connected layer, a second product obtained by multiplying the first activation function value and the second target feature is calculated, and the second product is used as the target comprehensive feature.

5. The method according to claim 2, characterized in that The step of inputting the target website data in the training data set into the multimodal analysis model to obtain a first target consistency loss value includes: Determining a first consistency loss value based on the intermediate layer output feature of the first multimodal sub-model and the second output feature; determining a second consistency loss value based on the first output feature and the second output feature; Determining a third consistency loss value based on the target comprehensive feature, the first output feature, and the second output feature; The sum of the first consistency loss value, the second consistency loss value and the third consistency loss value is used as the first target consistency loss value.

6. The method according to claim 5, characterized in that The determining a third consistency loss value based on the target comprehensive feature, the first output feature, and the second output feature comprises: For each target website data, calculating a first square value of a difference between the target comprehensive feature vector and the first output feature, and calculating a second square value of a difference between the target comprehensive feature vector and the second output feature; After calculating a first sum of the first square value and the second square value, calculating a second sum of the first sum of each target website data; After calculating a second quotient of the second sum and the number of eigenvalues, a third quotient of the second quotient and a second preset value is used as the third consistency loss value.

7. The method according to claim 1, characterized in that The determining of the probability value based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully connected layer, the bias term of the fully connected layer and the activation function includes: Concatenating the first comprehensive feature vector and the second comprehensive feature vector to obtain a concatenated comprehensive feature vector; After calculating a third product of the weight matrix and the concatenated comprehensive feature vector, calculating a third sum of the third product and the bias term; After calculating the second activation function value of the third sum value based on the activation function, the second activation function value is used as the probability value.

8. The method according to any one of claims 1 to 7, characterized in that: The calculating the cross entropy loss function value based on the probability value and the cross entropy loss function includes: Calculating a first difference between 1 and a third preset value, calculating a second difference between 1 and the probability value, and calculating a first derivative value of the probability value; calculating a second derivative value of the second difference; After calculating a fourth product of the third preset value and the first derivative value and a fifth product of the first difference and the second derivative value, calculating a fourth sum of the fourth product and the fifth product; The negative value of the fourth sum is used as the cross entropy loss function value.

9. A phishing website detection device, characterized in that: The device comprises: A category determination module is used to obtain the website data to be detected; input the website data to be detected into a phishing website detection model, and output a category corresponding to the website data to be detected, wherein the category is used to characterize whether the website to be detected is a phishing website; A model training module is used to input the target website data in the training data set into the multimodal analysis model to obtain a first comprehensive feature vector and a first target consistency loss value, and to input the suspicious website data in the training data set into the multimodal analysis model to obtain a second comprehensive feature vector and a second target consistency loss value; determine the probability value based on the first comprehensive feature vector, the second comprehensive feature vector, the weight matrix of the fully connected layer, the bias term of the fully connected layer, and the activation function; calculate the cross entropy loss function value based on the probability value and the cross entropy loss function, and then calculate the total loss function value based on the cross entropy loss function value, the first target consistency loss value, and the second target consistency loss value; calculate the optimized weight matrix and the optimized bias term based on the gradient descent algorithm and the total loss function value to realize the training of the multimodal analysis model and obtain the phishing website detection model.

10. A phishing website detection device, characterized in that: The invention comprises at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 8.