Phishing website detection method and device, electronic equipment and storage medium

By generating cross-feature and multimodal feature fusion models and combining discrete feature fusion strategies, the problem of low detection accuracy of phishing websites in the existing technology is solved, and higher detection accuracy and robustness are achieved.

CN120110784APending Publication Date: 2025-06-06CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510365969.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the detection of phishing websites, the recognition accuracy rate is significantly reduced in the face of complex attacks, making it difficult to effectively improve the detection accuracy.

Method used

By generating cross-features, multimodal feature fusion models and discrete feature fusion strategies, the basic information, source code position relationships and functional characteristics of the website to be detected and the target website are fused to improve the robustness and accuracy of the detection.

Benefits of technology

It significantly improves the accuracy of phishing website detection, especially when facing complex and mutated phishing attacks, reduces the false alarm rate and enhances detection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120110784A_ABST
    Figure CN120110784A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a phishing website detection method and device, electronic equipment and a storage medium, and the method comprises the steps: generating a cross feature based on to-be-detected basic information of a to-be-detected website and target basic information of a target website; respectively inputting the source code of the to-be-detected website and the source code of the target website into a multi-modal feature fusion model, and obtaining a first modal feature of the to-be-detected website and a second modal feature of the target website; and fusing the cross feature, the first modal feature and the second modal feature to obtain a fused feature, and determining a detection result for the to-be-detected website based on the fused feature and a discrete feature fusion strategy. According to the method, the cross features of the target website and the to-be-detected website are fused into the modal features, with the position information, of the target website and the to-be-detected website, so that codes of the modal features are richer, and the accuracy of phishing website detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security, and in particular to a phishing website detection method, device, electronic device and storage medium. Background Art

[0002] With the rapid development of the Internet, the task of phishing website detection has become particularly important. Phishing attacks not only damage customers' privacy and financial security, but also pose a great threat to corporate reputation and trust. Phishing websites disguise themselves as legitimate websites to trick customers into entering sensitive information, causing significant economic losses. Therefore, improving the detection accuracy of phishing websites, especially the ability to identify phishing websites in complex attack scenarios, has become an important goal in the current network security field.

[0003] In the related art, when detecting phishing websites, the detection is mainly carried out by extracting a single feature of the website to be detected and the legitimate website for comparison, for example, only by comparing the website icons of the website to be detected and the legitimate website. Obviously, this detection method based on a single feature has a significantly reduced recognition accuracy when facing complex attacks.

[0004] In summary, how to improve the accuracy of phishing website detection is an urgent problem to be solved. Summary of the invention

[0005] The embodiments of the present application provide a phishing website detection method, device, electronic device and storage medium to improve the accuracy of phishing website detection.

[0006] In a first aspect, an embodiment of the present application provides a phishing website detection method, comprising:

[0007] Generate cross-features based on the basic information of the website to be detected and the target basic information of the target website; the basic information to be detected is used to characterize the authenticity and security of the website to be detected; the target basic information is used to characterize the authenticity and security of the target website;

[0008] Inputting the source code of the website to be detected and the source code of the target website into the multimodal feature fusion model respectively, obtaining the first modal feature of the website to be detected and the second modal feature of the target website; the first modal feature includes the positional relationship and function of each element in the website to be detected in the website to be detected; the second modal feature includes the positional relationship and function of each element in the target website in the target website;

[0009] The cross-features, the first modal features and the second modal features are integrated to obtain fused features, and based on the fused features and the discrete feature fusion strategy, a detection result for the website to be detected is determined; the detection result is used to characterize whether the website to be detected is a phishing website of the target website.

[0010] In a second aspect, an embodiment of the present application provides a phishing website detection device, including:

[0011] A generating unit, configured to generate a cross feature based on basic information of the website to be detected and target basic information of the target website; the basic information to be detected is used to characterize the authenticity and security of the website to be detected; the target basic information is used to characterize the authenticity and security of the target website;

[0012] An acquisition unit is used to input the source code of the website to be detected and the source code of the target website into a multimodal feature fusion model respectively, and acquire a first modal feature of the website to be detected and a second modal feature of the target website; the first modal feature includes the positional relationship and function of each element in the website to be detected in the website to be detected; the second modal feature includes the positional relationship and function of each element in the target website in the target website;

[0013] A detection unit is used to fuse the cross-features, the first modal features and the second modal features to obtain fused features, and determine the detection results for the website to be detected based on the fused features and the discrete feature fusion strategy; the detection results are used to characterize whether the website to be detected is a phishing website of the target website.

[0014] In some embodiments, the acquisition unit is specifically used to:

[0015] Converting the source code of the website to be detected into a first model tree; the nodes in the first model tree are text elements in the source code of the website to be detected, or image elements in the source code of the website to be detected;

[0016] Based on the relative position of each node in the first model tree, at least one first sequence corresponding to the website to be detected is obtained; the elements contained in each first sequence are text feature vectors corresponding to text elements or image feature vectors corresponding to image elements that have the same parent node in the first model tree and are located in the same layer in the first model tree;

[0017] For each of the first sequences, the following iterative aggregation operation is performed until the first feature vector of the first parent node of the first layer in the first model tree is obtained:

[0018] Re-encode the elements in the first sequence by using a multi-head attention mechanism, and aggregate each of the re-encoded first sequences with a first parent node corresponding to the first sequence to obtain a first aggregation result;

[0019] Concatenate the aggregation result with the first parent node to obtain a first feature vector of the first parent node;

[0020] The first feature vector of the first parent node of the first layer in the first model tree is used as the first modal feature.

[0021] In some embodiments, the acquisition unit is specifically used to:

[0022] Performing a linear transformation on the first parent node to obtain a first transformation result;

[0023] Perform a dot product of the first transformation result and each element in the first sequence to obtain a first dot product sequence; the first dot product sequence includes at least one first element;

[0024] For each of the first elements, convert the first element into a second element by using a normalized exponential function; multiply the first element by the second element to obtain a first multiplication result;

[0025] A sequence formed based on at least one of the first multiplication results is used as the first aggregation result.

[0026] In some embodiments, the acquisition unit is specifically used to:

[0027] Converting the source code of the target website into a second model tree; the nodes in the second model tree are text elements in the source code of the target website, or image elements in the source code of the target website;

[0028] Based on the relative position of each node in the second model tree, obtaining at least one second sequence corresponding to the target website; the elements contained in each second sequence are text feature vectors corresponding to text elements or image feature vectors corresponding to image elements that have the same parent node in the second model tree and are located in the same layer in the second model tree;

[0029] For each of the second sequences, the following iterative aggregation operation is performed until the second feature vector of the second parent node of the first layer in the second model tree is obtained:

[0030] re-encode the elements in the second sequence by using a multi-head attention mechanism, and aggregate each re-encoded second sequence with a second parent node corresponding to the second sequence to obtain a second aggregation result;

[0031] Concatenate the aggregation result with the second parent node to obtain a second feature vector of the second parent node;

[0032] The second feature vector of the second parent node of the first layer in the second model tree is used as the second modal feature.

[0033] In some embodiments, the acquisition unit is specifically used to:

[0034] Performing a linear transformation on the second parent node to obtain a second transformation result;

[0035] Perform a dot product of the second transformation result and each element in the second sequence to obtain a second dot product sequence; the second dot product sequence includes at least one third element;

[0036] For each of the third elements, convert the third element into a fourth element by using a normalized exponential function; multiply the third element by the fourth element to obtain a second multiplication result;

[0037] A sequence formed based on at least one of the second multiplication results is used as the second aggregation result.

[0038] In some embodiments, the generating unit is specifically used for:

[0039] If the basic information to be detected includes a title to be detected, and the target basic information includes a target title, then the cross feature includes the title similarity between the title to be detected and the target title;

[0040] If the basic information to be detected includes an icon to be detected, and the target basic information includes a target icon, then the intersection feature includes an icon similarity between the icon to be detected and the target icon;

[0041] If the basic information to be detected includes a website certificate to be detected, and the target basic information includes a target website certificate, then the cross-feature includes the website certificate similarity between the website certificate to be detected and the target website certificate;

[0042] If the basic information to be detected includes the domain name of the website to be detected, and the target basic information includes the domain name of the target website, then the cross feature includes the edit distance between the domain name of the website to be detected and the domain name of the target website.

[0043] In some embodiments, the detection unit is specifically used for:

[0044] Determine the cosine similarity between the first modal feature and the second modal feature, and use the cosine similarity as a new feature;

[0045] Inputting the new features and the cross features into a deep neural network model to obtain fused features;

[0046] The determining of the detection result for the website to be detected based on the fusion feature, the first modal feature, the second modal feature and the discrete feature fusion strategy includes:

[0047] Splicing the first modal feature and the second modal feature to obtain a first splicing feature;

[0048] After performing a linear transformation on the first concatenated features, the linearly transformed first concatenated features are input into a fully connected layer to obtain a third modal feature;

[0049] Splicing the fusion feature and the third modality feature to obtain a second splicing feature;

[0050] The second concatenated feature is input into the fully connected layer, and the output of the fully connected layer is nonlinearly transformed to determine a detection result for the website to be detected.

[0051] In a third aspect, an embodiment of the present application provides an electronic device, including:

[0052] A memory for storing program instructions;

[0053] The processor is used to call the program instructions stored in the memory and execute the above-mentioned phishing website detection method according to the obtained program instructions.

[0054] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the above-mentioned phishing website detection method is implemented.

[0055] In the fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, which is stored in a computer-readable storage medium; when the processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device performs the above-mentioned phishing website detection method.

[0056] The embodiments of the present application provide a phishing website detection method, device, electronic device and storage medium. The present application makes the encoding of modal features richer by integrating the cross-features of the target website and the website to be detected into the modal features with location information of the target website and the website to be detected, retains the relationship between website elements in a structured manner, thereby forming a comprehensive feature representation, and also provides prior information for the modal features, further improving the detection effect. Thereby, potential phishing website features can be better captured, thereby improving the robustness and detection capability of the detection, especially when facing complex and variable phishing attacks, the false alarm rate can be significantly reduced. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 A schematic diagram of an application scenario of a phishing website detection method provided in an embodiment of the present application;

[0058] Figure 2 A flowchart of a phishing website detection method provided in an embodiment of the present application;

[0059] Figure 3 A schematic diagram of the structure of a DOM tree provided in an embodiment of the present application;

[0060] Figure 4 A schematic diagram of obtaining a first aggregation result provided in an embodiment of the present application;

[0061] Figure 5 A schematic diagram of the structure of an aggregator provided in an embodiment of the present application;

[0062] Figure 6 A schematic diagram of MLPD training provided in an embodiment of the present application;

[0063] Figure 7 A flowchart of another phishing website detection method provided in an embodiment of the present application;

[0064] Figure 8 A schematic diagram of the structure of an electronic device for a phishing website detection method in an embodiment of the present application;

[0065] Fig. 9 A schematic diagram of a hardware structure of an electronic device to which an embodiment of the present application is applied;

[0066] Fig.10 A schematic diagram of the hardware composition structure of a computing device using an embodiment of the present application. DETAILED DESCRIPTION

[0067] In order to make the purpose, technical solutions and advantages of this application clearer, this application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0068] It should be noted that the application scenarios described in the following embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. A person of ordinary skill in the art can know that with the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0069] The following first briefly introduces the application scenarios to which the technical solutions of the embodiments of the present application can be applied. It should be noted that the application scenarios introduced below are only used to illustrate the embodiments of the present application and are not limited. In specific implementation, the technical solutions provided by the embodiments of the present application can be flexibly applied according to actual needs.

[0070] like Figure 1 As shown, it is a schematic diagram of an application scenario of a phishing website detection method provided in an embodiment of the present application. The application scenario diagram includes a terminal device 110 and a server 120.

[0071] It should be noted that the phishing website detection method in the embodiment of the present application can be executed by an electronic device, which may be a server 120 or a terminal device 110, that is, the method can be executed by the server 120 or the terminal device 110 alone, or can be executed jointly by the server 120 and the terminal device 110. For example, when the server 120 and the terminal device 110 execute together, the object inputs the target website in the terminal device 110, and the terminal device 110 sends the target website to the server 120. The server 120 generates a website to be detected based on the target website, and generates a cross-feature based on the basic information to be detected of the website to be detected and the target basic information of the target website; then, the server 120 inputs the source code of the website to be detected and the source code of the target website into the multimodal feature fusion model respectively, and obtains the first modal feature of the website to be detected and the second modal feature of the target website; finally, the server 120 fuses the cross-feature, the first modal feature and the second modal feature to obtain the fused feature, and determines the detection result for the website to be detected based on the fused feature and the discrete feature fusion strategy, and the server 120 feeds the detection result back to the terminal device 110 so that the object can view it on the terminal device 110.

[0072] In an optional implementation, the terminal device 110 and the server 120 may communicate with each other via a communication network.

[0073] In an optional implementation, the communication network is a wired network or a wireless network.

[0074] It should be noted that Figure 1 The figure is only an example. In fact, the number of terminal devices and servers is not limited and is not specifically limited in the embodiments of the present application.

[0075] Figure 2 A flowchart of a phishing website detection method provided in an embodiment of the present application is shown. Figure 2 As shown, the method may include the following steps S21 to S23:

[0076] S21: Generate cross-features based on the basic information of the website to be detected and the target basic information of the target website.

[0077] The basic information to be detected is used to characterize the authenticity and security of the website to be detected; the target basic information is used to characterize the authenticity and security of the target website.

[0078] Specifically, the basic information of a website may be an icon, title, website domain name, website certificate, and other website identifiers that can be used to identify the authenticity and security of the website.

[0079] Among them, the website to be detected is generated based on the target website. In actual operation, in order to detect phishing websites, a similar website with some of the same content as the target website is generated based on the domain name, page and other contents of the target website. For example, a suspicious domain name can be generated based on the domain name of the target website's uniform resource locator (URL) through a domain name fuzzy algorithm, and then a website with a domain name that exists in the suspicious domain name is determined as the website to be detected. The domain name of the website to be detected has some of the same content as the domain name of the target website. Similarly, a website to be detected that contains something similar to the page can also be generated based on the page of the target website, and so on.

[0080] In the embodiment of the present application, the cross feature is a feature generated based on the difference between basic information of at least one target website and the website to be detected.

[0081] Specifically, if the basic information to be detected includes the title to be detected, and the target basic information includes the target title, then the cross feature includes the title similarity between the title to be detected and the target title;

[0082] For example, keywords may be extracted from the titles of the website to be detected and the target website respectively, and the number of repeated keywords between the two may be recorded as the title similarity.

[0083] If the basic information to be detected includes an icon to be detected, and the target basic information includes a target icon, then the cross feature includes an icon similarity between the icon to be detected and the target icon;

[0084] For example, a deep convolutional neural network (resnet) can be used to encode the icon, and the cosine similarity value between the icon of the website to be detected and the icon of the target website can be calculated as the icon similarity.

[0085] If the basic information to be detected includes the certificate of the website to be detected, and the target basic information includes the certificate of the target website, then the cross-feature includes the website certificate similarity between the certificate of the website to be detected and the certificate of the target website;

[0086] If the basic information to be detected includes the domain name of the website to be detected, and the target basic information includes the domain name of the target website, then the cross feature includes the edit distance between the domain name of the website to be detected and the domain name of the target website.

[0087] S22: inputting the source code of the website to be detected and the source code of the target website into the multimodal feature fusion model respectively, and obtaining the first modal feature of the website to be detected and the second modal feature of the target website.

[0088] The first modal feature includes the positional relationship and function of each element in the website to be detected; the second modal feature includes the positional relationship and function of each element in the target website.

[0089] The embodiment of the present application designs a phishing website detection method (MLPD) that combines location information and multimodal features. The method includes the above-mentioned multimodal feature fusion model (MLF-Loc). The multimodal feature fusion model uses the website's model (Document Object Model, DOM) tree to record the location information of each node in the DOM tree to generate more comprehensive modal features.

[0090] In the embodiment of the present application, when the first modal feature is generated according to the multimodal feature fusion model, an optional implementation method is as follows:

[0091] Step 1: Convert the source code of the website to be detected into a first model tree;

[0092] The nodes in the first model tree are text elements in the source code of the website to be detected, or image elements in the source code of the website to be detected.

[0093] Specifically, extract the source code of the website to be tested and convert it into Figure 3 The DOM tree structure shown.

[0094] Step 2: Based on the relative position of each node in the first model tree, obtain at least one first sequence corresponding to the website to be detected.

[0095] The elements included in each first sequence are text feature vectors corresponding to text elements or image feature vectors corresponding to image elements that have the same parent node in the first model tree and are located at the same layer in the first model tree.

[0096] Specifically, the DOM tree is traversed in breadth first, and the text elements in the DOM tree are input into a bidirectional encoder (Bidirectional Encoder Representations from Transformers, Bert) to obtain the text feature vector of the text; Crawling the URL under the tag, obtaining the image element, inputting it into the resnet model to obtain the image feature vector of the image. Recording the relative positions of the text element and the image element in the DOM tree, converting the text feature vector or the image feature vector into a sequence.

[0097] Assume that the DOM tree has H layers, and for the hth layer with the same parent node (assuming )'s i-th to j-th text elements or image elements constitute the first sequence:

[0098] Step 3: For each first sequence, perform the following iterative aggregation operation until the first feature vector of the first parent node of the first layer in the first model tree is obtained:

[0099] Re-encode the elements in the first sequence through a multi-head attention mechanism, and aggregate each re-encoded first sequence with the first parent node of the corresponding first sequence to obtain a first aggregation result;

[0100] Specifically, when the elements in the first sequence are re-encoded, the number of elements and the order of elements before and after encoding remain unchanged. For example, if the sequence before encoding is [A1 A2 A3], the sequence after encoding is [B1 B2 B3], where B1 is the element after A1 is re-encoded, B2 is the element after A2 is re-encoded, and B3 is the element after A3 is re-encoded.

[0101] In the embodiment of the present application, when the elements in the first sequence are re-encoded, an optional implementation is shown in the following formulas 1 and 2:

[0102]

[0103] Among them, W Q is the weight matrix of Q (query matrix), WK is the weight matrix of K (key matrix), W V is the weight matrix of V (value matrix). is the first sequence after recoding, softmax is the normalized exponential function, d k is the dimension of the key matrix.

[0104] In the embodiment of the present application, when obtaining the first aggregation result, an optional implementation method is as follows:

[0105] Perform a linear transformation on the first parent node to obtain a first transformation result; perform a dot product on the first transformation result and each element in the first sequence to obtain a first dot product sequence; the first dot product sequence includes at least one first element; for each first element, convert the first element into a second element by a normalized exponential function; multiply the first element by the second element to obtain a first multiplication result; and use a sequence based on at least one first multiplication result as a first aggregation result.

[0106] Specifically, first change the first parent node Perform a linear transformation and then compare it with the first sequence Each first element in is point-multiplied, and then passes through the softmax function to obtain the weight corresponding to each first element, that is, the second element. The weight is used to aggregate the i-th to j-th first elements into a feature vector, as shown in the following formulas 3 and 4:

[0107]

[0108]

[0109] Where W is the weight matrix for linear transformation of the first parent node, wigth is the weight sequence of the second element, and wigth n is the nth second element in the weight sequence, is the first aggregation result.

[0110] like Figure 4 As shown, it is a schematic diagram of obtaining a first aggregation result provided by an embodiment of the present application. The root node in the DOM tree is replaced by a special symbol, each rectangle numbered 1 is a text feature vector or an image feature vector, and the rectangle numbered 2 is the result of re-encoding the above text feature vector or image feature vector through a multi-head attention mechanism. The matrix numbered 3 is the first aggregation result. The rectangle in each dotted box is the feature vector and re-encoding result corresponding to the element of the same DOM tree layer under the same parent node. Figure 4 The method also includes an aggregator for aggregating each first sequence and its corresponding first parent node.

[0111] Step 4: Concatenate the aggregation result with the first parent node as the first feature vector of the first parent node; and use the first feature vector of the first parent node of the first layer in the first model tree as the first modal feature.

[0112] The first parent node of the first layer is the root node of the DOM tree.

[0113] An optional implementation is shown in the following formula 5:

[0114]

[0115] Among them, concat(*) is a vector concatenation operation; FFN(*) is a fully connected layer, is the first eigenvector of the first parent node.

[0116] like Figure 5 As shown, it is a structural schematic diagram of an aggregator provided in an embodiment of the present application. Figure 5 The rectangle numbered 1 is the first parent node that has not undergone linear transformation, and the rectangle numbered 2 is the first element in the first sequence. Figure 5 The first sequence in has 3 first elements. The rectangle numbered 3 is a linear transformation, the circle is the dot product operation of the first parent node and each first element after the linear transformation, and the rectangle numbered 4 refers to the operation of multiplying the first element and the corresponding second element and then adding them in the above formula 4. The rectangle numbered 5 is the first aggregation result. The rectangle numbered 6 is the fully connected layer, and the rectangle numbered 7 is the first eigenvector of the first parent node.

[0117] Similarly, for the target website, the following operations are performed to obtain the second aggregation result:

[0118] Step 1: Convert the source code of the target website into the second model tree;

[0119] The nodes in the second model tree are text feature vectors extracted from text elements in the source code of the target website, or image feature vectors extracted from image elements in the source code of the target website.

[0120] Step 2: based on the relative position of each node in the second model tree, obtaining at least one second sequence corresponding to the target website;

[0121] The elements contained in each second sequence are text feature vectors or image feature vectors that have the same parent node in the second model tree and are located at the same layer in the second model tree;

[0122] Step 3: For each second sequence, perform the following iterative aggregation operation until the second feature vector of the second parent node of the first layer in the second model tree is obtained:

[0123] Re-encode the elements in the second sequence through a multi-head attention mechanism, and aggregate each re-encoded second sequence with the second parent node of the corresponding second sequence to obtain a second aggregation result;

[0124] Specifically, a linear transformation is performed on the second parent node to obtain a second transformation result; the second transformation result is dot-multiplied with each element in the second sequence to obtain a second dot-product sequence; the second dot-product sequence includes at least one third element; for each third element, the third element is converted into a fourth element by a normalized exponential function; the third element is multiplied by the fourth element to obtain a second multiplication result; and a sequence composed of at least one second multiplication result is used as a second aggregation result.

[0125] Step 4: Concatenate the aggregation result with the second parent node as the second feature vector of the second parent node; and use the second feature vector of the second parent node of the first layer in the second model tree as the second modal feature.

[0126] In the above implementation, the embodiment of the present application converts the website page content into a DOM tree structure, obtains an element sequence using breadth-first traversal, and performs targeted encoding and aggregation on the text and images in the sequence. This method can not only capture the location information of each node of the website, but also retain the relationship between elements in a structured manner, thereby forming a comprehensive feature representation. This comprehensive feature representation realizes multimodal feature fusion with location information, enhances the correlation between features, and enables the detection system to more accurately identify complex phishing attacks.

[0127] S23: Fusing the cross feature, the first modal feature, and the second modal feature to obtain a fused feature, and determining a detection result for the website to be detected based on the fused feature and the discrete feature fusion strategy.

[0128] The detection result is used to indicate whether the website to be detected is a phishing website of the target website.

[0129] The embodiment of the present application designs an MLPD that combines location information and multimodal features. The strategy includes the above-mentioned discrete feature fusion strategy, which integrates the cross-features in S21 into the modal features in S22 to further improve the detection effect.

[0130] In the embodiment of the present application, when determining the detection result, an optional implementation method is as follows:

[0131] Step 1: Determine the cosine similarity between the first modal feature and the second modal feature, and use the cosine similarity as a new feature; input the new feature and the cross feature into the deep neural network model to obtain a fusion feature; concatenate the first modal feature and the second modal feature to obtain a first concatenated feature; perform a linear transformation on the first concatenated feature, and input the first concatenated feature after the linear transformation into a fully connected layer to obtain a third modal feature.

[0132] Specifically, as shown in the following formula 6:

[0133]

[0134] in, is the first modal feature, is the second modal feature, W ’ is the weight matrix, tanh(*) is the activation function, b is the bias vector, FFN(*) is the fully connected layer, It is the third modal feature.

[0135] Step 2: Concatenate the fused feature and the third modal feature to obtain a second concatenated feature; input the second concatenated feature into the fully connected layer, and perform a nonlinear transformation on the output of the fully connected layer to determine the detection result for the website to be detected.

[0136] Specifically, as shown in the following formula 7:

[0137]

[0138] Among them, concat(*) is a vector concatenation operation, FFN(*) is a fully connected layer; sigmoid(*), For the test results.

[0139] In the above implementation, the embodiment of the present application proposes a new discrete feature fusion idea to address the problem that the prior art relies too much on modal features in the feature fusion process. The cross-features of the target website and the website to be detected are integrated into the modal features of the target website and the website to be detected. This fusion not only enriches the encoding of the modal features, but also provides prior information for the modal features, thereby improving the robustness and detection capabilities of the system, especially in the face of complex and variable phishing attacks, it can significantly reduce the false alarm rate.

[0140] It should be noted that, for the MLPD including the above-mentioned MLF-Loc and discrete feature fusion strategy in the embodiment of the present application, training can be performed in the following manner:

[0141] Step 1: Data collection.

[0142] Collect phishing website data from various open source platforms as sample data. The sample data includes the source website and the suspicious website to be detected. Perform data preprocessing operations on the collected sample data in the following ways:

[0143] (1) Delete duplicate data and invalid links; remove noise data such as irrelevant fields and incomplete information;

[0144] (2) Standardize data formats, unify URL formats, and remove illegal characters;

[0145] (3) Integrate the sample data into { <origin i ,sample i ,y i >} format, where origin i For the source website, sample i is the suspicious website to be detected, y i is a label; when y i =1 means sample i is origin i phishing websites, which are used as positive samples of the model; y i =0 means sample i Not origin i The phishing websites are used as negative samples of the model.

[0146] Step 2: Obtain the modal features of the source website and suspicious website through MLF-Loc in MLPD.

[0147] Referring to the above S22, the source website and the suspicious website are input into the multimodal feature fusion model to obtain the first sample modal feature corresponding to the suspicious website and the second sample modal feature corresponding to the source website.

[0148] Step 3: Obtain sample detection results of source websites and suspicious websites through the discrete feature fusion strategy in MLPD.

[0149] Refer to S21 and S23 above to obtain the sample test results

[0150] Step 4: Update the parameters of MLPD through the cross entropy loss function.

[0151] As shown in the following formula 8:

[0152]

[0153] After calculating the above function values, the MLPD parameters can be trained through stochastic gradient descent.

[0154] like Figure 6As shown, it is a schematic diagram of an MLPD training provided in an embodiment of the present application. The target website and the website to be detected are respectively input into MLF-Loc to obtain Figure 6 The rectangle numbered 1 corresponds to the first modal feature, and the rectangle numbered 2 corresponds to the second modal feature. Determine the cosine similarity between the first modal feature and the second modal feature, and use the cosine similarity as a new feature, that is, Figure 6 The rectangle numbered 7 in . The rectangle numbered 3 refers to the title similarity, the rectangle numbered 4 refers to the icon similarity, the rectangle numbered 5 refers to the website certificate similarity, and the rectangle numbered 6 refers to the edit distance. After normalizing the cross feature and the new feature data, they are input into the deep neural network model (DNN) to obtain the fusion feature, that is, Figure 6 The rectangle numbered 8 in the figure; the first modal feature and the second modal feature are concatenated to obtain the first concatenated feature; after linear transformation of the first concatenated feature, the first concatenated feature after linear transformation is input into the fully connected layer (i.e. Figure 6 The third modal feature is obtained by Figure 6 The rectangle numbered 10 in the figure; concatenate the fusion feature and the third modal feature to obtain the second concatenated feature; input the second concatenated feature into the fully connected layer (i.e. Figure 6 The output of the fully connected layer is transformed nonlinearly to determine the detection result for the website to be detected.

[0155] like Figure 7 As shown, it is a flow chart of another phishing website detection method provided in an embodiment of the present application. Specifically, it includes S71 to S78:

[0156] S71: Collect sample data and perform data preprocessing;

[0157] S72: Obtaining the positional relationship of each element in the suspicious website and the source website in the website based on the DOM tree;

[0158] S73: extracting basic information of the suspicious website and the source website respectively, and generating cross features;

[0159] S74: Multimodal feature fusion based on location information;

[0160] S75: performing discrete feature fusion;

[0161] S76: Input the URL of the target website and extract the domain name of the URL;

[0162] S77: Generate the website to be detected using the domain name fuzzy algorithm;

[0163] S78: Use MLPD to detect whether the website to be detected is a phishing website of the target website.

[0164] In practice, the phishing website detection method in the embodiment of the present application can be implemented and used in multiple scenarios. Specific embodiments are listed as follows:

[0165] Example 1: Security enhancement for online financial services

[0166] In online financial services, phishing website detection technology is a key tool to protect customer accounts and funds. With the increasing prevalence of cyber attacks, hackers often obtain customers' login information, transaction passwords and other sensitive data through disguised web pages. These phishing websites are often very similar to regular financial websites and are difficult for ordinary customers to identify, causing customers to disclose personal information without knowing it. By applying phishing website detection technology, financial institutions can quickly identify and block these phishing websites, thereby effectively preventing customers from entering by mistake. In order to further enhance customers' security awareness, financial service platforms can also provide customers with security warnings and education on a regular basis to help customers understand common phishing methods and improve their ability to identify phishing websites.

[0167] Implementation step 1: Collect phishing website data from various open source platforms, pre-process the data, and organize it into MLF-Loc sample data.

[0168] Implementation step 2: Construct the MLPD in the embodiment of this application and use sample data to train the MLPD. The MLPD includes MLF-Loc with location information and discrete feature fusion strategy;

[0169] Implementation step three: The target enterprise provides its own legitimate website, which is called the target website; provides the domain name of the target website URL, uses the domain name fuzzy algorithm to generate the suspicious URL of the website to be detected, crawls the suspicious URL, and the resulting website is called the website to be detected; inputs the target website and the website to be detected into MLPD to detect whether the suspicious URL is a phishing website of the target website;

[0170] Implementation step four: Report and block identified phishing websites to prevent customers from being deceived.

[0171] Example 2: Enterprise network protection

[0172] In their daily work, corporate employees often come into contact with various online resources, which makes them easily enter malicious phishing websites, thereby leaking sensitive information or being infected with malware. By deploying phishing website detection technology, enterprises can monitor and filter access requests in real time, automatically identify suspicious websites and block access. In addition, enterprises can also conduct feature analysis based on the collected phishing website samples, conduct regular training for corporate employees, increase employees' vigilance against phishing websites, and further enhance the network security protection layer. Through this comprehensive security strategy, enterprises can not only protect their own data and assets, but also maintain their reputation and avoid economic losses caused by cyber attacks.

[0173] The implementation steps 1 and 2 are similar to those in the above-mentioned embodiment 1 and will not be described in detail here.

[0174] Implementation step three: Use the Internet Content Provider (ICP) database to obtain the domain names of enterprises across the country, build URLs based on the domain names and perform crawling. The crawled websites are enterprise websites, called target websites. For a certain enterprise website, use the domain name fuzzy algorithm to generate suspicious URLs and perform crawling. The crawled websites are called websites to be detected. Input the target website and the website to be detected into MLPD to detect whether the suspicious URL is a phishing website of the target website.

[0175] Implementation step four: Report the identified phishing websites and add them to the blacklist to prevent internal employees from accessing them.

[0176] In the above implementation, the embodiments of the present application take into account that current phishing website detection methods are mostly focused on a single feature set and lack effective integration of multimodal data, resulting in insufficient recognition accuracy in complex attack scenarios. In addition, related technologies often ignore the impact of location information on the relationship between features, which in turn affects the detection effect. The embodiments of the present application propose a phishing website detection method with location information, which records node location information by using the DOM tree of the website, and performs targeted encoding and aggregation of the node sequence to form a more comprehensive feature representation. This method enhances the correlation between features during the feature fusion process, so that the detection model can more effectively capture and identify the complex features of phishing websites, significantly improving the detection accuracy.

[0177] In addition, the detection of phishing websites by related technologies mostly relies on a single modal feature, fails to effectively utilize the cross-features of the target website and the website to be detected, and lacks analysis of comprehensive information other than the website text, which limits the performance and adaptability of the detection model. In order to overcome this shortcoming, the embodiment of the present application proposes a new discrete feature fusion idea, which aims to integrate the cross-features into the modal features of the target website and the website to be detected. By integrating these cross-features with the modal features, prior information is provided, thereby further improving the detection effect on the basis of feature fusion.

[0178] Based on the same inventive concept, the present application also provides a phishing website detection device, such as Figure 8 As shown, the phishing website detection device 8000 includes:

[0179] The generating unit 8001 is used to generate a cross feature based on the basic information of the website to be detected and the target basic information of the target website; the basic information to be detected is used to characterize the authenticity and security of the website to be detected; the target basic information is used to characterize the authenticity and security of the target website;

[0180] The acquisition unit 8002 is used to input the source code of the website to be detected and the source code of the target website into the multimodal feature fusion model respectively, and obtain the first modal feature of the website to be detected and the second modal feature of the target website; the first modal feature includes the position relationship and function of each element in the website to be detected in the website to be detected; the second modal feature includes the position relationship and function of each element in the target website in the target website;

[0181] The detection unit 8003 is used to fuse the cross features, the first modal features and the second modal features to obtain the fused features, and determine the detection results for the website to be detected based on the fused features and the discrete feature fusion strategy; the detection results are used to characterize whether the website to be detected is a phishing website of the target website.

[0182] In some embodiments, the acquisition unit 8002 is specifically used to:

[0183] Converting the source code of the website to be detected into a first model tree; the nodes in the first model tree are text elements in the source code of the website to be detected, or image elements in the source code of the website to be detected;

[0184] Based on the relative position of each node in the first model tree, at least one first sequence corresponding to the website to be detected is obtained; the elements contained in each first sequence are text feature vectors corresponding to text elements that have the same parent node in the first model tree and are located in the same layer in the first model tree, or image feature vectors corresponding to image elements;

[0185] For each first sequence, perform the following iterative aggregation operation until the first feature vector of the first parent node of the first layer in the first model tree is obtained:

[0186] Re-encode the elements in the first sequence through a multi-head attention mechanism, and aggregate each re-encoded first sequence with the first parent node of the corresponding first sequence to obtain a first aggregation result;

[0187] Concatenate the aggregation result with the first parent node as the first feature vector of the first parent node;

[0188] The first feature vector of the first parent node of the first layer in the first model tree is used as the first modal feature.

[0189] In some embodiments, the acquisition unit 8002 is specifically used to:

[0190] Perform a linear transformation on the first parent node to obtain a first transformation result;

[0191] Perform a dot product of the first transformation result and each element in the first sequence to obtain a first dot product sequence; the first dot product sequence includes at least one first element;

[0192] For each first element, convert the first element into a second element by using a normalized exponential function; multiply the first element by the second element to obtain a first multiplication result;

[0193] A sequence formed based on at least one first multiplication result is used as a first aggregation result.

[0194] In some embodiments, the acquisition unit 8002 is specifically used to:

[0195] Converting the source code of the target website into a second model tree; the nodes in the second model tree are text elements in the source code of the target website, or image elements in the source code of the target website;

[0196] Based on the relative position of each node in the second model tree, at least one second sequence corresponding to the target website is obtained; the elements contained in each second sequence are text feature vectors corresponding to text elements that have the same parent node in the second model tree and are located in the same layer in the second model tree, or image feature vectors corresponding to image elements;

[0197] For each second sequence, perform the following iterative aggregation operation until the second feature vector of the second parent node of the first layer in the second model tree is obtained:

[0198] Re-encode the elements in the second sequence through a multi-head attention mechanism, and aggregate each re-encoded second sequence with the second parent node of the corresponding second sequence to obtain a second aggregation result;

[0199] The aggregation result is concatenated with the second parent node to serve as the second feature vector of the second parent node;

[0200] The second feature vector of the second parent node of the first layer in the second model tree is used as the second modal feature.

[0201] In some embodiments, the acquisition unit 8002 is specifically used to:

[0202] Perform a linear transformation on the second parent node to obtain a second transformation result;

[0203] Perform a dot product of the second transformation result and each element in the second sequence to obtain a second dot product sequence; the second dot product sequence includes at least one third element;

[0204] For each third element, convert the third element into a fourth element by using a normalized exponential function; multiply the third element by the fourth element to obtain a second multiplication result;

[0205] A sequence formed based on the at least one second multiplication result is used as a second aggregation result.

[0206] In some embodiments, the generating unit 8001 is specifically used for:

[0207] If the basic information to be detected includes the title to be detected, and the target basic information includes the target title, then the cross feature includes the title similarity between the title to be detected and the target title;

[0208] If the basic information to be detected includes an icon to be detected, and the target basic information includes a target icon, then the cross feature includes an icon similarity between the icon to be detected and the target icon;

[0209] If the basic information to be detected includes the certificate of the website to be detected, and the target basic information includes the certificate of the target website, then the cross-feature includes the website certificate similarity between the certificate of the website to be detected and the certificate of the target website;

[0210] If the basic information to be detected includes the domain name of the website to be detected, and the target basic information includes the domain name of the target website, then the cross feature includes the edit distance between the domain name of the website to be detected and the domain name of the target website.

[0211] In some embodiments, the detection unit 8003 is specifically used to:

[0212] Determine the cosine similarity between the first modal feature and the second modal feature, and use the cosine similarity as a new feature;

[0213] Input new features and cross features into the deep neural network model to obtain fused features;

[0214] Based on the fusion feature, the first modality feature, the second modality feature and the discrete feature fusion strategy, the detection result for the website to be detected is determined, including:

[0215] Concatenate the first modal feature and the second modal feature to obtain a first concatenated feature;

[0216] After performing a linear transformation on the first concatenated feature, the linearly transformed first concatenated feature is input into a fully connected layer to obtain a third modal feature;

[0217] The fusion feature and the third modality feature are concatenated to obtain a second concatenated feature;

[0218] The second concatenated feature is input into the fully connected layer, and the output of the fully connected layer is nonlinearly transformed to determine the detection result for the website to be detected.

[0219] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present application. In one embodiment, the electronic device may be Figure 1 The terminal device 110 shown in FIG. 1 is a terminal device 110 shown in FIG. 1 . In this embodiment, the structure of the electronic device can be as follows: Fig. 9 As shown, it includes a memory 901 , a communication module 903 and one or more processors 902 .

[0220] The memory 901 is used to store computer programs executed by the processor 902. The memory 901 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system and programs required for running the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.

[0221] The memory 901 may be a volatile memory, such as a random-access memory (RAM); the memory 901 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); or the memory 901 may be any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 901 may be a combination of the above memories.

[0222] The processor 902 may include one or more central processing units (CPU) or a digital processing unit, etc. The processor 902 is used to implement the above-mentioned phishing website detection method when calling the computer program stored in the memory 901 .

[0223] The communication module 903 is used to communicate with terminal devices and other servers.

[0224] The specific connection medium between the memory 901, the communication module 903 and the processor 902 is not limited in the embodiment of the present application. Fig. 9 In the embodiment, the memory 901 and the processor 902 are connected via a bus 904. The bus 904 is connected to the processor 902 via a bus 904. Fig. 9 The connections between the other components are only for illustration and are not intended to be limiting. The bus 904 can be divided into an address bus, a data bus, a control bus, etc. For ease of description, Fig. 9 The diagram shows that only one thick line is used, but this does not mean that there is only one bus or only one type of bus.

[0225] The memory 901 stores a computer storage medium, and the computer storage medium stores computer executable instructions, and the computer executable instructions are used to implement the phishing website detection method of the embodiment of the present application. The processor 902 is used to execute the above-mentioned phishing website detection method. Based on the same inventive concept, the embodiment of the present application provides a computer-readable storage medium, and the computer program product includes: computer program code, when the computer program code runs on the computer, the computer executes any of the phishing website detection methods discussed above. Since the principle of solving the problem by the above-mentioned computer-readable storage medium is similar to that of the phishing website detection method, the implementation of the above-mentioned computer-readable storage medium can refer to the implementation of the method, and the repeated parts will not be repeated.

[0226] Refer to the following Fig.10 hereinafter, a computing device 1000 according to this embodiment of the present application is described. Fig.10 The computing device 1000 is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0227] like Fig.10 The computing device 1000 is in the form of a general computing device. The components of the computing device 1000 may include but are not limited to: at least one processing unit 1001, at least one storage unit 1002, and a bus 1003 connecting different system components (including the storage unit 1002 and the processing unit 1001).

[0228] Bus 1003 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a processor, or a local bus using any of a variety of bus architectures.

[0229] The storage unit 1002 may include a readable medium in the form of a volatile memory, such as a random access memory (RAM) 821 and / or a cache memory 1022 , and may further include a read-only memory (ROM) 1023 .

[0230] The storage unit 1002 may also include a program / utility 1025 having a set (at least one) of program modules 1024, such program modules 1024 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0231] The computing device 1000 may also communicate with one or more external devices 1004 (e.g., keyboards, pointing devices, etc.), one or more devices that enable a user to interact with the computing device 1000, and / or any device that enables the computing device 1000 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 1005. Furthermore, the computing device 1000 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 1006. Fig.10 As shown, the network adapter 1006 communicates with other modules for the computing device 1000 via the bus 1003. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the computing device 1000, including but not limited to: microcode, device drivers, redundant processors, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0232] The embodiment of the present application also provides a computer program product, and the method in the present application can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instruction is loaded and executed on a computer, the process or function described in the present application is executed in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, a core network device, an OAM or other programmable device.

[0233] A computer-readable storage medium can be implemented as a computer program product, that is, an embodiment of the present application also provides a computer-readable storage medium, which includes a computer program, and when the computer program is executed by a processor, it implements any one of the above-mentioned phishing website detection methods.

[0234] The computer program or instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium, for example, the computer program or instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, or a magnetic tape; it may also be an optical medium, such as a digital video disk; it may also be a semiconductor medium, such as a solid-state drive. The computer-readable storage medium may be a volatile or non-volatile storage medium, or may include both volatile and non-volatile types of storage media.

[0235] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0236] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0237] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0238] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0239] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for detecting phishing websites, characterized in that: include: Generate cross-features based on the basic information of the website to be detected and the target basic information of the target website; The basic information to be detected is used to characterize the authenticity and security of the website to be detected; the target basic information is used to characterize the authenticity and security of the target website; Inputting the source code of the website to be detected and the source code of the target website into the multimodal feature fusion model respectively, obtaining the first modal feature of the website to be detected and the second modal feature of the target website; the first modal feature includes the positional relationship and function of each element in the website to be detected in the website to be detected; the second modal feature includes the positional relationship and function of each element in the target website in the target website; The cross-features, the first modal features and the second modal features are integrated to obtain fused features, and based on the fused features and the discrete feature fusion strategy, a detection result for the website to be detected is determined; the detection result is used to characterize whether the website to be detected is a phishing website of the target website.

2. The method according to claim 1, characterized in that The step of inputting the source code of the website to be detected into the multimodal feature fusion model to obtain the first modal feature corresponding to the website to be detected includes: Converting the source code of the website to be detected into a first model tree; the nodes in the first model tree are text elements in the source code of the website to be detected, or image elements in the source code of the website to be detected; Based on the relative position of each node in the first model tree, at least one first sequence corresponding to the website to be detected is obtained; the elements contained in each first sequence are text feature vectors corresponding to text elements or image feature vectors corresponding to image elements that have the same parent node in the first model tree and are located in the same layer in the first model tree; For each of the first sequences, the following iterative aggregation operation is performed until the first feature vector of the first parent node of the first layer in the first model tree is obtained: Re-encode the elements in the first sequence by using a multi-head attention mechanism, and aggregate each of the re-encoded first sequences with a first parent node corresponding to the first sequence to obtain a first aggregation result; Concatenate the aggregation result with the first parent node to obtain a first feature vector of the first parent node; The first feature vector of the first parent node of the first layer in the first model tree is used as the first modal feature.

3. The method according to claim 2, characterized in that The step of aggregating each of the re-encoded first sequences with a first parent node corresponding to the first sequence to obtain a first aggregation result includes: Performing a linear transformation on the first parent node to obtain a first transformation result; Perform a dot product of the first transformation result and each element in the first sequence to obtain a first dot product sequence; the first dot product sequence includes at least one first element; For each of the first elements, convert the first element into a second element by using a normalized exponential function; multiply the first element by the second element to obtain a first multiplication result; A sequence formed based on at least one of the first multiplication results is used as the first aggregation result.

4. The method according to claim 1, characterized in that The step of inputting the source code of the target website into the multimodal feature fusion model to obtain the second modal feature corresponding to the target website includes: Converting the source code of the target website into a second model tree; the nodes in the second model tree are text elements in the source code of the target website, or image elements in the source code of the target website; Based on the relative position of each node in the second model tree, obtaining at least one second sequence corresponding to the target website; the elements contained in each second sequence are text feature vectors corresponding to text elements or image feature vectors corresponding to image elements that have the same parent node in the second model tree and are located in the same layer in the second model tree; For each of the second sequences, the following iterative aggregation operation is performed until the second feature vector of the second parent node of the first layer in the second model tree is obtained: re-encode the elements in the second sequence by using a multi-head attention mechanism, and aggregate each re-encoded second sequence with a second parent node corresponding to the second sequence to obtain a second aggregation result; Concatenate the aggregation result with the second parent node to obtain a second feature vector of the second parent node; The second feature vector of the second parent node of the first layer in the second model tree is used as the second modal feature.

5. The method according to claim 4, characterized in that The step of aggregating each re-encoded second sequence with a second parent node corresponding to the second sequence to obtain a second aggregation result includes: Performing a linear transformation on the second parent node to obtain a second transformation result; Perform a dot product of the second transformation result and each element in the second sequence to obtain a second dot product sequence; the second dot product sequence includes at least one third element; For each of the third elements, convert the third element into a fourth element by using a normalized exponential function; multiply the third element by the fourth element to obtain a second multiplication result; A sequence formed based on at least one of the second multiplication results is used as the second aggregation result.

6. The method according to claim 1, characterized in that: If the basic information to be detected includes a title to be detected, and the target basic information includes a target title, then the cross feature includes the title similarity between the title to be detected and the target title; If the basic information to be detected includes an icon to be detected, and the target basic information includes a target icon, then the intersection feature includes an icon similarity between the icon to be detected and the target icon; If the basic information to be detected includes a website certificate to be detected, and the target basic information includes a target website certificate, then the cross-feature includes the website certificate similarity between the website certificate to be detected and the target website certificate; If the basic information to be detected includes the domain name of the website to be detected, and the target basic information includes the domain name of the target website, then the cross feature includes the edit distance between the domain name of the website to be detected and the domain name of the target website.

7. The method according to any one of claims 1 to 6, characterized in that: The fusing the cross feature, the first modality feature and the second modality feature comprises: Determine the cosine similarity between the first modal feature and the second modal feature, and use the cosine similarity as a new feature; Inputting the new features and the cross features into a deep neural network model to obtain fused features; The determining of the detection result for the website to be detected based on the fusion feature, the first modal feature, the second modal feature and the discrete feature fusion strategy includes: Splicing the first modal feature and the second modal feature to obtain a first splicing feature; After performing a linear transformation on the first concatenated features, the linearly transformed first concatenated features are input into a fully connected layer to obtain a third modal feature; Splicing the fusion feature and the third modality feature to obtain a second splicing feature; The second concatenated feature is input into the fully connected layer, and the output of the fully connected layer is nonlinearly transformed to determine a detection result for the website to be detected.

8. A phishing website detection device, characterized in that: include: A generating unit, used for generating a cross feature based on the basic information of the website to be detected and the target basic information of the target website; The basic information to be detected is used to characterize the authenticity and security of the website to be detected; the target basic information is used to characterize the authenticity and security of the target website; An acquisition unit is used to input the source code of the website to be detected and the source code of the target website into a multimodal feature fusion model respectively, and acquire a first modal feature of the website to be detected and a second modal feature of the target website; the first modal feature includes the positional relationship and function of each element in the website to be detected in the website to be detected; the second modal feature includes the positional relationship and function of each element in the target website in the target website; A detection unit is used to fuse the cross-features, the first modal features and the second modal features to obtain fused features, and determine the detection results for the website to be detected based on the fused features and the discrete feature fusion strategy; the detection results are used to characterize whether the website to be detected is a phishing website of the target website.

9. An electronic device, characterized in that: include: A memory for storing program instructions; A processor is used to call the program instructions stored in the memory, and execute the steps included in the method according to any one of claims 1 to 7 according to the obtained program instructions.

10. A computer-readable storage medium, wherein a computer program is stored in the computer-readable storage medium, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.

11. A computer program product, characterized in that The method comprises a computer program stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Counterfeit website identification method and device, equipment and storage medium

    CN122437735A

  • Method, device and storage medium for identifying a counterfeit website

    CN122437735B