Method, device and storage medium for identifying a counterfeit website
Patent Information
- Application Number
- CN202610911834.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2046-06-24
AI Technical Summary
[0005]本发明的主要目的在于提供了一种仿冒网站识别方法、装置、设备和存储介质,旨在解决现有技术中,仿冒网站的识别方法缺乏对相似页面的准确判定能力,且在大规模检测场景下效率较低,存在识别准确性低和识别效率低的技术问题
本申请提供了一种仿冒网站识别方法、装置、设备和存储介质,上述方法包括:获取多个第一候选网站;基于预设的基准网站数字指纹,对多个第一候选网站进行相似性筛选,得到多个第二候选网站以及每个第二候选网站对应的相似性分值;基于预设的基准网站数字指纹,对每个第二候选网站进行安全性认证,得到每个第二候选网站对应的域名异常分值、证书异常分值和模型置信度;基于每个第二候选网站对应的相似性分值、域名异常分值、证书异常分值和模型置信度,确定每个第二候选网站对应的风险分值;基于每个第二候选网站对应的风险分值,确定多个第二候选网站中的仿冒网站。本申请中,基于预设的基准网站数字指纹,对候选网站进行相似性筛选以及安全性认证,确定候选网站对应的风险分值,进而基于候选网站对应的风险分值,确定仿冒网站。以此通过多重风险识别的方式,从海量网站中精准识别到仿冒网站,提高了对于仿冒网站的识别准确性以及识别效率。
Smart Images

Figure CN122437735B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data technology, and in particular to a method, apparatus, device, and storage medium for identifying counterfeit websites. Background Technology
[0002] With the development of the internet, a large number of counterfeit websites have emerged. These counterfeit websites steal users' private information and harm users' online assets, greatly affecting network security.
[0003] Currently, technologies for identifying counterfeit websites fall into the following categories. The first category is rule-based detection technology based on blacklists, domain rules, and registration information. This technology quickly identifies suspicious websites by comparing domain character similarity, brand word spelling, WHOIS / registration information, DNS resolution relationships, IP address attribution, and certificate issuing authority. The second category is webpage identification technology based on single-modal similarity; for example, comparing image similarity between webpage screenshots. The third category is phishing detection technology based on dynamic behavior analysis; this technology automatically accesses webpages, records behavioral information during the access process, and identifies potential risks.
[0004] However, the aforementioned methods for identifying counterfeit websites lack the ability to accurately determine similar pages and are inefficient in large-scale detection scenarios, exhibiting technical problems of low accuracy and low efficiency. Summary of the Invention
[0005] The main objective of this invention is to provide a method, apparatus, device, and storage medium for identifying counterfeit websites, aiming to solve the technical problems of low accuracy and low efficiency in existing methods for identifying counterfeit websites, which lack the ability to accurately determine similar pages and are inefficient in large-scale detection scenarios.
[0006] To achieve the above objectives, the present invention provides a method for identifying counterfeit websites, the method comprising: Multiple first candidate websites are obtained; the multiple first candidate websites are determined based on a pre-filtering process of preset websites. Based on a preset benchmark website digital fingerprint, the multiple first candidate websites are screened for similarity to obtain multiple second candidate websites and a similarity score for each second candidate website; the benchmark website digital fingerprint is determined based on the webpage features of the benchmark website, and the similarity score is used to characterize the similarity between the corresponding second candidate website and the benchmark website; the multiple second candidate websites are some of the multiple first candidate websites. Based on a preset benchmark website digital fingerprint, security authentication is performed on each second candidate website to obtain the domain anomaly score, certificate anomaly score, and model confidence score for each second candidate website. The domain anomaly score is used to characterize the domain authentication status of the corresponding second candidate website, the certificate anomaly score is used to characterize the certificate authentication status of the corresponding second candidate website, and the model confidence score is obtained by reasoning about the second candidate website based on a preset multimodal model. Based on the similarity score, domain name anomaly score, certificate anomaly score, and model confidence score corresponding to each second candidate website, the risk score corresponding to each second candidate website is determined. Based on the risk score corresponding to each of the second candidate websites, counterfeit websites are identified among the plurality of second candidate websites.
[0007] Optionally, the digital fingerprint of the benchmark website includes fingerprint clusters that correspond one-to-one with each webpage, and each fingerprint cluster includes a visual feature vector, a structural feature vector, and a semantic feature vector; The method, based on a preset benchmark website digital fingerprint, performs similarity screening on the multiple first candidate websites to obtain multiple second candidate websites and a similarity score for each second candidate website, including: Calculate the visual similarity score between the visual feature vector corresponding to each first candidate website and the visual feature vector corresponding to each fingerprint cluster; Calculate the structural similarity score between the structural feature vector corresponding to each first candidate website and the structural feature vector corresponding to each fingerprint cluster; Calculate the semantic similarity score between the semantic feature vector corresponding to each first candidate website and the semantic feature vector corresponding to each fingerprint cluster; The visual similarity score, the structural similarity score, and the semantic similarity score are weighted and summed to obtain the similarity score corresponding to each first candidate website; Based on the similarity score corresponding to each first candidate website, multiple second candidate websites are determined.
[0008] Optionally, determining multiple second candidate websites based on the similarity score corresponding to each first candidate website includes: The first candidate website with a similarity score greater than the first preset threshold is determined as the second candidate website; The visual feature vector, semantic feature vector, structural feature vector, and semantic summary of the first candidate website whose similarity score is less than or equal to the first preset threshold and greater than the second preset threshold are input into the preset multimodal model to obtain the risk identification result. The first candidate website, whose risk identification results indicate that there is a risk, is designated as the second candidate website.
[0009] Optionally, the benchmark website digital fingerprint includes a set of legitimate domain names and legitimate authentication information; The security authentication of each second candidate website is performed based on a preset benchmark website digital fingerprint, resulting in a domain anomaly score, certificate anomaly score, and model confidence level for each second candidate website, including: Based on the set of legitimate domain names, domain name authentication is performed on each of the second candidate websites to obtain the domain name anomaly score corresponding to each of the second candidate websites; Based on the legitimate authentication information, certificate authentication is performed on each of the second candidate websites to obtain the certificate anomaly score corresponding to each second candidate website; The visual feature vector, structural feature vector, semantic feature vector, domain name set, authentication information, and digital fingerprint of the benchmark website corresponding to each second candidate website are input into a preset multimodal model. Consistency inference is performed on each second candidate website through the multimodal model to obtain the model confidence level corresponding to each second candidate website.
[0010] Optionally, determining the risk score for each second candidate website based on its similarity score, domain anomaly score, certificate anomaly score, and model confidence score includes: The risk score for each second candidate website is obtained by weighted summing of the similarity score, domain name anomaly score, certificate anomaly score, and model confidence score.
[0011] Optionally, determining the counterfeit website among the plurality of second candidate websites based on the risk score corresponding to each second candidate website includes: Second candidate websites with risk scores greater than the third preset threshold are identified as high-risk counterfeit websites; The second candidate website with a risk score less than or equal to the third preset threshold and greater than or equal to the fourth preset threshold is identified as a medium-risk counterfeit website. The second candidate website with a risk score lower than the fourth preset threshold is identified as a non-counterfeit website.
[0012] Optionally, the method further includes: Extract the visual features of a preset benchmark website to obtain the visual feature vector corresponding to the benchmark website; Extract the structural features of a preset benchmark website to obtain the structural feature vector corresponding to the benchmark website; Extract the semantic features of a preset benchmark website to obtain the semantic feature vector corresponding to the benchmark website; Based on the visual feature vector, structural feature vector, semantic feature vector, set of legal domain names, legal authentication information, semantic summary and business semantic tags corresponding to the benchmark website, a digital fingerprint of the benchmark website is generated.
[0013] Furthermore, to achieve the above objectives, the present invention also provides a device for identifying counterfeit websites, the device comprising: The acquisition module is used to acquire multiple first candidate websites; the multiple first candidate websites are determined based on pre-filtering of preset websites; The filtering module is used to perform similarity filtering on the plurality of first candidate websites based on a preset benchmark website digital fingerprint, to obtain a plurality of second candidate websites and a similarity score corresponding to each second candidate website; the benchmark website digital fingerprint is determined based on the webpage features of the benchmark website, and the similarity score is used to characterize the similarity between the corresponding second candidate website and the benchmark website, and the plurality of second candidate websites are a portion of the plurality of first candidate websites; The authentication module is used to perform security authentication on each of the second candidate websites based on a preset benchmark website digital fingerprint, and to obtain the domain anomaly score, certificate anomaly score and model confidence score for each second candidate website. The domain anomaly score is used to characterize the domain authentication status of the corresponding second candidate website, the certificate anomaly score is used to characterize the certificate authentication status of the corresponding second candidate website, and the model confidence score is obtained by reasoning about the second candidate website based on a preset multimodal model. The first determining module is used to determine the risk score corresponding to each second candidate website based on the similarity score, domain name anomaly score, certificate anomaly score, and model confidence score corresponding to each second candidate website; The second determining module is used to determine the counterfeit website among the plurality of second candidate websites based on the risk score corresponding to each second candidate website.
[0014] To address the aforementioned technical problems, this application also provides a computer device that employs the following technical solution: The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the phishing website identification method proposed in any of the embodiments of this application.
[0015] To address the aforementioned technical problems, this application also provides a computer-readable storage medium, employing the technical solution described below: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the phishing website identification method proposed in any of the embodiments of this application.
[0016] Compared with the prior art, the embodiments of this application have the following main advantages: This application provides a method, apparatus, device, and storage medium for identifying counterfeit websites. The method includes: acquiring multiple first candidate websites; performing similarity screening on the multiple first candidate websites based on a preset benchmark website digital fingerprint to obtain multiple second candidate websites and a similarity score for each second candidate website; performing security authentication on each second candidate website based on the preset benchmark website digital fingerprint to obtain a domain name anomaly score, certificate anomaly score, and model confidence score for each second candidate website; determining a risk score for each second candidate website based on the similarity score, domain name anomaly score, certificate anomaly score, and model confidence score; and identifying counterfeit websites among the multiple second candidate websites based on the risk scores for each second candidate website. In this application, candidate websites are screened for similarity and authenticated for security based on a preset benchmark website digital fingerprint to determine the risk scores for the candidate websites, and then counterfeit websites are identified based on the risk scores for the candidate websites. This multi-risk identification method accurately identifies counterfeit websites from a massive number of websites, improving the accuracy and efficiency of counterfeit website identification. Attached Figure Description
[0017] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is an exemplary system architecture diagram of which the method for identifying counterfeit websites can be applied in this application; Figure 2 This is a flowchart illustrating the method for identifying counterfeit websites provided in this application embodiment; Figure 3 This is a schematic diagram of the similarity screening process provided in the embodiments of this application; Figure 4 This is a schematic diagram of the security authentication process provided in the embodiments of this application; Figure 5 This is a schematic diagram of the process for constructing a baseline website digital fingerprint provided in an embodiment of this application; Figure 6 This is an application flowchart for identifying counterfeit websites provided in the embodiments of this application; Figure 7 This is a schematic diagram of the structure of the counterfeit website identification device provided in the embodiments of this application; Figure 8This is a basic structural block diagram of a computer device provided in the embodiments of this application. Detailed Implementation
[0019] The method for identifying counterfeit websites provided in this invention is applied to a device for identifying counterfeit websites. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit this application. The terms "comprising" and "having," and any variations thereof, in the specification, claims, and accompanying drawings of this application are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or accompanying drawings of this application are used to distinguish different objects, not to describe a particular order.
[0020] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0021] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0022] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.
[0023] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social online platform software.
[0024] Terminal devices 101, 102, and 103 can be various electronic devices with displays and support web browsing, including but not limited to smartphones, tablets, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 players (Moving Picture Experts Group Audio Layer IV), laptops, and desktop computers, etc.
[0025] Server 105 can be a server that provides various services, such as a backend server that supports the pages displayed on terminal devices 101, 102, and 103.
[0026] It should be noted that the phishing website identification method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the phishing website identification device is generally set in the server / terminal device.
[0027] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0028] Please refer to Figure 2 The flowchart illustrates an embodiment of the phishing website identification method according to this application. The phishing website identification method provided in this application includes the following steps: S201, obtained multiple first-choice websites.
[0029] It should be understood that the aforementioned multiple first candidate websites were determined based on pre-filtering of preset websites.
[0030] Optionally, pre-filtering can be performed on websites that users have visited in the past. Specifically, this involves checking whether the domain names of historical websites contain similar brand words, character substitutions, homonyms, variations of pinyin or abbreviations; domain hierarchy depth; main domain registration age; IP address reputation; and whether there are abnormal redirect parameters, thereby filtering a massive number of historical websites to obtain the first candidate websites.
[0031] S202, based on the preset benchmark website digital fingerprint, perform similarity screening on the multiple first candidate websites to obtain multiple second candidate websites and the similarity score corresponding to each second candidate website.
[0032] It should be understood that the aforementioned digital fingerprint of the benchmark website is determined based on the webpage features of the benchmark website, and the aforementioned similarity score is used to characterize the similarity between the corresponding second candidate website and the benchmark website. Among them, the multiple second candidate websites are some of the multiple first candidate websites.
[0033] In this step, a digital fingerprint of a benchmark website is pre-constructed. This benchmark website can be understood as an official, legitimate website. Specific implementation methods for constructing the benchmark website's digital fingerprint can be found in subsequent embodiments. Based on the benchmark website's digital fingerprint, multiple first candidate websites are screened for similarity, thereby obtaining second candidate websites with high similarity to the benchmark website and a similarity score for each second candidate website.
[0034] S203, based on the preset benchmark website digital fingerprint, perform security authentication on each of the second candidate websites to obtain the domain name anomaly score, certificate anomaly score and model confidence level corresponding to each second candidate website.
[0035] It should be noted that the domain name anomaly score is used to characterize the domain name authentication status of the corresponding second candidate website, the certificate anomaly score is used to characterize the certificate authentication status of the corresponding second candidate website, and the model confidence score is obtained by reasoning about the second candidate website based on a preset multimodal model.
[0036] In this step, the digital fingerprint of the benchmark website is used to perform security authentication on each second candidate website, obtaining the domain anomaly score, certificate anomaly score, and model confidence score for each second candidate website. For specific implementation details, please refer to subsequent embodiments.
[0037] S204. Based on the similarity score, domain name anomaly score, certificate anomaly score, and model confidence score corresponding to each second candidate website, determine the risk score corresponding to each second candidate website.
[0038] In this step, after obtaining the similarity score, domain anomaly score, certificate anomaly score, and model confidence score for each second candidate website, the risk score for each second candidate website is determined based on these scores.
[0039] S205, based on the risk score corresponding to each second candidate website, determine the counterfeit website among the plurality of second candidate websites.
[0040] In this step, after determining the risk score for each second candidate website, counterfeit websites among the multiple second candidate websites are identified based on the risk score. For specific implementation details, please refer to subsequent embodiments.
[0041] In this embodiment, candidate websites are screened for similarity and undergo security authentication based on a preset benchmark website digital fingerprint to determine their corresponding risk scores. Then, based on these risk scores, counterfeit websites are identified. This multi-risk identification method accurately identifies counterfeit websites from a massive database, improving both the accuracy and efficiency of counterfeit website identification.
[0042] Optionally, the digital fingerprint of the benchmark website includes fingerprint clusters that correspond one-to-one with each webpage, and each fingerprint cluster includes a visual feature vector, a structural feature vector, and a semantic feature vector; The method, based on a preset benchmark website digital fingerprint, performs similarity screening on the multiple first candidate websites to obtain multiple second candidate websites and a similarity score for each second candidate website, including: Calculate the visual similarity score between the visual feature vector corresponding to each first candidate website and the visual feature vector corresponding to each fingerprint cluster; Calculate the structural similarity score between the structural feature vector corresponding to each first candidate website and the structural feature vector corresponding to each fingerprint cluster; Calculate the semantic similarity score between the semantic feature vector corresponding to each first candidate website and the semantic feature vector corresponding to each fingerprint cluster; The visual similarity score, the structural similarity score, and the semantic similarity score are weighted and summed to obtain the similarity score corresponding to each first candidate website; Based on the similarity score corresponding to each first candidate website, multiple second candidate websites are determined.
[0043] It should be understood that the digital fingerprint of the benchmark website includes fingerprint clusters that correspond one-to-one with each webpage. For example, the homepage, login page, activity page, payment page, and customer service page of the benchmark website each form a different fingerprint cluster. Each of the above fingerprint clusters includes a visual feature vector, a structural feature vector, and a semantic feature vector.
[0044] In this embodiment, for each first candidate website, an optional implementation is to use a headless browser to automatically load the first candidate website, capture webpage screenshots, HTML source code, DOM tree, page text and static resource reference relationships, and extract visual feature vectors, structural feature vectors and semantic feature vectors respectively.
[0045] Furthermore, the visual similarity score between the visual feature vector corresponding to each first candidate website and the visual feature vector corresponding to each fingerprint cluster is calculated; the structural similarity score between the structural feature vector corresponding to each first candidate website and the structural feature vector corresponding to each fingerprint cluster is calculated; and the semantic similarity score between the semantic feature vector corresponding to each first candidate website and the semantic feature vector corresponding to each fingerprint cluster is calculated. Then, the visual similarity scores, structural similarity scores, and semantic similarity scores are weighted and summed to obtain the similarity score for each first candidate website. Optionally, the similarity score can be determined using the following formula:
[0046] in, This represents the similarity score, where a, b, and c represent the weight values. Indicates visual similarity score, Indicates the structural similarity score. This represents the semantic similarity score.
[0047] Finally, based on the similarity score corresponding to each first candidate website, multiple second candidate websites are determined.
[0048] In this embodiment, based on a preset benchmark website digital fingerprint, multiple first candidate websites are screened for similarity to obtain second candidate websites that are more similar to the benchmark website, thereby improving the accuracy of identifying counterfeit websites.
[0049] Optionally, determining multiple second candidate websites based on the similarity score corresponding to each first candidate website includes: The first candidate website with a similarity score greater than the first preset threshold is determined as the second candidate website; The visual feature vector, semantic feature vector, structural feature vector, and semantic summary of the first candidate website whose similarity score is less than or equal to the first preset threshold and greater than the second preset threshold are input into the preset multimodal model to obtain the risk identification result. The first candidate website, whose risk identification results indicate that there is a risk, is designated as the second candidate website.
[0050] In this embodiment, for a first candidate website whose similarity score is greater than a first preset threshold, the first candidate website can be determined as a second candidate website. Optionally, the value range of the first preset threshold is 0.90-0.98, and in a specific embodiment, the first preset threshold is set to 0.95.
[0051] In this embodiment, for a first candidate website whose similarity score is less than or equal to a first preset threshold and greater than a second preset threshold, the visual feature vector, semantic feature vector, structural feature vector, and semantic summary of the first candidate website can be input into a preset multimodal model. The multimodal model then performs reasoning based on the input information to obtain the risk identification result.
[0052] Specifically: The multimodal model is input with a screenshot of the first candidate page, OCR text summary, DOM key node summary, and semantic summary. The multimodal model outputs a risk identification result. If the risk identification result indicates that the first candidate website poses a risk, then the first candidate website is designated as the second candidate website.
[0053] In this embodiment, a second candidate website similar to the benchmark website is determined based on the similarity score. For the first candidate website whose similarity score is less than or equal to the first preset threshold and greater than the second preset threshold, a multimodal model is introduced for further semantic verification to identify the risks of the first candidate website. The first candidate website with risks is then determined as the second candidate website, thereby accurately screening the first candidate website for similarity and obtaining a second candidate website similar to the benchmark website.
[0054] For example, please refer to Figure 3 , Figure 3 The "URLs to be systematically analyzed" in the list are the first candidate websites. Figure 3 In the illustrated application scenario, resource crawling and feature extraction are performed on the first candidate website to obtain its corresponding website digital fingerprint. This website digital fingerprint includes visual feature vectors, structural feature vectors, and semantic feature vectors. The website digital fingerprint is then compared with the official digital fingerprint (i.e., the baseline website digital fingerprint in this embodiment) to obtain the similarity score for each website digital fingerprint. Based on this similarity score, suspicious website digital fingerprints are determined, thus completing the similarity screening of the first candidate website and obtaining the second candidate website.
[0055] Optionally, the benchmark website digital fingerprint includes a set of legitimate domain names and legitimate authentication information; The security authentication of each second candidate website is performed based on a preset benchmark website digital fingerprint, resulting in a domain anomaly score, certificate anomaly score, and model confidence level for each second candidate website, including: Based on the set of legitimate domain names, domain name authentication is performed on each of the second candidate websites to obtain the domain name anomaly score corresponding to each of the second candidate websites; Based on the legitimate authentication information, certificate authentication is performed on each of the second candidate websites to obtain the certificate anomaly score corresponding to each second candidate website; The visual feature vector, structural feature vector, semantic feature vector, domain name set, authentication information, and digital fingerprint of the benchmark website corresponding to each second candidate website are input into a preset multimodal model. Consistency inference is performed on each second candidate website through the multimodal model to obtain the model confidence level corresponding to each second candidate website.
[0056] In this embodiment, the main domain, subdomains, intermediate domains in the redirection chain, and resource loading domain of the second candidate website can be precisely compared and fuzzily compared with the legal domains included in the legal domain set to obtain the domain anomaly score corresponding to each second candidate website. The fuzzy comparison can cover situations such as character replacement, hyphen insertion, spelling variations, brand word concatenation, and homographs in internationalized domains.
[0057] Extract the SSL / TLS certificate subject, organization name, alternative name, issuing authority, validity period, and certificate chain integrity of the second candidate website, compare them with the legitimate authentication information, and obtain the certificate anomaly score corresponding to the second candidate website.
[0058] The visual feature vector, structural feature vector, semantic feature vector, domain name set, and authentication information (including page screenshots, DOM key node summaries, form and interface information, domain information, and certificate summaries) corresponding to the second candidate website are input into the multimodal model. The differences between the website fingerprint of the second candidate website and the digital fingerprint of the benchmark website are also input into the multimodal model. The large multimodal model uses this information to perform consistency inference on the second candidate website, obtaining the model confidence score for the second candidate website.
[0059] Please see Figure 4 ,like Figure 4 As shown, the domain name legality is verified by performing digital fingerprinting on suspicious websites (which can be understood as second candidate websites), resulting in a domain name anomaly score Adomain; a certificate background check is performed, resulting in a certificate anomaly score Acert; and cross-modal consistency inference is performed to obtain the model confidence score Conf_mllm.
[0060] Optionally, determining the risk score for each second candidate website based on its similarity score, domain anomaly score, certificate anomaly score, and model confidence score includes: The risk score for each second candidate website is obtained by weighted summing of the similarity score, domain name anomaly score, certificate anomaly score, and model confidence score.
[0061] In this embodiment, the risk score corresponding to the second candidate website is determined based on the following formula:
[0062] Where R represents the risk score, λ1, , , For weight values, represents the similarity score, Adomain represents the domain name anomaly score, Acert represents the certificate anomaly score, and Conf_mllm represents the model confidence score.
[0063] In this embodiment, the domain name of the second candidate website is verified to be abnormal by using the set of legitimate domain names included in the digital fingerprint of the benchmark website; the certificate of the second candidate website is verified to be abnormal by using the legitimate authentication information included in the digital fingerprint of the benchmark website; and the second candidate website is verified to be a counterfeit website by using a multimodal model. This achieves the legitimacy authentication of the second candidate website and thus accurately identifies counterfeit websites.
[0064] Optionally, determining the counterfeit website among the plurality of second candidate websites based on the risk score corresponding to each second candidate website includes: Second candidate websites with risk scores greater than the third preset threshold are identified as high-risk counterfeit websites; The second candidate website with a risk score less than or equal to the third preset threshold and greater than or equal to the fourth preset threshold is identified as a medium-risk counterfeit website. The second candidate website with a risk score lower than the fourth preset threshold is identified as a non-counterfeit website.
[0065] In this embodiment, the counterfeit website is determined based on the relationship between the risk score, the third preset threshold, and the fourth preset threshold.
[0066] An optional implementation involves identifying second candidate websites with a risk score greater than a third preset threshold as high-risk counterfeit websites; identifying second candidate websites with a risk score less than or equal to the third preset threshold but greater than or equal to a fourth preset threshold as medium-risk counterfeit websites; and identifying second candidate websites with a risk score less than the fourth preset threshold as non-counterfeit websites.
[0067] Please see Figure 4 ,exist Figure 4 In the application scenario shown, risk fusion decision is made based on similarity score S1, domain name anomaly score Adomain, certificate anomaly score Acert, and model confidence score Conf_mllm to obtain the URL risk judgment result, and then the URL is determined as risky, risk-free, or partially risky.
[0068] The following section details how to construct a benchmark website's digital fingerprint: Optionally, the method further includes: Extract the visual features of a preset benchmark website to obtain the visual feature vector corresponding to the benchmark website; Extract the structural features of a preset benchmark website to obtain the structural feature vector corresponding to the benchmark website; Extract the semantic features of a preset benchmark website to obtain the semantic feature vector corresponding to the benchmark website; Based on the visual feature vector, structural feature vector, semantic feature vector, set of legal domain names, legal authentication information, semantic summary and business semantic tags corresponding to the benchmark website, a digital fingerprint of the benchmark website is generated.
[0069] It should be understood that the aforementioned benchmark websites can be interpreted as official websites, and the aforementioned visual features include, but are not limited to, page screenshots and HTML source code. The aforementioned structural features include, but are not limited to, the DOM tree structure and certificates. The aforementioned semantic features include, but are not limited to, page titles and body text. The aforementioned set of legitimate domain names represents the official set of legitimate domain names.
[0070] Please see Figure 5 ,like Figure 5 As shown, the official website's webpage screenshots, DOM tree structure, HTML source code, webpage text, and webpage certificate are considered official resources. Features are then extracted from these official resources to obtain the digital fingerprint corresponding to the official website, i.e., the baseline website's digital fingerprint. Optionally, official assets specifically include: a list of the official main domain and subdomains, a set of page URLs, corresponding complete webpage screenshots, HTML source code, DOM tree structure, page titles, and body text, etc.
[0071] To ensure the stability of digital fingerprints, data collection was performed on the same official website at different time windows, on different terminal viewpoints, and under different network environments to form a multi-view sample set. Dynamic advertising slots, timestamps, carousel blocks, and other easily fluctuating areas were weakened or masked to reduce the interference of occasional content on the official fingerprint.
[0072] On the one hand, convolutional neural networks, visual Transformers, or combinations thereof can be used to encode webpage screenshots and key area screenshots, while simultaneously extracting layout features such as page block distribution, button positions, input box density, navigation bar structure, and main color histogram to form a stable representation of the official website's appearance. This allows for the extraction of visual features from a benchmark website, resulting in a corresponding visual feature vector. On the other hand, methods such as tree edit distance, Tree-LSTM, graph neural networks, or Transformer encoders based on serialized DOM can be used to parse HTML and the DOM tree, extracting tag sequences, node levels, node attributes, etc., to obtain a structured DOM representation. This allows for the extraction of structural features from a benchmark website, resulting in a corresponding structural feature vector.
[0073] On the other hand, text such as page titles, navigation items, button text, brand words, login prompts, customer service prompts, form field descriptions, and disclaimers are extracted and supplemented with OCR recognition. Furthermore, a multimodal model can be used to perform semantic summarization of the official website's pages, generating descriptive metadata. These are then used to extract the semantic features of the benchmark website, resulting in the corresponding semantic feature vector.
[0074] On the other hand, the certificate subject name, alternative name (SAN), certificate authority, certificate chain, validity period, server ASN / CDN information, historical resolution information, etc., of the official website are collected as legitimate authentication information.
[0075] Furthermore, based on the visual feature vector, structural feature vector, semantic feature vector, set of legitimate domain names, legitimate authentication information, semantic summary, and business semantic tags corresponding to the benchmark website, a digital fingerprint of the benchmark website is generated. Optionally, the digital fingerprint of the benchmark website can be represented by the following formula: P = {Fv, Fd, Ft, Cb, Dallow, Sdesc} Where P represents the digital fingerprint of the benchmark website, Fv represents the visual feature vector, Fd represents the structural feature vector, Ft represents the semantic feature vector, Cb represents the legitimate authentication information, Dallow represents the set of legitimate domain names, and Sdesc represents the semantic summary and business semantic label generated by the multimodal model.
[0076] For a better understanding of the overall technical solution, please refer to [link / reference]. Figure 6 ,like Figure 6 As shown, screenshots, DOM tree structure, HTML source code, webpage text, and webpage certificates of the official website (i.e., the benchmark website) are collected as official assets of the official website. Feature extraction is performed on the above official resources to obtain the digital fingerprint corresponding to the official website, i.e., the digital fingerprint of the benchmark website.
[0077] A multi-round filtering strategy, namely the pre-filtering process in the above embodiments, is employed to screen a massive number of candidate URLs, obtaining the digital fingerprints of the URLs to be analyzed, which are the first candidate websites in the above embodiments. Similarity filtering is then performed on the digital fingerprints of the URLs to be analyzed to obtain fingerprints of suspicious counterfeit websites, which are the second candidate websites in the above embodiments. Key logic and background authentication are performed on the fingerprints of suspicious counterfeit websites to obtain key indicators of concern, namely, similarity score, domain name anomaly score, certificate anomaly score, and model confidence score in the above embodiments. These key indicators are then fused to classify suspicious risky URLs, that is, to determine suspicious risky URLs as risky URLs, risk-free URLs, and URLs with partial risk.
[0078] Further reference Figure 7As a response to the above Figure 2 To implement the method shown, this application provides an embodiment of a phishing website identification device 700, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0079] This invention provides a spoofing website identification device 700, which includes: The acquisition module 701 is used to acquire multiple first candidate websites; the multiple first candidate websites are determined based on pre-filtering of preset websites. The filtering module 702 is used to perform similarity filtering on the plurality of first candidate websites based on a preset benchmark website digital fingerprint, to obtain a plurality of second candidate websites and a similarity score corresponding to each second candidate website; the benchmark website digital fingerprint is determined based on the webpage features of the benchmark website, and the similarity score is used to characterize the similarity between the corresponding second candidate website and the benchmark website, and the plurality of second candidate websites are a portion of the plurality of first candidate websites; The authentication module 703 is used to perform security authentication on each of the second candidate websites based on a preset benchmark website digital fingerprint, and to obtain the domain anomaly score, certificate anomaly score and model confidence score corresponding to each second candidate website; the domain anomaly score is used to characterize the domain authentication status of the corresponding second candidate website, the certificate anomaly score is used to characterize the certificate authentication status of the corresponding second candidate website, and the model confidence score is obtained by reasoning about the second candidate websites based on a preset multimodal model; The first determining module 704 is used to determine the risk score corresponding to each second candidate website based on the similarity score, domain name anomaly score, certificate anomaly score and model confidence score corresponding to each second candidate website; The second determining module 705 is used to determine the counterfeit website among the plurality of second candidate websites based on the risk score corresponding to each second candidate website.
[0080] Optionally, the digital fingerprint of the benchmark website includes fingerprint clusters that correspond one-to-one with each webpage, and each fingerprint cluster includes a visual feature vector, a structural feature vector, and a semantic feature vector; The filtering module 702 is specifically used for: Calculate the visual similarity score between the visual feature vector corresponding to each first candidate website and the visual feature vector corresponding to each fingerprint cluster; Calculate the structural similarity score between the structural feature vector corresponding to each first candidate website and the structural feature vector corresponding to each fingerprint cluster; Calculate the semantic similarity score between the semantic feature vector corresponding to each first candidate website and the semantic feature vector corresponding to each fingerprint cluster; The visual similarity score, the structural similarity score, and the semantic similarity score are weighted and summed to obtain the similarity score corresponding to each first candidate website; Based on the similarity score corresponding to each first candidate website, multiple second candidate websites are determined.
[0081] Optionally, the filtering module 702 is further specifically used for: The first candidate website with a similarity score greater than the first preset threshold is determined as the second candidate website; The visual feature vector, semantic feature vector, structural feature vector, and semantic summary of the first candidate website whose similarity score is less than or equal to the first preset threshold and greater than the second preset threshold are input into the preset multimodal model to obtain the risk identification result. The first candidate website, whose risk identification results indicate that there is a risk, is designated as the second candidate website.
[0082] Optionally, the benchmark website digital fingerprint includes a set of legitimate domain names and legitimate authentication information; The authentication module 703 is specifically used for: Based on the set of legitimate domain names, domain name authentication is performed on each of the second candidate websites to obtain the domain name anomaly score corresponding to each of the second candidate websites; Based on the legitimate authentication information, certificate authentication is performed on each of the second candidate websites to obtain the certificate anomaly score corresponding to each second candidate website; The visual feature vector, structural feature vector, semantic feature vector, domain name set, authentication information, and digital fingerprint of the benchmark website corresponding to each second candidate website are input into a preset multimodal model. Consistency inference is performed on each second candidate website through the multimodal model to obtain the model confidence level corresponding to each second candidate website.
[0083] Optionally, the first determining module 704 is specifically used for: The risk score for each second candidate website is obtained by weighted summing of the similarity score, domain name anomaly score, certificate anomaly score, and model confidence score.
[0084] Optionally, the first determining module 705 is specifically used for: Second candidate websites with risk scores greater than the third preset threshold are identified as high-risk counterfeit websites; The second candidate website with a risk score less than or equal to the third preset threshold and greater than or equal to the fourth preset threshold is identified as a medium-risk counterfeit website. The second candidate website with a risk score lower than the fourth preset threshold is identified as a non-counterfeit website.
[0085] Optionally, the phishing website identification device 700 further includes: The first extraction module is used to extract the visual features of a preset benchmark website and obtain the visual feature vector corresponding to the benchmark website. The second extraction module is used to extract the structural features of a preset benchmark website and obtain the structural feature vector corresponding to the benchmark website. The third extraction module is used to extract the semantic features of a preset benchmark website and obtain the semantic feature vector corresponding to the benchmark website. The generation module is used to generate a digital fingerprint of the benchmark website based on the visual feature vector, structural feature vector, semantic feature vector, set of legal domain names, legal authentication information, semantic summary and business semantic tags corresponding to the benchmark website.
[0086] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 8 , Figure 8 This is a basic structural block diagram of the computer device in this embodiment. The computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are interconnected via a system bus. It should be noted that only the computer device 8 with components 81-83 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0087] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0088] The memory 81 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 81 may be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 81 may also be an external storage device of the computer device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 8. Of course, the memory 81 may include both the internal storage unit and its external storage device of the computer device 8. In this embodiment, the memory 81 is typically used to store the operating system and various application software installed on the computer device 8, such as the program code for spoofing website identification methods. In addition, the memory 81 can also be used to temporarily store various types of data that have been output or will be output.
[0089] In some embodiments, the processor 82 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 82 is typically used to control the overall operation of the computer device 8. In this embodiment, the processor 82 is used to run program code stored in the memory 81 or process data, for example, to run the program code for the phishing website identification method.
[0090] The network interface 83 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 8 and other electronic devices.
[0091] This application also provides another embodiment, namely, providing a computer-readable storage medium storing the phishing website identification program, which can be executed by at least one processor to cause the at least one processor to perform the steps of the phishing website identification method as described above.
[0092] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware online platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0093] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0094] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A method for identifying counterfeit websites, characterized in that, include: Obtain multiple first-choice websites; The plurality of first candidate websites are determined based on a pre-filtering process of preset websites; Based on the preset benchmark website digital fingerprint, the multiple first candidate websites are screened for similarity to obtain multiple second candidate websites and the similarity score corresponding to each second candidate website. The digital fingerprint of the benchmark website is determined based on the webpage features of the benchmark website. The similarity score is used to characterize the similarity between the corresponding second candidate website and the benchmark website. The multiple second candidate websites are some of the multiple first candidate websites. Based on the preset benchmark website digital fingerprint, security authentication is performed on each of the second candidate websites to obtain the domain name anomaly score, certificate anomaly score and model confidence level corresponding to each second candidate website; The domain name anomaly score is used to characterize the domain name authentication status of the corresponding second candidate website, the certificate anomaly score is used to characterize the certificate authentication status of the corresponding second candidate website, and the model confidence score is obtained by reasoning about the second candidate website based on a preset multimodal model. Based on the similarity score, domain name anomaly score, certificate anomaly score, and model confidence score corresponding to each second candidate website, the risk score corresponding to each second candidate website is determined. Based on the risk score corresponding to each of the second candidate websites, the counterfeit websites among the plurality of second candidate websites are determined; The benchmark website digital fingerprint includes a set of legitimate domain names and legitimate authentication information; The security authentication of each second candidate website is performed based on a preset benchmark website digital fingerprint, resulting in a domain anomaly score, certificate anomaly score, and model confidence level for each second candidate website, including: Based on the set of legitimate domain names, domain name authentication is performed on each of the second candidate websites to obtain the domain name anomaly score corresponding to each of the second candidate websites; Based on the legitimate authentication information, certificate authentication is performed on each of the second candidate websites to obtain the certificate anomaly score corresponding to each second candidate website; The visual feature vector, structural feature vector, semantic feature vector, domain name set, authentication information, and digital fingerprint of the benchmark website corresponding to each second candidate website are input into a preset multimodal model. Consistency inference is performed on each second candidate website through the multimodal model to obtain the model confidence level corresponding to each second candidate website.
2. The method according to claim 1, characterized in that, The benchmark website digital fingerprint consists of fingerprint clusters that correspond one-to-one with each webpage, and each fingerprint cluster includes visual feature vectors, structural feature vectors, and semantic feature vectors. The method, based on a preset benchmark website digital fingerprint, performs similarity screening on the multiple first candidate websites to obtain multiple second candidate websites and a similarity score for each second candidate website, including: Calculate the visual similarity score between the visual feature vector corresponding to each first candidate website and the visual feature vector corresponding to each fingerprint cluster; Calculate the structural similarity score between the structural feature vector corresponding to each first candidate website and the structural feature vector corresponding to each fingerprint cluster; Calculate the semantic similarity score between the semantic feature vector corresponding to each first candidate website and the semantic feature vector corresponding to each fingerprint cluster; The visual similarity score, the structural similarity score, and the semantic similarity score are weighted and summed to obtain the similarity score corresponding to each first candidate website; Based on the similarity score corresponding to each first candidate website, multiple second candidate websites are determined.
3. The method according to claim 2, characterized in that, The step of determining multiple second candidate websites based on the similarity score corresponding to each first candidate website includes: The first candidate website with a similarity score greater than the first preset threshold is determined as the second candidate website; The visual feature vector, semantic feature vector, structural feature vector, and semantic summary of the first candidate website whose similarity score is less than or equal to the first preset threshold and greater than the second preset threshold are input into the preset multimodal model to obtain the risk identification result. The first candidate website, whose risk identification results indicate that there is a risk, is designated as the second candidate website.
4. The method according to claim 1, characterized in that, The process of determining the risk score for each second candidate website based on its similarity score, domain anomaly score, certificate anomaly score, and model confidence score includes: The risk score for each second candidate website is obtained by weighted summing of the similarity score, domain name anomaly score, certificate anomaly score, and model confidence score.
5. The method according to claim 1, characterized in that, The step of determining counterfeit websites among the plurality of second candidate websites based on the risk score corresponding to each second candidate website includes: Second candidate websites with risk scores greater than the third preset threshold are identified as high-risk counterfeit websites; The second candidate website with a risk score less than or equal to the third preset threshold and greater than or equal to the fourth preset threshold is identified as a medium-risk counterfeit website. The second candidate website with a risk score lower than the fourth preset threshold is identified as a non-counterfeit website.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Extract the visual features of a preset benchmark website to obtain the visual feature vector corresponding to the benchmark website; Extract the structural features of a preset benchmark website to obtain the structural feature vector corresponding to the benchmark website; Extract the semantic features of a preset benchmark website to obtain the semantic feature vector corresponding to the benchmark website; Based on the visual feature vector, structural feature vector, semantic feature vector, set of legal domain names, legal authentication information, semantic summary and business semantic tags corresponding to the benchmark website, a digital fingerprint of the benchmark website is generated.
7. A device for identifying counterfeit websites, characterized in that, include: The acquisition module is used to acquire multiple first-choice candidate websites; The plurality of first candidate websites are determined based on a pre-filtering process of preset websites; The filtering module is used to perform similarity filtering on the plurality of first candidate websites based on a preset benchmark website digital fingerprint, to obtain a plurality of second candidate websites and a similarity score corresponding to each second candidate website; the benchmark website digital fingerprint is determined based on the webpage features of the benchmark website, and the similarity score is used to characterize the similarity between the corresponding second candidate website and the benchmark website, and the plurality of second candidate websites are a portion of the plurality of first candidate websites; The authentication module is used to perform security authentication on each of the second candidate websites based on a preset benchmark website digital fingerprint, and to obtain the domain name anomaly score, certificate anomaly score and model confidence score corresponding to each second candidate website. The domain name anomaly score is used to characterize the domain name authentication status of the corresponding second candidate website, the certificate anomaly score is used to characterize the certificate authentication status of the corresponding second candidate website, and the model confidence score is obtained by reasoning about the second candidate website based on a preset multimodal model. The first determining module is used to determine the risk score corresponding to each second candidate website based on the similarity score, domain name anomaly score, certificate anomaly score, and model confidence score corresponding to each second candidate website; The second determining module is used to determine the counterfeit website among the plurality of second candidate websites based on the risk score corresponding to each second candidate website; The benchmark website digital fingerprint includes a set of legitimate domain names and legitimate authentication information; The authentication module is specifically used for: Based on the set of legitimate domain names, domain name authentication is performed on each of the second candidate websites to obtain the domain name anomaly score corresponding to each of the second candidate websites; Based on the legitimate authentication information, certificate authentication is performed on each of the second candidate websites to obtain the certificate anomaly score corresponding to each second candidate website; The visual feature vector, structural feature vector, semantic feature vector, domain name set, authentication information, and digital fingerprint of the benchmark website corresponding to each second candidate website are input into a preset multimodal model. Consistency inference is performed on each second candidate website through the multimodal model to obtain the model confidence level corresponding to each second candidate website.
8. A computer device, characterized in that, The device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method for identifying counterfeit websites as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method for identifying counterfeit websites as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Counterfeit website detection method and device based on large language model
CN119341819A
Phishing website detection method and device, electronic equipment and storage medium
CN120110784A