Fraud website detection method and system
By acquiring and processing the multimodal features of fraudulent websites and utilizing multimodal search engines and domain name matching technology, the problem of misjudgment in fraudulent website detection in existing technologies is solved, and efficient and accurate identification of fraudulent websites is achieved.
Patent Information
- Application Number
- CN202511007747.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-22
- Publication Date
- 2025-09-16
AI Technical Summary
Existing fraud website detection methods are difficult to deal with new types of fraud websites, resulting in a high misjudgment rate and an inability to efficiently and accurately identify phishing websites that impersonate well-known brands.
By obtaining the multimodal features of the web page to be detected, including images, web documents and URL information, and preprocessing them, a multimodal retriever is used to retrieve similar brand information from the brand knowledge base, and domain name matching and comprehensive scoring are combined to determine whether it is a fraudulent website.
It achieves efficient and accurate detection of fraudulent websites, reduces misjudgments, improves detection accuracy, and can identify fraudulent websites that imitate well-known brands.
Smart Images

Figure CN120658504A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method and system for detecting fraudulent websites. Background Art
[0002] With the rapid development of the internet, online fraud methods have become more subtle and diverse, with phishing attacks becoming a major method of fraud. Attackers disguise themselves as legitimate websites or well-known brands to trick users into entering sensitive information such as usernames, passwords, and bank account numbers, resulting in severe financial losses and privacy risks for individuals, businesses, and governments. Since attacks typically target a small number of well-known brands, efficiently and accurately detecting fraudulent websites, especially phishing websites impersonating well-known brands, has become a key research topic in the field of cybersecurity.
[0003] Traditional methods for detecting fraudulent websites primarily include blacklists, rule-based matching, and machine learning models. Blacklist methods rely on lists of known fraudulent websites, but are inadequate for detecting newly emerging ones and prone to misjudgment. Rule-based matching methods identify fraudulent websites by analyzing features such as web page URLs, HTML structures, and keywords. However, due to their fixed rules, they struggle to cope with evolving fraudulent techniques, leading to misjudgments. Machine learning methods use trained models to distinguish fraudulent from legitimate websites. However, as fraudulent website technology continues to evolve, their appearance and content increasingly resemble legitimate websites, making them more susceptible to misjudgment. Therefore, how to address new types of fraudulent websites, improve detection accuracy, and avoid misjudgments has become a key research topic. Summary of the Invention
[0004] Based on the above-mentioned deficiencies in the prior art, the present application provides a method and system for detecting fraudulent websites to improve the accuracy of detecting fraudulent websites and avoid the problem of misjudgment.
[0005] In order to achieve the above objectives, this application provides the following technical solutions:
[0006] The first aspect of the present application provides a method for detecting fraudulent websites, comprising:
[0007] Obtaining multimodal features of the web page to be detected; wherein the multimodal features include at least images, web page documents, and website information;
[0008] Preprocessing the image, the web document, and the URL information respectively to obtain a logo image corresponding to the image, text corresponding to the web document, and domain name information corresponding to the URL information;
[0009] Based on the logo image and the text, a multimodal retriever is used to retrieve multiple brands similar to the webpage to be detected and their corresponding brand information from a brand knowledge base; wherein the multimodal retriever is preliminarily extracted from an open source database and an authoritative brand list using optical character recognition technology, and a modification task is added to the extracted data to construct the result;
[0010] Determining whether a target brand exists among the brands based on the text, the domain name information, and the brands and their corresponding brand information;
[0011] If the target brand exists among the brands, determining whether the webpage to be detected is a suspected fraudulent webpage based on the domain name list of the target brand;
[0012] If the webpage to be detected is a suspected fraudulent webpage, a comprehensive score is calculated based on the domain name information;
[0013] When the comprehensive score is less than a preset threshold, it is determined that the web page to be detected is a fraudulent web page.
[0014] Optionally, in the above-mentioned fraudulent website detection method, the pre-processing of the image, the webpage document, and the URL information to obtain a logo image corresponding to the image, text corresponding to the webpage document, and domain name information corresponding to the URL information includes:
[0015] Input the image into the object detection model to obtain multiple candidate landmark images and their corresponding confidence scores;
[0016] Extracting a maximum confidence score from all confidence scores, and using the candidate logo image corresponding to the maximum confidence score as the logo image corresponding to the image;
[0017] Filtering target text that meets preset requirements from the web document using regular expressions and tag filtering rules, and organizing and cleaning the target text to obtain text corresponding to the web document;
[0018] The website address information is parsed to obtain domain name information corresponding to the website address information.
[0019] Optionally, in the above-mentioned fraudulent website detection method, the step of retrieving multiple brands similar to the webpage to be detected and their corresponding brand information from a brand knowledge base using a multimodal retriever based on the logo image and the text includes:
[0020] Encoding the text using a text encoder to obtain an encoded text, and encoding the logo image using an image encoder to obtain an encoded image;
[0021] Performing a weighted summation on the encoded text and the encoded image to obtain a weighted vector, and normalizing the weighted vector to obtain a combined webpage code;
[0022] Encoding the brand name and brand logo in the brand knowledge base using the image encoder to obtain an encoded brand name and an encoded brand logo, and calculating a brand combination code based on the encoded brand name and the encoded brand logo;
[0023] The cosine similarity between the combined webpage code and the brand combination code is calculated, and a multimodal retriever is used to retrieve a plurality of brands having the highest matching cosine similarity and their corresponding brand information from the brand knowledge base.
[0024] Optionally, in the above-mentioned fraudulent website detection method, determining whether the target brand exists among the brands based on the text, the domain name information, and the brands and their corresponding brand information includes:
[0025] Inputting the text, the domain name information, and each brand and its corresponding brand information into a large language model for recognition to obtain a recognition result;
[0026] When the recognition result includes the target brand and its corresponding recognition data, determining that the target brand exists among the brands;
[0027] When the recognition result indicates that the target brand is not found, the logo image and each brand and its corresponding brand information are input into the large language model for recognition to obtain a target recognition result;
[0028] When the target recognition result is a result including the target brand and its corresponding recognition data, determining that the target brand exists among the brands;
[0029] When the recognition result is a result that the target brand is not found, it is determined that the target brand does not exist among the brands.
[0030] Optionally, in the above-mentioned fraudulent website detection method, judging whether the webpage to be detected is a suspected fraudulent webpage based on the domain name list of the target brand includes:
[0031] Determining whether the domain name corresponding to the web page to be detected matches a target domain name in a list of legal domain names of the target brand;
[0032] If the domain name corresponding to the web page to be detected does not match the target domain name in the legal domain name list of the target brand, then for each target domain name in the legal domain name list, the distance between the domain name corresponding to the web page to be detected and the target domain name is calculated using a distance algorithm;
[0033] Calculating the similarity between the web page to be detected and the target domain name based on the domain name corresponding to the web page to be detected, the target domain name, and the distance;
[0034] Extracting a maximum similarity from all similarities, and determining whether the maximum similarity is greater than a preset threshold;
[0035] If the maximum similarity is greater than the preset threshold, the webpage to be detected is determined to be a suspected fraudulent webpage;
[0036] If the maximum similarity is not greater than the preset threshold, it is determined that the web page to be detected is a fraudulent web page.
[0037] Optionally, in the above-mentioned fraudulent website detection method, calculating the comprehensive score based on the domain name information includes:
[0038] Querying resource data corresponding to the domain name information from an online resource library; wherein the resource data includes at least the number of search results, ranking position, content relevance, and official statement data;
[0039] Calculating the number of search results, the ranking position, the content relevance, and the scores of the official statement data respectively, and calculating the authority score and the maximum similarity score of the web page to be detected;
[0040] According to preset weights, the score of the number of search results, the score of the ranking position, the score of the content relevance, the score of the official statement data, the authority score and the score of the maximum similarity are weighted and summed to obtain a comprehensive score.
[0041] A second aspect of the present application provides a fraudulent website detection system, the detection system comprising: a pre-processing module, a multimodal brand recognition module, a domain name matching verification module, and a comprehensive evaluation module;
[0042] The preprocessing module is configured to obtain multimodal features of a web page to be detected, and to preprocess the image, web document, and URL information respectively to obtain a logo image corresponding to the image, text corresponding to the web document, and domain name information corresponding to the URL information; wherein the multimodal features include at least the image, the web document, and the URL information;
[0043] The multimodal brand recognition module is configured to retrieve, from a brand knowledge base, a plurality of brands similar to the web page to be detected and their corresponding brand information using a multimodal retriever based on the logo image and the text; wherein the multimodal retriever extracts brand data from an open source database and an authoritative brand list using optical character recognition technology in advance, and constructs the extracted data by adding a change task;
[0044] The domain name matching verification module is used to determine whether a target brand exists among the brands based on the text, the domain name information, and the brands and their corresponding brand information; and if the target brand exists among the brands, determine whether the webpage to be detected is a suspected fraudulent webpage based on the domain name list of the target brand;
[0045] The comprehensive evaluation module is used to calculate a comprehensive score based on the domain name information if the web page to be detected is a suspected fraudulent web page, and to determine that the web page to be detected is a fraudulent web page when the comprehensive score is less than a preset threshold.
[0046] Optionally, in the above-mentioned fraudulent website detection system, the preprocessing module performs preprocessing on the image, the web document, and the URL information respectively to obtain a logo image corresponding to the image, text corresponding to the web document, and domain name information corresponding to the URL information, specifically for:
[0047] Input the image into the object detection model to obtain multiple candidate landmark images and their corresponding confidence scores;
[0048] Extracting a maximum confidence score from all confidence scores, and using the candidate logo image corresponding to the maximum confidence score as the logo image corresponding to the image;
[0049] Filtering target text that meets preset requirements from the web document using regular expressions and tag filtering rules, and organizing and cleaning the target text to obtain text corresponding to the web document;
[0050] The website address information is parsed to obtain domain name information corresponding to the website address information.
[0051] Optionally, in the fraudulent website detection system described above, the multimodal brand recognition module uses a multimodal retriever to retrieve multiple brands similar to the webpage to be detected and their corresponding brand information from a brand knowledge base based on the logo image and the text, specifically for:
[0052] Encoding the text using a text encoder to obtain an encoded text, and encoding the logo image using an image encoder to obtain an encoded image;
[0053] Performing a weighted summation on the encoded text and the encoded image to obtain a weighted vector, and normalizing the weighted vector to obtain a combined webpage code;
[0054] Encoding the brand name and brand logo in the brand knowledge base using the image encoder to obtain an encoded brand name and an encoded brand logo, and calculating a brand combination code based on the encoded brand name and the encoded brand logo;
[0055] The cosine similarity between the combined webpage code and the brand combination code is calculated, and a multimodal retriever is used to retrieve a plurality of brands having the highest matching cosine similarity and their corresponding brand information from the brand knowledge base.
[0056] Optionally, in the fraudulent website detection system, the domain name matching verification module determines whether the target brand exists among the brands based on the text, the domain name information, and the brands and their corresponding brand information, specifically for:
[0057] Inputting the text, the domain name information, and each brand and its corresponding brand information into a large language model for recognition to obtain a recognition result;
[0058] When the recognition result includes the target brand and its corresponding recognition data, determining that the target brand exists among the brands;
[0059] When the recognition result indicates that the target brand is not found, the logo image and each brand and its corresponding brand information are input into the large language model for recognition to obtain a target recognition result;
[0060] When the target recognition result is a result including the target brand and its corresponding recognition data, determining that the target brand exists among the brands;
[0061] When the recognition result is a result that the target brand is not found, it is determined that the target brand does not exist among the brands.
[0062] Optionally, in the above-mentioned fraudulent website detection system, the domain name matching verification module is executed based on the domain name list of the target brand to determine whether the web page to be detected is a suspected fraudulent web page, specifically for:
[0063] Determining whether the domain name corresponding to the web page to be detected matches a target domain name in a list of legal domain names of the target brand;
[0064] If the domain name corresponding to the web page to be detected does not match the target domain name in the legal domain name list of the target brand, then for each target domain name in the legal domain name list, the distance between the domain name corresponding to the web page to be detected and the target domain name is calculated using a distance algorithm;
[0065] Calculating the similarity between the web page to be detected and the target domain name based on the domain name corresponding to the web page to be detected, the target domain name, and the distance;
[0066] Extracting a maximum similarity from all similarities, and determining whether the maximum similarity is greater than a preset threshold;
[0067] If the maximum similarity is greater than the preset threshold, the webpage to be detected is determined to be a suspected fraudulent webpage;
[0068] If the maximum similarity is not greater than the preset threshold, it is determined that the web page to be detected is a fraudulent web page.
[0069] Optionally, in the above-mentioned fraudulent website detection system, the comprehensive evaluation module calculates a comprehensive score based on the domain name information, specifically for:
[0070] Querying resource data corresponding to the domain name information from an online resource library; wherein the resource data includes at least the number of search results, ranking position, content relevance, and official statement data;
[0071] Calculating the number of search results, the ranking position, the content relevance, and the scores of the official statement data respectively, and calculating the authority score and the maximum similarity score of the web page to be detected;
[0072] According to preset weights, the score of the number of search results, the score of the ranking position, the score of the content relevance, the score of the official statement data, the authority score and the score of the maximum similarity are weighted and summed to obtain a comprehensive score.
[0073] The present application provides a method for detecting fraudulent websites. The method obtains multimodal features of a web page to be detected, wherein the multimodal features include at least images, web documents, and URL information. The images, web documents, and URL information are then preprocessed to obtain logo images corresponding to the images, text corresponding to the web documents, and domain name information corresponding to the URL information. Based on the logo images and text, a multimodal searcher is used to retrieve multiple brands similar to the web page to be detected and their corresponding brand information from a brand knowledge base. The multimodal searcher preliminarily extracts brand data from open source databases and authoritative brand lists using optical character recognition technology and constructs a task by adding changes to the extracted data. Based on the text, domain name information, and each brand and its corresponding brand information, a determination is made as to whether a target brand exists among each of the brands. If the target brand exists among each of the brands, a determination is made as to whether the web page to be detected is suspected fraudulent based on the domain name list of the target brand. If the web page to be detected is suspected fraudulent, a comprehensive score is calculated based on the domain name information. Finally, if the comprehensive score is less than a preset threshold, the web page to be detected is determined to be fraudulent. By extracting key information from web page screenshots, HTML documents and URLs, and combining brand identification, domain name matching verification and comprehensive scoring, efficient and accurate detection of fraudulent websites can be achieved, thus avoiding the problem of misjudgment. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without any creative work.
[0075] Figure 1 A schematic diagram of the structure of a fraudulent website detection system provided in an embodiment of the present application;
[0076] Figure 2 A flowchart of a method for detecting fraudulent websites provided in an embodiment of the present application;
[0077] Figure 3 A schematic flow chart of a multimodal feature preprocessing method provided in another embodiment of the present application;
[0078] Figure 4 A flowchart of a method for obtaining a brand and its brand information provided in another embodiment of the present application;
[0079] Figure 5A flowchart of a method for determining a target brand provided in another embodiment of the present application;
[0080] Figure 6 A flowchart of a method for determining a suspected fraudulent web page provided in another embodiment of the present application;
[0081] Figure 7 A flowchart of a method for calculating a comprehensive score provided in another embodiment of the present application;
[0082] Figure 8 A schematic structural diagram of a fraudulent website detection system provided in another embodiment of the present application. DETAILED DESCRIPTION
[0083] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0084] In this application, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.
[0085] The embodiment of the present application provides a method for detecting fraudulent websites, which is applied to a fraudulent website detection system to improve the accuracy of detecting fraudulent websites and avoid misjudgment.
[0086] Alternatively, as Figure 1 As shown, an embodiment of the present application provides a fraudulent website detection system, including: a preprocessing module, a multimodal brand recognition module, a domain name matching verification module and a comprehensive evaluation module.
[0087] The preprocessing module is used to obtain screenshots, HTML documents and URL information of the detected web page, and preprocess the screenshots, HTML documents and URL information.
[0088] The multimodal brand identification module consists of a multimodal retriever and a multimodal brand identifier. The multimodal brand identification module is used to determine whether the website under test is a fraudulent website. The multimodal retriever is used to retrieve the target brand range of the test webpage. The multimodal brand identifier is used to identify the brand of the test website.
[0089] The domain name matching verification module is used to verify whether the domain name of the web page to be detected is consistent with the legitimate domain name of the detected target brand, so as to further determine whether the web page is a fraudulent website.
[0090] The comprehensive evaluation module is used to prevent misjudgment of real websites. It combines the confidence score of the multimodal brand identifier for web page brand recognition and the web page domain similarity score to calculate a comprehensive score, thereby determining the final judgment result of the web page to be detected based on the comprehensive score.
[0091] Therefore, based on the above-mentioned fraudulent website detection system, the present application embodiment provides a fraudulent website detection method, such as Figure 2 As shown, the specific steps include:
[0092] S201: Obtain multimodal features of the web page to be detected.
[0093] Among them, multimodal features can include pictures, web documents and URL information, thereby comprehensively reflecting the properties of the web page from the three dimensions of vision, text and structure to ensure that the information extracted from the web page to be detected is complete and can accurately describe the characteristics of the web page to be detected, thereby providing a data basis for subsequent analysis and scoring of fraudulent websites.
[0094] In addition, the web page document is an HTML document, and the website address information is URL information.
[0095] S202: Pre-process the image, web document, and URL information respectively to obtain a logo image corresponding to the image, text corresponding to the web document, and domain name information corresponding to the URL information.
[0096] Among them, preprocessing methods can include model detection, rule screening and parsing methods, etc.
[0097] Optionally, in another embodiment of the present application, a specific implementation of step S202 is as follows: Figure 3 As shown, the specific steps include:
[0098] S301: Input the image into the target detection model to obtain multiple candidate logo images and their corresponding confidence scores.
[0099] It is understandable that the target detection model is trained based on a deep learning mechanism, such as Faster R-CNN or YOLO series, which are customized training for the basic network. Therefore, as long as the image is input into the target detection model, the target detection model will output a set of candidate landmark images and their corresponding confidence scores.
[0100] S302: Extract the confidence score with the maximum value from all the confidence scores, and use the candidate logo picture corresponding to the confidence score with the maximum value as the logo picture corresponding to the picture.
[0101] Specifically, in order to select the best logo image as the subsequent evaluation criterion, the confidence score with the maximum value is extracted from all the confidence scores, and the candidate logo image corresponding to the confidence score is taken as the extraction result, that is, as the logo image recorded as .
[0102] S303: Filter out target text that meets preset requirements from the web page document using regular expressions and tag filtering rules, and organize and clean up the target text to obtain text corresponding to the web page document.
[0103] It should be noted that in order to organize and clean the extracted text, thereby removing redundant spaces and special characters, thereby reducing the input content and improving the model processing efficiency, the HTML document of the website to be tested will be parsed first, and then regular expressions and tag filtering rules will be written to remove various tags in the HTML (such as 、 Layout tags, CSS style codes, JavaScript scripts (including tracking codes) and other irrelevant information are extracted to obtain the target text. The key text content is extracted from the target text by combining the tags and attributes in the tag filtering rules, including the web page title, meta description, text related to the logo, header and footer text, etc., to obtain the text corresponding to the web page document, which is recorded as .
[0104] S304: parse the website information to obtain the domain name information corresponding to the website information.
[0105] Specifically, the input UR information is parsed to extract the domain name part, and further extract the path, parameters and other information in the URL, recorded as . Used to check whether the domain name information exists in the specified domain name set to determine whether it is a legitimate website.
[0106] S203 , based on the logo image and text, a multimodal retriever is used to retrieve multiple brands similar to the web page to be detected and their corresponding brand information from a brand knowledge base.
[0107] Among them, the multimodal retriever pre-extracts brand data from open source databases (such as PhishTank, KnowPhish, etc.) and authoritative brand lists through optical character recognition technology (OCR technology), and adds change tasks to the extracted data to construct it.
[0108] It is understandable that in order to determine whether the web page to be detected is a fraudulent website, it is first necessary to determine whether the web page to be detected imitates a well-known brand. Therefore, it is necessary to use a multimodal search engine to identify the brand and further compare it with potential imitation brands one by one. Therefore, by searching the web page to be detected through the modal search engine, multiple brands similar to the web page to be detected and their corresponding brand information will be obtained, that is, the web page to be detected may contain information that imitates a well-known brand.
[0109] The brand knowledge base is a multimodal database that includes the brand names of potential target brands of fraudulent websites ( )、Brand Logo( ),domain name( ) and other basic information, which is built from open source databases and uses open source data such as PhishTank, KnowPhish, etc.
[0110] Optionally, in another embodiment of the present application, a specific implementation of step S203 is as follows: Figure 4 As shown, the specific steps include:
[0111] S401. Encode the text using a text encoder to obtain an encoded text, and encode the logo image using an image encoder to obtain an encoded image.
[0112] Specifically, if you want to know whether the web page to be tested has imitated a well-known brand, you first need to use a text encoder to Encode and generate a coding vector with representative information, that is, the encoded text The text encoder first uses a pre-trained large language model to directly call the text encoding layer of models such as GPT-4 and Claude3, and uses a multimodal large language model for unified processing. Then, a lightweight solution is adopted, using the Transformer model, and the input is cleaned HTML text (such as titles, descriptions, etc.) for training.
[0113] Then, we use the image encoder Logo picture Encode and get the encoded picture Among them, the image encoder Use existing models (e.g. CLIP-ViT or DINOv2 models).
[0114] S402: Perform weighted summation on the encoded text and the encoded image to obtain a weighted vector, and perform normalization processing on the weighted vector to obtain a combined web page code.
[0115] It is understandable that to ensure the consistency of the lengths of different codes in the vector space, it is necessary to first perform weighted summation of the coded text and coded image, and then perform normalization to obtain the combined web page code. .
[0116] Specifically, the normalized processing formula is:
[0117]
[0118] in, 、 Encoding and If no brand logo is detected in the web page to be detected, the web page encoding is calculated only by the text features (i.e., the encoding result of the HTML text of the web page to be detected), that is, there is no encoding processing of the logo image.
[0119] S403: Encode the brand name and brand logo in the brand knowledge base using an image encoder to obtain an encoded brand name and an encoded brand logo, and calculate a brand combination code based on the encoded brand name and the encoded brand logo.
[0120] Specifically, in order to retrieve brands that may have imitative behaviors with the grid to be detected from the brand knowledge base, it is necessary to and brand logo The two variables are encoded using the image encoder to obtain the encoded brand name and coded brand logo , and calculate the brand combination code , the specific calculation method is:
[0121]
[0122] in, 、 Encoding and If a brand has no associated logo, the brand code is calculated solely from text features. If a brand has multiple logos, the combined code for each logo is calculated separately. For example, for the three logos of Brand A, the combined codes are: BrandVec_A1 = α·v_name + β·v_logo1, BrandVec_A2 = α·v_name + β·v_logo2, and BrandVec_A3 = α·v_name + β·v_logo3. α is the weight for text encoding, β is the weight for image encoding, v_name is the brand name, and v_logo is the brand logo.
[0123] S404: Calculate the cosine similarity between the combined web page code and the brand combination code, and use a multimodal retriever to retrieve multiple brands with the highest cosine similarity matches and their corresponding brand information from the brand knowledge base.
[0124] It is understandable that in order to measure the degree of match between the webpage to be tested and the brands in the brand knowledge base, the cosine similarity between the combined webpage code and the brand combination code can be calculated first. Then, a dot product search method, i.e., a multimodal searcher, is used to find the top k brands with the highest cosine similarity matches from the brand knowledge base BKB.
[0125] Specifically, the search method is:
[0126]
[0127] in, represents a multimodal retriever, Indicates the web page to be detected. represents the brand knowledge base, Indicates the number of target brands expected to be acquired. Coding for web pages, Code the brand portfolio.
[0128] Therefore, after the above search operation, the search results finally obtained are k brands and their corresponding brand information including brand names and legal domain names.
[0129] S204: Determine whether the target brand exists in each brand based on the text, domain name information, and each brand and its corresponding brand information.
[0130] Specifically, the target brand refers to the brand that the webpage under inspection has imitated. Therefore, to determine the target brand from the brands obtained in step S203, a large language model can be used to determine whether the target brand exists among the brands based on the text, domain name information, and each brand and its corresponding brand information. This ensures recognition accuracy while optimizing computational effort and resource investment. If the target brand exists among the brands, it indicates that the webpage under inspection has imitated a well-known brand, so step S205 is executed.
[0131] Alternatively, if the target brand does not exist among the brands, it means that the webpage to be detected does not imitate well-known brands, and thus it can be further demonstrated that the webpage to be detected is a legitimate website, not a fraudulent website.
[0132] Optionally, in another embodiment of the present application, a specific implementation of step S204 is as follows: Figure 5 As shown, the specific steps include:
[0133] S501: Input text, domain name information, and each brand and its corresponding brand information into a large language model for recognition to obtain a recognition result.
[0134] It should be noted that in order to lock the target brand, the text of the web page to be detected, the domain name information, and the various brands and their corresponding brand information retrieved by the multimodal retrieval module need to be input into the large language model for identification, so that the corresponding prompt words are designed through the large language model, and a target brand is accurately determined by referring to the list of imitation brands. At the same time, the large language model is allowed to apply online information for retrieval, and the final brand recognition result is not limited to the list of potential imitation brands. In order to make the answer given by the large language model more logical and confident, a thinking chain is incorporated into the prompt words to guide reasoning. After analysis and processing by the large language model, the final output recognition result includes the extracted target brand name, confidence score and supporting evidence. Among them, the confidence score is a quantitative assessment of the accuracy of the identified target brand, with a value range of 0-1. The closer the value is to 1, the higher the recognition confidence. Supporting evidence is the basis for identifying the target brand, including the specific brand logo in HTML, the association characteristics between the domain name and the target brand, etc. The expression of the large language model is:
[0135]
[0136] in, is the detected target brand name, For the target brand domain name, Score for confidence, For supporting evidence.
[0137] It should also be noted that if the large language model fails to identify the target brand during analysis, it directly outputs "No relevant brand found," and thus executes step S503. If the large language model can identify the target brand during analysis, it directly outputs the recognition result, and thus executes step S502.
[0138] In addition, to ensure the reliability of the recognition results, a preset threshold is set. When the output confidence score is less than the preset threshold, it indicates that the large language model has a low degree of confidence in the current recognition result, and it will also enter the image recognition module, that is, execute step S503.
[0139] S502: When the recognition result includes the target brand and its corresponding recognition data, it is determined that the target brand exists among the brands.
[0140] S503: When the recognition result indicates that the target brand is not found, the logo image and each brand and its corresponding brand information are input into the large language model for recognition to obtain a target recognition result.
[0141] Specifically, if brand recognition based solely on text fails to identify the target brand, that is, when the recognition result is that the target brand is not found, or when the output confidence score is less than the preset threshold, the image brand recognition module will be enabled, and the multimodal large language model will be used to analyze the webpage screenshot. That is, the logo image and each brand and its corresponding brand information are input into the large language model for recognition. The target recognition result also includes the target brand name, confidence score and supporting evidence. Among them, the supporting evidence includes the specific characteristics, color, layout and other information of the brand logo in the webpage screenshot. The expression of the large language model at this time is:
[0142]
[0143] in, is the detected target brand name, For the target brand domain name, Score for confidence, For supporting evidence.
[0144] In addition, if the target brand cannot be identified from both the text and image information, it is considered that the webpage to be detected does not imitate a well-known brand and is determined not to be a fraudulent webpage, that is, step S505 is executed.
[0145] If the target brand can be identified through the image information, step S504 is executed.
[0146] S504: When the target recognition result includes the target brand and its corresponding recognition data, it is determined that the target brand exists among the brands.
[0147] S505: When the recognition result indicates that the target brand is not found, it is determined that the target brand does not exist among the brands.
[0148] It is understandable that when the identification result is that no target brand is found, it indicates that the web page to be detected is a legitimate web page.
[0149] S205: Determine whether the webpage to be detected is a suspected fraudulent webpage based on the domain name list of the target brand.
[0150] It should be noted that after the target brand is determined through the multimodal brand module, it is necessary to further check whether the domain name of the web page to be detected is consistent with the legal domain name of the detected target brand, so as to determine whether the web page to be detected is a fraudulent website, that is, based on the domain name list of the target brand, determine whether the web page to be detected is a suspected fraudulent web page. If the web page to be detected is determined to be a suspected fraudulent web page, it means that it cannot be accurately determined at this time that the web page to be detected is a fraudulent web page, and the final judgment process needs to be executed to avoid misjudgment, so step S206 is executed.
[0151] Optionally, if the web page to be detected is not a suspected fraudulent web page, it means that the web page to be detected is a legitimate web page.
[0152] Optionally, in another embodiment of the present application, a specific implementation of step S205 is as follows: Figure 6 As shown, the specific steps include:
[0153] S601: Determine whether the domain name corresponding to the webpage to be detected matches the target domain name in the legal domain name list of the target brand.
[0154] Specifically, a string matching algorithm is used to determine whether the domain name corresponding to the webpage being tested exactly matches a domain name in the list of legal domain names for the target brand in the offline knowledge base. If this indicates a suspicion of fraudulent activity on the website, further fuzzy matching analysis of the domain name should be performed, thus executing step S602. Alternatively, if the match is successful, the domain name is preliminarily deemed legal.
[0155] S602 : For each target domain name in the legal domain name list, calculate the distance between the domain name corresponding to the web page to be detected and the target domain name using a distance algorithm.
[0156] It should be noted that when the domain name corresponding to the webpage to be tested does not match the target domain name in the target brand's legal domain name list, it is necessary to use the Levenshtein distance algorithm to calculate the distance between the webpage domain name and each domain name in the legal domain name list, and further calculate the similarity score. Based on the similarity score, it is possible to determine whether the webpage to be tested is a fraudulent website imitating a well-known brand. Therefore, it is necessary to first use the distance algorithm to calculate the distance between the domain name corresponding to the webpage to be tested and the target domain name. The specific calculation formula is: .
[0157] in, is the target domain name, The domain name of the website to be tested.
[0158] S603: Calculate the similarity between the web page to be detected and the target domain name based on the domain name corresponding to the web page to be detected, the target domain name, and the distance.
[0159] Specifically, the formula for calculating the similarity between the web page to be detected and the target domain name is:
[0160]
[0161] in, and Target domain names for target brands and the domain name of the website to be tested The absolute value of the length, is the Levenshtein distance between two domain name strings.
[0162] S604: Extract the maximum similarity from all similarities, and determine whether the maximum similarity is greater than a preset threshold.
[0163] It should be noted that a brand domain name may include multiple subdomains or domain name variations, so the maximum similarity needs to be extracted from all similarities.
[0164] Specifically, in order to determine whether the website to be detected is a fraudulent website imitating a well-known brand, a threshold can be set, that is, to determine whether the maximum similarity is greater than the preset threshold. If the maximum similarity is greater than the preset threshold, it means that the website to be detected is still a suspected fraudulent website and needs to be marked as a suspected match, and a comprehensive score is calculated for final judgment, so step S605 is executed.
[0165] S605: Determine whether the web page to be detected is a suspected fraudulent web page.
[0166] Specifically, when the maximum similarity is greater than a preset threshold, it indicates that the webpage to be detected may be a suspected fraudulent webpage, and therefore further determination is required, so step S206 is executed.
[0167] S606: Determine whether the web page to be detected is a fraudulent web page.
[0168] Specifically, when the maximum similarity is not greater than a preset threshold, it can be considered that the website to be detected is a fraudulent website that imitates a well-known brand.
[0169] S206. Calculate a comprehensive score based on the domain name information.
[0170] It should be noted that in order to prevent misjudgment of real websites and avoid missing fraudulent websites with similar domain name spellings, when the web page to be detected is a suspected fraudulent web page, it is also necessary to calculate the comprehensive score based on the domain name information.
[0171] Optionally, in another embodiment of the present application, a specific implementation of step S206 is as follows: Figure 7 As shown, the specific steps include:
[0172] S701: Query resource data corresponding to the domain name information from an online resource library.
[0173] Among them, resource data includes at least the number of search results, ranking position, content relevance and official statement data.
[0174] Specifically, in order to reduce the risk of misjudging real websites and ensure that fraudulent websites with similar domain names to legitimate ones are not missed, the confidence score of the multimodal brand identifier for web page brand recognition and the similarity score of the web page domain name are combined to calculate a comprehensive score to identify fraudulent websites with similar domain name spellings. Therefore, it is necessary to use search engines to query information related to the domain name information of the web page to be detected from online resources, and the number of search results needs to be considered separately. , ranking position , content relevance and whether there is an official statement Four variables, and calculate the corresponding scores respectively to get the scores 、 、 、 In addition, use tools such as Moz's Domain Authority to calculate the authority score of the website. , while considering the domain similarity score .
[0175] S702: Calculate the number of search results, ranking position, content relevance, and scores of official statement data, and calculate the authority score and maximum similarity score of the web page to be detected.
[0176] Specifically, the formula for calculating the score of the number of search results is:
[0177]
[0178] in, is the number of current search results, The maximum number of search results during the test.
[0179] The formula for calculating the score for a ranking position is:
[0180]
[0181] in, is the number of search result pages, Ranking of target domains in search results.
[0182] The formula for calculating the score of official statement data is:
[0183]
[0184] If there is an official statement in the search results, it will get 100 points; if there is no official statement, it will get 0 points.
[0185] The formula for calculating the content relevance score is: or .in, , Confidence scores computed for the large language model in the multimodal brand identifier.
[0186] The formula for calculating the authority score of the web page to be tested is: , The domain authority of Moz.
[0187] The formula for calculating the maximum similarity score is: , Score for domain similarity.
[0188] S703: Based on preset weights, the scores of the number of search results, the ranking position, the content relevance, the official statement data, the authority, and the maximum similarity are weighted and summed to obtain a comprehensive score.
[0189] Specifically, the scores of the number of search results, the ranking position, the content relevance, the official statement data, the authority score and the maximum similarity score are weighted according to the preset weights. Perform weighted summation to obtain a comprehensive score :
[0190]
[0191] Among them, the preset weights satisfy .
[0192] S207: When the comprehensive score is less than a preset threshold, the web page to be detected is determined to be a fraudulent web page.
[0193] It's important to note that the composite score comprehensively estimates the visibility and brand relevance of the website being tested. If the domain name of the webpage being tested is indexed by mainstream search engines and receives a high score, it's likely a benign website. This is because benign websites typically have a long lifespan, high-quality content, and meet search engine algorithm requirements, making them easily indexed and receiving a high weight. However, most fraudulent websites have a short lifespan and are typically active only briefly, used to commit fraud or disseminate malicious content. Due to their short existence, low-quality content, and potential for user or security agency reporting, these websites are typically not indexed by search engines.
[0194] Specifically, based on the above overall judgment, if the domain name information is suspected to match a legitimate domain name and the comprehensive score is greater than the preset threshold, the website under test is considered to be a benign website. If the domain name is suspected to match a legitimate domain name, but the comprehensive score is less than the threshold, or the domain name does not fuzzily match the legitimate domain name, the detected webpage is determined to have a fraud risk.
[0195]
[0196] in, is the preset threshold of similarity, The preset threshold for the comprehensive score.
[0197] The present application provides a method for detecting fraudulent websites. The method obtains multimodal features of a web page to be detected, wherein the multimodal features include at least images, web documents, and URL information. The images, web documents, and URL information are then preprocessed to obtain logo images corresponding to the images, text corresponding to the web documents, and domain information corresponding to the URL information. Based on the logo images and text, a multimodal searcher is used to retrieve multiple brands similar to the web page to be detected and their corresponding brand information from a brand knowledge base. The multimodal searcher preliminarily extracts brand data from open source databases and authoritative brand lists using optical character recognition technology, and constructs the extracted data by adding change tasks. The searcher then determines whether a target brand exists among each brand based on the text, domain information, and each brand and its corresponding brand information. If the target brand exists among each brand, the searched web page is determined to be suspected of being fraudulent based on the target brand's domain name list. If the web page to be detected is suspected of being fraudulent, a comprehensive score is calculated based on the domain information. Finally, if the comprehensive score is less than a preset threshold, the searched web page is determined to be fraudulent. By extracting key information from web page screenshots, HTML documents and URLs, and combining brand identification, domain name matching verification and comprehensive scoring, efficient and accurate detection of fraudulent websites can be achieved, thus avoiding the problem of misjudgment.
[0198] Another embodiment of the present application provides a fraudulent website detection system, such as Figure 8 As shown, the detection system includes: a pre-processing module 801, a multimodal brand recognition module 802, a domain name matching verification module 803 and a comprehensive evaluation module 804.
[0199] Preprocessing module 801 is used to obtain multimodal features of the web page to be detected and to preprocess the image, web document, and URL information to obtain the logo image corresponding to the image, the text corresponding to the web document, and the domain name information corresponding to the URL information. The multimodal features include at least the image, web document, and URL information.
[0200] The multimodal brand identification module 802 is used to use a multimodal retriever to retrieve multiple brands and their corresponding brand information from a brand knowledge base based on logo images and text. The multimodal retriever pre-processes brand data from open source databases and authoritative brand lists using optical character recognition technology and then constructs a task by adding changes to the extracted data.
[0201] The domain name matching verification module 803 is used to determine whether the target brand exists among the brands based on the text, domain name information, and each brand and its corresponding brand information, and if the target brand exists among the brands, determine whether the web page to be detected is a suspected fraudulent web page based on the domain name list of the target brand.
[0202] The comprehensive evaluation module 804 is used to calculate a comprehensive score based on the domain name information if the web page to be detected is a suspected fraudulent web page, and to determine that the web page to be detected is a fraudulent web page when the comprehensive score is less than a preset threshold.
[0203] It should be noted that the specific working process of the above modules in the embodiment of the present application can refer to steps S201 to S207 in the above method embodiment, and will not be repeated here.
[0204] Optionally, in a fraudulent website detection system provided in another embodiment of the present application, a preprocessing module performs preprocessing on images, web documents, and URL information respectively to obtain a logo image corresponding to the image, text corresponding to the web document, and domain name information corresponding to the URL information, specifically for:
[0205] Input the image into the object detection model to obtain multiple candidate logo images and their corresponding confidence scores.
[0206] The confidence score with the maximum value is extracted from all the confidence scores, and the candidate logo image corresponding to the maximum confidence score is used as the logo image corresponding to the image.
[0207] Regular expressions and tag filtering rules are used to filter out target text that meets preset requirements from web page documents, and the target text is sorted and cleaned to obtain the text corresponding to the web page document.
[0208] Parse the URL information to obtain the domain name information corresponding to the URL information.
[0209] Optionally, in another embodiment of the present application, a fraudulent website detection system is provided, wherein the multimodal brand recognition module uses a multimodal retriever to retrieve multiple brands similar to the webpage to be detected and their corresponding brand information from a brand knowledge base based on logo images and text, specifically for:
[0210] The text is encoded using a text encoder to obtain an encoded text, and the logo image is encoded using an image encoder to obtain an encoded image.
[0211] The encoded text and the encoded image are weightedly summed to obtain a weighted vector, and the weighted vector is normalized to obtain the combined web page encoding.
[0212] An image encoder is used to encode brand names and brand logos in a brand knowledge base to obtain encoded brand names and encoded brand logos, and brand combination codes are calculated based on the encoded brand names and encoded brand logos.
[0213] The cosine similarity between the combined web page code and the brand combination code is calculated, and a multimodal retriever is used to retrieve multiple brands with the highest cosine similarity matching and their corresponding brand information from the brand knowledge base.
[0214] Optionally, in a fraudulent website detection system provided in another embodiment of the present application, the domain name matching verification module determines whether a target brand exists among various brands based on the text, domain name information, and various brands and their corresponding brand information, specifically for:
[0215] The text, domain name information, and each brand and its corresponding brand information are input into the large language model for recognition to obtain the recognition result.
[0216] When the recognition result is a result including the target brand and its corresponding recognition data, it is determined that the target brand exists among the brands.
[0217] When the recognition result is that the target brand is not found, the logo image and each brand and its corresponding brand information are input into the large language model for recognition to obtain the target recognition result.
[0218] When the target recognition result is a result including the target brand and its corresponding recognition data, it is determined that the target brand exists among the brands.
[0219] When the recognition result is that the target brand is not found, it is determined that the target brand does not exist among the brands.
[0220] Optionally, in a fraudulent website detection system provided in another embodiment of the present application, the domain name matching verification module determines whether the webpage to be detected is a suspected fraudulent webpage based on the domain name list of the target brand, specifically for:
[0221] Determine whether the domain name corresponding to the web page to be detected matches the target domain name in the target brand's legal domain name list.
[0222] If the domain name corresponding to the web page to be detected does not match the target domain name in the legal domain name list of the target brand, the distance between the domain name corresponding to the web page to be detected and the target domain name is calculated using a distance algorithm for each target domain name in the legal domain name list.
[0223] Based on the domain name corresponding to the web page to be detected, the target domain name and the distance, the similarity between the web page to be detected and the target domain name is calculated.
[0224] The maximum similarity is extracted from all similarities, and it is determined whether the maximum similarity is greater than a preset threshold.
[0225] If the maximum similarity is greater than a preset threshold, the web page to be detected is determined to be a suspected fraudulent web page.
[0226] If the maximum similarity is not greater than the preset threshold, it is determined that the web page to be detected is a fraudulent web page.
[0227] Optionally, in a fraudulent website detection system provided in another embodiment of the present application, the comprehensive evaluation module calculates a comprehensive score based on domain name information, specifically for:
[0228] Querying resource data corresponding to the domain name information from an online resource library; the resource data includes at least the number of search results, ranking position, content relevance, and official statement data;
[0229] Calculate the number of search results, ranking position, content relevance, and official statement data scores respectively, and calculate the authority score and maximum similarity score of the web page to be tested;
[0230] According to the preset weights, the scores of the number of search results, the ranking position, the content relevance, the official statement data, the authority score and the maximum similarity score are weighted and summed to obtain a comprehensive score.
[0231] It should be noted that the specific working processes of the various modules provided in the above embodiments of the present application can refer to the corresponding steps in the above method embodiments, and will not be repeated here.
[0232] It should also be noted that the fraudulent website detection system provided in the embodiment of the present application has the technical effects of any of the above embodiments, and the embodiments of the present application are not described in detail here.
[0233] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0234] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for detecting fraudulent websites, characterized in that: include: Obtaining multimodal features of the web page to be detected; wherein the multimodal features include at least images, web page documents, and website information; Preprocessing the image, the web document, and the URL information respectively to obtain a logo image corresponding to the image, text corresponding to the web document, and domain name information corresponding to the URL information; Based on the logo image and the text, a multimodal retriever is used to retrieve multiple brands similar to the webpage to be detected and their corresponding brand information from a brand knowledge base; wherein the multimodal retriever is preliminarily extracted from an open source database and an authoritative brand list using optical character recognition technology, and a modification task is added to the extracted data to construct the result; Determining whether a target brand exists among the brands based on the text, the domain name information, and the brands and their corresponding brand information; If the target brand exists among the brands, determining whether the webpage to be detected is a suspected fraudulent webpage based on the domain name list of the target brand; If the webpage to be detected is a suspected fraudulent webpage, a comprehensive score is calculated based on the domain name information; When the comprehensive score is less than a preset threshold, it is determined that the web page to be detected is a fraudulent web page.
2. The method according to claim 1, characterized in that The pre-processing of the image, the webpage document, and the URL information to obtain a logo image corresponding to the image, text corresponding to the webpage document, and domain name information corresponding to the URL information includes: Input the image into the object detection model to obtain multiple candidate landmark images and their corresponding confidence scores; Extracting a maximum confidence score from all confidence scores, and using the candidate logo image corresponding to the maximum confidence score as the logo image corresponding to the image; Filtering target text that meets preset requirements from the web document using regular expressions and tag filtering rules, and organizing and cleaning the target text to obtain text corresponding to the web document; The website address information is parsed to obtain domain name information corresponding to the website address information.
3. The method according to claim 1, characterized in that The method of retrieving multiple brands similar to the web page to be detected and their corresponding brand information from a brand knowledge base using a multimodal retriever based on the logo image and the text includes: Encoding the text using a text encoder to obtain an encoded text, and encoding the logo image using an image encoder to obtain an encoded image; Performing a weighted summation on the encoded text and the encoded image to obtain a weighted vector, and normalizing the weighted vector to obtain a combined webpage code; Encoding the brand name and brand logo in the brand knowledge base using the image encoder to obtain an encoded brand name and an encoded brand logo, and calculating a brand combination code based on the encoded brand name and the encoded brand logo; The cosine similarity between the combined webpage code and the brand combination code is calculated, and a multimodal retriever is used to retrieve a plurality of brands having the highest matching cosine similarity and their corresponding brand information from the brand knowledge base.
4. The method according to claim 1, wherein The determining whether the target brand exists in each of the brands according to the text, the domain name information, and each of the brands and their corresponding brand information includes: Inputting the text, the domain name information, and each brand and its corresponding brand information into a large language model for recognition to obtain a recognition result; When the recognition result includes the target brand and its corresponding recognition data, determining that the target brand exists among the brands; When the recognition result indicates that the target brand is not found, the logo image and each brand and its corresponding brand information are input into the large language model for recognition to obtain a target recognition result; When the target recognition result is a result including the target brand and its corresponding recognition data, determining that the target brand exists among the brands; When the recognition result is a result that the target brand is not found, it is determined that the target brand does not exist among the brands.
5. The method according to claim 1, characterized in that The step of determining whether the webpage to be detected is a suspected fraudulent webpage based on the domain name list of the target brand includes: Determining whether the domain name corresponding to the web page to be detected matches a target domain name in a list of legal domain names of the target brand; If the domain name corresponding to the web page to be detected does not match the target domain name in the legal domain name list of the target brand, then for each target domain name in the legal domain name list, the distance between the domain name corresponding to the web page to be detected and the target domain name is calculated using a distance algorithm; Calculating the similarity between the web page to be detected and the target domain name based on the domain name corresponding to the web page to be detected, the target domain name, and the distance; Extracting a maximum similarity from all similarities, and determining whether the maximum similarity is greater than a preset threshold; If the maximum similarity is greater than the preset threshold, the webpage to be detected is determined to be a suspected fraudulent webpage; If the maximum similarity is not greater than the preset threshold, it is determined that the web page to be detected is a fraudulent web page.
6. The method according to claim 5, characterized in that Calculating a comprehensive score based on the domain name information includes: Querying resource data corresponding to the domain name information from an online resource library; wherein the resource data includes at least the number of search results, ranking position, content relevance, and official statement data; Calculating the number of search results, the ranking position, the content relevance, and the scores of the official statement data respectively, and calculating the authority score and the maximum similarity score of the web page to be detected; According to preset weights, the score of the number of search results, the score of the ranking position, the score of the content relevance, the score of the official statement data, the authority score and the score of the maximum similarity are weighted and summed to obtain a comprehensive score.
7. A fraudulent website detection system, characterized in that: The detection system includes: a pre-processing module, a multi-modal brand recognition module, a domain name matching verification module and a comprehensive evaluation module; The preprocessing module is configured to obtain multimodal features of a web page to be detected, and to preprocess the image, web document, and URL information respectively to obtain a logo image corresponding to the image, text corresponding to the web document, and domain name information corresponding to the URL information; wherein the multimodal features include at least the image, the web document, and the URL information; The multimodal brand recognition module is configured to retrieve, from a brand knowledge base, a plurality of brands similar to the web page to be detected and their corresponding brand information using a multimodal retriever based on the logo image and the text; wherein the multimodal retriever extracts brand data from an open source database and an authoritative brand list using optical character recognition technology in advance, and constructs the extracted data by adding a change task; The domain name matching verification module is used to determine whether a target brand exists among the brands based on the text, the domain name information, and the brands and their corresponding brand information; and if the target brand exists among the brands, determine whether the webpage to be detected is a suspected fraudulent webpage based on the domain name list of the target brand; The comprehensive evaluation module is used to calculate a comprehensive score based on the domain name information if the web page to be detected is a suspected fraudulent web page, and to determine that the web page to be detected is a fraudulent web page when the comprehensive score is less than a preset threshold.
8. The system according to claim 7, characterized in that The preprocessing module performs preprocessing on the image, the web document, and the URL information respectively to obtain a logo image corresponding to the image, text corresponding to the web document, and domain name information corresponding to the URL information, specifically for: Input the image into the object detection model to obtain multiple candidate landmark images and their corresponding confidence scores; Extracting a maximum confidence score from all confidence scores, and using the candidate logo image corresponding to the maximum confidence score as the logo image corresponding to the image; Filtering target text that meets preset requirements from the web document using regular expressions and tag filtering rules, and organizing and cleaning the target text to obtain text corresponding to the web document; The website address information is parsed to obtain domain name information corresponding to the website address information.
9. The system according to claim 7, wherein: The multimodal brand recognition module is configured to retrieve multiple brands similar to the web page to be detected and their corresponding brand information from a brand knowledge base using a multimodal retriever based on the logo image and the text, specifically for: Encoding the text using a text encoder to obtain an encoded text, and encoding the logo image using an image encoder to obtain an encoded image; Performing a weighted summation on the encoded text and the encoded image to obtain a weighted vector, and normalizing the weighted vector to obtain a combined webpage code; Encoding the brand name and brand logo in the brand knowledge base using the image encoder to obtain an encoded brand name and an encoded brand logo, and calculating a brand combination code based on the encoded brand name and the encoded brand logo; The cosine similarity between the combined webpage code and the brand combination code is calculated, and a multimodal retriever is used to retrieve a plurality of brands having the highest matching cosine similarity and their corresponding brand information from the brand knowledge base.
10. The system according to claim 7, wherein: The domain name matching verification module is executed to determine whether the target brand exists among the brands based on the text, the domain name information, and the brands and their corresponding brand information, and is specifically used to: Inputting the text, the domain name information, and each brand and its corresponding brand information into a large language model for recognition to obtain a recognition result; When the recognition result includes the target brand and its corresponding recognition data, determining that the target brand exists among the brands; When the recognition result indicates that the target brand is not found, the logo image and each brand and its corresponding brand information are input into the large language model for recognition to obtain a target recognition result; When the target recognition result is a result including the target brand and its corresponding recognition data, determining that the target brand exists among the brands; When the recognition result is a result that the target brand is not found, it is determined that the target brand does not exist among the brands.
Citation Information
Patent Citations
Hybrid phishing website detection method and device, electronic equipment and storage medium
CN114070653A
Fraud website detection method and device based on multi-mode large language model
CN120185937A
Anti-phishing webpage detection system and method thereof
CN120200832A