Webpage processing method and device, computer equipment and storage medium

By obtaining brand features and web page type features, and combining web page recognition models and meta-information, the problem of low accuracy in traditional web page phishing recognition is solved, and more efficient web page phishing recognition is achieved.

CN120611072APending Publication Date: 2025-09-09TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410257797.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-06
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Traditional methods of identifying web page phishing through keyword matching technology have the problem of low accuracy.

Method used

By obtaining the brand characteristics and web page type characteristics of the target brand, the target web page is initially screened from the candidate web page set, the target web page recognition model is called to identify the web page type and brand, and the web page meta information is combined to perform counterfeiting identification to determine the web page counterfeiting results.

Benefits of technology

Improves the accuracy of phishing identification, ensures that the identified web pages are genuine and of the type associated with the target brand, and reduces the misidentification of counterfeit web pages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611072A_ABST
    Figure CN120611072A_ABST
Patent Text Reader

Abstract

The invention relates to a webpage processing method and device, computer equipment, a storage medium and a computer program product. The embodiment of the invention can be applied to various scenes such as cloud technology, artificial intelligence, intelligent traffic, auxiliary driving and the like. The method comprises the steps of determining a target webpage from a candidate webpage set based on brand features corresponding to a target brand and type features corresponding to a target webpage type; calling a target webpage identification model, and performing target webpage type identification processing and target brand identification processing on the target webpage based on the webpage text corresponding to the target webpage; under the condition that the identification result is that the target webpage belongs to the target webpage type and the target webpage is associated with the target brand, performing counterfeit identification processing on the target webpage based on webpage meta-information corresponding to the target webpage to obtain a webpage counterfeit identification result corresponding to the target webpage; and based on the webpage counterfeit identification result, determining a brand counterfeit identification result corresponding to the target brand. By adopting the method, the counterfeit identification accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a web page processing method, apparatus, computer equipment, storage medium, and computer program product. Background Art

[0002] With the continuous advancement of computer technology, the internet has become an indispensable part of people's lives and work, enabling users to access a vast array of web pages. Web pages serve as a bridge between brands and users, providing product information, service details, user guides, and more, helping users understand brands and answering questions. With the development of the internet, the need to identify counterfeit brand-related web pages has become increasingly urgent.

[0003] However, traditional methods usually use keyword matching technology to identify counterfeit products on brand-related web pages. However, various keywords that reflect product counterfeiting may appear in the web page content. Identification using keyword matching technology has certain limitations and is prone to missed detections, which in turn leads to low counterfeit identification accuracy. Summary of the Invention

[0004] Based on this, it is necessary to provide a web page processing method, apparatus, computer equipment, computer-readable storage medium and computer program product that can improve the accuracy of counterfeit identification in order to address the above technical problems.

[0005] This application provides a web page processing method, including:

[0006] Acquire target brands;

[0007] Determining a target webpage from a set of candidate webpages based on brand features corresponding to the target brand and type features corresponding to the target webpage type;

[0008] Invoking a target webpage recognition model to perform target webpage type recognition and target brand recognition on the target webpage based on the webpage text corresponding to the target webpage;

[0009] If the identification result shows that the target webpage belongs to the target webpage type and the target webpage is associated with the target brand, performing counterfeit identification processing on the target webpage based on webpage meta information corresponding to the target webpage to obtain a webpage counterfeit identification result corresponding to the target webpage;

[0010] Based on the webpage phishing identification result, a brand phishing identification result corresponding to the target brand is determined.

[0011] This application also provides a web page processing device, including:

[0012] Brand acquisition module, used to acquire target brands;

[0013] a webpage screening module for determining a target webpage from a set of candidate webpages based on brand characteristics corresponding to the target brand and type characteristics corresponding to the target webpage type;

[0014] A webpage content recognition module is used to call a target webpage recognition model and perform target webpage type recognition and target brand recognition on the target webpage based on the webpage text corresponding to the target webpage;

[0015] a webpage phishing identification module configured to, when an identification result indicates that the target webpage belongs to the target webpage type and is associated with the target brand, perform phishing identification processing on the target webpage based on webpage meta information corresponding to the target webpage, and obtain a webpage phishing identification result corresponding to the target webpage;

[0016] The brand counterfeiting identification module is configured to determine a brand counterfeiting identification result corresponding to the target brand based on the webpage counterfeiting identification result.

[0017] The present application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned web page processing method when executing the computer program.

[0018] The present application also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the web page processing method are implemented.

[0019] The present application also provides a computer program product, including a computer program, which implements the steps of the above-mentioned web page processing method when executed by a processor.

[0020] The webpage processing method, apparatus, computer device, storage medium, and computer program product described above obtain a target brand; determine a target webpage from a candidate webpage set based on brand features corresponding to the target brand and type features corresponding to the target webpage type; invoke a target webpage identification model to perform target webpage type identification and target brand identification on the target webpage based on the webpage text corresponding to the target webpage; if the identification result indicates that the target webpage belongs to the target webpage type and is associated with the target brand, perform counterfeit identification on the target webpage based on webpage meta-information corresponding to the target webpage to obtain a webpage counterfeit identification result corresponding to the target webpage; and determine a brand counterfeit identification result corresponding to the target brand based on the webpage counterfeit identification result. In this way, target webpages related to the target brand and belonging to the target webpage type are initially screened out from the candidate webpage set based on brand features and type features. The target webpage is then reviewed using the target webpage identification model to verify whether the target webpage is related to the target brand and whether it belongs to the target webpage type. The initial screening and review process accurately identifies target webpages related to the target brand and belonging to the target webpage type, thereby helping to improve the accuracy of subsequent counterfeit identification. Fake webpages try to mimic real webpages in their content, but webpage meta-information is generally difficult to counterfeit. Counterfeit identification of the target webpage based on its corresponding meta-information can improve the accuracy of counterfeit detection. Furthermore, determining the brand counterfeit detection result corresponding to the target brand based on the webpage counterfeit detection result can also improve the accuracy of counterfeit detection for the target brand. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0022] Figure 1 A diagram illustrating an application environment of a web page processing method in one embodiment;

[0023] Figure 2 1 is a flow chart of a web page processing method according to an embodiment;

[0024] Figure 3 A schematic diagram of the structure of a web page recognition model in one embodiment;

[0025] Figure 4 A schematic diagram of a process for determining a phishing identification result in one embodiment;

[0026] Figure 5A schematic diagram of a process for determining a webpage phishing identification result in another embodiment;

[0027] Figure 6 A schematic diagram of ICP filing information in one embodiment;

[0028] Figure 7 A schematic diagram of WHOIS filing information in one embodiment;

[0029] Figure 8 A schematic diagram of IP address information in one embodiment;

[0030] Figure 9 1. A flowchart of a method for identifying counterfeit web pages by scanning a QR code for authenticity verification in one embodiment;

[0031] Figure 10 2 is a flow chart of a method for identifying counterfeit web pages by scanning a QR code for authenticity verification in another embodiment;

[0032] Figure 11 is a structural block diagram of a web page processing device in one embodiment;

[0033] Figure 12 is a diagram of the internal structure of a computer device in one embodiment;

[0034] Figure 13 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0036] The embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc.

[0037] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to enable data computing, storage, processing, and sharing. Cloud technology involves big data.

[0038] Big data refers to a collection of data that cannot be captured, managed, and processed by conventional software tools within a certain timeframe. It is a massive, high-growth, and diverse information asset that requires new processing models to have stronger decision-making power, insight discovery, and process optimization capabilities. With the advent of the cloud era, big data has also attracted increasing attention. Big data requires special technologies to effectively process large amounts of data within a tolerable timeframe. Technologies applicable to big data include large-scale parallel processing databases, data mining, distributed file systems, distributed databases, cloud computing platforms, the Internet, and scalable storage systems. Through the method of this application, web pages related to the target brand and belonging to the target web page type can be identified from a large number of web pages, and it can be further determined whether the identified web page is a counterfeit web page, thereby determining the counterfeit results of the target brand.

[0039] The solutions provided in the embodiments of this application involve technologies such as machine learning based on artificial intelligence, and are specifically described through the following embodiments:

[0040] The web page processing method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. The terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other devices. The terminal 102 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car-mounted devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The server 104 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0041] It is understood that the terminal and the server can be used alone to execute the webpage processing method provided in the embodiment of the present application. The terminal and the server can also be used together to execute the webpage processing method provided in the embodiment of the present application.

[0042] For example, the server obtains a target brand and determines a target webpage from a candidate webpage set based on brand features corresponding to the target brand and type features corresponding to the target webpage type. The server then invokes a target webpage identification model and, based on the webpage text corresponding to the target webpage, performs target webpage type identification and target brand identification on the target webpage. If the identification result indicates that the target webpage belongs to the target webpage type and is associated with the target brand, the server then performs counterfeit identification on the target webpage based on the webpage meta-information corresponding to the target webpage, obtains a webpage counterfeit identification result corresponding to the target webpage, and determines a brand counterfeit identification result corresponding to the target brand based on the webpage counterfeit identification result.

[0043] In one embodiment, Figure 2 As shown, a web page processing method is provided, and the method is applied to a computer device as an example. The computer device can be a terminal or a server. It is understood that the method can be executed by the terminal or server itself alone, or can be implemented through interaction between the terminal and the server.

[0044] Step S202: Obtain the target brand.

[0045] The target brand refers to the brand to be identified as counterfeit. The target brand can be determined as needed. For example, the target brand can be specified by the user or randomly selected from existing brands.

[0046] Specifically, the computer device may obtain the target brand from a local device or another device, and determine the brand counterfeiting identification result corresponding to the target brand based on the relevant webpage of the target brand.

[0047] Step S204 : determining a target web page from the candidate web page set based on the brand characteristics corresponding to the target brand and the type characteristics corresponding to the target web page type.

[0048] Brand features are features that describe brand information, such as brand name, brand logo, brand slogan, brand keywords, etc.

[0049] The target webpage type refers to the type of webpage to be identified as counterfeit. The target webpage type can be determined as needed. For example, the target webpage type can be pre-set or user-specified.

[0050] Type features are features that describe the type of web page information. For example, type features can be web page type names, web page type keywords, web page type synonyms, etc.

[0051] The candidate webpage set includes multiple candidate webpages. Webpages related to the target brand and belonging to the target webpage type need to be identified from the candidate webpage set. Existing webpages can be obtained from various sources as candidate webpages to form the candidate webpage set. Alternatively, user-specified webpages can be used as candidate webpages to form the candidate webpage set.

[0052] Specifically, the computer device may obtain brand features corresponding to the target brand and type features corresponding to the target webpage type, and based on the brand features corresponding to the target brand and the type features corresponding to the target webpage type, perform a preliminary screening of the candidate webpage set, selecting as the target webpage a candidate webpage that matches the brand features and the type features in the candidate webpage set. For example, the brand features may be matched with the webpage content of the candidate webpage, and the type features may be matched with the webpage content of the candidate webpage, and the candidate webpage that successfully matches both may be selected as the target webpage.

[0053] Step S206 , calling the target webpage recognition model to perform target webpage type recognition processing and target brand recognition processing on the target webpage based on the webpage text corresponding to the target webpage.

[0054] The web page recognition model is an artificial intelligence model used to identify the web page type of a web page and the brands associated with or involved in the web page. The input data of the web page recognition model includes the web page text corresponding to the web page to be recognized, and the output data includes the recognition result of the web page to be recognized. The target web page recognition model refers to the trained web page recognition model. The target web page type recognition process refers to identifying whether the web page belongs to the target web page type. The target brand recognition process refers to identifying whether the web page is associated with the target brand. Based on the web page text and training labels corresponding to the training web pages, the initial web page recognition model undergoes supervised training to obtain the target web page recognition model.

[0055] In one embodiment, the input data of the web page recognition model include the web page text and the target brand corresponding to the web page to be recognized. The web page recognition model is used to identify the web page type of the web page and whether the web page is related to the target brand. Further, the web page recognition model can be a model corresponding to the target web page type, that is, the web page recognition model is specifically used to identify whether the web page belongs to the target web page type, and for the target web page type recognition process, the recognition result output by the web page recognition model is yes or no. In one embodiment, the input data of the web page recognition model include the web page text, the target brand and the target web page type corresponding to the web page to be recognized. The web page recognition model can be a model corresponding to the target web page type and the target brand, that is, the web page recognition model is used to identify whether the web page belongs to the target web page type and whether the web page is related to the target brand. The recognition result output by the web page recognition model is yes or no.

[0056] The webpage text corresponding to a webpage refers to the text content of the webpage. For example, the webpage text includes the webpage title, webpage body, webpage tags, webpage URL and other text related to the webpage.

[0057] Specifically, the computer device may invoke a target webpage recognition model to review the target webpage to determine whether the target webpage belongs to the target webpage type and whether the target webpage is related to the target brand. The computer device may obtain the target webpage recognition model locally or from another device, input the webpage text corresponding to the target webpage into the target webpage recognition model, and use the target webpage recognition model to perform target webpage type recognition and target brand recognition on the webpage text corresponding to the target webpage, thereby obtaining a recognition result corresponding to the target webpage.

[0058] Step S208 , when the identification result shows that the target webpage belongs to the target webpage type and is associated with the target brand, counterfeit identification processing is performed on the target webpage based on the webpage meta information corresponding to the target webpage to obtain a webpage counterfeit identification result corresponding to the target webpage.

[0059] Among them, web page text is usually information directly displayed on the web page. Web page meta information is usually information not directly displayed on the web page. Web page meta information is information related to web page registration. For example, web page meta information can be the registered web page registrant information. Since fake web pages will try their best to imitate real web pages in the page content, they can achieve a fake-real effect. However, web page text can be counterfeited, but web page meta information is usually difficult to counterfeit, such as ICP filing, WHOIS filing, IP address, etc., which are difficult to counterfeit. Therefore, the web page is subjected to counterfeit identification processing based on the web page meta information corresponding to the web page. Counterfeit identification processing of a web page refers to identifying whether the web page is counterfeit.

[0060] Specifically, when the target webpage identification model determines that the target webpage belongs to the target webpage type and is associated with the target brand, the target webpage can be further identified as counterfeit. The computer device obtains webpage meta information corresponding to the target webpage and, based on the webpage meta information, performs counterfeit identification on the target webpage to obtain a counterfeit identification result corresponding to the target webpage. For example, the target brand is compared with the brand registered in the webpage meta information. If the comparison is consistent, the target webpage is determined to be authentic; if the comparison is inconsistent, the target webpage is determined to be a counterfeit.

[0061] Step S210: determining a brand counterfeiting identification result corresponding to the target brand based on the webpage counterfeiting identification result.

[0062] Specifically, for a target brand, there may be multiple target web pages. Based on the webpage phishing identification results corresponding to each target web page, a brand phishing identification result corresponding to the target brand is determined. The brand phishing identification result corresponding to the target brand indicates the distribution of counterfeit web pages within the target web page type associated with the target brand. For example, the brand phishing identification result may include information such as the number of counterfeit web pages and the percentage of counterfeit web pages.

[0063] Understandably, the counterfeit identification results for target brands are beneficial for protecting brand equity, helping brands better assess the counterfeiting situation their products face and take appropriate action to protect their rights. These results also contribute to maintaining market order and protecting user rights.

[0064] In the webpage processing method described above, target webpages related to the target brand and belonging to the target webpage type are initially screened from a candidate webpage set based on brand and type characteristics. The target webpages are then reviewed using a target webpage identification model to determine whether they are related to the target brand and whether they belong to the target webpage type. This initial screening and review process allows for accurate identification of target webpages related to the target brand and belonging to the target webpage type, thereby improving the accuracy of subsequent counterfeit identification. Counterfeit webpages attempt to mimic authentic webpages in their content, but webpage meta-information is generally difficult to counterfeit. Counterfeit identification of target webpages based on their corresponding webpage meta-information can improve the accuracy of counterfeit identification for webpages. Furthermore, based on the webpage counterfeit identification results for the target webpage, the brand counterfeit identification results corresponding to the target brand are determined, further improving the accuracy of counterfeit identification for the target brand.

[0065] In one embodiment, determining a target webpage from a set of candidate webpages based on brand features corresponding to the target brand and type features corresponding to the target webpage type includes:

[0066] Obtaining a brand name set corresponding to a target brand as a brand feature corresponding to the target brand, and performing keyword matching on candidate web pages in the candidate web page set based on the brand feature to obtain an intermediate web page;

[0067] A type keyword corresponding to the target webpage type is obtained as a type feature corresponding to the target webpage type, and keyword matching is performed on the intermediate webpage based on the type feature to obtain the target webpage.

[0068] The brand name set corresponding to the target brand includes the brand names corresponding to the target brand. For example, the brand name set includes the full name, abbreviation, English name, and full pinyin of the target brand. The type keywords corresponding to the target webpage type are keywords that describe the webpage type. For example, type keywords include verbs, nouns, and exclusion words that describe the webpage type.

[0069] Candidate web pages are web pages to be matched. Intermediate web pages are candidate web pages that match brand features. Target web pages are web pages that match both brand features and category features.

[0070] Specifically, the computer device may obtain a set of brand names corresponding to the target brand as brand features corresponding to the target brand, perform keyword matching on candidate web pages in the candidate web page set based on the brand features, match the brand features with the web page text of the candidate web pages, and use the successfully matched candidate web pages as the intermediate web pages. It is understood that if the web page text of the candidate web page contains at least one brand name from the brand name set, the candidate web page may be used as the intermediate web page.

[0071] The intermediate webpage can be considered a webpage related to the target brand. Furthermore, the computer device can obtain a type keyword corresponding to the target webpage type as a type feature corresponding to the target webpage type. Based on the type feature, the computer device performs keyword matching on the intermediate webpage, matches the type feature with the webpage text of the candidate webpage, and selects the successfully matched intermediate webpage as the target webpage.

[0072] In the above embodiment, first, based on the brand name set corresponding to the target brand, target brand-related web pages are preliminarily screened out from the candidate web page set, and then based on the type keywords corresponding to the target web page type, web pages of the target web page type are preliminarily screened out from the target brand-related web pages. This allows for quick preliminarily screening out web pages related to the target brand and belonging to the target web page type.

[0073] In one embodiment, keyword matching is performed on the intermediate webpage based on the type feature to obtain the target webpage, including:

[0074] Obtaining an intermediate webpage that successfully matches the first and second keywords in the type feature and fails to match the third keyword in the type feature as the target webpage;

[0075] The type keywords corresponding to the target webpage type include first-category keywords, second-category keywords, and third-category keywords. First-category keywords are nouns associated with the target webpage type. For example, first-category keywords can be nouns describing the role and function of the target webpage type. Second-category keywords are verbs associated with the target webpage type. For example, second-category keywords can be verbs describing the role and function of the target webpage type. Third-category keywords are words unrelated to the target webpage type.

[0076] Specifically, keyword matching is performed on the intermediate web pages using multiple types of keywords to determine the target web page. The type features corresponding to the target web page type include first-category keywords, second-category keywords, and third-category keywords. When keyword matching is performed on the intermediate web pages based on the type features, the computer device obtains the intermediate web page that successfully matches both the first-category keywords and the second-category keywords in the type features, but fails to match the third-category keywords in the type features, as the target web page.

[0077] In the above embodiment, keyword matching is performed on intermediate web pages using multiple types of keywords to determine the target web page, which can improve the accuracy of web page screening and further help improve the accuracy of subsequent counterfeit identification.

[0078] In one embodiment, a candidate webpage set is screened based on the brand characteristics of a target brand to obtain webpages related to the target brand. Specifically, the full name, abbreviation, English name, and full pinyin spelling of the target brand can be expanded and used as brand characteristics of the target brand. Using these expanded words, keyword matching is performed on candidate webpages in the candidate webpage set to initially obtain webpages related to the target brand. Furthermore, based on the type keywords of the target webpage type, the target brand's related webpages are screened to obtain webpages related to the target brand and belonging to the target webpage type. Taking a QR code verification webpage as an example, QR code verification webpage keywords are set. These keywords are divided into three types: one type is related to scanning verification, such as anti-counterfeiting, verification, traceability code, backtracking code, etc.; one type is scanning behavior words, such as query, queried, first scan, successful scan, etc.; and the other type is white words, or exclusion words, which are generally unrelated to the QR code verification webpage. Keyword matching is then performed again on the target brand's related webpages using these three keywords. A successful match requires that both the first and second keywords must be matched, and the third keyword must not be matched. By using the scan code verification web page keywords, based on the relevant web pages of the target brand, the scan code verification web pages related to the target brand are obtained.

[0079] In one embodiment, the target webpage identification model is called to perform target webpage type identification and target brand identification on the target webpage based on the webpage text corresponding to the target webpage, including:

[0080] The web page text corresponding to the target brand and the target web page is input into the target web page recognition model; the feature extraction branch of the target web page recognition model is used to perform feature extraction processing on the web page text and the target brand to obtain the web page features corresponding to the target web page; the first prediction branch of the target web page recognition model is used to perform target web page type recognition processing on the web page features to obtain the target web page type recognition result; the second prediction branch of the target web page recognition model is used to perform target brand recognition processing on the web page features to obtain the target brand recognition result.

[0081] The webpage recognition model includes a feature extraction branch, a first prediction branch, and a second prediction branch. The feature extraction branch performs feature extraction on its input data. Feature extraction refers to extracting features from the input data. Information representing the essential attributes of the data is extracted from the raw data so that the model can better learn the inherent patterns of the data. For example, feature extraction can be performed using a convolutional neural network, a language model, or a BERT model. The input data of the feature extraction branch includes the webpage text corresponding to the target brand and target webpage, and the output data is the webpage features corresponding to the target webpage. The first prediction branch performs target webpage type identification on its input data. Target webpage type identification is used to determine whether the webpage belongs to the target webpage type. The input data of the first prediction branch includes the webpage features corresponding to the target webpage, and the output data is the target webpage type identification result corresponding to the target webpage. The second prediction branch performs target brand identification on its input data. Target brand identification is used to determine whether the webpage is related to the target brand. The input data of the second prediction branch includes the webpage features corresponding to the target webpage, and the output data is the target brand identification result corresponding to the target webpage.

[0082] Specifically, the computer device may input webpage text corresponding to the target brand and target webpage into the target webpage recognition model. After data processing by the model, the target webpage recognition model outputs a target webpage type recognition result and a target brand recognition result.

[0083] The web page text corresponding to the target brand and the target web page is input into the feature extraction branch of the target web page recognition model, the feature extraction branch performs feature extraction processing on the web page text corresponding to the target brand and the target web page, the feature extraction branch outputs the web page features corresponding to the target web page, the web page features corresponding to the target web page are input into the first prediction branch of the target web page recognition model for target web page type recognition processing, the first prediction branch outputs the target web page type recognition result, the web page features corresponding to the target web page are input into the second prediction branch of the target web page recognition model for target brand recognition processing, and the second prediction branch outputs the target brand recognition result.

[0084] In the above embodiment, the feature extraction branch of the target web page recognition model performs feature extraction on the web page text and the target brand to obtain web page features corresponding to the target web page. The first prediction branch of the target web page recognition model performs target web page type recognition on the web page features to obtain a target web page type recognition result. The second prediction branch of the target web page recognition model performs target brand recognition on the web page features to obtain a target brand recognition result. Performing different data processing using different prediction branches can ensure the accuracy and efficiency of data processing.

[0085] In one embodiment, reference Figure 3 The web page recognition model includes a BERT model (i.e., a feature extraction branch), a linear classifier 1 (i.e., a first prediction branch), and a linear classifier 2 (i.e., a second prediction branch). The web page URL, web page title, web page text, and target brand corresponding to the target web page are input into the web page recognition model. The features are extracted through the BERT model to obtain the web page features corresponding to the target web page. The web page features are input into linear classifier 1. Linear classifier 1 determines whether the target web page is of the target web page type. Linear classifier 1 outputs the target web page type recognition result. The web page features are input into linear classifier 2. Linear classifier 2 determines whether the target web page is a web page related to the target brand. Linear classifier 2 outputs the target brand recognition result. The linear classifier can be set as needed. For example, the linear classifier is a binary classification Softmax function.

[0086] It's understandable that URLs for specific types of web pages have distinct characteristics and formats. The web page title and body text represent the meaning and expression of the web page, and the target brand is one of the elements that require verification. By separately encoding and merging these three types of data, a highly effective verification model (i.e., a web page recognition model) can be obtained. The CLS and SEP tags in the model input data are special markers. The CLS tag is typically placed at the beginning of the input sequence. The output vector corresponding to the CLS tag (i.e., the output of the first hidden layer of the BERT model) represents the global information of the entire input sequence. This vector can be used for subsequent classification tasks. In these tasks, the CLS tag output vector captures the semantic information of the entire sentence, not just local information. The SEP tag is used to separate different parts of the input sequence. The SEP tag helps the BERT model understand sentence boundaries and the structure of the input data.

[0087] In one embodiment, the webpage processing method further includes:

[0088] Obtain a training sample set; the training samples in the training sample set are web page texts and training brands corresponding to web pages of known web page types and associated brands; input the first training sample in the training sample set into the initial web page recognition model to obtain the web page type prediction label corresponding to the first training sample, calculate the first model loss based on the web page type training label and the web page type prediction label corresponding to the first training sample, adjust the model parameters of the feature extraction branch and the first prediction branch in the initial web page recognition model based on the first model loss, and obtain a first intermediate web page recognition model; input the second training sample in the training sample set into the first intermediate web page recognition model to obtain the brand prediction label corresponding to the second training sample, calculate the second model loss based on the brand training label and the brand prediction label corresponding to the second training sample, adjust the model parameters of the feature extraction branch and the second prediction branch in the first intermediate web page recognition model based on the second model loss, and obtain a second intermediate web page recognition model; input the third training sample in the training sample set into the second intermediate web page recognition model to obtain the comprehensive prediction label corresponding to the third training sample, calculate the third model loss based on the comprehensive prediction label and the comprehensive training label corresponding to the third training sample, adjust the model parameters of the second intermediate web page recognition model based on the third model loss, and obtain a target web page recognition model.

[0089] The training sample set includes multiple training samples. The training samples are the web page text and training brands corresponding to the training web page. That is, the training web page is a web page with a known web page type and associated brand. The training brand corresponding to the training web page can be the brand associated with the training web page, or it can be a brand not associated with the training web page. The training samples have corresponding comprehensive training labels. The comprehensive training labels are the correct target web page type recognition results and brand recognition results corresponding to the training web page. That is, the comprehensive training labels include web page type training labels and brand prediction labels. It can be understood that the web page type training label is used to indicate whether the training web page belongs to the target web page type. The brand prediction label is used to indicate whether the training web page and the training brand are related or associated.

[0090] The training sample is input into the web page recognition model. The first prediction branch in the web page recognition model outputs a web page type prediction label, and the second prediction branch outputs a brand prediction label. The comprehensive prediction label includes the web page type prediction label and the brand prediction label. The web page type prediction label is the target web page type recognition result predicted by the model for the training web page. The brand prediction label is the brand recognition result predicted by the model for the training web page.

[0091] The initial web page recognition model refers to the web page recognition model to be trained. The initial web page recognition model is subjected to supervised training based on the first training sample in the training sample set to obtain a first intermediate web page recognition model. The first intermediate web page recognition model is subjected to supervised training based on the second training sample in the training sample set to obtain a second intermediate web page recognition model. The second intermediate web page recognition model is subjected to supervised training based on the third training sample in the training sample set to obtain a target web page recognition model.

[0092] It is understood that there may be multiple first training samples, multiple second training samples, and multiple third training samples. The multiple first training samples and the multiple second training samples may have overlapping training samples or different training samples. The multiple first training samples and the multiple third training samples may have overlapping training samples or different training samples. The multiple second training samples and the multiple third training samples may have overlapping training samples or different training samples. The training brands corresponding to different training samples may be the same or different.

[0093] Specifically, the computer device obtains a training sample set and performs a two-stage warm-up training on the initial webpage recognition model based on the training sample set to obtain the target webpage recognition model. The purpose of the two-stage warm-up training is to enable the model to learn in a way that mimics human thinking. Specifically, the model is expected to first learn to identify the target webpage type and then learn to determine whether the webpage is related to the input brand.

[0094] In the first stage, the prediction capabilities of the first prediction branch and the second prediction branch are trained separately. First, the prediction capability of the first prediction branch is trained. The first training sample in the training sample set is input into the initial web page recognition model. After data processing by the model, the first prediction branch of the model outputs the web page type prediction label corresponding to the first training sample. The web page type training label corresponding to the first training sample is obtained. The first model loss is calculated based on the web page type training label and the web page type prediction label corresponding to the first training sample. The first model loss is used to reflect the difference between the web page type training label and the web page type prediction label corresponding to the first training sample. The smaller the model loss, the smaller the difference. The first model loss is back-propagated to adjust the model parameters of the feature extraction branch and the first prediction branch in the initial web page recognition model. Through multiple model iterative training until the first end condition is met, the first intermediate web page recognition model is obtained. Then, the prediction capability of the second prediction branch is trained. The second training sample in the training sample set is input into the first intermediate web page recognition model. After data processing by the model, the model outputs the brand prediction label corresponding to the second training sample, obtains the brand training label corresponding to the second training sample, and calculates the second model loss based on the brand training label and the brand prediction label corresponding to the second training sample. The second model loss is used to reflect the difference between the brand training label and the brand prediction label corresponding to the second training sample. The smaller the model loss, the smaller the difference. The second model loss is back-propagated to adjust the model parameters of the feature extraction branch and the second prediction branch in the first intermediate web page recognition model. Through multiple model iterative training until the second end condition is met, the second intermediate web page recognition model is obtained.

[0095] In the second stage, the prediction capabilities of the first prediction branch and the second prediction branch are trained simultaneously. The third training sample in the training sample set is input into the second intermediate web page recognition model to obtain a comprehensive prediction label corresponding to the third training sample. A third model loss is calculated based on the comprehensive prediction label and comprehensive training label corresponding to the third training sample. The third model loss is used to reflect the difference between the comprehensive training label and the comprehensive prediction label corresponding to the third training sample. The smaller the model loss, the smaller the difference. The model parameters of the second intermediate web page recognition model are adjusted based on the third model loss. The model is trained through multiple iterations until the third end condition is met, thereby obtaining the target web page recognition model.

[0096] It is understood that the first, second, and third termination conditions may be the same or different. The termination conditions include, but are not limited to, at least one of: a model loss less than a preset loss value; a model iteration count greater than a preset iteration count; and a rate of change of model loss less than a preset rate of change.

[0097] In one embodiment, the prediction results of the model can be input into the loss function to calculate the model loss. The loss function of the third model loss is as follows:

[0098]

[0099] in, is the model input data The prediction results, represents the model parameters from the feature extraction branch to the first prediction branch, Represents the model parameters from the feature extraction branch to the second prediction branch. Representatives will Input the model, the data output by the first prediction branch, that is, the target web page type recognition result in the prediction label. Represents the value of the j-th dimension in the "Is it the target webpage type" label in the prediction label. It is the value of the j-th dimension in the "Is it the target webpage type" label in the one-hot format, that is, the value of the j-th dimension in the "Is it the target webpage type" label in the training label. is the value of the jth dimension of the "Is this a training brand-related webpage?" label in one-hot format. This is the value of the jth dimension of the "Is this a training brand-related webpage?" label in the training label. i represents the data subscript, m represents the number of samples in the current batch, j represents the category subscript, and c represents the number of categories. Here, both tasks are binary classification tasks, so c = 2. and It can be set according to actual needs. The goal of the training process is to reduce the loss function.

[0100] Referring to the loss function of the third model loss, the loss function of the first model loss is as follows:

[0101]

[0102] Referring to the loss function of the third model loss, the loss function of the second model loss is as follows:

[0103]

[0104] It can be understood that for different model losses, m can be the same or different.

[0105] In the above embodiment, the model parameters of the feature extraction branch and the first prediction branch in the web page recognition model are first adjusted based on the web page type training label and web page type prediction label corresponding to the first training sample. Then, the model parameters of the feature extraction branch and the second prediction branch in the web page recognition model are adjusted based on the brand training label and brand prediction label corresponding to the second training sample. Finally, the complete model parameters of the web page recognition model are adjusted based on the comprehensive prediction label and comprehensive training label corresponding to the third training sample. In this way, during the training process, the model can first learn to recognize what the target web page type is, and then learn to determine whether the web page is related to the input brand, thereby orderly improving the prediction ability of each prediction branch, and ultimately comprehensively ensuring the prediction ability of each prediction branch and ensuring the prediction effect of the model.

[0106] In one embodiment, obtaining a training sample set includes:

[0107] Acquire multiple unlabeled initial samples; perform deduplication processing on the multiple initial samples based on sample similarity between the initial samples to obtain multiple training samples; perform web page type labeling and brand labeling on the training samples to obtain comprehensive training labels corresponding to the multiple training samples; obtain a training sample set based on the multiple training samples and the comprehensive training labels corresponding to the multiple training samples.

[0108] The initial samples are samples without training labels. In other words, the initial samples are webpage text and training brands corresponding to webpages that are not labeled as "whether they are the target webpage type" or "whether they are associated with the training brand."

[0109] The sample similarity between two initial samples is based on the text similarity between the webpage texts in the two initial samples. Various similarity calculation algorithms can be used to calculate the text similarity between two webpage texts. For example, the edit distance between the two webpage texts can be calculated as text similarity; text features corresponding to the webpage texts can be extracted and the cosine distance or Euclidean distance between the two text features can be calculated as text similarity; and so on.

[0110] Web page type labeling refers to labeling whether the training web page corresponding to the training sample is of the target web page type. Brand labeling refers to labeling whether the training web page corresponding to the training sample is a web page related to the training brand.

[0111] Specifically, to improve the quality of the training sample set, a rich and diverse set of training samples can be obtained to form the training sample set. A computer device obtains multiple unlabeled initial samples and, based on the sample similarity between the initial samples, performs deduplication on the multiple initial samples. Initial samples with higher sample similarity are deduplicated, thereby obtaining a rich and diverse set of training samples. Furthermore, the training samples are annotated with web page type and brand to obtain comprehensive training labels corresponding to each training sample. The comprehensive training labels include web page type training labels and brand training labels. Finally, the multiple training samples and the comprehensive training labels corresponding to the multiple training samples are combined to form a training sample set.

[0112] In the above embodiment, multiple initial samples are first deduplicated based on sample similarity between them to obtain multiple training samples. These training samples are then labeled with web page type and brand to obtain comprehensive training labels corresponding to each of the multiple training samples. The training samples and their corresponding comprehensive training labels are then combined to form a training sample set. This ensures that the training sample set is small and precise, preventing repeated learning and misleading the model. This allows the model to quickly learn relevant knowledge for the target web page type and brand recognition tasks from this small and precise training set.

[0113] In one embodiment, the initial web page recognition model includes a BERT model, a first prediction branch, and a second prediction branch. The BERT model is pre-trained. It has a certain degree of common sense and has accumulated a large amount of contextual knowledge during its pre-training process. The pre-training task can be to predict the following text based on the previous text. For the review task, in order for the model to fully learn the particularity of the review task and the connection between it and the conventional knowledge learned in the past, an accurate and data-rich training set is essential. First, a large amount of web page text of unlabeled web pages is collected, and then based on the edit distance between the web page texts, web pages with high similarity are deduplicated to obtain a rich and diverse basic samples. The basic samples are manually labeled, and the labeling information includes whether it is a web page related to the training brand and whether it is the target web page type. Here, there are only two answers: yes or no. After manual labeling, a manual secondary review can also be performed to ensure the accuracy of the information. Accuracy means that the labeled information is accurate. The accurately labeled basic samples and the corresponding labeling labels form the training set of the web page recognition model.

[0114] In one embodiment, the web page meta information includes the website registration information corresponding to the target web page. Figure 4 As shown, based on the webpage meta information corresponding to the target webpage, the target webpage is subjected to counterfeit identification processing to obtain the webpage counterfeit identification result corresponding to the target webpage, including:

[0115] Step S402: performing a correlation test on the website filing party and the target brand in the website filing information to obtain the authenticity of the first webpage corresponding to the target webpage.

[0116] Step S404: Match the website filing party in the website filing information with a preset filing party set to obtain the second webpage authenticity corresponding to the target webpage.

[0117] Step S406: Determine a phishing identification result corresponding to the target webpage based on at least one of the first webpage authenticity and the second webpage authenticity.

[0118] Website filing information refers to information obtained through registration with relevant authorities. For example, this information can be ICP filing information or WHOIS filing information. This information records basic website information, such as the website's owner, domain name, website address, and creation date. Website filing information includes the website's filing entity. The website's filing entity refers to the website's owner, which is the organization, enterprise, department, or individual to which the website belongs.

[0119] The preset registrant set includes multiple pre-set, trusted, and legitimate website registrants. In other words, the preset registrant set acts as a whitelist. If the target webpage's website registrant is listed in the preset registrant set, it indirectly indicates that the target webpage's website registrant is trusted and legitimate, and the probability that the target webpage is not a counterfeit webpage is higher.

[0120] Association testing involves examining the degree of association and relevance between the website registrant and the target brand. For example, the textual similarity between the website registrant and the target brand can be calculated as the association. A correlation between the website registrant and the target brand indicates that the target webpage is legitimately registered, increasing the probability that the target webpage is genuine.

[0121] The authenticity of the target webpage can be determined based on the website registration information corresponding to the target webpage. The authenticity of the webpage is used to reflect the authenticity of the webpage. The higher the authenticity of the webpage, the greater the probability that the webpage is not a counterfeit webpage.

[0122] Specifically, after determining that the target webpage belongs to the target webpage type and is related to the target brand, first determine whether the target webpage is a non-counterfeit webpage based on the intelligent analysis strategy of non-counterfeit webpages. Specifically, determine whether the target webpage is a non-counterfeit webpage based on the website registration information of the target webpage.

[0123] The computer device obtains the website filing information corresponding to the target webpage, and can perform a correlation test between the website filing entity and the target brand in the website filing information to obtain the authenticity of the first webpage corresponding to the target webpage. For example, the higher the correlation obtained through the correlation test, the higher the authenticity of the first webpage. A phishing identification result corresponding to the target webpage can be determined based on the authenticity of the first webpage. For example, if the authenticity of the first webpage is greater than a webpage authenticity threshold, the phishing identification result is determined to be a non-phishing webpage.

[0124] The computer device may also match the website recorder in the website record information with a preset recorder set to obtain a second webpage authenticity corresponding to the target webpage. For example, if the match is successful, the second webpage authenticity is the first preset authenticity; if the match fails, the second webpage authenticity is the second preset authenticity, and the first preset authenticity is greater than the second preset authenticity. Based on the second webpage authenticity, a phishing identification result corresponding to the target webpage may be determined. For example, if the second webpage authenticity is greater than a webpage authenticity threshold, the phishing identification result indicates that the webpage is not phishing.

[0125] Of course, the webpage phishing identification result corresponding to the target webpage may also be determined based on the authenticity of the first webpage and the authenticity of the second webpage.

[0126] In one embodiment, the webpage meta information corresponding to the target webpage is obtained based on the URL corresponding to the target webpage.

[0127] In the above embodiment, the website registrant in the website filing information and the target brand are tested for correlation to obtain a first webpage authenticity corresponding to the target webpage. The website registrant in the website filing information is matched with a preset set of registrants to obtain a second webpage authenticity corresponding to the target webpage. Based on at least one of the first webpage authenticity and the second webpage authenticity, a phishing identification result corresponding to the target webpage is determined. In this way, the correlation between the website registrant and the target brand of the target webpage can reflect the legitimacy of the target webpage, and the matching between the website registrant and the preset set of registrants can reflect the legitimacy of the target webpage. Determining the phishing identification result for the target webpage based on at least one of the correlation and matching can ensure the accuracy of the phishing identification result.

[0128] In one embodiment, a correlation test is performed between the website filing party and the target brand in the website filing information to obtain the authenticity of the first webpage corresponding to the target webpage, including:

[0129] The website filing party and the target brand in the website filing information are pre-processed; when the similarity between the website filing party and the target brand after pre-processing is greater than the similarity threshold, the authenticity of the first web page corresponding to the target web page is determined to be a first preset authenticity; when the similarity between the website filing party and the target brand after pre-processing is less than or equal to the similarity threshold, the authenticity of the first web page corresponding to the target web page is determined to be a second preset authenticity; the first preset authenticity is greater than the second preset authenticity.

[0130] Among them, preprocessing refers to extracting keywords. For example, preprocessing can be to remove redundant and invalid words in the website registration party and target brand.

[0131] Specifically, when performing criticality detection, the website recorder and target brand in the website record information can be pre-processed separately, and then the similarity between the pre-processed website recorder and the target brand can be calculated, and the calculated similarity can be compared with the similarity threshold. If the calculated similarity is greater than the similarity threshold, it means that the website recorder and the target brand are highly similar and are highly likely to be related. Therefore, the first web page authenticity corresponding to the target web page is determined to be a first preset authenticity with a relatively large value. If the calculated similarity is less than or equal to the similarity threshold, the first web page authenticity corresponding to the target web page is determined to be a second preset authenticity with a relatively small value.

[0132] It can be understood that the similarity threshold can be set as needed.

[0133] In the above embodiment, when the similarity between the website filing entity and the target brand after pre-processing is greater than the similarity threshold, the authenticity of the first webpage corresponding to the target webpage is determined to be the first preset authenticity. When the similarity between the website filing entity and the target brand after pre-processing is less than or equal to the similarity threshold, the authenticity of the first webpage corresponding to the target webpage is determined to be the second preset authenticity, and the first preset authenticity is greater than the second preset authenticity. In this way, by comparing the similarity between the website filing entity and the target brand with the similarity threshold, the authenticity of the first webpage corresponding to the target webpage can be quickly determined.

[0134] In one embodiment, determining a phishing identification result corresponding to a target webpage based on at least one of the first webpage authenticity and the second webpage authenticity includes:

[0135] At least one of the first web page authenticity and the second web page authenticity is integrated to obtain a comprehensive web page authenticity corresponding to the target web page; when the comprehensive web page authenticity is greater than or equal to a web page authenticity threshold, the web page phishing identification result corresponding to the target web page is determined to be a non-phishing web page.

[0136] Specifically, the computer device may combine at least one of the first web page authenticity and the second web page authenticity to obtain a comprehensive web page authenticity corresponding to the target web page, compare the comprehensive web page authenticity with a web page authenticity threshold, and determine that the web page phishing identification result corresponding to the target web page is a non-phishing web page if the comprehensive web page authenticity is greater than or equal to the web page authenticity threshold. In other words, if the comprehensive web page authenticity is greater than or equal to the web page authenticity threshold, the target web page is determined to be a genuine web page.

[0137] Furthermore, if the comprehensive web page authenticity is less than the web page authenticity threshold, the phishing identification result for the target web page can be determined to be a phishing web page, or the phishing identification result for the target web page can be determined to be pending. If the phishing identification result for the target web page is pending, it means that it is impossible to accurately determine whether the target web page is a phishing web page, and other measures can be used to further identify phishing.

[0138] It is understandable that the web page authenticity threshold can be set as needed.

[0139] In the above embodiment, at least one of the first web page authenticity and the second web page authenticity is integrated to obtain the comprehensive web page authenticity corresponding to the target web page, and the comprehensive web page authenticity is compared with the web page authenticity threshold to quickly determine whether the target web page is a genuine web page.

[0140] In one embodiment, the webpage meta information includes at least one of the website registration information, network protocol address information, and website provider registration information corresponding to the target webpage. Figure 5 As shown, the web page processing method also includes:

[0141] Step S502: When the phishing identification result corresponding to the target webpage is other than a non-phishing webpage, a correlation test is performed between the website filing party in the website filing information and the target brand to obtain a first phishing degree corresponding to the target webpage.

[0142] Step S504: Perform reliability detection based on the filing time information in the website filing information to obtain a second webpage counterfeiting degree corresponding to the target webpage.

[0143] Website filing information includes the website filing entity and filing date. Filing date information refers to time-related information recorded during website filing registration. For example, filing date information could be the website creation date or the website expiration date. Reliability testing involves testing the reliability of a webpage based on its meta-information.

[0144] The webpage phishing degree of the target webpage can be determined based on the webpage meta information corresponding to the target webpage. The webpage phishing degree is used to reflect the degree of phishing of the webpage. The higher the webpage phishing degree, the greater the probability that the webpage is a phishing webpage.

[0145] Specifically, when the target web page is not identified as a non-counterfeit web page, it is possible to further determine whether the target web page is a counterfeit web page based on the intelligent analysis strategy for counterfeit web pages. Specifically, it is possible to determine whether the target web page is a counterfeit web page based on at least one of the website registration information, network protocol address information, and website provider registration information of the target web page.

[0146] Regarding the website filing information, a correlation test can be performed between the website filing entity and the target brand in the website filing information to determine the first webpage phishing degree corresponding to the target webpage. For example, the lower the correlation degree obtained through the correlation test, the higher the first webpage phishing degree. If the similarity between the website filing entity and the target brand is greater than a similarity threshold, the first webpage phishing degree is determined to be a first preset phishing degree. If the similarity between the website filing entity and the target brand is less than or equal to the similarity threshold, the first webpage phishing degree is determined to be a second preset phishing degree, with the first preset phishing degree being less than the second preset phishing degree. Furthermore, the second webpage phishing degree corresponding to the target webpage can also be determined based on the filing time information in the website filing information. For example, the smaller the time interval between the filing time and the current time in the filing time information, the newer the target webpage. However, phishing webpages generally have a short lifespan and are typically recently created. Therefore, the smaller the time interval between the filing time and the current time in the filing time information, the less reliable the webpage is, and the higher the second webpage phishing degree can be. For another example, if the current time does not exceed the expiration time in the filing time information, the second webpage counterfeiting degree is determined to be the third preset counterfeiting degree; if the current time has exceeded the expiration time in the filing time information, it indicates that the webpage may be unreliable, and the second webpage counterfeiting degree is determined to be the fourth preset counterfeiting degree, and the third preset counterfeiting degree is less than the fourth preset counterfeiting degree.

[0147] Step S506: Perform reliability detection based on the network protocol address information to obtain a third webpage phishing degree corresponding to the target webpage.

[0148] Step S508: Compare the current registration status in the website provider registration information with the preset registration status to obtain a fourth webpage phishing degree corresponding to the target webpage.

[0149] Internet Protocol address information refers to the website's Internet Protocol address (IP) information. This information is used to record information related to the website's IP address, such as the IP address, its location, and the time of the most recent IP address change.

[0150] Website provider registration information refers to the website provider's business registration information, namely the business registration information of the company or individual providing the website, as well as the business registration information of the website's owner. Website provider registration information includes the website provider's registration status, which refers to the website provider's operating status.

[0151] In one embodiment, web page meta information includes ICP filing information, WHOIS filing information, and IP address information. Referring to Figure 6, ICP filing information includes website domain name, unit name of the unit to which the website belongs, unit nature, website name, website filing number, review time, and website address. Referring to Figure 7, WHOIS filing information includes website domain name, website creation time, registration time, update time, expiration time, registrant, registered email (mailbox), and registered email malicious intensity. Figure 8 ,IP address information includes each IP address corresponding to the website, the geographical location, status, DNS (Domain Name System) resolution time, IP malicious level, IP attributes, malicious status, and details.

[0152] The current registration status is the registration status recorded in the website provider's registration information. The default registration status is the default registration status. The default registration status can be set as needed, for example, the default registration status is "Existing".

[0153] Specifically, with respect to the network protocol address information, a third webpage phishing degree corresponding to the target webpage is determined based on the network protocol address information. For example, the network protocol address information includes the time of the most recent IP address switch. Phishing webpages generally switch IP addresses frequently. Therefore, the smaller the time interval between the most recent IP address switch and the current time, the less reliable the webpage and the greater the third webpage phishing degree. For another example, the network protocol address information includes the location of the IP address. If the IP address is located within the country, the third webpage phishing degree is determined to be the fifth preset phishing degree. If the IP address is located outside the country, indicating that the webpage may be unreliable, the third webpage phishing degree is determined to be the sixth preset phishing degree, with the fifth preset phishing degree being less than the sixth preset phishing degree.

[0154] For the website provider registration information, the current registration status in the website provider registration information is compared with a preset registration status to determine a fourth webpage phishing degree corresponding to the target webpage. For example, registration statuses include cancellation, revocation, and continued, and the preset registration status is continued. If the current registration status is continued, the fourth webpage phishing degree is determined to be a seventh preset value phishing degree. If the registration status is cancellation or revocation, the fourth webpage phishing degree is determined to be an eighth preset phishing degree, and the seventh preset phishing degree is less than the eighth preset phishing degree.

[0155] Step S510: determining a phishing identification result corresponding to the target webpage based on at least two of the first phishing degree, the second phishing degree, the third phishing degree, and the fourth phishing degree.

[0156] Specifically, the webpage meta-information includes at least two of the website filing information, network protocol address information, and website provider registration information corresponding to the target webpage. The computer device can determine the webpage phishing identification result corresponding to the target webpage based on at least one of the first webpage phishing degree, the second webpage phishing degree, the third webpage phishing degree, and the fourth webpage phishing degree.

[0157] In the above embodiment, the corresponding phishing degrees are determined according to different types of web page meta information, and the phishing identification result corresponding to the target web page is determined based on the various phishing degrees, thereby ensuring the accuracy of phishing identification.

[0158] In one embodiment, determining a phishing identification result corresponding to a target webpage based on at least two of the first phishing degree, the second phishing degree, the third phishing degree, and the fourth phishing degree includes:

[0159] At least two of the first phishing degree, the second phishing degree, the third phishing degree, and the fourth phishing degree are integrated to obtain a comprehensive phishing degree corresponding to the target webpage; when the comprehensive phishing degree is greater than a phishing degree threshold, the phishing identification result corresponding to the target webpage is determined to be a phishing webpage.

[0160] Specifically, at least two of the first, second, third, and fourth phishing degrees are combined to obtain a comprehensive phishing degree. For example, the comprehensive phishing degree can be obtained by summing the various phishing degrees; or the comprehensive phishing degree can be obtained by weighted summing the various phishing degrees. The weights corresponding to the various phishing degrees can be set according to actual needs. The greater the comprehensive phishing degree, the greater the probability that the webpage is a phishing webpage. Therefore, if the comprehensive phishing degree is greater than or equal to a phishing degree threshold, the computer device can determine that the phishing identification result corresponding to the target webpage is a phishing webpage.

[0161] Furthermore, if the comprehensive phishing degree is less than the phishing degree threshold, the phishing identification result for the target webpage can be determined to be non-phishing, or the phishing identification result for the target webpage can be determined to be pending. If the phishing identification result for the target webpage is pending, it indicates that it is impossible to accurately determine whether the target webpage is a phishing webpage, and other measures can be used for phishing identification, such as manual verification.

[0162] It is understandable that the phishing degree threshold can be set according to actual needs.

[0163] In the above embodiments, at least one of the first web page forgery degree, the second web page forgery degree, the third web page forgery degree, and the fourth web page forgery degree is integrated to obtain a comprehensive web page forgery degree. By comparing the comprehensive web page forgery degree with the web page forgery degree threshold, it is possible to quickly determine whether the target web page is a forged web page.

[0164] In one embodiment, the web page content and the website address can be forged, but the meta-information of the web page, such as ICP filing, WHOIS filing, IP address, etc., is difficult to forge. The meta-information of the target web page as shown in Table 1 can be obtained through relevant network interfaces. Further, referring to Table 2, the web page meta-information of the target web page is parsed and feature-transformed to obtain 10 features. Further, first, based on the intelligent judgment strategy of the genuine web page, it is judged whether the target web page is a genuine web page. When the target web page cannot be identified as a genuine web page, then based on the intelligent judgment strategy of the forged web page, it is judged whether the target web page is a forged web page.

[0165] For the intelligent judgment strategy of the genuine web page, starting from a base score of 0 points, that is, true_score = 0, it is increased or decreased based on the following feature analysis situations. Feature 1 + Feature 2: Whether the ICP filing is associated with the target brand, and whether the WHOIS registrant is associated with the target brand. If one of them is associated, then true_score = true_score + 80; if neither is, then true_score remains unchanged. Here, the association includes two situations. One is that the ICP filing / WHOIS registrant is exactly the same as the target brand, and the other is that they are not exactly the same but the similarity between them is extremely high. Here, the similarity is measured by the edited distance after preprocessing. The preprocessing means removing common words such as "Co., Ltd." and "Branch" from the ICP filing / WHOIS registrant and the target brand, and then judging the edited distance between the two. If the edited distance / the maximum text length between the two < 0.2, it can be regarded as a high similarity, and the two are very likely to be associated. The threshold for the genuine web page is set to 80. If true_score ≥ 80, it is regarded as a genuine web page.

[0166] The intelligent analysis strategy for fake webpages starts with a base score of 0 (fake_score = 0), and increases or decreases based on the following feature analysis. Feature 1 + Feature 2: Whether the ICP registration is associated with the target brand, and whether the WHOIS registrant is associated with the target brand. If either is true, fake_score = fake_score + 60; if neither is true, fake_score remains unchanged. Feature 3: The number of days since WHOIS creation. If the number of days is less than 30, fake_score = fake_score + 10; if 30 < number of days < 180, fake_score = fake_score + 5. Feature 4: Whether the WHOIS is expired. If so, fake_score = fake_score + 10; otherwise, no change. Feature 5: Whether the filing entity is a business. If not, fake_score = fake_score + 20; otherwise, no change. Feature 6: Is the IP address located within China? If so, fake_score remains unchanged; otherwise, fake_score = fake_score + 10. Feature 7: The number of days since the last IP address switch. If the number of days is less than 30, fake_score = fake_score + 10; if 30 < number of days < 180, fake_score = fake_score + 5. Feature 8: Business status. If the business status is revoked or cancelled, fake_score = fake_score + 10. The threshold for fake webpages is set to 80; webpages with a fake_score ≥ 80 are considered fake.

[0167] Table 1

[0168]

[0169] Table 2

[0170]

[0171] In one embodiment, determining a brand counterfeiting identification result corresponding to a target brand based on the webpage counterfeiting identification result includes:

[0172] Based on the webpage phishing identification results of each of the multiple target webpages corresponding to the target brand, the number of non-phishing webpages and the number of phishing webpages corresponding to the target brand are determined; based on the number of non-phishing webpages and the number of phishing webpages, the brand phishing degree corresponding to the target brand is determined; and the brand phishing degree is used as the brand phishing identification result corresponding to the target brand.

[0173] Specifically, the computer device can identify multiple target web pages related to the target brand and belonging to the target web page type, and determine a phishing identification result corresponding to each target web page. Furthermore, based on the phishing identification results for each of the multiple target web pages corresponding to the target brand, the number of non-phishing web pages and the number of phishing web pages corresponding to the target brand can be determined. Based on the number of non-phishing web pages and the number of phishing web pages, the brand phishing degree corresponding to the target brand can be determined. For example, brand phishing degree = number of non-phishing web pages / (number of non-phishing web pages + number of phishing web pages). Ultimately, the brand phishing degree corresponding to the target brand is used as the brand phishing identification result corresponding to the target brand.

[0174] In the above embodiment, the brand counterfeiting identification result corresponding to the target brand is determined by integrating the webpage counterfeiting identification results of the plurality of target webpages corresponding to the target brand, which can improve the accuracy of brand counterfeiting identification.

[0175] In a specific embodiment, the method of the present application can be applied to the counterfeit identification scenario for the scan code verification web page. In the prior art, the online and offline counterfeit clues are deeply hidden, and the goods sales chain is extremely complex, which makes the brand counterfeit rate difficult to calculate or even impossible to calculate. The present application proposes a brand counterfeit rate calculation method based on the Internet scan code verification web page. It cuts in from the perspective of daily user operations such as scan code verification, digs out clues of true and false goods from the side, and automatically identifies and mines and judges in a wide range of Internet web pages. It can output the target brand counterfeit rate reflected by the scan code verification operation on the entire Internet. It has the characteristics of innovation, high precision, low cost, high detection, and support for unlimited target brand configurations. The detected counterfeit rate can be provided to the corresponding business partners for business behavior evaluation and rights protection operations on the one hand, and can also be provided to relevant departments as market environment assessment data on the other hand.

[0176] refer to Figure 9 The method of this application includes a target customer code scanning and verification web page identification module based on a policy model and a deep learning model, a genuine and fake product code scanning and verification web page identification module based on web page meta-information, and a counterfeit rate calculation module. The input of the calculation system is an unlimited range of Internet web pages and their web page texts, as well as the brand name (i.e., target customer, target brand) for which the counterfeit rate needs to be calculated. In actual applications, the brand name can be flexibly configured, such as adding, deleting, and modifying the brand name, and the calculation system runs periodically, with each input containing hundreds of millions of web pages. The output of the calculation system is the brand name for which the counterfeit rate needs to be calculated and its counterfeit rate.

[0177] The target customer QR code verification webpage identification module, based on a policy model, involves strategies for filtering target brand-related webpages and scanning QR code verification webpage screening. The target customer QR code verification webpage identification module, based on a deep learning model, involves customized model structure design, training set construction, and two-stage warm-up training. The policy model is used to initially screen target brand-related scanning QR code verification webpages, and then the deep learning model is used to verify the accuracy of the initial screening results.

[0178] For the target brand-related webpage screening strategy, a large number of webpages are screened based on the target brand's extended keywords to obtain target brand-related webpages. For the QR code verification webpage screening strategy, based on the QR code verification webpage keywords, a quick screening of target brand-related webpages is performed to select those that are subject to QR code verification.

[0179] For customized model design, a BERT model + two linear classifiers was used. The model's input data consists of the webpage URL, title, body text, and target brand. The model's output data indicates whether the webpage is a scan-to-verify page or a page related to the target brand. BERT possesses a certain degree of common sense and has accumulated a vast amount of contextual information during pre-training. To fully understand the specificity of the verification task and its connection to previously learned general knowledge, an accurate and data-rich training set is essential. Therefore, to construct the training set, a large number of web pages were collected, and based on edit distance, duplicates were removed from pages with high similarity to obtain the training set. The training set was then manually annotated, with information including whether the page is related to the training brand and whether it is a scan-to-verify page. Here, there are only two possible answers: yes or no.

[0180] For the two-stage warm-up training, the model is trained on two separate tasks. First, the model is trained to identify webpages for QR code verification based on the scan code verification information. Then, the model is trained to identify brand-related webpages based on the brand-related information. The two tasks are then combined. During training, the Adam optimization method is used, with a learning rate of 0.001 and linear decay. The goal of the two-stage warm-up training is to enable the model to learn in a way that mimics human thinking. The goal is to first teach the model to identify QR code verification webpages and then to determine whether a webpage is brand-related.

[0181] Based on actual business operations, a strategy model centered around target brands and webpage attributes was constructed to detect suspected target brand scan code verification webpages from the complex internet. This model boasts exceptional speed and the ability to filter out a large number of irrelevant webpages. Next, using the results of the first phase of review as a learning objective, a deep learning model was trained to enable it to identify webpage types and conduct secondary brand confirmation. This two-stage deep learning review model was used to review the first-stage data and filter out webpages that are scan code verification webpages and feature the target brand as a core product. This two-stage target webpage detection solution, formed by combining the above strategy model and deep learning model, boasts overall speed, the ability to process massive amounts of data, and high accuracy.

[0182] The module for identifying genuine and counterfeit products by scanning barcodes based on web page metadata involves parsing and characterizing web page metadata, as well as intelligent strategies for identifying genuine and counterfeit web pages. The module for calculating the counterfeit rate involves calculating the counterfeit rate.

[0183] For web page meta-information parsing and characterization, information such as the web page's ICP filing, WHOIS filing, IP address, and the business registration information of the ICP-filing company is obtained as web page meta-information. This web page meta-information is parsed and converted into features to obtain the web page meta-features. For example, the web page meta-information can be parsed and converted into features, as shown in Table 2. For intelligent identification of genuine web pages, a policy model for genuine web pages is designed based on web page meta-features. The web page authenticity (i.e., true_score) is calculated. If the web page authenticity corresponding to the target brand's scan code verification web page exceeds a certain threshold, the target brand's scan code verification web page is determined to be genuine. For intelligent identification of fake web pages, a policy model for fake web pages is designed based on web page meta-features. The web page counterfeiting score (i.e., fake_score) is calculated. If the web page counterfeiting score corresponding to the target brand's scan code verification web page exceeds a certain threshold, the target brand's scan code verification web page is determined to be counterfeit.

[0184] To calculate the counterfeit rate, we aggregate data by brand, obtaining all brands and the number of fake and real URLs for each brand. For brands with both fake and real URLs, we use the following formula to calculate the counterfeit rate: Brand counterfeit rate = Number of fake URLs for a brand / (Number of fake URLs for a brand + Number of real URLs for a brand).

[0185] Using webpage metadata such as ICP filings, WHOIS records, and IP location information, combined with the target brand's scan-to-verify webpages discovered in the previous steps, a comprehensive assessment is conducted to determine the authenticity of the scan-to-verify webpages. Data from the authentic and counterfeit webpages is then compared to calculate the corresponding counterfeit rate. This method of calculating counterfeit rates is unaffected by additional information such as region, product type, or user group, and offers the advantages of high accuracy and objectivity.

[0186] refer to Figure 10 The target brand and a large number of web pages are obtained using the URLs (i.e., website addresses), titles, and content. Using the first-stage strategy model and the target brand input by the customer, the system quickly identifies and extracts the scan-based authentication web pages of the customer's desired brand from a complex webpage landscape. A second-stage verification model is then used to reconfirm these web pages. The URLs, titles, and content of the target brand's scan-based authentication web pages identified in the first stage, along with the target brand, are input into the verification model to determine if they are the target brand's scan-based authentication web pages. The first-stage strategy model and the second-stage verification model ensure identification accuracy. The target brand's scan-based authentication web pages are then fed into the next stage for authenticity verification. Meta information is obtained for each web page, parsed, and characterized. These meta features are then fed into the strategy model for genuine web pages for scoring. The score is then used to determine if the web page is a genuine scan-based authentication web page. If the web page cannot be confirmed as authentic, the meta features are fed into the strategy model for fake web pages for scoring. The score is then used to determine if the web page is a fake scan-based authentication web page. Finally, based on the above real and fake web pages, the counterfeit rate of the target brand is calculated as the final output result.

[0187] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0188] Based on the same inventive concept, embodiments of the present application also provide a web page processing device for implementing the web page processing method described above. The implementation solution provided by this device is similar to the implementation solution described in the above method. Therefore, the specific limitations of one or more web page processing device embodiments provided below can be found in the above-mentioned limitations of the web page processing method and will not be repeated here.

[0189] In one embodiment, Figure 11 As shown, a web page processing device is provided, including: a brand acquisition module 1102, a web page screening module 1104, a web page content identification module 1106, a web page counterfeiting identification module 1108 and a brand counterfeiting identification module 1110, wherein:

[0190] The brand acquisition module 1102 is used to acquire the target brand.

[0191] The webpage screening module 1104 is configured to determine a target webpage from the candidate webpage set based on brand features corresponding to the target brand and type features corresponding to the target webpage type.

[0192] The webpage content identification module 1106 is configured to call a target webpage identification model and perform target webpage type identification and target brand identification on the target webpage based on the webpage text corresponding to the target webpage.

[0193] The webpage phishing identification module 1108 is configured to, when the identification result indicates that the target webpage belongs to the target webpage type and is associated with the target brand, perform phishing identification processing on the target webpage based on the webpage meta information corresponding to the target webpage, and obtain a webpage phishing identification result corresponding to the target webpage.

[0194] The brand counterfeiting identification module 1110 is configured to determine a brand counterfeiting identification result corresponding to a target brand based on the webpage counterfeiting identification result.

[0195] In one embodiment, the webpage screening module 1104 is further configured to:

[0196] Obtaining a brand name set corresponding to a target brand as a brand feature corresponding to the target brand, and performing keyword matching on candidate web pages in the candidate web page set based on the brand feature to obtain an intermediate web page;

[0197] A type keyword corresponding to the target webpage type is obtained as a type feature corresponding to the target webpage type, and keyword matching is performed on the intermediate webpage based on the type feature to obtain the target webpage.

[0198] In one embodiment, the webpage screening module 1104 is further configured to:

[0199] Obtaining an intermediate webpage that successfully matches the first and second keywords in the type feature and fails to match the third keyword in the type feature as the target webpage;

[0200] Among them, the first category of keywords are nouns associated with the target webpage type, the second category of keywords are verbs associated with the target webpage type, and the third category of keywords are words unrelated to the target webpage type.

[0201] In one embodiment, the webpage content identification module 1106 is further configured to:

[0202] Input the webpage text corresponding to the target brand and the target webpage into the target webpage recognition model;

[0203] Through the feature extraction branch of the target web page recognition model, feature extraction processing is performed on the web page text and the target brand to obtain the web page features corresponding to the target web page;

[0204] Performing target web page type identification processing on web page features through the first prediction branch of the target web page identification model to obtain a target web page type identification result;

[0205] The target brand recognition processing is performed on the web page features through the second prediction branch of the target web page recognition model to obtain the target brand recognition result.

[0206] In one embodiment, the webpage processing apparatus is further configured to:

[0207] Obtaining a training sample set; the training samples in the training sample set are web page texts and training brands corresponding to web pages of known web page types and associated brands;

[0208] Inputting a first training sample from the training sample set into the initial web page recognition model to obtain a web page type prediction label corresponding to the first training sample, calculating a first model loss based on the web page type training label and the web page type prediction label corresponding to the first training sample, and adjusting model parameters of a feature extraction branch and a first prediction branch in the initial web page recognition model based on the first model loss to obtain a first intermediate web page recognition model;

[0209] inputting a second training sample from the training sample set into the first intermediate web page recognition model to obtain a brand prediction label corresponding to the second training sample, calculating a second model loss based on the brand training label and the brand prediction label corresponding to the second training sample, and adjusting model parameters of a feature extraction branch and a second prediction branch in the first intermediate web page recognition model based on the second model loss to obtain a second intermediate web page recognition model;

[0210] The third training sample in the training sample set is input into the second intermediate web page recognition model to obtain the comprehensive prediction label corresponding to the third training sample, the third model loss is calculated based on the comprehensive prediction label and the comprehensive training label corresponding to the third training sample, and the model parameters of the second intermediate web page recognition model are adjusted based on the third model loss to obtain the target web page recognition model.

[0211] In one embodiment, the webpage processing apparatus is further configured to:

[0212] Obtain multiple unlabeled initial samples;

[0213] Based on the sample similarity between the initial samples, multiple initial samples are deduplicated to obtain multiple training samples;

[0214] Perform web page type labeling and brand labeling on the training samples to obtain comprehensive training labels corresponding to multiple training samples;

[0215] A training sample set is obtained based on the multiple training samples and the comprehensive training labels corresponding to the multiple training samples.

[0216] In one embodiment, the webpage meta information includes the website registration information corresponding to the target webpage. The webpage phishing identification module 1108 is also used to:

[0217] Perform a correlation test between the website filing party and the target brand in the website filing information to obtain the authenticity of the first webpage corresponding to the target webpage;

[0218] Matching the website filing party in the website filing information with a preset filing party set to obtain the second webpage authenticity corresponding to the target webpage;

[0219] A phishing identification result corresponding to the target webpage is determined based on at least one of the first webpage authenticity and the second webpage authenticity.

[0220] In one embodiment, the phishing identification module 1108 is further configured to:

[0221] Pre-process the website filing party and target brand in the website filing information;

[0222] When the similarity between the pre-processed website filing party and the target brand is greater than the similarity threshold, determining the first webpage authenticity corresponding to the target webpage as a first preset authenticity;

[0223] When the similarity between the pre-processed website recorder and the target brand is less than or equal to the similarity threshold, the first webpage authenticity corresponding to the target webpage is determined to be the second preset authenticity; the first preset authenticity is greater than the second preset authenticity.

[0224] In one embodiment, the phishing identification module 1108 is further configured to:

[0225] fusing at least one of the first webpage authenticity and the second webpage authenticity to obtain a comprehensive webpage authenticity corresponding to the target webpage;

[0226] When the comprehensive webpage authenticity is greater than or equal to the webpage authenticity threshold, the webpage phishing identification result corresponding to the target webpage is determined to be a non-phishing webpage.

[0227] In one embodiment, the webpage meta information includes at least one of the website registration information, network protocol address information, and website provider registration information corresponding to the target webpage. The webpage phishing identification module 1108 is further configured to:

[0228] When the phishing identification result corresponding to the target webpage is a result other than a non-phishing webpage, a correlation test is performed between the website filing party in the website filing information and the target brand to obtain a first phishing degree corresponding to the target webpage;

[0229] Performing a reliability test based on the filing time information in the website filing information to obtain the second webpage counterfeiting degree corresponding to the target webpage;

[0230] Perform reliability testing based on network protocol address information to obtain the third webpage counterfeiting degree corresponding to the target webpage;

[0231] Comparing the current registration status in the website provider's registration information with the preset registration status to obtain a fourth webpage counterfeiting degree corresponding to the target webpage;

[0232] A phishing identification result corresponding to the target webpage is determined based on at least two of the first phishing degree, the second phishing degree, the third phishing degree, and the fourth phishing degree.

[0233] In one embodiment, the phishing identification module 1108 is further configured to:

[0234] Combining at least two of the first phishing degree, the second phishing degree, the third phishing degree, and the fourth phishing degree to obtain a comprehensive phishing degree corresponding to the target webpage;

[0235] When the comprehensive phishing degree is greater than the phishing degree threshold, the phishing identification result corresponding to the target webpage is determined to be a phishing webpage.

[0236] In one embodiment, the brand counterfeit identification module 1110 is further configured to:

[0237] determining the number of non-phishing web pages and the number of phishing web pages corresponding to the target brand based on the phishing identification results of each of the plurality of target web pages corresponding to the target brand;

[0238] Determine the brand counterfeiting degree of the target brand based on the number of non-counterfeit web pages and the number of counterfeit web pages;

[0239] The brand counterfeiting degree is used as the brand counterfeiting identification result corresponding to the target brand.

[0240] The webpage processing device described above first preliminarily screens target webpages related to the target brand and belonging to the target webpage type from a candidate webpage set based on brand and type characteristics. It then verifies the target webpages using a target webpage identification model to determine whether they are related to the target brand and belong to the target webpage type. This initial screening and verification process accurately identifies target webpages related to the target brand and belonging to the target webpage type, thereby improving the accuracy of subsequent counterfeit identification. Counterfeit webpages attempt to mimic authentic webpages in their content, but webpage meta-information is generally difficult to counterfeit. Counterfeit identification of target webpages based on their corresponding webpage meta-information can improve the accuracy of counterfeit identification for webpages. Furthermore, based on the webpage counterfeit identification results for the target webpage, a brand counterfeit identification result corresponding to the target brand is determined, further improving the accuracy of counterfeit identification for the target brand.

[0241] Each module in the webpage processing device described above may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in the form of hardware, or may be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.

[0242] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 12 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data such as target web page recognition models and various preset data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a web page processing method is implemented.

[0243] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 13 As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with external terminals via wired or wireless means, and the wireless means can be achieved via Wi-Fi, mobile cellular networks, NFC (near-field communication), or other technologies. When executed by the processor, the computer program implements a webpage processing method. The display unit of the computer device is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.

[0244] Those skilled in the art will understand that Figure 12 、 13 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0245] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.

[0246] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0247] In one embodiment, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.

[0248] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0249] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processing units (GPUs), digital signal processors (DSPs), programmable logic devices (PLDs), data processing logic devices based on quantum computing, and the like.

[0250] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0251] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A web page processing method, characterized in that: The method comprises: Acquire target brands; Determining a target webpage from a set of candidate webpages based on brand features corresponding to the target brand and type features corresponding to the target webpage type; Invoking a target webpage recognition model to perform target webpage type recognition and target brand recognition on the target webpage based on the webpage text corresponding to the target webpage; If the identification result shows that the target webpage belongs to the target webpage type and the target webpage is associated with the target brand, performing counterfeit identification processing on the target webpage based on webpage meta information corresponding to the target webpage to obtain a webpage counterfeit identification result corresponding to the target webpage; Based on the webpage phishing identification result, a brand phishing identification result corresponding to the target brand is determined.

2. The method according to claim 1, characterized in that The determining of the target webpage from the candidate webpage set based on the brand feature corresponding to the target brand and the type feature corresponding to the target webpage type includes: Obtaining a brand name set corresponding to the target brand as a brand feature corresponding to the target brand, and performing keyword matching on candidate web pages in the candidate web page set based on the brand feature to obtain an intermediate web page; A type keyword corresponding to the target webpage type is obtained as a type feature corresponding to the target webpage type, and keyword matching is performed on the intermediate webpage based on the type feature to obtain the target webpage.

3. The method according to claim 2, characterized in that The step of performing keyword matching on the intermediate webpage based on the type feature to obtain the target webpage includes: Obtaining an intermediate webpage that successfully matches the first and second keywords in the type feature and fails to match the third keyword in the type feature as the target webpage; The first category of keywords are nouns associated with the target webpage type, the second category of keywords are verbs associated with the target webpage type, and the third category of keywords are words unrelated to the target webpage type.

4. The method according to claim 1, wherein The calling of the target webpage identification model to perform target webpage type identification processing and target brand identification processing on the target webpage based on the webpage text corresponding to the target webpage includes: Inputting the target brand and the webpage text corresponding to the target webpage into a target webpage recognition model; Performing feature extraction processing on the webpage text and the target brand through the feature extraction branch of the target webpage recognition model to obtain webpage features corresponding to the target webpage; Performing target web page type identification processing on the web page features through the first prediction branch of the target web page identification model to obtain a target web page type identification result; The target brand recognition process is performed on the webpage features through the second prediction branch of the target webpage recognition model to obtain a target brand recognition result.

5. The method according to claim 4, characterized in that The method further comprises: Obtaining a training sample set; wherein the training samples in the training sample set are web page texts and training brands corresponding to web pages of known web page types and associated brands; inputting a first training sample from the training sample set into an initial web page recognition model to obtain a web page type prediction label corresponding to the first training sample, calculating a first model loss based on the web page type training label and the web page type prediction label corresponding to the first training sample, and adjusting model parameters of a feature extraction branch and a first prediction branch in the initial web page recognition model based on the first model loss to obtain a first intermediate web page recognition model; inputting a second training sample from the training sample set into the first intermediate web page recognition model to obtain a brand prediction label corresponding to the second training sample, calculating a second model loss based on the brand training label and the brand prediction label corresponding to the second training sample, and adjusting model parameters of a feature extraction branch and a second prediction branch in the first intermediate web page recognition model based on the second model loss to obtain a second intermediate web page recognition model; Input the third training sample in the training sample set into the second intermediate web page recognition model to obtain the comprehensive prediction label corresponding to the third training sample, calculate the third model loss based on the comprehensive prediction label and the comprehensive training label corresponding to the third training sample, adjust the model parameters of the second intermediate web page recognition model based on the third model loss, and obtain the target web page recognition model.

6. The method according to claim 5, characterized in that The obtaining of the training sample set includes: Obtain multiple unlabeled initial samples; Based on the sample similarity between the initial samples, multiple initial samples are deduplicated to obtain multiple training samples; Performing web page type labeling and brand labeling on the training samples to obtain comprehensive training labels corresponding to the plurality of training samples; A training sample set is obtained based on the multiple training samples and the comprehensive training labels respectively corresponding to the multiple training samples.

7. The method according to claim 1, characterized in that The webpage meta information includes the website filing information corresponding to the target webpage; The performing counterfeit identification processing on the target webpage based on the webpage meta information corresponding to the target webpage to obtain a counterfeit identification result corresponding to the target webpage includes: Performing a correlation test between the website filing party in the website filing information and the target brand to obtain the authenticity of the first webpage corresponding to the target webpage; Matching the website filing party in the website filing information with a preset filing party set to obtain a second webpage authenticity corresponding to the target webpage; A phishing identification result corresponding to the target webpage is determined based on at least one of the first webpage authenticity and the second webpage authenticity.

8. The method according to claim 7, characterized in that The performing of correlation detection between the website filing party in the website filing information and the target brand to obtain the authenticity of the first webpage corresponding to the target webpage includes: Pre-processing the website filing party and the target brand in the website filing information; When the similarity between the pre-processed website filing party and the target brand is greater than a similarity threshold, determining that the first webpage authenticity corresponding to the target webpage is a first preset authenticity; When the similarity between the pre-processed website filing party and the target brand is less than or equal to the similarity threshold, the first webpage authenticity corresponding to the target webpage is determined to be a second preset authenticity; the first preset authenticity is greater than the second preset authenticity.

9. The method according to claim 7, characterized in that The determining, based on at least one of the first webpage authenticity and the second webpage authenticity, a webpage phishing identification result corresponding to the target webpage includes: fusing at least one of the first webpage authenticity and the second webpage authenticity to obtain a comprehensive webpage authenticity corresponding to the target webpage; When the comprehensive webpage authenticity is greater than or equal to the webpage authenticity threshold, the webpage phishing identification result corresponding to the target webpage is determined to be a non-phishing webpage.

10. The method according to claim 7, characterized in that The webpage meta information includes at least one of the website filing information, network protocol address information, and website provider registration information corresponding to the target webpage; The method further comprises: When the webpage phishing identification result corresponding to the target webpage is a result other than a non-phishing webpage, performing a correlation test between the website filing party in the website filing information and the target brand to obtain a first webpage phishing degree corresponding to the target webpage; Performing a reliability test based on the filing time information in the website filing information to obtain a second webpage counterfeiting degree corresponding to the target webpage; Performing reliability detection based on the network protocol address information to obtain a third webpage counterfeiting degree corresponding to the target webpage; Comparing the current registration status in the website provider registration information with the preset registration status to obtain a fourth webpage counterfeiting degree corresponding to the target webpage; A phishing identification result corresponding to the target webpage is determined based on at least two of the first phishing degree, the second phishing degree, the third phishing degree, and the fourth phishing degree.

11. The method according to claim 10, characterized in that The determining of a phishing identification result corresponding to the target webpage based on at least two of the first phishing degree, the second phishing degree, the third phishing degree, and the fourth phishing degree includes: fusing at least two of the first phishing degree, the second phishing degree, the third phishing degree, and the fourth phishing degree to obtain a comprehensive phishing degree corresponding to the target webpage; When the comprehensive webpage phishing degree is greater than a webpage phishing degree threshold, the webpage phishing identification result corresponding to the target webpage is determined to be a phishing webpage.

12. The method according to claim 1, characterized in that The determining, based on the webpage phishing identification result, a brand counterfeiting identification result corresponding to the target brand includes: determining the number of non-counterfeit web pages and the number of counterfeit web pages corresponding to the target brand based on the webpage phishing identification results of each of the plurality of target web pages corresponding to the target brand; determining a brand counterfeiting degree corresponding to the target brand based on the number of the non-counterfeit web pages and the number of the counterfeit web pages; The brand counterfeiting degree is used as a brand counterfeiting identification result corresponding to the target brand.

13. A web page processing device, characterized in that: The device comprises: Brand acquisition module, used to acquire target brands; a webpage screening module for determining a target webpage from a set of candidate webpages based on brand characteristics corresponding to the target brand and type characteristics corresponding to the target webpage type; A webpage content recognition module is used to call a target webpage recognition model and perform target webpage type recognition and target brand recognition on the target webpage based on the webpage text corresponding to the target webpage; a webpage phishing identification module configured to, when an identification result indicates that the target webpage belongs to the target webpage type and is associated with the target brand, perform phishing identification processing on the target webpage based on webpage meta information corresponding to the target webpage, and obtain a webpage phishing identification result corresponding to the target webpage; The brand counterfeiting identification module is configured to determine a brand counterfeiting identification result corresponding to the target brand based on the webpage counterfeiting identification result.

14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 12 are implemented.

15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.