Website identification method and device and electronic equipment

By detecting the protocols and matching the explicit features of the websites to be tested, combined with judging abnormal information in the hypertext markup data of the registration pages, the problem of low quality of the dataset of fraudulent websites is solved, and the accuracy of the recognition algorithm and the quality of the model training dataset are improved.

CN120675734APending Publication Date: 2025-09-19BEIJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510611991.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

The existing data sets on fraudulent websites are of low quality, making it difficult to collect active fraudulent websites on a large scale and continuously, resulting in insufficient generalization capabilities of the recognition algorithm.

Method used

By obtaining the uniform resource locator information of the website to be tested, protocol detection is performed to determine the survival status, and target feature data is extracted in the survival status. Feature matching is performed using explicit features, and further, preset abnormal information in the registration page hypertext markup data is used to determine whether it is a fraudulent website.

Benefits of technology

It improves the generalization ability of the recognition algorithm, improves the accuracy of judging whether the tested website is a fraudulent website, and enhances the quality of the model training data set.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120675734A_ABST
    Figure CN120675734A_ABST
Patent Text Reader

Abstract

The invention provides a website identification method and device and electronic equipment, and the method comprises the steps: obtaining uniform resource locator information of a to-be-detected website, carrying out the protocol detection processing of the to-be-detected website according to the uniform resource locator information, and determining the survival state of the to-be-detected website; determining target feature data of the to-be-tested website in response to the fact that the survival state of the to-be-tested website is a survival state; obtaining a preset dominant feature, and performing feature matching on the target feature data by using the dominant feature to obtain a matching result; in response to the fact that the matching result is successful matching, determining registration page hypertext marking data corresponding to the website to be tested; and in response to preset abnormal information existing in the registration page hypertext marking data, determining that the to-be-tested website is a target website, and recording the to-be-tested website. According to the method and the device, the accuracy of judging whether the website to be tested is the fraud-related website or not is improved, and the quality of a subsequent fraud-related website model training data set is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of network security, and in particular to a website identification method, device, and electronic device. Background Art

[0002] With the help of the Internet, fraudsters have more tools and newer fraud methods. Compared with telephone fraud and SMS fraud in traditional telecommunications networks, more and more fraudsters have chosen to commit fraud through websites. They will use various advanced network development technologies to quickly build fraud websites according to the purpose of fraud or social hot events, and then lure victims into fraud traps by impersonating regular institutions or opening fake projects; or combine telephone and text messages to impersonate acquaintances, public security, procuratorial and judicial personnel, and e-commerce logistics customer service to lead victims to the already built fraud websites and commit fraud.

[0003] An important basis for the research on fraudulent website identification methods is the quantity and quality of fraudulent website datasets. Using high-quality datasets can improve the generalization ability of recognition algorithms. However, most fraudulent website-related datasets in current research work have problems. Summary of the Invention

[0004] In view of this, the purpose of the present disclosure is to provide a website identification method, device and electronic device to solve or partially solve the above-mentioned problems.

[0005] Based on the above objectives, a first aspect of the present disclosure provides a website identification method, the method comprising:

[0006] Obtaining uniform resource locator information of the website to be tested, performing protocol detection processing on the website to be tested according to the uniform resource locator information, and determining the existence status of the website to be tested;

[0007] In response to the survival status of the website to be tested being an alive state, determining target feature data of the website to be tested;

[0008] Acquire a preset dominant feature, and perform feature matching on the target feature data using the dominant feature to obtain a matching result;

[0009] In response to the matching result being a successful match, determining registration page hypertext markup data corresponding to the website to be tested;

[0010] In response to the presence of preset abnormal information in the registration page hypertext markup data, the website to be tested is determined to be a target website, and the website to be tested is recorded.

[0011] Based on the same inventive concept, the second aspect of the present disclosure provides a website identification device, the device comprising:

[0012] a data acquisition module configured to acquire uniform resource locator information of a website to be tested, perform protocol detection processing on the website to be tested according to the uniform resource locator information, and determine the existence status of the website to be tested;

[0013] a target characteristic data determination module, configured to determine target characteristic data of the website to be tested in response to the survival state of the website to be tested being in the alive state;

[0014] a feature matching module configured to obtain a preset dominant feature, perform feature matching on the target feature data using the dominant feature, and obtain a matching result;

[0015] a data determination module configured to determine registration page hypertext markup data corresponding to the website to be tested in response to the matching result being a successful match;

[0016] The anomaly detection module is configured to determine that the website to be tested is a target website in response to the presence of preset anomaly information in the registration page hypertext markup data, and record the website to be tested.

[0017] Based on the same inventive concept, the third aspect of the present disclosure proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the website identification method as described above when executing the computer program.

[0018] Based on the same inventive concept, a fourth aspect of the present disclosure proposes a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the website identification method as described above.

[0019] From the above, it can be seen that the present disclosure proposes a website identification method, device and electronic device, which obtain the uniform resource locator information of the website to be tested, perform protocol detection processing on the website to be tested according to the uniform resource locator information, and determine the survival status of the website to be tested. In response to the survival status of the website to be tested being a survival state, the target feature data of the website to be tested is determined. When it is determined that the website to be tested is still alive, the subsequent operation of determining whether the website to be tested is a fraudulent website is continued. If the website to be tested has been deactivated, there is no need to continue the detection, thereby avoiding unnecessary detection operations and causing waste of resources. A preset explicit feature is obtained, and the target feature data is matched using the explicit feature to obtain a matching result. In response to the matching result being a successful match, it indicates that the website to be tested may be a fraudulent website. For further determination, the registration page hypertext tag data corresponding to the website to be tested is determined. If the registration page's hypertext markup contains pre-set anomaly information, the website under test can be identified as the target website, indicating that it is a fraudulent website. This information is then recorded for subsequent training of the fraudulent website identification model, where it can be used as training data. This improves the generalization capabilities of the recognition algorithm. Furthermore, through the dual determination of feature matching based on explicit features and the presence of anomaly information, the accuracy of the fraudulent website determination is further enhanced, thereby improving the quality of the subsequent model training dataset. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 This is a flow chart of a website identification method according to an embodiment of the present disclosure;

[0022] Figure 2 This is a flowchart of a website identification method according to another embodiment of the present disclosure;

[0023] Figure 3 This is a flowchart of explicit feature matching according to another embodiment of the present disclosure;

[0024] Figure 4 This is a flow chart of an algorithm for detecting abnormal information on a registration page according to another embodiment of the present disclosure;

[0025] Figure 5 This is a structural block diagram of a website identification device according to an embodiment of the present disclosure;

[0026] Figure 6Schematic diagram of the structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0027] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0028] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may also change accordingly.

[0029] The terms used in this disclosure are explained as follows:

[0030] URL: A Uniform Resource Locator (URL) is the address of a resource on the Internet, commonly known as a "website." It is a concise representation of the location and access method of a resource on the Internet.

[0031] HTML: Hypertext Markup Language (HTML) is the standard markup language used to create web pages. It defines the meaning and structure of web page content and is the foundational technology that makes up the web.

[0032] HTTP: Hypertext Transfer Protocol (HTTP) is a simple request-response protocol that typically runs on top of TCP. It specifies what messages a client may send to a server and what responses it may receive.

[0033] HTTPS: Hypertext Transfer Protocol Secure (HTTPS) is a HTTP channel with security as its goal. It ensures the security of the transmission process through transmission encryption and identity authentication based on HTTP.

[0034] DOM: Document Object Model (DOM) is a programming interface that allows scripting languages ​​such as JavaScript to interact with and manipulate web documents. DOM structures documents into a series of programmatically accessible nodes and objects, allowing developers to change the structure, style, and content of documents.

[0035] JavaScript: JavaScript is a lightweight, interpreted or just-in-time compiled programming language that is widely used in web development, mainly to enhance the interactivity and dynamism of web pages.

[0036] CSS: Cascading Style Sheets (CSS) is a computer language used to express the style of documents such as HTML (an application of Standard Generalized Markup Language) or XML (a subset of Standard Generalized Markup Language).

[0037] Icon hashing: Icon hashing is a technique used for information gathering and penetration testing. It calculates the hash value of a website's icon (usually a favicon.ico file) to identify and locate related assets.

[0038] DNS: Domain Name System (DNS) is a distributed database that maps domain names to IP addresses. It can convert domain names entered by users into IP address strings that can be directly and machine-readable, enabling convenient network interconnection.

[0039] FOFA: FOFA is a cyberspace mapping search engine that helps users find Internet assets on the public network, such as cameras, printers, databases, etc.

[0040] With the internet, fraudsters have more tools and newer methods. Compared to traditional telecom phone and text message scams, more and more fraudsters are using websites to carry out their fraud. They use various advanced web development technologies to quickly build fraudulent websites based on the purpose of the fraud or social hot spots. They then lure victims into the trap by impersonating legitimate institutions or setting up fake projects. Or they combine phone calls and text messages, impersonating acquaintances, public security, judicial personnel, or e-commerce logistics customer service representatives to lure victims to the fraudulent websites they have set up and commit fraud. Websites are a key vector for internet fraud. Promptly identifying and discovering fraudulent websites in the real world can effectively maintain cybersecurity and keep it clean. Promptly discovering and blocking these fraudulent websites can greatly improve the financial and information security of online users, which is of great practical significance.

[0041] An important basis for the research on fraudulent website identification methods is the quantity and quality of fraudulent website datasets. Using high-quality datasets can improve the generalization ability of recognition algorithms. However, most fraudulent website-related datasets in current research work have problems.

[0042] Ke et al. used a unit's internet egress traffic as experimental data and independently designed a dataset extraction method for three high-value domain names. This method extracted daily DNS domain access and HTTP log data and collected fraudulent websites based on three characteristics: short lifespan, domain hijacking, and social sharing. Chen et al. conducted extensive searches on Google search engines, Whoscall, and a fraud ranking website and collected a number of fraudulent websites. Bitaab et al. crawled relevant posts from the Reddit forum and automatically labeled a number of e-commerce fraud websites using a BERT model based on user comment sentiment. You et al. used manual collection and annotation to collect a number of fraudulent websites from a publicly available online dataset, supplementing the dataset with samples of real cyberattack cases provided by relevant departments. Yang et al. directly used a large number of telecommunications fraud samples collected for public security work.

[0043] Based on the existing technical methods, it can be seen that the methods for collecting fraud-related websites currently used in the construction of fraud-related datasets generally have problems such as non-disclosure, poor timeliness, and low quality. In particular, it is difficult to collect active fraud-related websites in the real cyberspace. Therefore, it is not possible to widely and continuously collect fraud-related websites on the Internet to build a dataset that can adapt to the work of fraud-related website identification.

[0044] Therefore, based on the above description, this embodiment proposes a website identification method, such as Figure 1 As shown, the method includes:

[0045] Step 101: Obtain the uniform resource locator information of the website to be tested, perform protocol detection processing on the website to be tested according to the uniform resource locator information, and determine the existence status of the website to be tested.

[0046] Step 102: In response to the survival status of the website to be tested being an alive state, determine target feature data of the website to be tested.

[0047] Step 103: Obtain a preset dominant feature, and use the dominant feature to perform feature matching on the target feature data to obtain a matching result.

[0048] Step 104 : In response to the matching result being a successful match, determining the registration page hypertext markup data corresponding to the website to be tested.

[0049] Step 105 : In response to the presence of preset abnormal information in the registration page hypertext markup data, determining that the website to be tested is a target website, and recording the website to be tested.

[0050] In specific implementation, the uniform resource locator information of the website to be tested is obtained, and protocol detection processing is performed on the website to be tested according to the uniform resource locator information to determine the survival status of the website to be tested, wherein the protocol detection means detecting the application layer protocol used by the website to be tested.

[0051] In this embodiment, the website to be tested is a web page. For web pages, HTTP protocol or HTTPS protocol is usually used.

[0052] When it is determined that the survival state of the website to be tested is alive, target feature data of the website to be tested is determined, wherein the target feature data is feature data extracted from website hypertext markup data of the website to be tested.

[0053] Obtain preset explicit features, wherein the explicit features are feature data pre-extracted from multiple known fraudulent websites. The explicit features represent explicit or inductive information in the HTML, text, etc. of a certain website or a certain type of website. It does not need to guarantee the comprehensive and accurate identification of fraudulent websites, but by using these feature rules, suspected fraudulent websites can be preliminarily screened out from the intelligence data source.

[0054] The target feature data is feature matched using the dominant feature to obtain a matching result, wherein the feature matching refers to matching and comparing the target feature data with the dominant feature to determine whether the target feature data and the dominant feature are identical or contain each other.

[0055] If the matching result is a successful match, it indicates that the website to be tested is a suspected fraudulent website, that is, it may be a fraudulent website. The website hypertext markup data of the website to be tested is obtained, and the registration page hypertext markup data corresponding to the website to be tested is determined according to the website hypertext markup data.

[0056] The registration page hypertext markup data is traversed to determine whether preset exception information exists in the registration page hypertext markup data. If the preset exception information exists in the registration page hypertext markup data, the website under test is determined to be the target website, i.e., the website under test is a fraudulent website. The website under test is recorded so that when a fraudulent website identification model is subsequently trained, the website under test can be input into the model as training data for training, thereby improving the generalization ability of the recognition algorithm. The preset exception information is pre-set information that can be used to determine that a website is a fraudulent website.

[0057] In this embodiment, the preset abnormal information includes at least one of the following: an invitation code, a referral code, a referrer, an account opening code, a transaction password, a withdrawal password, an ID number, or a real name. When the registration page hypertext markup data contains at least one of the preset abnormal information, the website to be tested can be determined to be a fraudulent website.

[0058] Through the above scheme, the uniform resource locator information of the website to be tested is obtained, and the protocol detection processing of the website to be tested is performed on the website to be tested according to the uniform resource locator information to determine the survival status of the website to be tested. In response to the survival status of the website to be tested being a survival state, the target feature data of the website to be tested is determined. When it is determined that the website to be tested is still alive, the subsequent operation of determining whether the website to be tested is a fraudulent website is continued. If the website to be tested has been deactivated, there is no need to continue the detection, so as to avoid unnecessary detection operations and waste of resources. The preset explicit feature is obtained, and the target feature data is matched with the explicit feature to obtain a matching result. In response to the matching result being a successful match, it means that the website to be tested may be a fraudulent website. For further determination, the registration page hypertext markup data corresponding to the website to be tested is determined. If the registration page's hypertext markup contains pre-set anomaly information, the website under test can be identified as the target website, indicating that it is a fraudulent website. This information is then recorded for subsequent training of the fraudulent website identification model, where it can be used as training data. This improves the generalization capabilities of the recognition algorithm. Furthermore, through the dual determination of feature matching based on explicit features and the presence of anomaly information, the accuracy of the fraudulent website determination is further enhanced, thereby improving the quality of the subsequent model training dataset.

[0059] In some embodiments, the explicit feature specifically includes at least one of the following: website title key text, hypertext markup resources, or website icon data. The website title key text includes website title keywords and website title key words. The hypertext markup resources include script data and cascading style sheets.

[0060] The website title is the content of the title attribute when the HTML language is used to create a web page, that is, the part of the HTML document. <title> The content within the tag is also the text displayed on the tab at the top of the browser, that is, one of the important visual elements that browser users obtain when visiting a website. A good website title usually summarizes the theme of the website or the core service content concisely and clearly, and can identify the main function or purpose of the website to a certain extent. Therefore, as an intuitive feature of the website, when developing a fraudulent website, fraudsters often put the theme of the website directly in the content of the title, such as in a counterfeit website, which can trick victims into believing that this is a genuine website. For the website title, this embodiment mainly performs feature extraction in two dimensions. One is to perform character-level segmentation on the title and count its word frequency. The other is to perform word segmentation on the title and count its word frequency.< / title>

[0061] The titles of the fraud-related websites that have been collected are split at the character level. Taking "China Cloud Data" as an example, it will be split into four characters: "中", "国", "云", and "数". Count the number of occurrences of all characters. If the statistical frequency of a certain character is high, it means that the keyword corresponding to this character is more likely to be included in the title of the fraud-related website, that is, the themes of the fraud-related websites developed by fraudsters tend to be close to the high-frequency characters. Therefore, when actively discovering from multi-channel data sources based on these high-frequency keywords, there is a greater probability of matching fraud-related websites from a vast number of websites.

[0062] The titles of the fraud-related websites that have been collected are split at the word level. For example, "China Cloud Data" is split into two words: "China" and "Cloud Data" after splitting. Count the frequency of occurrence of all words. The developers of fraud-related websites often use some high-frequency words to design the titles of the websites. When actively discovering from multi-channel data sources based on these high-frequency keywords, there is a greater probability of matching fraud-related websites from a vast number of websites. For example, using "Cloud Data" as a matching word can further discover other fraud-related websites such as "Cloud Data Fuyu", "Cloud Data Real Estate", and "Five Elements Cloud Data Coin".

[0063] The website HTML contains various elements, and the files among them are resources that need to be loaded into the browser when the web page content is loaded, including JavaScript, css, pictures, etc. During the development of fraud-related websites, fraudsters will use specific resources, and multiple fraud-related websites developed by the same fraud gang often adopt the same or similar HTML resources. Therefore, extracting the resources in the HTML of fraud-related websites as features, that is, taking hypertext markup resources as explicit features, can match more fraud-related websites from the intelligence data sources.

[0064] In HTML, JavaScript technology is used to add dynamic and interactive elements to web pages. It does not display content statically, but enables the page to respond to user operations, manipulate the page document, process asynchronous communication with the server, etc. Modern JavaScript also supports more complex application development, such as single-page applications (SPAs). In such applications, the entire website uses only one HTML page, and the update of content is completed by dynamically loading data through JavaScript, providing a smooth and fast-responsive user experience. Currently, more and more fraud websites have adopted the development mode of single-page applications. Therefore, JavaScript files can be used as an explicit feature for discovering fraud-related websites.

[0065] During the process of actively discovering fraud-related websites, the main focus on the use of JavaScript is the extraction type:

[0066] a. Identically named JavaScript resources. For example, when a fraud gang develops a fraudulent website, they often develop or purchase the same set of website source code templates. This way, the JavaScript file naming rules of the multiple websites they develop are consistent. Therefore, based on these JavaScript resources, multiple different fraudulent websites can be discovered from the intelligence data source.

[0067] b. Similar-named JavaScript resources. After a fraud gang develops or purchases a website source code template, they fine-tune the content and name of the JavaScript resources based on the business needs of different members or groups. Some of these resources use similar naming practices. Once the naming rules of JavaScript resources on fraudulent websites are discovered, regular expression matching is used to perform fuzzy matching within the HTML, allowing the discovery of multiple different fraudulent websites.

[0068] In HTML, CSS is a style design language used to describe the appearance of a document, primarily used in web page and user interface design. CSS allows precise control over the appearance of nearly every page element, including fonts, colors, background images, element sizing, and border styles. CSS also supports a range of visual and interactive features, such as website layout management and responsive design. Similar to JavaScript resources, websites built using the same web development framework often use the same CSS resource filename. For example, the front-end development framework uni-app is used for developing all front-end applications, allowing developers to write a single codebase and publish across multiple platforms. Some fraudsters are currently using this framework to rapidly develop the front-ends of fraudulent websites. Analysis of fraudulent websites built directly or with modifications to the uni-app framework revealed that the CSS resource ". / static / index.2da1efab.css" is used across all websites. Many other frameworks operate similarly, making it possible to proactively collect more fraudulent websites using CSS resources.

[0069] Website icon data iconhash, iconhash is a hashing algorithm used for icons (such as website favicons). It generates a unique fingerprint or identifier by hashing the visual content of the icon. Usually, each individually developed website contains an ico file in the HTML. When the website is accessed using a browser, the ico file will be loaded at the top of the browser to generate an icon. If the website is developed using the same framework and the ico file is not adjusted by itself, then its iconhash will be the same. Therefore, iconhash can be used to proactively discover fraudulent websites from intelligence data sources.

[0070] In this embodiment, the dominant features are feature data extracted from multiple known fraud-related websites. To improve the accuracy and universality of the dominant features, the multiple known fraud-related websites are collected from real cyberspace.

[0071] In this embodiment, the collection channel of the known fraud-related websites includes at least one of the following: telecommunications operator DNS domain name resolution data, cyberspace mapping intelligence platform or media release content.

[0072] The DNS domain name system is one of the core services of the Internet. When a user accesses a website corresponding to a domain name through a browser, the DNS converts the domain name into an IP address so that the browser can access the server to load the page resources. Operators have their own DNS resolution data. When the operator's users, including victims of telecommunications fraud, access domain names using networks such as home broadband, mobile networks, and data centers, the DNS server stores the domain name resolution records. Operators' services are spread across 31 provinces across China, and they possess a vast amount of domain name data, including various fraudulent websites visited by victims. While working in the relevant anti-fraud department of an operator, the author was able to leverage this large network capability to proactively discover fraudulent websites. Therefore, it is feasible to collect fraudulent websites through the operator's DNS resolution data.

[0073] Cyberspace mapping is a key technology in cybersecurity, designed to create detailed "maps" of cyberspace. This involves understanding and mapping the complex relationships between geographic, social, and cyberspace, with the goal of transforming the virtual and dynamically changing online environment into a dynamic, real-time, and effective diagram. FOFA, a cyberspace search engine launched by WhiteHat, uses cyberspace mapping to help researchers quickly match network assets. When searching using FOFA, it provides multiple website dimensions, including title, body, js_name, iconhash, header, and cname, which can be used to subsequently screen for fraudulent websites.

[0074] In some embodiments, performing protocol detection processing on the website to be tested according to the uniform resource locator information in step 101 to determine the existence status of the website to be tested specifically includes:

[0075] Step 1011, using Hypertext Transfer Protocol Security to execute the Uniform Resource Locator information to obtain a first execution result;

[0076] Step 1012: In response to the first operation result being a successful operation, determining that the survival status of the website to be tested is a live state;

[0077] Step 1013: In response to the first execution result being an execution failure, executing the uniform resource locator information using the hypertext transfer protocol to obtain a second execution result;

[0078] Step 1014: In response to the second operation result being a successful operation, determining that the survival status of the website to be tested is an alive state.

[0079] In specific implementation, the URL information is first run using the Hypertext Transfer Protocol Secure (HTTS) to obtain a first run result. If the first run result is a successful run, then the survival status of the website to be tested can be determined to be a live state.

[0080] If the first operation result is an operation failure, the URL information needs to be continued to be operated using the hypertext transfer protocol to obtain a second operation result. If the second operation result is an operation success, the survival state of the website to be tested is determined to be a survival state.

[0081] If the second running result is still a running failure, it can be determined that the website to be tested has been deactivated and no further subsequent processing is required.

[0082] For example, the URL of the website under test is domain-example.com. First, the HTTPS URL is checked for liveness, i.e., whether https: / / domain-example.com is live. If so, the website under test is alive. If there is no response, the HTTP URL is checked for liveness, i.e., whether http: / / domain-example.com is live. If so, the website under test is alive.

[0083] Through the above scheme, by using the Hypertext Transfer Protocol (HTTP) and the Hypertext Transfer Protocol (HTTP) to judge whether the website to be tested is alive, it avoids the situation where errors occur when relying solely on the Hypertext Transfer Protocol (HTTP), thereby improving the accuracy of judging the survival status of the website to be tested.

[0084] In some embodiments, in response to the survival status of the website to be tested being in the alive state, determining target feature data of the website to be tested in step 102 specifically includes:

[0085] Step 1021: in response to the existence status of the website to be tested being in an alive state, obtaining website hypertext markup data of the website to be tested;

[0086] Step 1022: Parse the website hypertext markup data to obtain a document object model tree.

[0087] Step 1023: extract features from the document object model tree to obtain target feature data of the website to be tested.

[0088] In specific implementation, if the survival state of the website to be tested is determined to be alive, the website hypertext markup data of the website to be tested is obtained by obtaining the text attribute of the response, and the returned data is the website hypertext markup data.

[0089] The website hypertext markup data is parsed using a parser to obtain a document object model tree, and then feature extraction is performed on the document object model tree to obtain target feature data of the website to be tested.

[0090] In some embodiments, obtaining a preset dominant feature in step 103 and performing feature matching on the target feature data using the dominant feature to obtain a matching result specifically includes:

[0091] Step 1031: Obtain preset explicit features, wherein the explicit features include at least one of the following: preset title key text, preset hypertext markup resources, and preset website icon data;

[0092] Step 1032: In response to the target feature data including the preset title key text, determining that the matching result is a successful match; or

[0093] Step 1033: In response to the target feature data including the preset hypertext markup resource, determining that the matching result is a successful match; or,

[0094] Step 1034 : In response to the icon hash value in the target feature data being the same as the icon hash value corresponding to the preset website icon data, determining that the matching result is a successful match.

[0095] During specific implementation, a preset explicit feature is obtained, wherein the explicit feature includes at least one of the following: a preset title key text, a preset hypertext markup resource, and a preset website icon data.

[0096] The target feature data is matched using the preset title key text, and the title attribute of the parsed object can be used to obtain the website title. If the website title in the target feature data includes the preset title key text, the matching result is determined to be a successful match.

[0097] If the target feature data includes a preset hypertext markup resource, the match result is determined to be a successful match. The ico file in the target feature data is obtained and the corresponding iconhash value is calculated, where the iconhash value is the icon hash value in the target feature data. If the icon hash value in the target feature data is the same as the icon hash value corresponding to the preset website icon data, the match result is determined to be a successful match.

[0098] In some embodiments, in response to the matching result being a successful match, determining the registration page hypertext markup data corresponding to the website to be tested in step 104 specifically includes:

[0099] Step 1041: in response to the matching result being a successful match, obtaining website hypertext markup data of the website to be tested and a preset first regular expression;

[0100] Step 1042: performing regular expression matching on the website hypertext markup data using the first regular expression to obtain a first matching result;

[0101] Step 1043: in response to the first matching result being a successful match, obtaining a registration page path, crawling the website to be tested using the registration page path, and obtaining registration page hypertext markup data corresponding to the website to be tested; or

[0102] Step 1044 : In response to the first matching result being a match failure, determining the script file of the website to be tested, and determining registration page hypertext markup data corresponding to the website to be tested based on the script file.

[0103] In specific implementation, if the target feature data is matched with the explicit feature, the matching result is a successful match, which means that the website to be tested is a suspected fraudulent website. The website to be tested is tested again to determine whether it is a fraudulent website.

[0104] Get a preset first regular expression. In this embodiment, the first regular expression is ["\']( / [^"\']+ / reg[^"\']*| / reg[^"\']*)(?<!\.js)(?<!\.css)(?<!\.png)(?<!\.jpg)(?<!\.jpeg)["\'].

[0105] The website hypertext markup data of the website to be tested is obtained, and the website hypertext markup data is subjected to regular matching using the first regular expression to match the registration page path of the website to be tested, thereby obtaining a first matching result.

[0106] If the first matching result is a successful match, indicating that the registration page path is matched, the registration page URL is concatenated, that is, the web page path of the web page to be tested is replaced with the registration page path, and then the website to be tested is crawled according to the registration page path to obtain the registration page hypertext markup data corresponding to the website to be tested.

[0107] If the first matching result is a match failure, it means that the registration page path cannot be matched, that is, the registration page path may exist in a script file, such as a single-page application. Therefore, the script file of the web page to be tested is determined, and the registration page hypertext markup data corresponding to the website to be tested is determined based on the script file.

[0108] In some embodiments, determining the script file of the website to be tested in step 1044 specifically includes:

[0109] Step A, obtaining a preset second regular expression, and using the second regular expression to perform regular matching on the website hypertext markup data to obtain a second matching result;

[0110] Step B: In response to the second matching result being a successful match, obtaining the script file of the website to be tested.

[0111] In the specific implementation, all script files are matched in the website hypertext markup data, specifically, a preset second regular expression is obtained. In this embodiment, specifically, the second regular expression is<script\s+[^> ]*src=["\']?([^"\'>]+\.js)["\']?[^>]*>.

[0112] The second regular expression is used to perform regular matching on the website hypertext markup data, that is, to match the script file path to obtain a second matching result. If the second matching result is a successful match, it means that the script file path is matched, and the script file can be obtained according to the script file path.

[0113] If the second matching result is a matching failure, it is determined that the website to be tested is not a fraud-related website.

[0114] In this embodiment, if multiple script files are matched, all script files are crawled according to the script file path.

[0115] In some embodiments, determining the registration page hypertext markup data corresponding to the website to be tested based on the script file in step 1044 specifically includes:

[0116] Step a, obtaining a preset third regular expression, and performing regular matching on the script file using the third regular expression to obtain a third matching result;

[0117] Step b: in response to the third matching result being a successful match, obtaining a registration page path;

[0118] Step c: crawling the website to be tested using the registration page path to obtain registration page hypertext markup data corresponding to the website to be tested.

[0119] In a specific implementation, a preset third regular expression is obtained, and the third regular expression is used to match the registration page path in all script files, that is, the script files are regularly matched using the third regular expression to obtain a third matching result. In this embodiment, the third regular expression is expressed as [path:]["\']( / [^"\']+ / reg[^"\']*| / reg[^"\']*)(?<!\.js)(?<!\.css)(?<!\.png)(?<!\.jpg)(?<!\.jpeg)["\'].

[0120] In response to the third matching result being a successful match, which indicates that the registration page path is matched in the script file, the registration page path is concatenated and the registration page is loaded through a browser simulation to obtain the registration page hypertext markup data.

[0121] Specifically, the web page path of the web page to be tested is replaced with the registration page path. Since dynamic loading of the website is required, the playwright tool is used to simulate the loading of the registration page by the browser, and the website to be tested is crawled according to the registration page path to obtain the registration page hypertext markup data corresponding to the website to be tested.

[0122] In some embodiments, in response to the presence of preset abnormal information in the registration page hypertext markup data, determining that the website to be tested is a target website and recording the website to be tested in step 105 specifically include:

[0123] Step 1051: In response to the presence of preset abnormality information in the registration page hypertext markup data, output the website information of the website to be tested and prompt information, wherein the prompt information is information prompting to check whether the website to be tested has abnormality;

[0124] Step 1052: Receive feedback information corresponding to the prompt information, and in response to the feedback information indicating that an abnormality exists, determine that the website to be tested is a target website, and record the website to be tested.

[0125] In specific implementation, if it is determined that preset abnormal information exists in the registration page hypertext markup data, the website information and prompt information of the website to be tested are output to prompt the user to check whether there is any abnormality in the website to be tested, that is, whether it involves fraud.

[0126] Receive feedback information corresponding to the prompt information, wherein the feedback information indicates whether the website under test is fraudulent after the user checks the website under test. If the feedback information indicates that there is an abnormality, determine that the website under test is the target website, that is, the website under test is a fraudulent website, and record the website under test.

[0127] Through the above scheme, after determining that there is preset abnormal information in the hypertext markup data of the registration page of the website to be tested, manual review and screening are performed to further improve the recognition accuracy of fraudulent websites.

[0128] Through the website identification method described in this disclosure, a collection of websites suspected of being involved in fraud is obtained by combining automatic crawling with manual collection. Different methods are used to screen for websites suspected of being involved in fraud, respectively. If the source is from automatic crawling, the website registration page abnormal input detection algorithm proposed in this invention is used to further narrow the scope of suspected websites, and then manual screening is performed. If the source is from manual collection, manual screening is performed. After active discovery and semi-automatic screening, a group of real and effective fraud-related websites will be obtained.

[0129] In this embodiment, based on existing fraud intelligence data and domestic telecom fraud websites continuously collected during the active discovery process of fraud websites, the dominant features of fraud websites are counted and summarized. Based on these feature rules, other fraud-related websites can be matched and collided from various intelligence and data sources. In other words, these feature rules should be included in the information data that can be obtained by automatic crawling methods and website information acquisition tools, or included in the mapping dimensions of various intelligence data sources.

[0130] At the same time, based on the extracted explicit features of the website, it is necessary to find multiple types of intelligence data sources from the real network space and analyze their feasibility of using them to collect fraud-related websites, so as to make full use of the features to discover more fraud-related websites.

[0131] Based on the same inventive concept, another embodiment of the present disclosure provides a website identification method, such as Figure 2 As shown, the method specifically includes:

[0132] Step 201: Extract explicit features of fraudulent websites.

[0133] Based on existing fraud intelligence data and the continuous collection of domestic telecom fraud websites during the proactive discovery process, we collect and summarize the dominant characteristics of fraudulent websites. Based on these characteristic rules, we can match and collide other fraudulent websites from various intelligence and data sources. In other words, these characteristic rules should be included in the information data that can be obtained by automatic crawling methods and website information acquisition tools, or included in the mapping dimensions of various intelligence data sources.

[0134] Step 202: Select a collection channel for fraudulent websites.

[0135] Based on the extracted explicit features of the website, it is necessary to find multiple types of intelligence data sources from the real network space and analyze their feasibility for collecting fraud-related websites, so as to make full use of the features to discover more fraud-related websites.

[0136] Step 203: Actively discover and screen fraudulent websites.

[0137] Based on the dominant features and collection channels obtained in steps 201 and 202, a collection of suspected fraudulent websites is obtained using a combination of automated crawling and manual collection. Different methods are used to screen for fraudulent websites, depending on the website acquisition method. If the source is automated crawling, the website registration page abnormal input detection algorithm proposed in this invention is used to further narrow the scope of suspected websites, followed by manual screening. If the source is manual collection, manual screening is performed. After active discovery and semi-automated screening, a group of authentic and valid fraudulent websites will be obtained.

[0138] In some embodiments, step 203 specifically includes discovering suspected fraudulent websites and screening fraudulent websites. Specifically,

[0139] In each fraudulent website collection channel, enter the original website domain name or website URL, and then use the web crawler and browser loading method to obtain the website's multi-dimensional information, such as Figure 3 As shown, the specific process is as follows:

[0140] Perform protocol detection and liveness checks on the entered domain name. Protocol detection refers to checking the application layer protocol used by the website. For web pages, HTTP or HTTPS is usually used. First, check whether the URL using the HTTPS protocol is alive. If it is alive, use that URL as the test URL and continue with the following steps. If there is no response, use the HTTP protocol URL to detect liveness. If it is alive, use that URL as the test URL and continue with the following steps. Otherwise, the website is considered inactive and no further processing is performed.

[0141] Get the website HTML. After completing the URL request and checking for activity, for any live websites, get the text attribute of the response. The returned value is the website's HTML document.

[0142] Parse the website HTML to form a DOM tree. After obtaining the website's HTML, use a parser to parse the HTML document and convert the result into a DOM tree, so that you can easily extract the required content from the web page HTML.

[0143] Different methods are used to obtain objects related to explicit features. For title keywords and key words, the website title needs to be matched and parsed by the object's title attribute to obtain the website title. For HTML resources, whether JavaScript or CSS, the way they are referenced in HTML varies depending on the development style, so the matching object is the entire HTML document. For the website iconhash, the ico file in the HTML needs to be obtained and the corresponding iconhash value needs to be calculated.

[0144] After crawling the web, we obtain the HTML document, website title, website text letter, and website ico file of a website. Then, we use the existing explicit features of fraudulent websites to perform feature matching in each part, and we can actively discover suspected fraudulent websites.

[0145] After actively discovering suspected fraudulent websites by using explicit features of websites through multiple channels, it is necessary to screen out confirmed fraudulent websites because explicit features cannot completely determine whether a website is fraudulent. Among the suspected fraudulent websites, there are many normal websites. For example, by actively discovering through the title keyword "capital", the official websites of regular financial websites such as "Zhenghai Capital" will be collected. To this end, an abnormal input detection algorithm for website registration pages is proposed to screen out fraudulent websites from suspected fraudulent websites as accurately as possible and eliminate normal websites. Figure 4 As shown, the process specifically includes:

[0146] Step A: Match the registration page path in the website HTML. In the front-end development of many websites, the link information is written in the front-end HTML. You can directly use regular expressions to match the registration page path. Use the regular expression: ["\']( / [^"\']+ / reg[^"\']*| / reg[^"\']*)(?<!\.js)(?<!\.css)(?<!\.png)(?<!\.jpg)

[0147] (?<!\.jpeg)["\'] matches the path of the registration page of the fraudulent website. If the match is successful, proceed to step B; if the match fails, proceed to step C.

[0148] Step B: Concatenate the registration page URL and crawl the registration page HTML. Once the registration page path is matched, replace the original path with the registration page path, crawl the registration page URL HTML, and proceed to step G.

[0149] Step C, match all JavaScript file paths in the website HTML. If the match fails in step 1.1, consider the case where the registration page path exists in the javascript file, such as a single page application.<script\s+[^> ]*src=["\']?([^"\'>]+\.js)["\']?[^>]*> to match. If the JavaScript file path is matched, proceed to step D; if the match fails, it is determined that the website is not a fraudulent website.

[0150] Step D: Crawl the JavaScript file contents. If one or more JavaScript files are found, crawl the contents of all JavaScript files.

[0151] Step E: Match the registration page path in all JavaScript file contents. Use [path:]["\']( / [^"\']+ / reg[^"\']*| / reg[^"\']*)(?<!\.js)(?<!\.css)(?<!\.png)(?<!\.jpg)(?<!\.jpeg)["\'] to match the registration page path in the JavaScript file. If the match is successful, proceed to step F. If the match fails, determine that the website is not a fraudulent website.

[0152] Step F: Concatenate the registration page path and simulate loading the registration page in a browser to obtain the HTML. After matching the registration page path in the JavaScript file, a similar substitution is performed to generate the registration page URL. Because dynamic loading of the website is required, the Playwright tool is used to simulate loading the registration page in a browser and then obtain the loaded HTML.

[0153] Step G: Detecting abnormal information in the registration page HTML. The HTML obtained in both Step B and Step F contains the text required on the registration page. Therefore, the HTML is checked to see if it contains the collected abnormal information. If a match is found, manual review and screening are performed. If a match is found, the website is deemed not to be fraudulent.

[0154] By using the abnormal input detection algorithm for website registration pages, the suspected fraudulent websites collected during the active discovery process can be automatically screened, effectively narrowing the scope of websites that need to be reviewed. Finally, the results produced by the algorithm are manually screened to collect websites that are confirmed to be fraudulent and exist in the real cyberspace.

[0155] It should be noted that the method of the embodiments of the present disclosure can be performed by a single device, such as a computer or server. The method of the embodiments of the present disclosure can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiments of the present disclosure, and the multiple devices will interact with each other to complete the method.

[0156] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0157] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure further provides a website identification device.

[0158] refer to Figure 5 , Figure 5 The website identification device of the embodiment specifically includes:

[0159] The data acquisition module 501 is configured to obtain the uniform resource locator information of the website to be tested, perform protocol detection processing on the website to be tested according to the uniform resource locator information, and determine the existence status of the website to be tested;

[0160] The target characteristic data determination module 502 is configured to determine the target characteristic data of the website to be tested in response to the survival status of the website to be tested being in the survival status;

[0161] The feature matching module 503 is configured to obtain a preset dominant feature, perform feature matching on the target feature data using the dominant feature, and obtain a matching result;

[0162] The data determination module 504 is configured to determine the registration page hypertext markup data corresponding to the website to be tested in response to the matching result being a successful match;

[0163] The anomaly detection module 505 is configured to determine that the website to be tested is a target website in response to the presence of preset anomaly information in the registration page hypertext markup data, and record the website to be tested.

[0164] In some embodiments, the data acquisition module 501 is specifically configured to:

[0165] Executing the uniform resource locator information using the Hypertext Transfer Protocol Secure (HTTS) to obtain a first execution result;

[0166] In response to the first operation result being a successful operation, determining that the survival state of the website to be tested is a survival state;

[0167] In response to the first operation result being an operation failure, executing the uniform resource locator information using a hypertext transfer protocol to obtain a second operation result;

[0168] In response to the second operation result being a successful operation, it is determined that the survival state of the website to be tested is an alive state.

[0169] In some embodiments, the target feature data determination module 502 is specifically configured to:

[0170] In response to the existence status of the website to be tested being an alive state, obtaining website hypertext markup data of the website to be tested;

[0171] Parsing the website hypertext markup data to obtain a document object model tree;

[0172] Feature extraction is performed on the document object model tree to obtain target feature data of the website to be tested.

[0173] In some embodiments, the feature matching module 503 is specifically configured to:

[0174] Obtaining preset explicit features, wherein the explicit features include at least one of the following: preset title key text, preset hypertext markup resources, and preset website icon data;

[0175] In response to the target feature data including the preset title key text, determining that the matching result is a successful match; or,

[0176] In response to the target feature data including the preset hypertext markup resource, determining that the matching result is a successful match; or,

[0177] In response to the icon hash value in the target feature data being the same as the icon hash value corresponding to the preset website icon data, the matching result is determined to be a successful match.

[0178] In some embodiments, the data determination module 504 is specifically configured to:

[0179] In response to the matching result being a successful match, obtaining website hypertext markup data of the website to be tested and a preset first regular expression;

[0180] Performing regular expression matching on the website hypertext markup data using the first regular expression to obtain a first matching result;

[0181] In response to the first matching result being a successful match, obtaining a registration page path, crawling the website to be tested using the registration page path, and obtaining registration page hypertext markup data corresponding to the website to be tested; or

[0182] In response to the first matching result being a matching failure, a script file of the website to be tested is determined, and registration page hypertext markup data corresponding to the website to be tested is determined according to the script file.

[0183] In some embodiments, the data determination module 504 is specifically configured to:

[0184] Obtaining a preset second regular expression, and performing regular matching on the website hypertext markup data using the second regular expression to obtain a second matching result;

[0185] In response to the second matching result being a successful match, a script file of the website to be tested is obtained.

[0186] In some embodiments, the data determination module 504 is specifically configured to:

[0187] Obtain a preset third regular expression, and perform regular matching on the script file using the third regular expression to obtain a third matching result;

[0188] In response to the third matching result being a successful match, obtaining a registration page path;

[0189] The website to be tested is crawled using the registration page path to obtain registration page hypertext markup data corresponding to the website to be tested.

[0190] In some embodiments, the anomaly detection module 505 is specifically configured to:

[0191] In response to the presence of preset abnormal information in the registration page hypertext markup data, outputting website information and prompt information of the website to be tested, wherein the prompt information is information prompting to check whether the website to be tested has abnormality;

[0192] Feedback information corresponding to the prompt information is received, and in response to the feedback information indicating that an abnormality exists, the website to be tested is determined to be a target website, and the website to be tested is recorded.

[0193] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0194] The apparatus of the above embodiment is used to implement the corresponding website identification method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0195] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the website identification method described in any of the above embodiments is implemented.

[0196] Figure 6 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.

[0197] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.

[0198] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.

[0199] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.

[0200] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).

[0201] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).

[0202] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.

[0203] The electronic device of the above embodiment is used to implement the corresponding website identification method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0204] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the website identification method described in any of the above embodiments.

[0205] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0206] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the website identification method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0207] It is understandable that before using the technical solutions of each embodiment of the present disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved will be informed to the user in an appropriate manner, and the user's authorization will be obtained.

[0208] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the disclosed technical solution based on the prompt message.

[0209] As an optional but non-limiting implementation, in response to a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0210] It is understandable that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0211] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0212] In addition, to simplify the description and discussion, and so as not to obscure the embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the purview of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0213] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0214] The embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A website identification method, characterized in that: include: Obtaining uniform resource locator information of the website to be tested, performing protocol detection processing on the website to be tested according to the uniform resource locator information, and determining the existence status of the website to be tested; In response to the survival status of the website to be tested being an alive state, determining target feature data of the website to be tested; Acquire a preset dominant feature, and perform feature matching on the target feature data using the dominant feature to obtain a matching result; In response to the matching result being a successful match, determining registration page hypertext markup data corresponding to the website to be tested; In response to the presence of preset abnormal information in the registration page hypertext markup data, the website to be tested is determined to be a target website, and the website to be tested is recorded.

2. The method according to claim 1, characterized in that The performing protocol detection processing on the website to be tested according to the uniform resource locator information to determine the existence status of the website to be tested includes: Executing the uniform resource locator information using the Hypertext Transfer Protocol Secure (HTTS) to obtain a first execution result; In response to the first operation result being a successful operation, determining that the survival state of the website to be tested is a survival state; In response to the first operation result being an operation failure, executing the uniform resource locator information using a hypertext transfer protocol to obtain a second operation result; In response to the second operation result being a successful operation, it is determined that the survival state of the website to be tested is an alive state.

3. The method according to claim 1, characterized in that In response to the survival state of the website to be tested being an alive state, determining target feature data of the website to be tested includes: In response to the existence status of the website to be tested being an alive state, obtaining website hypertext markup data of the website to be tested; Parsing the website hypertext markup data to obtain a document object model tree; Feature extraction is performed on the document object model tree to obtain target feature data of the website to be tested.

4. The method according to claim 1, wherein The obtaining of a preset dominant feature, and performing feature matching on the target feature data using the dominant feature to obtain a matching result, includes: Obtaining preset explicit features, wherein the explicit features include at least one of the following: preset title key text, preset hypertext markup resources, and preset website icon data; In response to the target feature data including the preset title key text, determining that the matching result is a successful match; or, In response to the target feature data including the preset hypertext markup resource, determining that the matching result is a successful match; or, In response to the icon hash value in the target feature data being the same as the icon hash value corresponding to the preset website icon data, the matching result is determined to be a successful match.

5. The method according to claim 1, wherein In response to the matching result being a successful match, determining the registration page hypertext markup data corresponding to the website to be tested includes: In response to the matching result being a successful match, obtaining website hypertext markup data of the website to be tested and a preset first regular expression; Performing regular expression matching on the website hypertext markup data using the first regular expression to obtain a first matching result; In response to the first matching result being a successful match, obtaining a registration page path, crawling the website to be tested using the registration page path, and obtaining registration page hypertext markup data corresponding to the website to be tested; or In response to the first matching result being a matching failure, a script file of the website to be tested is determined, and registration page hypertext markup data corresponding to the website to be tested is determined according to the script file.

6. The method according to claim 5, characterized in that The step of determining the script file of the website to be tested includes: Obtaining a preset second regular expression, and performing regular matching on the website hypertext markup data using the second regular expression to obtain a second matching result; In response to the second matching result being a successful match, a script file of the website to be tested is obtained.

7. The method according to claim 5, characterized in that The step of determining the registration page hypertext markup data corresponding to the website to be tested according to the script file includes: Obtain a preset third regular expression, and perform regular matching on the script file using the third regular expression to obtain a third matching result; In response to the third matching result being a successful match, obtaining a registration page path; The website to be tested is crawled using the registration page path to obtain registration page hypertext markup data corresponding to the website to be tested.

8. The method according to claim 1, characterized in that In response to the presence of preset abnormal information in the registration page hypertext markup data, determining that the website to be tested is a target website and recording the website to be tested includes: In response to the presence of preset abnormal information in the registration page hypertext markup data, outputting website information and prompt information of the website to be tested, wherein the prompt information is information prompting to check whether the website to be tested has abnormalities; Feedback information corresponding to the prompt information is received, and in response to the feedback information indicating that an abnormality exists, the website to be tested is determined to be a target website, and the website to be tested is recorded.

9. A website identification device, characterized in that: include: A data acquisition module is configured to acquire uniform resource locator information of a website to be tested, perform protocol detection processing on the website to be tested according to the uniform resource locator information, and determine the existence status of the website to be tested; a target characteristic data determination module, configured to determine target characteristic data of the website to be tested in response to the survival state of the website to be tested being in the alive state; a feature matching module configured to obtain a preset dominant feature, perform feature matching on the target feature data using the dominant feature, and obtain a matching result; a data determination module configured to determine registration page hypertext markup data corresponding to the website to be tested in response to the matching result being a successful match; The anomaly detection module is configured to determine that the website to be tested is a target website in response to the presence of preset anomaly information in the registration page hypertext markup data, and record the website to be tested.

10. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 8 is implemented.