A Website Homepage Recognition Method and Electronic Device Based on URL Features

Through the recognition method based on URL features, nested URLs are stripped and regular expressions are used to match URL domain names and keywords, the problem of insufficient recognition speed and accuracy of website homepages in the prior art is solved, and efficient and accurate recognition effect is achieved.

CN114201698BActive Publication Date: 2025-06-24NAT COMP NETWORK & INFORMATION SECURITY MANAGEMENT CENT +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202010981078.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-17
Publication Date
2025-06-24
Estimated Expiration
2040-09-17

AI Technical Summary

Technical Problem

The existing technology has problems in identifying website homepages that label data sets are time-consuming and labor-intensive, difficult to deal with nested URLs, and huge resources are consumed in extracting page content, resulting in insufficient recognition speed and accuracy.

Method used

Using a URL feature-based recognition method, we solve the problem of multi-layer nesting affecting web page homepage recognition by stripping the nested URL, using regular expressions to match URL domain names and keywords, and setting the length threshold after the "/" character.

Benefits of technology

It improves the speed and accuracy of website homepage recognition, reduces the need for manual labeling, reduces the false alarm rate, and saves network resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114201698B_ABST
    Figure CN114201698B_ABST
Patent Text Reader

Abstract

The present invention provides a method for identifying the home page of a website based on URL features and an electronic device, including removing the http: / / character or the https: / / character at the head of the URL to be identified, and obtaining a temporary variable t1 containing the http: / / character or the https: / / character; splitting the temporary variable t1 according to the " / " character and performing a validity judgment; if it cannot be split or can only be split into two parts and the second part is empty, then judge whether the temporary variable t1 contains a second-level, third-level or fourth-level domain name; if it can only be split into two parts, the second part is not empty and the length of the second part is less than the first threshold, then judge whether the second part contains a specific character; if the temporary variable t1 contains a second-level, third-level or fourth-level domain name or the second part contains a specific character, then judge that the URL to be identified is the home page URL. The present invention does not require training a classifier, manually annotating a large number of data sets and analyzing the URL page content, solves the situation where nested URLs cannot be identified by semantics, reduces the false alarm rate, saves manpower and network resources, and improves the recognition speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network communication technology, and in particular to a method and an electronic device for identifying a website homepage based on URL features. Background Art

[0002] With the continuous development of WEB technology, cloud computing and other technologies, the presentation of WEB homepages is also constantly changing. Many researchers conduct research based on specific tags and content of website homepages for the purpose of identifying fast-changing services, identifying malicious websites, discovering homologous IPs, collecting web page information, and studying web page layouts. For example, patent CN103812673A calculates the similarity of website homepages and content sampling statistics to screen mirror websites and quasi-mirror websites, avoid collecting identical content, save network resources and local resources, and improve service quality and efficiency.

[0003] Since there is no fixed naming method for the URL of the website homepage, and manual identification is time-consuming and labor-intensive, this has created a bottleneck for research based on the website homepage. Therefore, there is an urgent need for a method to automatically identify the website homepage to significantly improve the recognition speed.

[0004] Existing website identification work is mainly divided into two categories. One is based on URL string splitting and training using machine learning methods. For example, patent CN110855635A proposes a method for identifying malicious websites based on URLs, that is, identifying character combinations split from URLs through classification models; for example, patent CN101692639A uses the semantics of the URL main domain name and the structure of the entire URL to determine whether it is a pornographic site. The second category is to visit the URL and analyze and study it based on the content of the web page. For example, patent CN102332028A extracts the visual structure of the website page, TML tag information, link information, and text information to determine whether it is a bad website; for example, patent CN111428180A extracts binary word vectors from web pages, and uses semantic local sensitive hashing to represent web page content and identify similar web pages.

[0005] The above-mentioned method based on URL recognition has the following problems in home page recognition:

[0006] 1. Labeling datasets is time-consuming and laborious. In order to improve the accuracy of the classification model, a large number of URLs need to be manually labeled. During the labeling process, if there is no public dataset, each URL needs to be manually accessed for verification, which is very time-consuming.

[0007] 2. The URLs of many home pages are displayed in the form of nested URLs for reasons such as marking the source of access. It is difficult to deal with this situation solely through semantic analysis, which may cause serious misjudgments.

[0008] 3. Extracting page content consumes huge resources. Accessing the URL to extract pictures, links, text, page structure and other methods from the web page for analysis is very time-consuming and will consume a lot of network resources. This method is simply unable to cope with massive amounts of data.

[0009] To sum up, it is necessary to design a website homepage identification method that solves the above problems in order to meet the needs of identifying homologous IPs, collecting information, and studying page layout. Summary of the invention

[0010] To solve the above problems, the present invention provides a website homepage identification method and electronic device based on URL features, which identify the stripped nested URLs, use regular expressions to match URL domain names and match certain iconic keywords, and solve the problem of multi-layer nesting affecting web page homepage identification by setting a length threshold after the " / " character, thereby improving the speed and accuracy of website homepage identification.

[0011] The technical solution adopted by the present invention is as follows:

[0012] A method for identifying a website homepage based on URL features, the steps of which include:

[0013] 1) Remove the http: / / character or https: / / character from the header of the URL to be identified, and obtain a temporary variable t1 containing the http: / / character or https: / / character;

[0014] 2) Split the temporary variable t1 according to the " / " character and make a validity judgment;

[0015] 3) If it cannot be split or can only be split into two parts and the second part is empty, then determine whether the temporary variable t1 contains a second-level, third-level or fourth-level domain name; if it can only be split into two parts, the second part is not empty and the length of the second part is less than the first threshold, then determine whether the second part contains specific characters;

[0016] 4) If the temporary variable t1 contains a second-level, third-level or fourth-level domain name or the second part contains specific characters, it is determined that the URL to be identified is the homepage URL.

[0017] Furthermore, the method for determining whether the temporary variable t1 contains a second-level, third-level or fourth-level domain name includes: using a regular expression method.

[0018] Further, the first threshold is obtained by the following steps:

[0019] 1) Obtain n nested URL corresponding sample temporary variables t1 that can only be split into two parts and the second part is not empty, and calculate the total length N1 of the second part.

[0020] 2) Calculate the first threshold L1 = N1 / n.

[0021] Furthermore, determine whether the second part contains specific characters through the following strategy:

[0022] 1) Take the second part as the temporary variable t2;

[0023] 2) Determine whether the temporary variable t2 contains characters indicating the URL source;

[0024] 3) Determine whether the temporary variable t2 contains characters identifying the home page;

[0025] 4) If the length of the temporary variable t2 is less than the second threshold, determine whether the temporary variable t2 contains a web page suffix.

[0026] Furthermore, the method for determining whether the temporary variable t2 contains characters indicating the URL source includes: using the regular expression method; the characters indicating the URL source include: the src character or the from character.

[0027] Furthermore, the characters containing the identifier for the home page include: the index character or the homepage character.

[0028] Furthermore, the web page suffix includes: the html character or the jsp character.

[0029] Furthermore, obtain the second threshold through the following steps:

[0030] 1) Obtain m corresponding sample temporary variables t1 of URLs that can only be split into two parts and the second part is not empty, and calculate the total length N2 of the second part;

[0031] 2) Calculate the second threshold L2 = N2 / m.

[0032] A storage medium stores a computer program, wherein the computer program is configured to execute the above-mentioned method when running.

[0033] An electronic device includes a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer to execute the above-mentioned method.

[0034] Compared with the prior art, the present invention has the following advantages:

[0035] 1. There is no need to train a classifier, no need to manually label a large number of data sets, and only through regular expressions can the recognition be completed, saving manpower.

[0036] 2. Strip nested URLs, solve the situation where nested URLs cannot be recognized semantically, and reduce the false alarm rate.

[0037] 3. It is not necessary to collect and analyze the structure, layout, pictures, titles, etc. in the URL page, which greatly improves the recognition speed and saves network resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0039] The present invention will be further described in detail below through specific embodiments and the drawings.

[0040] In order to meet the needs of research such as identifying malicious websites, discovering homologous IPs, and collecting web page information, and at the same time to save network resources and improve the recognition speed, the present invention designs a method for identifying the home page of a website based on the URL, which mainly includes two parts: nested URL parsing and rule matching. The process is as Figure 1 shown

[0041] The specific steps are as follows:

[0042] Step 101: Remove the characters "http: / / " or "https: / / " at the head of the URL to be recognized.

[0043] Step 102: Determine whether there are still characters "http: / / " or "https: / / " in the URL to be recognized. If so, create a temporary variable t1, and t1 stores all the content after the characters "http: / / " or "https: / / ", and assign the temporary variable t1 to the URL to be recognized.

[0044] Step 103: Split the URL to be recognized according to the character " / ", and perform a validity judgment. If it cannot be split, jump to Step 104; if it is split according to the character " / ", and it can only be split into two parts and the second part is empty, remove the " / " character in the URL and jump to Step 104; if it is split according to the character " / ", and it can only be split into two parts, the second part is not empty, and the length of the second part is less than the threshold L1, assign the second part after splitting to the temporary variable t2 and jump to Step 105; otherwise, it belongs to the following two cases: 1) It can only be split into two parts, the second part is not empty, and the length of the second part is greater than or equal to the threshold L1; 2) It can be split into three parts or more. In these two cases, it is output that it is not the home page and the judgment ends.

[0045] Step 104: Use a regular expression to determine whether the URL to be recognized is a second-level, third-level, or fourth-level domain name. If so, output that it is the home page and end the judgment; otherwise, output that it is not the home page. The method of regular expression judgment is to construct a string that conforms to the format of the second-level, third-level, and fourth-level domain names, and then perform pattern matching.

[0046] Step 105: Use regular expressions to determine whether the temporary variable t2 contains characters such as src and from that indicate the URL source. If so, output that it is the home page and end the determination; otherwise, jump to Step 106. The method of regular expression determination is to perform string matching.

[0047] Step 106: Determine whether the temporary variable t2 contains characters such as index and homepage that identify the home page. If so, output that it is the home page and end the determination; otherwise, jump to Step 107.

[0048] Step 107: Determine whether the length of the temporary variable t2 is less than the threshold L2. If so, jump to Step 108; otherwise, output that it is not the home page and end the determination.

[0049] Step 108: Determine whether the temporary variable t2 contains web page suffixes such as "html" and "jsp". If so, output that it is the home page and end the determination; otherwise, output that it is not the home page and end the determination.

[0050] The threshold is calculated as follows:

[0051] The calculation of L1 is as follows: Manually verify n URLs that meet the conditions of "after splitting according to the character ' / ', it can only be split into two parts, and the second part is not empty" and are home pages. Calculate the total length of the second part of these n URLs, denoted as N1, and L1 = N1 / n.

[0052] The calculation of L2 is as follows: Manually verify m URLs that can extract t2 and are home pages. Calculate the total length of the t2 part of these m URLs, denoted as N2, and L2 = N2 / m.

[0053] Given that the url formats are similar, especially for URLs in the same field, taking n and m as 100 can meet the requirements.

[0054] Experimental data

[0055] The discovery of many phishing websites, website style planning, etc. are all based on the research of the website home page. For example, identifying phishing websites based on the logo and host name of the website home page. First, confirm whether the url is the home page, and then extract the logo, which can shorten the discovery time of phishing websites.

[0056] To reflect the technical advantages of this application, this application is compared with other methods that can be used to identify the website home page as follows:

[0057] 1) Compare with the method based on website content (CN102332028A):

[0058] The top 1170 URLs in the recognized Alexa dataset were selected, accessed, and the web page titles, structures, images, and links were extracted. Without further identifying the home page, the method of identifying based on website content took 6 hours. However, this patent does not require online identification. In an environment with macOS 10.15.5, 16GB of memory, and a 2.6GHz six-core Intel Core i7 processor, using this patent to identify the home page only takes 3 seconds. Therefore, the identification time can be reduced by at least 99.9%.

[0059] 2) Comparison with the method based on semantic structure (CN101692639A):

[0060] We collected 200,000 real web surfing records of laboratory personnel. Compared with the method based on semantic structure, the present invention can increase the recall rate by at least 2.22%.

[0061] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Those of ordinary skill in the art can modify or equivalently replace the technical solutions of the present invention without departing from the spirit and scope of the present invention. The protection scope of the present invention shall be subject to what is described in the claims.

Claims

1. A method for identifying the home page of a website based on URL features, the steps of which include: 1) Remove the "http: / / " character or "https: / / " character at the beginning of the URL to be identified, and obtain a temporary variable t1 containing the "http: / / " character or "https: / / " character; 2) Split the temporary variable t1 according to the " / " character and perform a validity check; 3) If it cannot be split or can only be split into two parts and the second part is empty, determine whether the temporary variable t1 contains a second-level, third-level, or fourth-level domain name; if it can only be split into two parts, the second part is not empty, and the length of the second part is less than the first threshold, determine whether the second part contains a specific character; 4) If the temporary variable t1 contains a second-level, third-level, or fourth-level domain name or the second part contains a specific character, determine that the URL to be identified is the home page URL.

2. The method according to claim 1, characterized in that The method for determining whether the temporary variable t1 contains a second-level, third-level, or fourth-level domain name includes: using the regular expression method.

3. The method according to claim 1, characterized in that, The first threshold is obtained through the following steps: 1) Obtain n corresponding sample temporary variables t1 of nested URLs that can only be split into two parts and the second part is not empty, and calculate the total length N1 of the second part; 2) Calculate the first threshold L1 = N1 / n.

4. The method according to claim 1, wherein The following strategy is used to determine whether the second part contains a specific character: 1) Use the second part as the temporary variable t2; 2) Determine whether the temporary variable t2 contains a character indicating the source of the URL; 3) Determine whether the temporary variable t2 contains a character identifying the home page; 4) If the length of the temporary variable t2 is less than the second threshold, determine whether the temporary variable t2 contains a web page suffix.

5. The method according to claim 4, wherein The method for determining whether the temporary variable t2 contains a character indicating the source of the URL includes: using the regular expression method; the characters indicating the source of the URL include: the "src" character or the "from" character.

6. The method according to claim 4, wherein The characters containing the identifier for the home page include: the "index" character or the "homepage" character.

7. The method according to claim 4, characterized in that The web page suffix includes: the "html" character or the "jsp" character.

8. The method according to claim 4, wherein The second threshold is obtained through the following steps: 1) Obtain m corresponding sample temporary variables t1 of URLs that can only be split into two parts and the second part is not empty, and calculate the total length N2 of the second part; 2) Calculate the second threshold L2 = N2 / m.

9. A storage medium, in which a computer program is stored, wherein, The computer program is set to execute the method described in any one of claims 1-8 when running.

10. An electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is set to run the computer program to execute the method described in any one of claims 1-8.

Citation Information

Patent Citations

  • Bad webpage recognition method based on URL

    CN101692639A

  • Webpage-oriented unhealthy Web content identifying method

    CN102332028A

  • Method for automatically recognizing multiple IP changes in website

    CN103812673A

  • Website top page presumption device, top page presumption method, program for this method, recording medium with the program recorded thereon

    JP2003186731A

  • Method And System For Mapping Domain Prefixes To Qualified URLs

    US20130067115A1