Website type identification method and device, equipment and storage medium

By grading the URL access records of the target website, building a tree-like directory structure and generating fingerprints, the problem of low accuracy of website type recognition is solved by simplifying the DOM structure, and achieving higher recognition accuracy and recognition rate.

CN120336654APending Publication Date: 2025-07-18TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410071172.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, when generating website fingerprints based on the DOM structure of the website, the accuracy of website type recognition for websites with a streamlined structure is low.

Method used

By obtaining the URL access record of the target website, performing directory grading, building a tree-like directory structure of the target website, and generating a website directory structure fingerprint based on the directory structure and weights, using this fingerprint to identify the website type.

Benefits of technology

It improves the accuracy and recognition rate of website type recognition, especially for websites with a streamlined structure, which can achieve better recognition results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336654A_ABST
    Figure CN120336654A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a website type identification method and device, equipment and a storage medium, and relates to the technical field of networks. Comprising the following steps: acquiring a URL access record of a target website; directory grading is conducted on URL links contained in the URL access record, a target website directory structure of the target website and directory weights of all levels of directories are obtained, and the target website directory structure is a tree structure formed by all levels of directories of the target website; generating a target website directory structure fingerprint of the target website based on the target website directory structure and the directory weight of each hierarchical directory; under the condition that the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is larger than a similarity threshold value, it is determined that the target website belongs to the target type, and the candidate website directory structure fingerprint is the website directory structure fingerprint of the known target type website. By adopting the scheme provided by the embodiment of the invention, the recognition rate and the recognition accuracy of the website type can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of network technologies, and particularly to a method, apparatus, device, and storage medium for identifying website types. Background Art

[0002] The Internet is filled with a large number of various websites, and users can obtain different types of services by accessing different types of websites.

[0003] In related technologies, by obtaining the Document Object Model (DOM) structure of a website, generating a website fingerprint based on the DOM structure, and using the similarity between the website fingerprint and the website fingerprints of known type websites, website type identification is achieved.

[0004] However, due to the relatively simple DOM structure of some websites, the accuracy of website type identification using website fingerprints is relatively low. Summary of the Invention

[0005] Embodiments of the present application provide a method, apparatus, device, and storage medium for identifying website types. The technical solutions are as follows:

[0006] On the one hand, embodiments of the present application provide a method for identifying website types, the method including:

[0007] Obtaining the Uniform Resource Locator (URL) access record of a target website, where the URL access record includes the URL links generated by accessing the target website;

[0008] Performing directory classification on the URL links included in the URL access record to obtain the target website directory structure of the target website and the directory weights of each level of directories, where the target website directory structure is a tree structure formed by each level of directories of the target website;

[0009] Generating a target website directory structure fingerprint of the target website based on the target website directory structure and the directory weights of each level of directories;

[0010] When the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than a similarity threshold, determining that the target website belongs to a target type, where the candidate website directory structure fingerprint is the website directory structure fingerprint of a known target type website.

[0011] On the other hand, embodiments of the present application provide an apparatus for identifying website types, the apparatus including:

[0012] A record acquisition module, configured to acquire URL access records of a target website, where the URL access records include URL links generated by accessing the target website;

[0013] A directory grading module, configured to grade the URL links included in the URL access records to obtain a target website directory structure of the target website and directory weights of each level directory, where the target website directory structure is a tree structure formed by each level directory of the target website;

[0014] A fingerprint generation module, configured to generate a target website directory structure fingerprint of the target website based on the target website directory structure and the directory weights of each level directory;

[0015] A determination module, configured to determine that the target website belongs to a target type when the similarity between the target website directory structure fingerprint and a candidate website directory structure fingerprint is greater than a similarity threshold, where the candidate website directory structure fingerprint is a website directory structure fingerprint of a known target type website.

[0016] On the other hand, an embodiment of the present application provides a computer device, where the computer device includes a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the above-mentioned website type recognition method.

[0017] On the other hand, an embodiment of the present application provides a computer-readable storage medium, where at least one instruction is stored in the readable storage medium, and the at least one instruction is loaded and executed by a processor to implement the website type recognition method as described in the above aspect.

[0018] On the other hand, an embodiment of the present application provides a computer program product, where the computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the website type recognition method provided in the above aspect.

[0019] Compared with the DOM structure, since the website directory structure can contain richer website structure information, in the embodiments of the present application, by obtaining a large number of URL links generated by accessing the target website and performing directory grading on the URL links, a tree structure formed by each level of directories in the target website and the directory weights of each level of directories are obtained, and then a website directory structure fingerprint is generated based on the website directory structure and the directory weights, and finally the website directory structure fingerprint is used to identify the website type of the target website; in addition, since the website directory structure is generated based on the URL access records, a better type recognition effect can also be achieved for websites with a streamlined website structure, which helps to improve the recognition rate and recognition accuracy of the website type. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 FIG. shows a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application;

[0021] Figure 2 FIG. shows a flowchart of a method for identifying a website type provided by an exemplary embodiment of the present application;

[0022] Figure 3 FIG. shows an implementation schematic diagram of a relationship chain diffusion mining process provided by an exemplary embodiment of the present application;

[0023] Figure 4 FIG. shows a flowchart of a method for identifying a website type provided by another exemplary embodiment of the present application;

[0024] Figure 5 is a schematic diagram of a URL access record shown in an exemplary embodiment of the present application;

[0025] Figure 6 is a flowchart of a link cleaning and word segmentation process shown in an exemplary embodiment of the present application;

[0026] Figure 7 is a schematic diagram of a URL link after data cleaning shown in an exemplary embodiment of the present application;

[0027] Figure 8 is a schematic diagram of a word segmentation result of a URL link shown in an exemplary embodiment of the present application;

[0028] Figure 9 is a flowchart of a directory grading process shown in an exemplary embodiment of the present application;

[0029] Figure 10 is a schematic diagram of a directory grading result corresponding to a URL link shown in an exemplary embodiment of the present application;

[0030] Figure 11It is a schematic diagram of the target website directory structure shown in an exemplary embodiment of the present application;

[0031] Figure 12 It is a schematic diagram of the directory weights corresponding to each hierarchical directory shown in an exemplary embodiment of the present application;

[0032] Figure 13 It is a schematic diagram of the binary conversion result of the directory shown in an exemplary embodiment of the present application;

[0033] Figure 14 It is a schematic diagram of the directory vector corresponding to the directory and the website vector corresponding to the target website shown in an exemplary embodiment of the present application;

[0034] Figure 15 It shows a flowchart of a method for identifying website types provided in another exemplary embodiment of the present application;

[0035] Figure 16 It is a structural block diagram of a device for identifying website types provided in another exemplary embodiment of the present application;

[0036] Figure 17 It shows a schematic diagram of the structure of a computer device provided in an exemplary embodiment of the present application. Detailed implementation manners

[0037] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0038] In the related art, the process of generating a website fingerprint based on the DOM structure of a website may include the following steps:

[0039] 1. Extract the Hyper Text Markup Language (HTML) file of the target website;

[0040] 2. Parse the DOM tree in the HTML file to obtain a DOM sequence, and the DOM sequence is composed of the node order in the DOM tree;

[0041] 3. Perform a hash operation on the DOM sequence to obtain the website fingerprint of the target website.

[0042] It can be seen from the above website fingerprint generation process that when the DOM structure is rich, the reliability of website type identification based on the website fingerprint is relatively high; on the contrary, the simpler the DOM structure, the lower the reliability of website type identification based on the website fingerprint.

[0043] In order to improve the recognition rate and accuracy of website type recognition, an embodiment of the present application provides a solution for generating a website directory structure fingerprint based on a URL access link and using the website directory structure fingerprint for website type recognition.

[0044] Please refer to Figure 1 , which shows a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application. The implementation environment includes a terminal 110 and a server 120.

[0045] The terminal 110 is an electronic device with the function of accessing Internet websites. The electronic device can be a smart phone, a tablet computer, a personal computer, a wearable device, a vehicle-mounted terminal, etc. In some embodiments, the terminal 110 can access websites through a native browser, a third-party browser, or a browser kernel built into an application.

[0046] The server 120 is a background server that provides website type recognition functions. The server 120 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0047] In some embodiments, the server 120 is the background server of the browser (or the browser kernel in the application) in the terminal 110. When using the browser in the terminal 110 to access a website, the browser sends the URL link to the server 120. In other embodiments, a security application is installed in the terminal 110, and the server 120 is the background server of the security application. The security application in the terminal 110 can send the URL link used when the terminal 110 accesses a website to the server 120.

[0048] It should be noted that during the process of collecting relevant data of users (such as URL access records) in this application, a prompt interface, pop-up window or voice prompt information can be displayed. The prompt interface, pop-up window or voice prompt information is used to prompt the user that their relevant data is being collected currently, so that this application only starts to execute the relevant steps of obtaining the user's relevant data after obtaining the confirmation operation issued by the user for the prompt interface or pop-up window. Otherwise (that is, when the confirmation operation issued by the user for the prompt interface or pop-up window is not obtained), the relevant steps of obtaining the user's relevant data are ended, that is, the relevant data of the user is not obtained. In other words, the information (including but not limited to user device information, URL access records), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions.

[0049] In a possible implementation manner, when it is necessary to identify the website type of a target website, the server 120 obtains the URL access record 121 related to the target website, constructs the website directory structure of the target website based on the URL link in the URL access record 121, and generates the website directory structure fingerprint 122 of the target website based on the website directory structure and the directory weights of each level directory in the structure. A website directory structure fingerprint library 123 is set in the server 120, and candidate website directory structure fingerprints of known target type websites are stored in the fingerprint library. When identifying the website type of the target website, the server 120 performs fingerprint matching between the website directory structure fingerprint 122 of the target website and the candidate website directory structure fingerprints in the website directory structure fingerprint library 123, and obtains the website type recognition result 124 of the target website based on the fingerprint matching result. Optionally, when it is identified that the target website belongs to the target type, the server 120 updates the website directory structure fingerprint 122 of the target website to the website directory structure fingerprint library 123.

[0050] In addition to implementing single-point recognition based on the website directory structure fingerprint, the server 120 can also perform website type association recognition based on the reference relationship, jump relationship and aggregation relationship between websites, and further update the website directory structure fingerprint library 123 based on the recognition result.

[0051] In some embodiments, when identifying a specific type of website, the server 120 can perform access prompts and interceptions on the URL access related to the specific type of website. For example, the server 120 can send the root URL of the identified risky website to the terminal 110, and the terminal 110 identifies the risky website access behavior of its own end based on the root URL and performs prompts or interceptions on the access behavior.

[0052] It should be noted that, in addition to the server 120 performing website type recognition, in other possible implementation manners, the terminal 110 can also perform website type recognition locally, or the terminal 110 and the server 120 cooperate to implement website type recognition (for example, the server 120 distributes website type recognition tasks to different terminals in a distributed manner and collects the recognition results of the terminals).

[0053] For the convenience of description, in the following embodiments, the example of the computer device executing the website type recognition method is used for illustration.

[0054] Figure 2 The flowchart of the website type recognition method provided by an exemplary embodiment of the present application is shown. This embodiment is described by taking this method being used in a computer device as an example, and the method includes the following steps.

[0055] Step 201, obtain the URL access record of the target website, where the URL access record includes the URL link generated by accessing the target website.

[0056] In this embodiment, the target website is the website for which website type recognition is required. Among them, the website type can be divided according to the website usage, such as shopping websites, social websites, game websites, news websites, video websites, etc.; or, the website type can be divided according to attributes, such as secure websites and risky websites divided according to security attributes. The specific division method of the website type in the embodiments of the present application is not limited.

[0057] In a possible implementation manner, the computer device filters out the URL link generated by accessing the target website from a large number of URL links according to the root URL of the target website, and obtains the URL access record of the target website.

[0058] For example, when the root URL of the target website is http: / / www.xxx.com, the URL access record obtained by the computer device may include http: / / www.xxx.com / topic / 245836.html, "http: / / www.xxx.com / ?entryScene=zhida_05_001&jump_from=1_13_18_00", "http: / / www.xxx.com / ?jump_from=1_05_37_01", etc.

[0059] In some embodiments, in order to improve the recognition accuracy, the data volume of the URL access record obtained by the computer device needs to be large enough to cover as many terminals as possible.

[0060] Optionally, when the data volume of the URL access records of the target website is less than the data volume threshold, the computer device can generate a large number of URL access records by randomly simulating access to the target website to achieve data augmentation.

[0061] Step 202: Perform directory classification on the URL links included in the URL access records to obtain the target website directory structure of the target website and the directory weights of each level of directories. The target website directory structure is a tree structure formed by each level of directories of the target website.

[0062] Since different URL links are used to access the content under different levels of directories in the target website, by identifying the directory information included in the URL links and performing directory classification on each URL link based on this directory information, the discretized directory classification information of the target website can be obtained.

[0063] Furthermore, based on the hierarchical relationship between directories represented by the directory classification information, the computer device can construct the target website directory structure of the target website. Among them, the website directory structure is a tree structure composed of directory nodes and can be used to represent the hierarchical relationship between directories represented by different directory nodes.

[0064] It should be noted that different from the DOM structure generated based on the HTML file of a single web page, the website directory structure is generated based on a large number of URL access records, involving different levels of web pages of the same website. Therefore, it can contain deeper and richer website structure information. And even if the DOM structure of the website is simplified, as long as there are access records to the website, an effective website directory structure fingerprint can be constructed.

[0065] To improve the accuracy of the website directory structure fingerprint constructed subsequently, for each level of directory in the website directory structure, the computer device also determines the directory weight of the directory.

[0066] Optionally, the computer device determines the directory weights corresponding to all levels of directories respectively, or determines the directory weights corresponding to some levels of directories respectively.

[0067] In some embodiments, the directory weight is determined based on the access conditions of each level of directory in the URL access records, so that the access characteristics of the website are incorporated into the website directory structure fingerprint generated subsequently.

[0068] Optionally, the directory weight is determined based on the access frequency of the directory, and the higher the access frequency, the higher the directory weight of the directory.

[0069] Step 203: Generate the target website directory structure fingerprint of the target website based on the target website directory structure and the directory weights of each level of directories.

[0070] Further, the computer device generates a target website directory structure fingerprint based on the target website directory structure and directory weight of the target website, so that the target website directory structure fingerprint can not only reflect the directory structure characteristics of the target website, but also reflect the access characteristics of the target website.

[0071] In some embodiments, the website directory structure fingerprint is a binary fingerprint of a fixed length (the value of each position is 0 or 1). For example, the network directory structure fingerprint is a 64-bit binary fingerprint.

[0072] In other embodiments, the website directory structure fingerprint can also be a real number fingerprint of a fixed length (the data of each position is a real number).

[0073] It should be noted that the length of the website directory structure fingerprint can be adjusted according to the recognition accuracy requirement and recognition rate requirement of the website type, and the embodiments of the present application do not limit the fingerprint length.

[0074] Step 204, in the case where the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than the similarity threshold, it is determined that the target website belongs to the target type, and the candidate website directory structure fingerprint is the website directory structure fingerprint of a known target type website.

[0075] In this embodiment, the computer device is provided with a website directory structure fingerprint library. The website directory structure fingerprint library contains candidate website directory structure fingerprints of known target type websites. Among them, the known target type websites can be identified by the method provided in the embodiments of the present application, or can be identified by other methods (such as manual identification); and, the generation method of the candidate website directory structure fingerprint is the same as the generation method of the above-mentioned target website directory structure fingerprint.

[0076] Optionally, the computer device can set multiple website target structure fingerprint libraries, which respectively correspond to different types of websites.

[0077] In some embodiments, the computer device calculates the similarity between the target website directory structure fingerprint and each candidate website directory structure fingerprint. In the case where the similarity is greater than the similarity threshold, the computer device determines that the target website belongs to the target type, and in the case where the similarity is less than the similarity threshold, the computer device determines that the target website does not belong to the target type.

[0078] Optionally, the computer device can use methods such as Hamming distance, Euclidean distance, Mahalanobis distance, cosine distance, etc. to determine the similarity of the website directory structure fingerprint. Correspondingly, the similarity threshold can be a Hamming distance threshold, Euclidean distance threshold, Mahalanobis distance threshold, cosine distance, etc., and the embodiments of the present application do not limit this.

[0079] Optionally, when it is recognized that the target website belongs to the target type, the computer device adds the target website directory structure fingerprint to the website directory structure fingerprint library for use in identifying the website types of other websites.

[0080] In summary, compared with the DOM structure, since the website directory structure can contain richer website structure information, in the embodiments of the present application, by obtaining a large number of URL links generated by accessing the target website and performing directory grading on the URL links, a tree structure formed by each level of directories in the target website and the directory weights of each level of directories are obtained, and then a website directory structure fingerprint is generated based on the website directory structure and the directory weights, and finally the website directory structure fingerprint is used to identify the website type of the target website; in addition, since the website directory structure is generated based on the URL access records, a good type recognition effect can also be achieved for websites with a streamlined website structure, which helps to improve the recognition rate and accuracy of website types.

[0081] After realizing single-point website type recognition through the above solution, in order to further improve the recognition rate of website types, the computer device determines other websites that have an associated relationship with the target website through the method of relationship chain diffusion mining.

[0082] In a possible implementation manner, the computer device determines that a website that has a reference relationship or a jump relationship with the target website belongs to the target type based on the website URL relationship chain.

[0083] In some embodiments, the website URL relationship chain can be generated through big data mining based on the historical website access records reported by the terminal.

[0084] Optionally, the website URL relationship chain is a graph with website URLs as nodes and website URL relationships between the website URLs as edges. Among them, the website URL relationships between the website URLs can include reference relationships and jump relationships.

[0085] In some embodiments, the computer device takes the target website as the center and performs reference relationship diffusion through the website URL relationship chain to determine other target type websites that have a reference relationship with the target website. Among them, performing reference relationship diffusion can determine the reference target website and other websites referenced by the target website.

[0086] Optionally, when the computer device performs reference relationship diffusion, the reference relationship layer number or the cumulative weight of the reference relationship is used as the stop diffusion condition. For example, the computer device takes the target website as the center and mines websites with a reference relationship layer number less than 4 layers with the target website, or the computer device takes the target website as the center and mines websites with a cumulative weight of the reference relationship greater than 0.5 with the target website. Among them, the cumulative weight of the reference relationship is the product of the weights of the reference relationships represented by the edges between the nodes.

[0087] In some other embodiments, the computer device centers around the target website and spreads the jump relationship through the URL relationship chain to determine other target-type websites that have a jump relationship with the target website. Among them, spreading the jump relationship can determine the websites that jump to the target website and the other websites that jump from the target website.

[0088] Optionally, when the computer device spreads the jump relationship, the number of jumps or the cumulative jump probability is used as the stop condition for spreading. For example, the computer device centers around the target website and mines websites with less than 3 jumps between the target website, or the computer device centers around the target website and mines websites with a cumulative jump probability greater than 0.5 between the target website. Among them, the cumulative jump probability is the product of the jump probabilities represented by the edges between nodes.

[0089] In another possible implementation, the computer device determines that websites belonging to the same aggregation medium as the target website belong to the target type, where the aggregation medium includes at least one of Internet Protocol (IP), Internet Content Provider (ICP), whois, and network devices.

[0090] Optionally, after determining and mining other websites through the above method, the computer device can generate the website directory structure fingerprint of other websites and further verify whether other websites belong to the target type based on this fingerprint.

[0091] In a schematic example, as Figure 3 shown, when it is recognized that the target website belongs to the target type, the computer device determines websites A and B that have a reference relationship with the target website through reference relationship spreading, determines websites C and D that have a jump relationship with the target website through jump relationship spreading, and determines websites E and F that belong to the same aggregation medium as the target website through aggregation relationship spreading. Finally, the computer device determines that the target website and websites A, B, C, D, E, and F all belong to the target type.

[0092] In a possible application scenario, the website type recognition method provided by the embodiments of the present application can be used to identify risky websites. When identifying risky websites, the computer device needs to first generate the website directory structure fingerprint of known risky websites through the website directory structure fingerprint generation method provided by the embodiments of the present application, so as to determine the target website with a similarity greater than the similarity threshold to the directory structure fingerprint of the known risky website as a risky website. With this solution, even if someone maliciously simplifies the DOM structure of a risky website, the risky website can be accurately identified based on the directory structure fingerprint.

[0093] Further, in the case where the target website is identified as a risky website, the computer device may discover other risky websites based on the reference relationship, jump relationship, or aggregation relationship between the target website and other websites.

[0094] Figure 4 FIG. 4 shows a flowchart of a method for identifying website types provided by another exemplary embodiment of the present application. This embodiment is described by taking the method being used in a computer device as an example. The method includes the following steps.

[0095] Step 401, obtain the URL access record of the target website. The URL access record includes the URL links generated by accessing the target website.

[0096] The implementation manner of this step may refer to the above step 201, and this embodiment will not elaborate here.

[0097] In a schematic example, the computer device obtains the URL access record of the target website "www.xxx.com" as Figure 5 shown.

[0098] Step 402, perform data cleaning and word segmentation processing on the URL links to obtain the word segmentation result of the URL links.

[0099] Since there are a large number of symbols, meaningless content, and duplicate content in the original URL links, directly performing directory classification on the original URL links may cause inaccurate classification results and overly complex classification results. To facilitate subsequent directory classification, the computer device first needs to perform data cleaning on the URL links and perform word segmentation processing on the URL links after data cleaning to obtain the word segmentation result of the URL links. Subsequently, the computer device performs directory classification based on the word segmentation result.

[0100] In a possible implementation manner, as Figure 6 shown, the computer device performing data cleaning and word segmentation processing may include the following steps.

[0101] Step 402A, based on the string filtering table, perform string filtering on the URL links. The string filtering table contains URL common element strings and the root website address of the target website.

[0102] Since there are usually a large amount of common content in the URL links, such as "http: / / ", "https: / / ", and such common content does not belong to the unique directory structure of the website. Therefore, during data cleaning, the computer device needs to identify the URL common element strings in the URL links based on the string filtering table and filter out the URL common element strings.

[0103] Among them, the URL general element string may include "http: / / ", "https: / / ", etc., and the embodiments of the present application do not limit this.

[0104] In addition, since the URL links generated by accessing websites will all contain the root website address of the website, in order to simplify the website directory structure, the computer device also needs to filter the root website address of the target website in the URL link.

[0105] In a schematic example, when the target website is www.xxx.com, for the URL link http: / / www.xxx.com / topic / 245836.html, the computer device filters its URL general element string and the root website address, and the filtered URL link is / topic / 245836.html.

[0106] Step 402B, replace the string of the target type in the filtered URL link with a substitute word.

[0107] Generally, the URL link will contain some strings without specific meanings, and such strings usually include numerical strings and alphanumeric mixed strings.

[0108] If the directory classification is directly based on the URL link containing such strings, it will cause the website directory structure to be too large (manifested as too many child nodes).

[0109] For example, the URL links generated by accessing different pages under the same website include http: / / www.xxx.com / topic / 245836.html and http: / / www.xxx.com / topic / 149367.html, and these two URL links are respectively used to access different pages under topic. If the directory classification is directly performed on this URL link, there will be two subdirectories, 245836.html and 149367.html, under the topic directory.

[0110] To avoid the above situation, the computer device identifies the string of the target type in the filtered URL link and replaces the string with a unified substitute word. Among them, different types of strings can correspond to different substitute words.

[0111] Optionally, for the numerical string in the URL link, the computer device uniformly replaces the numerical string with num; for the alphanumeric mixed string without actual meaning in the URL link, the computer device replaces the alphanumeric mixed string with mix.

[0112] In a schematic example, for the filtered URL links / topic / 245836.html and / topic / 149367.html, the computer device uniformly replaces the numeric type strings therein with num, and the URL links after data cleaning are all / topic / num.html.

[0113] It should be noted that the above target types are only for illustrative purposes, and the target types can be set according to requirements. The embodiments of the present application do not limit this.

[0114] In a schematic example, for Figure 5 the URL access records shown, after data cleaning, the URL links after data cleaning are as Figure 7 shown.

[0115] Step 402C: Using the target punctuation as the word segmentation position, perform word segmentation processing on the URL link after string replacement to obtain the word segmentation result of the URL link.

[0116] After completing data cleaning, the computer device performs word segmentation on the URL link to obtain multiple word segments corresponding to the URL link.

[0117] Usually, different directory levels are usually distinguished by characteristic punctuation. Therefore, in a possible implementation manner, the computer device can identify the target punctuation included in the URL link, and use the position of the target punctuation as the word segmentation position to perform word segmentation on the URL link.

[0118] Among them, the target punctuation may include predefined typical directory division punctuation. For example, the target punctuation may include " / ", "?", "=", "&", etc. The embodiments of the present application do not limit the specific type of the target punctuation.

[0119] For example, for the URL link / topic / num.html after data cleaning, the computer device performs word segmentation processing on it based on the target punctuation " / ", and the obtained word segmentation result includes two word segments: topic and num.html.

[0120] In a schematic example, for Figure 7 the URL link shown, perform word segmentation processing to obtain the word segmentation result of the URL link as Figure 8 shown.

[0121] Step 403: Based on the word segmentation results of each URL link, perform directory classification to obtain the link directory classification results of each URL link.

[0122] In a possible implementation, the computer device uses each word segment in the word segmentation result as a directory level to perform directory classification on the word segmentation result of the URL link, obtaining the link directory classification result corresponding to each URL link. Among them, the directory level corresponding to the word segment that is more forward in the word segmentation result is higher, and among adjacent word segments, the directory corresponding to the latter word segment is a sub-directory of the directory corresponding to the former word segment.

[0123] To enable the classification result to reflect the hierarchical relationship between different directories, when the computer device performs directory classification, it adds a corresponding hierarchical identifier to the directory and concatenates the upper-level directory and the current-level directory as the current-level link directory. As Figure 9 shown, this step may include the following sub-steps:

[0124] Step 403A, for the first word segment in the word segmentation result, concatenate the first-level identifier and the first word segment to obtain the first-level link directory.

[0125] For the first word segment in the word segmentation result, since there are no other word segments before the first word segment, the computer device concatenates the first word segment and the first-level identifier representing the first level to obtain the first-level link directory.

[0126] Among them, the hierarchical identifier can be represented by numbers. For example, the number 1 identifies the first-level directory, the number 2 identifies the second-level directory, and so on.

[0127] In a schematic example, for the word segmentation result "topic num.html", the computer device determines that the first-level link directory is 1topic.

[0128] Step 403B, for the nth word segment in the word segmentation result, concatenate the (n - 1)-level identifier, the (n - 1)th word segment, the n-level identifier, and the nth word segment to obtain the nth-level link directory, where n is an integer greater than or equal to 2.

[0129] For the second word segment and subsequent word segments, in addition to concatenating the word segment with the corresponding hierarchical identifier, the computer device also concatenates the previous word segment with the corresponding hierarchical identifier in the previous position to obtain the link directory corresponding to the current level. In other words, the second-level link directory is composed of the first-level and second-level directories, the third-level link directory is composed of the second-level and third-level directories, and so on.

[0130] In some embodiments, during the concatenation process, the concatenation result of the (n - 1)-level identifier and the (n - 1)th word segment is connected to the concatenation result of the n-level identifier and the nth word segment through a preset connection symbol. For example, the preset connection symbol can be "_" or "-", etc., and this embodiment does not limit this.

[0131] In a schematic example, for the word segmentation result "topic num.html", the computer device determines that the second-level link directory is 1topic_2num.html.

[0132] Step 403C: Determine each level of link directory as the link directory classification result of the URL link.

[0133] In a schematic example, based on Figure 8 the word segmentation result of the URL link shown, perform directory classification to obtain the directory classification results corresponding to each URL link as Figure 10 shown.

[0134] It should be noted that the computer device can also splice only based on the current word segmentation and the hierarchical identifier of its affiliated level as the link directory of the current level, and the embodiments of the present application do not make limitations.

[0135] Step 404: Based on the link directory classification results of each URL link, determine the target website directory structure of the target website and the directory weights of each level of directory.

[0136] In a possible implementation manner, after performing directory classification using the solution in the above steps, since the link directory of each level contains the information of the upper-level directory, the computer device can perform summary of the same directory for the link directory classification results of different URL links, and determine the target website directory structure of the target website based on the hierarchical relationship between different directories.

[0137] Since different URL links can cover the access situations of different levels of directories of the target website, as the URL access records are continuously enriched, the target website directory structure generated by the computer device will become more and more perfect and comprehensive.

[0138] In some embodiments, the computer device can continuously improve the website directory structure based on newly generated URL access records periodically.

[0139] In a schematic example, the computer device based on Figure 10 the link directory classification result shown, obtains the target website directory structure of the target website as Figure 11 shown.

[0140] In order to incorporate the access characteristics of different directories under the target website into the fingerprint, when determining the directory weights of each level of directory, the computer device counts the occurrence frequency of each level of directory in the link directory classification result, and thus determines the directory weights of each level of directory based on this occurrence frequency, where the directory weight has a positive correlation with the occurrence frequency.

[0141] In a possible implementation, the computer device traverses each hierarchical directory in the target website directory structure according to the breadth-first principle, and counts the total number of occurrences (the sum of the total number of accesses) of the hierarchical directory in the link directory classification result.

[0142] Optionally, the computer device determines the directory weight of the hierarchical directory as the occurrence frequency of the hierarchical directory.

[0143] Combined with Figure 10 the directory classification result shown, for the hierarchical directory 1topic, the computer device determines that the directory weight of this hierarchical directory is 15 + 8 + 14 + 50 = 87; for the hierarchical directory 1kj, the computer device determines that the directory weight of this hierarchical directory is 10 + 13 = 23.

[0144] Schematically, Figure 10 the directory weights corresponding to each hierarchical directory in Figure 12 are as shown.

[0145] To improve the accuracy of the determined weight, in addition to determining the weight based on the occurrence frequency, the computer device can also construct a correspondence between different vocabulary and word weights, and determine the word weight corresponding to the hierarchical directory from this correspondence, so as to fuse the occurrence frequency and the word weight to obtain the directory weight of the hierarchical directory.

[0146] Regarding the determination method of the correspondence between vocabulary and word weights, in some embodiments, the computer device can count the occurrence frequencies of hierarchical directories in the URLs of a large number of websites belonging to the target type and the URLs of websites not belonging to the target type (not limited to the URLs generated by accessing the target website), and determine the word weights of different hierarchical directories based on this occurrence frequency.

[0147] Optionally, determining the directory weight of the hierarchical directory may include the following steps:

[0148] 1. For the target hierarchical directory in the target website directory structure, determine the occurrence frequency of the target hierarchical directory in the link directory classification result.

[0149] 2. If the target hierarchical directory is included in the thesaurus, obtain the target word weight corresponding to the target hierarchical directory from the thesaurus; determine the directory weight of the target hierarchical directory based on the target word weight and the occurrence frequency.

[0150] The computer device determines whether there is a vocabulary corresponding to the target hierarchical directory in the thesaurus. If so, the computer device further obtains the target word weight corresponding to the target hierarchical directory. Optionally, the computer device performs a weighted process on the target word weight and the occurrence frequency to obtain the directory weight of the target hierarchical directory.

[0151] Among them, the weights corresponding to the target word weight and the occurrence frequency can be set in advance.

[0152] In some embodiments, before performing the weighting process, the computer device needs to normalize the occurrence frequency and the target word weight.

[0153] 3. In the case where the target hierarchical directory is not included in the thesaurus, determine the directory weight of the target hierarchical directory based on the occurrence frequency and the average word weight of the thesaurus.

[0154] If there is no vocabulary corresponding to the target hierarchical directory in the thesaurus, the computer device can obtain the average word weight of the vocabulary in the thesaurus, and perform a weighting process on the occurrence frequency and the average word weight to obtain the directory weight of the target hierarchical directory.

[0155] Step 405: Convert each hierarchical directory in the target website directory structure into a binary sequence, where the lengths of the binary sequences obtained by converting each hierarchical directory are the same.

[0156] In a possible implementation manner, the computer device can perform a hash process on each hierarchical directory to obtain a binary sequence corresponding to each hierarchical directory, and this binary sequence is a fixed-length 01 sequence.

[0157] For example, the computer device can convert each hierarchical directory into a 64-bit binary sequence through a hash function.

[0158] In a schematic example, the computer device Figure 12 performs a binary sequence conversion (converted into a 64-bit binary sequence) on the hierarchical directory shown in Figure 13 as shown.

[0159] To avoid the long-tail effect and reduce the computational amount, when the computer device performs a binary sequence conversion on the hierarchical directory, it first filters out the top m hierarchical directories based on the descending order of the directory weight, that is, filters out the top m hierarchical directories with larger directory weights, and performs a binary sequence conversion on the top m hierarchical directories.

[0160] For example, the computer device performs a binary sequence conversion on the top 100 hierarchical directories with directory weights.

[0161] For the hierarchical directories after the top m hierarchical directories, the computer device will not perform a binary sequence conversion. Correspondingly, the access characteristics of this part of the hierarchical directories will not be used to generate the website directory structure fingerprint.

[0162] Step 406: Use the directory weight to perform a weighting process on the binary sequence to obtain a directory vector corresponding to the hierarchical directory.

[0163] In a possible implementation, for each position in the binary sequence, the computer device determines the vector element corresponding to that position based on the value at that position and the directory weight, thereby obtaining a directory vector composed of vector elements. Among them, the vector length of the directory vector is the same as the sequence length of the binary sequence.

[0164] To enable the directory vector to reflect the access times of different directories, in some embodiments, for each position in the binary sequence, when the value of the position is 1, the computer device determines the directory weight as the vector element corresponding to that position.

[0165] When the value of the position is 0, the computer device determines the negative value of the directory weight as the vector element corresponding to that position.

[0166] Combined with Figure 13 the example in, for the binary sequence corresponding to the hierarchical directory 1topic, since the value of the 1st bit is 0 and the directory weight is 87, the computer device determines the vector element corresponding to the 1st bit as -87; since the value of the 5th bit is 1 and the directory weight is 87, the computer device determines the vector element corresponding to the 5th bit as 87.

[0167] Furthermore, the computer device determines the vector composed of vector elements as the directory vector corresponding to the hierarchical directory.

[0168] Schematically, for Figure 13 the binary sequences corresponding to each hierarchical directory in, after weighted processing, the obtained directory vector is as Figure 14 shown.

[0169] Step 407, generate the target website directory structure fingerprint of the target website based on the directory vectors corresponding to each hierarchical directory.

[0170] In a possible implementation, the computer device performs vector fusion on the directory vectors corresponding to each hierarchical directory to obtain the target website directory structure fingerprint of the target website. Among them, the target website target structure fingerprint is also in vector form and is consistent with the vector dimension of the directory vector.

[0171] Regarding the fusion method of the directory vectors, optionally, the computer device adds the directory vectors corresponding to each hierarchical directory to obtain the website vector corresponding to the target website.

[0172] Optionally, the website vector obtained by vector addition is a real number vector. Schematically, as Figure 14 shown, the website vector of the target website is (-228, -14, -72, -140, 14, 128, -142, -40, 118,...).

[0173] Optionally, the computer device may also perform a vector weighting operation on the directory vector according to the level corresponding to the hierarchical directory to obtain the website vector corresponding to the target website. The relationship between the level and the weight can be preset, and the weight of the intermediate-level directory is the largest. The embodiments of the present application do not limit the specific fusion method of the directory vector.

[0174] In some embodiments, the computer device determines the website vector as the target website directory structure fingerprint of the target website.

[0175] Since the search space of the real number vector is large, in order to reduce the computational complexity in the subsequent similarity calculation process and improve the recognition rate of website types, in other embodiments, the computer device normalizes the vector elements in the website vector to obtain the target website directory structure fingerprint of the target website, where the target website directory structure fingerprint is a binary sequence.

[0176] Among them, the normalization process is to convert the vector elements in the website vector to 0 or 1.

[0177] Regarding the method of normalization processing, in a possible implementation manner, when the vector element is greater than 0, the computer device converts the vector element to 1, and when the vector element is less than 0, the computer device converts the vector element to 0.

[0178] Of course, in other possible implementation manners, the computer device may also calculate the average value of each vector element and perform normalization processing on the vector elements based on the average value. The embodiments of the present application do not limit the specific method of normalization.

[0179] Step 408, in the case where the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than the similarity threshold, it is determined that the target website belongs to the target type, and the candidate website directory structure fingerprint is the website directory structure fingerprint of the known target type website.

[0180] In a possible implementation manner, when the website directory structure fingerprint is a binary sequence, the computer device may use the Hamming distance to represent the similarity. Correspondingly, the computer device determines the Hamming distance between the target website directory structure fingerprint and the candidate website directory structure fingerprint. In the case where the Hamming distance is less than the distance threshold, it is determined that the target website belongs to the target type; in the case where the Hamming distance is greater than the distance threshold, it is determined that the target website does not belong to the target type.

[0181] In this embodiment, before classifying the directory of the URL link, the computer device filters the general element string, filters the root website address, and replaces the specific type string of the URL link, so as to avoid excessive invalid elements from affecting the accuracy of subsequent directory classification, reduce unnecessary subsequent structural branches, and help improve the efficiency and accuracy of fingerprint generation.

[0182] In addition, the computer device converts the hierarchical directory into a binary sequence with a fixed length, and weights the binary sequence by using the directory weight corresponding to the hierarchical directory to obtain the directory vector of each hierarchical directory, which is convenient for subsequent vector fusion of the directory vectors to obtain the website vector of the target website. By normalizing the vector elements in the website vector, the search space of the vector is reduced, which helps to improve the efficiency of subsequent similarity calculation and helps to improve the recall rate of website type recognition.

[0183] Moreover, when determining the directory weight of the directory, in addition to based on the occurrence frequency of the directory, the word vector corresponding to the directory vocabulary is further combined to avoid the limitation caused by determining the weight only using the URL access record of the target website, and further improve the accuracy of the finally generated fingerprint.

[0184] Combined with the above embodiments, the complete process of the computer device for website type recognition is as Figure 15 shown.

[0185] Step 1501, determine the target website.

[0186] In some embodiments, the computer device determines the website suspected of having access risks as the target website.

[0187] Step 1502, obtain the URL access record of the target website.

[0188] Optionally, the computer device obtains the URL access record generated by accessing the target website from a large number of historical access records based on the root website address of the target website. In some embodiments, the computer device determines the historical access record containing the root website address of the target website as the URL access record of the target website.

[0189] Step 1503, perform data cleaning and word segmentation processing on the URL link in the URL access record.

[0190] After the computer device filters the general element string, filters the root website address, and replaces the specific type string of the URL link, it performs word segmentation processing on the URL link based on the target punctuation to obtain at least one word segment corresponding to the URL link.

[0191] In some embodiments, the computer device first filters the URL link based on a string filter table, where the string filter table contains the URL common element strings and the root web address of the target website. Then, the computer device replaces the strings of the target type in the filtered URL link with substitute words. Further, taking the target punctuation as the word segmentation position, the computer device performs word segmentation on the URL link after string replacement to obtain the word segmentation result of the URL link.

[0192] Step 1504, perform directory grading based on the word segmentation result.

[0193] Optionally, the computer device determines the parent-child relationship between different directories based on the order relationship before word segmentation in the word segmentation result, constructs a website directory structure based on the parent-child relationship, and thus determines the directory grading based on the website directory structure.

[0194] In order to improve the expression of the directory levels after directory grading, in some embodiments, for the first word segmentation in the word segmentation result, the computer device concatenates the first-level identifier and the first word segmentation to obtain the first-level link directory; for the word segmentations after the first word segmentation in the word segmentation result, the computer device concatenates the upper-level identifier, the previous word segmentation, the current-level identifier, and the current word segmentation to obtain the link directory of the current level. Further, the computer device determines the link directory grading result of the URL link for each level of link directory.

[0195] Step 1505, determine the directory weights of each level of directory.

[0196] In some embodiments, the computer device performs summary of the same directories on the link directory grading results of different URL links, and determines the target website directory structure of the target website based on the hierarchical relationship between different directories; thus, based on the occurrence frequency of each level of directory in the link directory grading result in the target website directory structure, the computer device determines the directory weight of each level of directory, and the directory weight is positively correlated with the occurrence frequency.

[0197] In some other embodiments, in order to further improve the accuracy of the determined directory weights, for the target level directory in the target website directory structure, the computer device determines the occurrence frequency of the target level directory in the link directory grading result, and when the target level directory is included in the thesaurus, the computer device obtains the target word weight corresponding to the target level directory from the thesaurus, so as to determine the directory weight of the target level directory based on the target word weight and the occurrence frequency. When the target level directory is not included in the thesaurus, the computer device determines the directory weight of the target level directory based on the occurrence frequency and the average word weight of the thesaurus. Wherein, the thesaurus contains the corresponding relationship between different words and word weights.

[0198] Step 1506: Obtain the top m directories according to the directory weights.

[0199] To avoid the long-tail effect, the computer device generates the website directory structure fingerprint of the target website only based on the top m directories with larger directory weights. For example, the computer device generates the website directory structure fingerprint only based on the directories with the top 100 directory weights.

[0200] Step 1507: Map the directories to binary sequences of a fixed length.

[0201] The computer device maps different-level directories to binary sequences of equal length through a hash function.

[0202] Step 1508: Weight the binary sequences based on the directory weights to obtain a real-number vector of the directories.

[0203] The computer device performs a weighting process on the binary sequences of the directories based on the directory weights, so that the real-number vector of the directories can reflect the access characteristics of the directories.

[0204] In some embodiments, for each position in the binary sequence, when the value of the position is 1, the computer device determines the directory weight as the vector element corresponding to the position; when the value of the position is 0, the computer device determines the negative value of the directory weight as the vector element corresponding to the position.

[0205] Furthermore, the computer device determines the vector composed of the vector elements as the directory vector (i.e., the real-number vector) corresponding to the hierarchical directory.

[0206] Step 1509: Generate a real-number vector of the target website based on the real-number vector of the directories.

[0207] The computer device performs vector fusion on the real-number vectors of different directories to obtain the real-number vector of the target website. Among them, the vector fusion can be vector addition or vector weighting.

[0208] Step 1510: Perform normalization processing on the real-number vector of the target website to obtain the website directory structure fingerprint of the target website.

[0209] To narrow the search space, the computer device performs normalization processing on the real-number vector of the target website to obtain the website directory structure fingerprint of the target website in binary form.

[0210] In some embodiments, when the vector element is greater than 0, the computer device converts the vector element to 1, and when the vector element is less than 0, the computer device converts the vector element to 0. In other possible implementation manners, the computer device can also calculate the average value of each vector element and perform normalization processing on the vector elements based on the average value.

[0211] Step 1511, calculate the Hamming distance between the fingerprint of the target website directory structure and the candidate website directory structure fingerprints in the fingerprint database.

[0212] In some embodiments, the computer device calculates the Hamming distance between the fingerprint of the target website directory structure and each candidate website directory structure fingerprint in the fingerprint database.

[0213] Step 1512, if the Hamming distance is less than the distance threshold, determine that the target website belongs to the target type.

[0214] In some embodiments, when there is at least one candidate website directory structure fingerprint whose Hamming distance from the target website directory structure fingerprint is less than the distance threshold, the computer device determines that the target website belongs to the target type.

[0215] Step 1513, if the Hamming distance is greater than the distance threshold, determine that the target website does not belong to the target type.

[0216] In some embodiments, when there is no candidate website directory structure fingerprint in the fingerprint database whose Hamming distance from the target website directory structure fingerprint is less than the distance threshold, the computer device determines that the target website does not belong to the target type.

[0217] To further improve the recognition accuracy of the website type, the computer device can comprehensively use the DOM structure fingerprint of the target website and the fingerprint of the target website directory structure for website type recognition.

[0218] To ensure the reliability of the DOM structure fingerprint, in a possible implementation, the computer device obtains the DOM structure of the target website and determines whether the structural complexity of the DOM structure meets the recognition requirements.

[0219] Among them, the structural complexity of the DOM structure can be determined based on at least one of the number of nodes, the breadth of nodes, and the depth of nodes in the DOM structure. Correspondingly, the recognition requirements can include at least one of a lower limit of the number of nodes, a lower limit of the breadth of nodes, and a lower limit of the depth of nodes.

[0220] In some embodiments, when the number of nodes in the DOM structure is greater than the number threshold, the breadth of nodes is greater than the breadth threshold, and the depth of nodes is greater than the depth threshold, the computer device determines that the structural complexity of the DOM structure of the target website meets the recognition requirements.

[0221] Furthermore, in the case where the similarity between the fingerprint of the target website directory structure and the candidate website directory structure fingerprint is greater than the similarity threshold, and the structural complexity of the DOM structure does not meet the recognition requirements, since the reliability of the DOM structure fingerprint determined based on the current DOM structure is relatively low, the computer device determines that the target website belongs to the target type.

[0222] Optionally, when the structural complexity of the DOM structure meets the recognition requirements, since the reliability of the DOM structure fingerprint determined based on the current DOM structure is relatively high, the computer device generates the target website DOM structure fingerprint of the target website based on the DOM structure.

[0223] When the similarity between the target website DOM structure fingerprint and the candidate website DOM structure fingerprint is greater than the similarity threshold, and the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than the similarity threshold, the computer device determines that the target website belongs to the target type, and the candidate website DOM structure fingerprint is the website DOM structure fingerprint of a known target type website.

[0224] Optionally, when the similarity between the target website DOM structure fingerprint and the candidate website DOM structure fingerprint is less than the similarity threshold, and the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is less than the similarity threshold, the computer device determines that the target website does not belong to the target type.

[0225] When the similarity between the target website DOM structure fingerprint and the candidate website DOM structure fingerprint is less than the similarity threshold, or the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is less than the similarity threshold, the computer device further performs website type recognition by other means.

[0226] In this embodiment, when the complexity of the DOM structure of the target website meets the recognition requirements, the final recognition result is determined by integrating the website type recognition results of the DOM fingerprint and the website directory structure fingerprint, further improving the accuracy of website type recognition.

[0227] Figure 16 It is a structural block diagram of a website type recognition device provided by an exemplary embodiment of the present application. The device includes:

[0228] A record acquisition module 1601, configured to acquire the URL access record of the target website, where the URL access record includes the URL link generated by accessing the target website;

[0229] A directory grading module 1602, configured to perform directory grading on the URL link included in the URL access record to obtain the target website directory structure of the target website and the directory weights of each level directory, where the target website directory structure is a tree structure formed by each level directory of the target website;

[0230] A fingerprint generation module 1603, configured to generate the target website directory structure fingerprint of the target website based on the target website directory structure and the directory weights of each level directory;

[0231] A determination module 1604, configured to determine that the target website belongs to a target type when the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than a similarity threshold, where the candidate website directory structure fingerprint is the website directory structure fingerprint of a known target type website.

[0232] Optionally, a directory grading module 1602, configured to:

[0233] Perform data cleaning and word segmentation processing on the URL link to obtain a word segmentation result of the URL link;

[0234] Perform directory grading based on the word segmentation results of each URL link to obtain a link directory grading result of each URL link;

[0235] Based on the link directory grading results of each URL link, determine the target website directory structure of the target website and the directory weights of each hierarchical directory.

[0236] Optionally, the directory grading module 1602 is specifically configured to:

[0237] For the first word in the word segmentation result, splice the first-level identifier and the first word to obtain a first-level link directory;

[0238] For the nth word in the word segmentation result, splice the (n - 1)th level identifier, the (n - 1)th word, the nth level identifier, and the nth word to obtain the nth level link directory, where n is an integer greater than or equal to 2;

[0239] Determine each level of link directory as the link directory grading result of the URL link.

[0240] Optionally, the directory grading module 1602 is specifically configured to:

[0241] Perform summary of the same directories for the link directory grading results of different URL links, and determine the target website directory structure of the target website based on the hierarchical relationship between different directories;

[0242] Based on the occurrence frequency of each hierarchical directory in the link directory grading result in the target website directory structure, determine the directory weight of each hierarchical directory, where the directory weight has a positive correlation with the occurrence frequency.

[0243] Optionally, the directory grading module 1602 is specifically configured to:

[0244] For the target hierarchical directory in the target website directory structure, determine the occurrence frequency of the target hierarchical directory in the link directory grading result;

[0245] When the target hierarchical directory is included in the thesaurus, obtain the target word weight corresponding to the target hierarchical directory from the thesaurus; determine the directory weight of the target hierarchical directory based on the target word weight and the occurrence frequency;

[0246] When the target hierarchical directory is not included in the thesaurus, determine the directory weight of the target hierarchical directory based on the occurrence frequency and the average word weight of the thesaurus;

[0247] Wherein, the thesaurus includes the corresponding relationship between different words and word weights.

[0248] Optionally, the directory classification module 1602 is specifically configured to:

[0249] Based on the string filtering table, perform string filtering on the URL link, and the string filtering table includes URL common element strings and the root website address of the target website;

[0250] Replace the string of the target type in the filtered URL link with a substitute word;

[0251] Taking the target punctuation as the word segmentation position, perform word segmentation processing on the URL link after string replacement to obtain the word segmentation result of the URL link.

[0252] Optionally, the fingerprint generation module 1603 is used to:

[0253] Convert each hierarchical directory in the target website directory structure into a binary sequence, wherein the lengths of the binary sequences obtained by converting each hierarchical directory are the same;

[0254] Use the directory weight to perform weighted processing on the binary sequence to obtain a directory vector corresponding to the hierarchical directory;

[0255] Generate the target website directory structure fingerprint of the target website based on the directory vectors corresponding to each hierarchical directory.

[0256] Optionally, the fingerprint generation module 1603 is specifically configured to:

[0257] For each position in the binary sequence, when the value at the position is 1, determine the directory weight as the vector element corresponding to the position;

[0258] When the value at the position is 0, determine the negative value of the directory weight as the vector element corresponding to the position;

[0259] Determine the vector composed of the vector elements as the directory vector corresponding to the hierarchical directory.

[0260] Optionally, the fingerprint generation module 1603 is specifically configured to:

[0261] Perform vector addition on the directory vectors corresponding to each hierarchical directory to obtain the website vector corresponding to the target website;

[0262] Perform normalization processing on the vector elements in the website vector to obtain the target website directory structure fingerprint of the target website, where the target website directory structure fingerprint is a binary sequence.

[0263] Optionally, the similarity is represented by the Hamming distance;

[0264] The determination module 1604 is configured to:

[0265] Determine the Hamming distance between the target website directory structure fingerprint and the candidate website directory structure fingerprint;

[0266] In the case that the Hamming distance is less than the distance threshold, determine that the target website belongs to the target type.

[0267] Optionally, the fingerprint generation module 1603 is specifically configured to:

[0268] Based on the descending order of the directory weights, screen out the top m hierarchical directories, where m is a positive integer;

[0269] Convert the top m hierarchical directories into the binary sequence.

[0270] Optionally, the device further includes:

[0271] The DOM acquisition module is configured to acquire the DOM structure of the target website;

[0272] The determination module 1604 is configured to:

[0273] In the case that the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than the similarity threshold, and the structural complexity of the DOM structure does not meet the recognition requirements, determine that the target website belongs to the target type.

[0274] Optionally, the device further includes:

[0275] The DOM fingerprint generation module is configured to:

[0276] In the case that the structural complexity of the DOM structure meets the recognition requirements, generate the target website DOM structure fingerprint of the target website based on the DOM structure;

[0277] The determining module 1604 is further configured to: determine that the target website belongs to the target type when the similarity between the DOM structure fingerprint of the target website and the DOM structure fingerprint of the candidate website is greater than the similarity threshold, and the similarity between the directory structure fingerprint of the target website and the directory structure fingerprint of the candidate website is greater than the similarity threshold, where the DOM structure fingerprint of the candidate website is the DOM structure fingerprint of a known website of the target type.

[0278] Optionally, the apparatus further includes:

[0279] A diffusion mining module, configured to determine that a website having a reference relationship or a jump relationship with the target website belongs to the target type based on the URL relationship chain; or,

[0280] Determine that a website belonging to the same aggregation medium as the target website belongs to the target type, where the aggregation medium includes at least one of IP, ICP, whois, and network devices.

[0281] In summary, compared with the DOM structure, since the website directory structure can contain richer website structure information, in the embodiments of the present application, by obtaining a large number of URL links generated by accessing the target website and performing directory grading on the URL links, a tree structure formed by each level of directories in the target website and the directory weights of each level of directories are obtained, so as to generate a website directory structure fingerprint based on the website directory structure and the directory weights, and finally use the website directory structure fingerprint to identify the website type of the target website; in addition, since the website directory structure is generated based on the URL access records, a better type recognition effect can also be achieved for websites with a simple website structure, which helps to improve the recognition rate and recognition accuracy of the website type.

[0282] Please refer to Figure 17 , which shows a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application.

[0283] The computer device 1700 includes a central processing unit (CPU) 1701, a system memory 1704 including a random access memory 1702 and a read-only memory 1703, and a system bus 1705 connecting the system memory 1704 and the central processing unit 1701. The computer device 1700 also includes a basic input / output system (I / O system) 1706 for transmitting information between various components in the computer, and a mass storage device 1707 for storing an operating system 1713, application programs 1714, and other program modules 1715.

[0284] The basic input / output system 1706 includes a display 1708 for displaying information and input devices 1709 such as a mouse, keyboard, etc. for user input of information. Both the display 1708 and the input devices 1709 are connected to the central processing unit 1701 through an input / output controller 1710 connected to the system bus 1705. The basic input / output system 1706 may also include an input / output controller 1710 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1710 also provides outputs to a display screen, printer, or other types of output devices.

[0285] The mass storage device 1707 is connected to the central processing unit 1701 through a mass storage controller (not shown) connected to the system bus 1705. The mass storage device 1707 and its associated computer-readable medium provide non-volatile storage for the computer device 1700. That is, the mass storage device 1707 may include a computer-readable medium (not shown) such as a hard disk or drive.

[0286] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes random access memory (RAM), read-only memory (ROM), flash memory or other solid-state storage technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, magnetic disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several. The above system memory 1704 and mass storage device 1707 may be collectively referred to as memory.

[0287] The memory stores one or more programs, the one or more programs are configured to be executed by one or more central processing units 1701, the one or more programs contain instructions for implementing the above method, and the central processing unit 1701 executes the one or more programs to implement the methods provided by the above various method embodiments.

[0288] According to various embodiments of the present application, the computer device 1700 can also run on a remote computer on the network through a network such as the Internet. That is, the computer device 1700 can be connected to the network 1712 through the network interface unit 1711 connected to the system bus 1705. Or rather, the network interface unit 1711 can also be used to connect to other types of networks or remote computer systems (not shown).

[0289] An embodiment of the present application also provides a computer-readable storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the website type identification method provided in the above embodiment.

[0290] Optionally, the computer-readable storage medium may include: ROM, RAM, solid state drives (SSDs), optical discs, etc. Among them, RAM may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM).

[0291] An embodiment of the present application provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the website type identification method described in the above embodiment.

[0292] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disc, etc.

[0293] The above are only optional embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for identifying website types, characterized in that, The method includes: Obtaining the URL access record of the target website, where the URL access record includes the URL links generated by accessing the target website; Performing directory classification on the URL links included in the URL access record to obtain the target website directory structure of the target website and the directory weights of each level of directories, where the target website directory structure is a tree structure formed by each level of directories of the target website; Generating a fingerprint of the target website directory structure of the target website based on the target website directory structure and the directory weights of each level of directories; When the similarity between the fingerprint of the target website directory structure and the fingerprint of the candidate website directory structure is greater than the similarity threshold, determining that the target website belongs to the target type, where the fingerprint of the candidate website directory structure is the fingerprint of the website directory structure of a known target type website.

2. The method according to claim 1, wherein The performing directory classification on the URL links included in the URL access record to obtain the target website directory structure of the target website and the directory weights of each level of directories includes: Performing data cleaning and word segmentation processing on the URL links to obtain the word segmentation results of the URL links; Performing directory classification based on the word segmentation results of each URL link to obtain the link directory classification results of each URL link; Determining the target website directory structure of the target website and the directory weights of each level of directories based on the link directory classification results of each URL link.

3. The method according to claim 2, wherein The performing directory classification based on the word segmentation results of each URL link to obtain the link directory classification results of each URL link includes: For the first word segment in the word segmentation results, concatenating the first-level identifier and the first word segment to obtain the first-level link directory; For the nth word segment in the word segmentation results, concatenating the (n - 1)-level identifier, the (n - 1)th word segment, the n-level identifier, and the nth word segment to obtain the nth-level link directory, where n is an integer greater than or equal to 2; Determining each level of link directory as the link directory classification result of the URL link.

4. The method according to claim 2, wherein The determining the target website directory structure of the target website and the directory weights of each level of directories based on the link directory classification results of each URL link includes: Performing summary of the same directories for the link directory classification results of different URL links, and determining the target website directory structure of the target website based on the hierarchical relationship between different directories; Determining the directory weights of each level of directories based on the occurrence frequency of each level of directory in the link directory classification results, where the directory weight has a positive correlation with the occurrence frequency.

5. The method according to claim 4, wherein The determining the directory weights of each level of directories based on the occurrence frequency of each level of directory in the link directory classification results includes: For the target level directory in the target website directory structure, determining the occurrence frequency of the target level directory in the link directory classification results; When the target hierarchical directory is included in the thesaurus, obtain the target word weight corresponding to the target hierarchical directory from the thesaurus; determine the directory weight of the target hierarchical directory based on the target word weight and the occurrence frequency; When the target hierarchical directory is not included in the thesaurus, determine the directory weight of the target hierarchical directory based on the occurrence frequency and the average word weight of the thesaurus; Wherein, the thesaurus includes the correspondence between different words and word weights.

6. The method according to claim 2, wherein The data cleaning and word segmentation processing of the URL link to obtain the word segmentation result of the URL link includes: Based on the string filtering table, perform string filtering on the URL link, and the string filtering table includes URL common element strings and the root URL of the target website; Replace the strings of the target type in the filtered URL link with substitute words; Using the target punctuation as the word segmentation position, perform word segmentation processing on the URL link after string replacement to obtain the word segmentation result of the URL link.

7. The method according to claim 1, characterized in that, The generating the target website directory structure fingerprint of the target website based on the target website directory structure and the directory weights of each hierarchical directory includes: Convert each hierarchical directory in the target website directory structure into a binary sequence, wherein the lengths of the binary sequences obtained by converting each hierarchical directory are the same; Perform weighted processing on the binary sequence using the directory weight to obtain the directory vector corresponding to the hierarchical directory; Generate the target website directory structure fingerprint of the target website based on the directory vectors corresponding to each hierarchical directory.

8. The method according to claim 7, wherein The performing weighted processing on the binary sequence using the directory weight to obtain the directory vector corresponding to the hierarchical directory includes: For each position in the binary sequence, when the value at the position is 1, determine the directory weight as the vector element corresponding to the position; When the value at the position is 0, determine the negative value of the directory weight as the vector element corresponding to the position; Determine the vector composed of the vector elements as the directory vector corresponding to the hierarchical directory.

9. The method according to claim 7, wherein The generating the target website directory structure fingerprint of the target website based on the directory vectors corresponding to each hierarchical directory includes: Perform vector addition on the directory vectors corresponding to each hierarchical directory to obtain the website vector corresponding to the target website; Perform normalization processing on the vector elements in the website vector to obtain the target website directory structure fingerprint of the target website, wherein the target website directory structure fingerprint is a binary sequence.

10. The method according to claim 9, wherein The similarity is represented by the Hamming distance; The determining that the target website belongs to the target type when the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than the similarity threshold includes: Determine the Hamming distance between the target website directory structure fingerprint and the candidate website directory structure fingerprint; When the Hamming distance is less than the distance threshold, determine that the target website belongs to the target type.

11. The method according to claim 7, wherein Converting each level directory in the target website directory structure into a binary sequence includes: Filtering to obtain the first m level directories based on the descending order of the directory weights, where m is a positive integer; Converting the first m level directories into the binary sequence.

12. The method according to any one of claims 1 to 11, characterized in that, The method further includes: Obtaining the DOM structure of the target website; When the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than the similarity threshold, determining that the target website belongs to the target type includes: When the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than the similarity threshold, and the structural complexity of the DOM structure does not meet the recognition requirements, determining that the target website belongs to the target type.

13. The method according to claim 12, wherein The method further includes: When the structural complexity of the DOM structure meets the recognition requirements, generating a target website DOM structure fingerprint of the target website based on the DOM structure; When the similarity between the target website DOM structure fingerprint and the candidate website DOM structure fingerprint is greater than the similarity threshold, and the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than the similarity threshold, determining that the target website is determined to be the target type, where the candidate website DOM structure fingerprint is the website DOM structure fingerprint of a known target type website.

14. The method according to any one of claims 1 to 11, characterized in that, After determining that the target website belongs to the target type, the method further includes: Based on the URL relationship chain, determining that websites having a reference relationship or a jump relationship with the target website belong to the target type; or, Determining that websites belonging to the same aggregation medium as the target website belong to the target type, where the aggregation medium includes at least one of IP, ICP, whois, and network devices.

15. An apparatus for identifying website types, characterized in that The device includes: A record acquisition module, configured to acquire a URL access record of a target website, where the URL access record includes a URL link generated by accessing the target website; A directory classification module, configured to perform directory classification on the URL links included in the URL access record to obtain a target website directory structure of the target website and the directory weights of each level directory, where the target website directory structure is a tree structure formed by each level directory of the target website; A fingerprint generation module, configured to generate a target website directory structure fingerprint of the target website based on the target website directory structure and the directory weights of each level directory; A determination module, configured to determine that the target website belongs to the target type when the similarity between the target website directory structure fingerprint and the candidate website directory structure fingerprint is greater than the similarity threshold, where the candidate website directory structure fingerprint is the website directory structure fingerprint of a known target type website.

16. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the website type recognition method according to any one of claims 1 to 14.

17. A computer-readable storage medium, characterized in that, At least one instruction is stored in the readable storage medium, and the at least one instruction is loaded and executed by a processor to implement the website type recognition method according to any one of claims 1 to 14.

18. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, the processor obtains the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to implement the website type recognition method according to any one of claims 1 to 14.