Website detection method and device, storage medium and electronic device

By extracting website association features and calculating fingerprint information from the set of URLs to be detected, the problem of poor malicious URL detection in existing technologies is solved, and accurate same-origin detection of batch URLs is achieved, thus improving detection efficiency.

CN119004137BActive Publication Date: 2026-05-15BEIJING HONGTENG INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING HONGTENG INTELLIGENT TECH CO LTD
Filing Date
2024-09-24
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies are ineffective at identifying and blocking newly emerging similar malicious URLs, making it difficult to improve the accuracy of identifying and blocking batches of similar URLs.

Method used

By determining the set of URLs to be detected, website association features are extracted, website fingerprint information is calculated, and same-origin URLs are detected based on the fingerprint information. Website-related features are used to characterize the correlation and similarity between URLs, thereby achieving accurate detection of batch URLs.

Benefits of technology

It improves the accuracy and efficiency of detecting same-origin URLs for batches of URLs to be detected, can quickly identify malicious URLs, and reduces the need to detect each URL individually.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119004137B_ABST
    Figure CN119004137B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a website detection method and device, a storage medium and an electronic device. The method comprises: determining a set of to-be-detected websites, extracting website correlation features of the to-be-detected websites in the set of to-be-detected websites, obtaining website correlation features corresponding to the set of to-be-detected websites, then calculating website feature values of the to-be-detected websites by using the website correlation features to obtain website fingerprint information, and performing homologous website detection processing on the set of to-be-detected websites based on the website fingerprint information to obtain homologous website detection results corresponding to the set of to-be-detected websites. The website correlation features between the to-be-detected websites in the set of to-be-detected websites are calculated to obtain website fingerprint information of each to-be-detected website. The homologous website detection processing is performed by using the website fingerprint information representing the similarity between the websites, so that the homologous website detection results of a batch of to-be-detected websites in the set of to-be-detected websites can be accurately determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, storage medium, and electronic device for detecting website addresses. Background Technology

[0002] Malicious website detection is a crucial technology in cybersecurity, designed to identify and block malicious websites that may endanger user security. Malicious websites are typically used to spread viruses, Trojans, phishing attacks, or other cyberattacks, potentially leading to issues such as personal information leaks and financial losses. Malicious website detection systems analyze multiple dimensions of features, including webpage content, domain characteristics, and IP addresses, employing machine learning models, blacklist databases, and static feature matching to identify and block these dangerous websites, thereby protecting users' cybersecurity. Summary of the Invention

[0003] This application provides a method, apparatus, computer storage medium, and electronic device for detecting website addresses. The technical solution is as follows:

[0004] In a first aspect, embodiments of this application provide a URL detection method, the method comprising:

[0005] Determine the set of URLs to be detected;

[0006] Website association features are extracted from the set of URLs to be detected to obtain the website-related features corresponding to the set of URLs to be detected.

[0007] Based on the website-related features, website feature values ​​are calculated for each URL to be detected to obtain website fingerprint information;

[0008] Based on the website fingerprint information, the same-origin URL detection process is performed on the set of URLs to be detected to obtain the same-origin URL detection results corresponding to the set of URLs to be detected.

[0009] Secondly, embodiments of this application provide a URL detection device, the device comprising:

[0010] The URL determination module is used to determine the set of URLs to be detected;

[0011] The feature extraction module is used to extract website association features from the set of URLs to be detected, and obtain the website-related features corresponding to the set of URLs to be detected.

[0012] The fingerprint calculation module is used to calculate website feature values ​​for each URL to be detected based on the relevant website features, and obtain website fingerprint information.

[0013] The URL detection module is used to perform same-origin URL detection processing on the set of URLs to be detected based on the website fingerprint information, and obtain the same-origin URL detection results corresponding to the set of URLs to be detected.

[0014] Thirdly, embodiments of this application provide a computer storage medium having multiple instructions adapted for loading and executing the methods described above by a processor.

[0015] Fourthly, embodiments of this application provide an electronic device, which may include: a memory and a processor; wherein the memory stores a computer program adapted to be loaded by the memory and to execute the above-described method.

[0016] The beneficial effects of the technical solutions provided in this application include at least the following:

[0017] The URL detection method provided in this application determines a set of URLs to be detected. It extracts website association features from these URLs to obtain website-related features corresponding to the set. Then, it uses these website-related features to calculate website feature values ​​for each URL to obtain website fingerprint information. Finally, it performs same-origin URL detection processing on the set of URLs based on the website fingerprint information to obtain the same-origin URL detection results. By calculating website feature values ​​from the website association features among the URLs in the set, website fingerprint information is obtained for each URL. Since website fingerprint information is calculated from the website-related features that are mutually related between the URLs, it can represent the degree of association or similarity between the URLs in the set. Therefore, by performing same-origin URL detection processing using website fingerprint information, which characterizes the similarity between URLs, the same-origin URL detection results for a batch of URLs in the set can be accurately determined. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a schematic flowchart of a URL detection method provided in an embodiment of this application;

[0020] Figure 2 This is a flowchart illustrating another URL detection method provided in an embodiment of this application;

[0021] Figure 3 This is a schematic diagram of the structure of a URL detection device provided in an embodiment of this application;

[0022] Figure 4 This is a schematic diagram of the structure of a fingerprint calculation module provided in an embodiment of this application;

[0023] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0024] To make the inventive objectives, features, and advantages of the embodiments of this application more apparent and understandable, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0025] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this application, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0026] In related technologies, malicious website identification techniques include static feature matching and machine learning models. These techniques detect whether a website is malicious by independently identifying individual URLs. However, many malicious website providers continuously update and change the domain names and other content of related malicious websites using wildcard domains, cloud services, and dynamic IP pools. This leads to a situation where, after an initial malicious website is detected and blocked, many more similar malicious websites immediately appear. Static feature matching relies on pre-set rules, and machine learning models require extensive training with real data. Therefore, these identification techniques suffer from poor performance when detecting these newly emerging similar malicious websites. Thus, improving the accuracy of detecting batches of similar websites is a pressing technical problem that needs to be solved.

[0027] To address the aforementioned technical problems, this application will be described in detail below with reference to specific embodiments.

[0028] In one embodiment, such as Figure 1 As shown, a URL detection method is proposed. This method can be implemented using a computer program and can run on a URL detection device based on the von Neumann architecture. This computer program can be integrated into applications or run as a standalone utility application.

[0029] Specifically, the execution subject of this URL detection method is an electronic device, and the URL detection method includes:

[0030] S101, Determine the set of URLs to be detected.

[0031] The set of URLs to be tested refers to the collection containing at least two URLs to be tested. After testing, the URLs in the set can be classified as either benign or malicious. Malicious URLs can include, but are not limited to, URLs involved in phishing scams, fraud, hacking, viruses, Trojans, or other illegal or unethical activities.

[0032] It is understood that the application scenarios of the URL detection method provided in this application embodiment may include scenarios such as actively detecting malicious URLs, and searching for more malicious URLs when a customer requests (such as requests from regulatory agencies, financial institutions, etc.) when a known malicious URL is known.

[0033] The electronic device used in the URL detection method provided in this application embodiment can be a server deployed in the backend. For example, the electronic device can be a server in the backend of a certain search engine.

[0034] In some embodiments, determining the set of URLs to be detected may involve determining a set of URLs to be detected that includes one URL of a known type and at least one URL of an unknown type; or, determining a set of URLs to be detected that includes at least two URLs of unknown types. Specifically, the URLs to be detected in the set of URLs to be detected may be sent by the user equipment or found by the electronic device.

[0035] S102, extract website association features from the set of URLs to be detected to obtain the website-related features corresponding to the set of URLs to be detected.

[0036] Website-related features refer to website characteristics with a high degree of correlation among the URLs in the set of URLs to be tested. For example, website-related features may include at least one of the following: domain name, certificate, title, icon, blacklist, resource content, and website structure. Taking domain name and title as examples, the similarity of the domain name's feature values ​​among the URLs to be tested is 88%, while the similarity of the title's feature values ​​among the URLs to be tested is 40%. Clearly, the similarity of the domain name's feature values ​​among the URLs to be tested is higher than that of the title's feature values. Therefore, the domain name is a website feature with a higher degree of correlation among the URLs to be tested, and thus, the domain name can be identified as the website-related feature corresponding to the set of URLs to be tested.

[0037] In some embodiments, feature extraction can be performed based on the set of URLs to be detected to obtain the reference website feature values ​​corresponding to each URL to be detected, and correlation feature detection processing can be performed based on all reference website feature values ​​to obtain the website-related features corresponding to the set of URLs to be detected.

[0038] In some embodiments, relevant feature extraction prompts can be generated based on each URL in the set of URLs to be detected. These prompts, along with the URLs to be detected, are then input into a large-scale relevant feature detection model to obtain website-related features corresponding to the set of URLs. The relevant feature extraction prompts instruct the large-scale relevant feature detection model to perform website association feature extraction processing on each URL to obtain the website-related features corresponding to the set of URLs. This large-scale relevant feature detection model can directly use a basic Large Language Model (LLM), or it can be a large language model trained on a basic LLM for website-related feature extraction scenarios. An LLM is an artificial intelligence content generation model designed to understand and generate human language; for example, an LLM can be a generative artificial intelligence (AIGC) model. Specifically, the training process of the large-scale relevant feature detection model can be as follows: obtaining a basic LLM, creating a relevant feature detection scenario adaptation module for website-related feature extraction scenarios and a large language generative module based on the basic LLM, and then forming an initial large-scale relevant feature detection model based on the large language generative module and the relevant feature detection scenario adaptation module. The set of sample URLs to be detected is input into the initial large-scale relevance feature detection model for at least one round of model training to obtain predicted website-related features. Based on the predicted website-related features and the corresponding sample website-related feature labels of the set of sample URLs to be detected, the comprehensive model loss is calculated. Based on the comprehensive model loss, the model parameters of the relevant feature detection scenario adaptation module in the initial large-scale relevance feature detection model are adjusted, while the model parameters of the large language generation module remain unchanged. This process continues until the model training termination condition is met, resulting in the large language generation module and the relevant feature detection scenario adaptation module. The models of the large language generation module and the relevant feature detection scenario adaptation module are then fused to obtain the trained large-scale relevance feature detection model.

[0039] S103, calculate the website feature value of the URL to be detected based on the website's relevant features to obtain the website fingerprint information.

[0040] Among them, website fingerprint information refers to the website representation information with specific meaning represented by the website-related characteristics of the URL to be detected. Each URL to be detected has a website fingerprint information.

[0041] In some embodiments, the original website feature value corresponding to the URL to be detected is obtained based on the website-related features, the comprehensive feature vector corresponding to the website-related features is determined based on the original website feature value, the dimensionality reduction processing is performed based on the comprehensive feature vector to obtain the similarity hash value corresponding to the URL to be detected, and the similarity hash value is used as the website fingerprint information of each URL to be detected.

[0042] S104. Based on the website fingerprint information, perform same-origin URL detection processing on the set of URLs to be detected, and obtain the same-origin URL detection results corresponding to the set of URLs to be detected.

[0043] Among them, the same-origin URL detection results can be used to display the same-origin relationship and / or non-same-origin relationship between each URL in the set of URLs to be detected.

[0044] In some embodiments, a similar hash value corresponding to each URL to be detected can be obtained from the website fingerprint information. A target URL to be detected is determined from the set of URLs to be detected. A target similar hash value corresponding to the target URL to be detected is determined from all similar hash values. The target URL to be detected is a URL of a known type. A first similar hash value matching the target similar hash value is determined from all similar hash values. A second similar hash value not matching the target similar hash value is determined. The URL to be detected corresponding to the first similar hash value is determined as a homologous similar URL corresponding to the target URL to be detected. The second similar hash value is determined as a non-homologous similar URL corresponding to the target URL to be detected. A homologous URL detection result including the target URL to be detected, homologous similar URLs, and non-homologous similar URLs is generated.

[0045] In this way, by using the same-origin URL detection results, it is possible to identify URLs in the set of URLs to be detected that share the same origin as the target URL, as well as URLs that do not share the same origin as the target URL. Specifically, the target URL can be a known malicious URL. Therefore, using the method of this application embodiment, malicious URL detection can be performed on a batch of URLs in the set of URLs to be detected, so as to accurately determine the presence of malicious and non-malicious URLs in the batch of URLs to be detected.

[0046] The URL detection method provided in this application determines a set of URLs to be detected. It extracts website association features from these URLs to obtain website-related features corresponding to the set. Then, it uses these website-related features to calculate website feature values ​​for each URL to obtain website fingerprint information. Finally, it performs same-origin URL detection processing on the set of URLs based on the website fingerprint information to obtain the same-origin URL detection results. By calculating website feature values ​​from the website association features among the URLs in the set, website fingerprint information is obtained for each URL. Since website fingerprint information is calculated from the website-related features that are mutually related between the URLs, it can represent the degree of association or similarity between the URLs in the set. Therefore, by performing same-origin URL detection processing using website fingerprint information, which characterizes the similarity between URLs, the same-origin URL detection results for a batch of URLs in the set can be accurately determined.

[0047] Please see below. Figure 2 , Figure 2 This is a flowchart illustrating another embodiment of the URL detection method proposed in this application.

[0048] Specifically, the execution subject of this URL detection method is an electronic device, and the URL detection method includes:

[0049] S201, Determine the set of URLs to be detected.

[0050] In some embodiments, determining the set of URLs to be detected may involve determining a set of URLs to be detected that includes one URL of a known type and at least one URL of an unknown type; or, determining a set of URLs to be detected that includes at least two URLs of unknown types. Specifically, the URLs to be detected in the set of URLs to be detected may be sent by the user equipment or found by the electronic device.

[0051] S202, feature extraction is performed based on the set of URLs to be detected to obtain the reference website feature values ​​corresponding to each URL to be detected.

[0052] In some embodiments, S202 may specifically involve: determining preset website features; and performing feature extraction processing on the URLs to be detected in the set of URLs to be detected based on the preset website features to obtain reference website feature values ​​corresponding to each URL to be detected.

[0053] The preset website features can be website features set based on experience to detect URL types. These preset website features can be of various types, including but not limited to domain names, certificates, titles, icons, blacklists, resource content, website structure, and other features. Specifically, a certificate refers to the website's registration certificate; an icon refers to an icon on the webpage; a blacklist includes blacklisted words in the webpage's text that match the preset blacklist; resource content can include at least one of webpage text, webpage images, and webpage code; and website structure refers to the layout structure of the website content, including page hierarchy, navigation menus, internal links, and information categorization.

[0054] The reference website feature value refers to the feature value of the preset website feature.

[0055] In some embodiments, preset website features can be obtained by querying a feature configuration table, which can be configured based on expert experience. The feature configuration table is used to store reference website features corresponding to different URL detection requirements. For example, URL detection requirements may include URL type detection, URL content detection, URL security detection, URL performance detection, URL legality detection, etc. Since this embodiment is applied to the URL type detection scenario, the reference website features corresponding to URL type detection are queried from the feature configuration table, and the reference website features corresponding to URL type detection are used as preset website features.

[0056] Furthermore, for each URL in the set of URLs to be detected, feature values ​​of preset website features are extracted from the website of each URL to obtain reference website feature values ​​corresponding to each URL. It can be understood that when the preset website features include multiple types of website features, reference website feature values ​​for the multiple website features included in the preset website features can be extracted from the website of each URL to be detected.

[0057] S203, based on the feature values ​​of all reference websites, perform correlation feature detection processing to obtain the website-related features corresponding to the set of URLs to be detected.

[0058] In some embodiments, for each preset website feature, there is a corresponding reference website feature value in websites with different URLs to be detected. The feature similarity ranking of each preset website feature among all types of preset website features can be determined by the reference website feature values ​​corresponding to each preset website feature in websites with different URLs to be detected. Based on the feature similarity ranking, a preset number of preset website features that rank first are taken as website-related features. The preset number can be flexibly set according to actual application. For example, the preset number can be set to 1, 2, or 3. For example, preset website features include domain name, certificate, title, icon, black keyword database, resource content, and website structure; website-related features include domain name, or website-related features include domain name and title, or website-related features include domain name, title, and black keyword database.

[0059] Specifically, the method for determining the feature similarity ranking is as follows: Among the reference website feature values ​​corresponding to each preset website feature, determine the number of similar feature values ​​corresponding to each preset website feature. Arrange the number of similar feature values ​​corresponding to each preset website feature in descending order to determine the feature similarity ranking of each preset website feature. Here, the number of similar feature values ​​refers to the number of similar reference website feature values ​​determined among the reference website feature values ​​corresponding to each preset website feature.

[0060] In this way, by referencing website feature values, at least one preset website feature with the highest correlation is selected from preset website features as the website-related feature to calculate more accurate website fingerprint information.

[0061] S204: Obtain the original website feature values ​​corresponding to each URL to be detected based on website-related features.

[0062] Among them, the original website feature value can refer to the feature value of the website-related features in the website of the URL to be detected.

[0063] In some embodiments, after determining the website-related features, feature values ​​of the website-related features can be extracted from the websites of each URL to be detected to obtain the original website feature values ​​corresponding to the URL to be detected. It is understood that when the website-related features include at least two types of related features, the original website feature values ​​of the multiple related features contained in the website-related features can be extracted from the websites of each URL to be detected.

[0064] For example, if the relevant feature of a website is its domain name, the feature value of the domain name extracted from URL A is abc.com, and abc.com is the original website feature value corresponding to URL A; the feature value of the domain name extracted from URL B is abd.com, and abd.com is the original website feature value corresponding to URL B; the feature value of the domain name extracted from URL C is bcd.com, and bcd.com is the original website feature value corresponding to URL C.

[0065] S205, determine the comprehensive feature vector corresponding to the relevant features of the website based on the original website feature values.

[0066] In some embodiments, performing S205 may specifically include the following steps:

[0067] A1, Perform feature filtering on the original website feature values ​​to obtain the target website feature values;

[0068] Specifically, determine the matching website features corresponding to the original website feature values. If the matching website features are of the first type, then perform feature filtering on the original website feature values ​​to obtain the target website feature values. If the matching website features are of the second type, then cancel the feature filtering on the original website feature values ​​to obtain the target website feature values, and use the original website feature values ​​as the target website feature values.

[0069] In this context, "matching website-related features" refers to extracting a specific website-related feature from the website of the URL to be detected, resulting in a single website feature value. Specifically, there are two types of related features: Type 1 features, Type 2 features, and Type 3 features. Type 1 features may include domain names, certificates, and website structure; Type 2 features may include resource content, icons, and black hat keywords. For website-related features like titles, which may yield one or more feature values, if the title is a first-level heading, it falls under Type 1 related features; if it's a second-level heading, it falls under Type 2 related features.

[0070] Given that the original website feature values ​​are the feature values ​​of the first type of related features, the target website feature values ​​can be obtained by performing feature filtering on the original website feature values. This can be done by: determining the feature extraction positions corresponding to the original website feature values, and identifying the original website feature values ​​corresponding to the feature extraction positions belonging to a first preset position as the target website feature values. The first preset position can be set based on experience; for example, the first preset position may include the website's main page position, website text position, etc. In this way, by filtering the target website feature values ​​from the original website feature values, the influence of uncritical feature values ​​on the calculation of the comprehensive feature vector can be avoided.

[0071] A2, based on the target website feature value, perform hash calculation to obtain the reference hash value corresponding to the target website feature value, and perform weight allocation processing on the target website feature value to obtain the reference weight;

[0072] Specifically, the hash value of each target website feature value is calculated to obtain a reference hash value corresponding to each target website feature value. The reference weight refers to the weight assigned to the target website feature value. This weight assignment process can include: determining the feature extraction position corresponding to the target website feature value within the URL to be detected; obtaining the target website-related features corresponding to the target website feature value; and inputting the feature extraction position, target website-related features, and target website feature value into a weight assignment model for weight allocation processing to obtain the reference weight. The feature extraction position is the location within the website of the URL to be detected where the target website feature value is extracted, and the target website-related features refer to the website-related features for which the target website feature value was extracted.

[0073] The weight allocation model can be created using a machine learning model. The training process for the weight allocation model may include the following steps:

[0074] Model creation: Specifically, an initial weight allocation model is created based on the machine learning model for the weight allocation scenario.

[0075] Obtaining sample data: Specifically, obtain the feature values ​​of the sample website, obtain the sample feature extraction positions corresponding to the feature values ​​of the sample website, obtain the relevant features of the sample website corresponding to the feature values ​​of the sample website, and generate sample data containing the feature values ​​of the sample website, the sample feature extraction positions, and the relevant features of the sample website.

[0076] Labeling the sample data: Specifically, based on the needs of the weight allocation scenario, an expert service is introduced to manually label the sample data with corresponding sample labels, including the sample weights corresponding to the feature values ​​of the sample websites.

[0077] Training the model: Input the sample data into the initial weight allocation model for at least one round of training to obtain the predicted weights corresponding to the sample data. Based on the predicted weights and sample weights, use the model loss function to determine the model loss value. Based on the model loss value, adjust the model parameters of the initial weight allocation model until the model training termination condition is met to obtain the weight allocation model.

[0078] Optionally, the model's training termination conditions may include, for example, the loss function value being less than or equal to a preset loss function threshold, or the number of iterations reaching a preset threshold. Specific training termination conditions can be determined based on actual circumstances and are not specifically limited here.

[0079] It should be noted that the machine learning models involved in one or more embodiments of this specification include, but are not limited to, fitting one or more of the following machine learning models: Convolutional Neural Network (CNN) model, Deep Neural Network (DNN) model, Recurrent Neural Networks (RNN) model, embedding model, Gradient Boosting Decision Tree (GBDT) model, Logistic Regression (LR) model, etc.

[0080] A3, based on the reference hash value and reference weight, performs vector mapping processing to obtain the reference feature vector corresponding to the feature value of the target website;

[0081] Specifically, for each target website feature value, its corresponding reference hash value is used to perform weight mapping on the reference weights to obtain the target weight. Based on all target weights, a reference feature vector corresponding to the target website feature value is generated. For example, the reference hash value can be in binary form. If the reference hash value corresponding to a certain target website feature value is 10110, and the reference weight corresponding to this target website feature value is 5, then the number 1 in the hash value can be mapped to a target weight of 5, and the number 0 in the hash value can be mapped to another target weight of -5. Therefore, using 10110 to map 5, a series of target weights 5, -5, 5, 5, -5 can be obtained, and the reference feature vector (5, -5, 5, 5, -5) can be generated.

[0082] A4, based on the reference feature vector, performs vector merging to obtain the comprehensive feature vector corresponding to the website's relevant features.

[0083] Among them, the comprehensive feature vector can refer to the data representation that characterizes the website-related features of each URL to be detected, and can be used to describe the attributes of the website-related features.

[0084] Specifically, each target website feature value corresponds to a reference feature vector. All reference feature vectors are summed to obtain a comprehensive feature vector corresponding to the website's relevant features.

[0085] S206. Based on the comprehensive feature vector, dimensionality reduction is performed to obtain the similarity hash value corresponding to each URL to be detected, and the similarity hash value is used as the website fingerprint information of each URL to be detected.

[0086] In some embodiments, a preset threshold can be used to map each element value in the comprehensive feature vector to obtain a target binary value, and a similar hash value corresponding to the URL to be detected can be generated based on the target binary value. Specifically, if an element value in the comprehensive feature vector is greater than or equal to the preset threshold, the element value can be mapped to a first target binary value; if an element value in the comprehensive feature vector is less than the preset threshold, the element value can be mapped to a second target binary value. Further, a similar hash value containing the first target binary value and / or the second target binary value is generated. The first target binary value and the second target binary value are different.

[0087] For example, the first target binary value is configured as 1, the second target binary value is configured as 0, and the comprehensive feature vector corresponding to a certain website's related features is (3, -3, 3, 5, -7). The preset threshold is 0. For the first element value 3, which is greater than 0, it is mapped to 1; for the second element value -3, which is less than 0, it is mapped to 0; for the third element value 3, which is greater than 0, it is mapped to 1; for the fourth element value 5, which is greater than 0, it is mapped to 1; and for the fifth element value -7, which is less than 0, it is mapped to 0. Therefore, the similarity hash value can be obtained as 10110.

[0088] S207, Obtain the similar hash value corresponding to each URL to be detected from the website fingerprint information.

[0089] Specifically, since the website fingerprint information of each URL to be detected is determined by the similar hash value corresponding to the URL to be detected, it is easy to obtain the similar hash value corresponding to each URL to be detected from the website fingerprint information.

[0090] S208, determine the target URL to be detected from the set of URLs to be detected, determine the target similar hash value corresponding to the target URL to be detected from all similar hash values, determine the first similar hash value that matches the target similar hash value from all similar hash values, and determine the URL to be detected corresponding to the first similar hash value as the same-origin similar URL corresponding to the target URL to be detected.

[0091] The target URL to be detected is a URL of a known type. For example, the target URL to be detected could be a known malicious URL.

[0092] Among them, "same-origin similar URLs" refers to malicious URLs that belong to the same origin as the target URL to be detected.

[0093] It is understandable that the type of the target URL to be detected can be known before S201 is executed, or it can be determined before S207 is executed, that is, during the execution of S201-S206, or after S206 is executed and before S207 is executed.

[0094] For example, a specific implementation of determining the target URL from the set of URLs to be detected can be: searching for the target URL from the set of URLs to be detected that has type marking information, which is used to indicate that the target URL is a URL of a known type and to indicate the type of the target URL.

[0095] For example, determining the first similar hash value that matches the target similar hash value from all similar hash values ​​can be done by determining the first similar hash value that is the same as the target similar hash value from all similar hash values.

[0096] For example, determining the first similar hash value matching the target similar hash value from all similar hash values ​​can be done by determining the first similar hash value that is similar to the target similar hash value from all similar hash values. Specifically, the hash value distance between the target similar hash value and each reference similar hash value is calculated, where each reference similar hash value is any similar hash value other than the target similar hash value. The reference similar hash value with a distance less than a first threshold is taken as the first similar hash value; where the hash value distance can be Hamming distance, Euclidean distance, or other distances, etc. Alternatively, the similarity between the target similar hash value and each reference similar hash value is calculated, and the reference similar hash value with a similarity greater than a second threshold is taken as the first similar hash value.

[0097] S209, determine the second similar hash value that does not match the target similar hash value from all similar hash values, and determine the URL to be detected corresponding to the second similar hash value as a non-same-origin similar URL to the target URL to be detected.

[0098] Among them, non-same-origin similar URLs refer to URLs that are neither from the same origin nor similar to the target URL to be detected.

[0099] It is understandable that similar URLs from different origins may be benign. For example, if the target URL to be detected is a fake bank website, a similar URL from different origins may be a website of an institution. In this case, the similar URL from different origins is a benign URL.

[0100] For example, determining the first similar hash value that does not match the target similar hash value from all similar hash values ​​can be done by determining the first similar hash value that is different from the target similar hash value from all similar hash values.

[0101] For example, determining the first similar hash value that does not match the target similar hash value from all similar hash values ​​can be done by determining the first similar hash value that is dissimilar to the target similar hash value from all similar hash values. Specifically, the hash value distance between the target similar hash value and each reference similar hash value is calculated, where the reference similar hash value is any similar hash value other than the target similar hash value among all similar hash values. The reference similar hash value with a distance greater than a third threshold is taken as the second similar hash value; where the hash value distance can be Hamming distance, Euclidean distance, or other distances, etc. Alternatively, the similarity between the target similar hash value and each reference similar hash value is calculated, and the reference similar hash value with a similarity less than a fourth threshold is taken as the second similar hash value.

[0102] S210, Generate the same-origin URL detection results corresponding to the set of URLs to be detected based on the target URL to be detected, same-origin similar URLs, and non-same-origin similar URLs.

[0103] Specifically, the same-origin URL detection results can include the target URL to be detected, similar URLs from the same origin, and similar URLs from different origins. For example, the set of URLs to be detected can include URL a, URL b, URL c, URL d, URL e, and URL f, where URL b is the target URL to be detected, URLs a, d, and f are determined to be similar URLs from the same origin as the target URL to be detected, and URLs c and e are determined to be similar URLs from different origins as the target URL to be detected. The same-origin URL detection results can be: the target URL to be detected - URL b is a known malicious URL, similar URLs from the same origin - URL a, similar URLs from the same origin - URL d, and similar URLs from the same origin - URL f are malicious URLs from the same origin as URL b, and similar URLs from different origins - URL c and similar URLs from different origins - URL e are URLs from different origins and not similar to URL b.

[0104] In the URL detection method provided in this application embodiment, after determining the set of URLs to be detected, feature extraction is performed based on the set of URLs to be detected to obtain reference website feature values ​​corresponding to each URL to be detected. Based on all reference website feature values, correlation feature detection processing is performed to obtain website-related features corresponding to the set of URLs to be detected. Therefore, at least one preset website feature with the highest correlation is selected from preset website features using the reference website feature values ​​as a website-related feature to calculate more accurate website fingerprint information. Then, based on the website-related features, the original website feature values ​​corresponding to each URL to be detected are obtained. Based on the original website feature values, a comprehensive feature vector corresponding to the website-related features is determined. Based on the comprehensive feature vector, dimensionality reduction processing is performed to obtain the similarity hash value corresponding to each URL to be detected. The similarity hash value is used as the website fingerprint information for each URL to be detected. Then, the similar hash value corresponding to each URL to be detected is obtained from the website fingerprint information. The target URL to be detected is determined from the set of URLs to be detected. The target similar hash value corresponding to the target URL to be detected is determined from all similar hash values. The first similar hash value that matches the target similar hash value is determined from all similar hash values. The URL to be detected corresponding to the first similar hash value is determined as a homologous similar URL to the target URL to be detected. The second similar hash value that does not match the target similar hash value is determined from all similar hash values. The URL to be detected corresponding to the second similar hash value is determined as a non-homologous similar URL to the target URL to be detected. Based on the target URL to be detected, homologous similar URLs, and non-homologous similar URLs, homologous URL detection results corresponding to the set of URLs to be detected are generated. Since website fingerprint information is calculated using the most relevant website features, and website fingerprint information is represented by similar hash values, by searching for similar hash values ​​that match known types of URLs to be detected, it is possible to quickly and accurately determine similar URLs from the same origin and similar URLs from different origins from the set of URLs to be detected, without needing to perform independent URL detection on batches of URLs to be detected. This achieves clustered classification detection of batches of URLs to be detected, and while accurately detecting malicious URLs, it also improves the detection efficiency of batches of URLs to be detected.

[0105] The following will combine Figure 3 This paper provides a detailed description of the URL detection device provided in the embodiments of this application. It should be noted that... Figure 3 The URL detection device shown is used to perform the functions described in this application. Figures 1-2 The methods shown in the embodiments are for illustrative purposes only, illustrating the parts relevant to the embodiments of this application. For specific technical details not disclosed, please refer to this application. Figures 1-2 The example shown.

[0106] Please see Figure 3This diagram illustrates the structure of a URL detection device according to an embodiment of this application. The URL detection device 1 can be implemented as all or part of a device through software, hardware, or a combination of both. According to some embodiments, the URL detection device 1 includes a URL determination module 11, a feature extraction module 12, a fingerprint calculation module 13, and a URL detection module 14, specifically used for:

[0107] URL determination module 11 is used to determine the set of URLs to be detected;

[0108] Feature extraction module 12 is used to extract website association features from the set of URLs to be detected, and obtain website-related features corresponding to the set of URLs to be detected;

[0109] Fingerprint calculation module 13 is used to calculate website feature values ​​for each URL to be detected based on the website-related features, and obtain website fingerprint information;

[0110] The URL detection module 14 is used to perform same-origin URL detection processing on the set of URLs to be detected based on the website fingerprint information, and obtain the same-origin URL detection result corresponding to the set of URLs to be detected.

[0111] Optionally, the feature extraction module 12 includes:

[0112] The first extraction unit is used to extract features based on the set of URLs to be detected, and obtain the reference website feature values ​​corresponding to each URL to be detected.

[0113] The second extraction unit is used to perform associated feature detection processing based on the feature values ​​of all the reference websites to obtain the website-related features corresponding to the set of URLs to be detected.

[0114] Optionally, the first extraction unit is specifically used for:

[0115] Determine the preset website characteristics;

[0116] Based on the preset website features, feature extraction processing is performed on the URLs to be detected in the set of URLs to be detected to obtain reference website feature values ​​corresponding to each URL to be detected.

[0117] Optionally, please see Figure 4 This is a schematic diagram of the structure of a fingerprint calculation module 13 provided in an embodiment of this application. The fingerprint calculation module 13 may include a feature acquisition unit 131, a vector determination unit 132, and a fingerprint calculation unit 133, specifically used for:

[0118] Feature acquisition unit 131 is used to acquire the original website feature value corresponding to each URL to be detected based on the website-related features;

[0119] Vector determination unit 132 is used to determine the comprehensive feature vector corresponding to the website-related features based on the original website feature values;

[0120] The fingerprint calculation unit 133 is used to perform dimensionality reduction processing based on the comprehensive feature vector to obtain the similar hash value corresponding to each URL to be detected, and to use the similar hash value as the website fingerprint information of each URL to be detected.

[0121] Optionally, the vector determination unit 132 includes:

[0122] The first determining unit is used to perform feature filtering processing on the original website feature values ​​to obtain target website feature values;

[0123] The second determining unit is used to perform hash calculation based on the target website feature value to obtain a reference hash value corresponding to the target website feature value, and to perform weight allocation processing on the target website feature value to obtain a reference weight;

[0124] The third determining unit is used to perform vector mapping processing based on the reference hash value and the reference weight to obtain the reference feature vector corresponding to the target website feature value.

[0125] The fourth determining unit is used to perform vector merging processing based on the reference feature vector to obtain the comprehensive feature vector corresponding to the website-related features.

[0126] Optionally, the second determining unit is specifically used for:

[0127] Determine the feature extraction position corresponding to the feature value of the target website in the URL to be detected;

[0128] Obtain the target website-related features corresponding to the target website feature values;

[0129] The feature extraction location, the relevant features of the target website, and the feature value of the target website are input into the weight allocation model for weight allocation processing to obtain the reference weight.

[0130] Optionally, the URL detection module 14 is specifically used for:

[0131] Obtain the similar hash value corresponding to each URL to be detected from the website fingerprint information;

[0132] The target URL to be detected is determined from the set of URLs to be detected. The target similar hash value corresponding to the target URL to be detected is determined from all the similar hash values. The first similar hash value matching the target similar hash value is determined from all the similar hash values. The URL to be detected corresponding to the first similar hash value is determined as a similar URL from the same source as the target URL to be detected.

[0133] From all the similar hash values, determine a second similar hash value that does not match the target similar hash value, and determine the URL to be detected corresponding to the second similar hash value as a non-same-origin similar URL to the target URL to be detected;

[0134] Based on the target URL to be detected, the same-origin similar URLs, and the non-same-origin similar URLs, generate the same-origin URL detection results corresponding to the set of URLs to be detected.

[0135] Please refer to Figure 5 , Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. For example, the electronic device in this embodiment may specifically be a server, and the electronic device may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 can be connected via the bus 150.

[0136] Processor 110 may include one or more processing cores. Processor 110 connects to various parts of the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 110 may integrate one or more of a central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately using a communication chip.

[0137] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (e.g., touch function, sound playback function, image playback function, etc.), instructions for implementing the various method embodiments described below, etc. The operating system may be the Android system, including systems deeply developed based on the Android system, the iOS system developed by Apple Inc., including systems deeply developed based on the iOS system, or other systems.

[0138] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to establish data communication between the third-party applications and the operating system. This would allow the operating system to obtain the current scenario information of the third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.

[0139] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 can be a touch display screen.

[0140] The touch display screen can be designed as a full-screen, curved screen, or irregularly shaped screen. It can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen; however, this application does not limit the specific design of the touch display screen.

[0141] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, Wireless Fidelity (WiFi) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.

[0142] exist Figure 5In the illustrated electronic device, the processor 110 can be used to call the program for the URL detection method stored in the memory 120, and specifically perform the following operations:

[0143] Determine the set of URLs to be detected;

[0144] Website association features are extracted from the set of URLs to be detected to obtain the website-related features corresponding to the set of URLs to be detected.

[0145] Based on the website-related features, website feature values ​​are calculated for each URL to be detected to obtain website fingerprint information;

[0146] Based on the website fingerprint information, the same-origin URL detection process is performed on the set of URLs to be detected to obtain the same-origin URL detection results corresponding to the set of URLs to be detected.

[0147] In some embodiments, when the processor 110 performs the step of extracting website association features from the set of URLs to be detected to obtain website-related features corresponding to the set of URLs to be detected, it specifically performs the following operations:

[0148] Based on the set of URLs to be detected, feature extraction is performed to obtain the reference website feature values ​​corresponding to each URL to be detected;

[0149] Based on the feature values ​​of all the reference websites, correlation feature detection processing is performed to obtain the website-related features corresponding to the set of URLs to be detected.

[0150] In some embodiments, when the processor 110 performs the step of extracting features based on the set of URLs to be detected to obtain the reference website feature values ​​corresponding to each URL to be detected, it specifically performs the following operations:

[0151] Determine the preset website characteristics;

[0152] Based on the preset website features, feature extraction processing is performed on the URLs to be detected in the set of URLs to be detected to obtain reference website feature values ​​corresponding to each URL to be detected.

[0153] In one embodiment, when the processor 110 performs the step of calculating website feature values ​​for each URL to be detected based on the website-related features to obtain website fingerprint information, it specifically performs the following operations:

[0154] Based on the aforementioned website-related features, obtain the original website feature values ​​corresponding to each URL to be detected;

[0155] Based on the original website feature values, determine the comprehensive feature vector corresponding to the relevant features of the website;

[0156] The dimensionality reduction process based on the comprehensive feature vector is used to obtain the similarity hash value corresponding to each URL to be detected, and the similarity hash value is used as the website fingerprint information of each URL to be detected.

[0157] In some embodiments, when the processor 110 performs the step of determining the comprehensive feature vector corresponding to the website-related features based on the original website feature values, it specifically performs the following operations:

[0158] The target website feature values ​​are obtained by performing feature filtering processing on the original website feature values;

[0159] Based on the target website feature value, a hash calculation is performed to obtain the reference hash value corresponding to the target website feature value, and a weight allocation process is performed on the target website feature value to obtain the reference weight;

[0160] Based on the reference hash value and the reference weight, a vector mapping process is performed to obtain the reference feature vector corresponding to the target website feature value;

[0161] Based on the reference feature vector, a vector merging process is performed to obtain the comprehensive feature vector corresponding to the website-related features.

[0162] In some embodiments, when the processor 110 performs the step of weighting the feature values ​​of the target website to obtain reference weights, it specifically performs the following operations:

[0163] Determine the feature extraction position corresponding to the feature value of the target website in the URL to be detected;

[0164] Obtain the target website-related features corresponding to the target website feature values;

[0165] The feature extraction location, the relevant features of the target website, and the feature value of the target website are input into the weight allocation model for weight allocation processing to obtain the reference weight.

[0166] In some embodiments, when the processor 110 performs the step of performing same-origin URL detection processing on the set of URLs to be detected based on the website fingerprint information to obtain the same-origin URL detection result corresponding to the set of URLs to be detected, it specifically performs the following operations:

[0167] Obtain the similar hash value corresponding to each URL to be detected from the website fingerprint information;

[0168] The target URL to be detected is determined from the set of URLs to be detected. The target similar hash value corresponding to the target URL to be detected is determined from all the similar hash values. The first similar hash value matching the target similar hash value is determined from all the similar hash values. The URL to be detected corresponding to the first similar hash value is determined as a similar URL from the same source as the target URL to be detected.

[0169] From all the similar hash values, determine a second similar hash value that does not match the target similar hash value, and determine the URL to be detected corresponding to the second similar hash value as a non-same-origin similar URL to the target URL to be detected;

[0170] Based on the target URL to be detected, the same-origin similar URLs, and the non-same-origin similar URLs, generate the same-origin URL detection results corresponding to the set of URLs to be detected.

[0171] This application also provides a computer-readable storage medium storing at least one instruction, which is executed by a processor to implement the URL detection method as described in the above embodiments.

[0172] This application also provides a computer program product that stores at least one instruction, which is loaded and executed by the processor to implement the URL detection method described in the above embodiments.

[0173] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.

[0174] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for detecting website addresses, characterized in that, The method includes: Determine the set of URLs to be detected; Website association features are extracted from the set of URLs to be detected to obtain the website-related features corresponding to the set of URLs to be detected. Based on the website-related features, website feature values ​​are calculated for each URL to be detected to obtain website fingerprint information; Based on the website fingerprint information, the same-origin URL detection process is performed on the set of URLs to be detected to obtain the same-origin URL detection result corresponding to the set of URLs to be detected. The step of calculating website feature values ​​for each URL to be detected based on the website-related features to obtain website fingerprint information includes: obtaining the original website feature values ​​corresponding to each URL to be detected based on the website-related features; performing feature filtering processing on the original website feature values ​​to obtain target website feature values; performing hash calculation on the target website feature values ​​to obtain reference hash values ​​corresponding to the target website feature values; performing weight allocation processing on the target website feature values ​​to obtain reference weights; performing vector mapping processing on the reference hash values ​​and the reference weights to obtain reference feature vectors corresponding to the target website feature values; performing vector merging processing on the reference feature vectors to obtain comprehensive feature vectors corresponding to the website-related features; mapping each element value in the comprehensive feature vectors to obtain target binary values ​​using a preset threshold; generating similar hash values ​​corresponding to the URLs to be detected based on the target binary values; and using the similar hash values ​​as the website fingerprint information for each URL to be detected.

2. The method according to claim 1, characterized in that, The step of extracting website association features from the set of websites to be detected to obtain website-related features corresponding to the set of websites to be detected includes: Based on the set of URLs to be detected, feature extraction is performed to obtain the reference website feature values ​​corresponding to each URL to be detected; Based on the feature values ​​of all the reference websites, correlation feature detection processing is performed to obtain the website-related features corresponding to the set of URLs to be detected.

3. The method according to claim 2, characterized in that, The step of extracting features based on the set of URLs to be detected to obtain reference website feature values ​​corresponding to each URL to be detected includes: Determine the preset website characteristics; Based on the preset website features, feature extraction processing is performed on the URLs to be detected in the set of URLs to be detected to obtain reference website feature values ​​corresponding to each URL to be detected.

4. The method according to claim 1, characterized in that, The step of assigning weights to the feature values ​​of the target website to obtain reference weights includes: Determine the feature extraction position corresponding to the feature value of the target website in the URL to be detected; Obtain the target website-related features corresponding to the target website feature values; The feature extraction location, the relevant features of the target website, and the feature value of the target website are input into the weight allocation model for weight allocation processing to obtain the reference weight.

5. The method according to claim 1, characterized in that, The process of performing same-origin URL detection on the set of URLs to be detected based on the website fingerprint information to obtain the same-origin URL detection results corresponding to the set of URLs to be detected includes: Obtain the similar hash value corresponding to each URL to be detected from the website fingerprint information; The target URL to be detected is determined from the set of URLs to be detected. The target similar hash value corresponding to the target URL to be detected is determined from all the similar hash values. The first similar hash value matching the target similar hash value is determined from all the similar hash values. The URL to be detected corresponding to the first similar hash value is determined as a similar URL from the same source as the target URL to be detected. From all the similar hash values, determine a second similar hash value that does not match the target similar hash value, and determine the URL to be detected corresponding to the second similar hash value as a non-same-origin similar URL to the target URL to be detected; Based on the target URL to be detected, the same-origin similar URLs, and the non-same-origin similar URLs, generate the same-origin URL detection results corresponding to the set of URLs to be detected.

6. A website address detection device, characterized in that, The device includes: The URL determination module is used to determine the set of URLs to be detected; The feature extraction module is used to extract website association features from the set of URLs to be detected, and obtain the website-related features corresponding to the set of URLs to be detected. The fingerprint calculation module is used to calculate website feature values ​​for each URL to be detected based on the relevant website features, and obtain website fingerprint information. The URL detection module is used to perform same-origin URL detection processing on the set of URLs to be detected based on the website fingerprint information, and obtain the same-origin URL detection result corresponding to the set of URLs to be detected; The step of calculating website feature values ​​for each URL to be detected based on the website-related features to obtain website fingerprint information includes: obtaining the original website feature values ​​corresponding to each URL to be detected based on the website-related features; performing feature filtering processing on the original website feature values ​​to obtain target website feature values; performing hash calculation on the target website feature values ​​to obtain reference hash values ​​corresponding to the target website feature values; performing weight allocation processing on the target website feature values ​​to obtain reference weights; performing vector mapping processing on the reference hash values ​​and the reference weights to obtain reference feature vectors corresponding to the target website feature values; performing vector merging processing on the reference feature vectors to obtain comprehensive feature vectors corresponding to the website-related features; mapping each element value in the comprehensive feature vectors to obtain target binary values ​​using a preset threshold; generating similar hash values ​​corresponding to the URLs to be detected based on the target binary values; and using the similar hash values ​​as the website fingerprint information for each URL to be detected.

7. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions, which are adapted to be loaded by a processor and executed as described in any one of claims 1 to 5.

8. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed as described in any one of claims 1 to 5.