Webpage risk identification method and device, electronic equipment and medium

By using a multi-feature data identification method based on web pages and operational data, and adjusting the second type of indicator data with the first type of indicator data to generate target risk data, the problem of low efficiency and accuracy of web page risk identification in existing technologies is solved, and more efficient and accurate web page risk identification is achieved.

CN114036367BActive Publication Date: 2026-02-13BEIJING DUSHANG SOFTWARE TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111381872.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-19
Publication Date
2026-02-13
Estimated Expiration
2041-11-19

AI Technical Summary

Technical Problem

Existing technologies for identifying web page risks are inefficient and inaccurate.

Method used

By using webpage data and operational data based on the target webpage, multiple types of feature data are identified. The first type of indicator data is used to adjust the second type of indicator data to generate target risk data to characterize the webpage risk situation.

Benefits of technology

It improves the efficiency and accuracy of webpage risk identification and reduces the impact of interference factors such as images on text content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114036367B_ABST
    Figure CN114036367B_ABST
Patent Text Reader

Abstract

The present disclosure provides a webpage risk identification method, device, equipment, medium and product, relates to the technical field of computers, in particular to the technical field of big data, intelligent identification and the like. The webpage risk identification method comprises: determining at least one type of feature data associated with a target webpage based on at least one of webpage data of the target webpage and operation data for the target webpage; determining first index data and second index data based on the at least one type of feature data; and adjusting initial risk data obtained from the second index data to obtain target risk data by using the first index data, wherein the target risk data represents a risk situation existing in the target webpage.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, in particular to the technical field of big data, intelligent identification and the like, and more particularly to a webpage risk identification method and device, an electronic device, a medium and a program product. BACKGROUND

[0002] In the related art, a webpage is usually identified to determine whether the webpage has a risk. For example, when it is identified that the webpage has illegal advertisements, it can be determined that the webpage has a risk. The webpage risk identification in the related art has low identification efficiency and low identification accuracy. SUMMARY

[0003] The present disclosure provides a webpage risk identification method and device, an electronic device, a storage medium and a program product.

[0004] According to an aspect of the present disclosure, a webpage risk identification method is provided, including: determining at least one type of feature data associated with a target webpage based on at least one of webpage data of the target webpage and operation data for the target webpage; determining first type of index data and second type of index data based on the at least one type of feature data; and adjusting initial risk data obtained from the second type of index data by using the first type of index data to obtain target risk data, wherein the target risk data represents a risk situation of the target webpage.

[0005] According to another aspect of the present disclosure, a webpage risk identification device is provided, including: a first determination module, a second determination module and an adjustment module. The first determination module is configured to determine at least one type of feature data associated with a target webpage based on at least one of webpage data of the target webpage and operation data for the target webpage. The second determination module is configured to determine first type of index data and second type of index data based on the at least one type of feature data. The adjustment module is configured to adjust initial risk data obtained from the second type of index data by using the first type of index data to obtain target risk data, wherein the target risk data represents a risk situation of the target webpage.

[0006] According to another aspect of the present disclosure, an electronic device is provided, including: at least one processor and a memory connected with the at least one processor in communication. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the webpage risk identification method described above.

[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the webpage risk identification method described above.

[0008] According to another aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the webpage risk identification method described above.

[0009] It should be understood that the contents described in this part are not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The accompanying drawings are used to better understand the present scheme, and do not constitute a limitation on the present disclosure. Among them:

[0011] Figure 1 The system architecture of the webpage risk identification and device according to an embodiment of the present disclosure is schematically shown;

[0012] Figure 2 The flowchart of the webpage risk identification method according to an embodiment of the present disclosure is schematically shown;

[0013] Figure 3 The principle diagram of the webpage risk identification method according to an embodiment of the present disclosure is schematically shown;

[0014] Figure 4 The schematic diagram of determining the first index threshold according to an embodiment of the present disclosure is schematically shown;

[0015] Figure 5 The block diagram of the webpage risk identification device according to an embodiment of the present disclosure is schematically shown; and

[0016] Figure 6 The block diagram of the electronic device for performing webpage risk identification to realize the embodiments of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0017] The exemplary embodiments of the present disclosure are described below in conjunction with the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, the description below omits the description of well-known functions and structures.

[0018] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. As used herein, the terms "comprises", "comprising", "includes", "including" and the like are specifically intended to be open-ended and to mean that other features, steps, operations, and / or components can be added.

[0019] All terms used herein (including technical and scientific terms) have the meanings commonly understood by one of ordinary skill in the art unless otherwise defined herein. It should be noted that the terms used herein are to be interpreted as having a meaning that is consistent with the understanding of that term by those having ordinary skill in the art and are to be interpreted not in an idealized or overly formal sense unless otherwise specifically defined herein.

[0020] In the case of using expressions similar to "at least one of A, B, and C, etc.", it should generally be interpreted to include any of the combinations of the items it is listed as well as the items listed alone (e.g., "a system having at least one of A, B, and C" should be interpreted to include a system having A alone, a system having B alone, a system having C alone, a system having A and B together, a system having A and C together, a system having B and C together, and / or a system having A, B, and C together, etc.).

[0021] Embodiments of the disclosure provide a webpage risk identification method, comprising: determining at least one type of feature data associated with a target webpage based on at least one of webpage data of the target webpage and operation data for the target webpage. Then, determining first type of index data and second type of index data based on the at least one type of feature data. Next, adjusting initial risk data obtained from the second type of index data to obtain target risk data by using the first type of index data, wherein the target risk data represents a risk situation of the target webpage.

[0022] Figure 1 The system architecture of the webpage risk identification and device according to an embodiment of the disclosure is schematically shown. It should be noted that, Figure 1 The shown is only an example of the system architecture to which the embodiments of the disclosure can be applied, to help those skilled in the art understand the technical content of the disclosure, but does not mean that the embodiments of the disclosure cannot be used in other devices, systems, environments or scenarios.

[0023] As Figure 1 The system architecture 100 according to the embodiment can include clients 101, 102, 103, a network 104 and a server 105, as shown. The network 104 is used as a medium to provide a communication link between the clients 101, 102, 103 and the server 105. The network 104 can include various connection types, such as wired, wireless communication links or fiber optic cables, etc.

[0024] A user can use the clients 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the clients 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).

[0025] The clients 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, desktop computers, etc. The clients 101, 102, 103 of the embodiments of the present disclosure can run application programs, for example.

[0026] The server 105 can be a server providing various services, such as a background management server supporting websites browsed by users using the clients 101, 102, 103 (only as an example). The background management server can analyze and process received user requests and other data, and feed back the processing results (such as web pages, information, or data, etc. obtained or generated according to user requests) to the clients. In addition, the server 105 can also be a cloud server, i.e. the server 105 has cloud computing functions.

[0027] It should be noted that the web page risk identification method provided by the embodiments of the present disclosure can be executed by the server 105. Correspondingly, the web page risk identification apparatus provided by the embodiments of the present disclosure can be arranged in the server 105. The web page risk identification method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the clients 101, 102, 103 and / or the server 105. Correspondingly, the web page risk identification apparatus provided by the embodiments of the present disclosure can also be arranged in a server or a server cluster different from the server 105 and capable of communicating with the clients 101, 102, 103 and / or the server 105.

[0028] In one example, the server 105 can obtain web page data for a target web page and operation data for the target web page from the clients 101, 102, 103 through the network 104, and process the web page data and the operation data to determine whether the target web page has a risk situation.

[0029] It should be understood that the number of clients, networks and servers in the system architecture is only illustrative. Any number of clients, networks and servers can be provided according to implementation needs. Figure 1

[0030] The embodiments of the present disclosure provide a web page risk identification method, which is described below in combination with the system architecture of Figure 1 Figures 2-4 ​​A webpage risk identification method according to an example embodiment of the present disclosure is described. The webpage risk identification method of the present embodiment may, for example, be performed by the server shown in Figure 1 Figure 1 The electronic device of the present embodiment may, for example, be the same as or similar to the server shown below.

[0031] Figure 2 A flowchart of a webpage risk identification method according to an example embodiment of the present disclosure is schematically shown.

[0032] As shown in Figure 2 The webpage risk identification method 200 of the present embodiment may, for example, include operations S210-S230.

[0033] At operation S210, at least one type of feature data associated with the target webpage is determined based on at least one of webpage data of the target webpage and operation data for the target webpage.

[0034] At operation S220, first type of index data and second type of index data are determined based on the at least one type of feature data.

[0035] At operation S230, the initial risk data derived from the second type of index data is adjusted by the first type of index data to obtain target risk data.

[0036] The webpage data of the target webpage may, for example, include attributes of the webpage, content categories of the webpage, content quantities of each content category, and the like. The content categories may, for example, include a text category, a picture category, and the like, and the content quantities of the text category may, for example, include a number of texts, and the content quantities of the picture category may, for example, include a number of pictures, and the like.

[0037] The operation data for the target webpage may, for example, include data obtained by a user operating the target webpage, and the operation may, for example, include a search operation, an operation of clicking on related content in the target webpage, and the like.

[0038] After obtaining the webpage data and the operation data, at least one of the webpage data and the operation data may be processed to obtain at least one type of feature data associated with the target webpage. The at least one type of feature data may include multiple types of feature data.

[0039] Then, the at least one type of feature data is processed to obtain multiple types of index data. The multiple types of index data may, for example, include first type of index data and second type of index data. The first type of index data may, for example, be used to represent that the target webpage is a safe webpage, and the second type of index data may, for example, be used to represent that the target webpage is a risk webpage.

[0040] ​Next, the initial risk data can be determined by the second type of index data, which can represent whether the target webpage has risks to some extent. Then, the initial risk data is adjusted by the first type of index data to obtain the target risk data representing whether the target webpage has risks. In an example, when the initial risk data obtained by the second type of index data represents that the target webpage is a risk webpage, if the first type of index data represents that the target webpage is a safe webpage, the target risk data obtained by adjusting the initial risk data by the first type of index data can represent that the target webpage is a safe webpage. In another example, when the initial risk data represents that the target webpage is a risk webpage, if the first type of index data cannot represent that the target webpage is a safe webpage, the target risk data obtained by adjusting the initial risk data by the first type of index data can represent that the target webpage is a dangerous webpage or a safe webpage.

[0041] According to an embodiment of the present disclosure, a plurality of types of feature data are obtained by processing webpage data and operation data, and then the first type of index data and the second type of index data are determined based on the plurality of types of feature data, and the initial risk data for the target webpage is determined by the second type of index data, and the target risk data representing whether the target webpage has risks is obtained by adjusting the initial risk data by the first type of index data. It can be understood that the technical solution of the embodiment of the present disclosure improves the identification efficiency and accuracy of identifying whether the target webpage has risks.

[0042] Figure 3 The principle diagram of the webpage risk identification method according to an embodiment of the present disclosure is schematically shown.

[0043] As shown in Figure 3 , for the target webpage 310, the webpage data 321 and the operation data 322 of the target webpage 310 are determined.

[0044] Exemplarily, based on the operation data 322 of the target webpage 310, the first type of feature data 331 is determined, which represents the search information associated with the target webpage 310.

[0045] For example, based on various raw data of the entire process of inputting the search term by the user to opening the target webpage 310, the first type of feature data 331 is determined. The raw data includes, for example, the search term, the target webpage consumption data on the day, and the target webpage 7-day consumption data.

[0046] For example, a keyword classification table is established, and fuzzy matching technology is used to match the input search keyword with the classification table to determine the search keyword as a risk keyword, a safe keyword, or an unknown keyword. In principle, the smaller the proportion of unknown keywords, the better. For each search record, the search record can be classified into risk traffic (corresponding to a risk keyword), safe traffic (corresponding to a safe keyword), or unknown traffic (corresponding to an unknown keyword) according to the type of search keyword. For the target webpage 310, which usually has multiple search records, the total number of risk traffic and the total number of safe traffic for the target webpage 310 can be counted every day. One traffic can be regarded as one search record.

[0047] The target webpage 310 usually has a promoted advertisement, and the first type of feature data 331, for example, reflects the promotion strategy of the advertiser, for example, indirectly reflects the purpose of the promotion advertisement of the advertiser.

[0048] Exemplarily, based on the webpage data 321 of the target webpage 310, the second type of feature data 332 is determined, which, for example, represents the attribute information of the target webpage 310.

[0049] The second type of feature data 332, for example, reflects information other than the text content of the target webpage 310. The information other than the text content, for example, includes the industry of the advertiser, the number of pictures in the webpage, the number of characters, the number of lead conversion components (a lead conversion component is, for example, a component for collecting user information, and the user can interact with the component and the background, and the lead conversion component, for example, includes an information submission component), the number of download components, the total number of business component categories (picture component category, video component category, payment component category, etc.), the number of approved webpages in the account of the advertiser, the number of rejected webpages, and the like. The second type of feature data 332 can be obtained from a database related to the target webpage 310, or can be obtained by analyzing the structural information of the target webpage 310 through crawling technology.

[0050] The second type of feature data 332, for example, reflects the degree of refinement and complexity of the target webpage 310, and can also reflect whether the target webpage 310 has the suspicion of bypassing the audit system in the technical level.

[0051] Exemplarily, based on the operation data 322 for the target webpage 310, the third type of feature data 333 is determined, which, for example, represents the interaction information associated with the target webpage 310.

[0052] The third type of feature data 333, for example, is obtained based on the data of the user interacting with the target webpage 310. The third type of feature data 333, for example, includes the webpage stay duration, whether a form is submitted, whether a phone call is made, whether a communication identifier is copied, whether a customer service is dialoged in a consultation tool, and the like.

[0053] The third type of feature data 333, for example, embodies the reaction of the user to the target webpage 310, and indirectly embodies the quality and credibility of the target webpage 310.

[0054] According to an embodiment of the present disclosure, the plurality of types of feature data are obtained based on the webpage data and the operation data, and the first type of index data and the second type of index data are obtained based on the plurality of types of feature data, which improves the richness of the index data. In addition, compared to webpage risk identification based on text content, which has interference factors such as text content being replaced by pictures, the embodiments of the present disclosure obtain feature data based on information other than the text content of the target webpage, which improves the effect of identifying webpage risks based on the feature data.

[0055] After obtaining the first type of feature data 331, the second type of feature data 332, and the third type of feature data 333, the first type of feature data 331, the second type of feature data 332, and the third type of feature data 333 can be abstracted to obtain the first type of index data 341 and the second type of index data 342.

[0056] The first type of index data 341, for example, is exempt type index data, which is, for example, threshold type index data, that is, a result obtained by calculating a certain feature data or a plurality of feature data for the target webpage 310 is taken as the first type of index data 341. When any one of the index data in the first type of index data 341 falls within a certain threshold range, it indicates that the index data is effective. After the index data is effective, the index data indicates that the target webpage 310 is a safe, risk-free, and legal webpage.

[0057] The first type of index data 341, for example, includes risk traffic proportion (for example, calculated from the total number of risk traffic and the total number of safe traffic in the first type of feature data), account pass page proportion, and the like.

[0058] The second type of index data 342, for example, is risk type index data, which can be threshold type or Boolean type index data. For a certain second type of index data 342, if the index data is Boolean type index data, it indicates that the index data is effective when the target webpage 310 has a certain feature. The effective index data indicates that the target webpage 310 is a risk webpage.

[0059] The second type of index data 342, for example, includes webpage daily conversion number (the number of user interactions through the webpage), daily consumption number, 7-day promotion consumption variance, number of used business components, number of Chinese characters, number of pictures, average height of pictures, average page dwell time, and the like.

[0060] For the first indicator threshold 351 corresponding to the first type of indicator data 341, when the first type of indicator data 341 includes multiple indicator data, each indicator data corresponds to a first indicator threshold. Similarly, for the second indicator threshold 352 corresponding to the second type of indicator data 342, when the second type of indicator data 342 includes multiple indicator data, each indicator data corresponds to a second indicator threshold.

[0061] For example, based on the first type of indicator data 341 and the first indicator threshold 351 corresponding to the first type of indicator data 341, the first weight 361 is determined. For example, taking the risk traffic proportion of the first type of indicator data 341 as an example, the first indicator threshold 351 is, for example, 70%, when the risk traffic proportion of the target webpage 310 is less than the first indicator threshold 351 (70%), it can be indicated that the target webpage 310 hits the exempt type of indicator, that is, the risk traffic proportion indicates that the target webpage 310 is a safe webpage, at this time, the first weight 361 is, for example, 0. When the risk traffic proportion is greater than or equal to the first indicator threshold 351 (70%), it can be indicated that the target webpage 310 does not hit the exempt type of indicator, that is, the risk traffic proportion indicates that the target webpage 310 is a risk webpage, at this time, the first weight 361 is, for example, 1.

[0062] Next, based on the second type of indicator data 342 and the second indicator threshold 352 corresponding to the second type of indicator data 342, the initial risk data 380 is determined.

[0063] For example, the second type of indicator data 342 includes multiple indicator data, based on the multiple indicator data and the multiple second indicator thresholds 352 corresponding to the multiple indicator data one by one, the target indicator data 362 is determined from the multiple indicator data, and based on the target indicator data 362 and the precision rate 371 for the target indicator data 362, the initial risk data 380 is determined.

[0064] For example, taking the second type of indicator data 342 including 3 indicator data as an example, the 3 indicator data includes, for example, webpage average stay time, 7-day promotion consumption variance, and Chinese character number. The second indicator threshold 352 corresponding to the webpage average stay time is, for example, a seconds, the second indicator threshold 352 corresponding to the 7-day promotion consumption variance is, for example, b, and the second indicator threshold 352 corresponding to the Chinese character number is, for example, c.

[0065] For example, if the average page dwell time of the target webpage 310 is less than the second index threshold 352 (a second), it means that the average page dwell time index is effective (i.e., it means that the target webpage 310 is at risk). If the 7-day promotion consumer variance of the target webpage 310 is greater than the second index threshold 352 (b), it means that the 7-day promotion consumer variance index is effective (i.e., it means that the target webpage 310 is at risk). If the number of Chinese characters of the target webpage 310 is less than the second index threshold 352 (c), it means that the number of Chinese characters index is effective (i.e., it means that the target webpage 310 is at risk).

[0066] Next, the index data representing that the target webpage 310 is at risk is taken as the target index data 362, for example, the target index data 362 includes the average page dwell time, the 7-day promotion consumer variance. Each target index data 362 corresponds to a precision rate 371, for example, the precision rate 371 is obtained from a plurality of historical webpages, and the specific process will be described later.

[0067] Based on the precision rate 371 of the target index data 362, the second weight 372 of the target index data 362 is determined. Then, based on the precision rate 371 and the second weight 372 of the target index data 362, the initial risk data 380 is determined.

[0068] For example, the plurality of target index data 362 (average page dwell time, 7-day promotion consumer variance) is sorted according to the precision rate, and the higher the precision rate, the higher the sorting. The second weight of the target index data 362 with higher sorting is greater. The initial risk data 380 is obtained from the following formula (1), for example.

[0069]

[0070] Wherein, p i is the precision rate of the i-th target index data; w i is the second weight of the i-th target index data; n is the number of target index data; P is the initial risk data 380 (also referred to as risk coefficient).

[0071] Then, the initial risk data 380 is adjusted by the first weight 361 to obtain the target risk data 390. For example, the first weight 361 is multiplied by the initial risk data 380 (risk coefficient) to obtain the target risk data 390.

[0072] It can be understood that when the target webpage 310 hits the first type of index data 341 (exempt type index), the first type of index data 341 represents that the target webpage 310 is a safe webpage, and at this time, the first weight 361 is for example 0, and the target risk data 390 obtained by multiplying the initial risk data 380 by the first weight 361 is also for example 0, indicating that the target webpage 310 is a safe webpage. When the target webpage 310 does not hit the first type of index data 341 (exempt type index), the first type of index data 341 fails to represent that the target webpage 310 is a safe webpage, and at this time, the first weight 361 is for example 1, and the value of the target risk data 390 obtained by multiplying the initial risk data 380 by the first weight 361 is for example the value of the initial risk data 380 (risk coefficient), at which time whether the target webpage 310 is at risk is determined based on the specific value of the initial risk data 380 (risk coefficient).

[0073] According to an embodiment of the present disclosure, the initial risk data for the target webpage is obtained by the second type of index data, and then the initial risk data is adjusted based on the first weight obtained by the first type of index data to obtain the target risk data, which improves the accuracy of the target risk data, and further realizes the accuracy of determining whether the target webpage is at risk based on the target risk data.

[0074] For each index data in the first type of index data, the recall rate and the precision rate for each index data are calculated, and the first index threshold of each index data is calculated based on the recall rate and the precision rate. The specific process is as follows.

[0075] For example, M historical webpages are obtained, and the M historical webpages include M1 risk historical webpages, and M and M1 are both integers greater than 0. Then, P initial thresholds are set for the first type of index data, and P is an integer greater than 0. The first type of index data includes for example a plurality of index data, and P initial thresholds can be set for each index data, and the number of initial thresholds set for each index data can be the same or different.

[0076] Based on the M1 risk historical webpages and the P initial thresholds, the recall rate for each index data in the first type of index data is determined. For example, each index data for the M1 risk historical webpages is compared with each of the P initial thresholds corresponding to the index data, to determine m1 risk historical webpages from the M1 risk historical webpages, and m1 is an integer greater than 0. Then, the ratio between m1 and M1 is determined as the recall rate for the index data.

[0077] For example, the first type of index data is risk traffic proportion, and multiple first index thresholds 60%, 65%, 70%, and 75% are set for the index data. M historical webpages are, for example, 50,000 historical webpages, of which 30,000 (M1) historical webpages are, for example, labeled as risk historical webpages, and the remaining 20,000 (M2) historical webpages are labeled as safe historical webpages, and M2 is an integer greater than 0. For the first index threshold 60%, for example, 25,000 (m1) of the 30,000 risk historical webpages have an index risk traffic proportion greater than 60%, and the recall rate is, for example, 25,000 / 30,000. For the first index threshold 65%, for example, 20,000 (m1) of the 30,000 risk historical webpages have an index risk traffic proportion greater than 65%, and the recall rate is 20,000 / 30,000. For the first index thresholds 70% and 75%, the same applies. The recall rate is shown in formula (2).

[0078]

[0079] wherein M1 represents the number of risk historical webpages; and m1 may, for example, represent the number of risk historical webpages that meet a certain index data in the first type of index data.

[0080] For each index data in the first type of index data, based on the M historical webpages and the P initial thresholds, the precision rate for the first type of index data is determined.

[0081] For example, the M historical webpages further include M2 safe historical webpages, each index data (first type of index data) for the M2 safe historical webpages is compared with each initial threshold in the P initial thresholds corresponding to the index data, to determine m2 safe historical webpages from the M2 safe historical webpages, and m2 is an integer greater than 0. Then, the ratio between the sum of m1 and m2 and M is determined as the precision rate for the index data.

[0082] Taking the risk traffic ratio as an example, multiple first indicator thresholds of 60%, 65%, 70%, and 75% are set for this indicator data. M historical web pages are, for example, 50,000 historical web pages, of which 30,000 (M1) are marked as risky historical web pages, and the remaining 20,000 (M2) are marked as safe historical web pages. For the first indicator threshold of 60%, for example, 10,000 (m1) of the 30,000 risky historical web pages have a risk traffic ratio greater than or equal to 60%, and for example, 15,000 (m2) of the 20,000 safe historical web pages have a risk traffic ratio less than 60%. Therefore, the accuracy rate of this risk traffic ratio indicator data is (10,000 + 15,000) / 50,000. The same applies to the first indicator thresholds of 65%, 70%, and 75%. The accuracy rate is shown in formula (3).

[0083]

[0084] Where M represents the number of historical web pages; m1 can represent, for example, the number of risky historical web pages that hit a certain indicator in the first type of indicator data; m2 can represent, for example, the number of safe historical web pages that did not hit a certain indicator in the first type of indicator data.

[0085] After obtaining the recall and precision for each metric in the first category of metric data, based on the recall and precision for the first category of metric data and the P initial thresholds corresponding to each metric, the first metric threshold for each metric is determined, as follows: Figure 4 As shown.

[0086] Figure 4 A schematic diagram illustrating the determination of a first index threshold according to an embodiment of the present disclosure is shown.

[0087] like Figure 4 As shown, the horizontal axis represents, for example, recall, and the vertical axis represents, for example, precision. For a specific metric in the first category of data, multiple thresholds for the first metric (60%, 65%, 70%, 75%) are set, and the recall and precision corresponding to each threshold are calculated. Then, the ROC (Receiver Operating Characteristic) curve is used to find the optimal threshold for the first metric. For example, as... Figure 4As shown, for a certain index data, a tangent line of a curve is used to determine 70% as the final first index threshold for the index data. It can be understood that the use of the ROC curve to determine the final first index threshold is only an example, and the embodiments of the present disclosure do not specifically limit the determination manner of the final first index threshold. The final first index threshold is used to determine the first weight mentioned above.

[0088] For each index data of the first type of index data, the higher the recall rate, the higher the credibility of the index data; the higher the precision rate, the better the identification performance of the index data.

[0089] According to the embodiments of the present disclosure, the first index threshold for the first type of index data is determined based on the recall rate and the accuracy rate, so as to determine the first weight for the target webpage based on the first index threshold, thereby improving the determination accuracy of the first weight.

[0090] For each index data in the second type of index data, the precision rate for each index data is calculated, and the second index threshold for each index data is calculated based on the precision rate. The specific process is as follows.

[0091] For example, N historical webpages are obtained, N being an integer greater than 0. Then, Q initial thresholds are set for the second type of index data, Q being an integer greater than 0. The second type of index data, for example, includes a plurality of index data, and Q initial thresholds can be set for each index data. The number of initial thresholds set for each index data can be the same or different. Next, based on the N historical webpages and the Q initial thresholds, the precision rate for each index data in the second type of index data is determined, and based on the precision rate of each index data and the Q initial thresholds, the second index threshold for each index data is determined. The N historical webpages may, for example, be the same as the M historical webpages described above.

[0092] For example, the N historical webpages include N1 risk historical webpages and N2 safe historical webpages, N1 and N2 being integers greater than 0.

[0093] Each index data (second type of index data) for the N1 risk historical webpages is compared with each initial threshold corresponding to each index data in the Q initial thresholds, so as to determine n1 risk historical webpages from the N1 risk historical webpages, n1 being an integer greater than 0.

[0094] Each index data (second type of index data) for the N2 safe historical webpages is compared with each initial threshold corresponding to each index data in the Q initial thresholds, so as to determine n2 safe historical webpages from the N2 safe historical webpages, n2 being an integer greater than 0.

[0095] The ratio between the sum of n1 and n2 and N is determined as the precision rate of the index data. The precision rate is shown in formula (4). After obtaining the precision rate of each index data (second type index data), the second weight mentioned above can be determined based on the precision rate.

[0096]

[0097] Wherein, N represents the number of historical webpages; n1 may represent the number of risk historical webpages that hit a certain index data in the second type index data, for example; n2 may represent the number of safe historical webpages that miss a certain index data in the second type index data, for example. It can be understood that the calculation process of the precision rate of a certain index data in the second type index data is similar to the calculation process of the precision rate of a certain index data in the first type index data shown in formula (3), and will not be described in detail here.

[0098] For a certain index data in the second type index data, a plurality of second index thresholds are set for the index data, and after the precision rate corresponding to each second index threshold is calculated by formula (4), the second index threshold corresponding to the maximum precision rate is determined as the final second index threshold for the index data. The higher the precision rate corresponding to the index data, the more important the index data is, and the greater the contribution of the index data to the risk identification of the webpage is.

[0099] According to an embodiment of the present disclosure, the second index threshold for the second type index data is determined based on the accuracy rate, so as to determine the initial risk data (risk coefficient) for the target webpage based on the second index threshold, thereby improving the determination accuracy of the initial risk data (risk coefficient).

[0100] In another embodiment of the present disclosure, for a plurality of historical webpages used to determine the first index threshold and the second index threshold, the number of historical webpages (for example, the number of historical webpages is increased) and the labeling information of historical webpages (the labeling information indicates that the historical webpages are risk historical webpages or safe historical webpages) can be adjusted by manual verification. The first index threshold, the second index threshold, the first weight, and the second weight are iteratively determined based on the adjusted plurality of historical webpages, thereby improving the accuracy of risk webpage identification.

[0101] Figure 5 A block diagram of a webpage risk identification apparatus according to an embodiment of the present disclosure is schematically shown.

[0102] As Figure 5 shown, the webpage risk identification apparatus 500 of the embodiment of the present disclosure, for example, includes a first determination module 510, a second determination module 520, and an adjustment module 530.

[0103] The first determining module 510 can be configured to determine at least one type of feature data associated with the target webpage based on at least one of the webpage data of the target webpage and the operation data for the target webpage. According to an embodiment of the present disclosure, the first determining module 510 may, for example, perform the operation S210 described above and will not be described here again. Figure 2 The operation S210 described above will not be described here again.

[0104] The second determining module 520 can be configured to determine the first type of index data and the second type of index data based on the at least one type of feature data. According to an embodiment of the present disclosure, the second determining module 520 may, for example, perform the operation S220 described above and will not be described here again. Figure 2 The operation S220 described above will not be described here again.

[0105] The adjusting module 530 can be configured to adjust the initial risk data obtained by the second type of index data to obtain the target risk data by using the first type of index data, where the target risk data represents a risk situation of the target webpage. According to an embodiment of the present disclosure, the adjusting module 530 may, for example, perform the operation S230 described above and will not be described here again. Figure 2 The operation S230 described above will not be described here again.

[0106] According to an embodiment of the present disclosure, the adjusting module 530 includes a first determining sub-module, a second determining sub-module, and an adjusting sub-module. The first determining sub-module is configured to determine a first weight based on the first type of index data and a first index threshold corresponding to the first type of index data. The second determining sub-module is configured to determine the initial risk data based on the second type of index data and a second index threshold corresponding to the second type of index data. The adjusting sub-module is configured to adjust the initial risk data by using the first weight to obtain the target risk data.

[0107] According to an embodiment of the present disclosure, the second type of index data includes a plurality of index data, and the second determining sub-module includes a first determining unit and a second determining unit. The first determining unit is configured to determine target index data from the plurality of index data based on the plurality of index data and a plurality of second index thresholds corresponding to the plurality of index data one by one. The second determining unit is configured to determine the initial risk data based on the target index data and a precision rate for the target index data.

[0108] According to an embodiment of the present disclosure, the second determining unit includes a first determining sub-unit and a second determining sub-unit. The first determining sub-unit is configured to determine a second weight for the target index data based on the precision rate for the target index data. The second determining sub-unit is configured to determine the initial risk data based on the precision rate for the target index data and the second weight.

[0109] According to an embodiment of the present disclosure, the first index threshold corresponding to the first type of index data is obtained by the following modules: a first obtaining module, configured to obtain M historical webpages, wherein the M historical webpages include M1 risk historical webpages, and M and M1 are integers greater than 0; a first setting module, configured to set P initial thresholds for the first type of index data, and P is an integer greater than 0; a third determining module, configured to determine a recall rate for the first type of index data based on the M1 risk historical webpages and the P initial thresholds; a fourth determining module, configured to determine a precision rate for the first type of index data based on the M historical webpages and the P initial thresholds; and a fifth determining module, configured to determine the first index threshold based on the recall rate for the first type of index data, the precision rate for the first type of index data, and the P initial thresholds.

[0110] According to an embodiment of the present disclosure, the third determining module includes a first comparison submodule and a third determining submodule. The first comparison submodule is configured to compare the first type of index data of the M1 risk historical webpages with each of the P initial thresholds to determine m1 risk historical webpages from the M1 risk historical webpages, and m1 is an integer greater than 0. The third determining submodule is configured to determine a ratio between m1 and M1 as the recall rate for the first type of index data.

[0111] According to an embodiment of the present disclosure, the M historical webpages further include M2 safe historical webpages, and M2 is an integer greater than 0. The fourth determining module includes a second comparison submodule and a fourth determining submodule. The second comparison submodule is configured to compare the first type of index data of the M2 safe historical webpages with each of the P initial thresholds to determine m2 safe historical webpages from the M2 safe historical webpages, and m2 is an integer greater than 0. The fourth determining submodule is configured to determine a ratio between a sum of m1 and m2 and M as the precision rate for the first type of index data.

[0112] According to an embodiment of the present disclosure, the second index threshold corresponding to the second type of index data is obtained by the following modules: a second obtaining module, configured to obtain N historical webpages, and N is an integer greater than 0; a second setting module, configured to set Q initial thresholds for the second type of index data, and Q is an integer greater than 0; a sixth determining module, configured to determine a precision rate for the second type of index data based on the N historical webpages and the Q initial thresholds; and a seventh determining module, configured to determine the second index threshold based on the precision rate for the second type of index data and the Q initial thresholds.

[0113] According to an embodiment of the present disclosure, the N historical webpages include N1 risky historical webpages and N2 safe historical webpages, N1 and N2 are integers greater than 0; the sixth determining module includes a third comparing sub-module, a fourth comparing sub-module and a fifth determining sub-module. The third comparing sub-module is configured to compare the second type index data for the N1 risky historical webpages with each of the Q initial thresholds, to determine n1 risky historical webpages from the N1 risky historical webpages, n1 is an integer greater than 0; the fourth comparing sub-module is configured to compare the second type index data for the N2 safe historical webpages with each of the Q initial thresholds, to determine n2 safe historical webpages from the N2 safe historical webpages, n2 is an integer greater than 0; and the fifth determining sub-module is configured to determine, as the precision rate for the second type index data, a ratio between a sum of n1 and n2 and N.

[0114] According to an embodiment of the present disclosure, the first determining module 510 includes at least one of a sixth determining sub-module, a seventh determining sub-module and an eighth determining sub-module. The sixth determining sub-module is configured to determine first type feature data based on the operation data for the target webpage, wherein the first type feature data represents search information associated with the target webpage; the seventh determining sub-module is configured to determine second type feature data based on the webpage data of the target webpage, wherein the second type feature data represents attribute information of the target webpage; and the eighth determining sub-module is configured to determine third type feature data based on the operation data for the target webpage, wherein the third type feature data represents interaction information associated with the target webpage.

[0115] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in the technical solution comply with relevant laws and regulations and do not violate public order and good customs.

[0116] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.

[0117] Figure 6 is a block diagram of an electronic device for performing webpage risk identification used to implement an embodiment of the present disclosure.

[0118] Figure 6A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device 600 is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the implementations of the present disclosure described and / or claimed in this document.

[0119] As shown in Figure 6 The device 600 includes a computing unit 601 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 602 or a computer program loaded into a random access memory (RAM) 603 from a storage unit 608. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0120] Various components in the device 600 are connected to the I / O interface 605, including an input unit 606, such as a keyboard, a mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, an optical disk, etc.; and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0121] The computing unit 601 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, and the like. The computing unit 601 performs various methods and processes described above, such as the webpage risk identification method. For example, in some embodiments, the webpage risk identification method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded onto the RAM 603 and executed by the computing unit 601, one or more steps of the webpage risk identification method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the webpage risk identification method by any other appropriate means, such as by means of firmware.

[0122] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0123] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, implements the functions / acts specified in the flowcharts and / or block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine and partially on a remote machine or entirely on a remote machine or server.

[0124] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0125] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0126] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0127] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.

[0128] It should be understood that the various forms of flow shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure can be achieved, which is not limited herein.

[0129] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.

Claims

1. A webpage risk identification method, comprising: determining at least one type of feature data associated with a target webpage based on webpage data of the target webpage and operation data for the target webpage; determining first type of index data and second type of index data based on the at least one type of feature data, wherein the first type of index data is used to represent that the target webpage is a safe webpage, and the second type of index data is used to represent that the target webpage is a risk webpage; determining a first weight based on the first type of index data and a first index threshold corresponding to the first type of index data; determining a target index data from the second type of index data representing that the target webpage has a risk based on a plurality of the second type of index data and a plurality of second index thresholds corresponding to the plurality of the second type of index data; determining a second weight for the target index data based on a precision rate for the target index data, wherein the second weight is positively correlated with the precision rate; determining initial risk data based on the precision rate for the target index data and the second weight; adjusting the initial risk data by using the first weight to obtain target risk data, wherein the target risk data represents a risk situation of the target webpage.

2. The method of claim 1, wherein, The first index threshold corresponding to the first type of index data is obtained by: obtaining M historical webpages, wherein the M historical webpages include M1 risk historical webpages, and M and M1 are integers greater than 0; setting P initial thresholds for the first type of index data, wherein P is an integer greater than 0; determining a recall rate for the first type of index data based on the M1 risk historical webpages and the P initial thresholds; determining a precision rate for the first type of index data based on the M historical webpages and the P initial thresholds; and determining the first index threshold based on the recall rate for the first type of index data, the precision rate for the first type of index data, and the P initial thresholds.

3. The method of claim 2, wherein, The determination of the recall rate for the first type of index data based on the M1 risk historical webpages and the P initial thresholds comprises: comparing the first type of index data for the M1 risk historical webpages with each of the P initial thresholds to determine m1 risk historical webpages from the M1 risk historical webpages, wherein m1 is an integer greater than 0; and determining a ratio between m1 and M1 as the recall rate for the first type of index data.

4. The method of claim 3, wherein, The M historical webpages further include M2 safe historical webpages, wherein M2 is an integer greater than 0; and the determination of the precision rate for the first type of index data based on the M historical webpages and the P initial thresholds comprises: comparing the first type of index data for the M2 safe historical webpages with each of the P initial thresholds to determine m2 safe historical webpages from the M2 safe historical webpages, wherein m2 is an integer greater than 0; and A ratio between a sum of the m1 and the m2 and the M is determined as a precision rate for the first type of index data.

5. The method of claim 1, wherein, A second index threshold corresponding to the second type of index data is obtained by: obtaining N historical webpages, N being an integer greater than 0; setting Q initial thresholds for the second type of index data, Q being an integer greater than 0; determining a precision rate for the second type of index data based on the N historical webpages and the Q initial thresholds; and determining the second index threshold based on the precision rate for the second type of index data and the Q initial thresholds. The N historical webpages include N1 risk historical webpages and N2 safe historical webpages, N1 and N2 both being integers greater than 0; and the determining of the precision rate for the second type of index data based on the N historical webpages and the Q initial thresholds includes:

6. The method of claim 5, wherein, comparing the second type of index data for the N1 risk historical webpages with each of the Q initial thresholds to determine n1 risk historical webpages from the N1 risk historical webpages, n1 being an integer greater than 0; comparing the second type of index data for the N2 safe historical webpages with each of the Q initial thresholds to determine n2 safe historical webpages from the N2 safe historical webpages, n2 being an integer greater than 0; and determining a ratio between a sum of the n1 and the n2 and the N as the precision rate for the second type of index data. The determining of the at least one type of feature data associated with the target webpage based on at least one of webpage data of the target webpage and operation data for the target webpage includes at least one of:

7. The method of any of claims 1-6, wherein, determining first type of feature data based on the operation data for the target webpage, wherein the first type of feature data represents search information associated with the target webpage; determining second type of feature data based on the webpage data of the target webpage, wherein the second type of feature data represents attribute information of the target webpage; and determining third type of feature data based on the operation data for the target webpage, wherein the third type of feature data represents interaction information associated with the target webpage.

8. A webpage risk identification apparatus, comprising: a first determining module configured to determine at least one type of feature data associated with a target webpage based on at least one of webpage data of the target webpage and operation data for the target webpage; a second determining module configured to determine first type of index data and second type of index data based on the at least one type of feature data, the first type of index data being used to represent that the target webpage is a safe webpage and the second type of index data being used to represent that the target webpage is a risk webpage; and an adjusting module, comprising: a first determining submodule configured to determine a first weight based on the first type of index data and a first index threshold corresponding to the first type of index data. ​ ​ The first determining unit is configured to determine, based on the second type index data and the second type index threshold corresponding to the second type index data, the second type index data representing the risk of the target webpage from the second type index data as the target index data The first determining subunit is configured to determine the second weight for the target index data based on the precision rate of the target index data, and the second weight is positively correlated with the precision rate. The second determining subunit is configured to determine the initial risk data based on the precision rate of the target index data and the second weight. The adjusting sub-module is configured to adjust the initial risk data by using the first weight to obtain target risk data, wherein the target risk data represents the risk of the target webpage.

9. The apparatus of claim 8, wherein, The first index threshold corresponding to the first type index data is obtained by the following modules: The first obtaining module is configured to obtain M historical webpages, wherein the M historical webpages include M1 risk historical webpages, and M and M1 are integers greater than 0. The first setting module is configured to set P initial thresholds for the first type index data, and P is an integer greater than 0. The third determining module is configured to determine the recall rate for the first type index data based on the M1 risk historical webpages and the P initial thresholds. The fourth determining module is configured to determine the precision rate for the first type index data based on the M historical webpages and the P initial thresholds. The fifth determining module is configured to determine the first index threshold based on the recall rate for the first type index data, the precision rate for the first type index data, and the P initial thresholds.

10. The apparatus of claim 9, wherein, The third determining module includes: The first comparison sub-module is configured to compare the first type index data of the M1 risk historical webpages with each of the P initial thresholds to determine m1 risk historical webpages from the M1 risk historical webpages, and m1 is an integer greater than 0. The third determining sub-module is configured to determine the ratio between m1 and M1 as the recall rate for the first type index data.

11. The apparatus of claim 10, wherein, The M historical webpages further include M2 safe historical webpages, and M2 is an integer greater than 0. The fourth determining module includes: The second comparison sub-module is configured to compare the first type index data of the M2 safe historical webpages with each of the P initial thresholds to determine m2 safe historical webpages from the M2 safe historical webpages, and m2 is an integer greater than 0.

12. The apparatus of claim 8, wherein, The fourth determining sub-module is configured to determine the ratio between the sum of m1 and m2 and M as the precision rate for the first type index data. The second index threshold corresponding to the second type index data is obtained by the following modules: The second obtaining module is configured to obtain N historical webpages, and N is an integer greater than 0. The second setting module is configured to set Q initial thresholds for the second type index data, and Q is an integer greater than 0. a sixth determining module, configured to determine a precision rate for the second type of index data based on the N historical webpages and the Q initial thresholds; and a seventh determining module, configured to determine the second index threshold based on the precision rate for the second type of index data and the Q initial thresholds.

13. The apparatus of claim 12, wherein, The N historical webpages include N1 risk historical webpages and N2 safe historical webpages, N1 and N2 are integers greater than 0; the sixth determining module includes: a third comparing sub-module, configured to compare the second type of index data for the N1 risk historical webpages with each of the Q initial thresholds to determine n1 risk historical webpages from the N1 risk historical webpages, n1 is an integer greater than 0; a fourth comparing sub-module, configured to compare the second type of index data for the N2 safe historical webpages with each of the Q initial thresholds to determine n2 safe historical webpages from the N2 safe historical webpages, n2 is an integer greater than 0; and a fifth determining sub-module, configured to determine a ratio between a sum of the n1 and the n2 and the N as the precision rate for the second type of index data.

14. The apparatus of any of claims 8-13, wherein, The first determining module includes at least one of: a sixth determining sub-module, configured to determine first type of feature data based on operation data for the target webpage, wherein the first type of feature data represents search information associated with the target webpage; a seventh determining sub-module, configured to determine second type of feature data based on webpage data of the target webpage, wherein the second type of feature data represents attribute information of the target webpage; and an eighth determining sub-module, configured to determine third type of feature data based on operation data for the target webpage, wherein the third type of feature data represents interaction information associated with the target webpage.

15. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-7.

16. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-7.

17. A computer program product, comprising a computer program which, when executed by a processor, implements the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Applet risk identification method and device

    CN112148603A