Method and System for Detecting Harmful Content on WEB Pages with Privacy Protection

By optimizing the hash function group in the learning-type Bloom filter, the problems of low detection efficiency of encrypted content and insufficient user privacy protection in the prior art are solved, and efficient and accurate detection of harmful information is achieved.

CN119577523BActive Publication Date: 2025-07-01NAT COMPUTER NETWORK & INFORMATION SECURITY MANAGEMENT CENT ZHEJIANG BRANCH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510130013.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-05
Publication Date
2025-07-01
Estimated Expiration
2045-02-05

AI Technical Summary

Technical Problem

The prior art is not efficient in detecting encrypted content, has poor accuracy, and cannot effectively protect user privacy, which poses security risks.

Method used

The hash function group in the learning Bloom filter is used to encrypt and calculate the keyword set, and the hash function group is optimized through machine learning algorithms and locally sensitive hash algorithms to achieve efficient detection of harmful information.

Benefits of technology

It improves the efficiency and accuracy of harmful information detection, protects user privacy, and reduces security risks during the detection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119577523B_ABST
    Figure CN119577523B_ABST
Patent Text Reader

Abstract

Method and system for detecting harmful content on a privacy-protected WEB page. The method includes: configuring a set of harmful information; sending an application request for accessing the network; training a learning-based Bloom filter with at least one pair of hash functions based on the application request, splitting each pair of hash functions after training, storing the split deep hash function and verification array in a database, and sending the split surface hash function in response to the application request; obtaining the response of the network page, generating and sending a set of surface hash values for harmful content detection; and performing corresponding processing in response to the detection result of the set of surface hash values. By constructing a Bloom filter and setting a pair of hash functions to detect the set of surface hash values of keywords, the present invention does not involve plaintext data, transmits irreversible detection results, and cannot restore plaintext data, thus protecting user privacy; and also improves the efficiency, effectiveness, and accuracy of detection by reducing the length of the verification array.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of harmful content detection of web page privacy data, and particularly relates to a method and system for detecting harmful content of a privacy-protected WEB page. Background Art

[0002] Currently, there are often WEB pages with harmful content on the Internet, and the harmful content includes but is not limited to harmful information such as porn, violence, gambling, and fraud. Some of the harmful content uses encryption protocols to encrypt access requests and backhaul data, making it difficult to conduct effective and accurate detection. In addition, there are two basic requirements for harmful content detection that protects user privacy: one is to ensure the privacy of the content accessed by the user, and the clear text content accessed by the user cannot be cracked by attackers or the detection server; the other is not to disclose the harmful information set of the detection server. Currently, the mainstream encrypted content detection mechanism uses federated learning means, which has a large computational overhead, cannot match the actual application scenarios in reality, and the detection efficiency and accuracy of encrypted content are not high. However, with the improvement of the security of encryption protocols, it is impossible to effectively discover harmful content in WEB pages through traffic detection on the network side, and thus it is impossible to timely remind users that the WEB pages they access have harmful content, posing a relatively large security risk.

[0003] The present invention provides a method, system, device and medium for detecting harmful content of a WEB page that protects user privacy, which is used to solve the problems of low detection efficiency of encrypted content, poor effectiveness and accuracy, and difficulty in effectively guaranteeing the security of user privacy when accessing WEB pages. Summary of the Invention

[0004] The present invention provides a method and system for detecting harmful content in WEB pages to protect user privacy, which is used to solve the problems of low detection efficiency of encrypted content, poor effectiveness and accuracy, and difficulty in effectively protecting the security of user privacy during traditional WEB page access. The present invention converts the data of the WEB page into a keyword set through a keyword extraction algorithm, encrypts and calculates the keyword set through a hash function group in a learning Bloom filter, and transmits the irreversible encrypted calculation result to the server for harmful content detection. The detection process on the server does not involve the plaintext data of the WEB page and cannot restore the plaintext data of the WEB page, achieving the purpose of protecting user privacy; in addition, the hash function group of the Bloom filter is optimized through machine learning algorithms and locality-sensitive hashing algorithms, and in the learning Bloom filter, harmful information is set as a key and non-harmful information is set as a non-key through a hash function heuristic search algorithm. On the one hand, the classification of multiple types of keys and keys, as well as non-keys and non-keys, is realized. By minimizing the conflict degree between keys and keys, as well as non-keys and non-keys in the Bloom filter, the search target of the search algorithm is effectively optimized, the length of the verification array is effectively reduced, and the detection efficiency is improved. At the same time, the cracking difficulty of the plaintext data in the WEB page accessed by the user is increased through the verification array and distance metric, and the time complexity of the detection algorithm is low, improving the effectiveness and accuracy of harmful information detection.

[0005] The purpose of the present invention and the technical problems to be solved are achieved by the following technical solutions.

[0006] The present invention provides a method for detecting harmful content in a WEB page to protect privacy. The detection method on the server includes: obtaining an access application request for accessing an external network web page, and constructing a learning Bloom filter with at least one pair of hash functions based on the access application request; splitting each pair of hash functions after training based on the harmful information set, storing the split deep hash function and the corresponding verification array after training in a database, and responding to the access application request by feeding back the split surface hash function; obtaining a set of surface hash values replied by the external network web page; performing harmful content detection on the set of surface hash values based on the split deep hash function and the corresponding verification array after training, and feeding back the harmful content detection result of the set of surface hash values.

[0007] As another optional implementation manner of the present invention, the performing harmful content detection on the set of surface hash values based on the split deep hash function and the corresponding verification array after training includes: obtaining a corresponding set of deep hash values based on the set of surface hash values and the split deep hash function; determining the data information of the quantity, category, and level of harmful information in the set of surface hash values based on the set of deep hash values and the corresponding verification array after training; generating a harmful content detection result based on the data information of the quantity, category, and level of harmful information in the set of surface hash values.

[0008] As another alternative embodiment of the present invention, for each pair of hash functions trained based on the harmful information set, the split deep hash function and the corresponding verified array after training are stored in a database. Responding to an access request, the feedback of the split surface hash function includes: training the deep hash function based on the surface hash values corresponding to the deep hash function according to the configured harmful information set and the bucket model based on locality-sensitive hashing; mapping the surface hash function values corresponding to the same category of harmful information into fixed positions of the corresponding verified array.

[0009] As another alternative embodiment of the present invention, for each pair of hash functions trained based on the harmful information set, before storing the split deep hash function and the corresponding verified array after training in a database and responding to an access request to feedback the split surface hash function, it includes: classifying the configured harmful information set and setting the level of each category of harmful information according to the classification; generating the corresponding verified array after training based on the category and level of the harmful information set and the training of the deep hash function.

[0010] As another alternative embodiment of the present invention, the mapping of the surface hash function values corresponding to the same category of harmful information into fixed positions of the corresponding verified array includes: iteratively training the surface hash function based on the corresponding verified array after training to establish a distance metric between the surface hash function values; setting the harmful information as keys and the non-harmful information as non-keys; in the domain corresponding to a surface hash function, obtaining the target hash function of the deep hash function based on the heuristic search algorithm with the optimization objective of minimizing the conflicts between keys and between non-keys; mapping different categories of harmful information to positions under the corresponding distance metric of the deep hash function based on the target hash function.

[0011] As another alternative embodiment of the present invention, in the domain corresponding to a surface hash function, obtaining the target hash function of the deep hash function based on the heuristic search algorithm with the optimization objective of minimizing the conflicts between keys and between non-keys includes: generating a set of standard Bloom filters based on the deep hash function of the learning Bloom filter and the keys and the heuristic search algorithm; verifying the existence of non-keys in the standard Bloom filters; if it is verified that the non-keys do not exist and the conflicts between non-keys are maximized, the training ends; otherwise, sequentially adjust the deep hash function of the standard Bloom filters until the adjusted deep hash function is the target hash function.

[0012] The present invention also provides a method for detecting harmful content on a WEB page while protecting privacy. The detection method on the client side includes: generating and sending an access application request; obtaining a split surface hash function trained based on a configured harmful information set; generating and sending a set of surface hash function values to be detected based on the surface hash function according to the response of an external network web page; and responding to the detection result of the set of surface hash function values to be detected.

[0013] As another alternative embodiment of the present invention, the generating and sending a set of surface hash function values to be detected based on the surface hash function according to the response of an external network web page includes: extracting valid text information corresponding to traffic according to the response of the external network web page; extracting the vocabulary to be detected based on a word segmentation algorithm from the valid text information to generate a set of vocabulary to be detected; performing text preprocessing on the set of vocabulary to be detected based on a keyword extraction algorithm; performing classification processing on the vocabulary in the set of vocabulary to be detected after text preprocessing, screening and updating the vocabulary set; and generating a set of surface hash function values based on the screened vocabulary set using the surface hash function.

[0014] The present invention also provides a system for detecting harmful content on a WEB page while protecting privacy. The detection system includes: a client, configured to generate and send an access application request; obtain a split surface hash function trained based on a configured harmful information set; generate a set of surface hash function values to be detected based on the surface hash function according to the response of an external network web page; and respond to the detection result of the set of surface hash function values to be detected. A server, configured to obtain an access application request for accessing an external network web page, construct a learning-based Bloom filter of at least one pair of hash functions based on the access application request; split each pair of hash functions trained based on the harmful information set, store the split deep hash function and the corresponding verification array after training in a database, and feedback the split surface hash function in response to the access application request; obtain a set of surface hash values of the external network web page response; perform harmful content detection on the set of surface hash values based on the split deep hash function and the corresponding verification array after training, and feedback the harmful content detection result of the set of surface hash values.

[0015] As another alternative embodiment of the present invention, the server includes: a first initialization module, configured to obtain an access application request for accessing an external network web page, and construct a learning-based Bloom filter of at least one pair of hash functions based on the access application request; split each pair of hash functions after training based on the harmful information set, store the split deep hash function and the corresponding verification array after training in a database, and feedback the split surface hash function in response to the access application request; a first detection module, configured to obtain a set of surface hash values replied by the external network web page; perform harmful content detection on the set of surface hash values based on the split deep hash function and the corresponding verification array after training, and feedback the harmful content detection result of the set of surface hash values. The client includes: a second initialization module, configured to generate and send an access application request; obtain the split surface hash function after training based on the configured harmful information set; a second detection module, configured to generate and send a set of surface hash function values to be detected based on the surface hash function according to the reply of the external network web page; and respond to the detection result of the set of surface hash function values to be detected.

[0016] Compared with the prior art, the present invention has obvious advantages and beneficial effects. Based on the above technical solutions, the present invention has at least one of the following advantages and effects:

[0017] 1. A method for detecting harmful content of a privacy-protected WEB web page provided by the present invention. The detection method of the server includes: obtaining an access application request for accessing an external network web page, and constructing a learning-based Bloom filter of at least one pair of hash functions based on the access application request; splitting each pair of hash functions after training based on the harmful information set, storing the split deep hash function and the corresponding verification array after training in a database, and feedbacking the split surface hash function in response to the access application request; obtaining a set of surface hash values replied by the external network web page; performing harmful content detection on the set of surface hash values based on the split deep hash function and the corresponding verification array after training, and feedbacking the harmful content detection result of the set of surface hash values. The detection method of the present invention is applied to the server. The surface hash function is encrypted before and after training by the hash function group in the learning-based Bloom filter, and the encrypted surface hash function value is transmitted to the server for harmful content detection. The detection process of the server does not involve the plaintext data of the WEB web page, and the plaintext data of the WEB cannot be restored, achieving the purpose of protecting user privacy, increasing the difficulty of cracking the plaintext data in the WEB web page accessed by the user. By using the verification array for verification, the time complexity of the detection algorithm is low, improving the effectiveness and accuracy of harmful information detection.

[0018] II. A method for detecting harmful content on a WEB page that protects privacy provided by the present invention. The detection method on the client side includes: generating and sending an access application request; obtaining a split surface hash function after training based on a configured harmful information set; generating and sending a set of surface hash function values to be detected based on the response of the external network web page using the surface hash function; and responding to the detection result of the set of surface hash function values to be detected. By generating an access application request for the user to access an external network web page on the client side and sending the generated access application request to the server to trigger the initialization of the server, the present invention obtains the surface hash function fed back by the server after training and splitting according to the configured harmful information set and deploys the surface hash function; converts the response data of the WEB page into a keyword set through a keyword extraction algorithm, performs hash calculation on the keyword set using the deployed surface hash function to generate a set of surface hash function values and encrypts them, and transmits the encrypted irreversible set of surface hash function values to the server for harmful content detection. The detection method on the client side enables the detection process on the server side not to involve the plaintext data of the WEB page and makes it impossible to restore the plaintext data of the WEB page, achieving the purpose of protecting user privacy; in addition, by generating the surface hash function values, it increases the difficulty of cracking the plaintext data in the WEB page accessed by the user and has a low time complexity for detection, improving the effectiveness and accuracy of the harmful information detection method on the client side.

[0019] III. A harmful content detection system for WEB pages that protects privacy provided by the present invention. The detection system includes: A client, which is used to generate and send an access application request; obtain a surface hash function split after being trained based on a configured harmful information set; generate a set of surface hash function values to be detected based on the surface hash function according to the response of an external network web page; and respond to the detection result of the set of surface hash function values to be detected. A server, which is used to obtain an access application request for accessing an external network web page, construct a learning-based Bloom filter of at least one pair of hash functions based on the access application request; split each pair of hash functions after being trained based on the harmful information set, store the split deep hash function and the corresponding verification array after training in a database, respond to the access application request and feedback the split surface hash function; obtain a set of surface hash values in the response of the external network web page; perform harmful content detection on the set of surface hash values based on the split deep hash function and the corresponding verification array after training, and feedback the harmful content detection result of the set of surface hash values.The detection system of the present invention converts the data of WEB web pages into a keyword set based on a keyword extraction algorithm on the client side, and generates and sends a set of surface hash function values to be detected to the server through a surface hash function split after training based on a configured set of harmful information. The server performs harmful content detection on the set of surface hash values based on the split deep hash function and the corresponding verification array after training; trains the deep hash function for the surface hash values corresponding to the deep hash function based on the configured set of harmful information and the bucket model of locality-sensitive hashing; iteratively trains the surface hash function based on the corresponding verification array after training to establish a distance metric between the surface hash function values; sets harmful information as keys and non-harmful information as non-keys; in the domain corresponding to a surface hash function, obtains the target hash function of the deep hash function based on the heuristic search algorithm with the optimization goal of minimizing the conflicts between keys and between non-keys; maps different categories of harmful information to positions under the distance metric corresponding to the deep hash function based on the target hash function to obtain and encrypt the feedback set of surface hash values for harmful content detection results, and by feeding back the irreversible harmful content detection results to the client, it is difficult for the client to crack the detection process verified by the server and the plaintext data of the configured set of harmful information, and it is impossible to completely crack the plaintext data of the configured set of harmful information, achieving the purpose of protecting user privacy; additionally, iteratively trains the hash function group of the Bloom filter through machine learning and optimizes it with the locality-sensitive hashing algorithm, and sets harmful information as keys and non-harmful information as non-keys by using the heuristic search algorithm for the hash function of the learning-based Bloom filter; on the one hand, it realizes the classification of multiple types of keys and keys, non-keys and non-keys, effectively optimizes the search target of the search algorithm by minimizing the conflicts between keys and between non-keys in the Bloom filter, effectively reduces the length of the verification array, increases the difficulty of cracking the plaintext data in the user's access to WEB web pages through the verification array, and the time complexity of this detection method is low, further improving the effectiveness and accuracy of harmful information detection.

[0020] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above structure and other purposes, features and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given and described in detail in conjunction with the drawings as follows. Brief Description of the Drawings

[0021] Figure 1 It is a schematic diagram of the detection process of a method for detecting harmful content in a WEB web page for protecting privacy in this embodiment applied to the server.

[0022] Figure 2 It is a schematic diagram of the detection process of a method for detecting harmful content in a WEB web page for protecting privacy in this embodiment applied to the client.

[0023] Figure 3 This is a schematic structural diagram of a method for detecting harmful content on a WEB page that protects privacy in this embodiment.

[0024] Figure 4 This is a schematic flowchart of the initialization stage of a method for detecting harmful content on a WEB page that protects privacy in this embodiment.

[0025] Figure 5 This is a schematic flowchart of the detection stage of a method for detecting harmful content on a WEB page that protects privacy in this embodiment.

[0026] Figure 6 This is a schematic structural diagram of a system for detecting harmful content on a WEB page that protects privacy in this embodiment.

[0027] Figure 7 This is a schematic structural diagram of a server for detecting harmful content on a WEB page that protects privacy in this embodiment.

[0028] Figure 8 This is a schematic structural diagram of a client for detecting harmful content on a WEB page that protects privacy in this embodiment.

[0029] Figure 9 This is a schematic structural diagram of an electronic device in this embodiment.

[0030] Explanation of the reference numerals in the drawings:

[0031] 200: Server 210: First initialization module

[0032] 220: First detection module 300: Client

[0033] 310: Second initialization module 320: Second detection module

[0034] 400: System for detecting harmful content on a WEB page 500: Electronic device

[0035] 510: Memory 520: Processor

[0036] 530: Computer-readable instructions Detailed implementation manners

[0037] To further elaborate on the technical means and effects adopted by the present invention to achieve the intended invention purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation manners, structures, features, and effects thereof proposed according to the present invention as follows:

[0038] The present invention provides a method for detecting harmful content on a WEB page that protects privacy, as Figures 1 to 9As shown, this detection method runs on the server side and interacts with the corresponding client. The method for detecting harmful content on the WEB page includes the following processes:

[0039] S1: Obtain an access application request for accessing an external network web page, and construct at least one learning-based Bloom filter for a pair of hash functions based on the access application request; split each pair of hash functions after training based on the harmful information set, store the split deep hash function and the corresponding verification array after training in the database, and feedback the split surface hash function in response to the access application request.

[0040] It should be noted that the server side includes at least one detection server, and each detection server may include at least one learning-based Bloom filter generation module and at least one learning-based Bloom filter detection module. In the present invention, after receiving a request from the client to access the external network, the detection server starts the initialization operation process. The detection server constructs a standard learning-based Bloom filter through the learning-based Bloom filter generation module. The learning-based Bloom filter is composed of two parts as a whole. One part is a hash function group composed of at least one pair of hash functions; the other part is the verification array corresponding to the pair of hash functions. As Figure 4As shown, the learning Bloom filter constructed by the detection server includes a hash function group pair of n (n represents an integer not less than 1) hash function pairs, and a corresponding verification array. Among them, each hash function pair includes a surface hash function and a deep hash function, and the surface hash function and the deep hash function of each hash function pair are related to each other. Before initialization, the detection server needs to configure the data of the harmful information set in the database, train the deep hash function group composed of n hash functions based on the harmful information set, and the surface hash function can be another corresponding surface hash function generated by transforming the initial deep hash function through a preset training relationship (such as the deep hash function value of the harmful information set corresponding to the deep hash function group). For example, the deep hash function group composed of n hash functions corresponds to the surface hash function group composed of n hash functions. The learning Bloom filter contains untrained deep hash functions and transformed surface hash functions. For example, the surface hash function can convert a keyword into a 512-bit fixed-length binary number, denoted as the surface hash function value; the deep hash function can convert the 512-bit surface hash function value into a 64-bit binary number, denoted as the deep hash function value. Overall, the combination of the surface hash function and the deep hash function can be regarded as a hash function that can convert an arbitrarily long string into 64-bit binary data, which is consistent with the hash function in the conventional Bloom filter in terms of the function of calculating the hash value. The detection server can configure the harmful information set and train the constructed learning Bloom filter based on the configured harmful information set, and generate a learning Bloom filter that meets the filtering function of the configured harmful information set and the corresponding verification array after training, and store them in the database of the detection server. The learning Bloom filter is composed of the above-mentioned hash function group pair of n hash function pairs. Among them, in response to the access request of the client to access the external network web page, the split surface hash function is fed back to the client for the initialization operation of client detection.

[0041] During the initialization of the detection server, the construction process of the learning Bloom filter includes three parts: constructing a verification array, training the surface hash function group, and training the deep hash function group. First, when constructing the verification array based on the configured harmful information set, the total number of harmful information in the harmful information set in the detection server and the total length of the verification array need to be considered. The total length λ of the verification array is preset based on the total number of harmful information, and the tester inputs the total length λ of the verification array into the parameter for configuring the total length of the verification array in the detection server. Secondly, in the training process of the surface hash function group, a hash function method using a hash function to construct a message authentication code (HMAC) can be adopted, and different surface hash functions are distinguished by their keys. Finally, the training of the deep hash function group is completed. The surface hash function trained by the learning Bloom filter of the present invention has the property of irreversibility, and the detection server cannot reverse the plaintext corresponding to the keyword hash value provided by the client through the surface hash function group. Generally, due to the high collision between non-harmful information of the surface hash function of this learning Bloom filter, on the basis that the detection server does not affect the execution operation process, there is no malicious training of the surface hash function to reduce the collision of non-harmful information of the client. Therefore, the number of plaintexts corresponding to a single non-keyword hash value found by the detection server through the collision method is large and there is no semantic consistency. Therefore, brute force cracking cannot accurately crack the keyword plaintext information of the client. The present invention can ensure that the client keywords of the user are not leaked, thereby protecting the user privacy. The present invention can effectively prevent the client from attacking the deep hash function constructed by the harmful information set by deploying the split deep hash function after training on the detection server.

[0042] In the present invention, before obtaining an access application request for accessing an external network web page, a set of harmful information to be detected can be configured by a staff member in the database of the detection server. The present invention realizes different harmful information detection requirements for the harmful information set based on different learning Bloom filters based on a verification array. For example, the number of harmful information in keywords can be detected by configuring a training learning Bloom filter; or the harmful information in keywords can be detected by configuring a training learning Bloom filter, classified according to different categories of harmful information such as "pornography", "violence", etc., and then the number of harmful information in each category is counted; or the harmful information in keywords can be detected by configuring a training learning Bloom filter, and the harmful information in different categories such as "pornography", "violence", etc. is divided into harmful levels of "weakly harmful", "relatively harmful", "harmful", and "strongly harmful" according to different harmful levels, and finally the number of harmful information in different harmful levels in each category is counted separately. When a rapid request for applying to access the external network is triggered in the detection server, in response to the access application request, a split surface hash function is fed back to complete the initialization operation process of the detection server.

[0043] S2: Obtain a set of surface hash values of the external network web page response; perform harmful content detection on the set of surface hash values based on the split deep hash function and the corresponding verification array after training, and feed back the harmful content detection result of the set of surface hash values.

[0044] It should be noted that the detection server uses the deep hash function group stored in the database to convert each keyword in the set of surface hash values into a deep hash value through the locally sensitive hashing algorithm corresponding to ( ≤n) deep hash functions obtained surface hash values. Among them deep hash functions correspond one-to-one with surface hash values.

[0045] In the present invention, after the set of surface hash values generates the set of deep hash values, each keyword is detected based on the mapping position and within the distance metric range through the stored verification array. By disjuncting each keyword's The Boolean value (such as 0 or 1) of a deep hash value in the verification array position of the index gives the harmful information detection result of the keywords in the surface hash value set. For example, if the Boolean value in the result is 1, it indicates that the word is harmful information. After detecting the harmfulness of the harmful information in the keywords, further check whether the range of the distance metric mapped by the local deep hash value set is located at the fixed position of the verification array. Further, the above fixed position can be preset as a "bucket" at a fixed position, and check whether the range mapped by the local deep hash value set is within the bucket. If a local deep hash value set is mapped within the "bucket" (the fixed position range of the verification array) within the distance metric range, the harmful information set of the classification or classification level to which the harmful information belongs can be determined. Detect all the harmful information in the keywords, and count the detection results to obtain the request currently sent by the client to the WEB web server, and extract the information on the number, category, and harmful level of the harmful information contained in the response data packet in the response sent back by the WEB web server. The detection server makes corresponding responses and processes to the access application request for accessing the external network web page according to the detection result.

[0046] As an optional implementation manner, the harmful content detection of the surface hash value set based on the split deep hash function and the corresponding verification array after training includes: obtaining the corresponding deep hash value set according to the surface hash value set based on the split deep hash function; determining the data information of the quantity, category, and level of the harmful information in the surface hash value set according to the deep hash value set based on the corresponding verification array after training; generating a harmful content detection result based on the data information of the quantity, category, and level of the harmful information in the surface hash value set.

[0047] It should be noted that in the detection server, after detecting whether the keyword is harmful information, it is possible to calculate the data information of the quantity, category, and level of harmful information in the keyword through the deep hash value set. For example, it is possible to distinguish different categories and different harmful levels of the harmful information in the determined keyword, as well as perform statistical operations on its quantity. After completing the detection of harmful information in the surface hash value submitted by the client, calculate, analyze, and count the classification, harmful level, and quantity of the detected harmful information, and record the data information of the category to which the detected harmful information belongs, the harmful level of the harmful information, and the quantity. Further, the detection result of harmful information can be generated through the data information of the quantity, category, and level of harmful information in the keyword represented by the surface hash value set. Specifically, during training, the verification array will be pre-segmented into multiple segments to correspond to different categories, levels, and quantities of harmful information. The present invention can check the values at the fixed positions indexed by the deep hash value in the verification array. If all these values correspond to 1, it is confirmed that the word is harmful information; otherwise, it is confirmed that the word is non-harmful information. For the confirmed harmful information, the category and level corresponding to the corresponding harmful information can also be determined by checking the position of the deep hash value in the interval. This corresponding relationship can refer to the bucket calculation method of locality-sensitive hashing, which will not be elaborated here. The present invention forms a team of hash functions in the learning Bloom filter through training, and combines locality-sensitive hashing (LSH) to adopt the bucket concept to accurately verify harmful information for the deep hash value converted from the surface hash value, making the calculation process of harmful information filtering and verification more efficient. When facing a large number of clients and a large amount of access data for each client, it is very necessary to efficiently filter harmful information through the learning Bloom filter trained through learning, and analyze and process the classification, harmful level, and quantity of harmful information.

[0048] As an optional implementation manner, each pair of hash functions after training based on the harmful information set is split, and the split deep hash function and the corresponding verification array after training are stored in the database. Responding to the access request, the feedback of the split surface hash function includes: training the deep hash function based on the surface hash value corresponding to the deep hash function according to the bucket model of locality-sensitive hashing based on the configured harmful information set; mapping the surface hash function values corresponding to the harmful information of the same category into the fixed positions of the corresponding verification array.

[0049] It should be noted that the constructed deep hash function group during training is not limited to including a group of hash functions of a standard Bloom filter, and it is not limited to constructing ( ≤n) locality-sensitive hash functions based on locality-sensitive hashing (LSH) with the distance metric d(p, q). It is determined by the detection server according to the configured harmful information. Specifically, the hash functions in the hash function group are actually composed of a pair of hash functions consisting of a surface hash function and a deep hash function. The deep hash function composed of each locality-sensitive hash function can be trained based on the surface hash value set, and after training, the corresponding deep hash value set can be obtained based on the input surface hash value set. The surface hash function converts a keyword into a 512-bit fixed-length binary number, called the surface hash function value; the deep hash function converts the 512-bit surface hash function value into a 64-bit binary number, called the deep hash function value. Overall, the combination of the surface hash function and the deep hash function can be regarded as a hash function that can convert an arbitrarily long string into 64-bit binary data, which is consistent with the hash function in the conventional Bloom filter in terms of the function of calculating the hash value. The constructed deep hash function group and the verification array are stored in the database of the detection server, and the constructed surface hash function group is sent back to the client. The client deploys the received surface hash function group to the detection tool to complete the initialization. The algorithm based on the fixed position of the locality-sensitive hash (LSH) in the present invention can make the operation efficiency of the learning Bloom filter much higher than the point indexing method of the tree structure. The learning Bloom filter of the present invention does not need to exactly know the plaintext information of the harmful information. Only by correspondingly mapping the hash value submitted by the client through the deep hash function, the value of the keyword plaintext vocabulary corresponding to the harmful information in the harmful information set in the surface hash value can be mapped to the fixed position of the verification array in the learning Bloom filter.

[0050] As an optional implementation manner, each pair of hash functions after training based on the harmful information set is split, and the split deep hash function and the corresponding verification array after training are stored in the database. Before responding to the access request and feedbacking the split surface hash function, it includes: classifying the configured harmful information set, and setting the level of each type of harmful information according to the classification; generating the corresponding verification array after training based on the category and level of the harmful information set and the training of the deep hash function.

[0051] It should be noted that when classifying the configured harmful information set, it is not limited to classifying harmful information according to different categories such as "pornography" and "violence", and setting the level of harmful information in each classification based on the harmful level. For example, the harmful information in the "violence" category is set to four harmful levels of "weakly harmful", "relatively harmful", "harmful", and "strongly harmful" based on the harmful level. Based on the category and level of the configured harmful information set, a verification array of the corresponding harmful information set is generated through the training of the deep hash function. Since the harmful information set in the detection server is divided into multiple different types and multiple different levels, a total of M harmful information sets of different types and different levels are preset. For each type and level of harmful information, a fixed position (fixed length range) in the verification array is divided as the "bucket" corresponding to each harmful information level in this category. The present invention correspondingly allocates the position (number of bits) of the bucket corresponding to the set in the verification array according to the number of elements mj (j = 1, 2, 3,..., M) in the j-th set of harmful information of different categories. The obtained corresponding surface hash value set can map 512-bit data within the same range to the same bucket at the preset fixed position in the verification array, so as to map the surface hash function values of the harmful information sets belonging to the same category or the same harmful level in different categories to the range of the same verification array in the database, and detect and verify whether it is located in the "bucket" representing a certain harmful information, realizing the verification and detection of harmful information in the surface hash value set composed of keywords through the verification array.

[0052] As an optional implementation manner, mapping the surface hash function values corresponding to the harmful information of the same category into the fixed position of the corresponding verification array includes: iteratively training the surface hash function based on the corresponding verification array after training, and establishing a distance metric between the surface hash function values; setting the harmful information as the key and the non-harmful information as the non-key; in the domain corresponding to a surface hash function, obtaining the target hash function of the deep hash function based on the heuristic search algorithm with the optimization goal of minimizing the conflicts between keys and between non-keys; mapping the harmful information of different categories to the positions under the corresponding distance metric of the deep hash function based on the target hash function.

[0053] It should be noted that before the training starts, it is first necessary to generate a distance metric d(p, q) based on the definition of distance metric (not limited to Euclidean distance), where p and q can be understood as data points composed of different keywords in two two-dimensional spaces mapped to the verification array. This distance metric d can be specified or randomly generated. The detection server needs to securely store this distance metric d. If the distance metric is leaked to the client, it will result in a decrease in the cost of brute-forcing the harmful information set in the detection server by the client. In the present invention, harmful information is set as keys, and non-harmful information is set as non-keys; and in the domain corresponding to the surface hash function, based on the heuristic search algorithm with the optimization objective of minimizing the conflicts between keys and keys, and between non-keys and non-keys, the target hash function of the deep hash function is obtained, and it is used as the deep hash function that is the most preferred target corresponding to the shallow hash function based on the conversion relationship.

[0054] In the present invention, for example, after generating the distance metric d, the optimal hash function group is searched through a combinatorial optimization algorithm. For each surface hash function value in the hash function group, the branch and bound method is used to find the optimal hash function in the entire set of surface hash functions within a limited time cost. Here, the preset surface hash function group is h_ i (x), (i = 1, 2, 3,..., n), where the objective function Maximize Z corresponding to the optimization of the search for the optimal hash function group by the combinatorial optimization algorithm of the i-th surface hash function is as follows in Equation (1): (1)

[0056] Where, , i and j represent positive integer sequences configured based on the harmful information set, M represents the total number of different types of harmful information in the configured harmful information set, represents the number of different types of harmful information sets, k represents the search depth of the configured harmful information, l represents the search depth of harmful information in the keywords, m j represents the number of elements in the j-th set of different types of harmful information sets, represents the harmful information set currently mapped into the bucket, represents the keyword set currently to be mapped into the bucket. The constraint condition of the combinatorial optimization algorithm for searching the optimal hash function group is an integer constraint, which runs in the entire set of hash functions for constructing the hash-based message authentication code (HMAC) in the feasible domain. After learning and iteration, the generated hash function set after training is the surface hash function group, so that the obtained surface hash function group can distinguish different types of keywords in terms of the distance metric d(p, q) for different keywords (data points) mapped to two two-dimensional spaces of the verification array.

[0057] In the present invention, when obtaining the corresponding deep hash value set according to the deep hash function of the trained learning Bloom filter based on the surface hash value set, there may be a situation where non-harmful information in the surface hash value set is mistakenly delivered to the position (bit number) of the bucket used for verifying the harmful information set, resulting in an example of "false positive" in the locality-sensitive hashing algorithm. A large number of false positives will pose a risk of misjudgment to the client, and the higher the level of harmful information, the more unacceptable the cost of misjudgment. In the present invention, the length of the verification array can be increased, and the distance between the buckets corresponding to each set in the verification array can be increased (the range corresponding to the fixed position), so that the probability of non-harmful information in the keyword being misjudged as harmful information and resulting in false positives is extremely low. As an optional implementation manner, for the harmful information set with a high harmful level, the corresponding fixed position (the bit number range corresponding to the bucket) in the verification array is appropriately increased according to the distance, so that the probability of non-harmful information in the keyword being misjudged as harmful information and resulting in false positives becomes extremely low, even negligible.

[0058] As an optional implementation manner, in the domain corresponding to a surface hash function, the target hash function of the deep hash function obtained based on the heuristic search algorithm with the optimization objective of minimizing the conflicts between keys and between non-keys includes: generating a set of standard Bloom filters based on the deep hash function of the learning Bloom filter and the heuristic search algorithm for keys; verifying the existence of non-keys in the standard Bloom filter; if it is verified that the non-keys do not exist and the conflicts between non-keys are maximized, the training ends; otherwise, the deep hash function of the standard Bloom filter is adjusted sequentially until the adjusted deep hash function is the target hash function.

[0059] It should be noted that the split deep hash function is used to obtain the corresponding deep hash value set according to the surface hash value set, and the verification array corresponding to the trained one is used to verify whether there is corresponding harmful content in the detected surface hash value set. For example, without limitation, the deep hash function of the standard Bloom filter can be adjusted sequentially until the adjusted deep hash function is the target hash function, until all the harmful information represented by the keys and the non-harmful keywords represented by the non-keys can maintain the minimum conflict degree. At this time, the output result of the deep hash function after adjustment and training is stored as the surface hash function. That is, the target hash function output after training is the deep hash function.

[0060] As an optional implementation manner, the learning Bloom filter iteratively trains the surface hash function by learning: during the iterative training process of the surface hash function, the split surface hash function is encrypted based on the hash message authentication code, and different surface hash functions are distinguished based on the encrypted key.

[0061] It should be noted that the constraint condition of the optimal hash function group combination optimization algorithm searched is an integer constraint, and it runs on the entire set of hash functions for constructing the hash-based message authentication code (HMAC) in the feasible domain. As an alternative implementation, the distance metric can be used as the key for encrypting the surface hash function with the hash message authentication code. Different surface hash functions can be distinguished based on the encrypted key. After learning and iterative training, the generated set of hash functions after training is the surface hash function group, and it can be encrypted based on the hash message authentication code. On the one hand, it ensures the confidentiality and security of the split surface hash functions during transmission in the encrypted state; on the other hand, it can better distinguish different keywords (data points) in the two two-dimensional spaces of the verification array for different categories of keywords based on the distance metric d(p, q). As a secure and confidential encryption, the transformation relationship between the hash function pairs composed of the deep hash function and the surface hash function can also be set as the key, and the trained surface hash function can be encrypted with this key and transmitted to the client for decryption and use. For example, the above transformation relationship is not limited to various confusion transformations, and the trained surface hash function is encrypted based on this transformation relationship to ensure the confidentiality of the split surface hash functions.

[0062] The present invention also provides a method for detecting harmful content on WEB pages to protect privacy. This detection method runs on the client and interacts with the corresponding server, such as Figure 2 shown, the method for detecting harmful content on WEB pages includes the following processes:

[0063] St1: Generate and send an access application request; obtain the surface hash functions split after training based on the configured harmful information set.

[0064] It should be noted that the server can include at least one detected client, and this detected client can be the smallest detection unit for mobile detection. When the detected client generates and sends an access application request, this access application request is not limited to including an access application request for accessing the external network generated on the smallest detection unit for mobile detection and sent to the server (such as Figure 3 ). The client of the present invention triggers the initialization operation of the detection server of the server by generating and sending an access application request, obtains the surface hash functions split after training based on the above-mentioned configured harmful information set, and completes the initialization operation process of the client detection method after deploying it behind the detection tool. The present invention directly deploys the surface hash functions split after training based on the configured harmful information set, improving the efficiency of client detection deployment, and the deployed surface hash functions can quickly and efficiently convert the extracted keyword set into a set of surface hash function values, enhancing the speed and efficiency of client detection.

[0065] St2: Generate and send a set of surface hash function values to be detected based on the response of an external network web page; in response to the detection result of the set of surface hash function values to be detected.

[0066] It should be noted that installing a browser plugin on each minimized client can extract the access traffic of the WEB pages accessed by the client, capture the traffic of the responses of the WEB pages, extract the set of keywords in the response traffic, and generate a set of surface hash function values to be detected corresponding to the set of keywords based on the surface hash function deployed in the initialization stage on the client. As an alternative implementation, the transformation relationship between the surface hash function and the deep hash function trained and split based on the configured set of harmful information is used as a key to encrypt the set of surface hash function values to be detected and send them to the detection server of the server to ensure the confidentiality of the set of surface hash function values to be detected during transmission. After the detection server generates a detection result, obtain the detection result of the set of surface hash function values to be detected in the response of the server, and the client performs a preset access operation in response to the detection result of the set of surface hash function values to be detected, such as allowing access or prohibiting access, to complete the detection of harmful content in the WEB pages accessed by the client. For example, the client detection method can access a WEB page server. The client sends a network web page access request to the external network web page server, and the request and response data capture module in the client detection tool captures and records the request sent by the client to the WEB page server and the traffic of the response returned by the WEB page server. Then, the keyword extraction module analyzes the captured traffic data packets, uses a word segmentation algorithm to extract the valid words involved in the traffic data packets, and filters and optimizes the set of extracted valid words through a keyword extraction algorithm to improve the detection efficiency. Store the set list formed by the above screening and optimization of keywords as a dictionary, which is called the surface hash value set of keywords, and send this surface hash value set to the detection server. After the harmful information detection and verification is performed based on the detection server, feedback the detection result of the set of surface hash function values to be detected to the client, and the client performs corresponding processing in response to the detection result of the set of surface hash function values to be detected.

[0067] In the present invention, the surface hash function trained by the learning Bloom filter maps the set of surface hash function values corresponding to harmful information of the same category to an approximate position defined by the distance metric d of the locality-sensitive hash function of the deep hash function. Since the client cannot know the parameter settings of the distance metric d in the locality-sensitive hash values used by the detection server, if the client tries to brute-force the harmful information database in the detection server through the surface hash function, it first needs to crack the measurement method selected in its deep hash function. When it is impossible to verify whether the result is in the corresponding "bucket" position, the client cannot brute-force the specific parameter value of the above distance metric d. When the accurate distance metric value cannot be determined, the harmful information set cannot be directly brute-forced. Due to the more accurate distance metric d and result verification based on a fixed position, the mutually related distance metric d and fixed position increase the complexity of cracking, causing the computational amount of cracking to increase exponentially. The client cannot crack the harmful information set through a brute-force traversal method, which can better protect the security of the harmful information set.

[0068] As an alternative implementation, generating and sending a set of surface hash function values to be detected based on the response of an external network web page according to the surface hash function includes: extracting valid text information corresponding to the traffic according to the response of the external network web page; extracting the vocabulary to be detected based on the word segmentation algorithm according to the valid text information to generate a set of vocabulary to be detected; performing text preprocessing on the set of vocabulary to be detected based on the keyword extraction algorithm; performing classification processing on the vocabulary in the set of vocabulary to be detected after text preprocessing, screening and updating the vocabulary set; generating a set of surface hash function values based on the screened vocabulary set according to the surface hash function.

[0069] It should be noted that, for example, the screening of valid words can be carried out based on the number of valid words, the categories of valid words, and the levels of valid words. As another alternative implementation, for a single document, the keywords in the single document can be regarded as a category of keywords, and statistical analysis can be carried out on the categories of harmful information in the keywords, the levels of harmful information in each category, the number, etc., and the detection results can be fed back. For multiple documents such as a certain news, its title words can be used as hot words to discover the focus of public opinion. Collect the corpus that needs to extract keywords, which is related to the hot words of the news public opinion focus, and perform text preprocessing such as text cleaning, spelling and error correction, unification of Chinese simplified and traditional characters, and unification of character encoding standards. Divide the preprocessed text into words and extract the valid text. Based on the valid text information, extract the words to be detected based on the word segmentation algorithm to generate a set of words to be detected. Commonly used word segmentation tools can choose "jieba" word segmentation, Chinese Natural Language Processing Library (SnowNLP), Chinese Lexical Analysis Toolkit (THULAC), and finally remove the repeated and stop valid words to complete the screening and optimization of the keywords to form a set list. In the present invention, the keyword set composed of the word set is not limited to being formed by a list using variable-length character encoding (utf-8).

[0070] The present invention also provides a WEB page harmful content detection system for protecting privacy, such as Figure 3 、 Figure 4 and Figure 6 shown, the WEB page harmful content detection system 400 includes at least one client 300 and at least one server 200. At least one client 300 is used to generate and send an access application request; obtain the surface hash function split after training based on the configured harmful information set; generate a set of surface hash function values to be detected based on the surface hash function according to the response of the external network web page; respond to the detection result of the set of surface hash function values to be detected. At least one server 200 is used to obtain the access application request for accessing the external network web page, construct a learning-based Bloom filter of at least one pair of hash functions based on the access application request; split each pair of hash functions after training based on the harmful information set, store the split deep hash function and the corresponding verification array after training in the database, respond to the access application request and feedback the split surface hash function; obtain the set of surface hash values of the external network web page response; perform harmful content detection on the set of surface hash values based on the split deep hash function and the corresponding verification array after training, and feedback the harmful content detection result of the set of surface hash values.

[0071] It should be noted that the server 200 includes at least one detection server. When the client 300 and the server 200 of the WEB page harmful content detection system 400 of the present invention are running, the data of the WEB page is converted into a keyword set through a keyword extraction algorithm. The hash function group in the learning Bloom filter can be used to perform encryption calculation on the keyword set, and the irreversible encryption calculation result can be transmitted to the detection server for harmful content detection. The detection process of the detection server does not involve the plaintext data of the WEB page and cannot restore the plaintext data of the WEB page, achieving the purpose of protecting user privacy; in addition, the hash function group of the Bloom filter is optimized through machine learning and locality-sensitive hashing algorithms, and in the heuristic search algorithm of the hash function in the learning Bloom filter, harmful information is set as the key and non-harmful information is set as the non-key. On the one hand, the classification of multiple types of keys and keys, non-keys and non-keys is realized. By increasing the conflict degree between keys and keys, and non-keys and non-keys in the Bloom filter, the search target of the search algorithm is effectively optimized, the length of the verification array is effectively reduced, the cracking difficulty of the plaintext data in the WEB page accessed by the user is increased, and the time complexity of the detection algorithm is low, improving the effectiveness and accuracy of harmful information detection.

[0072] As an alternative embodiment, as Figure 7 shown, the server 200 includes: a first initialization module 210, which is used for the initialization before server detection, obtains an access application request for accessing an external network web page, constructs a learning Bloom filter of at least one pair of hash functions based on the access application request; splits each pair of hash functions after training based on the harmful information set, stores the split deep hash function and the corresponding verification array after training in the database, and feeds back the split surface hash function in response to the access application request; a first detection module 220, which is used to obtain the set of surface hash values replied by the external network web page during server detection; performs harmful content detection on the set of surface hash values based on the split deep hash function and the corresponding verification array after training, and feeds back the harmful content detection result of the set of surface hash values. As Figure 8 shown, the client 300 includes: a second initialization module 310, which is used for the initialization before client detection, generates an access application request and sends it; obtains the split surface hash function after training based on the configured harmful information set; a second detection module 320, which is used to generate and send a set of surface hash function values to be detected based on the surface hash function according to the reply of the external network web page during client detection; responds to the detection result of the set of surface hash function values to be detected.

[0073] The present invention also provides an electronic device, which is applied to the server and the client, as Figure 9As shown, the electronic device 500 includes: a memory 510 for storing non-transitory computer-readable instructions 530; and a processor 520 for running the computer-readable instructions 530 such that when the computer-readable instructions 530 are executed by the processor 520, the above-described WEB page harmful content detection method is implemented.

[0074] The above technical solutions of the present invention at least include the following beneficial effects:

[0075] (1) The present invention realizes accurate and efficient harmful content detection through a learning Bloom filter. Compared with traditional methods such as homomorphic encryption, the present invention constructs a harmful information detection index with the same computational time cost and ensures the accuracy of the harmful information category, level, and quantity, and the time efficiency of each detection is increased by more than 1000 times, and a millisecond-level harmful information detection delay can be achieved, greatly improving the harmful information detection efficiency of keywords. Among them, the detection tool realizes the parsing of the plain text of the browser web page content by the client, and combines the word segmentation algorithm, keyword analysis algorithm, and surface hash function algorithm to realize the function requirements of keyword extraction, encryption, and secure transmission by the client. The learning Bloom filter constructed and trained by the present invention trains and optimizes its corresponding hash function through a machine learning algorithm, constructs and trains a deep hash function through locality-sensitive hashing, trains and constructs the surface function through a combinatorial optimization algorithm, and constructs a hash function group pair in the learning Bloom filter through a two-layer hash function group composed of a deep hash function and a shallow hash function.

[0076] (2) In the present invention, harmful information is set as the key in the improved learning Bloom filter, non-harmful information is set as non-key, and different types of harmful information are set as different types of words (such as harmful information and non-harmful information). By classifying harmful information (keys), the conflict degree among harmful information with the same type of keywords is high, and the conflict degree among harmful information with different types of keywords is low. It is precisely based on the characteristic of high conflict degree among harmful information that it can be applied to the detection and analysis of harmful information, enabling the Bloom filter to perform detection and verification with a 100% probability, thereby identifying harmful information in keywords. For example, in the detection and verification stage, the deep hash function realizes the verification of the position and structure values under the distance metric, improving the conflict degree between keys. At the same time, the non-harmful information (non-keys) of keywords in the deep hash function is mapped to the complement of the mapping position of all harmful information sets under the distance metric with a high probability, so as to enhance the high conflict degree among the non-harmful information (non-keys) of keywords. Thus, the hash function simultaneously achieves a high conflict degree among the non-harmful information (non-keys) of keywords and a high conflict degree among the harmful information (keys) of keywords within a relatively optimal time and space complexity. In the present invention, under the condition of high conflict among the non-harmful information of keywords and high conflict among the harmful information of keywords, the detection server cannot decipher the keyword set submitted by the client through brute force cracking, further protecting the privacy security of the client.

[0077] (3) In the present invention, a long verification array is used, so that the verification array of the improved learning Bloom filter can be more accurate in the verification of position and structure values, avoiding misjudging the non-harmful information in the keyword set as harmful information in the harmful information set after detection and verification. In addition to achieving the above precise detection (ensuring that the harmful information in the keyword set is determined to be harmful information in the harmful information set after detection and verification), it can distinguish different categories of harmful information with a high probability, greatly reducing the occurrence of the "false positive" situation where the probability of classification error in detection and verification is less than 1%, and making the non-harmful information in the keyword set be misidentified as harmful information in the harmful information set with an acceptable probability. In addition, in the present invention, the computational time cost of constructing and training the improved learning Bloom filter is highly acceptable, and there is no need to consider the problem of verification array transmission. Therefore, a long verification array can be used, and the distinguishing range of different categories of harmful information can be larger, improving the accuracy of detection and verification.

[0078] (4) The learning Bloom filter hash function group constructed by the present invention through two-layer hash functions protects the privacy of the client and the privacy of the detection server for the harmful information set. The communication data between the client and the detection server are all encrypted data and cannot be utilized by attackers. In the initialization stage, the data transmission of the surface layer hash function group can adopt the form of the key array corresponding to HMAC, so that attackers cannot analyze any transmission information between the detection server and the client from this key array; in the detection stage, the message sent by the client to the detection server is only the dictionary of the surface layer hash values. Without involving the distance metric of the detection server, the deep layer hash function group, the verification array, and their correlation relationships, attackers cannot obtain any useful information about the harmful information set from the dictionary of the surface layer hash values. Even if attackers obtain the detection server with the above distance metric, deep layer hash function group, verification array, and their correlation relationships, due to the irreversibility of the generated surface layer hash function, the security of the surface layer hash function is ensured and the keyword data information of the client cannot be cracked. Thus, the WEB network harmful content detection method and system of the present invention can be used to test whether the elements of the keyword set are members of the harmful information set, realizing secure encrypted harmful information detection, efficiently completing the detection task of the configured harmful information set, saving the time and space of filtering calculation, and having high accuracy.

[0079] The above are only the preferred embodiments of the present application and do not impose any form of limitation on the present application. Although the present application has been disclosed above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to the equivalent embodiments by using the disclosed technical content within the scope of the technical solution of the present application. However, as long as it does not depart from the content of the technical solution of the present application, any simple modification, equivalent change, and modification made to the above embodiments according to the technical essence of the present application still fall within the scope of the technical solution of the present application.

Claims

1. A privacy-protecting WEB page harmful content detection method, characterized in that: The server-side detection methods include: Obtain an access application request for accessing an external network web page, construct a learning Bloom filter of at least one hash function pair based on the access application request, and train the hash function pair based on the harmful information set; split each hash function pair trained based on the harmful information set, store the split deep hash function and the corresponding verification array after training in a database, and feedback the split surface hash function in response to the access application request; Obtain a surface hash value set generated by the client via an external network web page response; perform harmful content detection on the surface hash value set based on the split deep hash function and the corresponding verification array after training, and obtain the corresponding deep hash value set based on the surface hash value set based on the split deep hash function; determine the data information on the quantity, category, and level of harmful information in the surface hash value set based on the deep hash value set based on the corresponding verification array after training; generate a harmful content detection result based on the data information on the quantity, category, and level of harmful information in the surface hash value set, and feed back the harmful content detection result of the surface hash value set.

2. The method for detecting harmful content of a web page according to claim 1, characterized in that: The splitting is based on each hash function pair trained on the harmful information set, storing the split deep hash function and the corresponding verification array after training in a database, and responding to the access application request to feedback the split surface hash function includes: The deep hash function is trained based on the bucketing model of local sensitive hashing according to the configured harmful information set for the surface hash value corresponding to the deep hash function; Map the surface hash function value corresponding to the harmful information of the same category into the fixed position of the corresponding verification array.

3. The method for detecting harmful content of a web page according to claim 2, characterized in that: The splitting is based on each hash function pair trained based on the harmful information set, storing the split deep hash function and the corresponding verification array after training in a database, and responding to the access application request to feedback the split surface hash function before comprising: Classify the configured harmful information set, and set the level of each type of harmful information according to the classification; According to the category and level of the harmful information set, a corresponding verification array is generated based on the training of the deep hash function.

4. The method for detecting harmful content of a web page according to claim 2, characterized in that: The mapping of the surface hash function value corresponding to the harmful information of the same category into the fixed position of the corresponding verification array includes: Iteratively train the surface hash function based on the corresponding validation array after training, and establish a distance metric between the surface hash function values; Set harmful information as key and non-harmful information as non-key; In a domain corresponding to a surface hash function, a target hash function of a deep hash function is obtained based on a heuristic search algorithm based on minimizing conflicts between keys and between non-keys as an optimization target condition; Based on the target hash function, different categories of harmful information are mapped to the positions under the corresponding distance metric of the deep hash function.

5. The method for detecting harmful content of a web page according to claim 4, characterized in that: The target hash function of the deep hash function is obtained in a domain corresponding to a surface hash function according to a heuristic search algorithm based on minimizing conflicts between keys and between non-keys as an optimization target condition, including: Generate a set of standard Bloom filters based on the key and heuristic search algorithm according to the deep hash function of the learned Bloom filter; Verify the existence of non-keys in the standard Bloom filter; if it is verified that the non-key does not exist and the non-key-non-key conflict is maximized, the training ends; Otherwise, the deep hash function of the standard Bloom filter is adjusted in sequence until the adjusted deep hash function becomes the target hash function.

6. A privacy-protecting WEB page harmful content detection method, characterized in that: Client detection methods include: Generate and send access application requests; Get the surface hash function after training and splitting based on the configured harmful information set; Generate and send a set of surface hash function values ​​to be detected based on the surface hash function according to the response of the external network webpage; Obtain the detection result of the server in response to the set of surface hash function values ​​to be detected sent by the client.

7. The method for detecting harmful content of a web page according to claim 6, characterized in that: The step of generating and sending a surface hash function value set to be detected based on a surface hash function according to a response of an external network webpage includes: Extract valid text information corresponding to the traffic according to the response of the external network web page; extract the words to be tested based on the word segmentation algorithm according to the valid text information to generate a word set to be tested; Perform text preprocessing on the vocabulary set to be detected based on the keyword extraction algorithm; Classify the vocabulary set to be detected after text preprocessing, and update the vocabulary set after screening; The filtered vocabulary set is used to generate a surface hash function value set based on the surface hash function.

8. A privacy-protecting web page harmful content detection system, characterized in that: include: The client is used to generate and send access application requests; Get the surface hash function after training and splitting based on the configured harmful information set; Generate a set of surface hash function values ​​to be detected based on the surface hash function according to the response of the external network webpage; Obtaining the detection result of the server in response to the set of surface hash function values ​​to be detected sent by the client; The server is used to obtain an access application request for accessing an external network web page, construct a learning Bloom filter of at least one hash function pair based on the access application request, and train the hash function pair based on the harmful information set; split each hash function pair trained based on the harmful information set, store the split deep hash function and the corresponding verification array after training in a database, and feedback the split surface hash function in response to the access application request; Obtain the surface hash value set generated by the client through the external network web page response; Perform harmful content detection on the surface hash value set based on the split deep hash function and the corresponding verification array after training, and obtain the corresponding deep hash value set based on the surface hash value set based on the split deep hash function; determine the data information of the quantity, category and level of harmful information in the surface hash value set based on the deep hash value set and the corresponding verification array after training; Generate harmful content detection results based on the data information of the quantity, category and level of harmful information in the surface hash value set, and feed back the harmful content detection results of the surface hash value set.

9. The web page harmful content detection system according to claim 8, characterized in that: The server side includes: The first initialization module is used to obtain an access application request for accessing an external network web page, and construct a learning Bloom filter of at least one hash function pair based on the access application request; split each hash function pair trained based on the harmful information set, store the split deep hash function and the corresponding verification array after training in a database, and feedback the split surface hash function in response to the access application request; The first detection module is used to obtain a surface hash value set of an external network web page response; perform harmful content detection on the surface hash value set based on the split deep hash function and the corresponding verification array after training, and feed back the harmful content detection result of the surface hash value set; The client includes: The second initialization module is used to generate and send an access application request; Get the surface hash function after training and splitting based on the configured harmful information set; A second detection module, configured to generate and send a surface hash function value set to be detected based on a surface hash function according to a response of an external network webpage; In response to the detection result of the surface hash function value set to be detected.

Citation Information

Patent Citations

  • Pathogenic gene detection method based on privacy protection intersection calculation protocol

    CN111125736A

  • Content security identification method and device, storage medium and electronic equipment

    CN112600834A