A method and device for detecting pornographic websites based on keyword weights

By establishing a pornographic keyword lexicon and using logistic regression model and KNN algorithm combined with TextRank algorithm, the website text features are extracted for pornographic judgment, which solves the problem of how to effectively and automatically identify pornographic websites, achieves fast and accurate recognition efficiency, and effectively trains the algorithm on a small sample set.

CN113961855BActive Publication Date: 2025-05-27CHINA INTERNET NETWORK INFORMATION CENTER
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111098486.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-18
Publication Date
2025-05-27
Estimated Expiration
2041-09-18

AI Technical Summary

Technical Problem

How to effectively and automatically identify pornographic websites, reduce the negative impact on the healthy growth of adolescents' physical and mental health, and implement algorithm training on a small sample set to reduce the demand for computing resources.

Method used

By establishing a pornographic keyword lexicon, using logistic regression model and KNN algorithm, combined with TextRank algorithm, the website text features are extracted and pornographic judgments are made.

Benefits of technology

It realizes fast and accurate pornographic website recognition, reduces the workload of manual recognition, improves recognition efficiency, and can effectively train algorithms on small sample sets to reduce computing resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113961855B_ABST
    Figure CN113961855B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and device for detecting pornographic websites based on keyword weights. The method uses word frequency statistics to generate recognition weight values ​​of pornographic keywords that appear on both pornographic websites and non-pornographic websites, i.e., first recognition weights; then, for pornographic keywords that do not appear on non-pornographic websites, KNN is used to map TextRank weights to pornographic recognition weights, and then the pornographic recognition weights, i.e., second recognition weights, are calculated; then, text features of the website to be identified are extracted, including text length, hit keyword list length, and hit keyword weight mean, and a logistic regression model is used to identify whether the website is pornographic. The present invention can use web page text information to automatically identify pornographic content on the website. After the text content of the website is obtained through a crawler program, the method can quickly and effectively identify whether the website is pornographic, reduce the workload of manual identification, and improve identification efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of information technology, and particularly relates to a method and device for detecting pornographic websites based on keyword weights. Background Art

[0002] With the progress of technology and the development of society, the Internet has become an indispensable part of people's daily lives. People obtain various information, communicate with others, and purchase goods through the Internet. The Internet greatly facilitates people's lives, speeds up the flow of information, and reduces the cost of information flow. However, there are also some illegal and harmful information on the Internet, such as pornographic websites, which have brought serious negative impacts on the physical and mental health growth of people, especially teenagers. How to effectively and automatically identify pornographic websites and isolate the adverse effects of pornographic websites on teenagers has become a very realistic social problem. Summary of the Invention

[0003] The present invention proposes a method for automatically determining whether a website is pornographic by using web page text information. After obtaining the text content of a website through a crawler program, this method can quickly and effectively identify whether the website is pornographic, reduce the workload of manual identification, and improve the identification efficiency. At the same time, this method can also be used for the detection of gambling websites.

[0004] The technical solution adopted by the present invention is as follows:

[0005] A method for detecting pornographic websites based on keyword weights, comprising the following steps:

[0006] Establish a pornographic keyword library, which includes a list of pornographic keywords and a list of pornographic keyword weights;

[0007] Using the pornographic keyword library, extract the text features of the website to be judged, and use a logistic regression model to judge whether the website to be judged is pornographic.

[0008] Further, the establishment of the pornographic keyword library includes:

[0009] Download the text information of various websites from the Internet and label them as pornographic or non-pornographic, so as to generate a set of pornographic websites and a set of non-pornographic websites;

[0010] Manually extract and generate a list of pornographic keywords according to the text information in the pornographic websites;

[0011] Use word frequency statistics to generate recognition weights for pornographic keywords that appear in both pornographic websites and non-pornographic websites, which is called the first recognition weight;

[0012] For pornographic keywords that only appear on pornographic websites and not on non-pornographic websites, the mapping from TextRank weights to pornographic recognition weights is achieved through the KNN algorithm to obtain the second recognition weight;

[0013] The first recognition weight and the second recognition weight constitute a list of pornographic keyword weights.

[0014] Furthermore, assume that the word frequency statistical value of a certain pornographic keyword in the set of pornographic websites is A, and the word frequency statistical value in the set of non-pornographic websites is B. Then the first recognition weight of this pornographic keyword is A / (A + B).

[0015] Furthermore, the following steps are used to calculate the second recognition weight:

[0016] Use the TextRank algorithm to process the text information of pornographic websites to generate a list of TextRank weights of pornographic keywords;

[0017] Use the generated list of TextRank weights of pornographic keywords generated from the list of pornographic keywords with the first recognition weight to train the KNN algorithm. The first recognition weight is used as the output Y of the KNN algorithm, and the TextRank weight of the pornographic keyword is used as the input X of the KNN algorithm;

[0018] Use the generated KNN algorithm to generate the pornographic recognition weight, that is, the second recognition weight, for the list of pornographic keywords that do not appear on non-pornographic websites using their TextRank weight values.

[0019] Furthermore, the text feature values of the website to be discriminated include: text length, the length of the list of hit keywords, and the average weight of hit keywords.

[0020] Furthermore, the length of the list of hit keywords refers to the number of pornographic keyword entries included in the word segmentation of the website to be discriminated. Let the set of word segmentation entries of the website to be discriminated be N, and the set of pornographic keyword entries be M. Then the length of the list of hit keywords is the length of the set N∩M.

[0021] Furthermore, the average weight of hit keywords refers to the average weight of the keyword set in the list of hit keywords, and the calculation method is: ∑weights[N∩M] / length(N∩M), where ∑weights[N∩M] represents the sum of the weights of each keyword in the list of hit keywords, and length(N∩M) represents the length of the list of hit keywords.

[0022] A pornographic website detection device based on keyword weights using the above method, which includes:

[0023] The pornographic keyword library generation module is used to establish a pornographic keyword library, which includes a list of pornographic keywords and a list of pornographic keyword weights;

[0024] The pornographic website discrimination module is used to utilize the pornographic keyword library to extract the text features of the website to be discriminated, and use a logistic regression model to discriminate whether the website to be discriminated is pornographic.

[0025] The advantages and beneficial effects of the present invention are as follows:

[0026] 1) It is not necessary to collect a large number of samples, and algorithm training can be realized on a small sample set, reducing the difficulty of algorithm implementation.

[0027] 2) The calculation speed is fast, and it does not require powerful computing resources and can be deployed on a general server.

[0028] 3) This method can avoid omissions in the process of human recognition and further improve the accuracy of identifying pornographic websites. Description of the Drawings

[0029] Figure 1 It is a schematic diagram of the deployment method of the pornographic website identification method based on keyword weights of the present invention.

[0030] Figure 2 It is a flowchart for generating the recognition weight of pornographic keywords in the method of the present invention. Detailed Embodiments

[0031] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below through specific embodiments and drawings.

[0032] The main objects of the present invention are: 1) Obtain the text content of the website through a crawler program. 2) Identify whether the website text content is pornographic through the method of the present invention.

[0033] The present invention proposes a pornographic website identification method based on keyword weights, and the deployment method is as Figure 1 shown.

[0034] I) Generation of the pornographic keyword library:

[0035] The pornographic website identification server downloads the text information of various websites from the network through a crawler program, and labels the text of each website: pornographic, non-pornographic. Thus, a pornographic website set and a non-pornographic website set are generated.

[0036] According to the text information in the pornographic websites, a list of pornographic keywords is manually extracted and generated.

[0037] Perform word segmentation on the website text information obtained in step 1).

[0038] Using the word segmentation results, perform word frequency statistics on the pornographic keywords contained in the obtained collection of pornographic websites and the collection of non-pornographic websites. Suppose a certain keyword W1, the word frequency statistic in the collection of pornographic websites is A, and the word frequency statistic in the collection of non-pornographic websites is B.

[0039] According to the word frequency statistics generated in step 4), generate an identification weight for the pornographic keywords that appear in both pornographic websites and non-pornographic websites, which is called the first identification weight. This first identification weight is A / (A + B), where A is the word frequency statistic value of the pornographic keyword in the pornographic website, and B is the word frequency statistic value of the pornographic keyword in the non-pornographic website.

[0040] For the pornographic keywords that only appear in pornographic websites and do not appear in non-pornographic websites, calculate their identification weights according to steps 7)-9), which is called the second identification weight.

[0041] Use the TextRank algorithm to process the text information of pornographic websites and generate a TextRank weight list of pornographic keywords.

[0042] Use the list of pornographic keywords with the first identification weight generated in step 5) and the TextRank weight list of pornographic keywords generated in step 7) to train the KNN (K-Nearest Neighbor) algorithm. The first identification weight generated in step 5) is used as the output Y of the KNN algorithm, and the TextRank weight of the pornographic keywords generated in step 7) is used as the input X of the KNN algorithm.

[0043] Use the KNN algorithm model generated in 8) to generate its pornographic identification weight, that is, the second identification weight, for the list of pornographic keywords that do not appear in non-pornographic websites, using their TextRank weight values. The first identification weight and the second identification weight constitute the weight list of pornographic keywords, and thus the weight list of the entire list of pornographic keywords is obtained. The entire list of pornographic keywords and the weight list of pornographic keywords are used as the pornographic keyword library.

[0044] II) Discrimination of pornographic websites:

[0045] 1) The pornographic website identification server downloads the text content of the website to be judged from the network through a crawler program.

[0046] 2) Preprocess the text content of the website to be judged and obtain the length of the preprocessed text.

[0047] 3) Perform word segmentation on the obtained text content of the website to be judged.

[0048] 4) After obtaining the word segmentation, obtain the list of keywords included in the website, and use the pornographic keyword library containing the weight list generated in the previous steps to calculate the length of the hit keyword list and the average weight of the hit keywords;

[0049] Among them, the length of the hit keyword list refers to the number of pornographic keyword entries included in the word segmentation of the obtained website. The calculation method is: Let the set of entries after the website word segmentation be N, and the set of pornographic keyword entries be M. Then the length of the hit keyword list is the length of the set N∩M, which is represented by length(N∩M).

[0050] Among them, the average weight of the hit keywords refers to the average weight of the keyword set in the hit keyword list. The calculation method is: ∑weights[N∩M] / length(N∩M), where ∑weights[N∩M] represents the sum of the weights of each keyword in the hit keyword list, and length(N∩M) represents the length of the hit keyword list.

[0051] 5) For the website text features: text length, length of the hit keyword list, average weight of the hit keywords, use the logistic regression model to determine whether the website is pornographic.

[0052] The key points of the present invention mainly include:

[0053] I) Generation of the pornographic keyword library:

[0054] 1) Use word frequency statistics to generate the recognition weight value (i.e., the first recognition weight) of the pornographic keywords that appear in both pornographic websites and non-pornographic websites.

[0055] 2) Generate a set of TextRank weight values of pornographic keywords through TextRank.

[0056] 3) For the pornographic keywords that cannot generate the recognition weight through 1), that is, the pornographic keywords that do not appear in non-pornographic websites, use KNN (K-Nearest Neighbor Algorithm) to realize the mapping of the TextRank weight to the pornographic recognition weight, and then the pornographic recognition weight (i.e., the second recognition weight) can be calculated.

[0057] II) Discrimination of pornographic websites:

[0058] 1) Extract the website text feature values: text length, length of the hit keyword list, average weight of the hit keywords;

[0059] 2) Use the logistic regression model to determine whether the website is pornographic.

[0060] A specific embodiment is provided below. The process of generating the weights of pornographic keywords in this embodiment is as Figure 2 shown, and its steps are described as follows:

[0061] The crawler downloads the text information of the website and tags the website text as pornographic or non - pornographic.

[0062] According to the text information in pornographic websites, manually extract the list of pornographic keywords.

[0063] Segment the text content of the website.

[0064] Calculate the word frequency of keywords in the text of pornographic websites and the word frequency of keywords in the text of non - pornographic websites.

[0065] Determine whether pornographic keywords appear in both pornographic websites and non - pornographic websites.

[0066] If a pornographic keyword appears in both pornographic websites and non - pornographic websites, calculate the pornographic recognition weight of this keyword, that is, the first recognition weight. Suppose it appears A times in pornographic websites and B times in non - pornographic websites, then the pornographic recognition weight is A / (A + B). For pornographic keywords that only appear in pornographic websites and do not appear in non - pornographic websites, generate their pornographic recognition weights according to steps 7) - 9), that is, the second recognition weight.

[0067] Use the TextRank algorithm to process the text information of pornographic websites and generate a list of TextRank weights of pornographic keywords.

[0068] Use the first recognition weight generated in step 6) and the TextRank weight information to train the KNN algorithm. Among them, the TextRank weight information is the input of the algorithm, and the first recognition weight is the output of the algorithm.

[0069] Use the KNN algorithm model generated in step 8) to generate the recognition weights of pornographic keywords not covered in step 6), that is, the second recognition weight.

[0070] Then, this embodiment uses the generated pornographic keywords and their weights to determine whether the website is pornographic:

[0071] Use the existing text information of pornographic websites and non - pornographic websites to train the logistic regression model. The inputs of the model are the length of the website text, the length of the list of hit keywords, and the average weight of the hit keywords; the output of the model is the website category: pornographic / non - pornographic.

[0072] Use the trained logistic regression model in 1) to determine whether an unknown website is pornographic based on its text information. The inputs of the model are the length of the website text, the length of the list of hit keywords, and the average weight of the hit keywords; the output of the model is the category of this unknown website: pornographic / non - pornographic.

[0073] In this embodiment, the method of the present invention is experimentally verified by taking the discrimination of pornographic texts on English websites as an example. The text content of pornographic English websites and non-pornographic websites is crawled through a web crawler, and the websites are marked. 814 pornographic websites and 837 non-pornographic websites are obtained. According to the text information of pornographic websites, a list of pornographic keywords is manually extracted, with a total of 190 keywords. The text of pornographic websites is processed to generate the frequency A of pornographic keywords (including all 190 pornographic keywords, because the list of pornographic keywords is extracted from the text content of pornographic websites). The text of non-pornographic websites is processed to generate the frequency B of pornographic keywords, including 89 pornographic keywords. The recognition weights of the 89 calculated pornographic keywords are: A / (A + B). The TextRank algorithm is used to process the text of the collected pornographic websites to obtain the TextRank weight values of pornographic keywords. Using the weights of the generated 89 pornographic keywords and the TextRank weights of these 89 pornographic keywords, the KNN algorithm is trained. Among them, the TextRank weight value is the input of the KNN algorithm, and the recognition weight of pornographic keywords is the output of the algorithm. Using the obtained KNN model, the recognition weights of the remaining 101 pornographic keywords are generated. The TextRank weights of the remaining 101 pornographic keywords are the input of the model, and thus the recognition weight values of these 101 pornographic keywords are obtained. Thus, the entire list of pornographic keywords and their recognition weight values are obtained. The logistic regression model is trained using the existing text information of pornographic websites and non-pornographic websites. The model inputs are the average weight of the hit keywords, the length of the list of hit keywords on the website, and the length of the website text; the model output is the website category, pornographic / non-pornographic. 1) The trained logistic regression model is used to discriminate whether an unknown website is pornographic based on its text information. The model inputs are the length of the website text, the length of the list of hit keywords on the website, and the average weight of the hit keywords. The test results are shown in Table 1.

[0074] Table 1. Experimental Results

[0075] Precision Recall Test sample Non-pornographic 0.93 0.97 260 websites Pornographic 0.96 0.92 236 websites

[0076] Based on the same inventive concept, another embodiment of the present invention provides a pornographic website detection device based on keyword weights using the above method, which includes:

[0077] A pornographic keyword library generation module for establishing a pornographic keyword library, which includes a list of pornographic keywords and a list of pornographic keyword weights;

[0078] A pornographic website discrimination module for using the pornographic keyword library to extract the text features of the website to be discriminated and using a logistic regression model to discriminate whether the website to be discriminated is pornographic.

[0079] Based on the same inventive concept, another embodiment of the present invention provides an electronic device (such as a computer, a server, a smart phone, etc.), which includes a memory and a processor. The memory stores a computer program, and the computer program is configured to be executed by the processor. The computer program includes instructions for performing each step in the method of the present invention.

[0080] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, a magnetic disk, an optical disk). When the computer program stored in the computer-readable storage medium is executed by a computer, each step of the method of the present invention is implemented.

[0081] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and implement it accordingly. Those of ordinary skill in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification, and the protection scope of the present invention is subject to the scope defined by the claims.

Claims

1. A method for detecting pornographic websites based on keyword weights, characterized in that, it includes the following steps: Establish a pornographic keyword library, which contains a list of pornographic keywords and a list of pornographic keyword weights; Using the pornographic keyword library, extract the text features of the website to be judged, and use a logistic regression model to judge whether the website to be judged is pornographic; The establishment of the pornographic keyword library includes: Download the text information of various websites from the network and label them as pornographic or non-pornographic, so as to generate a set of pornographic websites and a set of non-pornographic websites; According to the text information in the pornographic websites, manually extract and generate a list of pornographic keywords; Use word frequency statistics to generate recognition weights for pornographic keywords that appear in both pornographic websites and non-pornographic websites, which are called the first recognition weights; For pornographic keywords that only appear in pornographic websites and do not appear in non-pornographic websites, use the KNN algorithm to map the TextRank weight to the pornographic recognition weight to obtain the second recognition weight; The first recognition weight and the second recognition weight constitute the list of pornographic keyword weights; The following steps are used to calculate the second recognition weight: Use the TextRank algorithm to process the text information of pornographic websites to generate a list of TextRank weights of pornographic keywords; Use the generated list of pornographic keywords with the first recognition weight and the generated list of TextRank weights of pornographic keywords to train the KNN algorithm. The first recognition weight is used as the output Y of the KNN algorithm, and the TextRank weight of the pornographic keyword is used as the input X of the KNN algorithm; Use the generated KNN algorithm to generate its pornographic recognition weight, that is, the second recognition weight, for the list of pornographic keywords that do not appear in non-pornographic websites, using its TextRank weight value; The text features of the website to be judged include: text length, length of the list of hit keywords, and average weight of hit keywords; The length of the list of hit keywords refers to the number of pornographic keyword entries contained in the word segmentation of the website to be judged. Let the set of word segmentation entries of the website to be judged be N, and the set of pornographic keyword entries be M, then the length of the list of hit keywords is the length of the set N∩M; the average weight of hit keywords refers to the average weight of the keyword set in the list of hit keywords, and the calculation method is: ∑weights[N∩M] / length(N∩M), where ∑weights[N∩M] represents the sum of the weights of each keyword in the list of hit keywords, and length(N∩M) represents the length of the list of hit keywords.

2. The method according to claim 1, characterized in that, Assume that the word frequency statistical value of a certain pornographic keyword in the set of pornographic websites is A, and the word frequency statistical value in the set of non-pornographic websites is B, then the first recognition weight of this pornographic keyword is A / (A + B).

3. A device for detecting pornographic websites based on keyword weights using the method described in claim 1 or 2, characterized in that, it includes: The pornographic keyword library generation module is used to establish a pornographic keyword library, and the pornographic keyword library includes a list of pornographic keywords and a list of weights of pornographic keywords; The pornographic website discrimination module is used to utilize the pornographic keyword library to extract the text features of the website to be discriminated, and use a logistic regression model to discriminate whether the website to be discriminated is pornographic.

4. An electronic device, characterized in that, it includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing the method according to claim 1 or 2.

5. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, and when the computer program is executed by a computer, the method according to claim 1 or 2 is implemented.

Citation Information

Patent Citations

  • Yellow-related and block-related website detection method based on mixed feature analysis

    CN112347244A