Method, device, electronic equipment and readable storage medium for identifying illegal website

By analyzing the characteristics of the request process and page source code, illegal websites are identified, solving the problems of low accuracy and low efficiency in existing technologies. This achieves efficient automated identification and reduces costs.

CN115766167BActive Publication Date: 2026-03-27BEIJING KNOWNSEC INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-08
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively identify illegal websites, especially when these websites use covert techniques to evade crawling. The low accuracy of identification and reliance on manual review result in inefficiency and high costs.

Method used

By sending website access requests, analyzing the request process and obtaining the page source code, request characteristics, IP characteristics, and HTML tag characteristics are extracted. These characteristics are then used to identify illegal websites and avoid directly obtaining the website's real content.

Benefits of technology

It improved the accuracy of identifying illegal websites, reduced reliance on manual review, increased identification efficiency, and lowered hardware and labor costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115766167B_ABST
    Figure CN115766167B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method and device for identifying illegal websites, electronic equipment and readable storage medium, and relate to the technical field of computer. The method comprises: obtaining page source code of a website to be identified corresponding to a website access request by sending the website access request; obtaining request features and IP features according to request process analysis; extracting features from the page source code to obtain code features in terms of HTML tags; and judging whether the website to be identified is an illegal website according to target features, wherein the target features are the code features, or at least one of the request features and the IP features and the code features. In this way, the identification accuracy of illegal websites can be effectively improved without obtaining the real and complete content of the website, and the identification efficiency can be improved without manual review.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computers, in particular to a method and device for identifying illegal websites, an electronic device and a readable storage medium. BACKGROUND

[0002] The Internet provides people with a large amount of information, and websites are an important portal for people to obtain information. Some unscrupulous people use websites to spread illegal and irregular content such as political reaction, pornography, terrorism and gambling in order to gain benefits, which is seriously detrimental to the health and safety of the network. In the early days of the Internet, web pages were presented in the form of static content, and through web scraping, complete page content could be obtained, and with manual review, illegal and irregular content could be easily identified and had no place to hide.

[0003] However, with the rapid development of Internet technology, unscrupulous people have become more and more sophisticated in hiding themselves and evading scraping. The development of technology has given these unscrupulous people an opportunity to not only spread illegal and irregular content faster and more widely, but also make it more difficult for people to identify and regulate. SUMMARY

[0004] The embodiments of the present application provide a method and device for identifying illegal websites, an electronic device and a readable storage medium, which can effectively improve the identification accuracy of illegal websites without obtaining the real and complete content of the website, and can improve the identification efficiency without manual review.

[0005] The embodiments of the present application can be implemented as follows:

[0006] In a first aspect, the embodiments of the present application provide a method for identifying illegal websites, the method comprising:

[0007] sending a website access request to obtain the page source code of a website to be identified corresponding to the website access request;

[0008] obtaining request features and IP features from the request process analysis;

[0009] extracting features from the page source code to obtain code features of HTML tags;

[0010] judging whether the website to be identified is an illegal website according to target features, wherein the target features are the code features, or at least one of the request features and the IP features and the code features.

[0011] In a second aspect, the embodiments of the present application provide a device for identifying illegal websites, the device comprising:

[0012] obtain a page source code of a website to be identified corresponding to the website access request by sending a website access request;

[0013] a first feature obtaining module configured to obtain a request feature and an IP feature according to request process analysis;

[0014] a second feature obtaining module configured to extract features from the page source code to obtain a code feature in terms of HTML tags;

[0015] a recognition module configured to determine whether the website to be identified is a rule violation website according to a target feature, wherein the target feature is the code feature, or at least one of the request feature and the IP feature and the code feature.

[0016] In a third aspect, an electronic device is provided, which includes a processor and a memory. The memory stores machine executable instructions which can be executed by the processor. The processor can execute the machine executable instructions to implement the rule violation website recognition method described in the foregoing embodiments.

[0017] In a fourth aspect, a readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the rule violation website recognition method described in the foregoing embodiments.

[0018] The rule violation website recognition method, device, electronic device and readable storage medium provided by the embodiments of the present application can obtain a page source code of a website to be identified corresponding to a website access request by sending a website access request, and then extract features from the page source code to obtain a code feature in terms of HTML tags. In addition, a request feature and an IP feature are obtained according to request process analysis. Finally, the code feature is taken as a target feature, or at least one of the request feature and the IP feature and the code feature is taken as the target feature, and whether the website to be identified is a rule violation website is determined according to the target feature. In this way, the identification accuracy of rule violation websites can be effectively improved in the case that the real and complete content of a website cannot be obtained, and the identification efficiency can be improved without manual review. BRIEF DESCRIPTION OF DRAWINGS

[0019] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the present application, and therefore should not be considered as a limitation to the scope. For those skilled in the art, other related drawings can also be obtained without creative labor.

[0020] Figure 1 The block schematic diagram of the electronic device provided by the embodiments of the present application is shown.

[0021] Figure 2 A flowchart of the method for identifying a non-compliant website provided in the embodiments of the present application is shown in the figure.

[0022] Figure 3 For Figure 2 A flowchart of the sub-steps included in step S120 is shown in the figure.

[0023] Figure 4 A block diagram of the device for identifying a non-compliant website provided in the embodiments of the present application is shown in the figure.

[0024] Icon: 100-electronic device; 110-memory; 120-processor; 130-communication unit; 200-device for identifying a non-compliant website; 210-code obtaining module; 220-first feature obtaining module; 230-second feature obtaining module; 240-identifying module. DETAILED DESCRIPTION

[0025] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments of the present application. The components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0026] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the claimed present application, but only represents selected embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative efforts fall within the scope of the present application.

[0027] It should be noted that the relational terms such as “first” and “second” and the like are merely used to distinguish one entity or action from another, and do not necessarily require or imply that there is any such actual relationship or order between these entities or actions. Moreover, the terms “include”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. Without more limitations, the element defined by the statement “including a…” does not exclude the presence of additional identical elements in the process, method, article or device including the element.

[0028] Currently, the identification of a non-compliant website is generally performed in the following two ways.

[0029] Method one: through the way of crawler, directly grab the web content, and then determine whether the grabbed web content includes the violation through artificial audit, if it includes, it is determined to be a violation website. However, the violation website often uses technical means to hide, and it is difficult to grab meaningful content, that is, the grabbed content is either random code or dynamic js and other anti-scraping content, which cannot be used for judgment. Moreover, the massive page content needs to be audited by artificial, which is prone to insufficient manpower and high cost.

[0030] Method two: instead of directly crawling the content, the behavior of people is simulated through the headless browser to obtain the content of the web page, and then the content is audited by artificial. In this way, the browser consumes more resources, and the efficiency of obtaining the content is extremely low. In the face of a large number of websites, a single machine is a performance disaster, and multiple machines are a cost disaster, and the artificial audit is high in cost and low in efficiency.

[0031] The present inventors have found that although the violation website escapes supervision through technical means, it has strong concealment. However, from a technical point of view, these illegal and irregular websites also expose many commonalities and have some common characteristics. The characteristics exhibited by these websites can be summarized from the request behavior, website IP, and obtainable web content. These characteristics can be used to effectively identify these illegal and irregular websites that try to escape supervision. Based on this, the embodiments of the present application provide a violation website identification method and device, an electronic device and a readable storage medium, which identify whether a website is a violation website according to the code characteristics of the website, or according to at least one of the request characteristics and the IP characteristics and the code characteristics of the website. In this way, the identification accuracy of the violation website can be effectively improved without obtaining the real and complete content of the website, and the identification efficiency can be improved without artificial audit.

[0032] It should be noted that the defects of the above solutions are the results obtained by the inventors after practice and careful research, and therefore, the discovery process of the above problems and the solutions proposed by the embodiments of the present application to solve the above problems should be the contributions made by the inventors to the present application in the process of the present application.

[0033] Some embodiments of the present application will be described in detail below with reference to the accompanying drawings. In the case of no conflict, the embodiments described below and the features in the embodiments can be combined with each other.

[0034] Please refer to Figure 1 , Figure 1A block schematic diagram of an electronic device 100 is provided in the embodiments of the present application. The electronic device 100 can be, but is not limited to, a computer, a server, etc. The electronic device 100 includes a memory 110, a processor 120 and a communication unit 130. The memory 110, the processor 120 and the communication unit 130 are electrically connected to each other directly or indirectly to realize data transmission or interaction. For example, the elements can be electrically connected to each other through one or more communication buses or signal lines.

[0035] The memory 110 is configured to store programs or data. The memory 110 can be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc.

[0036] The processor 120 is configured to read / write the data or programs stored in the memory 110 and perform corresponding functions. For example, the memory 110 stores a violation website identification apparatus 200. The violation website identification apparatus 200 includes at least one software function module stored in the memory 110 in the form of software or firmware. The processor 120 runs the software programs and modules stored in the memory 110, such as the violation website identification apparatus 200 in the embodiments of the present application, to perform various function applications and data processing, i.e., to realize the violation website identification method in the embodiments of the present application.

[0037] The communication unit 130 is configured to establish a communication connection between the electronic device 100 and other communication terminals through a network and to receive / transmit data through the network.

[0038] It should be understood that, Figure 1 The structure shown is only a structural schematic diagram of the electronic device 100. The electronic device 100 can further include more or fewer components than those shown in the embodiments of the present application or have a different configuration from that shown in the embodiments of the present application. Figure 1 The components shown in the embodiments of the present application can be realized in hardware, software or a combination thereof. Figure 1 Figure 1 Please refer to ,

[0039] Figure 2 , Figure 2 ​FIG. 1 is a schematic diagram of a flow of a method for identifying a non-compliant website according to an embodiment of the present application. The method can be applied to the electronic device 100 described above. The specific flow of the method for identifying a non-compliant website is described in detail below. In this embodiment, the method can include steps S110-S140.

[0040] In step S110, a page source code of a website to be identified corresponding to a website access request is obtained by sending the website access request.

[0041] In this embodiment, a domain name to be accessed or an IP address to be accessed can be determined in advance. The domain name or the IP address can be manually input by a user, can be selected by the user from a plurality of options, or can be sent by another device, and is not limited specifically herein. The domain name to be accessed or the IP address to be accessed is used to access the website to be identified. If the domain name to be accessed is determined, the IP address to be accessed can be determined according to the domain name to be accessed. The website access request for accessing the website to be identified can be sent according to the IP address to be accessed, so that the page source code of the website to be identified can be obtained directly.

[0042] The page source code can be obtained by a crawler based on the IP address to be accessed. If the website to be identified uses an anti-scraping technology, the hidden content can not be included in the obtained page source code.

[0043] In step S120, request features and IP features are analyzed according to a request process.

[0044] The request features and the IP features are analyzed according to the entire request process of sending the website access request and obtaining a response. The IP address is an IP address used when responding to the website access request.

[0045] In step S130, code features of HTML tags are obtained by performing feature extraction on the page source code.

[0046] The code features of the HTML tags can be obtained by performing feature extraction on the HTML (Hyper Text Markup Language) tags included in the page source code.

[0047] In step S140, it is determined whether the website to be identified is a non-compliant website according to target features.

[0048] The code feature can be taken as the target feature, or at least one of the request feature and the IP feature and the code feature can be taken as the target feature, and then whether the website to be identified is a violation website is identified according to the target feature. If the target feature does not include the request feature and the IP feature, step S120 can not be performed.

[0049] The embodiment of the present application extracts the website features according to the request process and the request result, and finally identifies whether the website to be identified is a violation website according to the extracted website features. This method has the efficiency of directly grabbing website content and the accuracy of identifying violation websites.

[0050] In the embodiment, the request feature can include at least one of the request response time, the response code and the jump number, and the IP feature can include the request response IP address and / or the home of the request response IP address. The features included in the request feature and the IP feature can be set according to actual needs, for example, the request feature can include the request response time, the response code and the jump number, and the IP feature can include the request response IP address and the home of the request response IP address.

[0051] Please refer to Figure 3 , Figure 3 For Figure 2 The sub-steps included in step S120 are shown in the schematic diagram. In the embodiment, step S120 can include sub-step S121 to sub-step S123. It should be noted that if some features are not included in the request feature or the IP feature, the step of obtaining the corresponding feature can not be performed.

[0052] In sub-step S121, the request response time, the response code and the jump number are obtained.

[0053] In sub-step S122, the request response IP address corresponding to the website access request is obtained.

[0054] In sub-step S123, the home of the request response IP address is obtained.

[0055] The request response time is the time length from sending the website access request to obtaining the response. Violation websites often have server resources shortage, overseas assumption, subjective evasion, etc., and the response time of the website is often longer, so the request response time can be extracted as a feature for identifying whether the website to be identified is a violation website.

[0056] The response code is a request response code, which can be used to judge whether the website jumps. Violation websites often jump pages on the homepage in order to avoid being grabbed.

[0057] The illegal website often jumps several times and finally jumps to a real destination website. Therefore, the number of jumps can also be extracted to determine whether a website is an illegal website.

[0058] The request response IP address is an IP address used for the response corresponding to the website access request. If a jump is made, the request response IP address and the IP address to be accessed are not the same IP address at this time, and the request response IP address is the IP address of the website after the final jump. Illegal websites of the same type often have the same IP address, although the sites are different. Therefore, the request response IP address can be extracted, and subsequent comparison with the IP address of a known illegal website can be performed to determine whether it is illegal from the IP address.

[0059] The request response IP address can also be queried through an open source IP library or through a third party information to obtain the location of the request response IP address. In order to evade monitoring, illegal websites are often assumed to be located abroad, and therefore the location of the request response IP address can also be used as a judgment factor for determining whether it is an illegal website.

[0060] Optionally, in the embodiment, the code features can include at least one type of features related to a title tag, a head tag, and a body tag, which can be set according to actual needs. For example, the code features can include features related to a title tag, a head tag, and a body tag.

[0061] In the case where the code features include the features related to the title tag, the features related to the title tag can be obtained in at least one of the following ways.

[0062] The total length of the page content is obtained according to the page source code. That is, the number of words of the page source code is obtained. The content of an illegal website is something that should not be seen, and in order to avoid being directly requested, the page content obtained by the current request is usually very little, and the real content is dynamically generated by front-end technology to be hidden.

[0063] It is determined whether the page source code includes a title tag. The purpose of a website is to be browsed by people, and in order to increase exposure, a website usually has a bright and upright title. The purpose of an illegal website is not SEO (Search Engine Optimization), and these websites often hide their <title>Tag.< / title>

[0064] In a case where the page source code includes a title tag, it is determined whether the title tag is normally encoded. The text content of the title tag can be directly or by using a third-party attack diagnosed for encoding to determine <title>Whether the tag is normal encoding. Some illegal websites contain< / title> <title>tags, but they are disguised in a way that ordinary people can not understand the code.< / title>

[0065] It is determined whether the encoding mode of the page source code is a preset encoding mode. The preset encoding mode can be a general encoding mode, such as the utf8 encoding mode. In order to avoid being detected, the content of a violation website is usually encoded in a rare encoding mode rather than a general encoding mode.

[0066] Based on the above description, the title tag related features can include at least one of the total length of the page content, whether the page source code includes a title tag, whether the title tag of the page source code is normally encoded, and whether the encoding mode of the page source code is a preset encoding mode. The specific obtaining manner of the title tag related features is determined by the specific features included in the title tag related features, and the title tag related features can be obtained by at least one of the above four ways.

[0067] In a case where the code features include the head tag related features, the head tag related features can be obtained by at least one of the following ways. Generally, a page includes only one head tag, but it is also possible that a page includes multiple head tags.

[0068] For each head tag in the page source code, it is determined whether the head tag contains a script tag. Violation websites often use JavaScript to dynamically generate violation content in the tag, so whether the head tag contains a script tag can be used as a reference factor for determining whether the website is a violation website.

[0069] In a case where the head tag contains a script tag, the first number of script tags contained in the head tag is analyzed. That is, the number of script tags included in each head tag in the page source code is obtained. The more JavaScript code is introduced, the more content is dynamically generated, and the more obvious the intention to avoid direct observation is. Therefore, the number of script tags included in the head tag can also be used to identify whether the website is a violation website.

[0070] The first quantity ratio is obtained by dividing the number of script tags contained in the head tag by the number of tags in the head tag excluding the meta tags. That is, the number of tags A1 in a head tag in the page source code excluding the meta tags is calculated first, and then the number of script tags B1 contained in the head tag is divided by the number of tags A1 calculated above to obtain the first quantity ratio.

[0071] The number of tags in the head tag excluding the meta tags is calculated. <meta> The meta tag only provides some meta information of the webpage and can be removed. A normal and rich webpage tends to be inclined to provide more information in the body tag. <style>或<link>引入样式、图片等内容,如果这些标签较少或没有,而<script>标签却占有极大比重,说明网页试图生成大量的动态内容,避免被抓取。因此,可以将所述第一数量占比作为违规网站识别的一个参考因素。

[0072] 计算得到该head标签中含有的script标签的标签内容长度在该head标签中除去meta标签后的标签内容长度中的第一长度占比。即,先计算出页面源代码中的一个head标签中除去meta标签之后的标签内容长度A2,然后将该head标签中含有的script标签的标签内容长度B2除以上述计算出的标签内容长度占比A2,得到所述第一长度占比。

[0073] 使用第一长度占比与使用第一数量占比的原因相同,除了<script>标签的数量外,内容长度占比也极为重要,为了隐藏真实的违规内容,违规网站往往需要通过复杂的JavaScript代码来实现逻辑。因此,也可以使用第一长度占比作为违规网站识别的一个参考因素。

[0074] 基于上述描述可知,所述head标签相关的特征可以包括所述页面源代码的各head标签中是否含有script标签、head标签中含有的script标签的数量、head标签中去meta标签后script标签数量占比、head标签中去meta标签后script标签内容长度占比四者中的至少一项。所述head标签相关的特征的具体获得方式由所述head标签相关的特征中所包括的具体特征确定,可以通过以上四种方式中的至少一种方式获得所述head标签相关的特征。

[0075] 在所述代码特征包括所述body标签相关的特征的情况下,可以通过以下至少一种方式获得所述body标签相关的特征。

[0076] 分析得到所述页面源代码中的body标签的内容长度。违规网站偏向于使用JavaScript动态生成网页内容,<body>标签的内容往往极少,甚至什么也没有。因此,可以将所述页面源代码中的body标签的内容长度作为判断网站是否为违规网站的一个参考因素。

[0077] 判断所述body标签内是否含有script标签。同网页的<head>标签,违规网站也可以在<body>标签中利用JavaScript动态生成违规内容。因此,可以将所述body标签内是否含有script标签作为判断网站是否为违规网站的一个参考因素。

[0078] 在所述body标签内含有script标签的情况下,分析得到所述body标签中含有的script标签的第二标签数量。引入的JavaScript代码越多,动态生成的内容越多,避免直接观察的意图越明显,因此可以所述body标签中含有的script标签的数量作为判断网站是否为违规网站的一个参考因素。

[0079] 计算得到所述body标签中含有的script标签的数量在所述body标签中除去style标签后的标签数量中的第二数量占比。即,先计算出body标签中除去style标签之后的标签数量A3,然后将body标签中含有的script标签的数量B3除以上述计算出的标签数量A3,得到所述第二数量占比。

[0080] 不同于<head>标签,正常网页的<body>标签中通常含有很多的标签,因<style>标签只为装饰内容,可以去掉<style>标签,如果<script>标签占比较大,或者都是<script>标签,则为违法违规网站的几率极高。因此,可以将所述第二数量占比作为违规网站识别的一个参考因素。

[0081] 计算得到所述body标签中含有的script标签的标签内容长度在所述body标签中除去style标签后的标签内容长度中的第二长度占比。即,先计算出body标签中除去style标签之后的标签内容长度A4,然后将body标签中含有的script标签的标签内容长度B4除以上述计算出的标签内容长度A4,得到所述第二长度占比。

[0082] 使用所述第二长度占比的原因与上述使用所述第二数量占比的原因相同,除了<script>标签的数量外,内容长度占比也极为重要,为了隐藏真实的违规内容,违规网站往往需要通过复杂的JavaScript代码来实现逻辑。因此,也可以使用第二长度占比作为违规网站识别的一个参考因素。

[0083] 基于上述描述可知,所述body标签相关的特征可以包括body标签的内容长度、body标签中是否含有script标签、body标签中含有的script标签的数量、body标签去style标签后script标签数量占比、body标签去style标签后script标签内容长度占比五者中的至少一项。所述body标签相关的特征的具体获得方式由所述body标签相关的特征中所包括的具体特征确定,可以通过以上五种方式中的至少一种方式获得所述body标签相关的特征。

[0084] 在获得目标特征的情况下,可以根据预设的判断规则及所述目标特征,识别所述待识别网站是否为违规网站。其中,所述预设的判断规则可以结合实际需求确定。比如,可以预先设置与所述目标特征中的各种特征对应的子规则,根据各条子规则确定出所述目标特征中的各种特征对应的违规网站概率,进而基于各种特征对应的违规网站概率计算得到违规网站总概率,最后将该违规网站总概率与预设概率进行比较,若大于,则确定待识别网站为违规网站,反之则确定是正常网站。其中,违规网站概率表示网站是违规网站的概率。

[0085] 或者,可以利用预先训练好的分类模型,根据所述目标特征,识别所述待识别网站是否为违规网站。其中,所述分类模型根据多个样本训练得到,每个所述样本中包括样本特征及所述样本特征对应的网站是否为违规网站。如此,可通过机器学习算法训练得到分类模型,利用利用该分类模型对违规网站进行识别。

[0086] 其中,在使用所述待访问IP地址访问的过程中,若未发生跳转,则所述待访问IP地址指向的网站即为待识别网站。若发生了跳转,可以将跳转后的网站作为待识别网站;也可以将所述待访问IP地址指向的网站以及最后跳转至的网址均作为所述待识别网站,在此情况下,若至少基于跳转后的网站的代码特征等确定跳转后的网站为违规网站,则所述待访问IP地址指向的网站也同时被确定为违规网站。

[0087] 在本实施例中,从请求行为、IP信息、源代码三个层面归纳出了违规网站普遍具有的各种类特征,进而可利用判断规则及机器学习手段,对违规网站进行有效的识别。本申请实施例可以在不能获取到网站真实完整内容(包括通过各种反爬手段直接或间接隐藏起来的内容)的前提下,有效提高违规网站识别率。上述方法不仅有利于强化监管部门对互联网违法违规内容的打击力度,同时,对企业应用来说,兼顾运作效率的同时可极大降低硬件成本。

[0088] 为了执行上述实施例及各个可能的方式中的相应步骤,下面给出一种违规网站识别装置200的实现方式,可选地,该违规网站识别装置200可以采用上述图1所示的电子设备100的器件结构。进一步地,请参照图4,图4为本申请实施例提供的违规网站识别装置200的方框示意图。需要说明的是,本实施例所提供的违规网站识别装置200,其基本原理及产生的技术效果和上述实施例相同,为简要描述,本实施例部分未提及之处,可参考上述的实施例中相应内容。在本实施例中,所述违规网站识别装置200可以包括:代码获得模块210、第一特征获得模块220、第二特征获得模块230及识别模块240。

[0089] 所述代码获得模块210,用于通过发送网站访问请求,获得与所述网站访问请求对应的待识别网站的页面源代码。

[0090] 所述第一特征获得模块220,用于根据请求过程分析得到请求特征及IP特征。

[0091] 所述第二特征获得模块230,用于对所述页面源代码进行特征提取,获得HTML标签方面的代码特征。

[0092] 所述识别模块240,用于根据目标特征,判断所述待识别网站是否为违规网站,其中,所述目标特征为所述代码特征,或者为所述请求特征及所述IP特征中的至少一项和所述代码特征。

[0093] 可选地,上述模块可以软件或固件(Firmware)的形式存储于图1所示的存储器110中或固化于电子设备100的操作系统(Operating System,OS)中,并可由图1中的处理器120执行。同时,执行上述模块所需的数据、程序的代码等可以存储在存储器110中。

[0094] 本申请实施例还提供一种可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现所述的违规网站识别方法。

[0095] 综上所述,本申请实施例提供一种违规网站识别方法、装置、电子设备及可读存储介质,通过发送网站访问请求,获得与所述网站访问请求对应的待识别网站的页面源代码,进而通过对该页面源代码进行提取特征,得到HTML标签方面的代码特征;以及根据请求过程分析得到请求特征及IP特征;最后,将所述代码特征作为目标特征,或者将所述请求特征及所述IP特征中的至少一项和所述代码特征作为目标特征,根据该目标特征识别该待识别网站是否是违规网站。如此,可在不能获取网站真实完整内容的情况下,有效提高违规网站的识别准确率,并且无需人工审核,可提高识别效率。

[0096] 在本申请所提供的几个实施例中,应该理解到,所揭露的装置和方法,也可以通过其它的方式实现。以上所描述的装置实施例仅仅是示意性的,例如,附图中的流程图和框图显示了根据本申请的多个实施例的装置、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段或代码的一部分,所述模块、程序段或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现方式中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和 / 或流程图中的每个方框、以及框图和 / 或流程图中的方框的组合,可以用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。

[0097] 另外,在本申请各个实施例中的各功能模块可以集成在一起形成一个独立的部分,也可以是各个模块单独存在,也可以两个或两个以上模块集成形成一个独立的部分。

[0098] 所述功能如果以软件功能模块的形式实现并作为独立的产品销售或使用时,可以存储在一个计算机可读取存储介质中。基于这样的理解,本申请的技术方案本质上或者说对现有技术做出贡献的部分或者该技术方案的部分可以以软件产品的形式体现出来,该计算机软件产品存储在一个存储介质中,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行本申请各个实施例所述方法的全部或部分步骤。而前述的存储介质包括:U盘、移动硬盘、只读存储器(ROM,Read-Only Memory)、随机存取存储器(RAM,Random Access Memory)、磁碟或者光盘等各种可以存储程序代码的介质。

[0099] 以上所述仅为本申请的可选实施例而已,并不用于限制本申请,对于本领域的技术人员来说,本申请可以有各种更改和变化。凡在本申请的精神和原则之内,所作的任何修改、等同替换、改进等,均应包含在本申请的保护范围之内。< / style>

Claims

1. A method for identifying illegal websites, characterized in that, The method includes: By sending a website access request, the source code of the page of the website to be identified corresponding to the website access request is obtained; The request characteristics and IP characteristics are obtained by analyzing the request process. Feature extraction is performed on the source code of the page to obtain code features related to HTML tags; Based on the target features, it is determined whether the website to be identified is a violation website, wherein the target features are the code features, or at least one of the request features and the IP features and the code features; When the code features include head tag-related features, the head tag-related features include: the percentage of the number of script tags contained in the head tag out of the total number of tags in the head tag after removing the meta tags; and the percentage of the length of the script tag content contained in the head tag out of the total length of the tag content in the head tag after removing the meta tags. When the code features include features related to the body tag, the features related to the body tag include: the percentage of the number of script tags contained in the body tag relative to the number of tags in the body tag after removing the style tag; and the percentage of the length of the content of the script tags contained in the body tag relative to the length of the content of the body tag after removing the style tag.

2. The method according to claim 1, characterized in that, The code features include at least one of the following: features related to the title tag, features related to the head tag, and features related to the body tag.

3. The method according to claim 2, characterized in that, When the code features include features related to the title tag, the features related to the title tag are obtained through at least one of the following methods: The total length of the page content is obtained from the page source code. Determine whether the page source code includes a title tag; If the page source code includes a title tag, determine whether the title tag is correctly encoded. Determine whether the encoding method of the page source code is the preset encoding method.

4. The method according to claim 2, characterized in that, When the code features include head tag-related features, the head tag-related features are obtained through at least one of the following methods: For each head tag in the page source code, determine whether the head tag contains a script tag; If the head tag contains script tags, analyze and obtain the first number of script tags contained in the head tag; The percentage of the number of script tags contained in the head tag out of the total number of tags in the head tag after removing meta tags is calculated. The length of the script tag content contained in the head tag is calculated as the first proportion of the length of the tag content after removing the meta tag in the head tag.

5. The method according to claim 2, characterized in that, When the code features include features related to the body tag, the features related to the body tag are obtained through at least one of the following methods: The length of the body tag content in the page source code was obtained through analysis; Determine whether the body tag contains a script tag; If the body tag contains a script tag, analyze and obtain the number of second tags containing script tags in the body tag; The number of script tags contained in the body tag is calculated as the second percentage of the total number of tags in the body tag after removing the style tag; The length of the script tag contained in the body tag is calculated as the second length ratio of the length of the tag content in the body tag after removing the style tag.

6. The method according to claim 1, characterized in that, The request characteristics include at least one of request-response time, response code, and number of redirects; the IP characteristics include the request-response IP address and / or the location of the request-response IP address; the step of analyzing the request process to obtain the request characteristics and IP characteristics includes: Obtain the request response time, response code, and number of redirects; Obtain the IP address of the request response corresponding to the website access request; Obtain the location of the IP address that triggered the request response.

7. The method according to any one of claims 1-6, characterized in that, The step of determining whether the website to be identified is a violation website based on target characteristics includes: Based on preset judgment rules and the target characteristics, identify whether the website to be identified is a violation website; or, Using a pre-trained classification model, the website to be identified is determined to be an illegal website based on the target features. The classification model is trained on multiple samples, and each sample includes sample features and whether the website corresponding to the sample features is an illegal website.

8. A device for identifying illegal websites, characterized in that, The device includes: The code acquisition module is used to obtain the page source code of the website to be identified corresponding to the website access request by sending a website access request; The first feature acquisition module is used to analyze the request process to obtain request features and IP features; The second feature acquisition module is used to extract features from the page source code to obtain code features related to HTML tags; The identification module is used to determine whether the website to be identified is a violation website based on the target features, wherein the target features are the code features, or at least one of the request features and the IP features and the code features; When the code features include head tag-related features, the head tag-related features include: the percentage of the number of script tags contained in the head tag out of the total number of tags in the head tag after removing the meta tags; and the percentage of the length of the script tag content contained in the head tag out of the total length of the tag content in the head tag after removing the meta tags. When the code features include features related to the body tag, the features related to the body tag include: the percentage of the number of script tags contained in the body tag relative to the number of tags in the body tag after removing the style tag; and the percentage of the length of the content of the script tags contained in the body tag relative to the length of the content of the body tag after removing the style tag.

9. An electronic device, characterized in that, The method includes a processor and a memory, the memory storing machine-executable instructions that can be executed by the processor, the processor executing the machine-executable instructions to implement the illegal website identification method according to any one of claims 1-7.

10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the illegal website identification method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Malicious webpage recognition model, recognition model establishing method, recognition method and system

    CN111259219A

  • Malicious website comprehensive evaluation method and system and storage medium

    CN115130104A