A network content classification method and device

By constructing a graph representation of the network content and using the graph neural network for classification, the problem of easy replication and dissemination within the Internet black industry is solved, and low-cost and efficient black industry content recognition and crackdown is achieved.

CN113722622BActive Publication Date: 2025-06-06SHANGHAI XINFANG INTELLIGENT SYST CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111026455.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-02
Publication Date
2025-06-06
Estimated Expiration
2041-09-02

AI Technical Summary

Technical Problem

How to establish a low-cost long-term mechanism to maintain a high-pressure crackdown on content in the Internet black industry, especially when it is easy to copy and disseminate in the black and gray industry.

Method used

By obtaining the web page URL to be classified from the network, crawling the web page content using the crawler engine, and saving it as an mhtml document, building a graph representation of the web page content, and using the graph neural network to classify and recognize the graph to achieve effective classification and recognition of network text and picture content.

Benefits of technology

It has achieved efficient identification and classification of content from Internet black industry, reduced its dependence on social human resources, provided a low-cost long-term mechanism to continuously crack down on content from Internet black industry.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113722622B_ABST
    Figure CN113722622B_ABST
Patent Text Reader

Abstract

The present application discloses a method for classifying network content, including: obtaining a web page URL to be classified from the network and writing it into a target URL document; crawling the content of each URL web page in the target URL document in the network through a crawler engine, and saving the content of each web page as an mhtml document, and writing it into an Internet content archive database; constructing a web page content graph representation corresponding to the web page URL according to the saved mhtml document; the web page graph representation includes a text graph and an image graph; classifying and identifying the constructed web page content graph representation graph, and using the classification and recognition results as the classification and recognition results of the web page URL corresponding to the web page content graph representation; wherein, when classifying and identifying the graph, the feature vector of the text graph and the feature vector of the image graph are determined by convolution and pooling operations, and the feature vector of the text graph and the feature vector of the image graph are spliced ​​as the feature vector of the web page content graph representation. The application of the present application can effectively classify network content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to network processing technology, and in particular to a method and device for classifying network content. Background Art

[0002] With the vigorous development of the Internet information technology industry, especially in the context of network intelligence and digitalization, illegal and criminal activities such as hacker attacks, online fraud, online pornography, and online theft are quietly growing. These black and gray production contents are on the edge of network supervision, disrupting people's normal work and life order, and breeding social instability. In recent years, relevant government departments have repeatedly carried out special operations to crack down on Internet black production contents, effectively cracking down on the spread of Internet black production contents. However, since Internet content itself is easy to copy and spread, Internet black production often revives after the end of special operations. At the same time, special operations will also occupy a large amount of social human resources, especially for related content that is spread by using pictures and text content displayed on pictures, which is more difficult to identify. How to establish a low-cost and long-term mechanism to maintain a high-pressure crackdown on Internet black production is a problem worth studying. Summary of the invention

[0003] The present application provides a method and device for classifying network content, which can effectively classify and identify text and image content on the network.

[0004] To achieve the above objectives, this application adopts the following technical solutions:

[0005] A method for classifying network content, comprising:

[0006] Get the URL of the web page to be classified from the network and write it into the target URL document;

[0007] The crawler engine crawls the content of each URL web page in the target URL document on the network, and saves the content of each web page as an mhtml document and writes it into the Internet content archive database;

[0008] According to the saved mhtml document, a web page content graph representation corresponding to the web page URL is constructed; wherein the web page graph representation includes a text graph and an image graph, the vertices in the text graph are text vertices, and the vertices in the image graph are image vertices;

[0009] A graph neural network is used to classify and recognize the graph representation of the web page content, and the classification and recognition results are used as the classification and recognition results of the web page URL corresponding to the graph representation of the web page content; wherein, when performing the graph classification and recognition, the feature vector of the text graph and the feature vector of the picture graph are determined through convolution and pooling operations, and the feature vector of the text graph and the feature vector of the picture graph are concatenated as the feature vector of the web page content graph representation.

[0010] Preferably, the writing into the target URL document comprises: filtering the web page URLs to be classified using a white list, and writing the web page URLs that pass the white list into the target URL document.

[0011] Preferably, crawling the contents of each URL web page in the target URL document on the network by using a crawler engine includes:

[0012] Each web page URL in the target URL document is input into a Redis task queue, and the crawler engine uses multiple crawler processes to read the web page URL from the Redis task queue respectively, and performs content loading and rendering.

[0013] Preferably, in the Internet content archive database, each mhtml document includes at least one of the following attributes: web page URL, web page host, web page domain, web page server IP address, and web page Title field content.

[0014] Preferably, for a text graph, the web page content graph representation corresponding to the web page URL is constructed including:

[0015] The webpage text content stored in the mhtml document is represented as an HTML tree; wherein the text content elements in the webpage content are configured as nodes of the HTML tree, and for pictures including text information in the webpage, the text embedded in the pictures is extracted and identified, and a text node is generated and added to the HTML tree;

[0016] A graph G=(V, E) is constructed using the nodes in the HTML tree and the relationships between the nodes; wherein each vertex v of G corresponds to a node in the HTML tree, and the edge e of G represents the topological relationship between the vertices.

[0017] Preferably, for the picture graph, the web page content graph representation corresponding to the web page URL is constructed including:

[0018] Take each image in the webpage as a vertex v in the image graph, and set all vertices as isolated vertices;

[0019] Each vertex is processed in turn according to the processing order set by the position relationship between the vertices. The specific processing includes:

[0020] Calculate the geometric center of each vertex and obtain its center position;

[0021] Select other vertices in turn to connect each of the vertices. If the line does not pass through any part of any image, an edge e is established between the two image vertices of the line. Otherwise, no edge is established between the two image vertices of the line.

[0022] Preferably, after expressing the webpage content stored in the mhtml document as an HTML tree and before constructing the graph G, the method further comprises:

[0023] Traverse each node in the HTML tree, if node n is empty and its child node is empty, delete node n; if node n is empty and node n has only one child node, replace node n with the child node.

[0024] Preferably, before constructing the web page content graph representation corresponding to the web page URL, the method further includes: rendering the web page content in the mhtml document in a browser, extracting coordinate information of each text node, and constructing content and attribute information of each text node in the HTML tree.

[0025] Preferably, constructing a graph G=(V, E) using the nodes in the HTML tree and the relationships between the nodes includes:

[0026] Correspond the node whose content in the HTML tree is not empty to a vertex v in the graph G, and assign the node content and attributes to the corresponding vertex;

[0027] According to the node with empty content and its child nodes in the HTML tree, a vertex v corresponding to the node with empty content is generated in the graph G;

[0028] The vertex association relationship in the graph G is constructed according to the node position relationship and the node hierarchy relationship in the HTML tree.

[0029] Preferably, generating a vertex v corresponding to the node whose content is empty in the graph G comprises:

[0030] For nodes with empty content in the HTML tree, when the content of their child nodes is not empty, the corresponding node and its child nodes are divided into a node group, and a new vertex is added to the graph G. The content of the new vertex is the set of node contents in the node group, and the position attribute is the intersection of the position sets of each node in the node group; the content and attributes of the new vertex are assigned to the node with empty content, and the node is regarded as a node with non-empty content;

[0031] For a node with empty content in the HTML tree, when its child nodes include a child node with empty content, a vertex corresponding to the child node with empty content is generated in the graph G, so that the child node is regarded as a child node with non-empty content, until all child nodes are regarded as having non-empty content.

[0032] Preferably, the construction of the vertex association relationship in the graph G includes:

[0033] The nodes with non-empty content in the HTML tree and their parent nodes establish edges between the two vertices in the graph G;

[0034] The nodes with non-empty content in the HTML tree and their corresponding child nodes establish edges between two vertices in the graph G;

[0035] All first-level nodes with non-empty content in the HTML tree establish fully connected relationships between the vertices in the graph G.

[0036] Preferably, before determining the feature vector of the text graph through convolution and pooling operations, the method further comprises:

[0037] Taking the vertices in the text graph representation as units, a word segmentation tool is used to segment the text content inside each vertex, and the word segmentation result is converted into a fixed-size word vector representation.

[0038] Preferably, the convolution operation includes:

[0039] In a text graph or an image graph, use each vertex x i and the first-order neighbor vertex x of the corresponding vertex j Generate the latest representation of the corresponding vertex Among them, x i is the vector representation of the vertex before convolution processing, x i ′ is the vector representation of the vertex after convolution processing, Θ is the linear layer of the pre-trained neural network module, α i,i and α i,j is the self-attention coefficient of the vertex, which is determined according to the content attention coefficient of the vertex and the position attention coefficient of the vertex, i is the vertex index, j is the first-order neighbor vertex index, A set of indices of the first-order neighbor vertices of the vertex with index i.

[0040] Preferably,

[0041]

[0042] in,

[0043]

[0044] β i,i and β i,j is the content attention coefficient of the vertex, γ i,i and, γ i,j is the position attention coefficient of the vertex, δ is the preset weight, r, b and Θ are the linear layers of the graph neural network determined by pre-training, and LeakyReLU is the nonlinear activation function.

[0045] Preferably, the pooling operation includes: using a projection function to project each vertex vector obtained by the convolution operation, selecting the first ε vertices whose vertex vectors have the largest projection values ​​among the vertices, applying maximum pooling and mean pooling operations to the vectors of the selected vertices, and splicing the result vectors of the maximum pooling and mean pooling to obtain a layer of pooling operation results.

[0046] Preferably, the convolution and pooling operations are performed serially N times. After each convolution operation, the result of this convolution operation is used as the input of the next convolution operation, and the pooling operation is performed on the result of this convolution operation; the pooling operation results of N layers after N pooling operations are summed to obtain the feature vector of the text image or picture image; wherein N is a preset positive integer.

[0047] Preferably, when performing graph classification and recognition, after the convolution and pooling operations, the feature vector represented by the web page content graph is used as the input of the MLP to perform graph classification and recognition.

[0048] A network content classification device, comprising: a capture unit, a crawler engine unit and a classification unit;

[0049] The crawling unit is used to obtain the URL of the web page to be classified from the network and write it into the target URL document;

[0050] The crawler engine unit is used to crawl the content of each URL web page in the target URL document in the network, and save the content of each web page as an mhtml document, and write it into the Internet content archive database; according to the saved mhtml document, a web page content graph representation corresponding to the web page URL is constructed; wherein the web page graph representation includes a text graph and an image graph, the vertices in the text graph are text vertices, and the vertices in the image graph are image vertices;

[0051] The classification unit is used to use a graph neural network to classify and identify the web page content graph representation, and use the classification and identification results as the classification and identification results of the web page URL corresponding to the web page content graph representation; wherein, when performing graph classification and identification, the feature vector of the text graph and the feature vector of the picture graph are determined through convolution and pooling operations, and the feature vector of the text graph and the feature vector of the picture graph are spliced ​​as the feature vector of the web page content graph representation.

[0052] As can be seen from the above technical solution, in this application, the web page URL to be classified is obtained from the network and written into the target URL document; the content of each URL web page in the target URL document is crawled in the network by a crawler engine, and the content of each web page is saved as an mhtml document and written into the Internet content archive database; wherein, the web page graph representation includes a text graph and an image graph, the vertices in the text graph are text vertices, and the vertices in the image graph are image vertices; the graph neural network is used to classify and identify the web page content graph representation, and the classification and identification results are used as the classification and identification results of the web page URL corresponding to the web page content graph representation. Through the above processing, the text and image content of the network can be effectively classified and identified. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 A schematic diagram of the basic process of the network content classification method in this application;

[0054] Figure 2 This is a processing diagram of the crawler engine in this application;

[0055] Figure 3 It is a schematic diagram of the specific model of the WEB-GNN method;

[0056] Figure 4 This is a basic structural diagram of the network content classification device in this application. DETAILED DESCRIPTION

[0057] In order to make the objectives, technical means and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings.

[0058] The present application provides an Internet content mining and classification method based on a graph neural network algorithm, constructs a W2G preprocessing algorithm for Internet web pages, reconstructs web page content into a graph representation, and realizes an efficient integration of web page content and structure; at the same time, the use of a classification algorithm based on a graph neural network can accurately identify Internet content, and has a good effect on the identification and classification of Internet black market content.

[0059] Figure 1 This is a basic flow chart of the network content classification method in this application. Figure 1 As shown, the method includes:

[0060] Step 101, obtain the URL of the web page to be classified from the network and write it into the target URL document.

[0061] In this step, a large number of web page URLs are collected from the Internet, and after filtering using a URL whitelist, the remaining URLs in the collected web pages excluding those in the filtered whitelist are written into the target URL document.

[0062] Step 102, crawling the content of each URL web page in the target URL document in the network through a crawler engine, and saving the content of each web page as an mhtml document, and writing it into an Internet content archive database.

[0063] In this step, a crawler engine is used to crawl web page content. Preferably, the crawler engine in this embodiment combines selenium crawler technology with container virtualization technology. The specific architecture is as follows: Figure 2 As shown. The target URLs are uniformly put into the Redis task queue, and then multiple crawler processes are used to read the webpage URLs from the Redis task queue respectively. Preferably, a crawler container cluster can be used to read the URLs from the Redis task queue respectively, and then operations such as content loading and rendering are performed, and finally the webpage content is saved in the Internet content archive database in the mhtml file format. Among them, in the Internet content archive database, the mhtml archive of a certain webpage can contain at least one of the following attributes: webpage URL, webpage host, webpage domain, webpage server IP address, webpage Title field content, etc.

[0064] The crawler engine given above has high compatibility with many web page contents. At the same time, the result of the crawler engine is the complete web page content at a specific time point, which can realize offline viewing of web pages and other operations. It provides technical support for manual verification of content classification results and has the characteristics of high availability, easy expansion, and lightweight.

[0065] Step 103, constructing a web page content graph representation corresponding to the web page URL according to the saved mhtml document.

[0066] The web page content graph includes a text graph and an image graph. The vertices of the text graph are the text content in the web page, and the vertices of the image graph are the images in the web page. This can distinguish between text and images, and classify and identify web page content from different perspectives of text and images.

[0067] For the text content in the web page, in this step, in order to obtain the text graph representation of the web page content, it is first necessary to represent the web page text content stored in the mhtml document as an HTML tree, and then use the nodes in the HTML tree and the relationships between the nodes to construct a graph G = (V, E).

[0068] Specifically, a web page is a specific information collection provided by a visited website and displayed in a web browser. A web page usually consists of many content elements. Its core element is one or more text collections. This application adopts the W2G (Webpage to Graph) web page preprocessing algorithm. Assuming that you want to parse the text content of a web page into a content graph G = (V, E), then the goal is to extract a set of content elements N = {n_{1}, n_{2}, ... n_{p}} from the HTML tree H of the 4eb page, each content element is a node n in the H tree, and use these content elements to construct a graph G = (V, E), where V is the set of all vertices v in the graph G, and E is the set of all edges e in the graph G. Each vertex v of G corresponds to an element of H. The edge e of G reflects the topological relationship between vertices. Here, the association relationship between the vertices of the graph G can be constructed by referring to the HTML hierarchy of each node in the H tree and the positional relationship presented.

[0069] In addition, due to the particularity of web page content, some text information may be included in the pictures on the web page. In order to enrich the text content in the web page, the present application preferably can extract and identify the text embedded in the pictures containing text information on the web page (for example, through OCR text recognition method), generate text nodes and add them to the HTML tree.

[0070] Next, we will introduce the process of building an HTML tree. HTML is called Hypertext Markup Language, which is an identification language. It includes a series of tags. Through these tags, the document format on the network can be unified, so that scattered Internet resources can be connected into a logical whole. HTML text is a descriptive text composed of HTML commands. HTML commands can describe text, graphics, animations, sounds, tables, links, etc. Each element can specify HTML attributes. Here, we focus on the BoundingClientRect of the element, that is, the location information of the browser window where the element is located, as well as the text display information and size. This is very important in the web page graph model.

[0071] Since there are some meaningless or empty elements in H, which are specifically manifested as elements that only contain attribute configuration information, the existence of these nodes increases the complexity of the graph. Therefore, preferably, H can be pruned to reduce the complexity of the tree, reduce the time for data preprocessing, and increase the operating efficiency of the algorithm. The following pruning strategy can be defined to traverse each node in the H tree and make the following judgments:

[0072] 1. If node n is empty and its child nodes are also empty, delete n.

[0073] 2. If node n is empty and has only one child node, replace n with its child node.

[0074] Finally, the construction process of text graph representation is described.

[0075] First, let's introduce the representation of text graphs. The text graph is represented by G = (V, E), which is constructed based on the HTML tree H obtained above. Among them, the node n in H corresponds to the vertex v in the graph G, so generally, the node n i and vertex v i There is a one-to-one mapping relationship between them. The properties of the vertex v and edge e of the graph G can be defined as shown in Table 1. Among them, v1 and v2 represent the two vertices connected by the edge e in the graph.

[0076]

[0077] Table 1

[0078] As shown in Table 1, the attributes of vertex v in the figure include vertex content and vertex coordinates; wherein the vertex content includes: the node text content of node n corresponding to vertex v; the vertex coordinates include at least one of the following: the vertical coordinate of node n corresponding to vertex v, the height of the element where vertex v corresponds to node n, and the width of the element where vertex v corresponds to node n;

[0079] The attributes of the edge e include at least one of the following: the horizontal distance between two vertices connected by edge e, the vertical distance between two vertices connected by edge e, the aspect ratio of a vertex connected by edge e, the height ratio of two vertices connected by edge e, and the width ratio of two vertices connected by edge e.

[0080] Next, the process of constructing the text graph G based on the HTML tree is described in detail.

[0081] In order to further reduce the complexity of graph G and enrich the relationship between nodes, it is preferred that the mhtml web page can be rendered in the browser to extract the coordinate information of each text node. In this processing method, the coordinate information of the nodes and vertices is the coordinate value of each node on the rendered web page. In graph G, for each node with non-empty content in the pruned HTML tree, the vertex of graph G is set, and the node attribute value is assigned to the corresponding vertex attribute.

[0082] As mentioned above, the construction of the relationship between vertices in the webpage graph G comprehensively considers the positional relationship of the nodes and the hierarchical relationship of the nodes. Preferably, for text nodes, after pruning, all nodes with non-empty content correspond to a vertex v in the graph G, and nodes with empty content must have child nodes with non-empty content, otherwise the node has been pruned. For these nodes with empty content after pruning, the following processing is performed to generate the corresponding vertices in the graph G:

[0083] When all child nodes of node A with empty content have non-empty content, the child nodes belonging to node A are divided into a node group, and all node groups constitute a node group set P. For each node group p∈P,

[0084] p={v′∈P|(v′ x =n||v′ h =h)|n,h∈N} (1)

[0085] Add a new vertex v' to graph G, corresponding to the group set element p. The content of the new vertex v' is the set of node contents in the group, and the position attribute is the intersection of the node position set in the group. In this way, a vertex corresponding to the empty node is generated in graph G, and the attributes of the vertex replace the attributes of node A. In this way, node A can be treated as a node with non-empty content. At the same time, as mentioned above, each node with non-empty content corresponds to a vertex in graph G, that is, all nodes of the tree after pruning correspond to a vertex in graph G;

[0086] When the child nodes of node A with empty content include child nodes with empty content, assuming that any child node with empty content is B, a vertex corresponding to child node B can be generated in the text graph G in the above manner, and then the attributes of the vertex replace the content and attributes of child node B, so that each child node B can be regarded as a node with non-empty content. In this way, node A can be regarded as a node whose child nodes are all non-empty content for processing.

[0087] At the same time, for the newly added vertex v' in the node group p, the vertices corresponding to the nodes in p in the graph G are connected to the newly added vertex v' respectively. That is to say, in the group p, the newly added vertex v' is the parent node, and the vertices corresponding to each node in the node group in the graph G are the child nodes. The nodes in the group are not related to each other, that is, there are no edges between the vertices corresponding to these nodes.

[0088] The above is the processing for nodes with empty content and their child nodes. For node C with non-empty content, the corresponding vertex in graph G is connected to the vertex of the parent node of node C in graph G, and the corresponding vertex in graph G is connected to the child node of node C. At the same time, all the first-level nodes with non-empty content are fully connected to all the vertices in graph G. Here, the nodes with non-empty content include the aforementioned nodes (such as the aforementioned node A) that replace the original node content and attributes by adding a vertex in graph G and returning it.

[0089] The above is the specific generation process of the text graph.

[0090] Next, we will introduce the process of generating an image map for images in a web page. Similar to the early processing of a text map, the mhtml web page can be rendered in a browser to extract the coordinate information of each image. Each image is regarded as a vertex v in the image map, and all vertices are set as isolated vertices. A processing order is set according to the positional relationship between the vertices, and each vertex is processed in turn according to the set processing order. Taking vertex X as an example, the specific processing may include:

[0091] Calculate the geometric center of vertex X and obtain its center position;

[0092] Select any vertex Y except vertex X and connect it to vertex X (that is, connect the geometric center positions of the two vertices). If the line does not pass through any part of any image, then establish an edge e between the two image vertices X and Y of the line. Otherwise, no edge is established between the two image vertices X and Y of the corresponding line.

[0093] When determining the processing order, the picture nodes may be selected from top to bottom and from left to right according to the positional relationship of the picture nodes.

[0094] For the picture graph, the properties of the vertices and edges in the picture graph established in the above manner are shown in Table 2. Among them, v1 and v2 represent two vertices connected by the edge e in the graph.

[0095]

[0096] Table 2

[0097] Step 104, using the constructed web page content graph representation as input, using a graph neural network to classify and recognize the graph, and using the classification and recognition results as the classification and recognition results of the web page URL corresponding to the web page content graph representation.

[0098] This application proposes an end-to-end web page classification method based on graph neural network WEB-GNN. The WEB-GNN method adopts an end-to-end modular architecture. In this process, the text graph in the web page content graph representation is convolved and pooled to obtain the feature vector of the text graph. The image graph is processed in the same way to obtain the feature vector of the image graph. The feature vector of the text graph and the feature vector of the image graph are then concatenated as the feature vector of the web page content graph representation. The MLP module is used to classify and recognize the feature vector of the web page content graph representation.

[0099] Preferably, the convolution and pooling operations can be performed multiple times, and after each convolution operation, the operation result is used as the input of the next convolution operation; at the same time, after each convolution operation, the pooling operation is performed using the result of the convolution operation.

[0100] Before performing convolution and pooling operations, for ease of processing, both the text map and the image map are first converted into vector representations.

[0101] Wherein preferably, for the text graph, the vertices in the graph can be used as units, and the text content inside each vertex can be segmented using a segmentation tool (such as jieba), and the segmentation result is converted into a fixed-size word vector representation. The above-mentioned processing draws on the relevant methods of natural language processing in deep learning, and uses word vector technology to pre-process the text. Specifically, the node v of the graph obtained in step 103 can be used as a unit, and the text inside v can be segmented using the jieba Chinese word segmentation tool. Subsequently, an open source word vector dictionary can be used to convert the segmentation result into a fixed-size word vector representation, and the content representation of v is the average value of the internal word vector. Finally, vertex v can be represented as a word vector of fixed dimension (such as 200 dimensions).

[0102] For the vector representation of each image vertex in the image graph, a pre-trained image model can be used to generate a vector representation of the image vertex. Specifically, a large-scale image dataset is used for training to obtain a deep learning model (for example, a resnet series model can be obtained based on the ImageNet image training set), so that when the input is a certain image, a fixed-dimensional image vector representation can be output. The dimension of the image vertex is consistent with the text vector, for example, both are 200 dimensions.

[0103] The following is a detailed description of the convolution operation. The graph convolution layer integrates the information of the vertex and its first-order neighbor vertices, and generates the latest representation of the vertex by fusing the local structural information. The vertex information propagation formula is:

[0104]

[0105] Among them, x′ i is the new representation of vertex v with index i after convolution. is a set of indices of the first-order neighbor vertices of the vertex with index i. Θ is a linear layer of a neural network module (e.g., an MLP linear layer), which can be obtained in advance through training based on the original training data. Its function is to transform and map the 200-dimensional vector of the vertex content into a low-order vector representation. Latest α i,i and α i,j is the self-attention coefficient of the vertex vector, which is the content attention coefficient of the vertex β i,j and the vertex position attention coefficient γ i,j The weighted sum of , the weight of the two is determined by the hyperparameter δ, which is defined as follows:

[0106] α i,i =δ*β i,i +(1-δ)*γ i,i

[0107] α i,j =δ*β i,j +(1-δ)*γ i,j (3)

[0108] β i,i and β i,j is the content attention coefficient of the vertex. || is the concatenation operation of the vector. b and Θ are linear layers, b is the same as the aforementioned Θ, and can be obtained in advance through training based on the original training data. LeakyReLU is a nonlinear activation function:

[0109]

[0110]

[0111] γ i,i and γ i,j is the position attention coefficient of the vertex. r and Θ are linear layers, r is the same as the aforementioned Θ, and can be obtained in advance through training based on the original training data. Intuitively, we hope that the convolution processing can capture the importance of nodes shown by the distribution of nodes in the rendering results, and give higher attention to important nodes. For example, for most web pages, the content title is usually located at the top of the web page, and the display effect is more prominent than other nodes. Therefore, the contribution of this node in web page recognition and classification is greater. The definition of γ is as follows:

[0112]

[0113]

[0114] Among them, p i is the position vector of the node with index i. The above is the processing of a convolution operation. The processing of the pooling operation is introduced below. Specifically, the projection function can be used to project each vertex vector obtained by the convolution operation, and the first ε vertices with the largest projection value of the vertex vector are selected from the vertices of the graph G, and the vectors of each selected vertex are used as the pooling operation results. Among them, ε<1.

[0115] In more detail, the main function of the pooling layer is to further reduce the complexity of the graph. Here, we introduce a learnable projection function, whose input is the vertex vector v and output is the projection value q. Here, the projection function is a linear function. According to the projection value of each vector, the vertices are sorted in descending order, and the vertices before ε are selected as the input of the next convolutional layer, and the vertices after (1-ε) are discarded.

[0116] After the above pooling operation, it is usually necessary to process the output vertex vector to obtain the graph representation vector. Based on this, an output layer can be added after each pooling layer. The output layer applies maximum pooling and mean pooling operations to the vertex vectors output by the pooling layer, and then concatenates the two vectors obtained after maximum pooling and mean pooling, and uses the concatenated result as the graph representation vector of this layer; for the convenience of description, the pooling operation and the processing of the output layer are collectively referred to as the pooling operation of one layer. When multiple convolution and pooling operations are performed, after the last layer of pooling operation is completed, the graph representation vectors output by each layer are added to obtain the feature vector of the graph.

[0117] As mentioned above, we obtain the feature vector representation of the web page text graph through convolution and pooling operations on the text graph, and obtain the feature vector representation of the web page image graph through convolution and pooling operations on the image graph. We concatenate the two to generate the feature vector representation of the web page image. The schematic diagram of the general model of the above convolution and pooling operations is as follows: Figure 3 As shown. Subsequently, the feature vector representation of the graph is used as the input of the MLP to obtain the final classification result.

[0118] In order to achieve better results for the above WEB-GNN method, the parameter values ​​of WEB-GNN can be obtained by pre-training. A specific model training method is given below. For example, the input of this model is the web page graph representation described above, and its details are shown in Table 2. The network parameter settings of this model are shown in Table 3, and the classification label settings of this model output are shown in Table 4.

[0119] Logo Dimensions describe x [num_nodes,num_node_features] Graph vertex vector matrix edge_index [2,num_edges] Graph Edge Index edge_attr [num_edges,num_edge_features] Graph edge attribute matrix y [1] Graph Category POS [num_nodes,num_node_features] Graph vertex position vector matrix

[0120] Table 2

[0121] name block input output conv1 GATConv 200 128 pool1 TopKPooling 128 0.8 conv2 GATConv 128 128 pool2 TopKPooling 128 0.8 conv3 GATConv 128 128 pool3 TopKPooling 128 0.8 lin1 Linear 512 256 lin2 Linear 256 64 lin3 Linear 64 6

[0122] Table 3

[0123] Label describe 0 Normal web page 1 Types of Gambling Scams 2 Types of Financial Fraud 3 Porn Types 4 Types of Phishing Scams 5 Other types of scams 6 Types of fake order fraud

[0124] Table 4

[0125] The Adam optimization algorithm can be used for training. Combined with the meaning of the initial value of the parameter, different values ​​are set for each parameter, and 10-fold cross validation and grid search are used to continuously fit the data, train the model, and output a stable training result model. In this example, the precision rate, recall rate, and F1-score are used to evaluate the model. The calculation formulas are shown in formula (5), formula (6), and formula (7) respectively.

[0126] Precision = TP / (TP+FP) (5)

[0127] Recall = TP / (TP+FN) (6)

[0128] F1-score=2*Precision*Recall / (Precision+Recall) (7)

[0129] Among them, TP represents the number of samples that are positive and the predicted results are positive, FP represents the number of samples that are negative and the predicted results are positive, TN represents the number of samples that are negative and the predicted results are negative, and FN represents the number of samples that are positive and the predicted results are negative.

[0130] The above method can be used to determine the high-quality model of WEB-GNN, thereby realizing web page classification based on graph representation.

[0131] After the comprehensive evaluation of the WEB-GNN model, the model persistence module of torch is used to serialize the model and save it to the server. Flask is used to build an API, and the model that meets business needs is deployed online in the form of an API interface to realize the mining and classification of Internet content. After obtaining the target web page, the Internet content crawler engine calls the API interface and inputs the data into the WEB-GNN model to realize the classification of web page content and output the predicted label corresponding to the target web page.

[0132] At this point, the basic process of the webpage content classification method in this application is completed.

[0133] The above is the specific implementation of the web page content classification method in this application. This application also provides a network content classification device that can be used to implement the above classification method. Figure 4 The basic structural diagram of the classification device is shown in Figure 2. Figure 4 As shown, the device includes: a grasping unit, a crawler engine unit and a classification unit.

[0134] A crawling unit is used to obtain the web page URL to be classified from the network and write it into the target URL document. A crawler engine unit is used to crawl the content of each URL web page in the target URL document on the network, and save the content of each web page as an mhtml document and write it into the Internet content archive database; based on the saved mhtml document, a web page content graph representation corresponding to the web page URL is constructed; wherein the web page graph representation includes a text graph and an image graph, the vertices in the text graph are text vertices, and the vertices in the image graph are image vertices. A classification unit is used to classify and recognize the web page content graph representation using a graph neural network, and use the classification and recognition results as the classification and recognition results of the web page URL corresponding to the web page content graph representation; wherein, when performing graph classification and recognition, the feature vectors of the text graph and the image graph are determined respectively by convolution and pooling operations, and the feature vectors of the text graph and the feature vectors of the image graph are concatenated as the feature vectors of the web page content graph representation.

[0135] As described above, the present application provides a method and device for Internet content mining and classification based on graph neural networks, and with this method as the core, a long-term Internet content mining and analysis platform is constructed to help relevant departments supervise Internet content and effectively combat Internet black content. This method parses the HTML source code of the web page and reconstructs it into a graph representation, thereby achieving efficient integration of web page structure and text content. Subsequently, the black web page is modeled based on the improved graph neural network method, and finally the mining, identification and analysis of black content are achieved. The platform operation results show that this method can significantly improve the efficiency of relevant departments in supervising Internet content, purify the network environment, and promote the healthy development of the Internet.

[0136] The applicant tested the classification method of this application by collecting Internet URL transmission records from the public network and filtering the target URL to be classified through a blacklist and whitelist filtering mechanism. Then, the resource content corresponding to the URL was crawled through an Internet content crawler engine, and an mhtml file was generated as the content archive of the URL. Finally, the WEB-GNN web page classification method proposed in this application was used to predict the mhtml file and obtain the Internet black industry content URL.

[0137] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for classifying network content. It is characterized in that include: Get the URL of the web page to be classified from the network and write it into the target URL document; The crawler engine crawls the content of each URL web page in the target URL document on the network, and saves the content of each web page as an mhtml document and writes it into the Internet content archive database; According to the saved mhtml document, a web page content graph representation corresponding to the web page URL is constructed; wherein the web page content graph representation includes a text graph and an image graph, the vertices in the text graph are text vertices, and the vertices in the image graph are image vertices; Using a graph neural network to classify and identify the web page content graph representation, and using the classification and identification results as the classification and identification results of the web page URL corresponding to the web page content graph representation; wherein, when performing the graph classification and identification, the feature vector of the text graph and the feature vector of the picture graph are determined by convolution and pooling operations, and the feature vector of the text graph and the feature vector of the picture graph are concatenated as the feature vector of the web page content graph representation; Wherein, for the text graph, the web page content graph representation corresponding to the web page URL is constructed including: The webpage text content stored in the mhtml document is represented as an HTML tree; wherein the text content elements in the webpage content are configured as nodes of the HTML tree, and for pictures including text information in the webpage, the text embedded in the pictures is extracted and identified, and a text node is generated and added to the HTML tree; Constructing a graph G=(V, E) using the nodes in the HTML tree and the relationships between the nodes; wherein each vertex v in G corresponds to a node in the HTML tree, and the edge e of G represents the topological relationship between the vertices; For the picture graph, the web page content graph representation corresponding to the web page URL is constructed including: Take each image in the webpage as a vertex v in the image graph, and set all vertices as isolated vertices; Each vertex is processed in turn according to the processing order set by the position relationship between the vertices. The specific processing includes: Calculate the geometric center of each vertex and obtain its center position; Select other vertices in turn to connect each of the vertices. If the line does not pass through any part of any image, an edge e is established between the two image vertices of the line. Otherwise, no edge is established between the two image vertices of the line.

2. The method according to claim 1, It is characterized in that For a text graph, after expressing the webpage content stored in the mhtml document as an HTML tree and before constructing the graph G, the method further comprises: Each node in the HTML tree is traversed. If node n is empty and its child node is empty, node n is deleted; if node n is empty and node n has only one child node, node n is replaced with the child node.

3. The method according to claim 2, It is characterized in that Before constructing the web page content graph representation corresponding to the web page URL, the method further includes: rendering the web page content in the mhtml document in a browser, extracting coordinate information of each text node, and constructing content and attribute information of each text node in an HTML tree.

4. The method according to claim 2, It is characterized in that The constructing graph G=(V, E) by using the nodes in the HTML tree and the relationships between the nodes includes: Correspond the node whose content in the HTML tree is not empty to a vertex v in the graph G, and assign the node content and attributes to the corresponding vertex; According to the node with empty content and its child nodes in the HTML tree, a vertex v corresponding to the node with empty content is generated in the graph G; The vertex association relationship in the graph G is constructed according to the node position relationship and the node hierarchy relationship in the HTML tree.

5. The method according to claim 4, It is characterized in that The step of generating a vertex v corresponding to the node whose content is empty in the graph G comprises: For nodes with empty content in the HTML tree, when the content of their child nodes is not empty, the corresponding node and its child nodes are divided into a node group, and a new vertex is added to the graph G. The content of the new vertex is the set of node contents in the node group, and the position attribute is the intersection of the position sets of each node in the node group; the content and attributes of the new vertex are assigned to the node with empty content, and the node is regarded as a node with non-empty content; For a node with empty content in the HTML tree, when its child nodes include a child node with empty content, a vertex corresponding to the child node with empty content is generated in the graph G, so that the child node is regarded as a child node with non-empty content, until all child nodes are regarded as having non-empty content.

6. The method according to claim 4, It is characterized in that The construction of the vertex association relationship in the graph G includes: The nodes with non-empty content in the HTML tree and their parent nodes establish edges between the two vertices in the graph G; The nodes with non-empty content in the HTML tree and their corresponding child nodes establish edges between two vertices in the graph G; All first-level nodes with non-empty content in the HTML tree establish fully connected relationships between the vertices in the graph G.

7. The method according to claim 1, It is characterized in that The convolution operation includes: In a text graph or an image graph, use each vertex x i and the first-order neighbor vertex x of the corresponding vertex j Generate the latest representation of the corresponding vertex α i,i =δ*β i,i +(1-δ)*γ i,i , α i,j =δ*β i,j +(1-δ)*γ i,j , Among them, x i is the vector representation of the vertex before convolution processing, x i ′ is the vector representation of the vertex after convolution processing, Θ is the linear layer of the pre-trained neural network module, α i,i and α i,j is the self-attention coefficient of the vertex, which is determined according to the content attention coefficient of the vertex and the position attention coefficient of the vertex, i is the vertex index, j is the first-order neighbor vertex index, is a set of indices of the first-order neighbor vertices of the vertex with index i, p i 、p j 、p k are the position vectors of nodes indexed i, j, and k respectively, β i,i and β i,j is the content attention coefficient of the vertex, γ i,i and, γ i,j is the position attention coefficient of the vertex, δ is the preset weight, r, b and Θ are the linear layers of the graph neural network determined by pre-training, and LeakyReLU is the nonlinear activation function; and / or, The pooling operation includes: using a projection function to project each vertex vector obtained by the convolution operation, selecting the vertex with the largest projection value of the vertex vector before ε among the vertices, applying maximum pooling and mean pooling operations to the vectors of each selected vertex, and splicing the result vectors of the maximum pooling and mean pooling to obtain a layer of pooling operation results; and / or, The convolution and pooling operations are performed serially N times. After each convolution operation, the result of the current convolution operation is used as the input of the next convolution operation, and the current pooling operation is performed on the result of the current convolution operation; the pooling operation results of the N layers after the N pooling operations are summed to obtain the feature vector of the text image or picture image; wherein N is a preset positive integer.

8. A network content classification device, It is characterized in that include: Crawling unit, crawler engine unit and classification unit; The crawling unit is used to obtain the URL of the web page to be classified from the network and write it into the target URL document; The crawler engine unit is used to crawl the content of each URL web page in the target URL document on the network, and save the content of each web page as an mhtml document, and write it into the Internet content archive database; according to the saved mhtml document, construct a web page content graph representation corresponding to the web page URL; wherein the web page content graph representation includes a text graph and an image graph, the vertices in the text graph are text vertices, and the vertices in the image graph are image vertices; The classification unit is used to classify and identify the web page content graph representation using a graph neural network, and use the classification and identification results as the classification and identification results of the web page URL corresponding to the web page content graph representation; wherein, when classifying and identifying the graph, the feature vector of the text graph and the feature vector of the picture graph are determined by convolution and pooling operations, and the feature vector of the text graph and the feature vector of the picture graph are concatenated as the feature vector of the web page content graph representation; Wherein, in the crawler engine unit, for the text graph, the web page content graph representation corresponding to the web page URL is constructed including: The webpage text content stored in the mhtml document is represented as an HTML tree; wherein the text content elements in the webpage content are configured as nodes of the HTML tree, and for pictures including text information in the webpage, the text embedded in the pictures is extracted and identified, and a text node is generated and added to the HTML tree; Constructing a graph G=(V, E) using the nodes in the HTML tree and the relationships between the nodes; wherein each vertex v in G corresponds to a node in the HTML tree, and the edge e of G represents the topological relationship between the vertices; For the picture graph, the web page content graph representation corresponding to the web page URL is constructed including: Take each image in the webpage as a vertex v in the image graph, and set all vertices as isolated vertices; Each vertex is processed in turn according to the processing order set by the position relationship between the vertices. The specific processing includes: Calculate the geometric center of each vertex and obtain its center position; Select other vertices in turn to connect each of the vertices. If the line does not pass through any part of any image, an edge e is established between the two image vertices of the line. Otherwise, no edge is established between the two image vertices of the line.

Citation Information

Patent Citations

  • Webpage-oriented unhealthy Web content identifying method

    CN102332028A

  • Content distribution method and device and storage medium

    CN111723295A