Webpage clustering method and device and electronic equipment

By extracting core text elements in the web page and using semantic extraction models based on pre-trained language models, combined with fuzzy clustering algorithm, the problem of low accuracy of web page clustering in the existing technology is solved, and a more efficient web page clustering effect is achieved.

CN120216803APending Publication Date: 2025-06-27CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510379505.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, the accuracy of clustering web pages is low, especially when processing high-dimensional and sparse data, it is susceptible to noise and outliers, and it is not able to effectively extract core elements and complex semantic relationships in web pages.

Method used

A web page clustering method is adopted to obtain multiple target web pages and extract their important target texts, and use the target semantics extraction model based on the pre-trained language model to extract the semantic information of the text, and then input the semantic information into the fuzzy clustering model for clustering to obtain more accurate clustering results.

Benefits of technology

By extracting core text elements in web pages and considering their semantic information, the accuracy of web page clustering is significantly improved, and the complex semantic relationships between data can be captured more effectively.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216803A_ABST
    Figure CN120216803A_ABST
Patent Text Reader

Abstract

The invention discloses a webpage clustering method and device and electronic equipment, and relates to the technical field of webpage processing. The method comprises the following steps: acquiring a plurality of target webpages, and extracting a target text in each target webpage to obtain a plurality of target texts; inputting each target text into a target semantic extraction model for processing to obtain semantic information of each target text; and inputting the semantic information of each target text into a target clustering model for processing to obtain a target clustering result of clustering the plurality of target webpages. Through the webpage clustering method and device, the problem that in the related technology, the accuracy of webpage clustering is low is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of web page processing, and in particular, to a web page clustering method, apparatus, and electronic device. Background Art

[0002] Massive information is spread on the network in various forms. How to quickly and accurately obtain useful data from this vast amount of information has become an important task. Moreover, web page clustering, as an important information processing technology, is widely used in fields such as asset mapping and search engine optimization.

[0003] In addition, clustering algorithms in related technologies, such as K-means, hierarchical clustering, DBSCAN, etc., although to a certain extent can classify and cluster web page data, often face the problem of insufficient accuracy in practical applications. This is mainly because these algorithms in related technologies are easily affected by noise and outliers when processing high-dimensional and sparse data, and fail to extract the core elements in web pages. In addition, clustering algorithms in related technologies generally rely on predefined similarity measurement methods and are difficult to fully capture the complex semantic relationships between data. Therefore, the clustering algorithms in related technologies result in relatively low clustering accuracy.

[0004] Regarding the problem of relatively low accuracy in clustering web pages in related technologies, no effective solution has been proposed yet. Summary of the Invention

[0005] The present invention provides a web page clustering method, apparatus, and electronic device to solve the problem of relatively low accuracy in clustering web pages in the prior art.

[0006] In a first aspect, the present application provides a web page clustering method, which includes: Obtain a plurality of target web pages, and extract target texts in each target web page to obtain a plurality of target texts, where the importance level of each target text is greater than a preset importance level; Input each target text into a target semantic extraction model for processing to obtain semantic information of each target text, where the target semantic extraction model is a model constructed based on a pre-trained language model; Input the semantic information of each target text into a target clustering model for processing to obtain a target clustering result for clustering the plurality of target web pages, where the target clustering model is constructed using a fuzzy clustering algorithm.

[0007] In a possible implementation manner, the extracting target texts in each target web page to obtain a plurality of target texts includes: Obtain the code information for constructing each target web page, and based on the code information for constructing each target web page, obtain multiple target documents, where each target document is used to construct each target web page; Using the Document Object Model, convert each target document into a tree structure to obtain multiple target trees; Obtain the target nodes in each target tree, and based on the target nodes in each target tree, extract the target text in each target web page to obtain the multiple target texts.

[0008] In one possible implementation, the obtaining the target nodes in each target tree includes: Obtain the tags of the target type in each target tree, and delete the tags of the target type from each target tree to obtain multiple target trees after deleting the tags, where the degree of correlation between the data content in the tags of the target type and the semantic information is less than a preset degree of correlation; Obtain multiple nodes from each target tree after deleting the tags, and delete the nodes of the first type and the nodes of the second type from the multiple nodes to obtain multiple target trees after deleting the nodes, where the nodes of the first type are nodes without child nodes, and the nodes of the second type are nodes with one child node; According to the similarity between the nodes in each target tree after deleting the nodes, perform grouping processing on the nodes in each target tree after deleting the nodes to obtain multiple processing results; Based on each processing result, obtain the target nodes in each target tree.

[0009] In one possible implementation, the inputting each target text into the target semantic extraction model for processing to obtain the semantic information of each target text includes: Input each target text into the target semantic extraction model for processing to obtain multiple feature vectors, where each feature vector is used to represent the semantic features of each target text; Perform dimensionality reduction processing on each feature vector to obtain multiple processed feature vectors; Based on each processed feature vector, obtain the semantic information of each target text.

[0010] In one possible implementation, the target semantic extraction model is obtained through the following method: Obtain multiple sample web pages, and determine whether there are a first web page and a second web page in the multiple sample web pages, where the first web page and the second web page are the same sample web page; If there are the first web page and the second web page in the multiple sample web pages, then delete the first web page or the second web page from the multiple sample web pages to obtain multiple sample web pages after deletion; Construct a training data set based on the multiple deleted sample web pages; Use the training data set to learn and train the original semantic extraction model to obtain the target semantic extraction model, where the type of the original semantic extraction model includes a pre-trained language model.

[0011] In a possible implementation, the target clustering model is obtained through the following steps: Construct an original clustering model using a fuzzy clustering algorithm, and use the original clustering model to cluster the sample web pages in the training data set to obtain a sample clustering result; Determine whether the accuracy rate of the sample clustering result is greater than a preset value; If the accuracy rate of the sample clustering result is greater than the preset value, use the original clustering model as the target clustering model; If the accuracy rate of the sample clustering result is not greater than the preset value, adjust the parameters of the original clustering model to obtain an adjusted clustering model; Based on the adjusted clustering model, obtain the target clustering model.

[0012] In a possible implementation, after inputting the semantic information of each target text into the target clustering model for processing to obtain a target clustering result for clustering the multiple target web pages, the method further includes: Determine the association relationship between the multiple target web pages according to the target clustering result, and obtain the dangerous web pages among the multiple target web pages, where the danger level of the dangerous web pages is higher than a preset danger level; According to the association relationship between the multiple target web pages, determine whether there are web pages associated with the dangerous web pages among the multiple target web pages; If there are no web pages associated with the dangerous web pages among the multiple target web pages, allow access to the web pages other than the dangerous web pages among the multiple target web pages; If there are web pages associated with the dangerous web pages among the multiple target web pages, prohibit access to the dangerous web pages and the web pages associated with the dangerous web pages among the multiple target web pages.

[0013] In a second aspect, the present application further provides a web page clustering device, including: A first acquisition unit, configured to acquire multiple target web pages, and extract target texts in each target web page to obtain multiple target texts, where the importance level of each target text is greater than a preset importance level; A first processing unit, configured to input each target text into a target semantic extraction model for processing to obtain semantic information of each target text, where the target semantic extraction model is a model constructed based on a pre-trained language model; A second processing unit, configured to input the semantic information of each target text into a target clustering model for processing to obtain a target clustering result for clustering the multiple target web pages, where the target clustering model is constructed using a fuzzy clustering algorithm.

[0014] In a possible implementation manner, the first acquisition unit includes: A first acquisition module, configured to acquire code information for constructing each target web page, and obtain multiple target documents based on the code information for constructing each target web page, where each target document is used to construct each target web page; A first conversion module, configured to use a document object model to convert each target document into a tree structure to obtain multiple target trees; A second acquisition module, configured to acquire target nodes in each target tree, and extract target texts in each target web page based on the target nodes in each target tree to obtain multiple target texts.

[0015] In a possible implementation manner, the second acquisition module includes: A first processing sub-module, configured to acquire tags of a target type in each target tree, and delete the tags of the target type from each target tree to obtain multiple target trees after deleting the tags, where the data content in the tags of the target type has a correlation degree less than a preset correlation degree with the semantic information; A second processing sub-module, configured to acquire multiple nodes from each target tree after deleting the tags, and delete nodes of a first type and nodes of a second type from the multiple nodes to obtain multiple target trees after deleting the nodes, where the nodes of the first type are nodes without child nodes, and the nodes of the second type are nodes with one child node; A third processing sub-module, configured to group the nodes in each target tree after deleting the nodes according to the similarity between the nodes in each target tree after deleting the nodes to obtain multiple processing results; a first acquisition sub-module, configured to acquire target nodes in each target tree based on each processing result.

[0016] In a possible implementation manner, the first processing unit includes: A first processing module, configured to input each target text into a target semantic extraction model for processing to obtain multiple feature vectors, where each feature vector is used to represent the semantic features of each target text; A second processing module, configured to perform dimensionality reduction processing on each feature vector to obtain a plurality of processed feature vectors; a third processing module, configured to obtain semantic information of each target text based on each processed feature vector.

[0017] In a possible implementation manner, a third processing unit is configured to obtain a plurality of sample web pages, and determine whether a first web page and a second web page exist in the plurality of sample web pages, where the first web page and the second web page are the same sample web page; A first deletion unit is configured to, if the first web page and the second web page exist in the plurality of sample web pages, delete the first web page or the second web page from the plurality of sample web pages to obtain a plurality of deleted sample web pages; A first construction unit is configured to construct a training data set according to the plurality of deleted sample web pages; A first training unit is configured to use the training data set to perform learning and training on an original semantic extraction model to obtain a target semantic extraction model, where the type of the original semantic extraction model includes a pre-trained language model.

[0018] In a possible implementation manner, the target clustering model is obtained through the following units: a fourth processing unit is configured to construct an original clustering model by using a fuzzy clustering algorithm, and perform clustering processing on the sample web pages in the training data set by using the original clustering model to obtain a sample clustering result; a first judgment unit is configured to judge whether the accuracy rate of the sample clustering result is greater than a preset value; a first determination unit is configured to, if the accuracy rate of the sample clustering result is greater than the preset value, use the original clustering model as the target clustering model; a first adjustment unit is configured to, if the accuracy rate of the sample clustering result is not greater than the preset value, adjust the parameters of the original clustering model to obtain an adjusted clustering model; a second determination unit is configured to obtain the target clustering model based on the adjusted clustering model.

[0019] In a possible implementation manner, the apparatus further includes: a fifth processing unit, configured to, after inputting the semantic information of each target text into the target clustering model for processing to obtain a target clustering result for clustering a plurality of target web pages, determine an association relationship between the plurality of target web pages according to the target clustering result, and obtain a dangerous web page among the plurality of target web pages, where the danger level of the dangerous web page is higher than a preset danger level; a second judgment unit is configured to judge whether there is a web page associated with the dangerous web page among the plurality of target web pages according to the association relationship between the plurality of target web pages; a sixth processing unit is configured to, if there is no web page associated with the dangerous web page among the plurality of target web pages, allow access to the web pages other than the dangerous web page among the plurality of target web pages; a seventh processing unit is configured to, if there is a web page associated with the dangerous web page among the plurality of target web pages, prohibit access to the dangerous web page and the web page associated with the dangerous web page among the plurality of target web pages.

[0020] In a third aspect, the present application further provides an electronic device, including one or more processors and a memory, where the memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the web page clustering method according to any one of the first aspect.

[0021] In a fourth aspect, the present application further provides a computer-readable storage medium storing a program, where the program executes the web page clustering method according to any one of the first aspect.

[0022] The beneficial effects of the present invention are as follows: Through the present application, the following steps are adopted: obtaining a plurality of target web pages, and extracting target texts in each target web page to obtain a plurality of target texts, where the importance level of each target text is greater than a preset importance level; inputting each target text into a target semantic extraction model for processing to obtain semantic information of each target text, where the target semantic extraction model is a model constructed based on a pre-trained language model; inputting the semantic information of each target text into a target clustering model for processing to obtain a target clustering result for clustering the plurality of target web pages, where the target clustering model is constructed using a fuzzy clustering algorithm, thus solving the problem of low accuracy in clustering web pages in the related art. By extracting the core text elements (i.e., target texts) in each target web page, using a semantic extraction model (i.e., the target semantic extraction model) constructed based on a pre-trained language model to extract the semantic information of the text, and then clustering the plurality of target web pages according to the semantic information of the text using a fuzzy clustering algorithm to obtain a clustering result (i.e., the target clustering result), compared with the prior art, the present solution can accurately extract the core elements in the web page, consider the semantic information of the core elements in the web page, and thus achieve the effect of improving the accuracy of clustering web pages. Description of the Drawings

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0024] Figure 1 It is a flowchart of a web page clustering method provided by an embodiment of the present application; Figure 2 It is a schematic diagram for determining the type of a child node provided by this embodiment; Figure 3 It is a flowchart of an optional web page clustering method provided by an embodiment of the present application; Figure 4 Schematic diagram of a web page clustering device provided by an embodiment of the present application; Figure 5 Schematic diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0025] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0026] In order to enable those skilled in the art of the present technology to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so as to implement the embodiments of the present application described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0028] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) involved in the present disclosure are all information and data authorized by the user or fully authorized by all parties. For example, an interface is provided between the present system and relevant users or institutions. Before obtaining relevant information, a request for acquisition needs to be sent to the aforementioned users or institutions through the interface, and after receiving the consent information feedback from the aforementioned users or institutions, the relevant information can be obtained.

[0029] For the convenience of description, some nouns or terms involved in the embodiments of the present application are described below: TF-IDF (Term Frequency-Inverse Document Frequency): TF-IDF is a statistical method commonly used in information retrieval and text mining to evaluate the importance of a term in a document collection or corpus.

[0030] BOW (Bag of Words): BOW is a text representation method widely used in natural language processing and information retrieval. It ignores the order and grammar of words and only cares about the frequency of word occurrences. A document is represented as a vector of the number of occurrences of words in a vocabulary.

[0031] K-means: It is a commonly used clustering algorithm for partitioning data points into K clusters, with each cluster represented by a center (centroid). The algorithm determines the cluster membership of each data point by minimizing the squared distance from the data point to its assigned centroid through an iterative optimization process.

[0032] DBSCAN is a density-based clustering algorithm. It can discover clusters of arbitrary shapes and identify noise points.

[0033] DOM (Document Object Model): The Document Object Model is a programming interface for representing and manipulating HTML and XML documents. DOM represents the document structure as a tree structure, where each node corresponds to a part of the document (such as elements, attributes, and text, etc.).

[0034] BERT: It is a pre-trained language model based on the Transformer architecture. By performing unsupervised pre-training (masked language model and next sentence prediction tasks) on a large-scale text corpus, it learns rich semantic representations.

[0035] FCM is a fuzzy clustering algorithm, similar to K-means, but the difference is that the membership degree of each data point belonging to each cluster is fuzzy, and each data point can partially belong to multiple clusters.

[0036] The present invention will be described below in conjunction with preferred implementation steps. Figure 1 It is a flowchart of a web page clustering method provided according to an embodiment of the present application, as Figure 1 shown. The method includes the following steps: Step S101, obtain multiple target web pages, and extract the target text in each target web page to obtain multiple target texts, where the importance degree of each target text is greater than a preset importance degree.

[0037] For example, the above-mentioned target text can be the core data extracted from web page data. That is, important core data (corresponding to each of the above-mentioned target texts) can be extracted from each web page (corresponding to each of the above-mentioned target web pages).

[0038] Step S102: Input each target text into the target semantic extraction model for processing to obtain the semantic information of each target text. Among them, the target semantic extraction model is a model constructed based on a pre-trained language model.

[0039] For example, the above-mentioned target semantic extraction model can be a BERT model. For instance, the important core data (corresponding to each of the above-mentioned target texts) extracted from each web page (corresponding to each of the above-mentioned target web pages) can be input into the BERT model (corresponding to the above-mentioned target semantic extraction model), and the semantic information of the important core data (corresponding to each of the above-mentioned target texts) extracted from each web page (corresponding to each of the above-mentioned target web pages) is output.

[0040] Step S103: Input the semantic information of each target text into the target clustering model for processing to obtain the target clustering result for clustering multiple target web pages. Among them, the target clustering model is constructed using a fuzzy clustering algorithm.

[0041] For example, website data can be classified using the fuzzy text clustering method. Specifically, a fuzzy clustering algorithm can be used, and based on the semantic information of the text, the web pages (corresponding to the above-mentioned multiple target web pages) are clustered to obtain the above-mentioned target clustering result.

[0042] Through the above steps S101 to S103, by extracting the core text elements (i.e., target texts) in each target web page, and using the semantic extraction model (i.e., target semantic extraction model) constructed based on the pre-trained language model to extract the semantic information of the text, and then according to the semantic information of the text, and using the fuzzy clustering algorithm to cluster multiple target web pages to obtain the clustering result (i.e., target clustering result). Compared with the prior art, the present solution can accurately extract the core elements in the web page, and consider the semantic information of the core elements in the web page, thereby achieving the effect of improving the accuracy of clustering the web pages.

[0043] Optionally, in the web page clustering method provided in the embodiments of the present application, extracting the target text in each target web page to obtain a plurality of target texts includes: obtaining the code information used to construct each target web page, and obtaining a plurality of target documents according to the code information used to construct each target web page, where each target document is used to construct each target web page; using the Document Object Model to convert each target document into a tree structure to obtain a plurality of target trees; obtaining the target nodes in each target tree, and extracting the target text in each target web page based on the target nodes in each target tree to obtain a plurality of target texts.

[0044] For example, in order to accurately extract and process the information in a web page, the HTML code (corresponding to the above-mentioned code information) can be parsed into a DOM tree structure (corresponding to the above-mentioned target tree). That is, using DOM (Document Object Model), an HTML document (corresponding to the above-mentioned target document) can be represented as a tree structure, where each tag, attribute, and text node is an object. Then, the key data nodes can be obtained from the DOM tree, and the text data (corresponding to the above-mentioned target text) can be extracted from the key data nodes.

[0045] In summary, by parsing the HTML code into a DOM tree structure, the accurate extraction of the core semantics of the web page can be realized, and the efficiency and accuracy of data extraction are improved.

[0046] Optionally, in the web page clustering method provided in the embodiments of the present application, obtaining the target nodes in each target tree includes: obtaining the tags of the target type in each target tree, and deleting the tags of the target type from each target tree to obtain a plurality of target trees after deleting the tags, where the degree of correlation between the data content in the tags of the target type and the semantic information is less than a preset degree of correlation; obtaining a plurality of nodes from each target tree after deleting the tags, and deleting the nodes of the first type and the nodes of the second type from the plurality of nodes to obtain a plurality of target trees after deleting the nodes, where the nodes of the first type are the nodes without child nodes, and the nodes of the second type are the nodes with one child node; performing grouping processing on the nodes in each target tree after deleting the nodes according to the similarity between the nodes in each target tree after deleting the nodes to obtain a plurality of processing results; and obtaining the target nodes in each target tree based on each processing result.

[0047] For example, after obtaining the DOM tree (corresponding to the above-mentioned target tree), the tags with insufficient data meaning and unnecessary content (corresponding to the above-mentioned tags of the target type) can be deleted: (1) Title tags: The tag list includes h6, h5, h4, h3, h2, h1, and head tags. These tags are usually used as page and chapter titles. The text of the title tags is generally short and no further semantic meaning can be extracted.

[0048] (2) Input tags: input tags, selection tags, form tags, button tags, etc. These tags are usually used for user data input and have no specific semantic meaning.

[0049] (3) Script and style tags: These tags are responsible for the page rendering effect and response design and do not involve data semantics.

[0050] (4) Other tags: Such as iframe, map, HTML comments. Although these elements occupy a certain amount of page space, they have no impact on semantic parsing.

[0051] After deleting the above tags, obtain the subsequent tag tree that can be processed and used (corresponding to the target tree after deleting the tags above).

[0052] Then perform child node deletion. For example, the following two types of child nodes can be deleted: (1) Single data nodes (corresponding to the first type of nodes above): As Figure 2 shown, the p tag is a single data node of the body tag. This type of node has no child nodes. And Figure 2 is a schematic diagram for determining the type of child nodes in this embodiment.

[0053] (2) Recursive single nodes (corresponding to the second type of nodes above): As Figure 2 shown, the table tag has only 1 child node, and this child node has only one child node. Recursively check the child nodes of its child nodes in a depth-first manner. If the number of child nodes of all its recursive child nodes below is 1, delete such nodes. According to the data mining results, the common manifestation forms of this type of node are advertisements or other data irrelevant to the theme form.

[0054] Then group similar nodes: After performing the node deletion operation, parameters such as the depth of the data node from the root node, tag name, child node count, child data node tag name, and attribute comparison can be considered to find similar data nodes.

[0055] First, initialize the similar node set dictionary sel of similarElementsList. Store all node elements with similar attributes in a list and traverse each node in a depth-first manner. Pass each node to be checked to the node similarity judgment function elmentsimilarCompare function to check whether the node is similar to the nodes in the similar element list. If similar, append the current node to the similar list of the corresponding node. If not similar, add the current node as a new node to the dictionary.

[0056] The check items for similarity judgment in the elmentsimilarCompare function include: the node depth based on recursive depth of the root node, the tag name (which determines the overall behavior of the node), the number of child nodes and the tag names of child nodes, and the number of attributes. After the similarity judgment is completed, when extracting nodes later, the nodes in the similar list in the dictionary are regarded as the same content and extracted in batches.

[0057] Finally, the steps for extracting data of the main data nodes (corresponding to the above-mentioned target nodes) are as follows: Sort the number of child nodes of each node in the sel dictionary in descending order, and select all nodes smaller than the node number threshold from the node list. The node number threshold is determined by adjusting hyperparameters. After data mining and analysis, the threshold is selected as 5, and this hyperparameter can be adjusted according to specific tasks. At the same time, because the data nodes with relevant data have more text content than other node elements in the DOM tree, calculate the total text content length of the node and its similar nodes, and obtain the node set with the largest text as the main data nodes, and extract the text data of the node list.

[0058] In summary, by deleting irrelevant tags and child nodes to optimize the data structure, the efficiency and accuracy of data processing are improved.

[0059] Optionally, in the web page clustering method provided in the embodiment of the present application, input each target text into the target semantic extraction model for processing, and the semantic information obtained for each target text includes: input each target text into the target semantic extraction model for processing to obtain a plurality of feature vectors, where each feature vector is used to represent the semantic features of each target text; perform dimensionality reduction processing on each feature vector to obtain a plurality of processed feature vectors; based on each processed feature vector, obtain the semantic information of each target text.

[0060] For example, the BERT model (corresponding to the above-mentioned target semantic extraction model) can be used for semantic feature extraction. Among them, the text data extracted (corresponding to the above-mentioned target text) can be preprocessed first, mainly by performing tokenization (adding special tokens at the head and end), padding, and encoding. Performing padding and truncation mainly ensures that each text data document has the same length of tokens. Encoding maps the tokens to integers, which is convenient for the BERT to process the document to extract semantic features later. The text data representation is obtained by feeding the processed text data into the model, and the text data representation is obtained from the output given by the penultimate Transformer encoder layer.

[0061] Then, feature dimensionality reduction is performed. Two feature extraction strategies are adopted to convert the high-dimensional representation of BERT into a fixed-size feature vector with a lower dimension, namely, max pooling and average pooling, and layer normalization is performed to ensure that the fixed-size vector representation has the characteristics of normal distribution stability. The layer normalization strategy can avoid the problem of covariate shift during the training process of the neural network and perform feature dimensionality reduction.

[0062] In summary, by using the BERT model to extract semantic features and simultaneously perform dimensionality reduction optimization, the representation ability of web page content can be improved.

[0063] Optionally, in the web page clustering method provided in the embodiment of the present application, the target semantic extraction model is obtained in the following manner: obtaining a plurality of sample web pages, and determining whether there are a first web page and a second web page among the plurality of sample web pages, where the first web page and the second web page are the same sample web page; if there are a first web page and a second web page among the plurality of sample web pages, then deleting the first web page or the second web page from the plurality of sample web pages to obtain a plurality of deleted sample web pages; constructing a training data set based on the plurality of deleted sample web pages; using the training data set to learn and train the original semantic extraction model to obtain the target semantic extraction model, where the type of the original semantic extraction model includes a pre-trained language model.

[0064] For example, when collecting website data, in order to ensure the generalization ability and representativeness of the training set and the test set, a sufficient number of diverse data samples can be collected. Search result pages obtained from mainstream search engines can be used to cover web pages in different fields and topics to ensure the representativeness and comprehensiveness of the training set and the test set. At the same time, the diversity of data samples in terms of language, format, structure, etc. is ensured to enhance the adaptability of the model. New web page data can be continuously obtained through web crawler technology and API interfaces. After collecting the original data (corresponding to the above-mentioned plurality of sample web pages), if it is determined that there are duplicate web page data (corresponding to the above-mentioned first web page and second web page) in the collected original data, then the original data can be de-duplicated and deleted, that is, the duplicate web page data (corresponding to the above-mentioned first web page and second web page) can be deleted, and then the de-duplicated website data can be aggregated together to form the above-mentioned training data set; if there is no duplicate web page data (corresponding to the above-mentioned first web page and second web page) in the collected original data, then no de-duplication process is required, and the collected original data can be directly used to form the training data set. Then, the training data set is used to train the original BERT model (corresponding to the above-mentioned original semantic extraction model) to obtain the finally trained BERT model (corresponding to the above-mentioned target semantic extraction model). And the BERT model is a pre-trained language model based on the Transformer architecture.

[0065] Through the above solution, by deduplicating the web page data in the original data, it is possible to avoid deviations during the model training process, thereby ensuring the accuracy of the semantic extraction model.

[0066] Optionally, in the web page clustering method provided in the embodiments of the present application, the target clustering model is obtained in the following manner: using a fuzzy clustering algorithm to construct an original clustering model, and using the original clustering model to cluster the sample web pages in the training data set to obtain a sample clustering result; determining whether the accuracy of the sample clustering result is greater than a preset value; if the accuracy of the sample clustering result is greater than the preset value, using the original clustering model as the target clustering model; if the accuracy of the sample clustering result is not greater than the preset value, adjusting the parameters of the original clustering model to obtain an adjusted clustering model; and obtaining the target clustering model based on the adjusted clustering model.

[0067] For example, the data that was previously labeled for the collected data (corresponding to the above-mentioned training data set) can be used to evaluate the clustering result (corresponding to the above-mentioned sample clustering result) as an evaluation metric. The best mapping between the obtained clusters is evaluated using the true labels, the classification result (corresponding to the above-mentioned sample clustering result) is evaluated by the classification accuracy and the similarity of the clustered data, and the corresponding hyperparameters are adjusted to select the initial clustering centers with the best accuracy and recall rate.

[0068] Through the above solution, the accuracy of the clustering model for clustering web page data can be improved.

[0069] Optionally, in the web page clustering method provided in the embodiments of the present application, after inputting the semantic information of each target text into the target clustering model for processing to obtain a target clustering result for clustering multiple target web pages, the method further includes: determining the association relationship between the multiple target web pages according to the target clustering result, and obtaining the dangerous web pages among the multiple target web pages, where the danger level of the dangerous web pages is higher than a preset danger level; judging whether there are web pages associated with the dangerous web pages among the multiple target web pages according to the association relationship between the multiple target web pages; if there are no web pages associated with the dangerous web pages among the multiple target web pages, allowing access to the web pages other than the dangerous web pages among the multiple target web pages; if there are web pages associated with the dangerous web pages among the multiple target web pages, prohibiting access to the dangerous web pages and the web pages associated with the dangerous web pages among the multiple target web pages.

[0070] For example, the above-mentioned dangerous web pages can be web pages in phishing websites, etc. For instance, according to the association relationship between web pages, web pages associated with the dangerous web page can be determined; and if no web pages associated with the dangerous web page are found, then the dangerous web page cannot be accessed at this time, and web pages other than the dangerous web page can be accessed; if web pages associated with the dangerous web page are found, then the dangerous web page and the web pages associated with the dangerous web page cannot be accessed at this time, and web pages other than the dangerous web page and the web pages associated with the dangerous web page can be accessed.

[0071] Through the above solution, the security of the data accessed by users can be guaranteed.

[0072] In addition, this embodiment relates to the field of data security, and a website clustering method based on automatic data node extraction in this embodiment aims to solve the limitations of existing website clustering algorithms in terms of data positioning, semantic text representation, and the accuracy of clustering results. Through diverse data sources and precise manual annotation, a dataset with wide coverage and representativeness is constructed, ensuring the accuracy and adaptability of the model when processing different types of web page data. Through DOM tree parsing, tag optimization, similar node grouping, and core node data extraction, the accurate extraction and streamlining of the core semantics of web pages are achieved, improving the efficiency and accuracy of data extraction. In the semantic extraction stage, the BERT model is used for semantic feature extraction while dimensionality reduction optimization is performed to improve the representation ability of web page content. The EFCM algorithm is used for fuzzy text clustering. High-dimensional data is processed through dimensionality reduction and matrix decomposition, and an adaptive distance metric is introduced. The distance metric method is adjusted according to the density and distribution of data points to more accurately capture the similarity and difference between data, improving the accuracy of clustering.

[0073] For example, Figure 3 is a flowchart of an optional web page clustering method provided according to an embodiment of the present application, as Figure 3 shown. The optional web page clustering method includes the following steps: Step 1: Dataset construction.

[0074] Step (1): Website data collection. To ensure the generalization ability and representativeness of the training set and test set, collect a sufficient number of diverse data samples. The search result pages obtained from mainstream search engines cover web pages in different fields and topics, ensuring the representativeness and comprehensiveness of the training set and test set. At the same time, ensure the diversity of data samples in terms of language, format, structure, etc., to enhance the adaptability of the model. Through web crawler technology and API interfaces, continuously obtain new web page data. After collecting the original data, deduplicate and delete the data, deleting duplicate web page data to avoid bias during model training.

[0075] Step (2): Clustering data annotation. To test the accuracy of the solution, 5%-10% of the data is extracted and some of the data is manually annotated in advance. The main tasks of annotation include: Subject classification: According to the theme of the web page content, the web page is divided into different categories. For example, technology, finance, entertainment, etc.

[0076] Tag annotation: Add tags to each web page content to describe its main content and features. For example, a news website can be annotated as "financial news", "technology news", etc.

[0077] Similarity annotation: Annotate the similarity for a part of web page pairs to evaluate the effect of the clustering algorithm. For example, two news articles describing the same event should be annotated as highly similar.

[0078] Step 2: Key data node extraction. In this step, all data nodes are located, and then the most useful and relevant nodes on the page are extracted. The format of the node determination object is {tag name, HTML content, tag depth, child node count}, where: Step (1): DOM tree parsing. To accurately extract and process the information in the web page, the HTML code is parsed into a DOM tree structure. The DOM tree represents the HTML document as a tree structure, where each tag, attribute, and text node is an object, eliminating the influence of some web pages that may contain incomplete or damaged tags.

[0079] Step (2) Tag deletion. After obtaining the DOM tree, delete the tags with insufficient data meaning and unnecessary content: Title tags: The tag list includes h6, h5, h4, h3, h2, h1, head tags. These tags are usually used as page and chapter titles. The title tags generally have short text and cannot extract further semantic meaning.

[0080] Input tags: Input tags, select tags, form tags, button tags, etc. These tags are usually used for user data input and have no specific semantic meaning.

[0081] Script and style tags: These tags are responsible for the page rendering effect and responsive design and do not involve data semantics.

[0082] Other tags: Such as iframe, map, HTML comments. Although these elements occupy a certain amount of page space, they have no impact on semantic parsing.

[0083] After deleting the above tags, a tag tree that can be processed subsequently is obtained.

[0084] Step (3) Child node deletion. Delete two types of child nodes: Monadic data node (corresponding to the node of the first type mentioned above): For example Figure 2 As shown, the p tag is the monadic data node of the body tag, and this type of node has no child nodes. And Figure 2 This is the schematic diagram for determining the type of child node in this embodiment.

[0085] Recursive monadic node (corresponding to the node of the second type mentioned above): For example Figure 2 As shown, the table tag has only 1 child node, and this child node has only one child node. Recursively check the child nodes of its child nodes in a depth-first manner. If the number of child nodes of all recursive child nodes below it is 1, delete this type of node. According to the data mining results, the common manifestation forms of this type of node are advertisements or other data irrelevant to the theme form.

[0086] Step (4): Group similar nodes.

[0087] After performing the node deletion operation, consider parameters such as the depth of the data node from the root node, label name, child node count, label name of the child data node, and attribute comparison to find similar data nodes.

[0088] First, initialize the similarElementsList dictionary sel of the similar node set. Store all node elements with similar attributes in a list, and traverse each node in a depth-first manner. Pass each node to be checked to the elmentsimilarCompare function for judging node similarity to check whether the node is similar to the nodes in the similar element list. If it is similar, append the current node to the similar list of the corresponding node. If it is not similar, add the current node as a new node to the dictionary.

[0089] The items for the elmentsimilarCompare function to judge similarity include: the node depth recursively based on the root node depth, label name (which determines the overall behavior of the node), the number of child nodes and the label name of the child nodes, and the number of attributes. After the similarity determination is completed, when extracting nodes later, the nodes in the similar list in the dictionary are regarded as the same content for batch extraction.

[0090] Step (5): Extract the data of the main data nodes. Sort the nodes in the node list in descending order according to the number of child nodes of each node in the sel dictionary. Select all nodes smaller than the node number threshold from the node list. The node number threshold is determined by adjusting the hyperparameter. After data mining and analysis, the selected threshold is 5, and this hyperparameter can be adjusted according to the specific task. At the same time, because the data nodes with relevant data have more text content than other node elements in the DOM tree, calculate the total text content length of the node and its similar nodes, and obtain the node set with the largest text as the main data nodes, and extract the text data of this node list.

[0091] Step 3: Use BERT to extract semantic features.

[0092] Step (1): Preprocess the text data extracted in Step 2, mainly performing tokenization (adding special tokens at the head and end), padding, and encoding. Performing padding and truncation mainly ensures that each text data document has the same length of tokens. Encoding maps the tokens to integers, which is convenient for subsequent processing of the documents by BERT to extract semantic features. The processed text data is fed forward into the model to obtain the text data representation, and the text data representation is obtained from the output given by the penultimate Transformer encoder layer.

[0093] Step (2): Feature dimensionality reduction. Two feature extraction strategies are adopted to convert the high-dimensional representation of BERT into a fixed-size feature vector of lower dimension, namely max pooling and mean pooling, and layer normalization is performed to ensure that the fixed-size vector representation has stable features with normal distribution. The layer normalization strategy can avoid the problem of covariate shift during the training process of the neural network and perform feature dimensionality reduction.

[0094] Step 4: Classify the website data using the fuzzy text clustering method.

[0095] Step (1): Perform fuzzy clustering using EFCM. Since the FCM algorithm has certain limitations in clustering high-dimensional data, the high-dimensional data is first converted into low-dimensional data, and the truncated singular value decomposition (TSVD) is used to reduce the dimension of the data D. The result of matrix decomposition is used as the data input. At this time, we assume a set of decomposed data is divided into M classes, which has a set of cluster centroids , and the sum of the distances from each data to the cluster center is expressed as:

[0096] where is the fuzziness, representing the membership degree of data d belonging to class x. Through distance calculation, c is a hyperparameter to control the fuzziness adjustment. F is used as the clustering training optimization objective, and the smaller it is, the better the clustering classification effect. The cluster centers are iteratively updated and the results are reallocated to continuously optimize the clustering results until the minimum clustering criterion is met. Finally, the website classification results are obtained according to the clustering results. Each cluster is regarded as a website category, which contains website text data with similar features and themes.

[0097] Step (2): Clustering optimization. Use the previously labeled data to evaluate the clustering results as an evaluation index, use the true labels to evaluate the best mapping between the obtained clusters, evaluate the classification results by the classification accuracy and the similarity of the clustered data, and perform corresponding hyperparameter adjustments to select the initial cluster centers with the best accuracy and recall rate.

[0098] In addition, during the practice process, the method provided in this embodiment can also be implemented and applied in multiple scenarios, as listed below: 1. Network Association Analysis Engine In the application scenario of the association analysis engine, the method proposed in this embodiment can provide an effective means to extract and analyze the association information in web pages through automatic data node extraction and website clustering techniques, helping users discover potential association relationships and patterns. By continuously monitoring and analyzing the association information in the Internet, a dynamic association graph can be further generated to help users intuitively understand complex association relationships and provide in-depth insights and decision-making support.

[0099] 2. Phishing Website Mining In the application scenario of phishing mining, the method provided in this embodiment can provide an efficient and intelligent solution. Through automatic data node extraction and optimized clustering algorithms, websites with similar semantic information are extracted and clustered from a large amount of web page data, further reducing the amount of data for phishing algorithm discrimination, promoting the improvement of phishing detection accuracy and efficiency, and helping users prevent online fraud behavior.

[0100] 3. Industry Domain Corpus Construction In the application scenario of domain corpus construction, the method provided in this embodiment can automatically locate the main content nodes related to a specific domain in web pages, intelligently identify and extract the text data in these nodes, ignore irrelevant information, use Transformer to further generalize semantic features, optimize the clustering algorithm, and classify similar website data documents into one category according to the similarity of text content, extracting and organizing high-quality domain text data from a large amount of web page data to construct a corpus with rich content and clear structure. This can not only improve the efficiency of data collection and processing, but also significantly enhance the quality and applicability of the corpus, providing data support for natural language processing tasks.

[0101] Moreover, the method provided in this embodiment has the following several advantages compared with the prior art: Traditional website data extraction methods often show great limitations when facing web pages with complex structures and variable formats, and it is difficult to accurately extract the information required by users. The method provided in this embodiment can intelligently identify and extract the key nodes in web pages and ignore irrelevant information. This advantage enables the system to efficiently process a large amount of web page data, quickly locate and extract useful information, significantly reducing manual intervention and the workload of data processing, and improving the overall efficiency and accuracy of information extraction.

[0102] Because traditional TF-IDF and BOW methods have significant deficiencies in capturing semantic information and context relationships in text, the clustering effect is poor. The method provided in this embodiment can better understand and represent the semantics of text by introducing BERT to optimize semantic representation, enabling the system to more accurately capture complex semantic relationships in web page text. The feature dimension reduction and optimization processes carried out make the feature representation more stable and efficient, which can not only improve the quality of text representation but also further significantly improve the effect of clustering analysis, enabling the system to better classify and organize web page data and provide high-quality clustering results for various application scenarios. The method provided in this embodiment performs fuzzy clustering through EFCM, taking into account the uncertainty of fuzzy factors. The adaptive fuzzy parameters can dynamically adjust the values of fuzzy parameters according to the characteristics of data distribution, improving the stability and reliability of clustering and enabling it to adapt to the characteristics of different data sets. At the same time, for the associated clustering task, it can also significantly improve the stability and accuracy of the clustering results. It also further considers the comprehensive factors of the distance from data points to the cluster centers and membership degrees, enabling the cluster centers to more accurately represent the central positions of each cluster and improving the accuracy of similarity measurement.

[0103] In data node extraction, by parsing HTML code into a DOM tree structure, the integrity and accuracy of tags are ensured, and the data structure is optimized by deleting irrelevant tags and child nodes, improving the efficiency and accuracy of data processing. The similar node grouping optimization used in this embodiment can help identify nodes with similar structures and contents in web pages. These nodes often contain the same type of information. By grouping them, it can be ensured that the extracted data is relevant and consistent, thereby improving the accuracy of data extraction. By calculating the text content lengths of nodes and their similar nodes, the most important node information is identified and extracted to ensure that the extracted data has the highest value.

[0104] In summary, the web page clustering method provided in the embodiments of the present application obtains multiple target web pages, extracts target texts in each target web page to obtain multiple target texts, where the importance level of each target text is greater than a preset importance level; inputs each target text into a target semantic extraction model for processing to obtain semantic information of each target text, where the target semantic extraction model is a model constructed based on a pre-trained language model; inputs the semantic information of each target text into a target clustering model for processing to obtain a target clustering result for clustering the multiple target web pages, where the target clustering model is constructed using a fuzzy clustering algorithm, thereby solving the problem of low accuracy in clustering web pages in the related art. By extracting the core text elements (i.e., target texts) in each target web page, using a semantic extraction model constructed based on a pre-trained language model (i.e., the target semantic extraction model) to extract the semantic information of the text, and then clustering the multiple target web pages according to the semantic information of the text using a fuzzy clustering algorithm to obtain a clustering result (i.e., the target clustering result), compared with the prior art, this solution can accurately extract the core elements in the web page, consider the semantic information of the core elements in the web page, and thus achieve the effect of improving the accuracy of clustering web pages.

[0105] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0106] The embodiments of the present application also provide a web page clustering device. It should be noted that the web page clustering device in the embodiments of the present application can be used to execute the web page clustering method provided in the embodiments of the present application. The following introduces the web page clustering device provided in the embodiments of the present application.

[0107] Figure 4 is a schematic diagram of the web page clustering device provided in the embodiments of the present application. As Figure 4 shown, the device includes: a first acquisition unit 401, a first processing unit 402, and a second processing unit 403.

[0108] Specifically, the first acquisition unit 401 is configured to obtain multiple target web pages, and extract target texts in each target web page to obtain multiple target texts, where the importance level of each target text is greater than a preset importance level; The first processing unit 402 is configured to input each target text into a target semantic extraction model for processing to obtain semantic information of each target text, where the target semantic extraction model is a model constructed based on a pre-trained language model; A second processing unit 403, configured to input the semantic information of each target text into a target clustering model for processing, so as to obtain a target clustering result for clustering multiple target web pages, where the target clustering model is constructed by using a fuzzy clustering algorithm.

[0109] In summary, the web page clustering device provided in the embodiment of the present application obtains multiple target web pages through a first acquisition unit 401, and extracts target texts in each target web page to obtain multiple target texts, where the importance level of each target text is greater than a preset importance level; a first processing unit 402 inputs each target text into a target semantic extraction model for processing to obtain the semantic information of each target text, where the target semantic extraction model is a model constructed based on a pre-trained language model; a second processing unit 403 inputs the semantic information of each target text into a target clustering model for processing to obtain a target clustering result for clustering multiple target web pages, where the target clustering model is constructed by using a fuzzy clustering algorithm, and solves the problem of low accuracy in clustering web pages in the related art. By extracting the core text elements (i.e., target texts) in each target web page, using a semantic extraction model constructed based on a pre-trained language model (i.e., the target semantic extraction model) to extract the semantic information of the text, and then clustering multiple target web pages according to the semantic information of the text by using a fuzzy clustering algorithm to obtain a clustering result (i.e., the target clustering result), compared with the prior art, the present solution can accurately extract the core elements in the web page, consider the semantic information of the core elements in the web page, and further achieve the effect of improving the accuracy of clustering web pages.

[0110] Optionally, in the web page clustering device provided in the embodiment of the present application, the first acquisition unit includes: a first acquisition module, configured to acquire code information for constructing each target web page, and obtain multiple target documents according to the code information for constructing each target web page, where each target document is used to construct each target web page; a first conversion module, configured to use a document object model to convert each target document into a tree structure to obtain multiple target trees; a second acquisition module, configured to acquire target nodes in each target tree, and extract target texts in each target web page based on the target nodes in each target tree to obtain multiple target texts.

[0111] Optionally, in the web page clustering device provided in the embodiments of the present application, the second acquisition module includes: a first processing sub-module, configured to acquire tags of a target type in each target tree, and delete the tags of the target type from each target tree to obtain a plurality of target trees after deleting the tags, where the degree of correlation between the data content in the tags of the target type and the semantic information is less than a preset degree of correlation; a second processing sub-module, configured to acquire a plurality of nodes from each target tree after deleting the tags, and delete the nodes of the first type and the nodes of the second type from the plurality of nodes to obtain a plurality of target trees after deleting the nodes, where the nodes of the first type are nodes without child nodes, and the nodes of the second type are nodes with one child node; a third processing sub-module, configured to perform grouping processing on the nodes in each target tree after deleting the nodes according to the similarity between the nodes in each target tree after deleting the nodes, to obtain a plurality of processing results; a first acquisition sub-module, configured to acquire target nodes in each target tree based on each processing result.

[0112] Optionally, in the web page clustering device provided in the embodiments of the present application, the first processing unit includes: a first processing module, configured to input each target text into a target semantic extraction model for processing to obtain a plurality of feature vectors, where each feature vector is used to represent the semantic features of each target text; a second processing module, configured to perform dimensionality reduction processing on each feature vector to obtain a plurality of processed feature vectors; a third processing module, configured to obtain the semantic information of each target text based on each processed feature vector.

[0113] Optionally, in the web page clustering device provided in the embodiments of the present application, the target semantic extraction model is obtained through the following units: a third processing unit, configured to acquire a plurality of sample web pages, and determine whether a first web page and a second web page exist in the plurality of sample web pages, where the first web page and the second web page are the same sample web page; a first deletion unit, configured to, if the first web page and the second web page exist in the plurality of sample web pages, delete the first web page or the second web page from the plurality of sample web pages to obtain a plurality of sample web pages after deletion; a first construction unit, configured to construct a training data set based on the plurality of sample web pages after deletion; a first training unit, configured to use the training data set to perform learning and training on an original semantic extraction model to obtain the target semantic extraction model, where the type of the original semantic extraction model includes a pre-trained language model.

[0114] Optionally, in the web page clustering device provided in the embodiments of the present application, the target clustering model is obtained through the following units: a fourth processing unit, configured to construct an original clustering model by using a fuzzy clustering algorithm, and perform clustering processing on the sample web pages in the training dataset by using the original clustering model to obtain a sample clustering result; a first determination unit, configured to determine whether the accuracy rate of the sample clustering result is greater than a preset value; a first determination unit, configured to use the original clustering model as the target clustering model if the accuracy rate of the sample clustering result is greater than the preset value; a first adjustment unit, configured to adjust the parameters of the original clustering model if the accuracy rate of the sample clustering result is not greater than the preset value to obtain an adjusted clustering model; a second determination unit, configured to obtain the target clustering model based on the adjusted clustering model.

[0115] Optionally, in the web page clustering device provided in the embodiments of the present application, the device further includes: a fifth processing unit, configured to, after inputting the semantic information of each target text into the target clustering model for processing to obtain a target clustering result for clustering multiple target web pages, determine the association relationship between the multiple target web pages according to the target clustering result, and obtain the dangerous web pages among the multiple target web pages, where the danger level of the dangerous web pages is higher than the preset danger level; a second determination unit, configured to determine whether there are web pages associated with the dangerous web pages among the multiple target web pages according to the association relationship between the multiple target web pages; a sixth processing unit, configured to allow access to the web pages other than the dangerous web pages among the multiple target web pages if there are no web pages associated with the dangerous web pages among the multiple target web pages; a seventh processing unit, configured to prohibit access to the dangerous web pages and the web pages associated with the dangerous web pages among the multiple target web pages if there are web pages associated with the dangerous web pages among the multiple target web pages.

[0116] The web page clustering device includes a processor and a memory. The above-mentioned first acquisition unit 401, first processing unit 402, second processing unit 403, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions.

[0117] The processor includes a kernel, and the kernel retrieves the corresponding program units from the memory. One or more kernels can be set, and the accuracy of clustering web pages can be improved by adjusting the kernel parameters.

[0118] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0119] The embodiments of the present invention provide a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the web page clustering method is implemented.

[0120] An embodiment of the present invention provides a processor for running a program, wherein when the program runs, the web page clustering method is executed.

[0121] As Figure 5 shown, an embodiment of the present invention provides an electronic device, the device includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, the following steps are implemented: obtaining a plurality of target web pages, and extracting target texts in each target web page to obtain a plurality of target texts, wherein the importance level of each target text is greater than a preset importance level; inputting each target text into a target semantic extraction model for processing to obtain semantic information of each target text, wherein the target semantic extraction model is a model constructed based on a pre-trained language model; inputting the semantic information of each target text into a target clustering model for processing to obtain a target clustering result for clustering the plurality of target web pages, wherein the target clustering model is constructed using a fuzzy clustering algorithm.

[0122] When the processor executes the program, the following steps are also implemented: extracting target texts in each target web page to obtain a plurality of target texts, including: obtaining code information for constructing each target web page, and based on the code information for constructing each target web page, obtaining a plurality of target documents, wherein each target document is used to construct each target web page; using the Document Object Model to convert each target document into a tree structure to obtain a plurality of target trees; obtaining target nodes in each target tree, and based on the target nodes in each target tree, extracting target texts in each target web page to obtain a plurality of target texts.

[0123] When the processor executes the program, the following steps are also implemented: obtaining target nodes in each target tree includes: obtaining tags of a target type in each target tree, and deleting the tags of the target type from each target tree to obtain a plurality of target trees after deleting the tags, wherein the degree of correlation between the data content in the tags of the target type and the semantic information is less than a preset degree of correlation; obtaining a plurality of nodes from each target tree after deleting the tags, and deleting nodes of a first type and nodes of a second type from the plurality of nodes to obtain a plurality of target trees after deleting the nodes, wherein the nodes of the first type are nodes without child nodes, and the nodes of the second type are nodes with one child node; performing grouping processing on the nodes in each target tree after deleting the nodes according to the similarity between the nodes in each target tree after deleting the nodes to obtain a plurality of processing results; based on each processing result, obtaining target nodes in each target tree.

[0124] When the processor executes the program, the following steps are also implemented: input each target text into the target semantic extraction model for processing to obtain the semantic information of each target text, including: input each target text into the target semantic extraction model for processing to obtain multiple feature vectors, where each feature vector is used to represent the semantic features of each target text; perform dimensionality reduction processing on each feature vector to obtain multiple processed feature vectors; based on each processed feature vector, obtain the semantic information of each target text.

[0125] When the processor executes the program, the following steps are also implemented: the target semantic extraction model is obtained in the following manner: obtain multiple sample web pages, and determine whether there are a first web page and a second web page in the multiple sample web pages, where the first web page and the second web page are the same sample web page; if there are a first web page and a second web page in the multiple sample web pages, then delete the first web page or the second web page from the multiple sample web pages to obtain multiple deleted sample web pages; based on the multiple deleted sample web pages, construct a training data set; use the training data set to perform learning and training on the original semantic extraction model to obtain the target semantic extraction model, where the type of the original semantic extraction model includes a pre-trained language model.

[0126] When the processor executes the program, the following steps are also implemented: the target clustering model is obtained in the following manner: use the fuzzy clustering algorithm to construct the original clustering model, and use the original clustering model to perform clustering processing on the sample web pages in the training data set to obtain a sample clustering result; determine whether the accuracy of the sample clustering result is greater than a preset value; if the accuracy of the sample clustering result is greater than the preset value, then use the original clustering model as the target clustering model; if the accuracy of the sample clustering result is not greater than the preset value, then adjust the parameters of the original clustering model to obtain an adjusted clustering model; based on the adjusted clustering model, obtain the target clustering model.

[0127] When the processor executes the program, the following steps are also implemented: after inputting the semantic information of each target text into the target clustering model for processing to obtain a target clustering result for clustering multiple target web pages, it further includes: based on the target clustering result, determine the association relationship between multiple target web pages, and obtain the dangerous web pages among the multiple target web pages, where the danger level of the dangerous web pages is higher than the preset danger level; according to the association relationship between multiple target web pages, determine whether there are web pages associated with the dangerous web pages among the multiple target web pages; if there are no web pages associated with the dangerous web pages among the multiple target web pages, then allow access to the web pages other than the dangerous web pages among the multiple target web pages; if there are web pages associated with the dangerous web pages among the multiple target web pages, then prohibit access to the dangerous web pages and the web pages associated with the dangerous web pages among the multiple target web pages.

[0128] The device in this article can be a server, PC, PAD, mobile phone, etc.

[0129] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with the following method steps: obtaining a plurality of target web pages, and extracting target text in each target web page to obtain a plurality of target texts, wherein the importance level of each target text is greater than a preset importance level; inputting each target text into a target semantic extraction model for processing to obtain semantic information of each target text, wherein the target semantic extraction model is a model constructed based on a pre-trained language model; inputting the semantic information of each target text into a target clustering model for processing to obtain a target clustering result for clustering the plurality of target web pages, wherein the target clustering model is constructed using a fuzzy clustering algorithm.

[0130] When executed on a data processing device, it is also adapted to execute a program initialized with the following method steps: extracting target text in each target web page to obtain a plurality of target texts, including: obtaining code information for constructing each target web page, and obtaining a plurality of target documents based on the code information for constructing each target web page, wherein each target document is used to construct each target web page; using the Document Object Model to convert each target document into a tree structure to obtain a plurality of target trees; obtaining target nodes in each target tree, and based on the target nodes in each target tree, extracting target text in each target web page to obtain a plurality of target texts.

[0131] When executed on a data processing device, it is also adapted to execute a program initialized with the following method steps: obtaining target nodes in each target tree includes: obtaining tags of a target type in each target tree, and deleting the tags of the target type from each target tree to obtain a plurality of target trees after tag deletion, wherein the degree of correlation between the data content in the tags of the target type and the semantic information is less than a preset degree of correlation; obtaining a plurality of nodes from each target tree after tag deletion, and deleting nodes of a first type and nodes of a second type from the plurality of nodes to obtain a plurality of target trees after node deletion, wherein the nodes of the first type are nodes without child nodes, and the nodes of the second type are nodes with one child node; performing grouping processing on the nodes in each target tree after node deletion according to the similarity between the nodes in each target tree after node deletion to obtain a plurality of processing results; based on each processing result, obtaining target nodes in each target tree.

[0132] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: input each target text into a target semantic extraction model for processing to obtain the semantic information of each target text, including: input each target text into the target semantic extraction model for processing to obtain a plurality of feature vectors, where each feature vector is used to represent the semantic features of each target text; perform dimensionality reduction processing on each feature vector to obtain a plurality of processed feature vectors; and based on each processed feature vector, obtain the semantic information of each target text.

[0133] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: the target semantic extraction model is obtained by the following means: obtaining a plurality of sample web pages and determining whether there are a first web page and a second web page among the plurality of sample web pages, where the first web page and the second web page are the same sample web page; if there are a first web page and a second web page among the plurality of sample web pages, then delete the first web page or the second web page from the plurality of sample web pages to obtain a plurality of deleted sample web pages; construct a training data set based on the plurality of deleted sample web pages; and use the training data set to perform learning and training on the original semantic extraction model to obtain the target semantic extraction model, where the type of the original semantic extraction model includes a pre-trained language model.

[0134] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: the target clustering model is obtained by the following means: using a fuzzy clustering algorithm to construct an original clustering model, and using the original clustering model to perform clustering processing on the sample web pages in the training data set to obtain a sample clustering result; determining whether the accuracy rate of the sample clustering result is greater than a preset value; if the accuracy rate of the sample clustering result is greater than the preset value, then use the original clustering model as the target clustering model; if the accuracy rate of the sample clustering result is not greater than the preset value, then adjust the parameters of the original clustering model to obtain an adjusted clustering model; and based on the adjusted clustering model, obtain the target clustering model.

[0135] When executed on a data processing device, it is also suitable for executing a program initialized with the following method steps: After inputting the semantic information of each target text into a target clustering model for processing to obtain a target clustering result for clustering multiple target web pages, it further includes: determining the association relationship between multiple target web pages according to the target clustering result, and obtaining dangerous web pages among the multiple target web pages, where the danger level of the dangerous web pages is higher than a preset danger level; judging whether there are web pages associated with the dangerous web pages among the multiple target web pages according to the association relationship between the multiple target web pages; if there are no web pages associated with the dangerous web pages among the multiple target web pages, allowing access to the web pages other than the dangerous web pages among the multiple target web pages; if there are web pages associated with the dangerous web pages among the multiple target web pages, prohibiting access to the dangerous web pages and the web pages associated with the dangerous web pages among the multiple target web pages.

[0136] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0137] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or a combination of multiple flows and / or blocks

[0138] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device realizes the functions specified in Figure 1 one or more of the flows Figure 1 or a combination of multiple flows and / or blocks

[0139] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 steps of the functions specified in one block or multiple blocks.

[0140] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0141] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0142] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0143] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of another identical element in the process, method, commodity or device comprising the element.

[0144] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0145] The above are only the embodiments of the present application and are not used to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A web page clustering method, characterized in that: The method includes: Acquire multiple target web pages, and extract target text from each target web page to obtain multiple target texts, wherein the importance of each target text is greater than a preset importance; Input each target text into a target semantic extraction model for processing to obtain semantic information of each target text, wherein the target semantic extraction model is a model built based on a pre-trained language model; The semantic information of each target text is input into a target clustering model for processing to obtain a target clustering result for clustering the plurality of target web pages, wherein the target clustering model is constructed using a fuzzy clustering algorithm.

2. The method according to claim 1, characterized in that The target text in each target web page is extracted to obtain multiple target texts, including: Obtaining code information used to construct each target webpage, and obtaining a plurality of target documents based on the code information used to construct each target webpage, wherein each target document is used to construct each target webpage; Using the document object model, each target document is converted into a tree structure to obtain multiple target trees; A target node in each target tree is obtained, and based on the target node in each target tree, a target text in each target web page is extracted to obtain the plurality of target texts.

3. The method according to claim 2, characterized in that The step of obtaining a target node in each target tree includes: Obtaining a label of a target type in each target tree, and deleting the label of the target type from each target tree, to obtain a plurality of target trees after the labels are deleted, wherein the relevance between the data content in the label of the target type and the semantic information is less than a preset relevance; Acquire multiple nodes from each target tree after deleting labels, and delete nodes of the first type and nodes of the second type from the multiple nodes to obtain multiple target trees after deleting nodes, wherein the nodes of the first type are nodes without child nodes, and the nodes of the second type are nodes with one child node; According to the similarity between the nodes in the target tree after each node is deleted, the nodes in the target tree after each node is deleted are grouped and processed to obtain multiple processing results; Based on each processing result, a target node in each target tree is obtained.

4. The method according to claim 1, characterized in that: The step of inputting each target text into the target semantic extraction model for processing to obtain the semantic information of each target text includes: Input each target text into the target semantic extraction model for processing to obtain multiple feature vectors, wherein each feature vector is used to represent the semantic features of each target text; Perform dimensionality reduction processing on each feature vector to obtain multiple processed feature vectors; Based on each processed feature vector, the semantic information of each target text is obtained.

5. The method according to claim 1, characterized in that The target semantic extraction model is obtained by: Acquire a plurality of sample web pages, and determine whether a first web page and a second web page exist in the plurality of sample web pages, wherein the first web page and the second web page are the same sample web pages; If the first web page and the second web page exist in the plurality of sample web pages, deleting the first web page or the second web page from the plurality of sample web pages to obtain a plurality of deleted sample web pages; Constructing a training data set based on the plurality of deleted sample web pages; The original semantic extraction model is trained by using the training data set to obtain the target semantic extraction model, wherein the type of the original semantic extraction model includes a pre-trained language model.

6. The method according to claim 5, characterized in that The target clustering model is obtained in the following way: Using a fuzzy clustering algorithm to construct an original clustering model, and using the original clustering model to cluster the sample web pages in the training data set to obtain a sample clustering result; Determine whether the accuracy of the sample clustering result is greater than a preset value; If the accuracy of the sample clustering result is greater than the preset value, the original clustering model is used as the target clustering model; If the accuracy of the sample clustering result is not greater than the preset value, adjusting the parameters of the original clustering model to obtain an adjusted clustering model; Based on the adjusted clustering model, the target clustering model is obtained.

7. The method according to claim 1, characterized in that After inputting the semantic information of each target text into the target clustering model for processing to obtain a target clustering result for clustering the plurality of target web pages, the method further includes: Determining the association relationship between the multiple target web pages according to the target clustering result, and obtaining dangerous web pages among the multiple target web pages, wherein the danger level of the dangerous web pages is higher than a preset danger level; According to the association relationship between the multiple target web pages, determining whether there is a web page associated with the dangerous web page among the multiple target web pages; If there is no web page associated with the dangerous web page among the multiple target web pages, access to web pages other than the dangerous web page among the multiple target web pages is allowed; If there is a webpage associated with the dangerous webpage among the multiple target webpages, access to the dangerous webpage and the webpage associated with the dangerous webpage among the multiple target webpages is prohibited.

8. A web page clustering device, characterized in that: include: A first acquisition unit is used to acquire a plurality of target web pages and extract a target text in each target web page to obtain a plurality of target texts, wherein the importance of each target text is greater than a preset importance; A first processing unit is used to input each target text into a target semantic extraction model for processing to obtain semantic information of each target text, wherein the target semantic extraction model is a model constructed based on a pre-trained language model; The second processing unit is used to input the semantic information of each target text into the target clustering model for processing to obtain a target clustering result of clustering the multiple target web pages, wherein the target clustering model is constructed using a fuzzy clustering algorithm.

9. An electronic device, characterized in that: It comprises one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the web page clustering method described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The storage medium stores a program, wherein the program executes the web page clustering method according to any one of claims 1 to 7.