Text and image feature fused bad website classification method, system and equipment and medium
The DistilBERT-BiLSTM-Attention model is used to process text and CLIP model images, and combined with multimodal comparison learning, the problems of difficult text semantic capture in bad website detection and weak generalization ability of CNN models in the existing technology are solved, and more efficient and accurate bad website detection is achieved.
Patent Information
- Application Number
- CN202510250471.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-04
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to effectively capture the sequence dependence and global semantics of text in poor website detection, and the CNN model requires a large amount of label data for training, and the generalization ability is weak.
The text is processed using the DistilBERT-BiLSTM-Attention model, and the deep context information is extracted through DistilBERT. BiLSTM captures the long-distance dependence of the sequence and focuses on key information through the Attention mechanism. At the same time, the CLIP model is used for image content processing, and the generalization ability of the model is improved through multimodal comparison learning and large-scale data pre-training.
It improves the accuracy of bad website detection and generalization capabilities of models, can make more efficient use of the complementary characteristics of text and images, and enhances the ability to classify complex websites.
Smart Images

Figure CN120145115A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of text image classification, and in particular relates to a method, system, device and medium for classifying malicious websites by integrating text and image features. Background Art
[0002] With the rapid development and popularization of the Internet, people's lives are becoming more and more closely connected to the Internet. The huge quantity makes the monitoring and management of websites more and more difficult, and many lawbreakers take the opportunity to insert malicious information into websites. At the same time, the rapid spread and wide coverage of Internet information have led to a sharp increase in the number of malicious websites, and the control difficulty has also increased accordingly, exposing loopholes in management. After some website domain names expire, their filing information fails to be cancelled in time, resulting in a large number of unmaintained domain names being exploited. These abandoned domain names are often tampered with into malicious websites, further exacerbating the deterioration of the network environment. Therefore, in order to more efficiently maintain the campus network environment, it is necessary to propose a method that can accurately detect malicious websites and actively discover malicious links in campus websites.
[0003] The technical solution closest to the present invention is a CNN model based on multi-modal features proposed by ALSAEDI M et al. in 2024 [ALSAEDI M, GHALEB F A, SAEED F, et al. Multi-Modal Features Representation-Based Convolutional Neural Network Model for Malicious Website Detection[J]. IEEE Access, 2024, 12: 7271-7284.]. This solution extracts text features from URLs and DNSs, extracts image features from web page images at the same time, finally combines the outputs of two CNN models, and classifies malicious websites through an artificial neural network classifier. CNN models are difficult to capture sequential dependencies and global semantics in text, especially performing poorly when dealing with long texts; in terms of image content processing, CNN-based models require a large amount of labeled data for training and have weak generalization ability.
[0004] In addition, Mudaliar et al. proposed an adult website detection method based on CNN and FastText in 2023 [WEN L, ZHANG M, WANG C, et al. MEDAL: A Multimodality-Based Effective DataAugmentation Framework for Illegal Website Identification[J / OL]. Electronics, 2024, 13(11).]. This method uses the FastText algorithm to classify the web page text content, and at the same time uses CNN to extract and classify the features of the web page images. Finally, the results of the image and text classifiers are combined through a weighted fusion technique, thus improving the accuracy of adult website detection. The adult website detection method based on CNN and FastText has some limitations in text content processing, unable to effectively capture word order, long-distance dependencies, and complex context semantics, nor can it learn deep features. Summary of the Invention
[0005] In order to overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to propose a bad website classification method, system, device and medium that fuse text and image features. First, a DistilBERT-BiLSTM-Attention model is established for text processing and classification; secondly, a CLIP model is used for image content processing; the strong generalization ability and accuracy of CLIP are ensured through multi-modal contrast learning and large-scale data pre-training; in terms of feature selection, for text content, features are extracted by selecting the web page title, text extracted by image OCR, and web page text content data; for image content, embedded web page images and web page screenshots are selected; the present invention makes full use of various contents of the web page, combines the complementary characteristics of text and image, improves the classification accuracy, and enhances the model's ability to handle complex websites.
[0006] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0007] A method for classifying malicious websites by integrating text and image features. First, obtain the image and text data of the website and construct a multi-modal website content dataset. Second, establish a DistilBERT-BiLSTM-Attention model for text processing and classification. DistilBERT extracts deep contextual information from the input text, BiLSTM captures long-distance dependencies in the sequence, and then the Attention mechanism focuses on the most important words or sentence parts to enhance the model's perception of key information. Finally, the classification result is output through the Sigmoid function. Then, use a malicious image classification model based on CLIP to process the image content and predict the classification of the web page image content. Ensure the accuracy of the output features through multi-modal contrast learning and large-scale data pre-training. Finally, in terms of feature selection, for text content, extract features by selecting the web page title, text extracted by image OCR, and web page text content data; for image content, select embedded web page images and web page screenshots. Make full use of various types of content on the web page, combine the complementary characteristics of text and images, and perform classification.
[0008] A method for classifying malicious websites by integrating text and image features specifically includes the following steps:
[0009] Step 1: Obtain the image and text data of the website and construct a multi-modal website content dataset;
[0010] Step 2: Based on the dataset obtained in Step 1, construct a malicious text classification model based on DistilBERT-BiLSTM-Attention, extract features and perform classification prediction by combining the web page title, web page text content, and OCR text content of embedded web page images to achieve accurate classification prediction of web page text content;
[0011] Step 3: According to the dataset obtained in Step 1, construct a malicious image classification model based on CLIP, and at the same time use embedded web page images and web page screenshots for feature analysis and classification to achieve accurate classification prediction of web page image content;
[0012] Step 4: According to the classification prediction methods in Step 2 and Step 3, effectively combine the prediction probabilities of text content and image content through logistic regression, balance the contribution degrees of web page content and web page images to the website classification result, obtain the website classification result, and achieve a more reliable and accurate classification decision.
[0013] The specific method of Step 1 is as follows:
[0014] 1.1 Obtaining web page text data: First, use the Scrapy crawler framework and the Selenium tool to obtain the image and text data of the website. Use the Selenium tool to dynamically render the website interface. After the website is loaded, obtain the browser instance. Collect the web page title through the browser instance, and at the same time obtain the tags. Collect the web page text content through the InnerText attribute of the tags. Then, use the regular expression ([a-z0-9,;′:|-_# / !\.\$\{\}\(\)]{10,}) to filter the collected web page text content and remove the CSS style code and JavaScript code in it.
[0015] 1.2 Obtaining image data: Directly obtain the web page screenshot through the browser instance, and at the same time obtain all the tags in the page, and preliminarily screen the images through the width and height attributes, retaining the images with pixels greater than 80*80. Finally, extract the embedded image content of the web page through the src attribute.
[0016] 1.3 Preprocessing of text data: First, perform word segmentation on the text content, then count and screen the stop words, and then delete the stop words from the word segmentation results of the text. Secondly, use the Jieba tool in Python to perform word segmentation on the web page text. Then, use TF-IDF (Term Frequency-Inverse Document Frequency) to encode the words. Sort the words according to the TF-IDF value from low to high, and extract the first several words as the preliminary stop word list. After obtaining the preliminary stop word list, perform manual screening to remove the words with high relevance to the web page theme and retain the words irrelevant to the actual content of the web page, so as to generate the stop word list of the web page content. Then, delete the stop words in the word segmentation results of the web page text, remove the irrelevant text information, and improve the quality of the text content. Finally, split the text content into text lines in units of several characters to meet the needs of subsequent data analysis and model training.
[0017] 1.4 Preprocessing of image data: First, perform custom scaling on the web page screenshot, and then intercept the area in the center of the image as the final image. For the embedded images on the web page, first filter the images that cannot be decoded through the Pillow tool in Python, and then scale the remaining images to the same size as the intercepted final image. Use OCR to scan and extract the text content from the embedded images on the web page as a text feature of the web page.
[0018] Finally, a multi-modal content dataset of the website is obtained.
[0019] The specific method of step 2 is as follows:
[0020] 2.1 For the web page text data obtained in step 1, select n segmented lines as the input of the DistilBERT model for vectorization representation to obtain the feature vector matrix X 1 =(x 1 ,x 2 ,…,x n ); when the number of lines of the web page text data is less than n lines, take all lines and obtain the feature vector matrix X 1 =(x 1 ,x 2 ,…,x m ), pad n - m zero vectors at the end to finally obtain the n-dimensional vector matrix X 1 ; meanwhile, for the website title and the text data in the web page embedded images extracted by OCR, after preprocessing their original text content, take one line for vectorization to obtain the feature vectors X 2 and X 3 ;
[0021] 2.2 Take the vectors X 1 , X 2 , X 3 as inputs and transmit them to the BiLSTM for processing to obtain the corresponding hidden states H 1 , H 2 , H 3 ; then, use the attention mechanism to further aggregate the hidden states H 1 , H 2 , H 3 to obtain the global feature vector. The specific calculation is as follows:
[0022] For the hidden state H 1 of the web page text content, H 1 =[h 1 ,h 2 ,...h i ...h n ,
[0023] First, calculate the attention scores. For each hidden state h i , calculate its correlation score e i with the context:
[0024]
[0025] where W and b are learnable parameters, and v is the attention weight vector;
[0026] Next, normalize the attention scores. Convert the scores e i to weights a i through the Softmax function:
[0027]
[0028] Finally, the global feature vector is obtained by weighted summation:
[0029]
[0030] For the website title and the text extracted by OCR, the corresponding input sequence is H = [h 1 , and there is only one hidden state. Therefore, the global feature vector is directly equal to the hidden state h 1 ; The three global feature vectors calculated from the web page text content, website title, and text extracted by OCR are concatenated to obtain the feature vector y, and then the feature vector y is input into the Sigmoid function to obtain the predicted probability P t of the text content.
[0031] The specific method of step 3 is as follows:
[0032] Use the CLIP framework, select ViT-L / 14@336px as the image encoder, select the input image and output the predicted probability to construct a CLIP-based bad image classification model; for the web page embedded image and the web page screenshot, use the CLIP-based bad image classification model to output the predicted probability respectively, and fuse the probabilities of the web page embedded image and the web page screenshot to obtain the predicted probability of the web page image content;
[0033] For the web page embedded image, select n images and use the CLIP-based bad image classification model to obtain a set of predicted probabilities, and take the average value of this set of predicted probabilities as the predicted probability P 1 of the web page embedded image;
[0034] For the website screenshot, directly obtain the predicted probability P 2 of the website screenshot through the CLIP model; then, weight and sum the predicted probabilities of the embedded image and the website screenshot to obtain the final predicted probability of the web page image:
[0035] P i = a 1 ·P 1 + a 2 ·P 2
[0036] where the weights a 1 and a 2 respectively represent the importance of the web page embedded image and the website screenshot for bad image classification; cut out a part of the website multi-modal content dataset in step 1 as the training set, and through the predicted probability collection of the web page screenshots and web page embedded images in the training set, that is, the predicted probabilities P of the web page embedded images obtained from N groups of different data1 and the predicted probability P of the website screenshot 2 , which is input into the logistic regression model to calculate the weight coefficient a' 1 and a' 2 , and perform normalization processing on it to obtain the final weight value:
[0037]
[0038] The specific method of step 4 is as follows: for the predicted probability P of the web page text content t and the predicted probability P of the image content i , they are fused by weighted summation to obtain the final predicted probability:
[0039] P = b 1 ·P t + b 2 ·P i
[0040] Then, set the prediction threshold th for final classification; the weights b 1 and b 2 are obtained by inputting the predicted probability collection of the text data and image data in the training set in step 3 into the logistic regression model for calculation and normalization processing; when the final predicted probability P obtained by fusing the text content and the web page content prediction probability is greater than th, it is considered that the website is a bad website; otherwise, it is considered that the website is a healthy website.
[0041] A bad website classification system that fuses text and image features, including:
[0042] A data preparation module, used for step 1, to obtain the images and text data of the website through the Scrapy crawler framework and the Selenium tool, construct a bad image classification model based on CLIP, which can obtain the rendered website page and batch obtain the web page content to achieve efficient data collection; through the data processing script, it can perform preprocessing such as screening and correction on the text content and image content collected by the crawler to improve the data quality and reliability;
[0043] A text content classification module, used for step 2, to implement a text classification method based on the DistilBERT - BiLSTM - Attention model through design, which can extract features and perform classification prediction by combining the web page title, web page text content, and OCR text content of the embedded images on the web page to achieve accurate classification prediction of the web page text content;
[0044] An image content classification module, for step 3, by constructing a CLIP-based bad image classification model, can simultaneously use in-page embedded images and web page screenshots for feature analysis and classification, and achieve accurate prediction of web page image content classification;
[0045] A classification decision module, for step 4, effectively combines the prediction probabilities of text content and image content through logistic regression, can fully balance the contribution degrees of web page content and web page images to the website classification result, and achieve a more reliable and accurate classification decision.
[0046] A bad website classification device that fuses text and image features, including:
[0047] A memory, used to store computer programs;
[0048] A processor, used to implement the bad website classification method that fuses text and image features described in steps 1 to 4 when executing the computer program.
[0049] A computer-readable storage medium, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement bad website classification that fuses text and image features based on the method described in steps 1 to 4.
[0050] Compared with the prior art, the present invention has the following advantages:
[0051] (1) In the prior art, the CNN model is difficult to capture the sequential dependencies and global semantics of text, especially performs poorly when dealing with long texts; while FastText cannot effectively capture word order, long-distance dependencies, and complex context semantics, nor can it learn deep features. The present invention uses text classification technology. First, the DistilBERT model is used to vectorize the text data, thereby solving the problem that the same word has different meanings in different contexts. Then, the BiLSTM model is used to capture the context semantic features of the text, and the attention mechanism is used to focus on key information. Finally, the classification result is output through the Sigmoid function. It solves the problem that the same word has different meanings in different positions and fully obtains the semantic information of the text content, improving the accuracy of bad text content classification.
[0052] (2) In terms of image content processing, existing CNN models require a large amount of labeled data for training and have weak generalization ability. The present invention uses image classification technology to classify images through the CLIP model by matching the semantics of natural language with the natural language description. The ViT-L / 14@336px is selected as the image encoder to extract image features, fully capturing the global and more detailed information of the image, and improving the accuracy of bad image classification.
[0053] (3) In terms of feature selection, the present invention uses a decision fusion method. For text content, three data, namely the web page title, the text extracted by image OCR, and the web page text content, are selected for feature extraction; for image content, the embedded images in the web page and the web page screenshots are selected. By making full use of various types of content on the web page and combining the complementary characteristics of text and image, the classification accuracy is improved, and the ability of the model to handle complex websites is enhanced.
[0054] In summary, the present invention fully mines text semantic information by designing a DistilBERT-BiLSTM-Attention text classification model, combines the CLIP model and ViT-L / 14@336px to enhance the global semantic understanding of images while taking into account image detail information, and adopts a strategy of fusing the prediction probabilities of text and image modalities, having the advantages of strong adaptability to complex scenarios and high classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] Figure 1 is a flowchart of the present invention.
[0056] Figure 2 is a structural diagram of a method for detecting bad websites by fusing text and image features of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The following further describes the present invention in detail with reference to the accompanying drawings and specific embodiments.
[0058] See Figure 1 That is Figure 2 , Step 1: Construct a multi-modal website content dataset. Through web crawler technology and data processing technology, an image and text content dataset of the website is obtained. The present invention adopts the Scrapy crawler framework based on the Python language. First, the Selenium tool is used to dynamically render the website interface. After the website is loaded, the browser instance is obtained. Through the browser instance, the web page title and web page screenshots are collected, and at the same time, the tags are obtained, and the web page text content is collected through its InnerText attribute. Then, the regular expression ([a-z0-9,;′:|-_# / !\.\$\{\}\(\)]{10,}) is used to filter the text content to remove the CSS style code and JavaScript code therein. Subsequently, all tags in the page are obtained through the browser instance, and the images are preliminarily screened through the width and height attributes, and the images with the short side greater than 80 pixels are retained. Finally, the embedded image content in the web page is extracted through the src attribute.
[0059] The processing process of the text content is as follows: first, tokenize the text content, then count and filter the stop words, and finally remove the stop words from the text tokenization results. First, use the Jieba tool in Python to tokenize the web page text, and then use TF-IDF (Term Frequency-Inverse Document Frequency) to encode the words. Sort the words according to the TF-IDF values from low to high, and extract the top 100 words as the preliminary stop word list. The encoding idea of TF-IDF is that the words with high frequencies in the document are more important, while the words that appear in multiple documents in the dataset are relatively unimportant, and the encoding value represents the importance degree of the word. After obtaining the preliminary stop word list, further filter it manually, remove the words with high relevance to the web page theme, and retain the words irrelevant to the actual content of the web page, so as to generate the stop word list of the web page content. Then, further filter the web page text in combination with the common Chinese stop word list to remove the irrelevant text information and improve the quality of the text content. Finally, split the text content into text lines with 64 characters as a unit to meet the needs of subsequent data analysis and model training.
[0060] The processing process of the image content is as follows: first, scale the short side of the web page screenshot to 336 pixels, and scale the long side proportionally, and then intercept the 336x336 area in the center of the image as the final image. For the images embedded in the web page, first filter the images with data errors, and then scale the remaining images to 336x336 pixels. Through the OCR technology, extract the text content from the images embedded in the web page as a text feature of the web page.
[0061] Finally, a multi-modal content dataset of the website is obtained.
[0062] Step 2: Build a bad text classification model based on DistilBERT-BiLSTM-Attention. For the web page text content, select n segmented lines as the input of the DistilBERT model for vectorization representation to obtain the feature vector matrix X 1 =(x 1 ,x 2 ,…,x n ). If the number of lines of the web page text content is less than n lines, take all the lines and obtain the feature vector matrix X 1 =(x 1 ,x 2 ,…,x m ), and supplement n-m zero vectors at the end to finally obtain an n-dimensional vector matrix X 1 . At the same time, for the website title and the text content in the images embedded in the web page extracted by OCR, due to their short lengths, only take one line of the preprocessed data for vectorization to obtain the feature vector X2 and X 3 。
[0063] Take the vector X 1 、X 2 、X 3 as inputs and transmit them to the BiLSTM for processing respectively to obtain the corresponding hidden states H 1 、H 2 、H 3 。Then, use the attention mechanism to further aggregate these hidden states to obtain the global feature vector. For the hidden state H 1 of the web page text content, first, calculate the attention scores. For each hidden state h i , calculate its correlation score e i with the context as follows:
[0064]
[0065] where W and b are learnable parameters, and v is the attention weight vector. Then, normalize the attention scores, and transform the scores e i into weights a i through the Softmax function as follows:
[0066]
[0067] Finally, perform a weighted sum to obtain the global feature vector:
[0068]
[0069] For the website title and the text extracted by OCR, its corresponding input sequence is H = [h 1 , which has only one hidden state. Therefore, the global feature vector is directly equal to the hidden state h 1 . Concatenate the three global feature vectors to obtain the feature vector y, and then input this feature vector into the Sigmoid function to obtain the final classification result of the text content.
[0070] Use the DistilBERT model to perform vectorized representation on the text data, thus solving the problem that the same word has different meanings in different contexts. Use the BiLSTM model to capture the context semantic features of the text, and focus on the key information through the attention mechanism. Finally, output the classification result through the Sigmoid function. Solve the problem that the same word has different meanings in different positions and fully obtain the semantic information of the text content, improving the accuracy of the classification of bad text content.
[0071] Step 3: Build a CLIP-based bad image classification model. Select a pre-trained CLIP model and choose ViT-L / 14@336px as the image encoder to extract image features, fully capture the global and more detailed information of the image, and improve the accuracy of bad image classification. For in-page images and web page screenshots, perform binary classification predictions respectively, that is, judge whether it is a bad image or a healthy image. For in-page images, select n images for classification prediction respectively, and take the average of the final prediction probabilities as the prediction probability P of the in-page image. 1 For website screenshots, directly obtain the prediction probability P through the CLIP model. 2 Then, sum the two prediction probabilities after weighting to obtain the final prediction probability of the web page image:
[0072] P = a 1 ·P 1 + a 2 ·P 2
[0073] Among them, the weights a 1 and a 2 respectively represent the importance of in-page images and web page screenshots for bad image classification. Through the prediction probabilities of web page screenshots and in-page images in the training set, input them into the logistic regression model, calculate the weight coefficients a' 1 and a' 2 , and perform normalization processing on them to obtain the final weight values:
[0074]
[0075] Step 4: Obtain the website classification result. For the prediction probabilities P t and P i of the web page text content and image content, use the method of weighted summation for fusion to obtain the final prediction probability:
[0076] P = b 1 ·P t + b 2 ·P i
[0077] Then, set the prediction threshold th for the final classification. The weights b 1 and b 2 are obtained by inputting the prediction probabilities of the text data and image data in the training set into the logistic regression model, calculating, and performing normalization processing. If the final prediction probability P is greater than th, then the website is considered a bad website; otherwise, the website is considered a healthy website.
[0078] In step 4, a decision fusion method is used. In terms of text content, three data, namely the web page title, the text extracted by image OCR, and the web page text content, are selected for feature extraction; in terms of image content, the embedded images in the web page and the web page screenshots are selected. By making full use of various contents of the web page and combining the complementary characteristics of text and image, the classification accuracy is improved and the ability of the model to handle complex websites is enhanced.
[0079] Experimental verification
[0080] The multi-modal website content dataset in step 1 is input into the method of the present invention for experiments, and the accuracy rate, precision rate, recall rate, and F1 value are used to comprehensively evaluate the classification results of website sensitive content.
[0081] Experimental results of different text models
[0082] The experimental results are shown in Table 1.1 below:
[0083] Table 1.1 Experimental results of different text models
[0084]
[0085]
[0086] The experimental results show that all indicators of the DistilBERT-BiLSTM-Attention model are better than the other text classification models.
[0087] Experimental results of text line numbers
[0088] To explore the performance of the DistilBERT-BiLSTM-Attention model when the number of lines of different web page text contents is different and determine the optimal number of input text lines, in this experiment, the number of text lines is set to 1, 5, 10, 15, 20, 25, and 30 respectively for experiments. The test results are shown in Table 1.2:
[0089] Table 1.2 Experiments with different text line numbers
[0090]
[0091] From the experimental data, it can be seen that when 20 lines of web page text are selected as the input, the performance of the DistilBERT-BiLSTM-Attention model is the best. Therefore, 20 lines of text are selected as the input of the model to obtain the best classification effect.
[0092] Experiments on different image models
[0093] The experimental results are shown in Table 1.3 below:
[0094] Table 1.3 Experimental results of different image models
[0095]
[0096]
[0097] As can be seen from the experimental results, the CLIP model outperforms the other image classification models in all metrics, verifying its effectiveness on the website image content dataset.
[0098] Experiment with different numbers of images
[0099] To explore the influence of different in-page images on the classification performance of the CLIP-based inappropriate image classification model, 1, 3, 5, 10, and 15 images were selected for testing. The specific experimental results are shown in Table 1.4 below:
[0100] Table 1.4 Experimental results of different multimodal models
[0101]
[0102] From the experimental results, it can be seen that when 5 in-page images are selected, the model has the best classification effect.
[0103] Experiment with different prediction thresholds
[0104] To explore the influence of different prediction thresholds on the final classification performance, the threshold range was set from 0.3 to 0.7 for the experiment. The experimental results are shown in Table 1.5:
[0105] Table 1.5 Experimental results of different prediction thresholds
[0106]
[0107]
[0108] It can be found from the experimental results that when the threshold is 0.5, the model has the best classification effect, and the F1 value reaches 96.82%. Therefore, 0.5 is selected as the threshold for classification decision to obtain the optimal classification result.
Claims
1. A bad website classification method integrating text and image features, characterized in that: Firstly, the image and text data of the website are obtained to construct a multimodal website content dataset. Secondly, the DistilBERT-BiLSTM-Attention model is established to classify text processing. DistilBERT is used to extract deep contextual information from the input text, BiLSTM is used to capture long-distance dependencies in the sequence, and the Attention mechanism is used to focus on the most important words or sentences to improve the model's perception of key information. Finally, the classification results are output through the Sigmoid function. Then, the CLIP-based bad image classification model is used to process image content and predict the classification of web page image content. The accuracy of output features is ensured through multimodal comparative learning and large-scale data pre-training. Finally, in terms of feature selection, the text content is extracted by selecting web page titles, text extracted by image OCR, and web page text content data. In terms of image content, embedded images and web page screenshots are selected. All kinds of content on the web page are fully utilized, and the complementary characteristics of text and image are combined for classification.
2. The bad website classification method integrating text and image features according to claim 1 is characterized in that: The specific steps include: Step 1: Obtain image and text data from the website and build a multimodal website content dataset; Step 2: Based on the data set obtained in step 1, a bad text classification model based on DistilBERT-BiLSTM-Attention is constructed, and feature extraction and classification prediction are performed based on web page titles, web page text content, and web page embedded image OCR text content to achieve accurate web page text content classification prediction; Step 3: Based on the data set obtained in step 1, a CLIP-based bad image classification model is constructed. At the same time, embedded images and web page screenshots are used for feature analysis and classification to achieve accurate web page image content classification prediction; Step 4: Based on the classification prediction methods of steps 2 and 3, the prediction probabilities of text content and image content are effectively combined through the logistic regression method, the contribution of web page content and web page images to the website classification results is balanced, the website classification results are obtained, and more reliable and accurate classification decisions are achieved.
3. The bad website classification method integrating text and image features according to claim 2 is characterized in that: The specific method of step 1 is: 1.1 Obtain web page text data: First, use the Scrapy crawler framework and Selenium tools to obtain the website's image and text data, use the Selenium tool to dynamically render the website interface, and after the website is loaded, obtain the browser instance; collect the web page title through the browser instance, and obtain the tag at the same time, and collect the web page text content through the tag InnerText attribute; then, use the regular expression ([a-z0-9,;′:|-_# / !\.\$\{\}\(\)]{10,}) to filter the collected web page text content and remove the CSS style code and JavaScript code; 1.2 Get image data: Get a screenshot of the web page directly through the browser instance, and get all the Tags, and preliminarily screen images through width and height attributes, retain images with pixels larger than 80*80, and finally extract embedded image content on the web page through the src attribute; 1.3 Text data preprocessing: First, segment the text content, then count and filter the stop words, and then delete the stop words from the text segmentation results; secondly, use Python's Jieba tool to segment the web page text, and then use TF-IDF (Term Frequency-Inverse Document Frequency) to encode the words; sort the words from low to high according to the TF-IDF value, and extract the first few words as a preliminary stop word list; After obtaining the preliminary stop word list, we manually screened and removed the words that were highly relevant to the web page topic, and retained the words that were irrelevant to the actual content of the web page, thus generating a stop word list for the web page content; then, we deleted the stop words in the web page text segmentation results, removed irrelevant text information, and improved the quality of the text content; finally, we segmented the text content into text lines based on a number of characters to meet the needs of subsequent data analysis and model training; 1.4 Image data preprocessing: First, the webpage screenshot is custom scaled, and then the area in the center of the image is captured as the final image; for images embedded in the webpage, the Python Pillow tool is used to filter out the images that cannot be decoded, and then the retained images are scaled to the same size as the captured final image. The text content is scanned and extracted from the embedded images in the webpage through OCR as a text feature of the webpage; Finally, a multimodal content dataset of the website is obtained.
4. The bad website classification method integrating text and image features according to claim 2 is characterized in that: The specific method of step 2 is: 2.1 For the webpage text data obtained in step 1, select the n rows after segmentation as the input of the DistilBERT model for vectorization representation, and obtain the feature vector matrix X1 = (x1, x2, ..., x n ); when the number of rows of web page text data is less than n rows, all rows are taken and the eigenvector matrix X1=(x1,x2,…,x m ), fill nm zero vectors at the end, and finally obtain an n-dimensional vector matrix X1; at the same time, for the website title and the text data in the embedded image of the web page extracted by OCR, after preprocessing the original text content, take a row for vectorization to obtain the feature vectors X2 and X3; 2.2 Transmit vectors X1, X2, and X3 as input to BiLSTM for processing, and obtain the corresponding hidden states H1, H2, and H3; then, use the attention mechanism to further aggregate the hidden states H1, H2, and H3 to obtain the global feature vector. The specific calculation is as follows: For the hidden state H1 of the web page text content, H1=[h1,h2,...h i ...h n ], First, calculate the attention score, for each hidden state h i , calculate its relevance score with the context e i : e i =v T fishy i +b) Among them, W and b are learnable parameters, and v is the attention weight vector; Next, normalize the attention score and convert the score e into i Converted to weight a i : Finally, the weighted summation is performed to obtain the global eigenvector: For the website title and the text extracted by OCR, the corresponding input sequence is H = [h1], with only one hidden state, so the global feature vector is directly equal to the hidden state h1; the three global feature vectors calculated by the web page text content, website title, and text extracted by OCR are concatenated to obtain the feature vector y, and then the feature vector y is input into the Sigmoid function to obtain the predicted probability P of the text content. t .
5. The bad website classification method integrating text and image features according to claim 2 is characterized in that: The specific method of step 3 is: Use the CLIP framework, select ViT-L / 14@336px as the image encoder, select the input image and output the predicted probability to build a CLIP-based bad image classification model; for web page embedded images and web page screenshots, use the CLIP-based bad image classification model to output the predicted probability respectively, and after fusing the probabilities of web page embedded images and web page screenshots, get the predicted probability of web page image content; For web page embedded images, select n images and use the CLIP-based bad image classification model to obtain a set of prediction probabilities, and take the average of the set of prediction probabilities as the prediction probability P1 of the web page embedded image; For website screenshots, the predicted probability P2 of the website screenshots is directly obtained through the CLIP model; then, the predicted probability of the embedded image and the predicted probability of the webpage screenshot are weighted and summed to obtain the predicted probability of the final webpage image: <h2 style=";text-align:left;direction:ltr">P<h2 style=";text-align:left;direction:ltr"> i <h2 style=";text-align:left;direction:ltr"> =a1·P1+a2·P2 Among them, the weights a1 and a2 respectively represent the importance of website embedded images and web page screenshots to the classification of bad images; a part of the website multimodal content dataset in step 1 is cut out as a training set, and the predicted probability collection of web page screenshots and web page embedded images in the training set, that is, the predicted probability P1 of web page embedded images and the predicted probability P2 of website screenshots obtained from N groups of different data, are input into the logistic regression model, and the weight coefficients a'1 and a'2 are calculated, and they are normalized to obtain the final weight value:
6. The bad website classification method integrating text and image features according to claim 2 is characterized in that: The specific method of step 4 is: for the predicted probability P of the webpage text content t and the predicted probability P of the image content i , and the weighted summation method is used to fuse and obtain the final prediction probability: P=b1·P t +b2·P i Then, the prediction threshold th is set for final classification; the weights b1 and b2 are obtained by inputting the predicted probability set of text data and image data in the training set in step 3 into the logistic regression model for calculation and normalization; when the final predicted probability P obtained by integrating the predicted probability of text content and web page content is greater than th, the website is considered to be a bad website; otherwise, the website is considered to be a healthy website.
7. A bad website classification system integrating text and image features based on the method of claims 1-6, characterized in that: include: The data preparation module is used in step 1 to obtain the image and text data of the website through the Scrapy crawler framework and Selenium tools, and to build a bad image classification model based on CLIP. It can obtain the rendered website pages and obtain the web page content in batches to achieve efficient data collection; through the data processing script, it can pre-process the text content and image content collected by the crawler, such as screening and correction, to improve the data quality and reliability; The text content classification module is used in step 2. By designing and implementing a text classification method based on the DistilBERT-BiLSTM-Attention model, it can combine web page titles, web page text content, and web page embedded image OCR text content for feature extraction and classification prediction, thus achieving accurate web page text content classification prediction; The image content classification module is used in step 3. By building a bad image classification model based on CLIP, it can simultaneously use web page embedded images and web page screenshots for feature analysis and classification to achieve accurate web page image content classification prediction; The classification decision module is used in step 4. It effectively combines the prediction probabilities of text content and image content through the method of logistic regression, and can fully balance the contribution of web page content and web page images to the website classification results, thus achieving more reliable and accurate classification decisions.
8. A bad website classification device integrating text and image features based on the method according to any one of claims 1 to 6, characterized in that: include: Memory for storing computer programs; A processor is used to implement the bad website classification method integrating text and image features as described in claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it is possible to implement the classification of bad websites by integrating text and image features based on the methods described in claims 1 to 6.
Citation Information
Cited By
Academic paper title grading device and method based on Bert
CN120996032A
A device and method for classifying titles of academic papers based on Bert
CN120996032B