XSS attack detection method based on multi-feature extraction
By combining Word2Vec and FastText word vectorization methods and weighted processing of L-IDF algorithms, the weighted feature matrix is generated, which solves the problem of inaccurate feature extraction in cross-site script attack detection, and significantly improves the detection accuracy and accuracy.
Patent Information
- Application Number
- CN202510391112.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-06-20
AI Technical Summary
The prior art has problems such as inaccurate feature extraction, neglecting word shape changes and insufficient low-frequency word processing capabilities in cross-site script attack detection, which affects the accuracy and efficiency of detection.
A method based on multi-feature extraction is adopted, combined with Word2Vec and FastText word vectorization methods to generate word vectors, and weighted processing is performed through the L-IDF algorithm to generate a weighted feature matrix as input to the deep learning model.
It significantly improves the accuracy, accuracy and robustness of detection, and can extract key information and semantics more accurately, improving the performance of the model.
Smart Images

Figure CN120185898A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an XSS attack detection method based on multi-feature extraction, belonging to the field of network security technology. Background Art
[0002] In the task of cross-site scripting attack detection, it is crucial to vectorize data and extract features. Vectorization and feature extraction can transform the original text data into a numerical form that computers can understand and process, providing input for subsequent machine learning and deep learning algorithms. However, there are many problems with traditional word embedding methods and feature extraction. For example, although Word2Vec can map words to a low-dimensional vector space and show the semantic relationships between words, its ability to handle out-of-vocabulary words and low-frequency words is limited, and it will ignore the morphological changes of words during the processing and cannot effectively handle these situations. Although FastText can handle out-of-vocabulary words and low-frequency words well and is more robust to word form changes, the generated word vectors are slightly insufficient in the accuracy of semantic information, and for longer text sequences, it may not be able to accurately extract key information and semantics. The TF-IDF algorithm does not consider the different distributions of words among various categories, resulting in the possible neglect of those words that play a significant role in category discrimination, thus affecting the classification effect of the model. These problems will affect the accuracy and integrity of feature representation in practical applications, reduce the performance of the model, and affect the accuracy and efficiency of cross-site scripting attack detection. Summary of the Invention
[0003] To solve the technical problems existing in the prior art, the present invention provides an XSS attack detection method based on multi-feature extraction that can significantly improve the detection accuracy, precision, and robustness.
[0004] To achieve the above object, the technical solution adopted by the present invention is an XSS attack detection method based on multi-feature extraction, which is specifically operated according to the following steps. a. For cross-site scripting detection, a data set containing real cross-site scripting attack samples is constructed. b. The data samples in the data set are decoded, cleaned, normalized, and processed, and then the processed data is tokenized through a regular expression matching algorithm to obtain a target data set. c. The Word2Vec and FastText word vectorization methods are used to vectorize the target data set, and then the data vectorized by the two methods is concatenated to obtain a word vector matrix; at the same time, the L-IDF algorithm is used to calculate the class frequency variance of the target data set to identify keywords with high distinctiveness, obtaining a calculation weight matrix, and the calculation weight matrix is weighted with the word vector matrix to obtain a weighted feature matrix. d. Input the weighted feature matrix into the deep learning model to perform XSS attack detection.
[0005] Preferably, in step c, the calculation formula of the L-IDF algorithm is as follows: , where is , is the class frequency variance of the word ; The calculation formula of is as follows: The specific calculation formula of the class frequency variance is as follows: , is the class frequency variance of the word , is the number of text categories, is the entire text library that contains the word in the number of documents, is in the category that contains the word in the number of documents.
[0006] Compared with the prior art, the present invention has the following technical effects: The XSS detection model based on multi-feature extraction of the present invention generates word vectors through FastText and Word2Vec, and combines the L-IDF algorithm weighting mechanism, which improves the attention to key features and significantly improves the detection accuracy, precision and robustness. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 is a schematic diagram of the detection process of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0008] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0009] As Figure 1 shown, an XSS attack detection method based on multi-feature extraction is specifically operated according to the following steps: a. For cross-site scripting detection, a data set containing real cross-site scripting attack samples is constructed; b. Perform data decoding, data cleaning, data standardization and processing on the data samples in the data set, and then perform word segmentation processing on the processed data through a regular expression matching algorithm to obtain a target data set; c. The Word2Vec and FastText word vectorization methods are used to vectorize the target data set, and then the data vectorized by the two methods are concatenated to obtain a word vector matrix. At the same time, the L-IDF algorithm is used to calculate the class frequency variance of the target data set to identify keywords with high recognition, obtaining a calculated weight matrix. The calculated weight matrix and the word vector matrix are weighted to obtain a weighted feature matrix. d. The weighted feature matrix is input into a deep learning model to detect XSS attacks.
[0010] Among them, in step c, the calculation formula of the L-IDF algorithm is as follows: , where is , is the class frequency variance of the word . The calculation formula of is as follows: The specific calculation formula of the class frequency variance is as follows: , is the class frequency variance of the word , is the number of text categories, is the entire text library that contains the word in the document count, is in the category that contains the word in the document count, The larger is, the greater the fluctuation of the word
[0011] In the present invention, for cross-site scripting detection, a data set containing real cross-site scripting attack samples is constructed. The data set is divided into two parts, malicious cross-site scripting and normal requests. Among them, the malicious sample part is sourced from XSS payload samples uploaded by individuals on Github and part is crawled from the XSSed website. XSSed is currently the largest online database specifically for collecting cross-site scripting attacks. The normal samples are mainly collected from the open directory website DMOZ, which covers web page information from various fields. Through multi-source sample collection, the diversity and authenticity of the data set in terms of content are ensured, thus providing more comprehensive and representative training data for the XSS detection model and enhancing the generalization ability and accuracy of the model.
[0012] In the dataset collection phase, sample data with XSS attacks and sample data of normal requests with noise were collected from XSSed websites, DMOZ websites, and Github. Due to the diversity and inconsistency of data sources, there are inevitably dirty data, redundant data, and invalid data during the collection process. These non-standard data may have a negative impact on the training effect and prediction accuracy of the model. Therefore, the collected data will be strictly cleaned to improve the quality of the dataset and lay a solid foundation for subsequent analysis and model construction.
[0013] Attackers often obfuscate cross-site scripting (XSS) data through various encoding means to avoid being recognized by detection systems. These encoding techniques are usually not used alone, and attackers often combine multiple methods, making XSS scripts more complex and difficult to understand, and may interfere with the normal parsing process of browsers. Due to the superposition of these encoding methods, the readability of attack scripts is significantly reduced, and traditional detection tools are difficult to accurately identify their malicious behaviors. Common encoding methods include URL encoding, Base64 encoding, HTML entity encoding, and Unicode encoding, etc. To effectively analyze and identify these encoded XSS attack scripts, they usually need to be decoded first. Table 3-1 shows the original cross-site scripting data under different encoding methods and their decoded sample data. Through decoding, the structure and content of attack scripts become clearer, effectively improving the parsability of the data and laying a foundation for further security analysis and detection.
[0014] In the dataset collection phase, sample data with XSS attacks and sample data of normal requests with noise were collected from XSSed websites, DMOZ websites, and Github. Due to the diversity and inconsistency of data sources, there are inevitably dirty data, redundant data, and invalid data during the collection process. These non-standard data may have a negative impact on the training effect and prediction accuracy of the model. Therefore, the collected data will be strictly cleaned to improve the quality of the dataset and lay a solid foundation for subsequent analysis and model construction.
[0015] After data cleaning, the cross-site scripting attack samples already have good readability and structured features, but still contain a large number of complex elements such as HTML tags, JavaScript events, and URL links. Although the syntax and structure of these elements have become more standardized, they are still difficult to directly use for feature extraction and model training. Therefore, it is necessary to further tokenize these data. The main purpose of tokenization is to decompose the original cross-site scripting data into small units with semantic meanings, so as to more accurately analyze the components of the script and improve the accuracy and comprehensiveness of feature extraction. Therefore, a regular expression matching algorithm is used to tokenize the standardized data.
[0016] After the data processing is completed, the semantic information of the text is captured and the low-frequency words and word inflections are effectively processed through two word vectorization methods, Word2Vec and FastText. The word vectors generated by Word2Vec and FastText are concatenated to integrate the advantages of both and enhance the expressive power of the word vectors. At the same time, the L-IDF algorithm is introduced to further extract features. L-IDF combines the frequency of a word in a document and the inverse document frequency in the entire corpus, and introduces a class frequency variance mechanism, which can effectively identify keywords with high distinctiveness. Finally, the weighted feature matrix is used as the input of the deep learning model.
[0017] The preprocessed text data is modeled using the gensim.models.Word2Vec library and gensim.models.FastText library in Python to train the Word2Vec model and FastText model respectively. The Word2Vec model adopts the Skip-gram method to generate a low-dimensional vector representation of each word by predicting its context words given a central word. The goal of the model is to predict the neighboring words of a central word within a specified context window. The FastText model disassembles each word into several sub-words, which enables FastText to handle out-of-vocabulary words, such as being more robust to misspelled words or new words. This makes the FastText model perform better in dealing with some language tasks with low-frequency words or large word inflections.
[0018] After the training is completed, the Word2Vec and FastText models generate word vectors for each word. To better provide inputs for subsequent deep learning models, we concatenate the word vectors generated by these two models to integrate the advantages of Word2Vec and FastText. In this way, the representation of each word not only inherits the semantic information capture ability of Word2Vec but also combines the advantages of FastText in dealing with low-frequency words and word form variations, thus enhancing the expressive power of word vectors. In addition, to further enhance the feature expression ability of the model, the L-IDF algorithm is used to extract features. The L-IDF algorithm combines the frequency of a word in a document and its inverse document frequency in the entire corpus and introduces a class frequency variance mechanism, enabling it to effectively identify keywords with high distinctiveness in XSS attacks and normal requests. Finally, the obtained word vector matrix is multiplied by the weight matrix obtained by L-IDF to obtain a weighted feature matrix. By integrating the advantages of multiple feature extraction methods, the weighted feature matrix can capture the deep semantic information of the text, enhance the processing ability for low-frequency and inflected words, and highlight keywords with high distinctiveness in different categories, thus fully mining potential features. This weighted feature matrix will be used as the input for the deep learning model, which can effectively detect performance.
[0019] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the scope of the present invention.
Claims
1. A XSS attack detection method based on multi-feature extraction, characterized in that: Please follow the steps below to do this: a. For cross-site scripting detection, a dataset containing real cross-site scripting attack samples was constructed; b. Decode, clean, standardize and process the data samples in the data set, and then segment the processed data using a regular expression matching algorithm to obtain the target data set; c. The Word2Vec and FastText word vectorization methods are used to vectorize the target data set, and then the data after word vectorization by the two methods are spliced to obtain the word vector matrix; at the same time, the L-IDF algorithm is used to calculate the class frequency variance of the target data set to identify keywords with high recognition, obtain the calculation weight matrix, and perform weighted processing on the calculation weight matrix and the word vector matrix to obtain the weighted feature matrix; d. Input the weighted feature matrix into the deep learning model to detect XSS attacks.
2. According to claim 1, a XSS attack detection method based on multi-feature extraction is characterized in that: In step c, the L-IDF algorithm calculation formula is as follows: ,in, for , For words The class-frequency variance of The calculation formula is as follows: , The specific calculation formula of class frequency variance is as follows: , For words The class-frequency variance of is the number of text categories, For the entire text library Contains words The number of documents, For the category Contains words The number of documents.