Attack data identification method based on Flink
Through the Flink-based attack data identification method, the problem of lack of a real-time detection framework for reflective XSS attacks in the prior art is solved, real-time detection and response to XSS attacks are realized, and the security and detection efficiency of web applications are improved.
Patent Information
- Application Number
- CN202411955668.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-28
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art lacks a dedicated real-time detection framework for reflective XSS attacks, resulting in the inability to respond immediately when an attack occurs.
Using Flink-based attack data identification method, real-time detection and response to XSS attacks are achieved through the steps of data collection and preprocessing, feature engineering, design and training of LSTM models, real-time detection and response.
By relying on Flink's stream processing advantages, real-time detection and response to XSS attacks are achieved, the window period for security vulnerabilities is reduced, the security of web applications is improved, complex malicious program patterns are automatically identified, and the accuracy and efficiency of detection are improved.
Smart Images

Figure CN119989345A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of XSS attack data identification, and in particular to an attack data identification method based on Flink. Background Art
[0002] With the development of the Internet, network attacks are becoming more and more complex and frequent, especially reflected XSS attacks. XSS attacks are common security vulnerabilities in Web applications. Attackers implant malicious scripts to steal user information or perform malicious operations. Reflected XSS is one of the more common forms and needs to be identified and defended in real time.
[0003] Traditional detection methods are usually lagging and cannot respond immediately when an attack occurs. Flink, as an efficient big data processing engine, can process large-scale data and provide low-latency data stream analysis. Tools such as Apache Storm and Apache Kafka provide basic capabilities in real-time data transmission and processing, but have limitations in complex pattern recognition and real-time classification.
[0004] Machine learning, especially deep learning, has been gradually applied in intrusion detection systems to achieve better detection results by capturing complex patterns in data; however, most solutions remain in the offline analysis stage and have limited real-time processing capabilities;
[0005] That is, the prior art has the following technical problems: lack of a dedicated real-time detection framework for reflected XSS attacks. Therefore, a Flink-based attack data identification method is proposed to address the above problems. Summary of the invention
[0006] In this embodiment, a Flink-based attack data identification method is provided to solve the problem of lack of a dedicated real-time detection framework for reflective XSS attacks in the prior art.
[0007] According to one aspect of the present application, a Flink-based attack data identification method is provided, and the Flink-based attack data identification method includes the following steps:
[0008] (1) Data collection and preprocessing;
[0009] (2) Setting up a feature engineering module to extract and select effective features;
[0010] (3) Design and train LSTM models;
[0011] (4) Conduct real-time detection and response.
[0012] Furthermore, in the step (1), data collection and preprocessing are performed, data is extracted and cleaned from the Web log, and a data set is generated by parsing, cleaning, and formatting the log.
[0013] Furthermore, in step (1), data collection and preprocessing include the following steps:
[0014] a. Configure the log collector to collect web request logs from different servers;
[0015] b. Decode URL encoding and special characters in the logs;
[0016] c. Extract request information;
[0017] d. Use regular expressions to filter out potential attack features;
[0018] e. Remove useless fields and retain key data;
[0019] f. Clean the text data and remove HTML and script tags;
[0020] g. Handle missing values and fill them with interpolation and averaging strategies;
[0021] h. Parse date and timestamps and convert them into standard formats;
[0022] i. Prepare word segmentation for specific fields using natural language processing tools;
[0023] j. Convert the word segmentation results into vector representation;
[0024] k. Label and annotate known attack samples in the training set;
[0025] 1. End.
[0026] Furthermore, in step (2), a feature engineering module is set up to extract feature data and prepare training data.
[0027] Furthermore, in step (2), extracting and selecting effective features includes the following steps:
[0028] a. Extract keywords and phrases from request parameters;
[0029] b. Implement the bag-of-words model to construct the initial feature matrix;
[0030] c. Enriching text representation through word embedding;
[0031] d. Generate grammatical features of the request and analyze the user input framework;
[0032] e. Extract statistical features from the request;
[0033] f. Calculate time-related features;
[0034] g. Use PCA to reduce feature dimensionality and reduce redundancy;
[0035] h. Apply L1 regularization to select important features and enhance model sparsity;
[0036] i. Implement the chi-square test to select the features most relevant to the target variable;
[0037] j. Standardize numerical features to ensure consistency of model input;
[0038] k. Combine historical attack samples to enhance feature interpretability;
[0039] l. Perform cluster analysis on features to identify potential attack patterns;
[0040] m. Realize feature combination and construct new complex features;
[0041] n. Optimize feature selection strategy through incremental analysis;
[0042] o. Convert the final feature set into model input format;
[0043] p. End.
[0044] Furthermore, in step (3), an LSTM model is designed and trained to identify complex time dependencies in the sequence.
[0045] Furthermore, in step (3), designing and training the LSTM model includes the following steps:
[0046] a. Define the LSTM model architecture, including the input layer, multiple LSTM layers, and the output layer;
[0047] b. Set the number of units in the LSTM layer to optimize the model’s memory capacity;
[0048] c. Introduce bidirectional LSTM to improve the ability of sequence feature extraction;
[0049] d. Use Batch Normalization to stabilize model training;
[0050] e. Select the loss function and compile the model;
[0051] f. Use Adam optimizer to accelerate model convergence;
[0052] g. Divide the data set into training set and validation set to ensure the accuracy of model evaluation;
[0053] h. Perform multiple rounds of training on the training set;
[0054] i. Use the validation set to monitor model performance and adjust parameters in real time;
[0055] j. Use transfer learning to fine-tune the model to adapt it to new attack patterns;
[0056] k. Verify the performance of the model on the test set and evaluate the accuracy and recall;
[0057] 1. End.
[0058] Furthermore, in step (4), the trained LSTM model is deployed on Flink to perform stream data analysis, detect and respond to XSS attacks in real time. When a potential threat is detected, the system quickly takes measures to reduce the risk and ensure the security of the Web application.
[0059] Furthermore, in step (4), the real-time detection includes the following steps:
[0060] a. Import the LSTM model as a Flink stream processing task;
[0061] b. Use Flink DataStream API to access real-time network request streams;
[0062] c. Call the feature engineering module to process each data stream block;
[0063] d. The prediction results are output in real time, marking the request as normal or suspicious;
[0064] e. Set attack detection thresholds and classify request status;
[0065] f. Analyze complex attack patterns through Flink’s CEP module;
[0066] g. Implement a real-time alert system to notify administrators via Webhook or email;
[0067] h. Support blocking function to immediately block suspicious request sources;
[0068] i. Periodic model update: Regularly feed data back to the training module to update the model;
[0069] j. Adjust the detection strategy and change the threshold according to the latest attack characteristics;
[0070] k. Integrate API and link with existing security systems;
[0071] l. Test the system response capability through simulation exercises and optimize the system;
[0072] m. End.
[0073] Through the above technical solution of this application, relying on the stream processing advantages of Flink, it is possible to detect and respond to XSS attacks in real time, reduce the window period for exploiting security vulnerabilities, improve the security of Web applications, and automatically identify complex malicious program patterns to improve the accuracy and efficiency of detection without relying on traditional feature engineering. Furthermore, this application realizes automated detection and response, reduces manual intervention, improves overall defense efficiency, and reduces security management costs through real-time blocking and alarm functions. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0075] Figure 1 This is a schematic diagram of the overall process of an embodiment of the present application;
[0076] Figure 2 A schematic diagram of the data collection and preprocessing process of an embodiment of the present application;
[0077] Figure 3 A schematic diagram of a process of extracting and selecting effective features according to an embodiment of the present application;
[0078] Figure 4 A schematic diagram of the process of designing and training an LSTM model according to an embodiment of the present application;
[0079] Figure 5 A schematic diagram of the real-time detection process of an embodiment of the present application. DETAILED DESCRIPTION
[0080] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.
[0081] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application described here. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0082] In the present application, the terms "upper", "lower", "left", "right", "front", "back", "top", "bottom", "inner", "outer", "middle", "vertical", "horizontal", "lateral", "longitudinal" and the like indicate positions or positional relationships based on the positions or positional relationships shown in the drawings. These terms are mainly used to better describe the present application and its embodiments, and are not used to limit the indicated devices, elements or components to have a specific orientation, or to be constructed and operated in a specific orientation.
[0083] In addition, some of the above terms may be used to express other meanings in addition to indicating orientation or positional relationship. For example, the term "on" may also be used to express a certain dependency or connection relationship in some cases. For those of ordinary skill in the art, the specific meanings of these terms in this application can be understood according to specific circumstances.
[0084] In addition, the terms "installed", "set", "provided with", "connected", "connected", and "socketed" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral structure; it can be a mechanical connection, or an electrical connection; it can be a direct connection, or an indirect connection through an intermediate medium, or it can be an internal connection between two devices, elements, or components. For those of ordinary skill in the art, the specific meanings of the above terms in this application can be understood according to specific circumstances.
[0085] See also Figure 1-5 As shown, a Flink-based attack data identification method comprises the following steps:
[0086] (1) Data collection and preprocessing;
[0087] (2) Setting up a feature engineering module to extract and select effective features;
[0088] (3) Design and train LSTM models;
[0089] (4) Conduct real-time detection and response.
[0090] Through the above technical solutions, this application relies on Flink's stream processing advantages to detect and respond to XSS attacks in real time, reduce the window period for exploiting security vulnerabilities, improve the security of Web applications, and automatically identify complex malicious program patterns to improve the accuracy and efficiency of detection without relying on traditional feature engineering. Furthermore, this application realizes automated detection and response, reduces manual intervention, and improves overall defense efficiency and reduces security management costs through real-time blocking and alarm functions.
[0091] In the step (1), data collection and preprocessing are performed to extract and clean data from Web logs, and a data set is generated by parsing, cleaning, and formatting the logs;
[0092] In step (1), data collection and preprocessing include the following steps:
[0093] a. Configure the log collector to collect web request logs from different servers;
[0094] b. Decode URL encoding and special characters in the logs;
[0095] c. Extract request information;
[0096] d. Use regular expressions to filter out potential attack features;
[0097] e. Remove useless fields and retain key data;
[0098] f. Clean the text data and remove HTML and script tags;
[0099] g. Handle missing values and fill them with interpolation and averaging strategies;
[0100] h. Parse date and timestamps and convert them into standard formats;
[0101] i. Prepare word segmentation for specific fields using natural language processing tools;
[0102] j. Convert the word segmentation results into vector representation;
[0103] k. Label and annotate known attack samples in the training set;
[0104] 1. End.
[0105] In the step (2), a feature engineering module is set up to extract feature data and prepare training data;
[0106] In the step (2), extracting and selecting effective features comprises the following steps:
[0107] a. Extract keywords and phrases from request parameters;
[0108] b. Implement the bag-of-words model to construct the initial feature matrix;
[0109] c. Enriching text representation through word embedding;
[0110] d. Generate grammatical features of the request and analyze the user input framework;
[0111] e. Extract statistical features from the request;
[0112] f. Calculate time-related features;
[0113] g. Use PCA to reduce feature dimensionality and reduce redundancy;
[0114] h. Apply L1 regularization to select important features and enhance model sparsity;
[0115] i. Implement the chi-square test to select the features most relevant to the target variable;
[0116] j. Standardize numerical features to ensure consistency of model input;
[0117] k. Combine historical attack samples to enhance feature interpretability;
[0118] l. Perform cluster analysis on features to identify potential attack patterns;
[0119] m. Realize feature combination and construct new complex features;
[0120] n. Optimize feature selection strategy through incremental analysis;
[0121] o. Convert the final feature set into model input format;
[0122] p. End.
[0123] In the step (3), an LSTM model is designed and trained to identify complex time dependencies in the sequence;
[0124] In step (3), designing and training the LSTM model includes the following steps:
[0125] a. Define the LSTM model architecture, including the input layer, multiple LSTM layers, and the output layer;
[0126] b. Set the number of units in the LSTM layer to optimize the model’s memory capacity;
[0127] c. Introduce bidirectional LSTM to improve the ability of sequence feature extraction;
[0128] d. Use Batch Normalization to stabilize model training;
[0129] e. Select the loss function and compile the model;
[0130] f. Use Adam optimizer to accelerate model convergence;
[0131] g. Divide the data set into training set and validation set to ensure the accuracy of model evaluation;
[0132] h. Perform multiple rounds of training on the training set;
[0133] i. Use the validation set to monitor model performance and adjust parameters in real time;
[0134] j. Use transfer learning to fine-tune the model to adapt it to new attack patterns;
[0135] k. Verify the performance of the model on the test set and evaluate the accuracy and recall;
[0136] 1. End;
[0137] In step (4), the trained LSTM model is deployed on Flink to perform stream data analysis, detect and respond to XSS attacks in real time. When a potential threat is detected, the system quickly takes measures to reduce the risk and ensure the security of the Web application.
[0138] In step (4), the real-time detection includes the following steps:
[0139] a. Import the LSTM model as a Flink stream processing task;
[0140] b. Use Flink DataStream API to access real-time network request streams;
[0141] c. Call the feature engineering module to process each data stream block;
[0142] d. The prediction results are output in real time, marking the request as normal or suspicious;
[0143] e. Set attack detection thresholds and classify request status;
[0144] f. Analyze complex attack patterns through Flink’s CEP module;
[0145] g. Implement a real-time alert system to notify administrators via Webhook or email;
[0146] h. Support blocking function to immediately block suspicious request sources;
[0147] i. Periodic model update: Regularly feed data back to the training module to update the model;
[0148] j. Adjust the detection strategy and change the threshold according to the latest attack characteristics;
[0149] k. Integrate API and link with existing security systems;
[0150] l. Test the system response capability through simulation exercises and optimize the system;
[0151] m. End.
[0152] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A Flink-based attack data identification method, characterized by: The Flink-based attack data identification method includes the following steps: (1) Data collection and preprocessing; (2) Setting up a feature engineering module to extract and select effective features; (3) Design and train LSTM models; (4) Conduct real-time detection and response.
2. The attack data identification method based on Flink according to claim 1 is characterized in that: In the step (1), data collection and preprocessing are performed, data is extracted and cleaned from the Web log, and a data set is generated by parsing, cleaning, and formatting the log.
3. The attack data identification method based on Flink according to claim 1 is characterized in that: In step (1), data collection and preprocessing include the following steps: a. Configure the log collector to collect web request logs from different servers; b. Decode URL encoding and special characters in the logs; c. Extract request information; d. Use regular expressions to filter out potential attack features; e. Remove useless fields and retain key data; f. Clean the text data and remove HTML and script tags; g. Handle missing values and fill them with interpolation and averaging strategies; h. Parse date and timestamps and convert them into standard formats; i. Prepare word segmentation for specific fields using natural language processing tools; j. Convert the word segmentation results into vector representation; k. Label and annotate known attack samples in the training set; 1. End.
4. The attack data identification method based on Flink according to claim 1 is characterized in that: In the step (2), a feature engineering module is set up to extract feature data and prepare training data.
5. The attack data identification method based on Flink according to claim 1 is characterized in that: In step (2), extracting and selecting effective features includes the following steps: a. Extract keywords and phrases from request parameters; b. Implement the bag-of-words model to construct the initial feature matrix; c. Enriching text representation through word embedding; d. Generate grammatical features of the request and analyze the user input framework; e. Extract statistical features from the request; f. Calculate time-related features; g. Use PCA to reduce feature dimensionality and reduce redundancy; h. Apply L1 regularization to select important features and enhance model sparsity; i. Implement the chi-square test to select the features most relevant to the target variable; j. Standardize numerical features to ensure consistency of model input; k. Combine historical attack samples to enhance feature interpretability; l. Perform cluster analysis on features to identify potential attack patterns; m. Realize feature combination and construct new complex features; n. Optimize feature selection strategy through incremental analysis; o. Convert the final feature set into model input format; p. End.
6. The attack data identification method based on Flink according to claim 1 is characterized in that: In the step (3), an LSTM model is designed and trained to identify complex time dependencies in the sequence.
7. The attack data identification method based on Flink according to claim 1 is characterized in that: In step (3), designing and training the LSTM model includes the following steps: a. Define the LSTM model architecture, including the input layer, multiple LSTM layers, and the output layer; b. Set the number of units in the LSTM layer to optimize the model’s memory capacity; c. Introduce bidirectional LSTM to improve the ability of sequence feature extraction; d. Use Batch Normalization to stabilize model training; e. Select the loss function and compile the model; f. Use Adam optimizer to accelerate model convergence; g. Divide the data set into training set and validation set to ensure the accuracy of model evaluation; h. Perform multiple rounds of training on the training set; i. Use the validation set to monitor model performance and adjust parameters in real time; j. Use transfer learning to fine-tune the model to adapt it to new attack patterns; k. Verify the performance of the model on the test set and evaluate the accuracy and recall; 1. End.
8. The attack data identification method based on Flink according to claim 1 is characterized in that: In step (4), the trained LSTM model is deployed on Flink to perform stream data analysis, detect and respond to XSS attacks in real time. When a potential threat is detected, the system quickly takes measures to reduce the risk and ensure the security of the Web application.
9. The attack data identification method based on Flink according to claim 1 is characterized in that: In step (4), the real-time detection includes the following steps: a. Import the LSTM model as a Flink stream processing task; b. Use Flink DataStream API to access real-time network request streams; c. Call the feature engineering module to process each data stream block; d. The prediction results are output in real time, marking the request as normal or suspicious; e. Set attack detection thresholds and classify request status; f. Analyze complex attack patterns through Flink’s CEP module; g. Implement a real-time alert system to notify administrators via Webhook or email; h. Support blocking function to immediately block suspicious request sources; i. Periodic model update: Regularly feed data back to the training module to update the model; j. Adjust the detection strategy and change the threshold according to the latest attack characteristics; k. Integrate API and link with existing security systems; l. Test the system response capability through simulation exercises and optimize the system; m. End.