Enhancing the effectiveness of firewall applications with machine learning techniques
The integration of machine learning algorithms in WAFs addresses false positives and new attack vectors, enhancing WAF efficiency by minimizing processing power and updating rules in real-time.
Patent Information
- Application Number
- PCT/TR2024/050171
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-29
- Publication Date
- 2025-07-03
AI Technical Summary
Traditional rule-based Web Application Firewalls (WAFs) produce false positives due to outdated or incorrect rule sets, leading to performance issues and administrative challenges, and deep learning methods require significant processing power and time, increasing costs.
Integrate machine learning algorithms like LightGBM and Logistic Regression for real-time anomaly detection, reducing false positives and identifying new attack vectors, while using simpler structures to minimize processing power requirements.
Reduces false positives and enhances WAF efficiency by updating rules in real-time, improving access performance and reducing computational costs.
Smart Images

Figure TR2024050171_03072025_PF_FP_ABST
Abstract
Description
[0001] Enhancing the Effectiveness of Firewall Applications with Machine Learning Techniques
[0002] Technical Field
[0003] This invention relates to the enhancement of the efficiency of traditional rule -based WAFs (Web Application Firewalls).
[0004] Prior Art
[0005] To protect against web application attacks, there is a need for a solution capable of analyzing HTTP / HTTPS traffic and proactively identifying potential threats. Web Application Firewalls (WAFs) are effective solutions that ensure the security of web applications by preventing various complex attack types, including Cross-Site Scripting (XSS), SQL (Structured Query Language) Injection, Remote Code Execution (RCE), and many others. Upon examination of their working principle, it is evident that commonly used WAF solutions are based on a set of rules. A rulebased WAF blocks attack traffic in accordance with specifically determined rules tailored to the web application.
[0006] If the rule sets defined in WAFs (Web Application Firewalls) are not current or are incorrect, it can adversely affect the performance of the web application, resulting in false positives - these are instances where requests are blocked due to erroneous rules defined in the WAF, even though they are not actual attacks. Additionally, the general default static configurations of WAFs do not suit the application portfolio of every organization. Therefore, application- specific exceptions are defined in the rule sets. Considering that organizations may have hundreds of applications, it is observed that while rule-based WAF solutions provide effective protection, they also have administrative challenges. In the current technique, the 'deep learning' method is used to increase the effectiveness of firewalls. Deep learning models require numerous computations due to their complex structures. A large number of computations necessitate significant processing power and time for training and anomaly prediction, which increases costs.
[0007] In the known technique, since the static WAF configurations used do not fit the application portfolio of every organization, requests that should not be blocked are being blocked by WAF rules, leading to the production of false positives.
[0008] Objective of the Invention
[0009] The objective of this invention is to obtain a system / method that performs real-time anomaly detection in web traffic, thereby reducing the frequent false positives generated by traditional rule-based WAFs (Web Application Firewalls).
[0010] Another objective of the invention is to obtain a system / method that integrates machine learning algorithms (like Lightgbm, Logistic Regression) for real-time web traffic anomaly detection, reduction of false positives, and effective identification of new attack vectors, thereby enhancing the efficiency of WAFs.
[0011] Another objective of the invention is to obtain a system / method that reduces processing power and thereby improves access performance to web applications.
[0012] The system developed to achieve these objectives includes:
[0013] A Network (100) that allows multiple devices to communicate with each other either wirelessly or through cables, and where users make requests over HTTP / HTTPS protocols, sharing hardware and software resources. WAF (200) designed to protect against potential attacks on web applications by analyzing HTTP / HTTPS web traffic of normal requests (101) and attack requests (102).
[0014] A data collection module (300) where allowed requests (201), blocked requests (202), and false positive requests (203) are accumulated for training machine learning models.
[0015] A data pre-processing module (400), where data filtering, converting to lowercase, stemming words, and other text mining techniques, followed by natural language processing techniques, are applied to allowed request (401), blocked requests (402), and false positive requests (403) by the WAF
[0016] (200).
[0017] A model training module (500), where the best model obtained (506) is integrated into the false positive testing module (600) to analyze live incoming request logs and identify rules that generate false positives.
[0018] A false positive testing module (600) which contains a false positive rule list (602) and is capable of request anomaly prediction (601) using blocked requests (204), false positive requests (205), and the best model (506).
[0019] Explanation of the Figures
[0020] The attached Figure 1 is a work flow chart describing the operation of the system.
[0021] The main elements expressed in the figures are given below as numbers and names.
[0022] (100) Network
[0023] (101) Normal request
[0024] (102) Attack request
[0025] (103) Normal request
[0026] (104) Attack request
[0027] (105) Normal request
[0028] (200) WAF
[0029] (201) Allowed request (202) Blocked request
[0030] (203) False positive request
[0031] (204) Blocked request
[0032] (205) False positive request
[0033] (206) Rule configuration file
[0034] (300) Data collection module
[0035] (301) Allowed request
[0036] (302) Blocked request
[0037] (303) False positive request
[0038] (400) Data Pre-processing module
[0039] (401) Allowed request
[0040] (402) Blocked request
[0041] (403) False positive request
[0042] (500) Model training module
[0043] (501) Vectorization
[0044] (502) Training data
[0045] (503) LightGBM model
[0046] (504) Logistic Regression model
[0047] (505) Test data
[0048] (506) Best model
[0049] (600) False Positive Testing Module
[0050] (601) Request Anomaly Prediction
[0051] (602) False positive rule list
[0052] Detailed Explanation of the Invention
[0053] The invention concerns increasing the efficiency of the traditional rule-based WAF (200) (Web Application Firewall - Web Application Firewall) by detecting anomalies in web traffic in real time, thus reducing the frequently produced false positives, integrating machine learning algorithms for real-time web traffic anomaly detection, reduction of false positives, and effective detection of new attack vectors, thereby enhancing the efficiency of WAF (200).
[0054] In the system of the invention, machine learning algorithms have simpler structures than “deep learning” algorithms. Due to their simpler structure, they use less hardware resources, resulting in an algorithm that is more advantageous in terms of cost and time compared to “deep learning “algorithms.
[0055] The best model (506) obtained with the Model Training Module (500) is integrated into the False Positive Testing Module (600) and analyzes live incoming request logs to identify rules that generate false positives. The results obtained from detection are recorded in the False Positive Rule List (602). Then, these requests are updated in the rule configuration file (206) in the WAF (200), thus preventing the production of false positives. In this way, the access performance to web applications is increased.
[0056] Allowed requests (201), blocked requests (202), and false positive requests (203) are collected in the data collection module (300) for training machine learning models. Other blocked requests (204) and false positive requests (205) coming live from the Network (100) are directed to the False Positive Testing Module (600) for anomaly detection with the trained best model (506). Here, the rules that false positive requests are caught by are kept in the False Positive Rule List (602). Then, in the WAF (200), these rules that generate false positives in the rule configuration file (206) are updated to prevent their production.
[0057] Normal requests (101) coming from the Network (100) to web applications, if not caught by the rules defined in the WAF (200), are allowed to pass to the web applications. These requests are named as allowed requests (201). Even though there is an attack request (102) in the Network (100), requests (202) can be blocked if caught by the rules in the rule configuration file (206) in the WAF (200).
[0058] Even though there is a normal request (103) in the Network (100), it can be blocked by the WAF (200), resulting in false positive requests (203).
[0059] Allowed requests (201), blocked requests (202), and false positive requests (203) are collected in the data collection module (300).
[0060] Inputs to the Data Pre-processing Module (400) (indicated by arrows 7, 8, and 9), which are allowed requests (401), blocked requests (402), and false positive requests (403) from the WAF (200), are subjected to text mining techniques such as data filtering, converting to lowercase, word stemming, and then natural language processing techniques. The data subjected to pre-processing activities is then sent to the Model Training Module (500) for vectorization.
[0061] The input data in the Model Training Module (500) are first subjected to vectorization (501) using the TF-IDF (Term Frequency-Inverse Document Frequency) method, converting from word format to numeric format for processing by machine learning models. The numeric format data is processed as training data (502). Training activities start with the processing of the training data (502).
[0062] Using the separated training data (502), training of the EightGBM model (503) and Eogistic Regression models (504) is conducted.
[0063] EightGBM is an open-source machine learning library that is one of the Gradient Boosting algorithms and is particularly effective on large datasets. Gradient Boosting, which uses decision trees, is a commonly used learning technique in the field of machine learning. Logistic Regression is a statistical method used to model the probability of two possible outcomes of a dependent variable. It is used to predict the likelihood of an event occurring in the presence of one or more independent variables.
[0064] After the training activities, accuracy scores are obtained using test data (505) not previously seen by the models to analyze their performance. The model with the highest score is recorded as the best model (506).
[0065] The best model (506) obtained in the Model Training Module (500) is sent to the False Positive Testing Module (600). Attack requests (104) and normal requests (105) coming from the live network (100) are sent to the WAF (200).
[0066] Blocked requests (204) and false positive requests (205), which are normal yet blocked in the WAF (200), are sent to the False Positive Testing Module (600). Request anomaly predictions (601) are made through the best model (506) obtained in the Model Training Module (500).
[0067] Requests falling below a certain threshold value are considered as false positives. The rules causing false positives, determined by the Request Anomaly Prediction (601) process, are sent to the False Positive Rule List (602).
[0068] The rules recorded in the False Positive Rule List (602) are taken into consideration and transmitted to the rule configuration file (206) found in the WAF (200). In the WAF (200), the rules that cause the production of false positives, as defined in the rule configuration file (206), are updated to prevent the production of false positives.
[0069] The Network (100) is a system that enables multiple devices to communicate with each other either wired or wirelessly, and allows users to make requests over HTTP / HTTPS protocols, thereby sharing hardware and software resources. Normal request (101) refers to normal requests coming from the Network (100) that reach the web application without being caught by any attack rule in the WAF (200).
[0070] Attack request (102) refers to requests coming from the Network (100) to the WAF (200) and are blocked by the WAF (200), to be used in the training data set.
[0071] Normal request (103) refers to requests coming from the Network (100) to the WAF (200) and, although not an attack, are blocked by the WAF (200) and identified as false positives.
[0072] Attack request (104) refers to requests coming from the Network (100) to the WAF (200), which are blocked by the WAF (200), and the anomaly score will be calculated by the trained model.
[0073] Normal request (105) refers to requests coming from the Network (100) to the WAF (200), which are blocked by the WAF (200), and the anomaly score will be calculated by the trained model.
[0074] WAF (200) is a type of security system designed to protect against potential attacks on web applications by analyzing HTTP / HTTPS web traffic of normal requests (101) and attack requests (102) coming from the Network (100).
[0075] Allowed request (201) refers to requests that are allowed to pass to the applications by the WAF (200) and are to be used in model training.
[0076] Blocked request (202) refers to attack requests that are blocked by the WAF (200) and are to be used in model training.
[0077] False positive request (203) refers to normal requests that are blocked by the WAF (200) and are to be used in model training as false positives. Blocked request (204) refers to attack requests that are blocked by the WAF (200) and on which anomaly prediction will be performed by the trained machine learning model.
[0078] False positive request (205) refers to normal requests that are blocked by the WAF (200) and on which anomaly prediction will be performed by the trained machine learning model.
[0079] Data collection module (300) is the place where allowed requests (201), blocked requests (202), and false positive requests (203) are collected for training machine learning models.
[0080] Allowed request (301) refers to requests that are allowed to pass to the applications by the WAF (200) and are to be used in model training.
[0081] Blocked request (302) refers to attack requests that are blocked by the WAF (200) and are to be used in model training.
[0082] False positive request (303) refers to normal requests that are blocked by the WAF (200) and are to be used in model training as false positives.
[0083] Data pre-processing module (400) is the place where text mining techniques such as data filtering, converting to lowercase, word stemming, and then natural language processing techniques are applied to the allowed requests (401), blocked requests (402), and false positive requests (403) by the WAF (200).
[0084] Allowed request (401) refers to requests that are allowed to pass to the applications by the WAF (200) and are to be used in model training. Blocked request (402) refers to attack requests that are blocked by the WAF (200) and are to be used in model training.
[0085] False positive request (403) refers to normal requests that are blocked by the WAF (200) and are to be used in model training as false positives.
[0086] Model training module (500) ensures that the best model (506) obtained is integrated into the false positive testing module (600) to analyze live incoming request logs and identify rules that generate false positives.
[0087] Vectorization (501) ensures the conversion of data coming from the data preprocessing module (400) into a numeric format for training artificial intelligence models.
[0088] Training data (502) consists of a part of the current request data and is used as training data (502).
[0089] LightGBM model (503) is where the anomaly score prediction is made using request data.
[0090] Logistic regression model (504) is where the anomaly score prediction is made using request data.
[0091] Test data (505) is used as a part of the current request data.
[0092] The best model (506) is the model with the highest accuracy score in the test results. It is also used as input to the false positive testing module (600).
[0093] False positive testing module (600) contains the false positive rule list (602) and enables the prediction of request anomalies (601) using blocked requests (204), false positive requests (205), and the best model (506) coming from the WAF (200). Request anomaly prediction (601) calculates the anomaly score of live requests coming from the Network (100) to the WAF (200) and then to the false positive testing module (600), identifying the rules that generate false positives.
[0094] The system of the invention operates with the following workflow comprises these operational steps;
[0095] Transmission of a normal request (101) from the Network (100) to the WAF (200) and the request allowed by the WAF (200) (SI),
[0096] Transmission of an attack request (102) from the Network (100) to the WAF (200) and the request blocked by the WAF (200) (S2),
[0097] Transmission of a normal request (103) from the Network (100) to the WAF (200) and the request blocked by the WAF (200) (S3),
[0098] Transmission of the allowed request (201) from the WAF (200) to the data collection module (300) (S4),
[0099] Transmission of the blocked request (202) from the WAF (200) to the data collection module (300) (S5),
[0100] Transmission of the false positive request (203) from the WAF (200) to the data collection module (300) (S6),
[0101] Transmission of the allowed requests (401) from the data collection module (300) to the data pre-processing module (400) (S7),
[0102] Transmission of the blocked requests (402) from the data collection module (300) to the data preprocessing module (400) (S8),
[0103] Transmission of the false positive requests (403) from the data collection module (300) to the data preprocessing module (400) (S9),
[0104] Transmission of the allowed requests (401) from the data pre-processing module (400) to the model training module (500) after applying natural language processing techniques (S10),
[0105] Transmission of the blocked requests (402) from the data pre-processing module (400) to the model training module (500) after applying natural language processing techniques (Sil), Transmission of the false positive requests (403) from the data preprocessing module (400) to the model training module (500) after applying natural language processing techniques (S12),
[0106] Separation and transmission of vectorized data (501) as training data (502)
[0107] (S13),
[0108] Transmission of the separated training data (502) for training the LightGBM model (503) and Logistic Regression models (504) (S14),
[0109] Submission of the trained LightGBM model (503) for calculating the accuracy score on the test data (505) (S15),
[0110] Submission of the trained Logistic Regression Model (504) for calculating the accuracy score on the test data (505) (S16),
[0111] Transmission of the model with the highest accuracy score using the test data (505) to the best model (506) (S17),
[0112] Transmission of the best model (506) for request anomaly prediction (601) to the false positive testing module (600) (S18),
[0113] Transmission of the attack request (104) and the blocked request (204) blocked by the WAF (200) from the Network (100) to the WAF (200) (S19), Transmission of the normal request (105) and the false positive request (205) allowed by the WAF (200) from the Network (100) to the WAF (200)
[0114] (S20),
[0115] Transmission of the blocked requests (204) from the WAF (200) to the false positive testing module (600) for request anomaly prediction (601) by the best model (506) (S21),
[0116] Transmission of the false positive requests (205) from the WAF (200) to the false positive testing module (600) for request anomaly prediction (601) by the best model (506) (S22),
[0117] Transmission of the rules identified as false positives by the request anomaly prediction (601) to the false positive rule list (602) (S23),
[0118] Sending the false positive rule list (602) to the rule configuration file (206) in the WAF (200) for updating the defined rules in the WAF (200) (S24).
Claims
C L A I M S1. A system related to enhancing the efficiency of traditional rule-based WAFs, characterized in that; including,A network (100) that enables multiple devices to communicate with each other either wired or wirelessly, allowing users to make requests over HTTP / HTTPS protocols, thereby sharing hardware and software resources,A web application firewall (WAF) (200) designed to protect web applications against potential attacks by analysing HTTP / HTTPS web traffic of normal requests (101) and attack requests (102) coming from the Network (100),A data collection module (300) where allowed requests (201), blocked requests (202), and false positive requests (203) are accumulated for training machine learning models,A data pre-processing module (400), where data filtering, converting to lowercase, word stemming, and other text mining techniques, followed by natural language processing techniques, are applied to allowed requests (401), blocked requests (402), and false positive requests (403) permitted by the WAF (200),A model training module (500) that integrates the best model obtained (506) into the False Positive Detection Module (600) to analyze live incoming request logs and identify rules that generate false positives,A False Positive Testing Module (600) which contains a false positive rule list (602) and is capable of predicting request anomalies (601) using blocked requests (204), false positive requests (205), and the best model (506) coming from the WAF (200).
2. The Network (100) mentioned in claim 1 and characterized in that; including,A normal request (101) coming from the Network (100), which reaches the web application without being intercepted by any attack rule in the WAF (200),An attack request (102) originating from the Network (100) and blocked by the WAF (200),A normal request (103) coming from the Network (100) to the WAF (200), which, despite not being an attack, is identified as a false positive and blocked by the WAF (200),An attack request (104) coming from the Network (100) to the WAF (200), which will be evaluated for an anomaly score by the trained model and is blocked by the WAF (200),A normal request (105) coming from the Network (100) to the WAF (200), which will also be evaluated for an anomaly score by the trained model and is blocked by the WAF (200).
3. The WAF (200) mentioned in claim 1 and characterized in that; including,Allowed requests (201) which are permitted by the WAF (200) to pass to applications and will be used in model training,Blocked requests (202) which are attacks and blocked by the WAF (200), to be used in model training,False positive requests (203) which are normal but blocked by the WAF (200) and will be used in model training,Blocked requests (204) which are attacks and blocked by the WAF (200), on which anomaly prediction will be performed by the trained machine learning model, andFalse positive requests (205) which are normal but blocked by the WAF (200), on which anomaly prediction will be performed by the trained machine learning model.
4. The data collection module (300) mentioned in claim 1 and characterized in that; including,Allowed requests (301) which are permitted by the WAF (200) to pass to applications and will be used in model training,Blocked requests (302) which are attacks and blocked by the WAF (200), to be used in model training, andFalse positive requests (303) which are normal but blocked by the WAF (200) and will be used in model training.
5. The data pre-processing module (400) mentioned in claim 1 and characterized in that; including,The requests that are allowed to pass to applications by WAF (200) and the allowed request (401) to be used in model training,Blocked request (402), which is an attack and blocked by WAF (200) and will be used in model training, andFalse positive request (403) which, albeit normal, is blocked by WAF (200) and will be used in model training.
6. The model training module (500) mentioned in claim 1 and characterized in that; it includes:Vectorization (501) which converts data coming from the Data Pre-processingModule (400) into a numeric format for the training of artificial intelligence models,Training data (502) which forms a part of the current request data and is used as training data (502),The LightGBM model (503), a trained model using request data for predicting anomaly scores for requests,The logistic regression model (504), a trained model using request data for predicting anomaly scores for requests,Test data (505) used as a part of the current request data for calculating the accuracy scores of the trained models, andThe best model (506) with the highest accuracy score in the test results, used as input for the false positive testing module (600).
7. The false positive testing module (600) mentioned in claim 1 and characterized in that; it includes,Request anomaly prediction (601) that enables the determination of rules generating false positives by calculating the anomaly score of live requests coming from the network (100) to the WAF (200) and then to the false positive detection module (600) from the live web traffic, andThe false positive rule list (602) which modifies the identified false positive rules and defines them in the rule configuration file in the WAF (200) to prevent the production of false positives.
8. The invention is a method that enhances the efficiency of traditional rule-based WAFs, characterized in that; including the following process steps,Transmission of a normal request (101) from the Network (100) to the WAF (200) and the transmission of the request allowed by the WAF (200) (SI),Transmission of an attack request (102) from the Network (100) to the WAF (200) and the transmission of the request blocked by the WAF (200) (S2),Transmission of a normal request (103) from the Network (100) to the WAF (200) and the transmission of the request blocked by the WAF (200) (S3),Transmission of the allowed request (201) from the WAF (200) to the data collection module (300) (S4),Transmission of the blocked request (202) from the WAF (200) to the data collection module (300) (S5),Transmission of the False Positive request (203) from the WAF (200) to the data collection module (300) (S6),Transmission of the allowed requests (401) from the data collection module (300) to the data pre-processing module (400) (S7),Transmission of the blocked requests (402) from the data collection module (300) to the data pre-processing module (400) (S8),Transmission of the false positive requests (403) from the data collection module (300) to the data pre-processing module (400) (S9),Transmission of the allowed requests (401) from the data pre-processing module (400) to the model training module (500) after applying natural language processing techniques (S10),Transmission of the blocked requests (402) from the data pre-processing module (400) to the model training module (500) after applying natural language processing techniques (Sil),Transmission of the false positive requests (403) from the data pre-processing module (400) to the model training module (500) after applying natural language processing techniques (S12),Transmission of vectorized data (501), separated as training data (502) (S13), Transmission of the training data (502) for training the LightGBM model (503) and Logistic Regression models (504) (S14),Submission of the trained LightGBM model (503) to the test data (505) for calculating accuracy scores (S15),Submission of the trained logistic regression model (504) to the test data (505) for calculating accuracy scores (S16),Transmission of the model with the highest accuracy score using the test data (505) to the best model (506) (S17),Transmission of the best model (506) for request anomaly prediction (601) to the false positive testing module (600) (S18),Transmission of the attack request (104) and the blocked request (204) blocked by the WAF (200) from the Network (100) to the WAF (200) (S19),Transmission of the normal request (105) and the false positive request (205) allowed by the WAF (200) from the Network (100) to the WAF (200) (S20), Transmission of the Blocked requests (204) from the WAF (200) to the FalsePositive testing module (600) for Request Anomaly Prediction (601) by the Best model (506) (S21),Transmission of the false positive requests (205) from the WAF (200) to the false positive testing module (600) for request anomaly prediction (601) by the Best model (506) (S22),Transmission of the rules identified as false positives by the Request Anomaly Prediction (601) to the false positive rule list (602) (S23),Sending the false positive rule list (602) to the rule configuration file (206) in the WAF (200) for updating the defined rules in the WAF (200) (S24).
Citation Information
Patent Citations
Utilizing web applicaton firewall and machine learning to detect command and control
US20210377295A1
Layer 7 network attack detection using machine learning feature contribution
US20230135755A1
Attack categorization based on machine learning feature contribution
US20230144836A1
Content-based deep learning for inline phishing detection
US20230353595A1