Traffic filtering method and device based on CC attack characteristics
By performing feature extraction and Laplace eigenmap dimensionality reduction on historical HTTP request traffic samples, combined with a deep penalty generative adversarial network and a multi-layer game model, the shortcomings of existing CC attack defense methods in detection accuracy and resource utilization are addressed, achieving efficient and flexible CC attack defense.
Patent Information
- Application Number
- CN202510847206.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-24
AI Technical Summary
Existing CC attack defense methods have low detection accuracy when facing complex and changeable attack behaviors, accidentally damaging normal user traffic, and are unable to effectively capture the temporal correlation patterns and behavioral characteristics in attack traffic, resulting in resource waste or insufficient defense.
By extracting features and performing Laplace eigenmap dimensionality reduction on historical HTTP request traffic samples, a comprehensive set of HTTP request feature vectors is constructed. A deep penalty-based generative adversarial network is used to analyze CC attack features. An attention mechanism and denoising penalty constraints are introduced to build a multi-layer game model for adaptive defense and achieve hierarchical filtering processing.
It improves the accuracy and flexibility of CC attack detection, optimizes resource utilization efficiency, and reduces the impact on normal users. It shows significant advantages, especially in low-frequency CC attacks and mixed Flash Crowd traffic environments.
Smart Images

Figure CN120358098B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of traffic filtering, and in particular to a traffic filtering method and device based on CC attack characteristics. Background Art
[0002] With the widespread use of internet applications and the rapid development of network services, CC (Challenge Collapsar) attacks, a highly concealed, low-bandwidth DDoS attack method, have posed a serious threat to various network services. CC attacks simulate normal user behavior, sending a large number of seemingly legitimate HTTP requests in a low-frequency, distributed manner, consuming server resources and preventing legitimate users from accessing the target website or service. Existing defense measures, such as request frequency limiting, IP blacklisting, and access frequency threshold control, primarily rely on static rules or simple statistical features for filtering. However, these methods have low detection accuracy when faced with complex and changing attack behaviors and are prone to accidentally impacting legitimate user traffic. This is particularly true in FlashCrowd (sudden high-traffic) scenarios, where false positive rates are high and it is difficult to distinguish malicious traffic from highly concurrent legitimate access.
[0003] As attackers upgrade their technology, new CC attack variants emerge one after another, such as low-frequency CC attacks, hybrid CC attacks, and distributed CC attacks. These attacks evade detection by adjusting request frequency, changing request parameters, or mixing in legitimate requests. Existing defense methods based on statistical features lack in-depth analysis of the correlation between HTTP requests and cannot effectively capture the temporal correlation patterns and behavioral characteristics in attack traffic. Traditional defense systems usually use static defense rules and cannot adaptively adjust defense strategies according to changes in attack intensity, resulting in excessive defense and waste of resources or insufficient defense and security risks. Summary of the Invention
[0004] The present invention provides a traffic filtering method and device based on CC attack characteristics. The present invention can dynamically adjust defense parameters according to the attack situation. Compared with static defense rules, it is more flexible and adaptable, and can optimize resource utilization efficiency while ensuring security.
[0005] In a first aspect, the present invention provides a traffic filtering method based on CC attack features, the traffic filtering method based on CC attack features comprising:
[0006] Perform feature extraction and Laplace eigenmap dimensionality reduction on historical HTTP request traffic samples to obtain a set of reduced-dimensional feature vectors;
[0007] Perform CC attack feature analysis based on the dimension-reduced feature vector set to obtain a CC attack feature parameter set;
[0008] Performing an attack risk assessment on the real-time HTTP request to be filtered based on the CC attack feature parameter set to obtain an attack possibility score;
[0009] Performing time-varying parameter defense analysis based on the attack likelihood score to obtain an adaptive defense rule set;
[0010] Perform hierarchical filtering processing on the real-time HTTP request to be filtered according to the adaptive defense rule set to obtain filtered safe traffic.
[0011] In a second aspect, the present invention provides a traffic filtering device based on CC attack features, the traffic filtering device based on CC attack features comprising:
[0012] The mapping and dimensionality reduction module is used to perform feature extraction and Laplace feature mapping dimensionality reduction on historical HTTP request traffic samples to obtain a set of reduced-dimensional feature vectors;
[0013] A feature analysis module, configured to perform CC attack feature analysis based on the dimension-reduced feature vector set to obtain a CC attack feature parameter set;
[0014] A risk assessment module, configured to perform an attack risk assessment on the real-time HTTP request to be filtered based on the CC attack feature parameter set to obtain an attack possibility score;
[0015] a defense analysis module, configured to perform time-varying parameter defense analysis based on the attack likelihood score to obtain an adaptive defense rule set;
[0016] The hierarchical filtering module is used to perform hierarchical filtering processing on the real-time HTTP requests to be filtered according to the adaptive defense rule set to obtain filtered safe traffic.
[0017] In the technical solution provided by the present invention, a comprehensive set of HTTP request feature vectors is constructed by multi-dimensionally extracting and fusing temporal features, session behavior features, and content features from historical HTTP request traffic samples. Compared with traditional methods based on only a single or a small number of statistical features, this method can more comprehensively characterize the characteristic patterns of HTTP requests. The present invention introduces Laplace eigenmaps to reduce the dimensionality of the HTTP request feature vector set, converting the original high-dimensional discrete features into a continuous feature space, effectively preserving the topological relationship and similarity structure between HTTP requests. This not only reduces computational complexity but also enhances the expressiveness of features, enabling the system to better capture temporal correlation patterns in CC attacks. The present invention uses a deep penalized generative adversarial network to perform CC attack feature analysis. Through adversarial learning between the generator and the discriminator, the characteristic patterns of CC attacks are automatically learned, which has stronger feature learning capabilities than traditional machine learning methods. At the same time, the introduced attention mechanism can automatically focus on the key features that distinguish CC attacks, improving the interpretability and accuracy of the model. The denoising penalty constraint introduced by the present invention constrains the gradient norm of the discriminator near the real data, so that the model output remains stable when the input undergoes small changes, significantly enhancing the model's adaptability to HTTP request variations and effectively dealing with attackers' attempts to evade detection by changing request parameters and adjusting request frequency. The multi-layer game model constructed by the present invention regards CC attack defense as a dynamic game process between attackers and defenders. By solving the optimal control equation, the optimal defense strategy is obtained, and the defense parameters can be dynamically adjusted according to the attack situation. Compared with static defense rules, it is more flexible and adaptable, and can optimize resource utilization efficiency while ensuring security. The hierarchical filtering processing mechanism implemented by the present invention processes requests at different levels according to the attack probability score. Combined with multiple defense measures such as JavaScript challenge verification, CAPTCHA verification, and dynamic resource allocation, it can adopt differentiated defense strategies for requests of different risk levels, ensuring system security while minimizing the impact on normal users. It has significant advantages in dealing with low-frequency CC attacks and mixed Flash Crowd traffic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0019] Figure 1 Schematic diagram of a method for filtering traffic based on CC attack features according to an embodiment of the present invention;
[0020] Figure 2 Schematic diagram of an embodiment of a traffic filtering device based on CC attack features in an embodiment of the present invention. DETAILED DESCRIPTION
[0021] An embodiment of the present invention provides a method and apparatus for traffic filtering based on CC attack signatures. The terms "first," "second," "third," "fourth," and so on (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the terms used in this way are interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0022] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 In one embodiment of the present invention, a method for filtering traffic based on CC attack characteristics includes:
[0023] Step S101: Perform feature extraction and Laplace eigenmap dimensionality reduction on historical HTTP request traffic samples to obtain a set of reduced-dimensional feature vectors;
[0024] It is understandable that the execution subject of the present invention may be a traffic filtering device based on CC attack signatures, or a terminal or a server, which is not limited here. The embodiment of the present invention is described by taking a server as the execution subject as an example.
[0025] Specifically, historical HTTP request traffic samples were preliminarily cleaned and formatted. Predefined rules were used to normalize missing values, outliers, and ambiguous formats in the original request fields, constructing a preprocessed HTTP request record set with a unified structure and consistent temporal order. Based on this standardized request record set, the request sequence was divided into time windows according to request timestamps. Within each time window, temporal behavioral features highly correlated with CC attacks were extracted, such as request frequency fluctuations per unit time, the mean and variance of request intervals, and the peak frequency occurrence and distribution density. This formed an HTTP request temporal feature set that describes request burstiness and periodicity. To uncover attack indicators based on user behavior chains, all preprocessed request records were categorized and aggregated according to their session identifiers to construct a set of logically continuous HTTP sessions. Within each session, session behavioral features that measure interaction patterns were extracted, including session duration, number of requests within a single session, URL path repetition, resource request type distribution entropy, and session stability. This resulted in a session behavior feature set that reflects the stability and suspiciousness of user behavior. To capture the content perturbation strategies used by attackers to construct malicious traffic, the field content of preprocessed HTTP request records is parsed to extract information such as the path depth and character entropy of the request URL, the complexity of the User-Agent and Referer fields, the number of cookie structure levels, and the nesting level and variability of request parameters. This allows the construction of an HTTP content feature set for measuring content anomaly. These three dimensional features—the HTTP request timing feature set, the session behavior feature set, and the HTTP content feature set—are vectorized and aligned in dimensional space before being concatenated and fused to form a unified set of HTTP request feature vectors, covering potential manifestations of CC attacks at the temporal, behavioral, and content levels. Based on the HTTP request feature vector set, a graph structure approach is introduced to model the relationships between feature vectors. By constructing a similarity weight matrix and a node degree matrix between request samples, a Laplacian matrix is obtained. The characteristic equation is then solved to obtain a low-dimensional embedding vector. This allows the original high-dimensional HTTP request feature vector set to be projected into a continuous low-dimensional space that preserves its geometric structure, generating a reduced-dimensional feature vector set.
[0026] In this embodiment, a graph relationship reflecting the inter-sample correlation structure is constructed by taking a set of HTTP request feature vectors as input and measuring the feature similarity between each pair of HTTP requests. This similarity is calculated by comparing the relative proximity of requests across multiple dimensions, such as access frequency, time interval, request path, parameter complexity, and session behavior. The similarity relationships between all samples are aggregated into a weight matrix, representing the connection strength between each pair of requests. Based on this weight matrix, the total connection strength of each request node is calculated to form a degree diagonal matrix, thereby characterizing the distribution density structure of HTTP requests in the overall feature space. Based on these two types of matrices, a Laplacian matrix representing the entire request graph structure is constructed. This matrix reflects the local adjacency relationships and global distribution trends between samples. Its construction process does not change the original request data content, but effectively preserves the structural connections between them. The Laplacian matrix is subjected to feature analysis, i.e., its structural properties are decomposed to obtain feature datasets reflecting the internal variation patterns of the request set. These datasets reflect the most important direction of variation between samples. To compress the feature space dimensionality and improve subsequent model processing efficiency, the most representative principal components are selected from the feature data. Specifically, the set of feature dimensions that best preserves the distribution relationships between the data is selected. These dimensions represent the directions that best reflect the structural differences in attack behavior. The original set of HTTP request feature vectors is mapped into the new feature space formed by these representative dimensions, completing the dimensionality reduction transformation and generating a set of reduced-dimensional feature vectors. In this new feature space, each request data is represented in a lower-dimensional form while retaining its key structural properties and potential attack characteristics. This reduces the complexity of the entire data set while still retaining sufficient discriminatory power.
[0027] Step S102: Perform CC attack feature analysis based on the dimension-reduced feature vector set to obtain a CC attack feature parameter set;
[0028] Specifically, the reduced-dimensionality HTTP request feature vectors are fed into the input layer of a deep-penalized generative adversarial network model for feature loading. By initializing the connection weights between the input nodes and the reduced-dimensional vectors, each low-dimensional feature vector is accurately mapped into an initial representation that the network can process, resulting in initial network features that reflect the actual HTTP request structure. These initial network features are then fed into the generator module of the deep-penalized generative adversarial network. This module utilizes a multi-layer, fully connected neural network structure and introduces nonlinear activation functions such as ReLU or tanh. This module performs high-dimensional feature mapping and nonlinear transformations on the initial features, generating intermediate mapping features that capture the differential representation of authentic and forged requests at the feature level in the latent space. These intermediate mapping features are then fed into the discriminator module of the deep-penalized generative adversarial network for feature classification learning. The discriminator also employs a multi-layer, fully connected structure. Its core function is to map input features into probabilistic responses of attack and non-attack. Through reinforcement learning based on the distribution differences of sample features, it outputs a feature response map that reflects the contribution of each input dimension to the classification decision. To improve the model's ability to identify key features of CC attacks, an attention mechanism layer is embedded in the network structure. This layer constructs an attention weight matrix and calculates the global and local importance scores for each dimension of the feature response graph to obtain the final feature weight distribution, thereby highlighting the main discriminative features in the attack behavior. At the same time, to enhance the stability of the model in the face of request variations and feature perturbations, an input gradient-based denoising penalty mechanism is introduced. This mechanism constrains the gradient norm of the discriminator output with respect to the input features, calculates the gradient regularization term, and combines it with the feature weight distribution to generate a feature expression with robustness enhancement capabilities. Feature dimensions with weights significantly higher than the set threshold are selected from the robustness enhancement features, and combined with the corresponding boundary values to form a CC attack feature parameter set. This set includes parameters such as the upper limit of the attack request frequency, the request path entropy threshold, the session duration range, and the User-Agent structure complexity limit.
[0029] In this embodiment, the discriminator structure in a deep penalized generative adversarial network is extensively trained, enabling it to not only effectively distinguish input features but also quantify the response sensitivity of each feature dimension. During the training process, a gradient calculation is performed on each set of input features. This calculation calculates the degree of change in the discriminator output when the current input undergoes a small perturbation. This constructs a feature gradient matrix, in which each column represents the intensity of the impact of a change in a feature dimension on the output. The norm of the feature gradient matrix is calculated to determine the magnitude of the change corresponding to each feature dimension, reflecting the model's dependence on the input of each dimension. To improve the model's adaptability to traffic variability, a traffic state analysis mechanism is introduced. This mechanism determines a quantitative indicator of traffic variability based on the degree of fluctuation in traffic characteristics within the current time period. Based on this indicator, the constraint strength of the penalty mechanism is dynamically adjusted, enabling the model to enhance its ability to discriminate key features when attack features change significantly, while maintaining a stable response when normal traffic fluctuations are minimal. After completing the above analysis, the calculated feature sensitivity information is integrated with the existing feature weight distribution to form a comprehensive penalty term, which is then added to the training objective. This allows the model to automatically suppress over-reliance on unstable features during optimization while enhancing robust learning of core features. After optimization, the model outputs a set of features that exhibit increased stability and discrimination, known as robustness-enhancing features. Only a subset of these features consistently demonstrate high weight and significant impact across multiple rounds of training and perturbation testing. To this end, a weight threshold is set, and features exceeding this threshold are filtered across all dimensions. For these highly weighted features, their numerical ranges in the training set are statistically analyzed, and the upper and lower bounds with high frequency of occurrence are extracted as the behavioral boundaries for that dimension. This constitutes a set of CC attack feature parameters, which includes the specific feature index and its corresponding weight, as well as the upper and lower bounds of its recurring occurrence in historical attack data.
[0030] Step S103: Perform attack risk assessment on the real-time HTTP request to be filtered based on the CC attack feature parameter set to obtain an attack possibility score;
[0031] Specifically, a structured feature extraction operation is performed on each real-time HTTP request to be filtered. This operation extracts information such as request frequency, parameter complexity, path structure, number of header fields, and source IP distribution characteristics from the request to construct a preliminary real-time HTTP request feature vector. The real-time HTTP request feature vector is then Laplace mapped and projected into the low-dimensional space used by the discriminant model. This reduces the dimensionality of the real-time feature vector, allowing it to be aligned and compared with the learned attack signature parameter set at the same feature scale and structure. To effectively associate the real-time features with the extracted CC attack signature parameters, a feature similarity assessment model is constructed based on the aforementioned CC attack signature parameter set. This model comprises several core modules, including a key feature dimension identifier, a feature boundary interval detector, and a multi-scale similarity measurement engine. This model rapidly determines the degree of similarity between the real-time request and existing attack samples along key attack feature dimensions and outputs a comprehensive feature matching score. When a real-time feature exceeds the upper or lower bounds of known attack signatures in several key dimensions, or exhibits high overlap with attack signatures in multiple dimensions, the similarity score is significantly increased, reflecting its potential risk. To improve the contextual accuracy of attack identification, we supplement this by incorporating behavioral information from the session to which the request belongs, rather than relying solely on feature matching scores. This includes metrics such as request frequency trends, session duration, path repetitiveness, and parameter structure volatility. This assists in assessing the request context from a user behavior perspective. The feature matching score and session behavior characteristics are fed into the comprehensive risk assessment module, where they are combined into a final attack likelihood score through weighted fusion, rule reinforcement, and statistical analysis.
[0032] The index information of all key feature dimensions is extracted from the CC attack feature parameter set to construct a key feature dimension index set. This set indicates which reduced-dimensional features are decisive in identifying attack behaviors. The reduced-dimensional feature vector corresponding to the real-time HTTP request to be evaluated is matched with the index set. The feature values of all specified dimensions are extracted from the original reduced-dimensional vector to construct a real-time key feature vector. This vector retains only the core feature information with the highest correlation with the CC attack and excludes the interference of secondary or invalid dimensions on the scoring results. Based on the threshold and weight information in the above real-time key feature vector and the CC attack feature parameter set, a feature similarity evaluation model is constructed. This model is used to quantify the degree of match between the current request and the attack pattern in each key dimension. In the model design, the relationship between each dimension's eigenvalue and its corresponding upper and lower bounds is measured using a boundary determination function. If the dimension's value falls within the upper and lower bounds, the match is considered good, and the function output is 1. If the value exceeds the bounds, the value is gradually decayed to the minimum acceptable threshold in a linear or nonlinear manner, depending on the degree of deviation from the bounds, resulting in an output close to 0 or even a negative value, indicating the degree to which the feature dimension negatively incentivizes abnormal behavior. Furthermore, to reflect the influence of each dimension in attack signature identification, a per-dimensional feature weight provided in the attack signature parameter set is introduced. These weights are used to weight the boundary matching results of each dimension in the final scoring phase. During the matching calculation phase, each eigenvalue of the real-time key feature vector is substituted into the feature similarity evaluation model, and boundary matching results for all dimensions are calculated. These results are then multiplied by the corresponding feature weights to form a weighted score vector. Finally, all weighted scores are summed and normalized to map the final score to a uniform score range, forming the feature matching score. The higher the score, the closer the real-time request characteristics are to the characteristic patterns of historical CC attack behaviors. Otherwise, the behavior deviates from the attack characteristics.
[0033] Step S104: Perform time-varying parameter defense analysis based on the attack likelihood score to obtain an adaptive defense rule set;
[0034] Specifically, an attack situation vector is established to reflect the current network security situation. This vector is constructed based on attack likelihood scores collected within a continuous time window and comprises three core elements: attack intensity, attack variability, and attack distribution characteristics. Attack intensity represents the proportion of high-risk requests per unit time, attack variability reflects the dynamic drift of attack behavior in the feature space, and attack distribution characteristics characterize the spread of malicious requests across different sessions, source IP addresses, or paths. This situation vector provides a multi-dimensional characterization of attack trends, complexity, and concentration, serving as fundamental input for building dynamic defense strategies. To model the behavioral relationship between attack and defense, a multi-player game model is constructed based on this attack situation vector. Three key players are defined: a CC attacker, a defense controller, and a resource scheduler. The attacker attempts to maximize the impact of the attack while minimizing costs. The defense controller is responsible for developing response strategies to minimize resource loss and misjudgment costs. The resource scheduler, however, achieves an optimal balance between security and service performance within the constraints of system resources. Based on this game structure, optimization objective functions are defined for each participant. For example, the attacker's goal is to maximize the request pass rate and minimize the challenge cost; the defender focuses on reducing the false alarm rate, shortening the response time, and controlling the risk of service interruption; and the resource scheduler balances the fairness of the processing queue with the efficient allocation of security resources. After clearly defining the respective game objectives, the attack situation vector is introduced into the game model as a dynamic input that influences the decision variables of each party. A set of time-dependent parameter update rules is constructed. These rules guide the system to automatically adjust the frequency, threshold, and execution resources of defense responses as the attack intensity or behavioral variability increases. Based on this, optimal control equations are introduced, and these update rules are mathematically solved and optimized. Through iterative calculations and strategy simulations, the defense strategy combination with the optimal defense effectiveness, optimal resource utilization, and minimum response latency under the current situation is obtained. Ultimately, a set of optimal defense strategies is obtained. This strategy set includes the complexity and triggering mechanism of JS challenge verification, the suspicion threshold used for session behavior analysis, and the dynamic allocation ratio of system resources such as processors, bandwidth, and cache. To enable the policy set to be finely executed according to different levels of attack risk, a structured hierarchical combination process is performed on it. Multiple defense levels are divided according to the attack possibility score interval. A corresponding policy combination is bound to each level. For example, a low score corresponds to lightweight verification and resource preservation, a medium score triggers behavior analysis and delay control, and a high score initiates strong challenge verification and resource restriction, ultimately resulting in an adaptive defense rule set.
[0035] Step S105: Perform hierarchical filtering on the real-time HTTP request to be filtered according to the adaptive defense rule set to obtain filtered safe traffic.
[0036] Specifically, each real-time HTTP request to be filtered is assigned a defense level based on its attack likelihood score. Combined with the mapping between the multi-level scoring intervals and defense strategies pre-set in the adaptive defense rule set, the request is labeled with different security risk levels, resulting in an HTTP request with a defense level identifier. This identifier is used to trigger defense processes of varying strengths and strategies in subsequent stages. For requests with a score exceeding the first threshold, meaning they are initially identified as posing a certain risk but not yet meeting the rejection criteria, the system performs JavaScript challenge validation on them. This involves embedding specific JavaScript execution logic to test whether the client has normal browser parsing and response capabilities. Based on the execution results, a JavaScript verification result is generated, which is used to further determine whether the request behavior conforms to a legitimate interaction pattern. If a request score not only exceeds the first threshold but also fails JavaScript challenge validation, indicating significant automation or scripting, a CAPTCHA verification process is performed. This process guides the user through graphic recognition or interactive click operations to confirm that they are a real human user, resulting in a final verification response. The verification result is not only used to determine whether the current request passes or fails, but is also used, along with the JavaScript challenge results, to update the system's IP reputation database mechanism. The system dynamically adjusts the reputation score of a source IP based on its verification performance over time. Frequent verification failures will result in a downward adjustment in reputation, while long-term normal performance will gradually increase the reputation score. The updated IP reputation database is fed back into the adaptive defense rule set in real time, ensuring that defense rules always make policy decisions based on the latest credibility data during dynamic execution, forming a self-learning and self-adjusting defense closed loop. At the same time, fine-grained processor resource, memory buffer, and network bandwidth allocation policies are implemented based on the risk level and verification status of each level-tagged request, combined with the current node resource load. In the resource scheduling engine, high-risk requests are assigned a lower priority and their usage is limited, while requests that pass verification or have lower scores are given normal or weighted resource guarantees, effectively suppressing resource consumption for potential attacks while ensuring system performance. Combining the resource scheduling results with the verification process output, one of three response strategies is executed for each pending request: if the score is low or all verifications have been passed, the request is allowed to pass normally; if the score is medium and some verifications fail, the response delay mechanism is activated, with exponential backoff or queuing to delay the response; if the score is extremely high and all verifications fail, the request is directly rejected, thereby outputting the final filtered secure traffic.
[0037] In an embodiment of the present invention, the present invention constructs a comprehensive set of HTTP request feature vectors by performing multi-dimensional extraction and fusion of temporal features, session behavior features, and content features on historical HTTP request traffic samples. Compared with traditional methods based only on a single or a small number of statistical features, the present invention can more comprehensively characterize the characteristic patterns of HTTP requests. The present invention introduces Laplace eigenmaps to perform dimensionality reduction processing on the HTTP request feature vector set, converting the original high-dimensional discrete features into a continuous feature space, effectively preserving the topological relationship and similarity structure between HTTP requests, not only reducing the computational complexity but also enhancing the expressive power of features, enabling the system to better capture temporal correlation patterns in CC attacks. The present invention uses a deep penalized generative adversarial network to perform CC attack feature analysis. Through adversarial learning between the generator and the discriminator, the characteristic patterns of CC attacks are automatically learned, which has stronger feature learning capabilities than traditional machine learning methods. At the same time, the introduced attention mechanism can automatically focus on the key features that distinguish CC attacks, improving the interpretability and accuracy of the model. The denoising penalty constraint introduced by the present invention constrains the gradient norm of the discriminator near the real data, so that the model output remains stable when the input undergoes small changes, significantly enhancing the model's adaptability to HTTP request variations and effectively dealing with attackers' attempts to evade detection by changing request parameters and adjusting request frequency. The multi-layer game model constructed by the present invention regards CC attack defense as a dynamic game process between attackers and defenders. By solving the optimal control equation, the optimal defense strategy is obtained, and the defense parameters can be dynamically adjusted according to the attack situation. Compared with static defense rules, it is more flexible and adaptable, and can optimize resource utilization efficiency while ensuring security. The hierarchical filtering processing mechanism implemented by the present invention processes requests at different levels according to the attack probability score. Combined with multiple defense measures such as JavaScript challenge verification, CAPTCHA verification, and dynamic resource allocation, it can adopt differentiated defense strategies for requests of different risk levels, ensuring system security while minimizing the impact on normal users. It has significant advantages in dealing with low-frequency CC attacks and mixed Flash Crowd traffic environments.
[0038] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0039] Perform data preprocessing on historical HTTP request traffic samples to obtain a preprocessed HTTP request record set;
[0040] Extracting timing features from the preprocessed HTTP request record set to obtain an HTTP request timing feature set;
[0041] The pre-processed HTTP request record set is grouped by session ID to obtain an HTTP session set;
[0042] Extracting session behavior features from the HTTP session set to obtain a session behavior feature set;
[0043] Extract content features from the preprocessed HTTP request record set to obtain an HTTP content feature set;
[0044] Merge the HTTP request timing feature set, session behavior feature set, and HTTP content feature set to obtain the HTTP request feature vector set;
[0045] Perform Laplace eigenmap dimensionality reduction on the HTTP request feature vector set to obtain a reduced-dimensional feature vector set.
[0046] Specifically, we preprocess historical HTTP request traffic samples to unify field formats, timestamp standards, character encoding, and missing value completion strategies. We also remove meaningless fields, empty requests, duplicate records, and illegal character interference items, resulting in a preprocessed HTTP request record set with a neat data format, clear fields, and a complete timeline. We also ensure that the dataset contains accurate basic field information such as request time, session ID, URL path, request parameters, request header, and source IP. On this basis, we perform time series analysis on each request record, sorting all request records by timestamp and dividing the analysis interval into sliding time windows or fixed time periods. We extract a series of temporal features describing request temporal behavior, including dimensions such as request frequency per unit time, mean, standard deviation, maximum and minimum values of request intervals, request peak position, request density, and burstiness. We also identify abnormally dense request segments through time distribution functions and probability histogram modeling, forming a metric expression for each request in the temporal dimension, thereby constructing an HTTP request temporal feature set. Preprocessed request records are grouped according to their session identification fields. All requests with the same session ID or source IP address and consecutive request intervals below a set threshold are considered part of the same user behavior chain, forming a logical HTTP session set. Based on the session-level data structure, session behavior features containing complete behavioral information are extracted, including the total number of requests per session, the repetition rate of request paths, the diversity of URL paths within a session, the duration of a single session, the request density distribution, the number of resource access types, and session stability change indicators. These session behavior features reveal patterns in which potential attackers mimic normal user access behavior, such as low-frequency but persistent probing behavior, periodic bursts of high-density swiping behavior, or probing behavior with a single path but continuously varying parameters. This results in a set of session behavior features that indicate attack behavior. Content features are extracted from the preprocessed HTTP request record set, including content-related indicators such as the hierarchical structure depth of the URL path, path string entropy, the number and variability of parameter names and parameter values, the number of request header fields, the nesting level and number of fields in the Cookie structure, the complexity of the User-Agent string, the path jump length of the Referer, and content encoding tag anomalies. By analyzing the content structure and semantic layer, request types containing highly disguised, parameter contaminated, or feature-ambiguous behaviors can be effectively identified. The combination of regular template libraries and keyword blacklist and whitelist strategies can further enhance the semantic recognition ability of content features, making content feature integration an important analytical support for covering the covert expression of attack requests.The HTTP request timing feature set, session behavior feature set, and HTTP content feature set are uniformly encoded and aligned in the feature dimensional space. These features are then merged into a unified set of HTTP request feature vectors through feature vector concatenation. Each vector in this set fully captures the temporal behavior, contextual structure, and content complexity of the request. A Laplacian eigenmapping method is introduced to reduce the dimensionality of this high-dimensional feature set. This method constructs a similarity graph between HTTP request features, treating all request samples as nodes in the graph. The similarity between samples is calculated using a Gaussian kernel or Euclidean distance function to construct weighted edges, forming a weight matrix. A Laplacian matrix is then constructed based on the node connectivity, leveraging the graph structure to preserve the local geometric relationships between features. Within this graph structure, a feature mapping operation is performed. By calculating the eigenvectors of the Laplacian matrix, the original high-dimensional feature vector set is projected into a set of low-dimensional spaces while maintaining the similarity relationships of the original data in the graph structure. This reduces redundant information while preserving the discriminative structure and distribution characteristics of the attack behavior, forming a reduced-dimensional feature vector set useful for subsequent CC attack identification, feature adversarial training, and strategy optimization analysis.
[0047] In a specific embodiment, the step of performing Laplace eigenmap dimensionality reduction on the HTTP request feature vector set to obtain the reduced-dimensional feature vector set may specifically include the following steps:
[0048] Calculate HTTP request similarity based on the HTTP request feature vector set to obtain a weight matrix;
[0049] Perform metric calculation on the weight matrix to obtain the degree diagonal matrix;
[0050] Perform matrix operations on the degree diagonal matrix and the weight matrix to obtain the Laplace matrix;
[0051] Perform eigendecomposition calculation on the Laplace matrix to obtain the eigenvalue set and eigenvector set;
[0052] Sort and filter the eigenvector set according to the eigenvalue set to obtain the main eigenvector subset;
[0053] The HTTP request feature vector set is projected into the feature space composed of the main feature vector subset to obtain a reduced-dimensional feature vector set.
[0054] Specifically, the HTTP request similarity is calculated based on the HTTP request feature vector set. A similarity function is constructed between samples. Commonly used functions are Gaussian radial basis function, cosine similarity function or Manhattan distance function. The numerical relationship between each pair of requests is obtained through function operation, and a symmetrical similarity weight matrix is formed. Each element in the matrix represents the degree of similarity between the corresponding two HTTP requests. The larger the value, the more similar they are at the feature level. In practical applications, in order to improve the sparsity of the weight matrix and reduce the computational complexity, a local adjacency mechanism is introduced. Only the similarity between each request and the nearest samples in its feature space is calculated, while other positions are set to zero, thereby constructing a weight matrix with practical discriminative ability and sparse efficiency. The weight matrix is supplemented with a metric structure, that is, the connection strength of each request node is calculated. This step is completed by constructing a degree diagonal matrix. This matrix is a diagonal matrix. Each diagonal element represents the sum of all the connection edge weights of the corresponding request in the weight matrix, representing its total connection strength with other requests in the similarity graph. This degree diagonal matrix, along with the previously constructed weight matrix, is used to generate the Laplacian matrix within the graph structure. This matrix represents the overall transformation of the entire request feature graph within the graph space. By subtracting the weight matrix from the degree diagonal matrix, the resulting Laplacian matrix reflects the differences in connectivity between local and global nodes within the graph structure. Mathematically, it exhibits properties such as symmetry and positive semidefiniteness, making it suitable for feature embedding and dimensionality reduction modeling. The Laplacian matrix is subjected to eigendecomposition, which involves calculating the set of eigenvalues and eigenvectors within the matrix. This decomposition identifies the smoothest feature transformation directions within the graph, which are the principal component directions within the graph structure. All eigenvalues are sorted by size, with the smallest eigenvalue corresponding to the direction with the slowest variation and strongest local consistency within the sample structure. The corresponding eigenvector forms the basis of the embedding space. To achieve effective dimensionality reduction, the first several eigenvectors corresponding to the optimal eigenvalues are selected based on the preset target dimension or through spectral spacing analysis to construct a subset of principal eigenvectors. The original HTTP request feature vector set is projected into the low-dimensional space formed by the feature vector subset. Each request vector in the high-dimensional space is re-expressed in the low-dimensional principal component space through feature coordinate transformation, and a feature expression matrix after dimensionality reduction is formed.
[0055] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0056] Input the reduced dimension feature vector set into the input layer of the deep penalty generative adversarial network for feature loading to obtain the initial features of the network;
[0057] The generator in the deep penalty generative adversarial network performs nonlinear transformation on the initial features of the network to obtain intermediate mapping features;
[0058] The intermediate mapping features are input into the discriminator in the deep penalty-generated adversarial network to learn CC attack features and obtain the feature response map;
[0059] The feature importance of the feature response map is calculated through the attention mechanism layer in the deep penalty generative adversarial network to obtain the feature weight distribution;
[0060] The denoising penalty constraint is calculated based on the feature weight distribution and gradient information to obtain the robustness enhancement feature. The feature dimensions and their boundary values whose weights exceed the threshold are extracted from the robustness enhancement feature to obtain the CC attack feature parameter set.
[0061] Specifically, the set of reduced-dimensionality feature vectors of HTTP requests is fed into the input loading layer of a deep penalized generative adversarial network. This loading layer, connected to the input layer of the generator network, is a linear transformation layer that performs a preliminary rescaling and embedding encoding on the input features to ensure they meet the computational format and dimensionality requirements of subsequent neural network layers. The output serves as the network's initial features. The network's initial features are then fed into the generator, which consists of multiple, fully connected layers, specifically a four-layer structure: the first layer is the input embedding layer, the second and third layers are hidden layers, and the fourth layer is the output layer. Reluctant linear unit (ReLU) is used as a nonlinear activation function between layers, and the output layer uses the Tanh function to constrain the range of generated feature values. The generator's input can include both the actual reduced-dimensional features and random noise can be introduced during training to improve the generator's generalization ability. The generator performs a nonlinear mapping transformation on the input features and outputs intermediate mapped features. This intermediate feature vector is essentially a representation in the latent feature space that captures deep features not explicitly encoded in the original request but that are important for model discrimination. The intermediate mapped features are input into the discriminator module of a deep penalized generative adversarial network (DPGAN). The discriminator consists of a four-layer fully connected neural network: the first layer is the input perception layer, the second and third layers are the discriminative hidden layers, and the fourth layer is the probabilistic output layer. It uses a sigmoid function to output attack probability predictions. The training goal is to distinguish real features from generated features. In this discrimination process, the DPGAN introduces the concept of a feature response map. As each batch of training samples passes through the discriminator, the responses generated by the activation of neurons in the hidden layer for each feature dimension are recorded. This generates a response map that reflects the influence of each feature on the final classification output, providing input for feature interpretation and importance analysis. To achieve the goals of feature selection and feature compression, the DPGAN embeds an attention mechanism layer within the discriminator structure. This layer, placed after the second or third hidden layer of the discriminator, constructs a weight matrix to assign importance to each dimension in the feature response map. The weight value for each dimension is then normalized using the Softmax function and output. The attention layer is trained simultaneously with the discriminator backbone network. Its goal is to maximize discrimination accuracy while focusing the model's attention on the feature dimensions that best distinguish CC attacks from legitimate requests. This generates a feature weight distribution, with higher weights indicating a greater contribution to attack discrimination. To improve the model's adaptability to variant attacks in real-world scenarios, the Deep Penalized Generative Adversarial Network incorporates a denoising penalty mechanism, optimizing model robustness by stabilizing the discriminator's gradient response.In each round of training, small perturbations are applied to the input features, and the gradient of the discriminator output with respect to the input is calculated. A constraint is then applied to the gradient norm to form a denoising penalty term, which is added as an additional loss to the main loss function. The optimization goal is to mitigate large fluctuations in the model's output under perturbations, thereby maintaining output discriminant stability in the face of malicious request parameter randomization and path obfuscation attack strategies. After completing this training process, the deep penalized generative adversarial network model is capable of responding to features from different types of HTTP requests. During deployment, the trained discriminator and attention layer are directly used to infer and evaluate new input request features, generating corresponding feature response maps and weight distributions. Based on this, the attention weights are filtered, and a preset threshold is set. Dimensions with weights greater than this threshold are extracted from all feature dimensions and marked as key attack discriminative features. Based on the value range of this feature dimension in the corresponding training samples, the maximum and minimum values, or upper and lower quantile boundaries, are extracted to form a set of CC attack feature parameters. Each record in this set contains the feature dimension number, feature name, feature weight, upper and lower limit values.
[0062] In a specific embodiment, the step of calculating the denoising penalty constraint based on the feature weight distribution and gradient information to obtain the robustness enhancement feature, and extracting the feature dimensions and their boundary values whose weights exceed the threshold from the robustness enhancement feature to obtain the CC attack feature parameter set may specifically include the following steps:
[0063] Perform gradient calculation on the discriminator in the deep penalty generative adversarial network to obtain the feature gradient matrix, and perform norm calculation on the feature gradient matrix to obtain the gradient norm value;
[0064] The penalty coefficient is dynamically adjusted according to the gradient norm value and the current flow variability index to obtain an adaptive penalty coefficient;
[0065] Calculate the denoising penalty term based on the adaptive penalty coefficient and feature weight distribution to obtain the constrained model loss function, which includes the adversarial loss and the denoising penalty term;
[0066] The discriminator is trained using the constrained model loss function to output robustness-enhanced features.
[0067] Feature dimensions with weight values greater than a preset threshold are screened from the robustness enhancement features and their boundary values are extracted to obtain a CC attack feature parameter set. The CC attack feature parameter set includes feature dimension index, feature weight value, feature upper limit value, and feature lower limit value.
[0068] Specifically, a deep penalized generative adversarial network system consisting of a generator and a discriminator is constructed. The discriminator not only undertakes the adversarial task of distinguishing real from fake features but also requires the ability to stabilize its gradient response to input perturbations. The network structure consists of an input feature loading layer, a generator module, a discriminator module, an attention mechanism module, and a denoising penalty calculation module, all of which are organically connected through a standard deep learning framework. During the core training phase of the model, the discriminator performs a gradient calculation operation. Specifically, after inputting a reduced feature vector, the partial derivatives of the discriminator output with respect to the input features are calculated dimension by dimension, resulting in a feature gradient matrix. This matrix records the sensitivity of the discriminator's response to each dimension of the input feature. A larger value indicates a stronger influence of the feature on the model output and also indicates that the feature is more susceptible to input perturbations. After the feature gradient matrix is calculated, it is normed. The L2 norm of the gradient of each row (i.e., each sample) along each feature dimension is calculated to obtain the overall gradient response strength of each sample, which is then summarized as a set of quantitative stability indicators, called gradient norms. To enable the discriminator to dynamically adapt to input perturbations and data changes, and to account for the changing trends in the current network traffic environment, a traffic variability index is introduced as a control factor. This index is calculated based on the statistical variation range or standard deviation of the request feature vector in a sliding window and measures the fluctuation of the HTTP request structure over a short period of time. This index is combined with the previously calculated gradient norm value to dynamically adjust the denoising penalty strength in the current training round according to a preset functional relationship. This results in a penalty control variable, the adaptive penalty coefficient, that automatically changes with the training stage and traffic conditions. This coefficient directly influences the weight of the penalty term in the model loss function, automatically adjusting the model's sensitivity to input perturbations. Based on the adaptive penalty coefficient and combined with the feature weight distribution in the discriminator, the gradient strength of the discriminant output with respect to the input is weighted to form a denoising penalty term. The denoising penalty term is superimposed on the original adversarial loss function to form a new constrained model loss function that incorporates structural regularization. This loss function preserves the discriminator's ability to distinguish real from generated data while also constraining the model's local gradient behavior, thereby enhancing its invariance to small perturbations and its ability to detect attacks from unknown variants. This composite loss function, used during backpropagation and gradient optimization, automatically adjusts the network weight structure, ensuring that the discriminator not only achieves accurate discrimination in the training set but also maintains strong generalization robustness in the test set and real-world network environments. After several rounds of optimization training, the discriminator's output feature response will have enhanced resistance to perturbations, known as robustness-enhancing features.These features carry the discriminator's stable learning expression under multiple rounds of perturbations. Based on this, the feature weight distribution output by its embedded attention mechanism is sorted and analyzed, and a system-wide stability threshold is set. All feature dimensions with weights above this threshold are selected, and these features are identified as key inputs to the discriminant model. Simultaneously, the range of each key feature dimension in attack and normal samples is traced back in the training set, and the minimum and maximum values of each dimension are extracted, or the upper and lower quartiles are selected as the value boundaries of the dimension. This constructs a complete set of CC attack feature parameters, including feature dimension index, feature weight, and feature upper and lower limits. When this parameter set is applied to network security defense scenarios, its structure can serve as the basis for determining rule-based defense strategies. The input is the real-time extracted HTTP request features, and the output is whether the request meets the attack performance determination conditions of a certain feature dimension. It can also be combined with the scoring model to form a step-by-step risk assessment system. In model integration scenarios, this parameter set is also shared with other deep models to improve the discrimination weighting mechanism in multi-model fusion strategies. At the same time, in cloud platform deployments or edge gateway systems, this set is cached as a feature rule table to accelerate rapid matching and pre-judgment in low-resource environments, thereby improving the response efficiency of the overall defense system.
[0069] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0070] Perform feature extraction on the real-time HTTP request to be filtered to obtain a real-time HTTP request feature vector, and perform Laplace mapping on the real-time HTTP request feature vector to obtain a real-time feature vector after dimensionality reduction;
[0071] A feature similarity evaluation model is constructed based on the CC attack feature parameter set, and the reduced-dimensional real-time feature vector is input into the feature similarity evaluation model for feature matching calculation to obtain the feature matching score.
[0072] A comprehensive risk assessment is performed based on the feature matching score and the session behavior information of the real-time HTTP requests to be filtered to obtain an attack possibility score.
[0073] Specifically, a traffic processing framework for real-time request analysis is established. This framework can capture HTTP request packets at the network entry layer and quickly decode their content structure so that the request data can be structured and processed immediately. Key fields in the request content, such as URL path, parameter string, request header field, Cookie field, User-Agent identifier, Referer field, source IP address, and request timestamp, will be uniformly extracted and mapped into structures with numerical and vector processing capabilities. Subsequently, the above information is converted into original feature vectors that can be input into the artificial intelligence model through pre-processing operations such as standardization and normalization to form real-time HTTP request feature vectors. In order to maintain consistency with the feature space in the training phase, Laplace mapping processing is performed on the real-time feature vector. Based on the existing Laplace projection subspace, the input vector is projected using its feature vector transformation matrix to ensure that the mapped result has the same dimensionality reduction scale and semantic space as the training sample. During the modeling process, Laplacian eigenmaps construct a weight matrix, a degree matrix, and a Laplacian matrix based on the similarity graph between historical HTTP requests. Eigendecomposition is then used to extract the principal component directions and construct a projection matrix. During deployment, these matrices are cached in the online system to support low-latency computation. After receiving real-time request input, a set of linear transformations rapidly converts the original vector into a reduced-dimensional real-time feature vector, reducing the computational complexity during model inference and enhancing the ability to express nonlinear attack behavior structures. After feature dimensionality reduction, these vectors are used as input to invoke a pre-built feature similarity evaluation model. This model is instantiated based on a set of CC attack feature parameters generated offline. Its structure includes an index table of key feature dimensions, a list of corresponding feature weights, upper and lower boundary values for each feature dimension, and a set of interpretable matching functions. The evaluation model extracts a subset of key dimensions from the real-time feature vector and applies a boundary matching function to calculate the boundary belongingness of each feature dimension. This determines whether the feature value falls within the attack sample boundary interval. If it does, a full score is assigned; if it does, a penalty score is assigned based on the distance beyond the boundary. The matching scores of all dimensions are then weighted and summed using the corresponding weight factors, and normalized to obtain a feature matching score. This score directly quantifies the degree of similarity between the current request and a typical attack sample in terms of high-dimensional feature distribution. The behavioral feature information of the session in which the request is located is introduced to construct a complete risk assessment input set. The current request is clustered into sessions, and the session entity to which it belongs is defined based on the source IP, User-Agent, and time window relationship of the request. Behavioral statistical features are then extracted from the session, including dimensions such as session life cycle length, request frequency fluctuation range, path repetition rate, resource request type change rate, and parameter structure mutation times. These behavioral features, combined with the feature matching score, are fed into the comprehensive risk assessment module as input.This module is usually implemented as a lightweight scoring network or decision tree model in the system architecture. Its structure includes an input layer, a feature normalization layer, a weighted rule base layer, and an output scoring function module. The model uses weighted addition or a multi-layer perception mechanism to integrate and evaluate the feature matching score and the session behavior score, and outputs an attack possibility score between 0 and 1. The closer the score is to 1, the higher the attack risk.
[0074] In a specific embodiment, the execution step constructs a feature similarity evaluation model based on the CC attack feature parameter set, and inputs the reduced-dimensional real-time feature vector into the feature similarity evaluation model to perform feature matching calculation. The process of obtaining the feature matching score may specifically include the following steps:
[0075] Extract the feature dimension index from the CC attack feature parameter set to obtain a key feature dimension index set, and extract the feature value of the corresponding dimension from the real-time feature vector after dimensionality reduction based on the key feature dimension index set to obtain a real-time key feature vector;
[0076] Building a feature similarity evaluation model , where x is the real-time key feature vector, i represents the index, and w i is the feature weight value of the corresponding dimension in the CC attack feature parameter set, B i (x i ) is the boundary judgment function, n is the number of dimensions of the key feature dimension index set, and the boundary judgment function B i (x i ) is defined as when x i When B falls between the upper and lower limits of the characteristic i (x i )=1, otherwise B i (x i ) is reduced to a value between 0 and -1 according to the distance from the boundary, and S(x) represents the feature matching score;
[0077] Each dimension value xi of the real-time key feature vector is substituted into the feature similarity evaluation model to obtain the boundary matching results of each dimension, and the boundary matching results of all dimensions are weighted and normalized to obtain the feature matching score.
[0078] Specifically, during system operation, a pre-built set of CC attack feature parameters is stored in a cache or database in a structured table format. Each parameter record consists of four fields: feature dimension index, feature weight, upper and lower feature limits. This parameter set is generated using the attention weight outputs trained in a deep penalized generative adversarial network and a boundary extraction module, resulting in high discrimination and low redundancy. When a match is required for a real-time HTTP request, all feature dimension indices are extracted from this parameter set to form a key feature dimension index set. This set is an integer vector identifying the feature numbers from the full reduced-dimensional feature space that are required for matching. The reduced-dimensional feature vector of the input real-time HTTP request is obtained from this low-dimensional vector, which is a set of low-dimensional vectors previously reduced by the Laplace mapping module. Based on the key feature dimension index set, the corresponding eigenvalues are extracted from the reduced-dimensional vector by index, forming a new vector structure, the real-time key feature vector. To ensure processing efficiency, this extraction operation is performed using sparse matrix mapping or index-selected tensors to avoid the time overhead of loop structures. Construct a feature similarity evaluation model, which is a structured matching model with interpretability and adjustability. It is not a black box discriminator in the form of a deep neural network, but a scoring function for rule expression. Its core structure consists of the following parts: the input is a real-time key feature vector x, whose dimension is n; the internal structure includes a set of boundary judgment functions B i (x i ), each B i According to the value x of the i-th feature i Whether it falls between the upper and lower boundaries for segmented scoring, if x i is within the predefined boundary interval, then B i (x i )=1, indicating that the feature fully meets the attack feature performance; if x i If it exceeds the boundary, linear attenuation will be performed according to the distance it exceeds the upper and lower limits, so that B i (x i ) output value drops from 0 to −1, indicating that the feature is outside the safety boundary and has a certain degree of deviation. The reduction is defined as a linear or exponential function based on the maximum deviation ratio. The boundary judgment result of each dimension will be combined with the feature weight of the dimension. Perform weighted processing to indicate the influence of the feature in the overall score. In the scoring execution phase, the model performs the boundary judgment function B on each dimension i. i (x i ), get the matching result, and compare the result with the weight of the feature Multiplying them together forms a weighted matching score. After summing the weighted matching scores for all dimensions, this sum is divided by the sum of all feature weights and normalized to obtain the final feature matching score S(x). Its value range is between −1 and 1. Values closer to 1 indicate that the current request closely matches the attack pattern in multiple key feature dimensions. Values closer to 0 or negative numbers indicate that the behavior deviates from the attack sample feature distribution range, thus providing a quantitative basis for the subsequent risk assessment system.
[0079] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0080] Construct an attack situation vector based on the attack probability score. The attack situation vector includes the attack intensity, attack variability, and attack distribution characteristics within the current time window.
[0081] Construct a set of game participants based on the attack situation vector, which includes CC attackers, defense controllers, and resource schedulers;
[0082] The game optimization goal is defined according to the set of game participants, and time-varying parameter analysis is performed based on the attack situation vector and the game optimization goal to obtain the dynamic parameter update rules;
[0083] By solving the optimal control equation, the dynamic parameter update rules are optimized to obtain the optimal defense strategy set, which includes JavaScript challenge verification strength, session behavior analysis threshold, and resource allocation ratio.
[0084] The optimal defense strategy set is hierarchically combined to obtain an adaptive defense rule set. The adaptive defense rule set contains multiple defense levels, and each defense level corresponds to a different attack possibility score interval and a corresponding combination of defense measures.
[0085] Specifically, an attack situation vector describing the current attack status is constructed using the attack likelihood score as input. The attack likelihood score, output by an upstream deep discriminant network, feature matching model, or hybrid decision engine, represents the attack risk level of a single HTTP request. By setting a fixed-length sliding time window and counting all scoring results within the current time window, a multidimensional vector reflecting the overall security situation is constructed. The attack intensity is measured by the proportion of requests with scores above a certain high-risk threshold. Attack variability is calculated using the standard deviation, entropy, or dispersion coefficient of the fluctuation range of key features in high-scoring requests. Attack distribution characteristics are modeled based on the degree of dispersion of the IP source, path distribution, user agent, or session dimensions of high-scoring requests. Metrics such as cluster entropy and frequency dispersion coefficient are used to comprehensively assess whether the attack represents a centralized outbreak or distributed, scattered disguise. Using the attack situation vector as input, mathematical representations of three types of game participants are constructed: CC attackers, defense controllers, and resource schedulers. As a potential adversary, CC attackers employ tactics such as request frequency adjustment, behavioral obfuscation, signature perturbation, and IP switching, aiming to evade defense detection, maintain a high pass rate, and maximize the target server's resources. The defense controller, representing the security module's response execution core, employs strategies such as verification method selection (e.g., whether to enable JavaScript challenges), analysis granularity adjustment (e.g., session determination threshold), and risk response frequency. The resource scheduler is responsible for system-level resource balancing. Its strategies include CPU allocation, memory cache adjustment, bandwidth throttling, and request queuing strategies. Its optimization goal is to maintain system service capacity and stability while ensuring effective defenses. After defining the roles of the game participants, a multi-objective optimization function is constructed based on their strategy space and the attack situation vector, forming the optimization objective expression for the game problem. The CC attacker's objective function is defined as "maximizing the request pass rate minus the attack cost," the defense controller's objective function as "a weighted combination of maximizing attack detection accuracy and minimizing the false alarm rate," and the resource scheduler's objective function as "maximizing resource load balancing and minimizing system processing latency." To achieve dynamic adaptation, the attack situation vector is input into the game modeler. Combined with the game objective function, a time-varying parameter analysis is performed. By calculating indicators such as risk weight fluctuations, attack mode transition rates, and response effect feedback, dynamic adjustment rules for control variables are generated. Specifically, for each dimension of policy parameters (such as verification strength, analysis threshold, and resource quota), a set of adjustment models that respond to changes in the attack situation are established, such as linear regression functions, exponential response models, state machine switching strategies, or LSTM sequence prediction models. These adjustment models are jointly optimized using optimal control equations. The equation inputs include the current attack situation vector, the output value of the previous round of policies, the policy change cost function, and system resource constraints. The optimization goal is to maximize the attack interception rate and minimize the verification cost while satisfying the constraints of minimizing service interruption and resource overrun.The optimal control solution is implemented using dynamic programming, Lagrange multiplier methods, Bellman equations, or reinforcement learning methods, such as Q-learning or DDPG algorithms, to form a gradual approximation of the policy space in continuous time. The system ultimately outputs a set of verified optimal control strategies, known as the optimal defense policy set. This policy set includes at least three core variables: JavaScript challenge verification strength (e.g., whether enabled, enabled percentage, and challenge script complexity), session behavior analysis thresholds (e.g., lower limit for suspicion score), and resource allocation ratios (e.g., maximum bandwidth allowed for high-risk requests and percentage of CPU cores). To enhance the flexibility and scalability of the policy system, the optimal defense policy set is restructured to form a standardized hierarchical control table structure, creating an adaptive defense rule set. This rule set establishes multiple attack probability scoring intervals, mapping them to policy combinations. Each interval is defined as a defense level, with higher levels indicating stronger policies and more significant resource intervention. Each record in the rule set consists of a scoring interval, trigger conditions, verification strategy, resource control scheme, and recovery mechanism. Policy iteration is achieved through scheduled updates, policy evaluation, or manual administrator adjustments.
[0086] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0087] Based on the attack possibility score of the real-time HTTP request to be filtered, the real-time HTTP request to be filtered is divided into defense levels to obtain HTTP requests with level tags;
[0088] Performing JavaScript challenge verification on the HTTP requests with the level mark and the scores exceeding the first threshold, and obtaining JavaScript verification results;
[0089] Performing CAPTCHA verification on HTTP requests with level tags whose scores exceed a second threshold and whose JavaScript verification results fail, obtaining a verification response result, where the second threshold is greater than the first threshold;
[0090] Update the reputation score of the request source IP based on the JavaScript verification result and the verification response result, obtain the updated IP reputation database, and feed the updated IP reputation database back to the adaptive defense rule set for rule update;
[0091] Dynamically allocate processor resources, memory buffers, and network bandwidth to HTTP requests with level tags to obtain resource scheduling results;
[0092] According to the resource scheduling results and the verification status, the real-time HTTP request to be filtered is passed, delayed, or rejected to obtain the filtered safe traffic.
[0093] Specifically, a real-time scoring-driven hierarchical decision-making module is established. This module takes as input an attack likelihood score, derived by an artificial intelligence recognition model through inference of HTTP request feature vectors and session behavior data. The output value, ranging from 0 to 1, indicates the probability that the request represents a CC attack. The system sets two scoring thresholds: the first threshold identifies mildly suspicious requests, while the second threshold, which is higher than the first, identifies highly suspicious requests. Based on the score, each request is classified into multiple defense levels, such as low risk (score < first threshold), medium risk (between the two thresholds), and high risk (score > second threshold). A defense level tag is appended to the request object, forming a set of HTTP requests with this level tag. JavaScript challenge validation is performed on all requests with medium or high risk tags (i.e., scores exceeding the first threshold). This validation is implemented by dynamically injecting a JavaScript script on the server. Upon receiving the response, the client parses and correctly executes the script, which can, for example, construct a specific cookie, submit signature verification parameters, or perform mathematical logic tasks. The system then compares the client's response with the original logic to determine whether the client has the browser parsing and execution capabilities. This process does not rely on human interaction, is transparent to normal users, and has a certain degree of interception capability against simulation tools and scripts. Requests that pass verification are marked as "Verification Successful," while requests that fail verification or do not respond are recorded as "JS Verification Failed." For requests whose scores exceed the second threshold and fail JavaScript verification, a higher-level human-machine verification mechanism—CAPTCHA—is implemented. This verification mechanism requires the client to display interactive interfaces such as image recognition, character input, and dragging puzzle pieces, requiring the user to perform real-world verification tasks. The system records verification success or failure based on the returned results and generates a verification response record, while retaining session context to prevent duplicate verification. The entire verification process must incorporate timeout control, exception return handling, and verification status synchronization mechanisms to ensure stable performance even in the face of large-scale verification traffic. To implement dynamic reputation feedback and policy self-learning, a reputation scoring mechanism for the request source IP address is established. After each verification, the IP reputation score is updated based on the verification results. If an IP address consistently passes verification, its reputation value is increased; conversely, if it fails repeatedly, its reputation value is decreased. The reputation calculation model utilizes exponential smoothing or a score-backward update method with a memory window to maintain responsiveness and robustness against sudden changes. Updated IP reputation information is stored in an IP reputation database, which should utilize a high-concurrency key-value store (such as Redis or an in-memory database). This database is synchronized in real time to the adaptive defense rule set via a feedback interface, used to adjust the initial score for subsequent requests, verification thresholds, and policy level recommendations. For example, IPs with high reputations can be temporarily exempted from verification; IPs with extremely low reputations can be preemptively blacklisted.In terms of policy response, a resource-aware scheduling mechanism is introduced. Based on HTTP requests with level tags, the system performs dynamic allocation of processor resources, memory buffers, and network bandwidth to requests, taking into account the current system load status and resource pool capacity. The scheduling strategy uses a priority queue-based resource allocation model or a dynamic weighted allocation approach to prioritize resources for requests that have passed verification or are low-risk, while implementing bandwidth restrictions, buffer shrinkage, or queue delays for medium- and high-risk requests. Under high load conditions, an elastic degradation strategy is implemented to defer processing of high-risk requests to ensure the continuity of the system's core business. The resource scheduling results and verification status are combined to determine the final handling method for each HTTP request to be filtered. If resources are sufficient and verification is passed, the request is directly released and enters the business logic layer. If verification fails but system resources are sufficient, the request enters the delayed response queue and undergoes exponential back-off queuing. If verification fails and resources are limited, or the IP reputation is extremely low, a rejection response is directly returned, or the request is directed to a closed logic processing module such as an error page or sandbox environment. This multi-path response mechanism constitutes the final secure traffic screening output process, ensuring that the system can dynamically adapt to high-risk traffic bursts, achieving the triple goals of traffic quality filtering, precise resource allocation, and service availability assurance.
[0094] The above describes the traffic filtering method based on CC attack characteristics in the embodiment of the present invention. The following describes the traffic filtering device based on CC attack characteristics in the embodiment of the present invention. Figure 2 In one embodiment of the present invention, a traffic filtering device based on CC attack characteristics includes:
[0095] A mapping and dimensionality reduction module 201 is used to perform feature extraction and Laplace feature mapping dimensionality reduction on historical HTTP request traffic samples to obtain a set of reduced-dimensionality feature vectors;
[0096] A feature analysis module 202 is configured to perform CC attack feature analysis based on a set of reduced-dimensional feature vectors to obtain a set of CC attack feature parameters;
[0097] The risk assessment module 203 is used to perform attack risk assessment on the real-time HTTP request to be filtered based on the CC attack feature parameter set to obtain an attack possibility score;
[0098] A defense analysis module 204 is configured to perform time-varying parameter defense analysis based on the attack likelihood score to obtain an adaptive defense rule set;
[0099] The hierarchical filtering module 205 is configured to perform hierarchical filtering on the real-time HTTP requests to be filtered according to the adaptive defense rule set to obtain filtered safe traffic.
[0100] Through the collaborative cooperation of the above-mentioned components, the present invention constructs a comprehensive set of HTTP request feature vectors by performing multi-dimensional extraction and fusion of temporal features, session behavior features, and content features on historical HTTP request traffic samples. Compared with traditional methods based only on a single or a small number of statistical features, it can more comprehensively characterize the characteristic patterns of HTTP requests. The present invention introduces Laplace eigenmaps to perform dimensionality reduction processing on the HTTP request feature vector set, converting the original high-dimensional discrete features into a continuous feature space, effectively preserving the topological relationship and similarity structure between HTTP requests, not only reducing the computational complexity, but also enhancing the expressive power of the features, enabling the system to better capture the temporal correlation patterns in CC attacks. The present invention uses a deep penalty generative adversarial network to perform CC attack feature analysis. Through adversarial learning between the generator and the discriminator, it automatically learns the characteristic patterns of CC attacks, and has stronger feature learning capabilities than traditional machine learning methods. At the same time, the introduced attention mechanism can automatically focus on the key features that distinguish CC attacks, improving the interpretability and accuracy of the model. The denoising penalty constraint introduced by the present invention constrains the gradient norm of the discriminator near the real data, so that the model output remains stable when the input undergoes small changes, significantly enhancing the model's adaptability to HTTP request variations and effectively dealing with attackers' attempts to evade detection by changing request parameters and adjusting request frequency. The multi-layer game model constructed by the present invention regards CC attack defense as a dynamic game process between attackers and defenders. By solving the optimal control equation, the optimal defense strategy is obtained, and the defense parameters can be dynamically adjusted according to the attack situation. Compared with static defense rules, it is more flexible and adaptable, and can optimize resource utilization efficiency while ensuring security. The hierarchical filtering processing mechanism implemented by the present invention processes requests at different levels according to the attack probability score. Combined with multiple defense measures such as JavaScript challenge verification, CAPTCHA verification, and dynamic resource allocation, it can adopt differentiated defense strategies for requests of different risk levels, ensuring system security while minimizing the impact on normal users. It has significant advantages in dealing with low-frequency CC attacks and mixed Flash Crowd traffic environments.
[0101] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, apparatuses and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0102] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a traffic filtering device based on CC attack signatures (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0103] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A traffic filtering method based on CC attack characteristics, characterized in that: include: Perform feature extraction and Laplace eigenmap dimensionality reduction on historical HTTP request traffic samples to obtain a set of reduced-dimensional feature vectors; Based on the set of reduced-dimensional feature vectors, CC attack feature analysis is performed to obtain a set of CC attack feature parameters; specifically, the method comprises: inputting the set of reduced-dimensional feature vectors into the input layer of a deep penalty generative adversarial network for feature loading to obtain network initial features; performing nonlinear transformation on the network initial features through the generator in the deep penalty generative adversarial network to obtain intermediate mapping features; inputting the intermediate mapping features into the discriminator in the deep penalty generative adversarial network to perform CC attack feature learning to obtain a feature response map; performing feature importance calculation on the feature response map through the attention mechanism layer in the deep penalty generative adversarial network to obtain a feature weight distribution; performing gradient calculation on the discriminator in the deep penalty generative adversarial network to obtain a feature gradient matrix, and Performing a norm calculation on the feature gradient matrix to obtain a gradient norm value; dynamically adjusting the penalty coefficient according to the gradient norm value and the current traffic variability index to obtain an adaptive penalty coefficient; calculating a denoising penalty term based on the adaptive penalty coefficient and the feature weight distribution to obtain a constrained model loss function, wherein the constrained model loss function includes an adversarial loss and a denoising penalty term; optimizing and training the discriminator using the constrained model loss function to output a robustness enhancement feature; screening feature dimensions whose weight values are greater than a preset threshold from the robustness enhancement feature and extracting their boundary values to obtain a CC attack feature parameter set, wherein the CC attack feature parameter set includes a feature dimension index, a feature weight value, a feature upper limit value, and a feature lower limit value; Performing an attack risk assessment on the real-time HTTP request to be filtered based on the CC attack feature parameter set to obtain an attack possibility score; Performing time-varying parameter defense analysis based on the attack likelihood score to obtain an adaptive defense rule set; Perform hierarchical filtering processing on the real-time HTTP request to be filtered according to the adaptive defense rule set to obtain filtered safe traffic.
2. The traffic filtering method based on CC attack characteristics according to claim 1 is characterized in that: The feature extraction and Laplace feature map dimensionality reduction are performed on the historical HTTP request traffic samples to obtain a set of reduced-dimensional feature vectors, including: Perform data preprocessing on historical HTTP request traffic samples to obtain a preprocessed HTTP request record set; Extracting timing features from the preprocessed HTTP request record set to obtain an HTTP request timing feature set; Grouping the preprocessed HTTP request record set by session ID to obtain an HTTP session set; Extracting session behavior features from the HTTP session set to obtain a session behavior feature set; Extracting content features from the pre-processed HTTP request record set to obtain an HTTP content feature set; Merging the HTTP request timing feature set, the session behavior feature set, and the HTTP content feature set to obtain an HTTP request feature vector set; Perform Laplace eigenmap dimensionality reduction on the HTTP request feature vector set to obtain a reduced-dimensionality feature vector set.
3. The traffic filtering method based on CC attack characteristics according to claim 2 is characterized in that: The performing Laplace eigenmap dimensionality reduction on the HTTP request feature vector set to obtain a reduced-dimensionality feature vector set includes: Perform HTTP request similarity calculation based on the HTTP request feature vector set to obtain a weight matrix; Performing metric calculation on the weight matrix to obtain a degree diagonal matrix; Performing a matrix operation on the degree diagonal matrix and the weight matrix to obtain a Laplace matrix; Performing eigendecomposition calculation on the Laplace matrix to obtain an eigenvalue set and an eigenvector set; Sorting and screening the eigenvector set according to the eigenvalue set to obtain a main eigenvector subset; The HTTP request feature vector set is projected onto the feature space formed by the main feature vector subset to obtain a reduced-dimensionality feature vector set.
4. The traffic filtering method based on CC attack characteristics according to claim 1 is characterized in that: The attack risk assessment of the real-time HTTP request to be filtered based on the CC attack feature parameter set is performed to obtain an attack possibility score, including: Performing feature extraction on the real-time HTTP request to be filtered to obtain a real-time HTTP request feature vector, and performing Laplace mapping processing on the real-time HTTP request feature vector to obtain a real-time feature vector after dimensionality reduction; Building a feature similarity evaluation model based on the CC attack feature parameter set, and inputting the reduced-dimensional real-time feature vector into the feature similarity evaluation model to perform feature matching calculation to obtain a feature matching score; A comprehensive risk assessment is performed based on the feature matching score and the session behavior information of the real-time HTTP request to be filtered to obtain an attack possibility score.
5. The traffic filtering method based on CC attack characteristics according to claim 4 is characterized in that: The step of constructing a feature similarity evaluation model based on the CC attack feature parameter set and inputting the reduced-dimensional real-time feature vector into the feature similarity evaluation model to perform feature matching calculation to obtain a feature matching score includes: Extracting feature dimension indexes from the CC attack feature parameter set to obtain a key feature dimension index set, and extracting feature values of corresponding dimensions from the reduced real-time feature vector according to the key feature dimension index set to obtain a real-time key feature vector; Building a feature similarity evaluation model , where x is the real-time key feature vector, i represents the index, and w i is the feature weight value of the corresponding dimension in the CC attack feature parameter set, B i (x i ) is the boundary judgment function, n is the number of dimensions of the key feature dimension index set, and the boundary judgment function B i (x i ) is defined as when x i When B falls between the upper and lower limits of the characteristic i (x i )=1, otherwise B i (x i ) is reduced to a value between 0 and -1 according to the distance from the boundary, and S(x) represents the feature matching score; Each dimension value xi of the real-time key feature vector is substituted into the feature similarity evaluation model to obtain the boundary matching result of each dimension, and the boundary matching results of all dimensions are weighted and normalized to obtain the feature matching score.
6. The traffic filtering method based on CC attack characteristics according to claim 1 is characterized in that: The time-varying parameter defense analysis is performed based on the attack likelihood score to obtain an adaptive defense rule set, including: Constructing an attack situation vector based on the attack possibility score, wherein the attack situation vector includes attack intensity, attack variability, and attack distribution characteristics within the current time window; Constructing a game participant set based on the attack situation vector, wherein the game participant set includes a CC attacker, a defense controller, and a resource scheduler; Defining a game optimization goal according to the set of game participants, and performing time-varying parameter analysis based on the attack situation vector and the game optimization goal to obtain a parameter dynamic update rule; Optimizing the parameter dynamic update rule by solving the optimal control equation to obtain an optimal defense strategy set, wherein the optimal defense strategy set includes JavaScript challenge verification strength, session behavior analysis threshold, and resource allocation ratio; A hierarchical combination is performed on the optimal defense strategy set to obtain an adaptive defense rule set, wherein the adaptive defense rule set includes multiple defense levels, each defense level corresponding to a different attack possibility score interval and a corresponding defense measure combination.
7. The traffic filtering method based on CC attack characteristics according to claim 1 is characterized in that: The step of performing hierarchical filtering on the real-time HTTP request to be filtered according to the adaptive defense rule set to obtain filtered safe traffic includes: Classifying the real-time HTTP requests to be filtered into defense levels according to the attack possibility scores of the real-time HTTP requests to be filtered, and obtaining HTTP requests with level marks; Performing JavaScript challenge verification on the HTTP requests with the level mark and the scores exceeding the first threshold, to obtain a JavaScript verification result; Performing CAPTCHA verification on the HTTP requests with the level mark, the requests having a score exceeding a second threshold and failing the JavaScript verification result, to obtain a verification response result, wherein the second threshold is greater than the first threshold; Update the reputation score of the request source IP based on the JavaScript verification result and the verification response result to obtain an updated IP reputation database, and feed the updated IP reputation database back to the adaptive defense rule set for rule update; Dynamically allocating processor resources, memory buffers, and network bandwidth to the HTTP requests with level tags to obtain resource scheduling results; According to the resource scheduling result and the verification pass status, the real-time HTTP request to be filtered is passed, delayed responded or rejected to obtain filtered safe traffic.
8. A traffic filtering device based on CC attack characteristics, characterized in that: For implementing the traffic filtering method based on CC attack features according to any one of claims 1 to 7, the traffic filtering device based on CC attack features comprises: The mapping and dimensionality reduction module is used to perform feature extraction and Laplace feature mapping dimensionality reduction on historical HTTP request traffic samples to obtain a set of reduced-dimensional feature vectors; A feature analysis module, configured to perform CC attack feature analysis based on the dimension-reduced feature vector set to obtain a CC attack feature parameter set; A risk assessment module, configured to perform an attack risk assessment on the real-time HTTP request to be filtered based on the CC attack feature parameter set to obtain an attack possibility score; a defense analysis module, configured to perform time-varying parameter defense analysis based on the attack likelihood score to obtain an adaptive defense rule set; The hierarchical filtering module is used to perform hierarchical filtering processing on the real-time HTTP requests to be filtered according to the adaptive defense rule set to obtain filtered safe traffic.
Citation Information
Patent Citations
CC attack detection method and CC attack detection device
CN114499917A
Artificial intelligence enhanced distributed denial of service attack defense method and system
CN119865343A