Network threat detection method and device based on fusion of rule engine and machine learning
By integrating rule engines and machine learning into a network threat detection method, a closed-loop architecture is constructed to solve problems such as insufficient ability to identify unknown attacks, high false alarm rate, and data fragmentation in existing systems. This achieves efficient network security protection, adapts to high-speed network environments, and improves detection accuracy and response efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUJIAN POLYTECHNIC OF WATER CONSERVANCY & ELECTRIC POWER
- Filing Date
- 2026-04-08
- Publication Date
- 2026-05-15
AI Technical Summary
Existing network security protection systems are insufficient in their ability to identify unknown attacks and attack variants, have a high false alarm rate, cannot adapt to high-speed network scenarios, have a low degree of automation in security response, have fragmented data in system modules, lack long-term self-iterative capabilities in detection models, and lack real-time security situation awareness and attack tracing capabilities.
We adopt a network threat detection method based on the fusion of rule engine and machine learning, and construct a five-layer integrated distributed technical architecture with a closed loop of data collection, intelligent detection, policy execution, visual tracing, and data support. Through parallel detection by rule engine and machine learning dual engines, we differentiate the allocation of decision weights, set up a dual verification mechanism, and design adaptively adjustable dynamic decision thresholds and automated handling strategies to achieve a closed-loop self-learning model driven by data throughout the entire process.
It improves the ability to identify unknown attacks and attack variants, reduces false alarm rates, improves detection accuracy and response efficiency, adapts to high-speed network environments, breaks down data silos, provides end-to-end network security protection, simplifies operation and maintenance processes, and lowers the threshold and cost of operation and maintenance.
Smart Images

Figure CN122053250A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer network security technology, including intelligent network traffic analysis, network threat detection and automated protection technology, specifically involving a network threat detection method and device based on the fusion of rule engine and machine learning. Background Technology
[0002] With the deepening of digital transformation, the business scale carried by the network environments of various entities such as enterprises, government agencies, and educational institutions continues to expand. Network attack methods are showing a trend towards diversification, automation, and rapid iteration, constantly increasing the difficulty and requirements of network security protection. Currently, the mainstream basic protection technologies in the field of network security protection mainly include three categories: traditional firewalls, intrusion detection systems (IDS), and intrusion prevention systems (IPS). Traditional firewalls primarily implement network access control based on static rules such as IP addresses and ports, serving as a fundamental component of network boundary protection. IDS identifies known network attack behaviors and generates alerts through feature matching, achieving attack monitoring. IPS adds basic attack blocking capabilities to IDS. These three technologies, as core basic means of network security protection, have been widely applied across various industries.
[0003] Meanwhile, with the development of artificial intelligence technology, machine learning technology has also been initially applied in the field of cybersecurity. Anomaly identification schemes based on network traffic feature modeling have emerged, capable of identifying network traffic that deviates from the normal baseline, providing a new technical path for detecting unknown network anomalies. However, most current machine learning-related applications use a single algorithm for simple modeling, failing to effectively integrate with traditional rule-based detection technologies. While they can cover the identification of some unknown anomalies, they generally suffer from a high false alarm rate and lack the ability to integrate with existing security equipment, thus failing to form an end-to-end protection loop from detection to response.
[0004] Current mainstream network security protection systems based on fixed rules have several technical shortcomings in practical applications. At the threat identification level, these systems rely entirely on manually configured fixed feature rule bases, and can only identify known network attack behaviors predefined in the rule base. They are completely unable to identify new attack methods not covered by the rule base, such as zero-day attacks and attack variants, resulting in serious lag in protection and an inability to cope with rapidly evolving network attack situations.
[0005] At the data collection level, the network traffic collection dimensions of existing protection systems are relatively simple, mostly extracting only basic network features such as source IP, destination IP, and port, without comprehensively analyzing deeper features that can reflect traffic anomalies, such as packet length, transmission rate, protocol distribution, and session timing. At the same time, in high-speed network scenarios of 10Gbps and above, existing collection solutions generally suffer from high parsing latency (average latency exceeding 200ms) and high system resource consumption (CPU utilization exceeding 70%), which cannot adapt to the real-time protection requirements of high-bandwidth business scenarios.
[0006] At the security response and handling level, most existing protection systems adopt a handling mode that relies primarily on manual response and secondarily on automated response. Most attack incidents require manual intervention by operation and maintenance personnel, resulting in long response delays (average response time exceeding 30 minutes) and making it easy for the damage from attacks to spread horizontally. At the same time, existing solutions lack graded handling strategies based on the attack risk level, and most adopt simple handling actions such as indiscriminate IP blocking, which can easily cause normal business interruption due to wrongful blocking, and cannot achieve a balance between protection capabilities and business availability.
[0007] At the system architecture level, most of the functional modules of the existing protection system are independent of each other, lacking a unified data analysis and management platform. Traffic data, alarm data, handling data, and operation and maintenance data cannot be linked for analysis, forming serious "data silos". Operation and maintenance personnel need to switch between multiple independent systems, which not only increases operation and maintenance costs but also reduces the efficiency of handling security incidents.
[0008] In addition, existing single machine learning detection solutions generally have a false positive rate exceeding 15%, and the model performance cannot be iteratively optimized with the accumulation of attack samples, resulting in a continuous decline in the generalization ability against new attacks over time. At the visualization and operational support level, most existing protection systems only provide basic log query functions, lacking core capabilities such as real-time security situation awareness, attack chain tracing, and risk trend early warning. Operations personnel cannot intuitively grasp the overall network security status, nor can they quickly locate the attack source or reconstruct the complete attack path, failing to achieve end-to-end network security protection encompassing pre-incident warning, in-incident handling, and post-incident tracing. Summary of the Invention
[0009] To address the shortcomings and deficiencies of existing technologies, this invention provides a network threat detection method and device based on the fusion of rule engines and machine learning. It aims to solve industry pain points in traditional network security protection solutions, such as the lack of ability to identify unknown attacks and attack variants within fixed rule systems, high false positive rates with single machine learning detection schemes, poor adaptability to high-speed network scenarios, insufficient automation of security responses, fragmented data across system modules, and the lack of long-term self-iterative capabilities in detection models. This invention constructs a five-layer integrated distributed technical architecture with a closed-loop process encompassing data collection, intelligent detection, policy execution, visual tracing, and data support. The core design employs a parallel detection approach using both rule engines and machine learning, differentiating decision weights for known attacks and unknown anomalies: in known attack scenarios, priority is given to ensuring identification accuracy, allocating higher weight to rule-based detection results; in unknown anomaly scenarios, priority is given to ensuring the detection of novel attacks, allocating higher weight to machine learning detection results. Meanwhile, this invention employs a dual verification mechanism for unknown abnormal traffic identified by machine learning. It completes secondary verification by expanding the rule base and using historical behavior data from corresponding sessions. This improves the detection rate of unknown and variant attacks while keeping the false alarm rate consistently below 1% (verified based on over 100,000 real network traffic samples). The invention also includes an adaptively adjustable dynamic decision threshold, an automated handling strategy based on quantitative grading of attack threat levels, and a data-driven, closed-loop self-learning model system. The machine learning detection module uses a dual-algorithm fusion architecture to extract statistical and temporal features of the traffic for anomaly identification, further enhancing detection accuracy and generalization capabilities. The solution is compatible with existing mainstream network security equipment and natively adaptable to high-speed and even ultra-high-speed network environments. It achieves full-process network security protection, including pre-emptive warning, in-process handling, and post-incident tracing. The overall detection accuracy remains stable above 98%, and the automated handling latency is controlled at the millisecond level, significantly reducing the threshold and cost of network security operation and maintenance. It can be widely applied to security protection scenarios in various wired and wireless network environments, including enterprises, government agencies, and educational institutions.
[0010] The specific technical solution adopted by this invention to solve its technical problem is as follows:
[0011] A network threat detection method based on the fusion of rule engine and machine learning includes the following steps:
[0012] Collect traffic data, log data, and security device alarm data from the target network environment, clean and standardize the collected data, and extract standardized feature data for detection.
[0013] The standardized feature data is input into the rule engine and the machine learning detection model in parallel to obtain the rule matching result, the rule matching score, and the machine learning comprehensive anomaly probability, respectively.
[0014] Based on whether the rule matching result matches the preset core threat rule, the attack scenario to which the current traffic belongs is divided. The attack scenario includes known attack scenarios and unknown abnormal scenarios.
[0015] Based on the attack scenario to which the current traffic belongs, differentiated dynamic weights are assigned to the rule matching score and the comprehensive anomaly probability of machine learning, and the weighted fusion is used to obtain the total decision score; among them, the weight of the rule matching score is greater than the weight of the comprehensive anomaly probability of machine learning in the known attack scenario, and the weight of the comprehensive anomaly probability of machine learning is greater than the weight of the rule matching score in the unknown anomaly scenario.
[0016] Initial attack determination is made based on the total decision score and dynamic decision threshold. For traffic that is initially determined to be an attack and belongs to an unknown abnormal scenario, dual verification is performed by matching the extended rule base and the historical behavior data of the corresponding session.
[0017] The final attack determination is completed by combining the preliminary attack judgment and the results of dual verification. The traffic that is finally determined to be an attack is classified into threat levels, and automated handling actions are executed according to the preset graded handling strategy corresponding to the threat level.
[0018] Furthermore, the dynamic decision threshold adopts an adaptive adjustment mechanism: for traffic that is initially judged to be an attack and belongs to an unknown abnormal scenario, when the number of samples that pass dual verification reaches 5 consecutive times, the dynamic decision threshold is lowered by 10%; when the system's daily false alarm rate exceeds 5%, the dynamic decision threshold is raised by 15%.
[0019] Furthermore, the execution logic of the dual verification is as follows: for traffic initially determined to be an attack in an unknown abnormal scenario, a secondary matching is performed by calling an extended rule library covering non-core threat rules, abnormal behavior rules, and historical attack association rules, and at the same time, historical access behavior data of the corresponding session is retrieved for anomaly comparison. The final attack determination is confirmed based on the results of the secondary matching and comparison.
[0020] Furthermore, the machine learning detection model adopts a dual-algorithm fusion architecture. It extracts features from different dimensions of standardized feature data through two machine learning sub-models, outputs anomaly probabilities separately, and then fuses them with weights to obtain a comprehensive machine learning anomaly probability. The two machine learning sub-models include a random forest model and an XGBoost model. The random forest model is used to extract statistical features and output the first anomaly probability, while the XGBoost model is used to extract temporal features and output the second anomaly probability. The fusion weight is dynamically adjusted based on the accuracy of the validation set: when the accuracy of the random forest model is higher than that of the XGBoost model, it is assigned 60% weight, and otherwise, it is assigned 40% weight.
[0021] Furthermore, the threat level classification is based on the quantification of the threat level of the attack behavior, and different threat levels correspond to different handling strategies: low-risk attacks only trigger alarm notifications, medium-risk attacks trigger temporary blocking, and high-risk attacks trigger permanent blocking and network-wide collaborative protection.
[0022] Furthermore, the standardized feature data covers the basic attribute features, traffic statistics features, and session temporal features of network sessions, providing a unified feature input for rule matching and machine learning detection.
[0023] Furthermore, it also includes a closed-loop self-optimization step for the machine learning model: collecting attack detection data, handling log data and manually reviewed and labeled data from the entire system process, selecting effective samples to incrementally train and optimize the machine learning detection model, injecting adversarial perturbation samples generated based on the FGSM algorithm during the training process to enhance the model's generalization ability, completing the model's autonomous iterative update, setting the optimization cycle to once a week, and triggering emergency optimization when the model's accuracy drops by more than 3%.
[0024] Furthermore, it also includes an attack tracing step: constructing an attack propagation directed graph based on an improved graph computing algorithm. The improved graph computing algorithm is the PageRank algorithm, which introduces attack behavior time weights (weights of 1.0 in the last 24 hours and weights that decay to 0.5 in the last 24-72 hours) and success probability weights (dynamically calculated based on historical attack success rates). This identifies attack-related nodes and groups, reconstructs the complete attack chain, and enables rapid location of the attack source and attack path.
[0025] And, a network threat detection device based on the fusion of rule engine and machine learning, comprising:
[0026] The data acquisition module is used to collect traffic data, log data, and security device alarm data of the target network environment, clean and standardize the collected data, and extract standardized feature data for detection.
[0027] The dual-engine detection module has a built-in rule engine unit and a machine learning detection unit. It is used to input the standardized feature data into the rule engine and the machine learning detection model in parallel to obtain the rule matching result, the rule matching score, and the machine learning comprehensive anomaly probability, respectively.
[0028] The scenario-based fusion decision module is used to classify attack scenarios based on whether the rule matching result hits the preset core threat rules. Differentiated dynamic weights are assigned to the rule matching score and the comprehensive anomaly probability of machine learning according to the attack scenario, and the weighted fusion is used to obtain the total decision score. Among them, the weight of the rule matching score is greater than the weight of the comprehensive anomaly probability of machine learning in known attack scenarios, and the weight of the comprehensive anomaly probability of machine learning is greater than the weight of the rule matching score in unknown anomaly scenarios.
[0029] The dual verification and attack determination module is used to make preliminary attack determination based on the total decision score and dynamic decision threshold, perform dual verification on traffic that is initially determined to be an attack in unknown abnormal scenarios, and complete the final attack determination by combining the preliminary determination and dual verification results.
[0030] The tiered handling module is used to classify the threat level of traffic that is ultimately determined to be an attack and to execute automated handling actions according to the corresponding tiered handling strategy.
[0031] Furthermore, the device adopts a distributed deployment architecture, realizes multi-node traffic collaborative processing through a consistent hash load balancing strategy, configures a master-slave redundancy mechanism for core nodes, with a master-slave switchover time of no more than 500ms, and the cluster size can be elastically expanded to 100+ nodes to ensure stable system operation under high load scenarios.
[0032] Compared to existing technologies, this invention and its preferred solution, through a detection architecture that integrates a rule engine and machine learning dual engines, combined with a scenario-based dynamic weight allocation mechanism, effectively solves the core pain points of traditional protection solutions: insufficient ability of fixed rules to identify unknown attacks and difficulty in controlling the false positive rate of single machine learning detection. While ensuring accurate identification of known attacks, it also achieves effective detection of unknown attacks and attack variants, balancing the detection rate and accuracy of threat detection. The dual verification mechanism for unknown abnormal traffic identified by machine learning completes the review of attack judgment through multi-dimensional secondary verification, further reducing the risk of false positives, avoiding interference with normal business operations, and improving the reliability of threat detection results. The accompanying adaptive dynamic decision threshold mechanism can dynamically calibrate attack judgment standards according to changes in the network attack landscape, adapting to the evolution trend of new attack variants, avoiding the protection lag caused by fixed thresholds, and ensuring the long-term adaptability of the protection system. The constructed automated handling system based on quantitative grading of attack threat levels can match differentiated protection strategies according to the threat level of attack behavior, balancing network security protection capabilities with the operational availability of business systems, while significantly improving security. The efficient handling of incidents shortens the attack response window, and multi-device collaborative protection prevents the lateral spread of attack damage. The designed full-process data-driven model closed-loop self-learning optimization mechanism can autonomously iterate and update the model based on all data collected during system operation, continuously optimizing detection performance and avoiding the generalization attenuation problem of static detection models over time, ensuring the continuous effectiveness of the protection system. The integrated distributed architecture adopted connects the data links of data collection, threat detection, policy execution, and operation and maintenance management, breaking the data silo problem commonly found in traditional protection systems. It possesses excellent scalability and compatibility, adapting to network environments of different sizes and bandwidths. It can seamlessly interface with existing mainstream network security equipment, significantly reducing the deployment and migration costs of the system. The accompanying attack tracing and network-wide security situation awareness capabilities help operations and maintenance personnel intuitively grasp the overall network security status, quickly locate the source of attacks, and reconstruct the complete attack chain. This effectively simplifies the operation process of network security operations and maintenance, lowers the threshold and workload of operations and maintenance work. According to actual testing, the overall detection accuracy of the system reaches 98.5%, the false alarm rate is controlled within 0.8%, and the average latency of automated processing is 420ms. Attached Figure Description
[0033] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments:
[0034] Figure 1 This is a general framework diagram of an embodiment of the present invention. Detailed Implementation
[0035] To make the features and advantages of the present invention more apparent and understandable, specific embodiments are described below in detail:
[0036] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0037] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0038] To address the pain points in the current cybersecurity protection field, such as the lack of ability to identify unknown and variant attacks by traditional fixed-rule protection systems, insufficient data collection dimensions and performance, low degree of automation in security response, severe data silos, high false alarm rate and lack of self-iteration capability of single machine learning detection schemes, and weak visualization and attack tracing capabilities, this invention proposes a network threat detection method and device based on the fusion of rule engines and machine learning, aiming to achieve:
[0039] Construct a dual threat identification system that accurately identifies known attacks and effectively detects unknown attacks, while keeping the false alarm rate at an extremely low level, thus filling the gap in the existing system's ability to identify zero-day attacks and attack variants.
[0040] Efficient collection and standardized processing of multi-dimensional traffic characteristics and multi-source heterogeneous data solves the problems of single collection dimensions, high latency in high-speed traffic parsing, and large resource consumption in existing solutions, and is natively adapted to 10Gbps and above high-speed and even ultra-high-speed network environments.
[0041] Build an automated graded response system based on the attack risk level to achieve millisecond-level automated handling of attack behavior, balancing protection capabilities and business availability, and avoiding accidental interruption of normal business.
[0042] Build an integrated intelligent analysis platform to connect the entire data chain from data collection, threat identification, policy execution, data storage, and operation and maintenance management, realize closed-loop data flow, and break down the data silos of existing systems;
[0043] The design of the rule engine combines a fusion detection mechanism with machine learning, and is equipped with a closed-loop model self-learning and iteration system to solve the problems of high false alarm rate and detection performance degradation over time of single machine learning models.
[0044] It enables real-time visualization of the entire network security situation, rapid tracing of attack behavior across the entire chain, and early warning of potential risks, significantly reducing the threshold and operating costs of network security operations and maintenance.
[0045] To achieve the above objectives, this solution constructs a five-layer integrated distributed technical architecture: data acquisition layer, intelligent analysis layer, strategy execution layer, visualization layer, and data storage layer. Each layer achieves end-to-end data linkage through standardized API interfaces, forming a closed-loop protection system covering the entire lifecycle: multi-source data acquisition, dual-engine fusion detection, hierarchical automated handling, full-dimensional situational awareness, and model self-iterative optimization. The solution adopts a B / S architecture, with the backend developed using a hybrid Python+Go language. Go is used for modules with high real-time requirements, such as data acquisition, rule engine matching, and multi-threaded concurrent processing, while Python is used for modules with high algorithmic flexibility, such as feature engineering, machine learning model training and inference, and attack attribution graph analysis, balancing system real-time performance with algorithmic iteration flexibility. The frontend uses Vue3+ElementPlus for interactive visualization, and the core database uses a hybrid storage architecture of PostgreSQL+Redis+Neo4j. The system is deployed in a distributed manner, supporting elastic node expansion and natively adaptable to high-speed network environments of 10Gbps and above. It is also compatible with existing firewalls, WAFs, IPS, and other network security devices, eliminating the need to replace existing user security assets and significantly reducing deployment and migration costs.
[0046] The main innovations of this solution include: using a rule engine to ensure accurate identification of known attacks; using a random forest + XGBoost dual-algorithm machine learning model to achieve generalized detection of unknown and variant attacks; balancing detection detection rate and false alarm rate through a scenario-based dynamic weight fusion mechanism and a dual verification mechanism; supporting model self-learning iteration with a closed-loop data process to ensure continuous optimization of detection capabilities as attack samples accumulate; achieving rapid closed-loop handling of attacks through a graded automated response system driven by attack risk levels, balancing protection effectiveness and business availability; and ultimately building an intelligent network security protection system that balances accuracy, generalization, real-time performance, and availability, achieving full-process network security protection including pre-event warning, in-event handling, and post-event tracing.
[0047] To facilitate a complete understanding of the technical solution of this application, the underlying standards, publicly available datasets, compatible interface specifications, and core terminology of this solution are explained as follows:
[0048] 1. Basic Protocols and Standards: This system's network traffic parsing is based on the TCP / IP protocol suite and the IEEE 802.3 Ethernet standard; third-party device integration is compatible with SNMPv3 Simple Network Management Protocol and RESTful API standard.
[0049] 2. Model Training Dataset: The publicly available datasets used for training the machine learning model in this application are the KDD Cup 99 network traffic dataset and the UNSW-NB15 network attack dataset, supplemented by the latest CSE-CIC-IDS2018 / 2019 attack dataset to cover new threat samples such as ransomware and IoT attacks;
[0050] 3. Definitions of core terms:
[0051] Rule engine: A component embedded in an application to separate business rules from business code, enabling flexible configuration and modification of business rules;
[0052] Machine learning: a multidisciplinary field that uses algorithms to enable computers to learn patterns from data and make predictions and analyses of unknown data;
[0053] Attack attribution: The technical process of locating the source of an attack, reconstructing the attack path, and identifying the attack methods by correlating and analyzing relevant data on network attack behaviors;
[0054] Incremental training: Based on the pre-trained base model, supplementary training is performed using newly added sample data, eliminating the need for full retraining. This improves the efficiency of model iteration while retaining the original model's detection capabilities.
[0055] like Figure 1 As shown below, the main components of this solution will be further illustrated and described:
[0056] 1. Data Acquisition Layer
[0057] The data acquisition layer is the core data input foundation of the solution in this embodiment. It solves the problems of single acquisition dimension and high latency in high-speed traffic parsing in existing technologies. Its output standardized multi-source data is the only input source for the intelligent analysis layer. It adopts a distributed multi-source lightweight acquisition scheme, which includes three sub-modules: network traffic acquisition module, log acquisition agent, and device interface docking module.
[0058] Network traffic collection module
[0059] This module is based on the libpcap library and is a secondary development. It optimizes packet capture and parsing algorithms for high-speed traffic scenarios, supports simultaneous parallel monitoring of multiple network cards, and solves the traffic loss problem of traditional single network card packet capture.
[0060] As a preferred implementation, this module can extract 12 core traffic features, including: source IP, destination IP, source port, destination port, protocol type, packet length, transmission rate, protocol distribution, connection frequency, session duration, uplink / downlink traffic ratio, and retransmission packet ratio. It covers three core dimensions: ① basic session attributes, ② traffic statistics features, and ③ time-series behavior features. Compared with traditional collection schemes that only extract basic IP and port attributes, the feature support is greatly improved, avoiding detection blind spots caused by a single dimension.
[0061] Simultaneously, a lightweight LZ4 lossless compression algorithm is introduced to compress the raw traffic metadata in real time before transmission. According to actual tests, in a 10Gbps line-rate traffic scenario, in an Intel Xeon Gold 6248 server environment, the end-to-end latency from traffic capture to compression and transmission to the detection node can be controlled within 100ms. The peak CPU resource usage of the business-side collection agent is ≤30%, which significantly reduces the impact of the collection process on the performance of the business server.
[0062] As a preferred implementation, for 40Gbps / 100Gbps ultra-high-speed network scenarios, this module introduces the NVIDIA BlueField-3DPU (Data Processing Unit) hardware acceleration solution. This offloads computationally intensive operations such as traffic parsing and feature extraction to the dedicated DPU processing core, eliminating the need to occupy host CPU resources and reducing host CPU utilization to below 15%. Simultaneously, a traffic sharding parallel processing mechanism is adopted, hashing the raw traffic according to session ID and distributing it to different CPU cores or DPU cores for parallel processing, increasing single-node traffic processing capacity by more than 3 times. An adaptive sampling algorithm is designed, automatically enabling a 1:10 sampling ratio when the real-time traffic rate exceeds 20Gbps. During sampling, key session data packets containing attack characteristics are prioritized for retention through DDoS feature library matching and abnormal connection entropy calculation, ensuring that core detection features are not lost and achieving low-latency, low-resource-consumption detection in ultra-high-speed network scenarios.
[0063] Log collection agent
[0064] This module is a lightweight and deployable program that can be deployed in batches on various nodes such as servers, terminals, and network devices. It collects various types of log data, including system logs, application logs, device operation logs, and security alarm logs. It supports automatic parsing and standardized processing of mainstream log formats such as syslog and json, solving the problem of inconsistent log formats from multiple sources and the inability to perform linked analysis.
[0065] Device interface module
[0066] This module provides a standardized RESTful API interface, compatible with the SNMP simple network management protocol, and can seamlessly connect with existing security and network devices such as firewalls, WAFs, IPSs, switches, and routers to achieve real-time synchronization of device status data, alarm data, and handling logs. All collected multi-source data is standardized and then uniformly pushed to the intelligent analysis layer for further processing, while simultaneously being synchronized to the data storage layer for persistent storage.
[0067] 2. Intelligent Analysis Layer
[0068] The intelligent analysis layer serves as the decision-making center of this embodiment and is the core innovative improvement of this application. It addresses the problems of insufficient threat identification capabilities and high false alarm rates in existing technologies. Its input is standardized multi-source data output from the data acquisition layer, and the output attack judgment and level labeling results serve as the direct basis for the policy execution layer's actions. This layer employs a rule engine combined with a machine learning fusion detection scheme, comprising four sub-modules: feature extraction, rule matching, machine learning detection, and model self-learning. Data from each module is linked in real time to form a detection closed loop. Dynamic weight fusion utilizes the Sigmoid function to quantify the scenario confidence: given the attack scenario weight coefficient W... rule =0.7 + 0.2 * Sigmoid (rule matching score - 0.5), weight coefficient W for unknown and abnormal scenarios. ml =0.7 + 0.2 * Sigmoid (machine learning anomaly probability - 0.5).
[0069] Feature extraction module
[0070] This module performs deep feature extraction and preprocessing on the standardized multi-source data pushed by the acquisition layer to generate standardized feature vectors, providing a unified data foundation for subsequent rule matching and machine learning detection, and solving the problem of inconsistent multi-source data formats that cannot be directly used for detection.
[0071] As a preferred implementation, this module processes the collected raw features according to the following process to generate a standardized feature vector with a unified input:
[0072] (1) Data cleaning
[0073] First, invalid data is removed, including checksum errors, abnormal data packets exceeding the protocol length, and duplicate log reports from the same session. Second, missing values are filled: continuous features are filled with the mean within the same session's time window, and discrete features are filled with the mode of the same feature dimension. Sessions with a missing value rate exceeding 30% are directly judged as invalid sessions (this threshold is determined based on statistics from 3 million+ real network session data) and do not enter the subsequent detection process to avoid invalid data interfering with detection accuracy.
[0074] (2) Feature classification and coding
[0075] The 12 core traffic features are divided into 5 discrete features and 7 continuous features. Differentiated coding methods are used for different feature types to eliminate the influence of units and value ranges on the detection results.
[0076] Discrete features: divided into two categories and encoded separately, resulting in a total of 20 dimensions.
[0077] ① The high cardinality discrete features consist of 4 items: source IP, destination IP, source port, and destination port (with a very large value range, not suitable for one-hot encoding). Each item is mapped to a 4-dimensional vector using hash encoding, and the 4 items together have 16 dimensions.
[0078] ② The low cardinality discrete feature consists of 1 item: protocol type (containing only 4 types of values: TCP, UDP, ICMP, and unknown protocol), which is mapped to a 4-dimensional vector using one-hot encoding;
[0079] There are 7 continuous features: packet length, transmission rate, protocol distribution, connection frequency, session duration, uplink / downlink traffic ratio, and retransmission packet percentage. Among them, packet length and transmission rate are statistical features within a 5-minute sliding time window (each item has two dimensions: maximum value and mean value). Protocol distribution is the percentage of 4 types of protocols within the window (4 dimensions in total). The remaining 4 items are single-dimensional features, with a total of 12 original dimensions. Each dimension of all continuous features is individually normalized using min-max and mapped to the [0,1] interval.
[0080] (3) Construction of feature vectors
[0081] The encoded features are concatenated in a fixed order: discrete encoding result → normalized continuous features, to generate a 32-dimensional standardized feature vector (20 dimensions for discrete encoding and 12 dimensions for continuous features), which serves as the unified input for the subsequent rule matching module and machine learning detection module.
[0082] The standardized feature vectors are input synchronously and in parallel into the rule matching module and the machine learning detection module. The two modules process independently without dependency, which can further improve detection efficiency.
[0083] Rule matching module
[0084] This module is the core detection unit for known attacks. It has more than 2,000 industry-standard threat rules built in, and also supports user-defined rules and rule blacklists and whitelists. It can quickly and accurately identify known network attacks such as port scanning, DDoS attacks, SQL injection, XSS cross-site scripting, and weak password brute-force attacks.
[0085] As a preferred implementation, this module employs the AC automaton multi-pattern matching algorithm to optimize rule matching efficiency, achieving a rule matching response time of ≤50ms for a single traffic entry. The rule base is divided into core rules and non-core rules based on risk level. Core rules are mandatory rules covering high-risk known attacks, while non-core rules are auxiliary rules covering probing attacks and abnormal behaviors, allowing users to adjust them as needed. After rule matching is completed, the rule matching results are output, including whether a rule was hit, the number of hit rules, the highest risk level of the hit rules, and the rule matching score P. rule (Range 0-1, the higher the danger level and the greater the number of hits, the higher the P) rule The closer it is to 1), the more it is pushed to the fusion decision module and the machine learning detection module simultaneously.
[0086] The above core threat rules can be categorized based on the CVSS general vulnerability scoring system, and attack rules corresponding to high-risk vulnerabilities with CVSS scores ≥ 7.0 can be included in the core rule base.
[0087] Machine learning detection module
[0088] This module is the core detection unit for unknown attacks and variant attacks. It integrates the random forest and XGBoost dual algorithm models and adopts a dynamic weight fusion mechanism to solve the problem that a single algorithm cannot simultaneously cover statistical feature anomalies and temporal feature anomalies, thus achieving effective detection of unknown attacks.
[0089] As a preferred implementation, this module adopts a detection scheme that combines random forest and XGBoost algorithms with dynamic weights from a rule engine and a machine learning engine. The specific process is as follows:
[0090] (1) Dual-algorithm single-model inference: The standardized feature vectors are input into the pre-trained random forest model and XGBoost model respectively: The random forest model is responsible for extracting the statistical regularity of traffic features (such as abnormal connection frequency, abnormal data packet length distribution) and outputting the first anomaly probability for the statistical feature anomaly. (Range 0-1); The XGBoost model focuses on temporal feature analysis (such as traffic mutation trends, abnormal session durations, and abnormal access timing), and outputs a second anomaly probability for temporal feature anomalies. (Range 0-1).
[0091] (2) Single model result fusion: The overall anomaly probability of the machine learning model is calculated by weighted summation: P ml =0.6*P rf + 0.4*P xgbThe weighting is based on the following: Firstly, statistical anomalies in network attack behavior have a wider coverage, while time-series anomalies are more suitable for scenarios such as sudden DDoS attacks and brute-force attacks, which can balance the detection capabilities of the two types of features and avoid missed detections due to a single feature dimension; secondly, traffic statistical features contribute more to the identification of known attack variants, while time-series features are more sensitive to the identification of new and unknown attacks, which can balance the generalization ability and identification accuracy of the model at the same time. The weight values are determined by optimizing the F1 score on the validation set (adjusted to 0.7 / 0.3 when the F1 score of the random forest model is higher than that of the XGBoost model).
[0092] (3) Setting and basis of core parameters for dual models:
[0093] Random Forest Sub-model: Tree depth 15-20 layers (optimal tree depth determined by 5-fold cross-validation is 18 layers), minimum number of samples per leaf node ≥ 5, feature subset sampling rate 0.7. If the tree depth is less than 15 layers, the nonlinear correlation of traffic features cannot be fully extracted, leading to a decrease in the accuracy of complex attack identification; if the tree depth is greater than 20 layers, overfitting is likely to occur, reducing the generalization ability to unknown variant attacks and increasing inference time; this parameter setting can fully extract feature correlations while avoiding overfitting.
[0094] XGBoost sub-model: learning rate 0.05, maximum tree depth 8, L2 regularization coefficient 0.1, using 5-fold cross-validation to optimize parameters, and using regularization strategy to avoid model overfitting and ensure model stability under different sample distributions.
[0095] (4) Model pre-training scheme: The KDD Cup 99 and UNSW-NB15 mixed network attack datasets are used. The training set and test set are divided in a 7:3 ratio. The SMOTE oversampling algorithm (adjusting the ratio of attack samples to normal samples to 1:3) is used to solve the problem of imbalance between attack samples and normal samples, and to ensure the model's ability to identify minority attack samples.
[0096] (5) Attack scenario determination and dynamic weight allocation of dual engines: Based on the rule matching results, the traffic is scenario-based, the weights of the rule engine and the machine learning engine are allocated differently, and the total decision score is calculated:
[0097] If the rule matching result matches at least one core threat rule, it is determined to be a known attack scenario, and the total decision score is: S. total = 0.8*P rule + 0.2*P ml This weight allocation prioritizes ensuring the accuracy of identifying known attacks and reduces the false positive rate;
[0098] If the rule matching result is no match for any core threat rules, only a match for non-core rules, or no match for any rules, but the machine learning overall anomaly probability P ml > 0.7, judged as an unknown abnormal traffic scenario, total decision score: S total =0.3*P rule + 0.7*P ml This weight allocation prioritizes the detection capability of unknown attacks, avoiding the omission of new or variant attacks not covered by the rule base.
[0099] (6) Initial attack determination: If the total decision score S total If the value is ≥ 0.92, the traffic is determined to be attack traffic and enters the two-factor authentication and attack level marking process; otherwise, it is determined to be normal traffic, and only the access log is recorded without triggering subsequent processing.
[0100] Dual verification and dynamic threshold adjustment module
[0101] This module is the core innovation of this solution. It addresses the pain point of false positives when machine learning detects unknown attacks by performing a secondary verification on unknown abnormal traffic that is initially determined to be an attack. At the same time, it balances the detection detection rate and false positive rate by adjusting the dynamic threshold. Its input is the total decision score output by the fusion decision module and the preliminary attack judgment result. The output is the preliminary basis for the final attack judgment result and attack level label.
[0102] As a further preferred implementation method, the specific execution flow of this module is as follows:
[0103] (1) Triggering conditions: The dual verification mechanism is triggered only when the rule matching result does not hit any core threat rules, belongs to an unknown abnormal scenario, and is initially determined to be attack traffic based on the total decision score; the hit traffic in the known attack scenario does not need to trigger dual verification and directly enters the attack level marking process.
[0104] (2) Secondary verification method: The rule engine performs extended rule base matching for the abnormal traffic. The extended rule base includes industry-standard non-core threat rules, user-defined abnormal behavior rules, and related feature rules of historical attack events. At the same time, it retrieves the source IP and historical behavior data of the session corresponding to the traffic and matches historical attack records and abnormal access records.
[0105] (3) Verification result processing: If the secondary verification hits at least one extended rule or matches a historical abnormal behavior record, the verification is passed, the attack judgment result is maintained, and the attack level marking process is entered; if the secondary verification does not hit any extended rule and there is no historical abnormal behavior record, it is judged as a suspected false alarm, the sample is marked as a sample to be reviewed, and pushed to the visualization display layer for operation and maintenance personnel to review. Blocking actions are not performed temporarily to avoid misjudgment affecting normal business.
[0106] (4) Threshold dynamic adjustment mechanism: The decision threshold is adaptively calibrated based on system operation data. The basic threshold is determined by ROC curve analysis of 100,000+ attack samples, and the initial threshold is set to 0.92. When the system has 5 consecutive samples with "the core rule of the rule engine is not matched, the machine learning model identifies it as abnormal, and the secondary verification is passed", it is judged that a new attack variant has appeared, and the decision threshold is automatically lowered by 10% to improve the detection rate of unknown threats. When the system's daily false alarm rate exceeds 5%, the decision threshold is automatically raised by 15% to tighten the attack judgment standard and ensure the detection accuracy. The threshold adjustment step size is fixed at 0.03, and the number of threshold adjustments per day does not exceed 2 to avoid frequent threshold fluctuations that lead to unstable detection results.
[0107] Attack level marking module
[0108] This module calculates the threat entropy value of verified attack traffic using the entropy method. Combined with preset quantitative standards, it marks the attack traffic into three levels: low risk, medium risk, and high risk. It also includes key information such as the attack source IP, attack type, attack time, affected assets, and attack session details, and pushes it to the policy execution layer.
[0109] As a further preferred embodiment, the threat entropy value is calculated as follows:
[0110] Using single-IP attack behavior as the statistical unit, three core evaluation indicators were selected as input parameters for entropy value calculation: x1 = number of attack attempts per unit time (unit time is 5 minutes), x2 = number of hits on core threat rules, and x3 = number of historical abnormal behavior matches. After normalizing the three indicators, the weight of each indicator was calculated using the entropy method, and the comprehensive threat entropy value H of the attack behavior was finally obtained. The calculation formula is as follows:
[0111] H = w1×x1' + w2×x2' + w3×x3'
[0112] Where x1', x2', and x3' are normalized index values, w1, w2, and w3 are the weights of each index calculated by the entropy method, and the sum of the weights is 1; the threat entropy value H ranges from 0 to 1, and the higher the value, the higher the threat level of the attack behavior.
[0113] As a preferred implementation method, the attack level quantification and grading standard is as follows:
[0114] Low-risk attack: Single IP attack attempts < 5 times / minute, no core rules are triggered, threat entropy value < 0.3, corresponding to scanning and probing behavior with no substantial harm;
[0115] Medium-risk attack: Single IP attack attempts ≥ 5 times / minute, or triggering 1 core rule, threat entropy value 0.3-0.7, corresponding to penetration attempts with clear attack intent but unsuccessful;
[0116] High-risk attack: Attack attempts by a single IP address ≥ 20 times per minute, or triggering 2 or more core rules, with a threat entropy value > 0.7, corresponding to attack behaviors that have caused harm or seriously threaten core assets.
[0117] Model self-learning module
[0118] This module is a new core module added to this application. It solves the problems of existing technology models lacking self-learning ability and performance decaying over time. Through closed-loop data collection throughout the entire process, it enables the model to autonomously iterate and optimize, continuously improving detection capabilities.
[0119] As the preferred implementation method, the specific execution method of this module is as follows:
[0120] (1) Closed-loop collection of training data: Through a unified intelligent analysis platform, closed-loop data of the entire system process is collected, including: attack alarm data of the rule matching module, anomaly detection data of the machine learning detection module, handling log data of the strategy execution layer, and verification and annotation data (false alarm samples and missed alarm samples) of the operation and maintenance personnel of the visualization display layer, which solves the problem of disconnect between model training data and actual operation data in the existing technology.
[0121] (2) Sample screening and preprocessing: A sliding time window mechanism is adopted, with a default window size of 7 days. Only valid samples from the most recent 3 months are retained for incremental training to avoid the decline in the generalization ability of the model due to historical outdated data. The collected samples are screened to remove duplicate and invalid samples. The samples are divided into positive samples (confirmed attack samples, including attack samples that hit the rules and unknown attack samples confirmed by the operation and maintenance personnel) and negative samples (confirmed normal samples, including false alarm samples marked by the operation and maintenance personnel and normal traffic samples that did not trigger any alarms).
[0122] (3) Adversarial training enhancement: 5% of adversarial perturbation samples are injected into the training set. Specifically, the FGSM algorithm is used to add random perturbations within ±10% of the continuous features (such as data packet length and transmission rate) of positive samples, and to add a small amount of noise to the discrete features to simulate the feature changes of attack variants and improve the model's generalization ability to attack variants.
[0123] (4) Incremental training and optimization: Based on the pre-trained basic model, incremental training is performed using the latest selected samples. The training parameters are set as follows: learning rate 0.01 (1 / 5 of the initial training learning rate), and training rounds up to 20 rounds. An early stopping strategy is adopted, and training is terminated when the accuracy of the validation set does not improve for 5 consecutive rounds to avoid overfitting. At the same time, L1 regularization is used to prune the model features, remove redundant features that contribute less than 0.1% to the detection results, and control the model overfitting coefficient (measured by the difference between the accuracy of the training set and the validation set) to within 0.1.
[0124] (5) Model update and canary release: After incremental training is completed, the performance of the new model is verified. The new model is allowed to go online only when the accuracy of the new model on the test set is ≥98% and the false positive rate is ≤1%. The canary release mechanism is adopted for the online release. First, 10% of the traffic is diverted to the new model for parallel inference. The consistency of the detection results of the new and old models is compared in real time (the deviation rate must be ≤2%). After running continuously for 24 hours without any abnormalities, the full replacement is completed to ensure the stability of the system operation.
[0125] 3. Strategy Execution Layer
[0126] The strategy execution layer is the core of the security handling solution in this embodiment, which solves the problems of low automation and lack of hierarchical handling strategies in the existing technology. Its input is the attack level labeling result output by the intelligent analysis layer, and its output is the automated handling action and the network-wide collaborative protection strategy. It adopts an automated hierarchical response scheme and includes three sub-modules: strategy management module, automated handling module, and device linkage module.
[0127] Strategy Management Module
[0128] This module allows users to configure differentiated handling strategies based on the attack risk level, and also has built-in preset default strategies. The strategies can be flexibly modified, saved, and started / stopped, solving the problem that the indiscriminate handling of existing technologies can easily lead to business interruption.
[0129] As a preferred implementation method, the default hierarchical handling strategy is as follows:
[0130] Low-risk attacks: Only generate alarm information and push it to operation and maintenance personnel via SMS, email, system pop-ups, etc., without performing any blocking actions to avoid misjudgment and impact on normal business;
[0131] Medium-risk attacks: Automatically execute temporary IP blocking (default blocking duration is 2 hours, which can be dynamically adjusted according to the historical attack frequency of the IP, and automatically extended to 4 hours when the number of historical attacks is ≥5), corresponding port blocking and handling actions, and push alarm information to operation and maintenance personnel. After the handling actions are executed, the handling log is automatically recorded.
[0132] High-risk attacks: Immediately implement permanent IP blocking, abnormal traffic limiting, and network isolation measures for affected servers. Simultaneously, coordinate with existing firewalls, WAFs, IPSs, and other security devices to synchronize the response strategy to all security nodes across the network, achieving network-wide collaborative protection and preventing the attack from spreading laterally.
[0133] Automated processing module
[0134] This module achieves millisecond-level response based on a multi-threaded concurrent processing mechanism. It sets handling priorities for attacks of different danger levels, with high-risk attack handling threads having the highest priority. This ensures that the total latency from the intelligent analysis layer outputting the attack level label to the handling action taking effect on the corresponding security device is ≤500ms. In an Intel Xeon Silver 4210 server environment, the average latency is 380ms. No manual intervention is required, minimizing the attack response window and reducing the harm of attacks. At the same time, it records the execution details and results of all handling actions, forming a standardized handling log, which is synchronized to the data storage layer and sent back to the model self-learning module of the intelligent analysis layer as training data for model iteration.
[0135] Equipment linkage module
[0136] This module provides a standardized SDK interface that can seamlessly integrate with users' existing network security and network devices to achieve full network synchronization of response actions. This solves the problems of incomplete response from a single node and easy lateral spread of attacks in existing technologies. It also supports full-link backtracking of response actions, and operations and maintenance personnel can view the policy synchronization status and execution results of all devices through a visual interface.
[0137] 4. Visual Presentation Layer
[0138] The visualization layer is the core of human-computer interaction in this embodiment, which solves the problems of weak visualization capabilities, difficulty in attack tracing, and high operation and maintenance costs in existing technologies. It adopts a large-screen visualization + multi-dimensional query solution, which includes four sub-modules: real-time situation large screen, attack tracing module, risk analysis module, and log query module.
[0139] Real-time situation display screen
[0140] This module uses visual charts such as line charts, pie charts, topology diagrams, and heat maps to display core security information in real time, including the overall network traffic status, total number of attacks, distribution of attack types, distribution of affected assets, and execution status of response actions. It supports customizable screen refresh frequency (default refresh every 5 seconds), allowing operations and maintenance personnel to intuitively grasp the overall security status of the entire network through the screen.
[0141] Attack attribution module
[0142] This module, based on graph computing technology, constructs an attack tracing graph, which can quickly locate the source IP of the attack, reconstruct the attack path, identify the attack group, and locate the affected assets, providing a basis for emergency response and post-event review by operation and maintenance personnel.
[0143] As a preferred implementation, the attack attribution graph is constructed as follows: A directed graph of attack propagation is built based on an improved PageRank algorithm. The specific improvements are: time weights and success probability weights are added to the random jump factor of the traditional PageRank algorithm, assigning higher weights to nodes that have occurred recently and have a high probability of success, thus better reflecting the temporal and harmful characteristics of network attacks; the nodes of the directed graph are network entities such as IP addresses, ports, protocols, and assets, and the edges represent access behaviors between entities, with edge weights dynamically calculated based on attack frequency and success probability; the Louvain community discovery algorithm is used to identify associated attack IP groups, and the complete attack chain is reconstructed through time-series correlation analysis, including the entire process of scanning and probing → penetration attempt → privilege escalation → lateral movement → data theft; the Neo4j graph database is used to store and render the graph in real time, supporting graph rendering with 100,000 nodes, and the end-to-end attribution time is ≤3 seconds. The node similarity determination of the Louvain community discovery algorithm can be calculated based on two dimensions: attack behavior feature matching and time window overlap. The time weight T(i) is calculated as T(i) = exp(-0.01*(current time-attack time / 3600)), where λ=0.01 is the decay coefficient and the time unit is hours.
[0144] As a further preferred implementation, the improved PageRank algorithm calculation formula is as follows:
[0145]
[0146] in: denoted as , where is the attack weight value of node i in the directed graph of attack propagation. A higher value indicates a higher probability that the node is the source of the attack or a core attack node. d is the damping coefficient of the traditional PageRank algorithm, with a fixed value of 0.85. T(i) is the time weight of node i. The closer the attack behavior corresponding to the node occurs to the current time, the closer T(i) is to 1, and vice versa. S(i) is the attack success probability weight of node i. The more times the attack behavior corresponding to the node successfully penetrates, the closer S(i) is to 1, and vice versa. , These are the adjustment coefficients for the time weight and the success probability weight, respectively. Preferred =0.4, =0.6; j is the upstream node pointing to node i, L(j) is the number of out-degrees of node j, and w(j,i) is the weight of the access edge from node j to node i. The weight is dynamically calculated according to the attack frequency.
[0147] Risk Analysis Module
[0148] This module performs statistical analysis on historical attack data, alarm data, and response data to generate weekly / monthly / quarterly security risk trend reports. Based on the LSTM time series prediction algorithm, it predicts potential network security risks from attack time series data and provides early warnings for high-frequency attack targets and common attack types, with a prediction accuracy of over 85%, thus achieving proactive protection.
[0149] Log query module
[0150] This module supports searching and collecting logs, detection logs, handling logs, and operation and maintenance logs by multiple dimensions such as time, attack type, risk level, asset IP, and source IP. It supports fuzzy search, precise filtering, export, and printing of logs. The response time for a single log search is ≤100ms, and it supports 1000+ concurrent query requests per second, simplifying log auditing operations for operation and maintenance personnel.
[0151] 5. Data storage layer
[0152] The data storage layer is the core of the end-to-end data support in this embodiment, running through all levels of the system. It provides data storage and fast retrieval services for data collection, detection and analysis, processing and execution, and visualization. It adopts a hybrid storage solution of PostgreSQL + Redis + distributed file storage, and designs differentiated storage methods for different types of data. This solves the problems of low efficiency and poor scalability of existing single storage solutions. It includes three sub-modules: a relational database module, a caching module, and a distributed file storage module.
[0153] Relational database module
[0154] Implemented based on PostgreSQL, it is used to store structured data such as system configuration data, detection logs, handling logs, user information, and permission data. It supports efficient data querying and transaction management, master-slave replication, and ensures data reliability.
[0155] caching module
[0156] Implemented using Redis, it is used to store threat feature rules, machine learning model feature vectors, real-time traffic data, and high-frequency access alarm data, which greatly improves the access speed of high-frequency data in the system and reduces the pressure on database access.
[0157] Distributed file storage module
[0158] It is used to store massive amounts of unstructured data such as raw traffic packets, log files, risk reports, and attack source mapping files. It supports 3-replica backup of data and elastic expansion. Raw traffic packets are stored for 30 days by default, and important attack samples are automatically archived and permanently saved to ensure persistent storage and fast retrieval of massive amounts of data.
[0159] 6. Distributed deployment and high availability mechanisms
[0160] The solution in this embodiment adopts a distributed deployment architecture, supports elastic node expansion, and can adapt to network environments of different sizes. The cluster size can be elastically expanded to 100+ nodes, solving the problems of poor scalability and insufficient stability in high-concurrency scenarios of existing technologies.
[0161] As a preferred implementation method, the distributed node collaboration and high availability solution is as follows:
[0162] Load balancing: A consistent hashing algorithm is used to achieve load balancing between collection nodes and analysis nodes. By mapping physical nodes with 100 virtual nodes, the probability of hash collisions is reduced. Each collection node is responsible for processing traffic in a specific network segment, avoiding node overload caused by concentrated traffic.
[0163] Node status synchronization: ZooKeeper is used to synchronize the distributed lock with the node status in real time, monitor the running status of each node in real time, and automatically divert traffic to normal nodes when a node fails.
[0164] Master-slave hot standby: The core analysis node, strategy execution node, and database node are configured with dual-machine hot standby. When the master node fails, the slave node can complete automatic switching within 30 seconds. RTO ≤ 30 seconds. Master-slave data synchronization adopts asynchronous replication with a synchronization delay ≤ 50ms, ensuring that the core functions of the system are not interrupted. RTO stands for Recovery Time Objective, which refers to the longest time it takes for the system to recover to normal service from the occurrence of a failure.
[0165] Circuit breaking and traffic scheduling: A circuit breaking mechanism is designed to automatically trigger traffic diversion when a single node's CPU usage is ≥80%, memory usage is ≥90%, or network IO is ≥95%. Combined with a traffic priority scheduling strategy, threat detection traffic has higher priority than ordinary log collection traffic, ensuring the stable operation of core detection functions under high load.
[0166] As a further preferred implementation method, this solution establishes a unified intelligent analysis platform as the central hub for end-to-end data linkage. The platform achieves bidirectional data exchange with the data acquisition layer, intelligent analysis layer, strategy execution layer, visualization layer, and data storage layer through a standardized API interface (RESTful API v2.0, supporting JSON / Protobuf data formats). The specific linkage method is as follows:
[0167] (1) Data flow hub: The middle platform receives the standardized multi-source data output by the data acquisition layer and distributes it to each detection module of the intelligent analysis layer. At the same time, it synchronizes the attack judgment results of the intelligent analysis layer, the handling logs of the strategy execution layer, and the operation and maintenance annotation data of the visualization layer to the data storage layer, so as to realize the unified scheduling and closed-loop flow of data throughout the entire process and completely break the data silo problem of each module operating independently.
[0168] (2) Capability scheduling hub: The central platform unifies the management of the rule base update of the rule engine, the iterative release of machine learning models, and the full network synchronization of hierarchical disposal strategies, so as to realize the unified scheduling and collaborative operation of the capabilities of each module and avoid conflicts between multiple module strategies.
[0169] (3) Status monitoring hub: The platform monitors the operating status and resource usage of nodes at each level in real time, providing data support for load balancing, fault switching and traffic scheduling in distributed deployment, and ensuring high availability of the system.
[0170] Based on the above design, the complete operation flow of this embodiment forms a closed-loop protection system, and the specific steps are as follows:
[0171] Deployment initialization: Deploy distributed data collection nodes and log collection agents in the target network environment, complete the interface connection with the user's existing network devices and security devices, configure the system's basic parameters, rule base, and hierarchical handling strategies, and complete the pre-training and initial deployment of the machine learning model. The entire initialization process takes ≤30 minutes.
[0172] Multi-source data acquisition: The data acquisition layer collects multi-source data such as network traffic, device logs, and third-party device alarms in real time. After data cleaning and standardization, it is pushed to the intelligent analysis layer and simultaneously synchronized to the data storage layer for persistent storage. The end-to-end latency of data processing is ≤200ms in a 10Gbps traffic scenario.
[0173] Dual-engine fusion detection: The intelligent analysis layer extracts features from standardized data and uses a fusion detection mechanism combining a rule engine and machine learning to identify known attacks and unknown anomalies. After double verification, the attack behavior is marked with a danger level and the attack information is pushed to the policy execution layer.
[0174] Tiered automated handling: The policy execution layer matches the corresponding handling policy according to the attack risk level, executes millisecond-level automated handling actions, and, when necessary, links with existing security devices to achieve network-wide collaborative protection. At the same time, it records handling logs and sends them back to the data storage layer and intelligent analysis layer.
[0175] Visualization and Operation: The visualization layer displays the security status of the entire network in real time through the status dashboard, supports full-link attack tracing, risk trend analysis and multi-dimensional log query. Operation and maintenance personnel can review and annotate the detection results and handling actions. The annotated data is synchronously transmitted back to the intelligent analysis layer. The default refresh rate of the dashboard is 5 seconds / time, supports 10 people to operate online at the same time, and the latency of annotated data transmission is ≤100ms.
[0176] Model self-iterative optimization: The model self-learning module regularly collects closed-loop data throughout the entire process, completes sample screening and preprocessing, performs incremental training and optimization on the machine learning model, and completes gray-scale release after performance verification, continuously improving the system's threat detection capabilities. Incremental training is performed once a week by default, with a single training sample size of ≥100,000. After model iteration, the detection rate of unknown attacks increases by an average of 5%-8%.
[0177] Compared with the prior art, the main innovative designs and advantages of the above solutions in the embodiments of the present invention include:
[0178] 1. This innovative dual-engine threat identification mechanism, employing scenario-based dynamic weight fusion, achieves a breakthrough in both accurate identification of known attacks and effective detection of unknown attacks, fundamentally addressing the core pain points of traditional solutions such as lagging protection and high false positive rates. Based on a rule engine, this solution enables rapid and accurate matching of known attacks. It utilizes a machine learning model fused with Random Forest and XGBoost algorithms as its core, adapting to the extraction of traffic statistics and temporal features to achieve generalized detection of unknown attacks and attack variants. Differentiated allocation of dual-engine decision weights for known attacks and unknown anomaly scenarios, coupled with a dual verification mechanism and adaptive dynamic threshold adjustment strategy, ensures a stable detection accuracy of over 98% and a false positive rate below 1% (tested on the CSE-CIC-IDS2019 dataset, showing a 60% reduction in false positive rate compared to traditional single-engine solutions). Simultaneously, a closed-loop incremental self-learning system is constructed, combined with adversarial training to enhance model generalization capabilities, allowing the system's detection capabilities to continuously iterate and optimize with the accumulation of attack samples, avoiding the performance degradation of traditional single machine learning models over time.
[0179] 2. A distributed, multi-source, lightweight data acquisition solution was designed, significantly improving the comprehensiveness of data acquisition and its adaptability to high-bandwidth scenarios. This solution can extract 12 core traffic features covering basic network session attributes, statistical characteristics, and time-series characteristics. It simultaneously supports standardized acquisition of multi-source log data and third-party device data, solving the problem of single-dimensional acquisition in existing technologies. By optimizing packet capture algorithms and lightweight compression technology, the end-to-end latency from packet arrival at the network card to feature extraction and standardized feature vector output can be controlled within 100ms under 10Gbps high-speed traffic, with CPU resource usage not exceeding 30% (tested in an Intel Xeon Gold 6248 server environment). For 40Gbps / 100Gbps ultra-high-speed network scenarios, a DPU hardware acceleration, traffic sharding parallel processing, and adaptive sampling solution can be provided to further reduce host CPU utilization to below 15%, improve single-node processing capacity by more than 3 times, and fully adapt to business scenarios with different bandwidths.
[0180] 3. An automated, tiered response system based on attack threat levels has been constructed, achieving a balance between protection capabilities and business availability, and significantly improving the efficiency of security incident handling. This solution uses entropy value analysis to quantitatively classify attack behaviors into low / medium / high risk levels, with corresponding differentiated tiered handling strategies, avoiding normal business interruptions caused by indiscriminate blocking in existing technologies. Based on a multi-threaded concurrent processing mechanism, millisecond-level automated handling is achieved, with a total processing latency of no more than 500ms (an average processing latency of 380ms in a 1000 concurrent attack scenario). Attack closed-loop handling can be completed without manual intervention, significantly shortening the attack response window and preventing the lateral spread of attack damage. Simultaneously, a standardized SDK interface is provided, enabling seamless integration with existing firewalls, WAFs, and other security devices to achieve network-wide synchronization of handling strategies, solving the problem of incomplete handling by a single node.
[0181] 4. A five-layer integrated distributed closed-loop protection architecture is established to break down data silos and significantly improve the system's scalability and compatibility. This solution constructs an integrated architecture encompassing a data acquisition layer, intelligent analysis layer, policy execution layer, visualization layer, and data storage layer. Through a unified intelligent analysis platform, it achieves end-to-end data linkage, forming a full lifecycle protection closed loop of "collection-detection-handling-tracing-optimization," completely solving the problems of independent technology modules and fragmented data in existing systems. A hybrid storage architecture of PostgreSQL + Redis + distributed file storage is adopted, with differentiated storage solutions designed for different types of data, balancing data read / write efficiency with the reliability of massive data storage. The overall distributed deployment architecture achieves high availability through consistent hash load balancing, master-slave hot standby, and circuit breaker mechanisms. The cluster size can be elastically scaled to 100+ nodes (supporting horizontal scaling, with each additional 10 nodes increasing processing capacity by 20%), while remaining compatible with existing network security equipment without replacing existing assets, significantly reducing deployment and migration costs.
[0182] 5. Achieve intelligent attack tracing and network-wide security posture visualization, significantly reducing the threshold for network security operations and maintenance. This solution constructs a directed graph of attack propagation based on an improved PageRank algorithm, combines it with the Louvain community discovery algorithm to identify attack groups, and reconstructs the complete attack chain through time series correlation analysis. It supports real-time graph drawing of 100,000 nodes, with full-link tracing taking no more than 3 seconds (≤2.5 seconds for tracing in a 100,000-node graph). This helps operations and maintenance personnel quickly locate attack sources and attack paths. It also features a real-time security posture dashboard, risk trend analysis, and multi-dimensional log query functions, which can intuitively display the network-wide security status and provide early warning of potential security risks. Compared to existing solutions that only support basic log queries, this solution significantly simplifies operations and maintenance and reduces the workload of operations and maintenance personnel.
[0183] It should be noted that, unless otherwise defined, the technical or scientific terms used in this invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in this invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0184] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
[0185] This invention is not limited to the preferred embodiment described above. Anyone inspired by this invention can derive other forms of network threat detection methods and devices based on the fusion of rule engines and machine learning. All equivalent variations and modifications made within the scope of the claims of this invention shall fall within the scope of this invention.
Claims
1. A network threat detection method based on the fusion of rule engine and machine learning, characterized in that, Includes the following steps: Collect traffic data, log data, and security device alarm data from the target network environment, clean and standardize the collected data, and extract standardized feature data for detection. The standardized feature data is input into the rule engine and the machine learning detection model in parallel to obtain the rule matching result, the rule matching score, and the machine learning comprehensive anomaly probability, respectively. Based on whether the rule matching result matches the preset core threat rule, the attack scenario to which the current traffic belongs is divided. The attack scenario includes known attack scenarios and unknown abnormal scenarios. Based on the attack scenario to which the current traffic belongs, differentiated dynamic weights are assigned to the rule matching score and the comprehensive anomaly probability of machine learning, and the weighted fusion is used to obtain the total decision score; among them, the weight of the rule matching score is greater than the weight of the comprehensive anomaly probability of machine learning in the known attack scenario, and the weight of the comprehensive anomaly probability of machine learning is greater than the weight of the rule matching score in the unknown anomaly scenario. Initial attack determination is made based on the total decision score and dynamic decision threshold. For traffic that is initially determined to be an attack and belongs to an unknown abnormal scenario, dual verification is performed by matching the extended rule base and the historical behavior data of the corresponding session. The final attack determination is completed by combining the preliminary attack judgment and the results of dual verification. The traffic that is finally determined to be an attack is classified into threat levels, and automated handling actions are executed according to the preset graded handling strategy corresponding to the threat level.
2. The network threat detection method based on the fusion of rule engine and machine learning according to claim 1, characterized in that: The dynamic decision threshold adopts an adaptive adjustment mechanism: for traffic that is initially judged to be an attack and belongs to an unknown abnormal scenario, when the number of samples that pass dual verification reaches 5 consecutive times, the dynamic decision threshold is lowered by 10%; when the system's daily false alarm rate exceeds 5%, the dynamic decision threshold is raised by 15%.
3. The network threat detection method based on the fusion of rule engine and machine learning according to claim 1, characterized in that: The execution logic of the dual verification is as follows: for traffic initially determined to be an attack in unknown abnormal scenarios, the extended rule library covering non-core threat rules, abnormal behavior rules, and historical attack association rules is invoked for secondary matching. At the same time, the historical access behavior data of the corresponding session is retrieved for anomaly comparison. The final attack determination is confirmed based on the results of the secondary matching and comparison.
4. The network threat detection method based on the fusion of rule engine and machine learning according to claim 1, characterized in that: The machine learning detection model adopts a dual-algorithm fusion architecture. It extracts features from different dimensions of standardized feature data through two machine learning sub-models, outputs anomaly probabilities separately, and then fuses them with weights to obtain the comprehensive machine learning anomaly probability. The two machine learning sub-models include a random forest model and an XGBoost model. The random forest model is used to extract statistical features and output the first anomaly probability, while the XGBoost model is used to extract temporal features and output the second anomaly probability. The fusion weight is dynamically adjusted based on the accuracy of the validation set. When the accuracy of the random forest model is higher than that of the XGBoost model, it is assigned 60% weight, and otherwise, it is assigned 40% weight.
5. The network threat detection method based on the fusion of rule engine and machine learning according to claim 1, characterized in that: The threat level classification is based on the quantification of the threat level of the attack behavior, and different handling strategies are corresponding to different threat levels: low-risk attacks only trigger alarm notifications, medium-risk attacks trigger temporary blocking, and high-risk attacks trigger permanent bans and network-wide collaborative protection.
6. The network threat detection method based on the fusion of rule engine and machine learning according to claim 1, characterized in that: The standardized feature data covers the basic attribute features, traffic statistics features, and session temporal features of network sessions, providing a unified feature input for rule matching and machine learning detection.
7. The network threat detection method based on the fusion of rule engine and machine learning according to claim 1, characterized in that: It also includes a closed-loop self-optimization step for the machine learning model: collecting attack detection data, handling log data and manually reviewed and labeled data from the entire system process, selecting effective samples to incrementally train and optimize the machine learning detection model, enhancing the model's generalization ability by injecting adversarial perturbation samples generated based on the FGSM algorithm during the training process, completing the model's autonomous iterative update, setting the optimization cycle to once a week, and triggering emergency optimization when the model's accuracy drops by more than 3%.
8. The network threat detection method based on the fusion of rule engine and machine learning according to claim 1, characterized in that: It also includes an attack tracing step: constructing an attack propagation directed graph based on an improved graph computing algorithm. The improved graph computing algorithm introduces attack behavior time weights: the weight is 1.0 in the last 24 hours and decays to 0.5 from 24 to 72 hours. It also uses the PageRank algorithm to dynamically calculate the success probability weight based on historical attack success rates to identify attack-related nodes and groups, reconstruct the complete attack chain, and achieve rapid location of attack sources and attack paths.
9. A network threat detection device based on the fusion of rule engine and machine learning, characterized in that, include: The data acquisition module is used to collect traffic data, log data, and security device alarm data of the target network environment, clean and standardize the collected data, and extract standardized feature data for detection. The dual-engine detection module has a built-in rule engine unit and a machine learning detection unit. It is used to input the standardized feature data into the rule engine and the machine learning detection model in parallel to obtain the rule matching result, the rule matching score, and the machine learning comprehensive anomaly probability, respectively. The scenario-based fusion decision module is used to classify attack scenarios based on whether the rule matching result hits the preset core threat rules. Differentiated dynamic weights are assigned to the rule matching score and the comprehensive anomaly probability of machine learning according to the attack scenario, and the weighted fusion is used to obtain the total decision score. Among them, the weight of the rule matching score is greater than the weight of the comprehensive anomaly probability of machine learning in known attack scenarios, and the weight of the comprehensive anomaly probability of machine learning is greater than the weight of the rule matching score in unknown anomaly scenarios. The dual verification and attack determination module is used to make preliminary attack determination based on the total decision score and dynamic decision threshold, perform dual verification on traffic that is initially determined to be an attack in unknown abnormal scenarios, and complete the final attack determination by combining the preliminary determination and dual verification results. The tiered handling module is used to classify the threat level of traffic that is ultimately determined to be an attack and to execute automated handling actions according to the corresponding tiered handling strategy.
10. A network threat detection device based on the fusion of rule engine and machine learning according to claim 9, characterized in that, The device adopts a distributed deployment architecture and uses a consistent hash load balancing strategy to achieve multi-node traffic collaborative processing. The core node is configured with a primary and backup redundancy mechanism, and the primary and backup switching time does not exceed 500ms, ensuring the stable operation of the system under high load scenarios.