Distributed threat intelligent honeynet trapping method
By building a distributed threat intelligent honeynet, using the kill chain model and ATT&CK framework to accurately match attack tactics, and dynamically assemble the bait environment, the existing honeynet trapping method solves the problem of insufficient detection capabilities in the power system, and achieves efficient threat capture and identification.
Patent Information
- Application Number
- CN202510449593.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-05
AI Technical Summary
The existing honeynet trapping methods do not have multiple detection capabilities covering the network side to the application side in the power system, and cannot be deeply confused and disguised, resulting in a low success rate of trapping and unable to effectively deal with complex power business scenarios.
The distributed threat intelligent honeynet trapping method is adopted, and by constructing traps to obfuscate attack targets, combining attack counter and traceability technology, using the kill chain model and the ATT&CK framework to accurately match attack tactics, dynamically assemble the bait environment and monitoring modules, and construct a modular sandbox to improve deception capabilities.
It realizes efficient capture and identification of large-scale distributed threats to the power system, improves the success rate of trapping, enhances the ability to counter attackers, and adapts to a variety of power business scenarios.
Smart Images

Figure CN120433958A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security defense technology, and in particular to a distributed threat intelligent honeynet trapping method. Background Art
[0002] The Chinese patent "CN114157467A Distributed Switchable Industrial Control Honeynet Trapping Method" is a distributed switchable industrial control honeynet trapping method that captures intruders' attack behaviors on industrial systems and traces the intruder's identity through behavioral feature library comparison and customized Trojan horse technology. It includes the following steps: (1) The third-generation network data acquisition engine (Colasoft PacketCapture Engine, CSPCE) and the underlying driver (supporting both Windows and Linux platforms) collect data; (2) The data packets and data packet timestamps in the PCAP file are read and imported; (3) The storage engine (Colasoft Storage, CSStorage) is used to store the data; (4) The feature recognition technology and behavioral pattern recognition technology are used to perform retrospective analysis on the data.
[0003] Chinese patent application "CN103561003A: A Honeynet-Based Collaborative Active Defense Method" discloses a honeynet-based collaborative active defense method, comprising the following steps: data preprocessing, correlation analysis, threat analysis, vulnerability analysis, and blacklist generation. This method uses attacker information from distributed honeynets across different subnets as a data source. Employing a collaborative defense approach, this method analyzes the correlation, threat, and vulnerability between different subnets and attackers. Ultimately, based on a fusion analysis, it predicts a personalized list of attackers most likely to attack each subnet, i.e., a highly predictive blacklist. This method not only boasts a high defense rate, hit rate, real-time performance, and predictability, but also achieves very low missed and false positive rates. Furthermore, because attacker information is captured through the honeynet, there is no need for ordinary users to report malicious attackers, thus preventing disruption to their normal communications and infringing on their privacy. This method boasts a high defense rate, hit rate, collaborative defense rate, and a low collaborative failure rate. It provides predictive defense and can also assist in the defense of other subnets within the same enterprise network. This method also fully considers the differences and similarities between different subnets, developing a unique proactive defense strategy for each subnet. This reduces the length of blacklists for vulnerable subnets, reduces firewall load, and improves network transmission efficiency. Furthermore, this method utilizes a distributed, highly interactive honeynet, which collects data with high reliability and controllability, is low-cost, and requires no user reporting.
[0004] The Chinese patent "CN115150124A Deception Defense System" discloses a deception defense system, which includes a simulation subsystem; the simulation subsystem includes: a simulation capability module for generating a honeypot; a danger perception module for monitoring the network after the honeypot is buried and obtaining monitoring data; an intelligent analysis engine module for comprehensively collecting evidence of the attack process; a tracing and countermeasure module for obtaining attack data of the attack; the simulation capability module includes a customized simulation unit, which is used to generate a honeypot customized according to the actual business of the target customer. The present invention customizes the honeypot according to the actual business of the customer to meet the customer's personalized needs; supports probe and direct connection deployment modes, selects the deployment method according to the customer's network status, designs different honeynets, and improves the matching degree between the honeynet and the customer's network; the present invention supports single-point and distributed deployment modes, and can be deployed on the external network, network boundary or intranet, improving the environmental adaptability of the deception defense system. This patent generates honeypots customized according to the customer's actual business through customized simulation units, meeting the customer's personalized needs; supports probe and direct connection deployment modes, and can choose the deployment method according to the customer's network status and design different honeynets, thereby improving the matching degree between the honeynet and the customer's network and improving the practicality of the deception defense system; supports single-point and distributed deployment modes, and can be deployed on the external network, network boundary or intranet, improving the environmental adaptability of the deception defense system.
[0005] In actual use, the honeynet trapping method in the existing technology does not have the polymorphic construction technology of detection modules that covers multiple detection capabilities from the network side to the application side. It does not support the deployment of complex trapping capabilities for various typical power business scenarios. It cannot be deeply confused with the real power business system and network environment, and cannot be highly disguised as a large-scale smart honeynet with a low trapping success rate. To address the above problems, a distributed threat smart honeynet trapping method is provided. Summary of the Invention
[0006] The purpose of the present invention is to provide a distributed threat intelligent honeynet trapping method to solve the problems raised in the above background technology. To achieve the above purpose, the present invention provides the following technical solution: a distributed threat intelligent honeynet trapping method, comprising the following steps:
[0007] S1. Based on attack deception defense technology, traps are constructed along hacker routes to confuse their targets, accurately perceive hacker attack behaviors, and guide and isolate attack traffic into the intelligent honeynet.
[0008] S2. Combine attack countermeasures and attack tracing to accurately obtain the hacker's network identity and fingerprint information;
[0009] S3. From the attacker's perspective, design various attack scenarios, precisely set traps and defenses at each stage based on the kill chain model, accurately match the attacker's attack tactics with the ATT&CK framework, and perceive known and unknown attack methods based on behavioral analysis.
[0010] S4. Support the verification of various relevant key technologies and realize the capture function of large-scale distributed threats to the power system.
[0011] Preferably, the kill chain model operation process further includes the following steps:
[0012] S1. Obtain attack events and record in detail the threat actor, attack source, attack target, attack start time, attack end time, and event risk level. Record the attacker's detailed operation steps, display SSH and RDP operation records in video format, and download and analyze the files uploaded by the attacker.
[0013] S2. Analyze attack behaviors and display detailed attack behavior records of all monitored sandboxes in chronological order. Locate attackers more quickly based on information about attackers, attack sources, attack targets, isolated sandboxes, attack time, attack types, and attacker methods. Add identification and display of scanning tools and process escapes to attack behaviors. Record disguised proxies, half-open connections of honeypot extension nodes, and UDP scanning events. Filter attack steps using the ATT&CK framework.
[0014] Preferably, the process steps of the ATT&CK framework further include:
[0015] S1: The client executes the attachment sent in the phishing email, achieving persistence and privilege escalation through the phishing email, stealing the target system's account and password, and then roaming the entire network through network environment discovery and lateral movement, ultimately stealing the target's network data.
[0016] S2. Penetration testing uses high-frequency techniques in ATT&CK to conduct security testing on the demand side, develop security solutions, and strengthen enterprise security. In the red-blue confrontation, both the blue and red teams can use the ATT&CK framework. The red team can perform phishing and watering hole attacks, while the blue team can use the techniques in the framework to audit and protect the system website in advance.
[0017] S3: Create a hacker profile based on the attack behavior and view the attacker's detailed identity information, including the attack source IP, the attacker's real IP, internal network IP, public network IP, network ID, device fingerprint, operating system, and browser used. The Threat Intelligence Center links the corresponding content and displays the corresponding information in the intelligence linkage data module for real people who match the same device fingerprint;
[0018] S4. Utilize hacker profiling to trace the source of attacks. This technology uses a variety of techniques to capture the attacker's virtual identity, accurately locates multi-dimensional aggregation, restores the attack process, traces the attacker's identity, forms a hacker profiling, and finds the threat at its root. ZombieCookie utilizes a specific algorithm to calculate the device fingerprint information of a specific attacker and uses WebRTC technology to obtain the attacker's detailed IP information.
[0019] S5. Targeted attacks are targeted and countered. We leverage the ability to reversely control attackers for multi-platform countermeasures, covering common operating systems such as Windows, macOS, and Android. We capture the attacker's real IP address, device fingerprint information, and social ID information, reversely control the attacker's host, and reversely target common hacker attack plug-ins. This includes browser countermeasures, scanner countermeasures, and git countermeasures. Once an attacker accesses a web honeypot, we exploit browser vulnerabilities to extract sensitive information from the attacker's host and reversely control the attacker's host.
[0020] S6. Use time series graphics to display the attack path and attack chain of the captured attack behavior. Using Kill Chain as a reference model and the ATT&CK framework as a technical supplement, we correspond to each step taken by the attacker and utilize sandbox business simulation, vulnerability camouflage, virtual systems, desensitized false data, and behavior records to display the attack path and attack chain of the captured attack behavior in time series graphics.
[0021] S7. Build a modular sandbox in the form of building blocks. Use services as the basic resources for building sandboxes. Sandboxes can arbitrarily select one or more services for combination. When building sandboxes, existing templates, custom web classes, and system service sandboxes can be arbitrarily combined. Use different services, different pages, and multiple service combinations in different scenarios to enhance the deception capability of the smart honeynet.
[0022] Preferably, the algorithm flow of the kill chain model includes:
[0023] S1. Count the types of attacks contained in the original data;
[0024] S2. Classify or define each attack type into a certain stage of the kill chain model based on its characteristics. Next, preprocess the data and construct a feature vector.
[0025] S3. Eliminate all redundant fields and define the distance calculation method for each dimension of features;
[0026] S4. Add the kill chain model stage attributes of each attack to the feature vector as features;
[0027] S5. Using a filtering unsupervised feature selection method to screen feature vectors and construct feature vectors that best reflect the spatial structure correlation of the data set can effectively reduce the resources consumed by system operation;
[0028] S6. By using the new feature vector and field similarity measurement, an improved spectral clustering algorithm is used to mine the existing network kill chain model attacks from the data;
[0029] S7. Three different kill chain models are proposed, and the Markov model is used for derivation and analysis to realize the attack prediction function of the network kill chain model.
[0030] Preferably, each field of the feature vector reflects different attributes of the kill chain model attack process, a i and a j Represent two pieces of data respectively and j>i, details are as follows:
[0031] Time reflects the relationship between attack events. As time goes by, the correlation between two attack events gradually weakens. The similarity measurement is defined as:
[0032]
[0033] The attacker's IP address originates from the same network segment. The same attacker has similarity in the source IP or destination IP in the IDS alarm log. The similarity metric is defined as:
[0034]
[0035] Where M = max{H(a i .sIP,a j .sIP),H(a i .sIP,a j .dIP),H(a i .dIP,a j .sIP),H(a i .dIP,a j .dIP)}, the H function is the binary of two IP addresses, indicating the same number of bits from left to right, sIP refers to the source IP address, and dIP refers to the destination IP address;
[0036] If the attacker uses different tools to successfully intrude into the server, and the two data have the same HTTP request method, server port number, client port number, client environment, and HTTP response code, the similarity metric of these fields can be defined as:
[0037]
[0038] The same attacker will be in the same area on the map. Based on the size of the regional communication in the two data, the similarity measure can be defined as:
[0039] F Locate =(a i .country&a j .country)*0.1+(a i .province&a j .province)*0.2+(a i .city&a j ,city)*0.7;
[0040] Judging from the characteristics of the kill chain, the result of the attack method in the previous alarm log may be the prerequisite for launching the attack method in the startling alarm log. The similarity measurement can be defined as:
[0041]
[0042] Preferably, the feature scoring function comprises a feature selection algorithm, wherein the importance of each feature of the feature selection algorithm is evaluated by minimizing the reconstruction coefficient between the weight matrix of all features and a single feature, and the feature selection algorithm comprises:
[0043] Given a dataset matrix X = {1, ...,} ∈ g, search for a feature subset of size m that contains the most information. The structure represented by the data points in the m-dimensional space well preserves the intrinsic structure of the dataset in the original d-dimensional space.
[0044] The local geometric structure of the data set is modeled to construct a simple and effective neighbor graph A, which is generated by all candidate features within a specific neighborhood. No additional parameters are required, and the weight matrix can be calculated according to the following formula:
[0045]
[0046] Among them, A ij is the number of neighborhood samples of sample i.
[0047] Preferably, the overall system architecture of the network kill chain model detection platform is divided into a presentation layer, a business logic layer and a data storage layer from top to bottom.
[0048] Preferably, the presentation layer is mainly composed of the system front-end pages, including the situation awareness homepage, the data display page, and the search page, wherein the data display page display effect is displayed in the form of data visualization, including statistical charts, and all of the presentation layers are designed and implemented in the form of web pages.
[0049] Preferably, the business logic layer includes an algorithm module and a Web background module. The algorithm module is responsible for implementing the algorithm flow of the kill chain model detection and prediction model of the network, including raw data preprocessing, attack event classification, unsupervised feature selection, kill chain model detection, and kill chain model prediction sub-modules. The Web background module performs data interaction, and the data processing part performs type conversion and normalization on the raw data, and uses the model call to pass the running results of each algorithm module to the Web background module. The Web background module stores the data in the database, and the Web background module performs data transmission work that interacts with the front end. The data storage layer is composed of a data storage module, and the data storage module completes the data reading and writing functions, and directly interacts with the Web background module upward.
[0050] Preferably, the business logic layer includes a data preprocessing module, a feature selection module, a killchain model detection module, and a network killchain model prediction module.
[0051] Compared with the prior art, the present invention has the following beneficial effects:
[0052] Through dynamic assembly of decoy environments and monitoring modules, and polymorphic module reconstruction, the fundamental elements of the intelligent honeynet (sandboxes, decoys, etc.) are service-oriented, ensuring maximum flexibility through modular design. When deploying deception resources, different sandbox types can be randomly arranged based on existing templates or customized methods, invoking a variety of camouflage and obfuscation capabilities suitable for typical power application scenarios, effectively countering existing attacker anti-entrapment techniques and tactics. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 This is a diagram of the intelligent honeynet trapping technology of the present invention;
[0054] Figure 2 This is the overall algorithm flow chart of the kill chain detection and prediction model of the present invention;
[0055] Figure 3 This is the overall architecture diagram of the kill chain detection platform of the present invention. DETAILED DESCRIPTION
[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0057] See also Figures 1 to 3 The present invention provides a technical solution: a distributed threat intelligent honeynet trapping method, comprising the following steps:
[0058] S1. Based on attack deception defense technology, traps are constructed along hacker routes to confuse their targets, accurately perceive hacker attack behaviors, and guide and isolate attack traffic into the intelligent honeynet, thereby protecting real assets in power business scenarios and recording and analyzing attack behaviors.
[0059] S2. Combine attack countermeasures and attack tracing to accurately obtain the hacker's network identity and fingerprint information, so as to collect attack evidence and trace the source;
[0060] S3. From the attacker's perspective, design various attack scenarios, precisely set traps and defenses at each stage based on the kill chain model, accurately match the attacker's attack tactics with the ATT&CK framework, and perceive known and unknown attack methods based on behavioral analysis.
[0061] S4. Support the verification of various relevant key technologies and realize the capture function of large-scale distributed threats to the power system.
[0062] In this embodiment, the kill chain model operation process also includes the following steps:
[0063] S1. Obtain attack events and record in detail the threat actor, attack source, attack target, attack start time, attack end time, and event risk level. Record the attacker's detailed operation steps, display SSH and RDP operation records in video format, and download and analyze the files uploaded by the attacker.
[0064] S2. Analyze attack behavior, displaying detailed attack behavior records for all monitored sandboxes in chronological order. Attackers can be quickly located based on the attacker, attack source, attack target, isolated sandbox, attack time, attack type, and attacker tactics. Scanning tool identification and process escape are also displayed at the attack behavior level. Disguised proxies, half-open connections of honeypot extension nodes, and UDP scanning events are recorded, and attack steps within the ATT&CK framework are filtered. The "killchain" (KILL CHAIN) model is a threat intelligence-driven defense model used to guide and identify all activities an attacker must complete to intrude on a network. The "killchain" model includes seven steps: reconnaissance, weaponization, payload delivery, exploitation, installation, command and control, and action.
[0065] In this embodiment, the ATT&CK framework process steps also include:
[0066] S1: The client executes the attachment sent in the phishing email, achieving persistence and privilege escalation through the phishing email, stealing the target system's account and password, and then roaming the entire network through network environment discovery and lateral movement, ultimately stealing the target's network data.
[0067] S2. Penetration testing uses high-frequency techniques in ATT&CK to conduct security testing on the demand side, develop security solutions, and strengthen enterprise security. In the red-blue confrontation, both the blue and red teams can use the ATT&CK framework. The red team can perform phishing and watering hole attacks, while the blue team can use the techniques in the framework to audit and protect the system website in advance.
[0068] S3: Create a hacker profile based on the attack behavior and view the attacker's detailed identity information, including the attack source IP, the attacker's real IP, internal network IP, public network IP, network ID, device fingerprint, operating system, and browser used. The Threat Intelligence Center links the corresponding content and displays the corresponding information in the intelligence linkage data module for real people who match the same device fingerprint;
[0069] S4. Utilize hacker profiling to trace the source of attacks. This technology uses a variety of techniques to capture the attacker's virtual identity, accurately locates multi-dimensional aggregation, restores the attack process, traces the attacker's identity, forms a hacker profiling, and finds the threat at its root. ZombieCookie utilizes a specific algorithm to calculate the device fingerprint information of a specific attacker and uses WebRTC technology to obtain the attacker's detailed IP information.
[0070] S5. Targeted attacks are targeted and countered. We leverage the ability to reversely control attackers for multi-platform countermeasures, covering common operating systems such as Windows, macOS, and Android. We capture the attacker's real IP address, device fingerprint information, and social ID information, reversely control the attacker's host, and reversely target common hacker attack plug-ins. This includes browser countermeasures, scanner countermeasures, and git countermeasures. Once an attacker accesses a web honeypot, we exploit browser vulnerabilities to extract sensitive information from the attacker's host and reversely control the attacker's host.
[0071] S6. Use time series graphics to display the attack path and attack chain of the captured attack behavior. Using Kill Chain as a reference model and the ATT&CK framework as a technical supplement, we correspond to each step taken by the attacker and utilize sandbox business simulation, vulnerability camouflage, virtual systems, desensitized false data, and behavior records to display the attack path and attack chain of the captured attack behavior in time series graphics.
[0072] S7. Build a modular sandbox with building blocks. Use services as the basic resources for building sandboxes. Sandboxes can arbitrarily select one or more services to combine. When building sandboxes, existing templates, custom web classes, and system service sandboxes can be arbitrarily combined. The deception capability of the smart honeynet can be enhanced by using a variety of service combinations in different services, different pages, and different scenarios.
[0073] The cybersecurity risk management framework ATT&CK (Adversarial Tactics, Techniques, and Common Knowledge) describes the attack tactics and techniques used at each stage of an attack from the attacker's perspective. ATT&CK (Adversarial Tactics, Technigues and Common Knowledge) is a model that describes the techniques used at each stage of an attack from the attacker's perspective. Common application scenarios include network red-blue confrontation simulations, network security penetration testing, network defense gap assessments, and network threat intelligence collection. Existing deception defense technologies have the problem of single functions and simple configurations in the deployment of deception resources (such as sandboxes and decoys), which are easily identified and bypassed by attackers' constantly upgraded counter-reconnaissance methods. From the perspective of improving the camouflage, obfuscation, coverage, and sweetness of deception resources, this paper studies a set of dynamic assembly methods for decoy environments and monitoring modules, as well as polymorphic module reconstruction methods. The modular assembly of the decoy environment is achieved by servitizing the fundamental elements of the smart honeynet (sandboxes, decoys, etc.). By combining different services, pages, and multiple services in typical power business scenarios, attackers' attack techniques and tactics can be effectively countered. For example, common internet-based attacks often involve VPNs. This modular sandbox management approach automatically generates a VPN sandbox combined with a fake business system sandbox, luring attackers through the VPN sandbox and directing attack traffic into the honeynet to capture and isolate the attack. By combining and correlating detection modules from the traffic side to the application side, attack decoy capabilities covering layers 4 to 7 are built. By setting fake IP addresses and ports, or URLs of interest to attackers on real business systems, attackers are lured into scanning and brute-forcing attacks. Preemptive response techniques are then used to establish a connection with the attacker, capturing them, or directing attack traffic into the smart honeynet to isolate and analyze the attack.
[0074] In this embodiment, the kill chain model algorithm process includes:
[0075] S1. Count the types of attacks contained in the original data;
[0076] S2. Classify or define each attack type into a certain stage of the kill chain model based on its characteristics. Next, preprocess the data and construct a feature vector.
[0077] S3. Eliminate all redundant fields and define the distance calculation method for each dimension of features;
[0078] S4. Add the kill chain model stage attributes of each attack to the feature vector as features;
[0079] S5. Using a filtering unsupervised feature selection method to screen feature vectors and construct feature vectors that best reflect the spatial structure correlation of the data set can effectively reduce the resources consumed by system operation;
[0080] S6. By using the new feature vector and field similarity measurement, an improved spectral clustering algorithm is used to mine the existing network kill chain model attacks from the data;
[0081] S7. Propose three different kill chain models and use the Markov model for derivation and analysis to realize the attack prediction function of the network kill chain model;
[0082] The data used is derived from a company's IDS alarm log. Each object in the dataset represents an HTTP packet flow between a client and a server. Because existing security monitoring software records network attack logs in different formats and contents, the feature vectors constructed using this invention's raw data are somewhat unique. However, according to the field descriptions of the feature vectors in the Moore dataset
[57] and the KDD99 dataset, compared to the most common IP quintuple, the feature vectors constructed in this invention can fully display detailed information about network attacks. The fields included are as described in the table:
[0083] Field illustrate LocalDateTime Date and time RequestMethod HTTP request method IP Client and server IP addresses SeverPort Server port number ClientPort Client port number ClientEnv Client environment HTTPCode HTTP response codes Locate Geographical location Event The stage of the kill chain where the attack occurred
[0084] Attack events are a key field in IDS alert logs. They use a rule-based approach to identify the attack event type for each data entry. The network kill chain is a multi-stage conceptual model, with each stage being a general term for a specific type of attack event. There are 21 attack types captured from IDS alert logs. The table below describes the kill chain stages for each attack, based on their characteristics:
[0085]
[0086] For each field in the feature vector, a sample distance, also known as a similarity metric, must be calculated. Therefore, it is necessary to normalize some non-numeric fields. The following provides similarity metric definitions for all fields relevant to killchain analysis. This paper subsequently employs feature selection to identify the m-dimensional features that best preserve the spatial structure of the original data. This approach also provides a valuable reference for constructing feature vectors and similarity metrics for different datasets.
[0087] In this embodiment, each field of the feature vector reflects different attributes of the kill chain model attack process. i and a jRepresent two pieces of data respectively and j>i, details are as follows:
[0088] Time reflects the relationship between attack events. As time goes by, the correlation between two attack events gradually weakens. The similarity measurement is defined as:
[0089]
[0090] The attacker's IP address originates from the same network segment. The same attacker has similarity in the source IP or destination IP in the IDS alarm log. The similarity metric is defined as:
[0091]
[0092] Where M = max{H(a i .sIP,a j .sIP),H(a i .sIP,a j .dIP),H(a i .dIP,a j .sIP),H(a i .dIP,a j .dIP)}, the H function is the binary of two IP addresses, indicating the same number of bits from left to right, sIP refers to the source IP address, and dIP refers to the destination IP address;
[0093] If the attacker uses different tools to successfully intrude into the server, and the two data have the same HTTP request method, server port number, client port number, client environment, and HTTP response code, the similarity metric of these fields can be defined as:
[0094]
[0095] The same attacker will be in the same area on the map. Based on the size of the regional communication in the two data, the similarity measure can be defined as:
[0096] F Locate =(a i .country&a j .country)*0.1+(a i .province&a j .province)*0.2+(a i .city&a j ,city)*0.7
[0097] Judging from the characteristics of the kill chain, the result of the attack method in the previous alarm log may be the prerequisite for launching the attack method in the startling alarm log. The similarity measurement can be defined as:
[0098]
[0099] Although the dimension of the feature vector constructed for the IDS alarm log in the present invention is not high, it is still necessary to perform feature selection for three reasons: First, the feature vector constructed by the present invention is derived from a specific IDS alarm log, and different alarm logs have inconsistent data collection standards. It can be seen intuitively that the IDS alarm log relied on in the present invention is collected at the application layer of the network protocol, namely the HTTP protocol, but some IDS alarm logs collect data at the TCP protocol layer and contain certain statistical information, which shows that the content covered by the data set may not be the same. Secondly, the construction of the feature vector depends on the subjective experience of the researchers, and in this case it is very likely to contain many irrelevant features. Finally, although the dimension of the feature vector constructed in the present invention is not high, for the engineering implementation of the kill chain detection platform, reducing the dimension of the feature vector can not only reflect the comprehensiveness of the kill chain detection problem analysis, but also help to reduce the computational complexity and the timeliness of the system response. Therefore, it is necessary to adopt a suitable unsupervised feature selection method to screen and simplify the feature vector.
[0100] In this embodiment, the feature scoring function includes a feature selection algorithm. The importance of each feature of the feature selection algorithm is evaluated by minimizing the reconstruction coefficient between the weight matrix of all features and a single feature. The feature selection algorithm includes:
[0101] Given a dataset matrix X = {1, ...,} ∈ g, search for a feature subset of size m that contains the most information. The structure represented by the data points in the m-dimensional space well preserves the intrinsic structure of the dataset in the original d-dimensional space.
[0102] The local geometric structure of the data set is modeled to construct a simple and effective neighbor graph A, which is generated by all candidate features within a specific neighborhood. No additional parameters are required, and the weight matrix can be calculated according to the following formula:
[0103]
[0104] Among them, A ij is the number of neighborhood samples of sample i;
[0105] Given the local structure of the data represented by the neighborhood graph, most existing filter-based feature selection algorithms use a quadratic function to evaluate the importance of each feature, such as formulas (2-2) and (2-3). However, this scoring function has at least three disadvantages [2: 1) it cannot handle the case where all sample attribute values are equal; 2) it lacks scale invariance; 3) the neighborhood graph is insensitive to changes in the value of each attribute. These disadvantages will greatly reduce its performance in feature selection;
[0106] The above algorithm has unique advantages in unsupervised feature selection:
[0107] (1) No parameters. It can be seen that this method only involves one parameter, namely the neighborhood size, but relatively speaking, Laplace score, LGD score, LLE score and MCFS (Multi Cluster Feature Selection) also need to set the neighborhood size. In addition, these methods often require additional hyperparameters, such as the kernel width of the Gaussian kernel function in Laplace score, regularization parameters in MCFS and LLE score, etc. Since IDS alarm log data does not contain label information, the choice of parameters often does not conform to the actual situation. This is also one of the main advantages of the present invention's use of a parameter-free unsupervised feature selection algorithm.
[0108] (2) This method evaluates the importance of each feature by considering both feature relevance and redundancy. As a result, the redundancy of the selected feature subset is greatly reduced, and features that reflect the overall dataset structure and more relevant information are given greater weight.
[0109] (3) The global optimal solution of the optimization problem in the formula is easy to obtain, so greedy search or local optimality can be avoided.
[0110] (4) It is scale invariant. If the graph structure is the same when y = 26, then the scores of y and y are equal. Laplace scoring has different scores for such features (B01), and LGD scoring also has this problem (B2). (5) It can distinguish features with the same score value. There is no meaningful correlation between the local structures of equal-valued features, which only makes the reconstruction weight smaller. However, both Laplace scoring and LGD scoring give better scores to such features.
[0111] Based on the aforementioned algorithm, the killchain detection platform is deployed on hosts or clusters requiring killchain detection. Based on Spark and a web framework, the system provides log data analysis, historical data statistics, and detailed data retrieval. The homepage provides an overall situational awareness page, with visualizations displaying multi-level statistical information on recent cyberattacks. The system's functionality is described below:
[0112] (1) Interactive Functionality: The killchain detection platform needs to design a reasonable interactive interface to allow users to use the entire system more conveniently. Interaction is the most direct function for users to communicate with the system. The interactive interface needs to highlight important information to the user. At the same time, the user's operation should be concise, clear, and logical. If there is critical information, a warning should be issued in the form of a pop-up window.
[0113] (2) Data processing function: The killchain detection platform needs to implement two data processing functions. The first is that non-technical users may make some mistakes. For example, when performing keyword searches, redundant characters or incorrect string formats appear in the search strings passed by the front-end page. These situations need to be handled separately to ensure a good user experience and the continuation of subsequent operations. The second is that the implementation of the algorithm has specific format requirements for the original data. Most of the algorithms and processes involved in the system include complex data type conversions and operations based on mathematical statistics. Only by processing each piece of data completely and correctly can the system finally reflect the actual effect of the theoretical algorithm.
[0114] (3) Algorithm Implementation: The killchain detection and prediction model is the core of the entire system. The system sequentially implements processes such as data preprocessing, unsupervised feature selection algorithm, improved spectral clustering algorithm, and Markov model analysis. The IDS alarm log serves as the system's data input. All algorithm implementations in this invention do not involve the process of persisting the model locally. At specified intervals, when a new alarm log is generated in the IDS, the algorithm module performs killchain detection and prediction on the log data and stores the final abnormal value in the database.
[0115] (4) Data storage: The implementation and operation of the killchain detection platform are data-driven, and data storage is crucial. Data is the raw material of the entire system. This requires designing a reasonable database table structure to store data. To maintain high query efficiency, optimization operations such as indexing can also be implemented. Disaster recovery must also be considered when storing data. Regularly backing up system data ensures that data is not lost in the event of a problem.
[0116] In this embodiment, the overall system architecture of the network kill chain model detection platform is divided into a presentation layer, a business logic layer, and a data storage layer from top to bottom.
[0117] In this embodiment, the presentation layer is mainly composed of the system front-end pages, including the situation awareness homepage, the data display page, and the search page. The data display page displays the effect in the form of data visualization, including statistical charts. All presentation layers are designed and implemented in the form of web pages.
[0118] In this embodiment, the business logic layer includes an algorithm module and a Web background module. The algorithm module is responsible for implementing the algorithm flow of the network kill chain model detection and prediction model, including raw data preprocessing, attack event classification, unsupervised feature selection, kill chain model detection, and kill chain model prediction sub-modules. The Web background module performs data interaction. The data processing part performs type conversion and normalization on the raw data, and uses the model call to pass the running results of each algorithm module to the Web background module. The Web background module stores the data in the database. The Web background module performs data transmission work for interaction with the front end. The data storage layer is composed of the data storage module. The data storage module completes the data reading and writing functions and directly interacts with the Web background module upward.
[0119] In this embodiment, the business logic layer includes a data preprocessing module, a feature selection module, a killchain model detection module, and a network killchain model prediction module. The data preprocessing module primarily reads raw data from IDS alarm logs and performs similarity calculations on each field in the data. Since IDS alarm logs are stored in file format, the first step in data preprocessing is to read the file. The read data undergoes data fusion to remove incomplete or continuous residual data. At the same time, attack events are classified according to the characteristics of the network kill chain. The resulting source data is stored in the database, avoiding the time-consuming process of reading external files each time. For ease of distinction, we will refer to this raw log text data as raw data, and the data after feature selection as source data. For each sample data item in the source data, a similarity metric is calculated between that data item and each other data item. Finally, for fields with large or small distance values, the values of these fields are normalized, and the final results are stored in the database. Since similarity calculations between data items are already completed in the data preprocessing module, the feature selection algorithm module primarily implements an unsupervised feature selection algorithm, and the detailed algorithm is not detailed here. Since the implementation of the algorithm requires the application of a large number of mathematical formulas, this module uses the MLlib library and Breeze library under the Spark platform to encapsulate and call certain operations.
[0120] The class diagram for the feature selection module's functionality is shown below. The classes and their methods are detailed below: The ReadData and SaveData classes in the Dao layer implement database-related operations; the DataObject class represents each data object; and the Algorithm1 class encapsulates methods related to the feature selection algorithm. The kill chain detection module clusters the source data generated by the data preprocessing module to form a set of attack sequences. Each attack sequence is a list of alert IDs associated with a particular attack. Similar to the implementation of the feature selection module, the clustering algorithm relies on various mathematical libraries. The kill chain detection module detects a set of kill chain sequences from the source data. Each kill chain contains multiple attacks.
[0121] The class diagram for the network kill chain detection module shows the functional implementation. The classes and their methods are detailed below: The ReadData and SaveData classes in the Dao layer implement database-related operations; the DataObject class represents each data object; the Util class integrates utility methods required for calculations; and the Values class primarily stores parameter values and intermediate values during the algorithm's execution. The Algorithm2 object, written in Scala, encapsulates the methods and processes of the improved spectral clustering algorithm. The network kill chain prediction module mines multiple kill chains from the source data. This module primarily calculates the probability of an attack occurring at each stage of the existing kill chain data and, combined with Markov model analysis, predicts future kill chain attacks. Accordingly, the kill chain prediction algorithm relies on a mathematical library. Although the kill chain prediction module's output only contains a few predicted values, this is done to ensure system standardization.
[0122] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are merely preferred examples of the present invention and are not intended to limit the present invention. Various changes and modifications may be made to the present invention without departing from the spirit and scope of the present invention, and such changes and modifications fall within the scope of the present invention. The scope of protection claimed in the present invention is defined by the appended claims and their equivalents.
Claims
1. A distributed threat intelligent honeynet trapping method, characterized in that: The steps include: S1. Based on attack deception defense technology, traps are constructed along hacker routes to confuse their targets, accurately perceive hacker attack behaviors, and guide and isolate attack traffic into the intelligent honeynet. S2. Combine attack countermeasures and attack tracing to accurately obtain the hacker's network identity and fingerprint information; S3. From the attacker's perspective, design various attack scenarios, precisely set traps and defenses at each stage based on the kill chain model, accurately match the attacker's attack tactics with the ATT&CK framework, and perceive known and unknown attack methods based on behavioral analysis. S4. Support the verification of various relevant key technologies and realize the capture function of large-scale distributed threats to the power system.
2. The distributed threat intelligent honeynet trapping method according to claim 1, characterized in that: The killchain model operation process also includes the following steps: S1. Obtain attack events and record in detail the threat actor, attack source, attack target, attack start time, attack end time, and event risk level. Record the attacker's detailed operation steps, display SSH and RDP operation records in video format, and download and analyze the files uploaded by the attacker. S2. Analyze attack behaviors and display detailed attack behavior records of all monitored sandboxes in chronological order. Locate attackers more quickly based on information about attackers, attack sources, attack targets, isolated sandboxes, attack time, attack types, and attacker methods. Add identification and display of scanning tools and process escapes to attack behaviors. Record disguised proxies, half-open connections of honeypot extension nodes, and UDP scanning events. Filter attack steps using the ATT&CK framework.
3. The distributed threat intelligent honeynet trapping method according to claim 1, characterized in that: The ATT&CK framework process steps also include: S1: The client executes the attachment sent in the phishing email, achieving persistence and privilege escalation through the phishing email, stealing the target system's account and password, and then roaming the entire network through network environment discovery and lateral movement, ultimately stealing the target's network data. S2. Penetration testing uses high-frequency techniques in ATT&CK to conduct security testing on the demand side, develop security solutions, and strengthen enterprise security. In the red-blue confrontation, both the blue and red teams can use the ATT&CK framework. The red team can perform phishing and watering hole attacks, while the blue team can use the techniques in the framework to audit and protect the system website in advance. S3: Create a hacker profile based on the attack behavior and view the attacker's detailed identity information, including the attack source IP, the attacker's real IP, internal network IP, public network IP, network ID, device fingerprint, operating system, and browser used. The Threat Intelligence Center links the corresponding content and displays the corresponding information in the intelligence linkage data module for real people who match the same device fingerprint; S4. Utilize hacker profiling to trace the source of attacks. This technology uses a variety of techniques to capture the attacker's virtual identity, accurately locates multi-dimensional aggregation, restores the attack process, traces the attacker's identity, forms a hacker profiling, and finds the threat at its root. ZombieCookie utilizes a specific algorithm to calculate the device fingerprint information of a specific attacker and uses WebRTC technology to obtain the attacker's detailed IP information. S5. Targeted attacks are targeted and countered. We leverage the ability to reversely control attackers for multi-platform countermeasures, covering common operating systems such as Windows, macOS, and Android. We capture the attacker's real IP address, device fingerprint information, and social ID information, reversely control the attacker's host, and reversely target common hacker attack plug-ins. This includes browser countermeasures, scanner countermeasures, and git countermeasures. Once an attacker accesses a web honeypot, we exploit browser vulnerabilities to extract sensitive information from the attacker's host and reversely control the attacker's host. S6. Use time series graphics to display the attack path and attack chain of the captured attack behavior. Using Kill Chain as a reference model and the ATT&CK framework as a technical supplement, we correspond to each step taken by the attacker and utilize sandbox business simulation, vulnerability camouflage, virtual systems, desensitized false data, and behavior records to display the attack path and attack chain of the captured attack behavior in time series graphics. S7. Build a modular sandbox in the form of building blocks. Use services as the basic resources for building sandboxes. Sandboxes can arbitrarily select one or more services for combination. When building sandboxes, existing templates, custom web classes, and system service sandboxes can be arbitrarily combined. Use different services, different pages, and multiple service combinations in different scenarios to enhance the deception capability of the smart honeynet.
4. The distributed threat intelligent honeynet trapping method according to claim 1, characterized in that: The algorithm flow of the killchain model includes: S1. Count the types of attacks contained in the original data; S2. Classify or define each attack type into a certain stage of the kill chain model based on its characteristics. Next, preprocess the data and construct a feature vector. S3. Eliminate all redundant fields and define the distance calculation method for each dimension of features; S4. Add the kill chain model stage attributes of each attack to the feature vector as features; S5. Using a filtering unsupervised feature selection method to screen feature vectors and construct feature vectors that best reflect the spatial structure correlation of the data set can effectively reduce the resources consumed by system operation; S6. By using the new feature vector and field similarity measurement, an improved spectral clustering algorithm is used to mine the existing network kill chain model attacks from the data; S7. Three different kill chain models are proposed, and the Markov model is used for derivation and analysis to realize the attack prediction function of the network kill chain model.
5. The distributed threat intelligent honeynet trapping method according to claim 4, characterized in that: Each field of the feature vector reflects different attributes of the kill chain model attack process, a i and a j Represent two pieces of data respectively and j>i, details are as follows: Time reflects the relationship between attack events. As time goes by, the correlation between two attack events gradually weakens. The similarity measurement is defined as: The attacker's IP address originates from the same network segment. The same attacker has similarity in the source IP or destination IP in the IDS alarm log. The similarity metric is defined as: Where M = max{H(a i .sIP,a j .sIP),H(a i .sIP,a j .dIP),H(a i .dIP,a j .sIP),H(a i .dIP,a j .dIP)}, the H function is the binary of two IP addresses, indicating the same number of bits from left to right, sIP refers to the source IP address, and dIP refers to the destination IP address; If the attacker uses different tools to successfully intrude into the server, and the two data have the same HTTP request method, server port number, client port number, client environment, and HTTP response code, the similarity metric of these fields can be defined as: The same attacker will be in the same area on the map. Based on the size of the regional communication in the two data, the similarity measure can be defined as: F Locate =(a i .country&a j .country)*0.1+(a i .province&a j .province)*0.2+(a i .city&a j ,city)*0.7; Judging from the characteristics of the kill chain, the result of the attack method in the previous alarm log may be the prerequisite for launching the attack method in the startling alarm log. The similarity measurement can be defined as:
6. The distributed threat intelligent honeynet trapping method according to claim 5, characterized in that: The feature scoring function includes a feature selection algorithm, wherein the importance of each feature of the feature selection algorithm is evaluated by minimizing the reconstruction coefficient between the weight matrix of all features and a single feature, and the feature selection algorithm includes: Given a dataset matrix X = {1, ...,} ∈ g, search for a feature subset of size m that contains the most information. The structure represented by the data points in the m-dimensional space well preserves the intrinsic structure of the dataset in the original d-dimensional space. The local geometric structure of the data set is modeled to construct a simple and effective neighbor graph A, which is generated by all candidate features within a specific neighborhood. No additional parameters are required, and the weight matrix can be calculated according to the following formula: Among them, A ij is the number of neighborhood samples of sample i.
7. The distributed threat intelligent honeynet trapping method according to claim 1, characterized in that: The overall system architecture of the network kill chain model detection platform is divided into a presentation layer, a business logic layer, and a data storage layer from top to bottom.
8. The distributed threat intelligent honeynet trapping method according to claim 1, characterized in that: The presentation layer is mainly composed of the system front-end pages, including the situation awareness homepage, the data display page, and the search page. The display effect of the data display page is displayed in the form of data visualization, including statistical charts. All of the presentation layers are designed and implemented in the form of web pages.
9. The distributed threat intelligent honeynet trapping method according to claim 1, characterized in that: The business logic layer includes an algorithm module and a Web background module. The algorithm module is responsible for implementing the algorithm flow of the kill chain model detection and prediction model of the network, including raw data preprocessing, attack event classification, unsupervised feature selection, killchain model detection, and kill chain model prediction submodules. The Web background module performs data interaction, and the data processing part performs type conversion and normalization on the raw data, and uses the model call to pass the running results of each algorithm module to the Web background module. The Web background module stores the data in the database, and the Web background module interacts with the front end in data transmission. The data storage layer is composed of the data storage module, which completes the data reading and writing functions and directly interacts with the Web background module.
10. The distributed threat intelligent honeynet trapping method according to claim 1, characterized in that: The business logic layer includes a data preprocessing module, a feature selection module, a kill chain model detection module, and a network kill chain model prediction module.
Citation Information
Patent Citations
Cooperative type active defense method based on honeynets
CN103561003A
Distributed switchable industrial control honey net trapping method
CN114157467A
Defense system for cheating
CN115150124A
Cited By
System and method for comprehensively capturing network threats based on multi-dimensional mimicry trapping
CN120729621A
Important activity network security attack preposed early warning method
CN120785661A
Active defense system and method based on multi-protocol dynamic simulation and distributed trapping
CN120856472A