Automatic attack traceability countering process method

Through unsupervised learning models and dynamic honeypot technology, the weaknesses of the traditional network security defense system in advanced threat tracing are resolved, accurate identification and in-depth analysis of unknown attacks are achieved, and the automated response capabilities of network security are improved.

CN120639494APending Publication Date: 2025-09-12STATE GRID HENAN INFORMATION & TELECOMM CO
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511041016.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-28
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Traditional network security defense systems are unable to cope with advanced persistent threats, zero-day vulnerability exploits and new attacks. They have weak threat perception and tracing capabilities, single countermeasures that easily interfere with legitimate business, lack flexibility and deceptiveness, and are unable to achieve self-awareness and enemy awareness in offensive and defensive confrontations.

Method used

Through unsupervised learning models, real-time traffic on the edge side is self-learned and modeled to identify unknown suspicious traffic. A dynamic honeypot environment is deployed for deep interaction and behavior capture to extract attack indicators and form a highly credible attack evidence chain.

Benefits of technology

It achieves accurate identification, in-depth analysis and effective tracing of edge-side attacks, and improves the automated analysis and response capabilities for advanced and unknown attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639494A_ABST
    Figure CN120639494A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network security, and particularly discloses an automatic attack traceability countering process method, which comprises the following steps of: firstly, dynamically associating an isolated early warning event into a global attack context by fusing and analyzing multi-source heterogeneous data such as network traffic, equipment logs and external intelligence and then utilizing a threat knowledge graph technology; and performing intelligent prediction on the next intention of the attacker based on the topological relation of the atlas. Based on this, an instant spoofing environment that is highly matched with the predicted target is dynamically arranged and generated, and attack traffic is non-inductively redirected into the controlled environment. And finally, through transparent takeover and deep interaction of the redirection traffic, comprehensive capture and structural analysis of real behaviors, used tools and techniques and tactics of the attacker are realized. Therefore, while the attack traceability accuracy is improved, the initiative and flexibility of the countering means are enhanced, and an active defense system is constructed for network security defense.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of network security technology, and more specifically, to an automated attack tracing and countermeasure process method. Background Art

[0002] With the rapid development of information technology and the increasing complexity of cyberspace, network security threats are evolving at an unprecedented speed and scale. Traditional network security defense systems, such as firewalls and intrusion detection systems, primarily rely on known attack signatures and rules for passive defense. These systems struggle to cope with the dynamic evolution of attacker behavior during an attack. They are increasingly unable to cope with advanced persistent threats (APTs), zero-day vulnerability exploits, and various new and covert attack methods. Attackers are often able to bypass traditional defenses, lurking within intranets for extended periods and moving laterally, posing significant risks to the critical information assets of enterprises and organizations.

[0003] Existing technologies for attack tracing and countermeasures attempt to combine threat intelligence, log analysis, and security orchestration automation and response (SOAR) technology. These technologies aggregate multi-source logs and threat intelligence to identify potential malicious activity and automatically execute defensive actions such as host isolation and IP blocking based on pre-set scripts. However, these solutions still have significant limitations: First, threat perception and tracing capabilities are weak, often limited to isolated analysis of single alert events. This makes it difficult to construct a complete attack chain through deep correlation of heterogeneous data sources such as network traffic and device logs, and even more difficult to predict an attacker's subsequent behavior based on existing information. Second, countermeasures are limited and rigid, primarily relying on passive defenses such as blocking and isolation. These methods can easily expose defensive intentions and interfere with legitimate business access due to misjudgments, lacking flexibility and deception. Third, they lack the ability to continuously interact with attackers to obtain high-value intelligence such as their techniques, tactics, toolchains, and identity profiles. Furthermore, they are significantly inadequate in guiding attackers to specific environments for in-depth analysis and evidence collection. This makes it difficult to achieve both self-awareness and adversary-level understanding in attack and defense confrontations, and consequently, fails to provide effective data support for the long-term iteration of security capabilities.

[0004] Therefore, we look forward to an optimized automated attack tracing and countermeasure process solution. Summary of the Invention

[0005] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides an automated attack tracing and countermeasure process method, which performs self-learning modeling on the real-time traffic on the edge side through an unsupervised learning model, and intelligently identifies unknown suspicious traffic that deviates from normal behavior patterns in a manner that does not rely on attack samples. Once high-threat traffic is detected, a dynamic honeypot environment will be automatically orchestrated and deployed, and the suspicious traffic will be introduced into an isolated analysis environment through traffic redirection technology. In the honeypot, malicious traffic is deeply interacted and behavior captured to extract high-value attack indicators. Finally, the acquired attack indicators are aggregated and associated with the source information of the traffic to form a complete and highly reliable attack evidence chain. In this way, accurate identification, in-depth analysis and effective tracing of edge-side attacks can be achieved, thereby effectively improving the edge network's automated analysis and response capabilities to advanced and unknown attacks.

[0006] According to one aspect of the present application, a method for automated attack tracing and countermeasure process is provided, which includes:

[0007] Conduct multi-source threat awareness on raw network traffic byte streams, device log text, and external threat intelligence to obtain early warning events;

[0008] The threat knowledge graph is updated based on the warning event to obtain an updated threat knowledge graph;

[0009] In the updated threat knowledge graph, the next hop prediction is performed with the attacker IP node in the warning event as the starting point to obtain the predicted next target;

[0010] Performing real-time deception environment dynamic orchestration based on the predicted next-hop target to obtain surviving honeypot information and redirection instructions;

[0011] Based on the attacker's session information, survival honeypot information and redirection instructions, traffic transparent redirection and interaction takeover are performed to obtain structured attacker behavior logs, captured malware and attacker portrait fragments.

[0012] Compared with the existing technology, the automated attack tracing and countermeasure process method provided by this application uses an unsupervised learning model to self-learn and model the real-time traffic on the edge side, and intelligently identifies unknown suspicious traffic that deviates from normal behavior patterns in a way that does not rely on attack samples. Once high-threat traffic is detected, a dynamic honeypot environment will be automatically orchestrated and deployed, and the suspicious traffic will be introduced into an isolated analysis environment through traffic redirection technology. In the honeypot, malicious traffic is deeply interacted and behavior captured to extract high-value attack indicators. Finally, the acquired attack indicators are aggregated and associated with the source information of the traffic to form a complete and highly reliable attack evidence chain. In this way, accurate identification, in-depth analysis and effective tracing of edge-side attacks can be achieved, thereby effectively improving the edge network's automated analysis and response capabilities to advanced and unknown attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.

[0014] Figure 1 The present invention provides a flowchart of an automated attack source tracing and countermeasure process method according to an embodiment of the present application.

[0015] Figure 2 Schematic diagram of data flow of the automated attack tracing and countermeasure process method according to an embodiment of the present application.

[0016] Figure 3 This is a flowchart of sub-step S1 of the automated attack tracing and countermeasure process method according to an embodiment of the present application.

[0017] Figure 4 This is a flowchart of sub-step S14 of the automated attack source tracing and countermeasure process method according to an embodiment of the present application.

[0018] Figure 5 This is a flowchart of sub-step S142 of the automated attack source tracing and countermeasure process method according to an embodiment of the present application.

[0019] Figure 6 This is a flowchart of sub-step S4 of the automated attack tracing and countermeasure process method according to an embodiment of the present application. DETAILED DESCRIPTION

[0020] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.

[0021] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.

[0022] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.

[0023] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.

[0024] It is worth noting that in this application, all actions to obtain data are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.

[0025] In response to the technical problems described in the above background technology, the present application proposes an automated attack tracing and countermeasure process method, which uses an unsupervised learning model to self-learn real-time traffic on the edge side, and intelligently identifies unknown suspicious traffic that deviates from normal behavior patterns in a way that does not rely on attack samples. Once high-threat traffic is detected, a dynamic honeypot environment will be automatically orchestrated and deployed, and suspicious traffic will be introduced into an isolated analysis environment through traffic redirection technology. In the honeypot, malicious traffic is deeply interacted and behavior captured to extract high-value attack indicators. Finally, the acquired attack indicators are aggregated and associated with the source information of the traffic to form a complete, highly reliable attack evidence chain. In this way, accurate identification, in-depth analysis and effective tracing of edge-side attacks can be achieved, thereby effectively improving the edge network's automated analysis and response capabilities to advanced, unknown attacks.

[0026] Figure 1 The present invention provides a flowchart of an automated attack source tracing and countermeasure process method according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the automated attack source tracing and countermeasure process method according to an embodiment of the present application. Figure 1 and Figure 2 As shown, the automated attack tracing and countermeasure process method includes the following steps: S1, performing multi-source threat perception on the original network traffic byte stream, device log text and external threat intelligence to obtain early warning events; S2, updating the threat knowledge graph based on the early warning events to obtain an updated threat knowledge graph; S3, in the updated threat knowledge graph, performing next-hop prediction with the attacker IP node in the early warning event as the starting point to obtain the predicted next target; S4, performing real-time dynamic orchestration of the deception environment based on the predicted next-hop target to obtain survival honeypot information and redirection instructions; S5, performing transparent traffic redirection and interactive takeover based on the attacker session information, survival honeypot information and redirection instructions to obtain structured attacker behavior logs, captured malware and attacker portrait fragments.

[0027] In the above-mentioned automated attack tracing and countermeasure process method, the step S1 performs multi-source threat perception on the original network traffic byte stream, device log text and external threat intelligence to obtain early warning events. It should be understood that although the original network traffic byte stream can reflect real-time interaction, it lacks contextual information, and the device log text is scattered across various terminals and devices, with heterogeneous formats and isolated existence, while the external threat intelligence contains known attack features but cannot cover unknown attacks. Based on this, the present application integrates the multi-dimensional information of the original network traffic byte stream, device log text and external threat intelligence to extract suspicious signals with security value from multi-source heterogeneous data, integrate the advantages of different data sources, and convert scattered information into structured early warning events, providing initial and reliable input for subsequent attack path prediction and countermeasures. Among them, Figure 3 Flowchart of sub-step S1 of the automated attack source tracing and countermeasure process method according to an embodiment of the present application. Figure 3 As shown, the step S1 includes the following steps: S11, after traffic cleaning and preprocessing the original network traffic byte stream, performing session reorganization on it to obtain a detailed session record; S12, performing normalized parsing on the device log text to obtain a normalized log; S13, performing matching analysis based on the detailed session record and the normalized log with the external threat intelligence to obtain a matching analysis result; S14, in response to the matching analysis result being a successful match, outputting a discovery object; in response to the matching analysis result being an unsuccessful match, inputting the detailed session record into a trained unsupervised anomaly detection model to obtain a discovery object; S15, generating an event for the discovery object to obtain the warning event.

[0028] Specifically, in step S11, after cleaning and preprocessing the original network traffic byte stream, the session is reorganized to obtain a detailed record of the session. It should be understood that since there is a large amount of interference information in the original network traffic byte stream, such as error packets, duplicate packets and protocol redundant data, and it exists in the form of fragmented data packets, it cannot directly reflect the complete communication process. Therefore, the present application removes noise data through traffic cleaning, preprocesses the unified data format, such as timestamps and protocol fields, and then reorganizes the session based on the five-tuple, which can restore the complete interaction timing, message content and behavior logic of the two communicating parties, and provide a complete context for subsequent analysis. In this way, a structured session detailed record containing the start and end time of the session, the communication entity identifier, the interaction message sequence and the payload content is obtained, eliminating the interference caused by data fragmentation and noise, ensuring that subsequent matching with threat intelligence and anomaly detection can be carried out based on a complete and accurate communication process, and improving the quality and availability of basic data.

[0029] Specifically, step S12 performs normalized parsing on the device log text to obtain a normalized log. It should be understood that device log texts come from various types of devices from different manufacturers, such as firewalls, servers, and intrusion detection systems, and have problems with format heterogeneity, confusing field naming, and information redundancy. By normalizing the time format, aligning field names, and extracting key information through normalized parsing, the problem of cross-device logs being unable to be correlated and analyzed can be solved. In this way, a normalized log with unified fields and consistent format is obtained, which achieves semantic unification of logs from different devices, enables them to be effectively associated with detailed session records, eliminates information fragmentation caused by format differences, lays the foundation for multi-source data fusion analysis, and improves the utilization efficiency of log data.

[0030] Specifically, step S13 performs a matching analysis based on the session details and the normalized logs with the external threat intelligence to obtain a matching analysis result. Specifically, the external threat intelligence aggregates information such as known malicious entities and attack characteristics. By performing a multi-dimensional matching of the IP, payload, interaction behavior in the session details and the connection behavior and port information in the normalized logs with the characteristics in the threat intelligence, the parts of the session and log that match the known threats are accurately located, and an analysis result containing matching success / failure, matching threat intelligence entries, and corresponding data locations is generated. This provides a clear basis for the subsequent determination of the discovery object, greatly improving the accuracy and efficiency of identifying known threats and avoiding the omission of known threats.

[0031] Specifically, in step S14, in response to the matching analysis result being a successful match, the discovery object is output; in response to the matching analysis result being an unsuccessful match, the detailed record of the session is input into the trained unsupervised anomaly detection model to obtain the discovery object. Specifically, a successful match indicates the existence of a known threat, and determining the discovery object directly based on the matched threat intelligence can ensure that known threats are accurately captured; when the match is unsuccessful, unknown threats cannot be excluded, and the unsupervised anomaly detection model can identify abnormal behaviors that deviate from the normal pattern by learning the timing patterns of normal sessions, such as changes in packet size and connection frequency, thereby discovering unknown threats, thereby achieving comprehensive coverage of known and unknown threats and avoiding missing new attacks due to reliance on known intelligence. Among them, Figure 4 Flowchart of sub-step S14 of the automated attack source tracing and countermeasure process method according to an embodiment of the present application. Figure 4 As shown, the step S14 includes the following steps: S141, extracting a network traffic timing input vector from the session detail record; S142, extracting a network traffic timing pattern feature coding vector from the network traffic timing input vector; S143, inputting the network traffic timing pattern feature coding vector into the trained unsupervised anomaly detection model to obtain an anomaly score; S144, determining whether to output the discovered object based on a comparison between the anomaly score and a preset threshold.

[0032] More specifically, the step S141 extracts the network traffic timing input vector from the session detail record. It should be understood that although the session detail record presents the complete communication interaction process, the timing information therein exists in an unstructured form, such as the number and size of data packets at different time points and other scattered data, which cannot be directly processed by the model and need to be converted into a structured vector form to capture the changing pattern in the time dimension. Based on this, the present application extracts key indicators that reflect the changes in network traffic over time from the session detail record, such as the number of data packets per second, the mean and variance of the data packet size, the frequency of connection establishment, etc., and combines them in a time series order to form a vector of fixed dimension, so that the timing features can be effectively parsed by subsequent models. In this way, a standardized network traffic timing input vector is obtained, which completely retains the dynamic change information of the traffic in the time dimension, eliminates the format differences of the original data, and provides structured and computable basic data for the subsequent extraction of timing pattern features, thereby ensuring the effectiveness of the feature extraction process.

[0033] Figure 5 Flowchart of sub-step S142 of the automated attack source tracing and countermeasure process method according to an embodiment of the present application. Figure 5As shown, the step S142 includes the steps of: S1421, performing network traffic local time series feature extraction based on one-dimensional convolution coding on the network traffic time series input vector to obtain a sequence of network traffic local time series pattern feature coding vectors; S1422, performing network traffic time series pattern feature information transmission on the sequence of network traffic local time series pattern feature coding vectors to obtain the network traffic time series pattern feature coding vector.

[0034] In a specific example of the present application, the step S1421 performs a one-dimensional convolution coding-based network traffic local time series feature extraction on the network traffic time series input vector to obtain a sequence of network traffic local time series pattern feature coding vectors. Specifically, through a one-dimensional convolution operation, a sliding window calculation is performed on the network traffic time series input vector using convolution kernels of different sizes to capture feature combinations and change patterns within different local time ranges, such as burst features within a short window and trend features within a slightly longer window, and convert the original time series data into a sequence of network traffic local time series pattern feature coding vectors that can reflect the local time series pattern. Each vector corresponds to the key features of a specific local time window, effectively extracting the local correlation patterns hidden in the original data, enhancing the ability to characterize subtle time series changes, and enabling subsequent time series integration to be based on more meaningful local features, thereby improving the accuracy of overall feature extraction.

[0035] In a specific example of the present application, step S1422 includes: first, extracting the maximum eigenvalue of each network traffic local time series pattern feature encoding vector in the sequence of the network traffic local time series pattern feature encoding vector as a network traffic time series significant identifier, which is expressed as follows:

[0036] λ j =max(x j (t j ))

[0037] Among them, t j is the timestamp of the jth network traffic local temporal pattern feature encoding vector in the sequence of the network traffic local temporal pattern feature encoding vector, x j is the eigenvalue of the jth network traffic local time series pattern feature coding vector in the sequence of the network traffic local time series pattern feature coding vector, max is the maximum value, λ j The network traffic temporal salient identifier corresponding to the j-th network traffic local temporal pattern feature encoding vector.

[0038] That is, by capturing the most prominent signal in the feature encoding vector of each local temporal pattern of network traffic, the maximum eigenvalue is used as a proxy indicator of the importance or activation intensity of the network traffic pattern within the corresponding local time window. This gives the system the ability to identify key local events, such as sudden spikes in the number of traffic packets and drastic fluctuations in packet size. This allows subsequent processing of temporal features to not only rely on the natural passage of time, but also focus on important patterns at the content level, providing a basis for differentiated temporal information processing. The generated quantifiable network traffic temporal saliency identifier can effectively distinguish the importance of different local temporal patterns, allowing key local traffic patterns to be clearly marked, laying the foundation for differentiated retention of important information during the subsequent temporal decay propagation process.

[0039] Then, based on the timestamp of each network traffic local time series pattern feature coding vector in the sequence of the network traffic local time series pattern feature coding vector, the time span of each network traffic local time series pattern feature coding vector is calculated, which is expressed as follows:

[0040] Δt j =t current -t j

[0041] Among them, t current The timestamp corresponding to the current network traffic local temporal pattern feature encoding vector, Δt j is the time span corresponding to the j-th network traffic local temporal pattern feature encoding vector.

[0042] That is, the discretely distributed local temporal pattern feature coding vectors of network traffic are converted into a time measurement with continuous physical meaning. Specifically, by accurately calculating the time intervals between each local temporal pattern feature coding vector of network traffic and relative to the reference time point, the discrete information that originally only reflects the sequence position is converted into a continuous quantitative indicator reflecting the actual passage of time, providing a physical basis for constructing a temporal decay model that conforms to natural laws and ensures that the propagation and decay of temporal information can fit the objective characteristics of the passage of time. The specific time interval values ​​between each local temporal pattern feature coding vector of network traffic are obtained. These values ​​convert the discrete local feature sequence into structured information with a continuous time dimension, enabling the subsequent temporal decay propagation process to perform precise influence modulation based on the actual time span, avoiding attenuation model distortion caused by time measurement ambiguity.

[0043] Finally, based on the time span of each network traffic local time series pattern feature coding vector in the sequence of the network traffic local time series pattern feature coding vector and the network traffic time series significant identifier, the sequence of the network traffic local time series pattern feature coding vector is subjected to time series attenuation propagation to obtain the network traffic time series pattern feature coding vector, which is expressed as follows:

[0044]

[0045] v f =LSTM(H)

[0046] Where β>0 is a learnable decay coefficient that controls the rate of time decay (which is optimized by backpropagation), N(i) is the neighborhood set of the current network traffic local temporal pattern feature encoding vector, and λ k is the network traffic time series significant identifier corresponding to the kth network traffic local time series pattern feature encoding vector, which is used to enhance the contribution of important neighbors. exp is the logarithmic function value with the natural constant e as the base, exp(-β(Δt j ) 2 ) is the Gaussian decay kernel (decays more slowly for recent features), w j (t) is the time series attenuation factor, H is the sequence distribution of the feature encoding vector of the local time series pattern of network traffic, h1,h2,h j ,h n are the 1st, 2nd, jth and nth network traffic local time series pattern feature encoding vectors in the sequence distribution of network traffic local time series pattern feature encoding vectors, LSTM is the LSTM encoder, and v f Encode vectors for network traffic temporal pattern features.

[0047] Specifically, a unified modulation framework is constructed that integrates temporal physical properties and content importance to achieve precise aggregation of network traffic time series information. By combining the natural decay law corresponding to time spans with the importance modulation corresponding to network traffic time series salient identifiers, the influence of historical local time series patterns is rationally attenuated over time, while highly significant key patterns maintain a more persistent influence during propagation. This integrates the long-term dependencies of local time series patterns and captures the core laws in traffic time series. The result is a precisely distilled network traffic time series pattern feature encoding vector. This not only condenses the key information of each local time series pattern but also quantifies the sustained influence of key patterns at different time points through the time series decay propagation mechanism. This effectively characterizes the essential characteristics of network traffic in the temporal dimension, provides highly discernible input for subsequent unsupervised anomaly detection models, and improves the ability to identify complex and hidden traffic anomaly patterns.

[0048] More specifically, step S143 inputs the network traffic temporal pattern feature encoding vector into the trained unsupervised anomaly detection model to obtain an anomaly score. Specifically, the unsupervised anomaly detection model analyzes the network traffic temporal pattern feature encoding vector, calculates its deviation from the normal traffic feature distribution, and outputs a quantitative score reflecting the degree of anomaly. This effectively distinguishes between normal and abnormal traffic patterns, with a high score corresponding to a high degree of suspicion. This provides a reliable quantitative basis for discovering unknown threats and improves the ability to identify new and hidden attacks.

[0049] During implementation, the training process of the unsupervised anomaly detection model is as follows: First, training data is collected and preprocessed. Raw network traffic byte streams are collected from the enterprise during normal business hours, covering traffic data from different business scenarios to ensure data diversity and representativeness. The collected raw traffic is cleaned to remove error packets, duplicate frames, and link layer redundant information, while retaining valid transport layer data such as TCP and UDP. Sessions are then reassembled according to the five-tuple structure: source IP, destination IP, source port, destination port, and protocol. The fragmented data packets are restored into complete sessions, each containing start and end times, interaction message sequences, and payload features. The reassembled normal sessions are segmented into fixed time windows, such as a 10-second window. Time series feature indicators are extracted from each window, including the number of packets per second, average packet length, packet length variance, and connection state switching frequency. These indicators are arranged in chronological order to form a network traffic time series input vector with uniform dimensions, thus constructing a training dataset of 100,000 vectors.

[0050] Next, preparations for feature encoding layer training are performed. For each network traffic time series input vector in the training dataset, local time series features are extracted through a one-dimensional convolution operation: A sliding convolution operation is performed on the vector using multiple convolution kernels of different sizes. Each convolution kernel outputs a feature response within a corresponding local window, forming multiple sets of local time series pattern feature sequences. For each set of local feature sequences, the maximum eigenvalue of each local feature is extracted as a temporal salient identifier, and the time span of adjacent local features is calculated based on the timestamp, for example, a window interval of 1 second. Local features are integrated through a temporal decay propagation mechanism, where the influence of the local feature corresponding to the salient identifier on subsequent time windows decays exponentially with increasing time spans. For example, 80% of the weight is retained for a time span of 1 second, 60% for a span of 2 seconds, and the weight decays to 0. The influence of non-salient features decays even faster. Ultimately, all local features are integrated into a fixed-length network traffic time series pattern feature encoding vector to form the training sample set.

[0051] Subsequently, the model architecture was designed and training iterations were carried out. An autoencoder was used as the foundational architecture for the unsupervised anomaly detection model, consisting of an encoder and a decoder. The encoder consists of a three-layer fully connected neural network, sequentially mapping the 64-dimensional feature encoding vector to 32-dimensional and 16-dimensional hidden spaces, ultimately outputting an 8-dimensional core feature vector, achieving a compressed representation of normal traffic patterns. The decoder, with three symmetrically connected layers, gradually reconstructs the 8-dimensional core feature vector into a 64-dimensional vector. By minimizing the mean squared error between the input vector and the reconstructed vector, the model learns the reconstruction patterns of normal traffic. During training, the feature encoding vector sample set was split into a training set and a validation set with an 8:2 ratio. The initial learning rate was set to 0.001, the batch size was set to 128, and the Adam optimizer was used to minimize the reconstruction loss. During training, the average reconstruction error of the validation set was calculated every 100 iterations. Training was terminated when the validation error decreased by less than 0.0001 for 10 consecutive iterations, indicating that the model has fully learned the temporal patterns of normal traffic, such as the frequency of HTTP requests from office terminals and the fluctuation range of packet sizes in file transfers.

[0052] Finally, the model is evaluated and optimized. 1,000 normal samples are randomly sampled from the validation set, and the distribution characteristics of their reconstruction errors are calculated to determine the error threshold of the normal mode. For example, the 99.9% percentile is used as the initial anomaly judgment benchmark. At the same time, a small number of labeled abnormal samples are introduced, such as simulated port scans and malicious payload transmission traffic, to test the model's reconstruction error for abnormal samples. If the average error of abnormal samples is significantly higher than the normal threshold, such as more than three times the normal threshold, the model performance meets the standard. If some normal samples are misclassified as abnormal, that is, the reconstruction error exceeds the threshold, retraining is performed by adding training data for the corresponding scenario; if the recognition rate of abnormal samples is insufficient, the convolution kernel size and attenuation coefficient are adjusted to enhance the model's ability to capture local burst features. Ultimately, a stable unsupervised anomaly detection model is formed, which can quantify the degree of deviation of input traffic from the normal mode through reconstruction error, providing a reliable basis for identifying unknown threats.

[0053] More specifically, the step S144 determines whether to output the discovered object based on the comparison between the anomaly score and the preset threshold. It should be understood that in order to avoid misjudgment or missed judgment due to subjective judgment and to ensure that only anomalies that reach a certain degree of suspicion are identified as potential threats, this application determines to set a reasonable threshold based on the score distribution of normal and abnormal traffic in historical data. By comparing the anomaly score with the threshold, truly threatening abnormal behaviors are screened out, and highly suspicious discovered objects are accurately output, which avoids omissions due to excessively high thresholds and reduces false alarms due to excessively low thresholds, thereby ensuring the accuracy and effectiveness of threat discovery. It is worth mentioning that in response to the anomaly score being greater than the preset threshold, the discovered object is determined to be output.

[0054] Specifically, in step S15, an event is generated for the discovered object to obtain the warning event. It should be understood that the discovered object is scattered threat-related information, lacks structured key elements, and cannot be directly used for threat knowledge graph updates and subsequent processes. Through event generation, it is converted into a structured event containing elements such as time range, subject information, behavior description, threat level, and related entities, which can meet the standard format requirements for subsequent processing. The generated warning event has a standardized structure and complete information, accurately carries the core information of the threat, and can be directly used as input for threat knowledge graph updates, ensuring that subsequent attack path prediction, countermeasures arrangement and other steps can be carried out based on reliable and structured data, thereby improving the consistency and effectiveness of the entire attack tracing and countermeasure process.

[0055] In the above-mentioned automated attack tracing and countermeasure process method, the step S2 updates the threat knowledge graph based on the warning event to obtain an updated threat knowledge graph. In a specific example of the present application, the step S2 includes: integrating the entities and relationships in the warning event as new nodes and edges into the threat knowledge graph to obtain the updated threat knowledge graph. It should be understood that the threat knowledge graph, as a structured carrier for storing attack entities and relationships, needs to be dynamically iterated with new threat events to reflect the latest attack situation. The warning event contains new entities and relationships between entities such as attacker IP, target port, and attack behavior. If it is not integrated in time, the graph will not be able to support subsequent accurate predictions due to information lag. Specifically, the present application extracts entities and relationships between entities in the warning event, integrates them into the existing graph as new nodes and edges, realizes real-time updating of the graph, and forms an updated threat knowledge graph containing the latest threat information. The new nodes and edges fill the information gaps in the original graph, making the relationship between attack entities more complete, and providing more comprehensive topological relationship support for subsequent attacker behavior prediction.

[0056] In the above-mentioned automated attack tracing and countermeasure process method, in the updated threat knowledge graph, the attacker IP node in the warning event is used as the starting point to perform a next hop prediction to obtain the predicted next target. In a specific example of the present application, the step S3 includes: taking the attacker IP node in the warning event as the starting point, and performing a probability walk in the updated threat knowledge graph based on the weighted path search of the historical frequency to obtain the predicted next target. It should be understood that since attackers usually follow a certain behavioral path in the network, such as from scanning ports to attempting to log in to lateral movement, the frequency of historical paths in the updated threat knowledge graph reflects behavioral preferences. Therefore, this application is based on the attacker IP node, traversing all reachable paths in the graph, assigning weights to the paths according to the historical frequency of occurrence. The higher the frequency, the greater the weight. The access probability of each potential target is calculated through a probability walk, and the attacker's most likely next attack target is accurately located, providing a clear direction for the subsequent arrangement of the dynamic deception environment, and improving the pertinence of the countermeasures.

[0057] In the above-mentioned automated attack tracing and countermeasure process method, the step S4 performs dynamic orchestration of an instant deception environment based on the predicted next-hop target to obtain surviving honeypot information and redirection instructions. Specifically, a fixed honeypot environment is difficult to match the attacker's diverse attack targets and is easy to identify, while dynamic orchestration can quickly build a deception environment that fits the scenario based on the predicted target, thereby improving the success rate of trapping. At the same time, accurate redirection instructions need to be generated to ensure that the attack traffic is introduced without feeling. Specifically, the present application parses the type of predicted target, such as the service type corresponding to the database port, selects an adapted trapping template, injects dynamic scenarios, such as simulating the configuration information of a real database, generates a surviving honeypot through containerized deployment, and generates redirection rules that direct the attack traffic to the honeypot, thereby quickly deploying a deception environment that is highly matched with the predicted target. The honeypot has a high degree of simulation and is not easy to be detected. The redirection instructions ensure the accurate introduction of the attack traffic, laying the foundation for subsequent behavior capture and analysis. Among them, Figure 6 Flowchart of sub-step S4 of the automated attack source tracing countermeasure process method according to an embodiment of the present application. Figure 6 As shown, the step S4 includes the following steps: S41, parsing the predicted next-hop target to obtain the type of honeypot that needs to be orchestrated; S42, selecting a trapping template based on the type of honeypot that needs to be orchestrated to obtain a trapping template; S43, dynamically injecting the trapping template into the scenario to obtain a final honeypot deployment list to be deployed; S44, performing containerized deployment based on the final honeypot deployment list to be deployed to obtain surviving honeypot information and redirection instructions.

[0058] Specifically, step S41 parses the predicted next-hop target to obtain the honeypot type to be orchestrated. That is, core attributes are extracted from the predicted next-hop target, such as the service type corresponding to the port and the target system type. This clarifies the specific service and environment characteristics that the honeypot needs to simulate, and determines a honeypot type that is highly compatible with the predicted target. This ensures that subsequently deployed honeypots can accurately simulate the target environment, increasing their attractiveness to attackers and their success rate in trapping them, and providing a precise basis for subsequent template selection.

[0059] Specifically, the step S42 selects a trapping template based on the type of honeypot that needs to be arranged to obtain a trapping template. It should be understood that different types of honeypots need to have specific basic configurations, service simulation logic and interaction rules. Building directly from scratch will lead to low efficiency and prone to configuration vulnerabilities. The trapping template contains the basic framework and core functions of the corresponding type of honeypot, which can serve as the basis for rapid deployment. Therefore, according to the determined honeypot type, this application screens out a trapping template from the template library that contains the basic service programs, default configurations, and interactive response logic required for this type of honeypot. The template already has basic service simulation capabilities, which reduces the cost of repeated development, while ensuring the integrity and standardization of the basic functions of the honeypot, and providing a standardized framework for subsequent dynamic scenario injection.

[0060] Specifically, in step S43, dynamic scenario injection is performed on the trapping template to obtain the final honeypot deployment list to be deployed. It should be understood that dynamic scenario injection can add details close to the real environment and enhance the deceptiveness and simulation of the honeypot. Based on this, the present application injects dynamic scenario information related to the predicted target on the basis of the trapping template, such as false user operation records, simulated system vulnerabilities, and personalized configuration parameters, so that the honeypot presents behavioral characteristics that conform to the real scene, forming a detailed list that can be directly deployed. The final honeypot deployment list to be deployed, which includes dynamic scenarios, is obtained in this way. The honeypot not only has basic service functions, but also can simulate the detailed characteristics of the real environment, greatly reducing the probability of being discovered by attackers and improving the success rate of trapping.

[0061] Specifically, step S44 performs containerized deployment based on the final honeypot deployment list to obtain surviving honeypot information and redirection instructions. Specifically, this application relies on containerized deployment technology to quickly deploy honeypot instances according to the deployment list, obtain the survival status and network information of the honeypot, and generate network forwarding rules that direct attack traffic from the predicted target to the honeypot. The redirection instructions ensure that the attack traffic enters the honeypot without any sense, providing a stable environment support for the subsequent capture and analysis of the attacker's behavior.

[0062] In the above-mentioned automated attack tracing and countermeasure process method, step S5 performs transparent traffic redirection and interactive takeover based on the attacker's session information, survival honeypot information and redirection instructions to obtain structured attacker behavior logs, captured malware and attacker portrait fragments. It should be understood that transparent traffic redirection can introduce traffic into the honeypot without being noticed by the attacker. At the same time, relying solely on the honeypot to passively receive traffic cannot fully record the attacker's interactive behavior. It is necessary to capture details by taking over the interactive process. If there is a lack of structured records and malware capture, it will be difficult to form an effective attack evidence chain and attacker portrait, and it will be impossible to support subsequent tracing and countermeasures. Based on this, the present application ensures traffic directionality based on the attacker's session information, combines survival honeypot information and redirection instructions to achieve seamless forwarding of traffic to the honeypot, records all the attacker's operational behaviors through interactive takeover, extracts malware samples, and aggregates information to form attacker portrait fragments, providing complete data for tracing analysis.

[0063] In summary, the automated attack tracing and countermeasure process method based on the embodiment of the present application is explained, which uses an unsupervised learning model to self-learn real-time traffic on the edge side, and intelligently identifies unknown suspicious traffic that deviates from normal behavior patterns in a manner that does not rely on attack samples. Once high-threat traffic is detected, a dynamic honeypot environment will be automatically orchestrated and deployed, and suspicious traffic will be introduced into an isolated analysis environment through traffic redirection technology. In the honeypot, malicious traffic is deeply interacted and behavior captured to extract high-value attack indicators. Finally, the acquired attack indicators are aggregated and associated with the source information of the traffic to form a complete, highly reliable attack evidence chain. In this way, accurate identification, in-depth analysis and effective tracing of edge-side attacks can be achieved, thereby effectively improving the edge network's automated analysis and response capabilities to advanced, unknown attacks.

[0064] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in the present invention are merely illustrative and non-limiting, and should not be construed as necessarily possessed by each embodiment of the present invention. Furthermore, the specific details of the above embodiments are provided for illustrative purposes and to facilitate understanding, and are not intended to be limiting. These details do not necessarily limit the present invention to being implemented using these specific details.

[0065] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0066] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be encompassed therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

[0067] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units stated in the system claims can also be implemented by one unit through software or hardware.

[0068] Finally, it should be noted that the above description has been provided for purposes of illustration and description. Furthermore, the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to be limiting. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art will appreciate that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. An automated attack tracing and countermeasure process method, characterized in that: include: Conduct multi-source threat awareness on raw network traffic byte streams, device log text, and external threat intelligence to obtain early warning events; The threat knowledge graph is updated based on the warning event to obtain an updated threat knowledge graph; In the updated threat knowledge graph, the next hop prediction is performed with the attacker IP node in the warning event as the starting point to obtain the predicted next target; Performing real-time deception environment dynamic orchestration based on the predicted next-hop target to obtain surviving honeypot information and redirection instructions; Based on the attacker's session information, survival honeypot information and redirection instructions, traffic transparent redirection and interaction takeover are performed to obtain structured attacker behavior logs, captured malware and attacker portrait fragments.

2. The automated attack tracing and countermeasure process method according to claim 1 is characterized in that: Multi-source threat awareness is performed on raw network traffic byte streams, device log text, and external threat intelligence to obtain early warning events, including: After cleaning and preprocessing the original network traffic byte stream, reorganizing the session to obtain a detailed session record; Performing normalized parsing on the device log text to obtain a normalized log; Performing a matching analysis based on the detailed session record and the normalized log with the external threat intelligence to obtain a matching analysis result; In response to the matching analysis result being a successful match, outputting a discovery object; in response to the matching analysis result being an unsuccessful match, inputting the session detail record into a trained unsupervised anomaly detection model to obtain a discovery object; An event is generated for the discovered object to obtain the warning event.

3. The automated attack tracing and countermeasure process method according to claim 2 is characterized in that: In response to the matching analysis result being unsuccessful, inputting the session details record into a trained unsupervised anomaly detection model to obtain a discovery object, including: extracting a network traffic timing input vector from the session detail record; Extracting a network traffic time series pattern feature encoding vector from the network traffic time series input vector; Inputting the network traffic time series pattern feature encoding vector into the trained unsupervised anomaly detection model to obtain an anomaly score; Based on a comparison between the anomaly score and a preset threshold, it is determined whether to output the found object.

4. The automated attack tracing and countermeasure process method according to claim 3 is characterized in that: Extracting a network traffic time series pattern feature encoding vector from the network traffic time series input vector includes: Performing one-dimensional convolutional coding-based network traffic local time series feature extraction on the network traffic time series input vector to obtain a sequence of network traffic local time series pattern feature coding vectors; The network traffic temporal pattern feature information is transferred to the sequence of the network traffic local temporal pattern feature coding vectors to obtain the network traffic temporal pattern feature coding vectors.

5. The automated attack tracing and countermeasure process method according to claim 4 is characterized in that: The method of transferring network traffic temporal pattern feature information to the sequence of the network traffic local temporal pattern feature coding vectors to obtain the network traffic temporal pattern feature coding vectors includes: Extracting the maximum eigenvalue of each network traffic local time series pattern feature coding vector in the sequence of the network traffic local time series pattern feature coding vectors as a network traffic time series significant identifier; Calculating the time span of each network traffic local time series pattern feature coding vector based on the timestamp of each network traffic local time series pattern feature coding vector in the sequence of the network traffic local time series pattern feature coding vector; Based on the time span of each network traffic local time series pattern feature coding vector in the sequence of the network traffic local time series pattern feature coding vector and the network traffic time series significant identifier, the sequence of the network traffic local time series pattern feature coding vector is subjected to time attenuation propagation to obtain the network traffic time series pattern feature coding vector.

6. The automated attack tracing and countermeasure process method according to claim 1 is characterized in that: The threat knowledge graph is updated based on the warning events to obtain an updated threat knowledge graph, including: The entities and relationships in the warning event are integrated into the threat knowledge graph as new nodes and edges to obtain the updated threat knowledge graph.

7. The automated attack tracing and countermeasure process method according to claim 6 is characterized in that: In the updated threat knowledge graph, the next hop prediction is performed starting from the attacker IP node in the warning event to obtain the predicted next target, including: Taking the attacker IP node in the warning event as the starting point, a weighted path search based on historical frequency is performed on the updated threat knowledge graph to perform a probabilistic walk to obtain the predicted next target.

8. The automated attack tracing and countermeasure process method according to claim 1 is characterized in that: Based on the predicted next hop target, a real-time deception environment is dynamically orchestrated to obtain surviving honeypot information and redirection instructions, including: Parsing the predicted next-hop target to obtain the type of honeypot that needs to be orchestrated; Selecting a trapping template based on the honeypot type that needs to be arranged to obtain a trapping template; Performing dynamic scenario injection on the trapping template to obtain a final honeypot deployment list to be deployed; Containerized deployment is performed based on the final honeypot deployment list to be deployed to obtain surviving honeypot information and redirection instructions.

Citation Information

Cited By

  • Network attack active countering method and device based on dynamic intelligent honeypot

    CN121077821A