Hacker threat traffic data tracing method, system and device and medium
By building traffic fingerprints and combining improved probability packet marking algorithms and lightweight GRU models, the problem of weak anti-obfuscation capabilities of existing traffic traceability methods is solved, and the attack paths are accurately tracked and cross-domain attacks are determined without sharing data, which improves network security protection capabilities.
Patent Information
- Application Number
- CN202510665068.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-29
AI Technical Summary
The existing traffic traceability methods have weak anti-obfuscation capabilities, and attackers can easily bypass detection through encryption, tunneling or camouflage technology, making the traceability process difficult.
By obtaining the original traffic data flow, extracting traffic characteristics and building traffic fingerprints, using the lightweight Transformer model to generate compressed fingerprint vectors, combining the improved probability packet marking algorithm to track the attack path, and training the lightweight GRU model at the edge nodes, the central server performs model parameter aggregation, defines the attack similarity function to determine cross-domain attacks, and achieves accurate traceability.
On the premise of not sharing original data, cross-domain collaboration is achieved to protect data privacy, improve the accuracy of attack path reconstruction and the accuracy of cross-domain attack judgments, and form a complete tracing and protection closed loop, reducing losses caused by attacks.
Smart Images

Figure CN120567463A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network security technology, and in particular to a method, system, device and medium for tracing traffic data of hacker threats. Background Art
[0002] With the rapid development of internet technology and the widespread adoption of network applications, cyberattack techniques continue to evolve, and hacker attack methods are becoming more covert, complex, and cross-domain. Traditional network security defenses rely primarily on technologies such as firewalls, intrusion detection systems, and security audits. However, in the face of new advanced persistent threats, these can only identify attack behaviors but struggle to accurately trace their source. Traffic data tracing, a key component of network security protection, is crucial for revealing attacker identities, identifying attack chains, and preventing similar attacks. Common traffic tracing methods currently include IP backtracking, log analysis, and packet marking. While these methods offer advantages in specific scenarios, they also face numerous challenges.
[0003] However, traditional traffic tracing methods have several shortcomings. IP backtracking technology relies heavily on network infrastructure to support reverse path verification, resulting in high deployment costs and difficulty in combating IP spoofing attacks. Traditional log analysis requires storing raw user traffic data, posing a serious risk of privacy breaches and violating the principle of data minimization. Most critically, existing tracing methods lack anti-obfuscation capabilities, allowing attackers to easily circumvent detection through encryption, tunneling, or camouflage techniques, making tracing difficult. Summary of the Invention
[0004] In view of the above existing problems, the present invention is proposed.
[0005] Therefore, the present invention provides a method, system, device and medium for tracing traffic data of hacker threats to solve the problem that the existing tracing methods have weak anti-obfuscation capabilities and attackers can easily bypass detection through encryption, tunneling or camouflage technology, making the tracing process difficult.
[0006] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0007] In a first aspect, the present invention provides a method for tracing traffic data of hacker threats, comprising:
[0008] Obtaining an original traffic data stream, extracting features of the original traffic data stream and constructing a traffic fingerprint;
[0009] Based on the traffic fingerprint, tracking the attack path using a first algorithm, and determining the attack propagation trajectory through the attack path;
[0010] According to the attack propagation trajectory, a second algorithm is used to analyze the similarity between different attack events to determine cross-domain attack behavior;
[0011] Based on the cross-domain attack behavior determination result and the attack propagation trajectory, the attack source is obtained.
[0012] As a preferred solution of the method for tracing the traffic data source of hacker threats described in the present invention, the step of obtaining the traffic fingerprint includes:
[0013] Collect raw traffic data streams;
[0014] Extracting traffic statistical features and time series features from the original traffic data stream;
[0015] Optimizing the timing characteristics;
[0016] Splicing the traffic statistical features and the optimized time series features to form a fused feature vector;
[0017] The fused feature vector is input into an encoder to generate a compressed fingerprint vector, which is used as the traffic fingerprint.
[0018] As a preferred solution of the method for tracing the source of traffic data of hacker threats described in the present invention, the step of tracing the attack path by the first algorithm includes:
[0019] Inserting tags into data packets passing through the router with a preset probability;
[0020] When the data packet arrives at the receiving end, all tags carried by the data packet are collected to form a tag set;
[0021] Filtering out forged tags by verifying the legitimacy of the hash chain of each tag in the tag set;
[0022] The remaining legal tags in the tag set are sorted based on their timestamps, and the routers that the data packet passes through in sequence are collected to reconstruct the attack path.
[0023] The beneficial effects of this preferred technical solution are: the first algorithm ensures the accurate reconstruction of the attack path, effectively prevents attackers from tampering with the tag information through the hash verification mechanism, and achieves accurate attack path tracking while ensuring network performance. The accuracy of attack path reconstruction is significantly improved, and at the same time, the increase in network load is controlled within a reasonable range.
[0024] As a preferred solution of the method for tracing the source of traffic data of hacker threats described in the present invention, the step of analyzing the similarity between different attack events by the second algorithm includes:
[0025] Using the collected traffic data locally at the edge node, a lightweight gated recurrent unit model is trained, so that the lightweight gated recurrent unit model learns the mapping relationship between the traffic fingerprint and the attack type;
[0026] Send the local training model parameters of the edge node to the central server;
[0027] Define attack similarity function;
[0028] A similarity threshold is set. When the attack similarity calculated by the attack similarity function is greater than the similarity threshold, it is determined to be a cross-domain attack initiated by the same attacker.
[0029] The beneficial effects of this preferred technical solution are: it realizes cross-domain collaboration without sharing original data, effectively protects data privacy, and at the same time establishes a mapping relationship between traffic fingerprints and attack types by training a lightweight GRU model, providing technical support for cross-domain attack judgment.
[0030] As a preferred solution of the method for tracing the source of traffic data of hacker threats described in the present invention, security protection measures are taken after obtaining the attack source information, and the security protection measures include:
[0031] Block the network connection of the attack source in real time;
[0032] Based on the characteristics of the attack source and the attack propagation trajectory, security reinforcement is performed on the attacked network area;
[0033] Storing the attack source information, the attack propagation trajectory and the attack type characteristics in a security knowledge base;
[0034] Send warning information to other network nodes.
[0035] The beneficial effects of this preferred technical solution are: the protection measures based on the traceability results form a complete traceability and protection closed loop, and through real-time blocking, regional reinforcement, knowledge base construction and early warning distribution, it minimizes the losses caused by the attack, improves the security protection capabilities of the entire network environment, and realizes the transformation from passive defense to active prevention.
[0036] As a preferred solution of the method for tracing the source of traffic data of hacker threats described in the present invention, wherein: the attack similarity function Expressed as:
[0037]
[0038] Where, is the cosine similarity, which is used to measure the two traffic fingerprints and the degree of similarity between them; is the time decay factor, where t1 and t2 are the time when the two attacks occurred, respectively, and λ is the decay coefficient, which is used to control the influence of time on similarity.
[0039] As a preferred solution of the method for tracing the source of traffic data of hacker threats described in the present invention, the central server performs weighted averaging based on the data volume proportion of each edge node and updates the global model parameters;
[0040] Among them, the global model parameter update formula is:
[0041]
[0042] Where |D| is the total amount of data, θ global is the global model parameter, is the global model parameter updated after the t+1th iteration, N is the total number of edge nodes participating in federated learning, |Di| is the amount of local data of node i, are the local model parameters obtained after local training of node i.
[0043] In a second aspect, the present invention provides a traffic data tracing system for hacker threats, comprising a fingerprint construction module, an attack path tracing module, an attack similarity analysis module, and an attack source determination module;
[0044] The fingerprint building module is responsible for acquiring the original traffic data stream, extracting features from the acquired original traffic data stream, and building a traffic fingerprint;
[0045] The attack path tracing module uses a first algorithm to trace the attack path based on the constructed traffic fingerprint and determines the attack propagation trajectory;
[0046] The attack similarity analysis module uses a second algorithm to analyze the similarities between different attack events based on the attack propagation trajectory and determines cross-domain attack behavior;
[0047] The attack source determination module obtains the attack source based on the determination result of the cross-domain attack behavior and the attack propagation trajectory.
[0048] In a third aspect, the present invention provides an electronic device, comprising:
[0049] memory and processor;
[0050] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method for tracing the traffic data of hacker threats are implemented.
[0051] In a fourth aspect, the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method for tracing the traffic data of hacker threats.
[0052] Compared with the prior art, the present invention has the following beneficial effects:
[0053] By extracting traffic statistics (packet length distribution, time interval variance) and temporal features (bidirectional flow periodicity, entropy gradient), and applying wavelet transform technology to deeply explore features at different frequency components, the system comprehensively presents traffic characteristics. Furthermore, by utilizing lightweight Transformer model encoding and its attention mechanism, the system efficiently compresses feature vector dimensions and enhances anti-obfuscation capabilities. Even in the face of interference from encryption, tunneling, or camouflage technologies, it can accurately identify original traffic characteristics, laying a solid foundation for traceability analysis.
[0054] The improved probabilistic packet marking algorithm addresses the shortcomings of traditional algorithms. The marking probability p is carefully set to balance marking efficiency and network overhead. The marking structure incorporates router identification, timestamps, and a hash value based on a pre-shared private key to ensure tag integrity and authenticity, resisting tampering by attackers. During path reconstruction, the receiver verifies the hash chain to filter out forged tags and reconstructs the attack path by sorting legitimate tags based on timestamps, significantly improving tracking accuracy and algorithm security.
[0055] The multi-source correlation traceability model utilizes a federated learning architecture to address the data privacy risks of traditional centralized learning. Edge nodes use a lightweight GRU model to process local data and learn the mapping between traffic fingerprints and attack types. The central server aggregates parameters based on data volume and optimizes the global model. For traceability decisions, an attack similarity function is defined that comprehensively considers traffic fingerprint similarity and attack timing correlation. A threshold is set to determine cross-domain attacks, enabling precise traceability and adapting to complex network attack environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0057] Figure 1 The figure is a schematic diagram of the overall process of a method for tracing traffic data of hacker threats according to an embodiment of the present invention.
[0058] Figure 2 A flow chart illustrating the flow of traffic fingerprints in a method for tracing traffic data of hacker threats according to an embodiment of the present invention.
[0059] Figure 3 The present invention is a flowchart illustrating an attack source acquisition process of a method for tracing traffic data of hacker threats according to an embodiment of the present invention. DETAILED DESCRIPTION
[0060] To make the above-mentioned objects, features, and advantages of the present invention more clearly understood, the following detailed description of the specific embodiments of the present invention is given in conjunction with the accompanying drawings. It is obvious that the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary persons in this field without creative work should fall within the scope of protection of the present invention.
[0061] Example 1, with reference to Figure 1 , as one embodiment of the present invention, provides a method for tracing traffic data of hacker threats, comprising:
[0062] S100: Obtaining the original traffic data stream, extracting the features of the original traffic data stream and constructing a traffic fingerprint;
[0063] S200: Based on the traffic fingerprint, the attack path is tracked using a first algorithm, and the attack propagation trajectory is determined through the attack path;
[0064] S300: Based on the attack propagation trajectory, a second algorithm is used to analyze the similarities between different attack events to determine cross-domain attack behavior;
[0065] S400: Based on the cross-domain attack behavior determination result and the attack propagation trajectory, the attack source is obtained.
[0066] It should be noted that with the continuous evolution of cyberattack technology, hacker attack methods are becoming more covert, complex, and cross-domain. Traditional traffic tracing methods (such as IP backtracking, log analysis, and packet marking) face challenges such as high dependence on infrastructure, high computational overhead, privacy risks, and weak anti-obfuscation capabilities. Especially in the current network environment, attackers can easily evade detection through encryption, tunneling, or camouflage technologies (such as Tor and VPNs), making tracing extremely difficult.
[0067] Therefore, to address the aforementioned network security tracing issues, a complete hacker threat traffic data tracing method was constructed through steps S100-S400. This method first constructs a traffic fingerprint with anti-obfuscation capabilities through multi-dimensional feature extraction and a lightweight Transformer model; then, an improved probabilistic packet marking algorithm is used to accurately track the attack path; then, a multi-source correlation tracing model is used to determine cross-domain attack behavior while protecting data privacy; finally, the attack source is accurately located based on the tracing results. Through the organic combination of this series of technical means, the ability to trace the source of various attacks, especially encryption and tunnel camouflage attacks, has been significantly improved, providing strong technical support for network security protection.
[0068] Example 2, reference Figures 1 to 3 , is an embodiment of the present invention. Based on the above embodiments, a method for tracing the traffic data of hacker threats is provided.
[0069] In the embodiment of the present application, the specific implementation of obtaining the original traffic data stream, extracting the features of the original traffic data stream and constructing the traffic fingerprint in step S100 is as follows:
[0070] First, the original traffic data stream F={p1,p2,…,p n} Among them, p i Represents data packets, which contain rich information such as source address, destination address, data content, etc.
[0071] Then, traffic statistics and timing features are extracted from the original traffic data stream. Traffic statistics include packet length distribution μL and time interval variance σL2. Packet length distribution reflects the distribution of data packet lengths. Different types of network applications and attack behaviors often have specific packet length distribution patterns; time interval variance is used to measure the degree of change in the time interval between data packet arrivals, which can reflect the stability of network traffic. Timing features include bidirectional flow periodicity Δtvar and entropy gradient Hentropy. Bidirectional flow periodicity reflects the time law of bidirectional data transmission in network connections; entropy gradient is used to measure the uncertainty change trend of traffic data. Combining these features, a statistical feature vector is formed.
[0072] Next, we optimize the time series features. We process the original traffic data stream F using wavelet transform technology and set the scale parameter to 5 to obtain richer time series features.
[0073] Subsequently, the traffic statistical features SF and the time series features TF obtained by wavelet transform are concatenated to form a fused feature vector SF⊕TF.
[0074] Finally, the fused feature vector is input into the encoder of the lightweight Transformer model to generate the compressed fingerprint vector V F =Transformer Encoder (S F ⊕T F This fingerprint encoding method based on the lightweight Transformer model can not only effectively compress the dimension of the feature vector and reduce computational complexity, but also make full use of the model's self-attention mechanism to better capture the correlation between features, thereby generating a fingerprint vector that is highly representative and anti-aliasing.
[0075] In the embodiment of the present application, steps A1 to A5 are used to obtain the traffic fingerprint:
[0076] A1: Collects raw traffic data streams;
[0077] A2: Extract traffic statistics and time series features from the original traffic data stream;
[0078] A3: Optimize timing characteristics;
[0079] A4: Combine the traffic statistics features and the optimized time series features to form a fused feature vector.
[0080] A5: Input the fused feature vector into the encoder to generate a compressed fingerprint vector, which is used as the traffic fingerprint.
[0081] In an optional embodiment, the traffic statistical features extracted in step A2 may also include: packet arrival rate, traffic burst index and protocol distribution ratio, etc., to further improve the characterization capability of the features. Specifically, the packet arrival rate λ pkt Defined as the number of data packets arriving per unit time, through the formula λ pkt =N pkt / Δt calculation, where N pkt is the total number of data packets within the observation time window Δt. This feature can effectively identify abnormal traffic such as DDoS flood attacks, because such attacks are usually manifested as an abnormally high packet arrival rate in a short period of time. The traffic burst index B is obtained by calculating the coefficient of variation of the packet interval time, that is, B = σt / μt, where σt is the standard deviation of the packet interval time and μt is the average packet interval time. In experimental analysis, it was found that the B value of normal business traffic is usually in the range of 0.5-1.5, while the B value of malicious scanning and covert channel communication can reach more than 3.0, which is significantly distinguishable. Protocol distribution ratio R proto By counting the proportion of packets of different protocol types, we can form a vector R proto=[RTCP, RUDP, RICMP, ...]. Testing has shown that the enhanced statistical feature vector formed by combining these extended features with the original features improves the detection of low-intensity but persistent covert penetration attacks.
[0082] In another optional embodiment, the optimization of timing features in step A3 can also use methods such as Fourier transform or autoencoder to adapt to different network traffic characteristics and traceability requirements. When using the Fourier transform method, a discrete Fourier transform is applied to the original traffic data stream F. By selecting the main frequency components in the power spectrum as features, the periodic pattern of the traffic is captured, which is particularly suitable for detecting command and control (C&C) communications with obvious periodicity.
[0083] It should be noted that this fingerprint encoding method based on a lightweight Transformer model not only effectively compresses the dimension of the feature vector and reduces computational complexity, but also fully utilizes the model's self-attention mechanism to better capture the correlation between features, thereby generating a highly representative and anti-obfuscation fingerprint vector. Even in the face of interference methods such as encryption, tunneling, or camouflage, this fingerprint vector can still effectively identify the original traffic characteristics, providing a reliable basis for subsequent traffic tracing. Compared with traditional feature extraction and encoding methods, this method shows significant advantages in both tracing accuracy and anti-interference capabilities.
[0084] In the embodiment of the present application, the steps of tracing the attack path by using the first algorithm in step S200 include B1 to B4:
[0085] B1: inserts a mark into the data packets passing through the router with a preset probability;
[0086] Specifically, in B1, each time a packet passes through a router Ri, it inserts a marker Mi with probability p. Setting probability p to 0.1, we simulate scenarios with varying network sizes and traffic loads to study the relationship between marker coverage and network load.
[0087] The structure design of the tag Mi contains multiple key information, and its specific form is:
[0088] M i = <ID i ,t stamp ,H(ID i ||t stamp ||K priv )>;
[0089] Among them, the router ID iIt is a unique identifier for each router, used to distinguish different routers. In the network topology, each router has its specific position and function, which is identified by its ID. i The router node that the data packet passes through can be accurately identified. Timestamp t stamp Records the time when the data packet is marked at the router. Timestamp plays a key role in the path reconstruction process. It can help determine the order in which the data packet passes through each router, thereby restoring the data packet's transmission path. Hash value H(IDi||t stamp || Kpriv) is obtained by hashing the router ID, timestamp, and pre-shared private key Kpriv using the SHA-3 hash function. The pre-shared private key Kpriv is a private key pre-assigned to each router in the network. Only legitimate routers have this key. The hash value is used to ensure the integrity and authenticity of the tag information and prevent attackers from tampering with the tag content. If the attacker does not have the private key K priv , it is impossible to generate a legal hash value and thus it is impossible to forge a valid tag.
[0090] B2: When the data packet arrives at the receiving end, all the tags carried by the data packet are collected to form a tag set;
[0091] Specifically, in B2, when a data packet arrives at the receiving end (such as the target server, border gateway or security monitoring system), the receiving end will collect all the tags carried by the data packet to form a tag set {M1, M2, ..., M k In a real network environment, due to the use of a probabilistic labeling mechanism, a data packet may carry multiple labels from different routers, or may not carry any label at all. A data packet may pass through n routers during transmission, but since each router is labeled with a probability of p = 0.1, the expected number of labels obtained is approximately 0.1n. To improve the accuracy of path reconstruction, the receiver usually collects label information from multiple data packets that belong to the same attack flow or session. The receiver maintains a label buffer, continuously collects labels from newly arrived data packets, and triggers the path reconstruction process periodically (e.g., every 30 seconds) or when an attack behavior is detected.
[0092] B3: Filter out forged tags by verifying the legitimacy of the hash chain of each tag in the tag set;
[0093] Specifically, in B3, due to the complexity of the network environment, there may be cases where an attacker forges a tag. In order to ensure the accuracy of the reconstructed path, the receiver needs to verify the legitimacy of the tag. The specific method is to verify the legitimacy of the hash chain. For each tag M i , the receiver uses the known pre-shared private key K priv For the ID i and tstamp Perform the same SHA-3 hash operation, and then compare the calculated hash value with the hash value carried in the tag. The verification process can be expressed as:
[0094] H′=SHA-3(ID i ||t stamp ||K priv );
[0095] If the two match, the tag is considered legitimate; if they don't, the tag is considered forged and filtered out of the tag set. This effectively eliminates attacker-forged tags and improves the accuracy of path reconstruction. In practical applications, the SHA-3 hash algorithm is chosen for its high security, strong collision resistance, and fast computation speed, making it suitable for efficient implementation on network equipment.
[0096] B4: Sort the remaining legal tags in the tag set based on their timestamps, collect the routers that the data packets pass through in sequence, and reconstruct the attack path.
[0097] Specifically, in B4, after filtering out the forged tags, the receiver will use the timestamp t in the tag to stamp Sort the remaining legal tags. The timestamp reflects the order in which the data packet passes through each router. By sorting the tags in ascending order of timestamp, we can determine the routers that the data packet passes through. The sorted tag sequence can be expressed as {Msort1, Msort2, ..., Msort k}, the corresponding router sequence is {R1, R2, ..., R k Finally, by connecting these routers in order, we can reconstruct the attack path Path = [R1→R2→...→R k ].
[0098] In practice, due to the probabilistic nature of router markings, it may be impossible to obtain complete path information for a single packet. To address this issue, marking information is typically collected from multiple packets within the same attack flow. Through label aggregation analysis within a time window (e.g., 5 minutes), a more complete attack path is gradually constructed. Furthermore, network topology knowledge is leveraged for path verification and completion. When a jump is detected in the path (i.e., a missing intermediate router), the possible intermediate nodes can be inferred based on known network connectivity.
[0099] It should be noted that this improved probabilistic packet marking algorithm, through reasonable marking rules and effective path reconstruction methods, can accurately track network attack paths while ensuring network performance, providing strong support for network security protection. Furthermore, the use of pre-shared private keys and hash functions for tag verification significantly enhances the algorithm's security, effectively resisting forgery and tampering attempts by attackers.
[0100] In an optional embodiment, the mark in step B1 may also include router load information and link status indicators, so as to evaluate the impact of the attack on network performance while tracing the attack path, identify resource consumption hotspots caused by the attack, and assist network administrators in resource scheduling and defense strategy formulation.
[0101] In another optional implementation, steps B3-B4 can employ a distributed verification and collaborative path reconstruction approach, suitable for large-scale network environments. Multiple receiving nodes independently verify and sort the collected signatures, then report the results to a central coordinator, which integrates the path segments and globally optimizes them to generate a complete attack path.
[0102] In the embodiment of the present application, the steps of analyzing the similarity between different attack events by the second algorithm in step S300 include C1 to C4:
[0103] C1: Utilize the collected traffic data locally on the edge node to train a lightweight gated recurrent unit model, enabling it to learn the mapping relationship between traffic fingerprints and attack types.
[0104] Specifically, in step C1, each edge node (such as a financial institution's data center or a power grid's substation monitoring system) is responsible for local data processing and model training. Each edge node deploys a lightweight GRU (Gated Recurrent Unit) model, which has excellent time series data processing capabilities and can effectively learn the mapping relationship between traffic fingerprints and attack types. The core structure of the GRU model includes an update gate and a reset gate.
[0105] In actual implementation, the model adopts a two-layer GRU structure, with 64 hidden units in the first layer, 32 hidden units in the second layer, and finally a fully connected layer mapped to the attack type space.
[0106] C2: Send the local training model parameters of the edge node to the central server;
[0107] The central server performs weighted averaging based on the data volume proportion of each edge node and updates the global model parameters;
[0108] Among them, the global model parameter update formula is:
[0109]
[0110] Where |D| is the total amount of data, θ global is the global model parameter, is the global model parameter updated after the t+1th iteration, N is the total number of edge nodes participating in federated learning, |Di| is the amount of local data of node i, are the local model parameters obtained after local training of node i.
[0111] C3: Define the attack similarity function;
[0112] Attack similarity function Expressed as:
[0113]
[0114] Where, is the cosine similarity, which is used to measure the two traffic fingerprints and the degree of similarity between them; is the time decay factor, where t1 and t2 are the time when the two attacks occurred, respectively, and λ is the decay coefficient, which is used to control the influence of time on similarity.
[0115] C4: Set a similarity threshold. When the attack similarity calculated by the attack similarity function is greater than the similarity threshold, it is determined to be a cross-domain attack initiated by the same attacker.
[0116] In an optional implementation, the lightweight gated recurrent unit model in step C1 can be replaced with a temporal convolutional network (TCN-Attention) architecture based on an attention mechanism. This architecture combines the local feature extraction capabilities of a temporal convolutional network with the global dependency capture capabilities of an attention mechanism. Specifically, this architecture uses multiple dilated convolutional layers to extract multi-scale temporal features, followed by a self-attention layer to highlight key features, and finally maps them to the attack type space through a fully connected layer.
[0117] In the embodiment of the present application, after obtaining the attack source information in step S400, security protection measures are taken. The security protection measures include D1 to D4:
[0118] D1: Block the network connection of the attack source in real time;
[0119] Specifically, in step D1, the present invention first deploys a dynamic access control list (DACL) on network edge devices (such as firewalls and edge routers) based on the attack source's IP address, port, and protocol characteristics to block all inbound and outbound traffic from the attack source. For attack sources that hide their real IP addresses (such as those using proxies or botnets), the attack traffic characteristics (such as TTL values, TCP / IP stack fingerprints, traffic timing patterns, etc.) are analyzed to generate more refined feature matching rules for blocking.
[0120] To ensure the effectiveness of blocking measures, a layered defense strategy is implemented. Blocking is not only performed at the network perimeter, but also secondary blocking measures are deployed at key nodes within the internal network, forming a defense-in-depth approach. Blocking effectiveness is continuously monitored, with the effectiveness of blocking rules regularly evaluated (e.g., every 5 minutes). Blocking strategies are dynamically adjusted based on the attack source's adversarial behavior (e.g., changes in IP addresses or variations in attack signatures).
[0121] D2: Based on the characteristics of the attack source and the attack propagation trajectory, the security of the attacked network area is reinforced;
[0122] Specifically, in step D2, we first analyze the attack propagation trajectory and determine the affected network areas and key nodes. Security reinforcement measures include three levels:
[0123] Network layer reinforcement: Adjust the network segmentation configuration in the affected areas and implement stricter access control policies. For example, for DDoS attacks, increase traffic scrubbing capabilities; for lateral movement attacks, strengthen internal network isolation; for data exfiltration attacks, strengthen outbound traffic monitoring.
[0124] Host-level hardening: Deploy host protection measures at key nodes along the attack path. Based on attack signatures, the system generates host protection strategies, including updating antivirus rules, hardening system configurations, and applying the latest security patches.
[0125] Application-layer hardening: Implement application-level protection measures for the target application systems. For example, enhance input validation for web applications, adjust session management policies, and enable stricter authentication requirements.
[0126] After implementing reinforcement measures, ongoing verification testing is conducted to evaluate their effectiveness and ensure they remain effective without impacting business continuity. In a real-world case, a financial institution, after experiencing a targeted attack, successfully resisted three subsequent waves of attacks within 24 hours using this security reinforcement method, ensuring stable business operations.
[0127] D3: Store attack source information, attack propagation trajectory, and attack type characteristics into the security knowledge base;
[0128] Specifically, in step D3, the present invention structures the attack-related information and stores it in a security knowledge base to provide a reference for subsequent security analysis and protection. The stored content includes:
[0129] Attack source signature archive: This records basic information such as the attack source's IP address, geographic location, autonomous system number (ASN), domain name, activity time period, and known associated organizations, as well as advanced features such as TTP (tactics, techniques, and procedures) signatures, tool fingerprints, and attack patterns. Attack propagation trajectory records: This records in detail the attack path, the time each node was attacked, the duration of the attack, the propagation speed, and the lateral movement pattern.
[0130] This invention uses a graph database to store this information, supporting efficient correlation analysis and pattern mining. The knowledge base also implements an automatic update mechanism, automatically updating relevant records when new attack variants or defense bypass techniques are discovered. Furthermore, the knowledge base supports bidirectional conversion with threat intelligence sharing standards such as STIX and TAXII, facilitating data exchange with external threat intelligence platforms.
[0131] D4: Send warning information to other network nodes.
[0132] Specifically, in step D4, based on the attack source analysis results, structured warning information is generated and sent to potentially affected network nodes through multiple channels. The warning mechanism includes:
[0133] Alerts are categorized into four levels: "Urgent," "High," "Medium," and "Low," based on attack severity and confidence. These levels employ different notification methods and response requirements. Urgent alerts trigger automatic SMS and phone notifications, requiring the security team to respond within 15 minutes. Low-level alerts, on the other hand, are sent only via email or notification, allowing for a 24-hour response.
[0134] Rather than broadcasting all alerts to all nodes, this invention intelligently determines the alert's reach based on network topology, asset information, and attack signatures. For example, when an attack targeting a specific database vulnerability is detected, alerts are prioritized for nodes hosting that database, with secondary alert ranges determined based on network connectivity. This targeted alert mechanism avoids alert fatigue and improves response efficiency.
[0135] In an optional implementation, the network connection blocking in step D1 can also utilize an intent-based automated response mechanism. The security team predefines response intent for different attack types and automatically selects the optimal blocking strategy based on the current network status and business priorities. For example, when a suspicious but non-deterministic attack is detected during peak business hours, a partial blocking strategy is selected, limiting the bandwidth of the suspicious traffic rather than completely blocking it to avoid disrupting business operations. During non-business hours or when a high-confidence attack is detected, a more stringent complete blocking strategy is implemented.
[0136] In another optional implementation, the security knowledge base in step D3 can integrate a self-supervised learning mechanism to continuously optimize the extraction and classification of attack features. By regularly analyzing the differences between newly added attack samples and historical data, new attack variants and feature evolution trends can be automatically identified. For example, if traditional DDoS attack features are detected gradually blending into application layer patterns, new feature templates for hybrid DDoS attacks can be automatically generated and detection rules updated. This self-evolutionary capability ensures that the protection system maintains its ability to combat new attacks, reducing reliance on manual updates.
[0137] In summary, by extracting traffic statistical features (packet length distribution, time interval variance) and temporal features (bidirectional flow periodicity, entropy gradient), and applying wavelet transform technology to deeply explore features at different frequency components, this method comprehensively presents traffic characteristics. Furthermore, by utilizing lightweight Transformer model encoding and its attention mechanism, the feature vector dimensions are efficiently compressed, enhancing anti-obfuscation capabilities. Even in the face of interference from encryption, tunneling, or camouflage technologies, original traffic characteristics can be accurately identified, laying a solid foundation for source tracing analysis. The improved probabilistic packet marking algorithm addresses the shortcomings of traditional algorithms. The marking probability p is carefully set to balance marking efficiency and network overhead. The marking structure incorporates router identifiers, timestamps, and hash values based on pre-shared private keys to ensure label integrity and authenticity, resisting tampering. During path reconstruction, the receiver verifies the hash chain to filter out forged labels and reconstructs the attack path by sorting legitimate labels based on timestamps, significantly improving tracking accuracy and algorithm security. The multi-source correlation traceability model utilizes a federated learning architecture to address the data privacy risks of traditional centralized learning. Edge nodes use a lightweight GRU model to process local data and learn the mapping between traffic fingerprints and attack types. The central server aggregates parameters based on data volume and optimizes the global model. For tracing decisions, an attack similarity function is defined that comprehensively considers traffic fingerprint similarity and attack timing correlation. A threshold is set to determine cross-domain attacks, achieving precise tracing and adapting to complex network attack environments.
[0138] In Example 3, the above is a schematic scheme of a method for tracing the source of traffic data of a hacker threat. It should be noted that the technical scheme of the system for tracing the source of traffic data of a hacker threat and the technical scheme of the method for tracing the source of traffic data of a hacker threat are based on the same concept. For details not described in detail in the technical scheme of the system for tracing the source of traffic data of a hacker threat in this embodiment, please refer to the description of the technical scheme of the method for tracing the source of traffic data of a hacker threat.
[0139] This embodiment also provides a traffic data tracing system for hacker threats, including a fingerprint construction module, an attack path tracing module, an attack similarity analysis module, and an attack source determination module;
[0140] The fingerprint construction module is responsible for obtaining the original traffic data stream, extracting features from the obtained original traffic data stream, and constructing the traffic fingerprint;
[0141] The attack path tracing module uses the first algorithm to track the attack path based on the constructed traffic fingerprint and determines the attack propagation trajectory;
[0142] The attack similarity analysis module uses the second algorithm to analyze the similarities between different attack events based on the attack propagation trajectory and determine cross-domain attack behavior;
[0143] The attack source determination module obtains the attack source based on the determination results of cross-domain attack behavior and the attack propagation trajectory.
[0144] This embodiment also provides an electronic device suitable for tracing the source of traffic data threatening hacker threats, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute computer-executable instructions to implement the method for tracing the source of traffic data threatening hacker threats proposed in the above embodiment.
[0145] This embodiment also provides a storage medium on which a computer program is stored. When the program is executed by a processor, the method for tracing the traffic data source of hacker threats proposed in the above embodiment is implemented.
[0146] The storage medium proposed in this embodiment and the method for tracing the traffic data of hacker threats proposed in the above embodiment belong to the same inventive concept. For technical details not described in detail in this embodiment, please refer to the above embodiment, and this embodiment has the same beneficial effects as the above embodiment.
[0147] Through the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented with the help of software and necessary general hardware, and of course can also be implemented by hardware. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory (FLASH), hard disk or optical disk, etc., including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods of various embodiments of the present invention.
[0148] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A method for tracing traffic data of hacker threats, characterized by: include: Obtaining an original traffic data stream, extracting features of the original traffic data stream and constructing a traffic fingerprint; Based on the traffic fingerprint, tracking the attack path using a first algorithm, and determining the attack propagation trajectory through the attack path; According to the attack propagation trajectory, a second algorithm is used to analyze the similarity between different attack events to determine cross-domain attack behavior; Based on the cross-domain attack behavior determination result and the attack propagation trajectory, the attack source is obtained.
2. The method for tracing the traffic data source of hacker threats according to claim 1, characterized in that: The steps of obtaining the traffic fingerprint include: Collect raw traffic data streams; Extracting traffic statistical features and time series features from the original traffic data stream; Optimizing the timing characteristics; Splicing the traffic statistical features and the optimized time series features to form a fused feature vector; The fused feature vector is input into an encoder to generate a compressed fingerprint vector, which is used as the traffic fingerprint.
3. The method for tracing the traffic data source of hacker threats according to claim 2, characterized in that: The steps of tracing the attack path using the first algorithm include: Inserting tags into data packets passing through the router with a preset probability; When the data packet arrives at the receiving end, all tags carried by the data packet are collected to form a tag set; Filtering out forged tags by verifying the legitimacy of the hash chain of each tag in the tag set; The remaining legal tags in the tag set are sorted based on their timestamps, and the routers that the data packet passes through in sequence are collected to reconstruct the attack path.
4. The method for tracing traffic data of hacker threats according to claim 3, characterized in that: The second algorithm step of analyzing the similarity between different attack events includes: Using the collected traffic data locally at the edge node, a lightweight gated recurrent unit model is trained, so that the lightweight gated recurrent unit model learns the mapping relationship between the traffic fingerprint and the attack type; Send the local training model parameters of the edge node to the central server; Define attack similarity function; A similarity threshold is set. When the attack similarity calculated by the attack similarity function is greater than the similarity threshold, it is determined to be a cross-domain attack initiated by the same attacker.
5. The method for tracing the flow data source of hacker threats according to claim 4, characterized in that: After obtaining the attack source information, security protection measures are taken, the security protection measures including: Block the network connection of the attack source in real time; Based on the characteristics of the attack source and the attack propagation trajectory, security reinforcement is performed on the attacked network area; Storing the attack source information, the attack propagation trajectory and the attack type characteristics in a security knowledge base; Send warning information to other network nodes.
6. The method for tracing the flow data source of hacker threats according to claim 5, characterized in that: The attack similarity function Expressed as: Where, is the cosine similarity, which is used to measure the two traffic fingerprints and the degree of similarity between them; is the time decay factor, where t1 and t2 are the time when the two attacks occurred, respectively, and λ is the decay coefficient, which is used to control the influence of time on similarity.
7. The method for tracing the traffic data source of hacker threats according to claim 6, characterized in that: The central server performs weighted averaging based on the data volume proportion of each edge node and updates the global model parameters; Among them, the global model parameter update formula is: Where |D| is the total amount of data, θ global is the global model parameter, is the global model parameter updated after the t+1th iteration, N is the total number of edge nodes participating in federated learning, |Di| is the amount of local data of node i, are the local model parameters obtained after local training of node i.
8. A traffic data tracing system for hacker threats, applying the method according to any one of claims 1 to 7, characterized in that: It includes fingerprint construction module, attack path tracing module, attack similarity analysis module, and attack source determination module; The fingerprint building module is responsible for acquiring the original traffic data stream, extracting features from the acquired original traffic data stream, and building a traffic fingerprint; The attack path tracing module uses a first algorithm to trace the attack path based on the constructed traffic fingerprint and determines the attack propagation trajectory; The attack similarity analysis module uses a second algorithm to analyze the similarities between different attack events based on the attack propagation trajectory and determines cross-domain attack behavior; The attack source determination module obtains the attack source based on the determination result of the cross-domain attack behavior and the attack propagation trajectory.
9. An electronic device comprising: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the traffic data tracing method for hacker threats described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method for tracing traffic data of hacker threats as described in any one of claims 1 to 7.
Citation Information
Cited By
Data security and privacy protection method for distribution automation system
CN121441649A