Network attack tracing method and device in attack and defense confrontation scene

By deploying honeypots in offensive and defensive confrontation scenarios and analyzing attack behavior data using dynamic clustering algorithms, the problems of slow response and inaccurate positioning of traditional traceability methods are solved, and fast and accurate attacker traceability and path positioning are achieved, and detailed traceability reports are generated.

CN120238343APending Publication Date: 2025-07-01STATE GRID CORP NORTHEAST DIVISION +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510369584.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Traditional cyber attack traceability methods are insufficient in offensive and defensive confrontation scenarios, and cannot provide immediate analysis and feedback, and cannot effectively support fast response and precise positioning of attackers.

Method used

Deploy honeypots in the offensive and defensive confrontation scenario to obtain attack behavior data, build a feature vector set by analyzing the time distribution characteristics, access frequency characteristics and field interaction characteristics, and use the clustering algorithm of the dynamic adjustment mechanism to classify, generate feature groups, and locate the attacker's source and path.

Benefits of technology

It significantly improves the real-time and accuracy of cyber attack tracing, can quickly and accurately locate the attacker's source and path, and generates detailed traceability reports to support network security protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238343A_ABST
    Figure CN120238343A_ABST
Patent Text Reader

Abstract

The invention discloses a network attack tracing method and device in an attack and defense confrontation scene, and the method comprises the steps: deploying a honeypot in a target network according to a potential attack mode in the attack and defense confrontation scene, so as to simulate a plurality of attacked targets, obtaining attack behavior data according to an attacker access record, the attack behavior data comprises a source IP address, a protocol field and a behavior sequence of communication flow; analyzing the attack behavior data, extracting features, and constructing a feature vector set; classifying the feature vector set by a clustering algorithm based on a dynamic adjustment mechanism, dynamically initializing the position of a clustering center by calculating the mean value of time distribution features, the highest frequency of access frequency features and the correlation of field interaction features, calculating the similarity from each vector in the feature vector set to the clustering center, generating feature groups, and performing clustering on the feature groups; extracting a central feature of each feature group; and positioning an attacker source, determining an attack behavior path and generating a traceability report according to the center features of the feature groups and the attack behavior data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical fields of network security and network management, and more specifically, relates to a network attack traceability method and device in an attack and defense confrontation scenario. Background Art

[0002] Traditional traceability technologies, including log analysis and intrusion detection systems, usually adopt a post - analysis method to trace the source of network attacks. Traditional network attack traceability methods have obvious limitations in actual operation, especially in a rapidly changing attack and defense confrontation environment. Since they rely on the analysis of events that have already occurred, it is difficult to meet the requirements of the modern network security field for traceability work in terms of response speed and accuracy. In such a highly adversarial scenario, attackers often use complex technical means to launch attacks quickly, while traditional traceability methods cannot provide immediate analysis and feedback, thus unable to effectively support the needs of rapid response and precise positioning of attackers. Summary of the Invention

[0003] A honeypot - based traceability method designed to solve the false positives of traditional logs and intrusion detection methods can accurately provide the attacker's IP, highlighting its superiority in the attack and defense confrontation scenario. The present invention provides a network attack traceability method and device in an attack and defense confrontation scenario.

[0004] The present invention adopts the following technical solutions.

[0005] The first aspect of the present invention provides a network attack traceability method in an attack and defense confrontation scenario, including the following steps:

[0006] Deploy a honeypot in the attack and defense confrontation scenario to simulate various attacked targets and obtain attack behavior data, where the attack behavior data includes an attacker - side data set and an attacked - side data set;

[0007] Analyze the time - distribution characteristics and access - frequency characteristics of the attacker - side data set, as well as the field - interaction characteristics of the attacked - side data set, and construct a feature vector set;

[0008] Classify the feature vector set using a clustering algorithm based on a dynamic adjustment mechanism, dynamically initialize the position of the clustering center by calculating the mean of the time - distribution characteristics, the highest frequency of the access - frequency characteristics, and the correlation of the field - interaction characteristics, and calculate the similarity of each vector in the feature vector set to the clustering center to generate feature groups, and extract the central features of each feature group;

[0009] According to the central features of the feature groups, combined with the attack behavior data, locate the attacker's source, determine the attack behavior path, and generate a traceability report.

[0010] Preferably, deploy honeypots in the attack and defense confrontation scenario to simulate various attacked targets, including:

[0011] Select honeypot nodes in the attack and defense confrontation scenario, build a honeypot environment using virtual servers, and allocate independent computing resources;

[0012] Pre-configure different operating environments in the honeypot environment to simulate different types of attacked targets; configure network protocols and set communication protocol parameters to simulate network interaction behaviors;

[0013] Configure the open ports of the honeypot environment, including database services and remote login services, to provide different attack targets;

[0014] Collect and monitor attack behavior data. The attack-side dataset includes source IP addresses, time distribution characteristics, access frequency characteristics, communication traffic characteristics, and protocol fields; the attacked-side dataset includes target IP addresses, service ports, field interaction characteristics, and protocol fields.

[0015] Preferably, analyze the time distribution characteristics and access frequency characteristics of the attack-side dataset, and the field interaction characteristics of the attacked-side dataset, and construct a feature vector set; including:

[0016] Calculate the time interval between adjacent attack behaviors in the attack-side dataset, segment the time series using a sliding window method, and perform numerical mapping on the time interval data based on adaptive scale transformation to obtain time distribution characteristics;

[0017] Count the number of accesses from the same source IP in the attack-side dataset within a set time window, and divide the access frequency into intervals using a hierarchical bucketing method to obtain access frequency characteristics;

[0018] Analyze the communication protocol fields of the attacked-side dataset, construct a co-occurrence matrix, and calculate the semantic distance between fields as follows to obtain field interaction characteristics:

[0019]

[0020] In the formula, D(f i , f j ) represents the semantic distance between field f i and field f j ; f i and f j represent two different fields in the communication protocol; f k represents an intermediate field affecting the relationship between f i and f j ; P(f k |f i ) represents f k and f iThe probability of simultaneous occurrence; P(f k |f j ) represents the probability of simultaneous occurrence of f k and f j ;

[0021] The normalized time distribution feature, access frequency feature, and field interaction feature are combined to generate a feature vector set.

[0022] Preferably, the dynamic initialization of the clustering center position by calculating the mean of the time distribution feature, the highest frequency of the access frequency feature, and the correlation of the field interaction feature includes:

[0023] Based on the highest frequency of the access frequency feature, outliers in the feature vector set are screened, the distribution density of the access frequency feature is calculated, and a threshold F th is set to remove outliers beyond the set range:

[0024] F th = μ F + λσ F

[0025] In the formula, F th represents the set threshold; μ F represents the mean of the access frequency feature; λ is a constant coefficient with a value range greater than 1; σ F represents the degree of dispersion of the eigenvalue;

[0026] In the screened feature vector set, the mean μ T of the time distribution feature is calculated, and samples with time intervals falling within the set interval [μ T - σ T , μ T + σ T are used as candidate center points; where μ T represents the mean of the time distribution feature; σ T is the standard deviation of the time distribution feature, representing the fluctuation range of the time feature;

[0027] Based on the correlation of the field interaction feature, the center point distribution is optimized, the similarity between the field interaction feature vectors of the candidate center points is calculated, and the center point selection is optimized based on the spectral clustering method:

[0028]

[0029] In the formula, S i,j represents the similarity between the field interaction feature vector of the i-th candidate center point and the field interaction feature vector of the j-th candidate center point; represents the field interaction feature vector of the i-th candidate center point; represents the field interaction feature vector of the j-th candidate center point; The norm of the field interaction feature vector representing the i-th candidate center point; The norm of the field interaction feature vector representing the j-th candidate center point;

[0030] Determine the final clustering center, calculate the density value of each candidate center point, and select N c points whose density values meet the set criteria as the final clustering centers to complete the dynamic initialization:

[0031]

[0032] In the formula, D i represents the density value of the i-th candidate center point; σ represents the scale controlling the similarity square term; represents the weight of the similarity between two candidate center points.

[0033] Preferably, the highest access frequency for calculating the distribution density of the access frequency feature is calculated by the following formula:

[0034]

[0035] In the formula, P(F) represents the probability density of the access frequency; K(x) represents the kernel function; h represents the smoothing parameter; N represents the total number of attack behavior samples; F represents a specific value of the access frequency feature; F i represents the access frequency feature of the i-th sample.

[0036] Preferably, the similarity between each vector in the feature vector set and the clustering center is calculated by the following formula:

[0037]

[0038] In the formula, W i,j represents the similarity between the feature vector V i and the clustering center C j ; d i,j represents the distance metric between the feature vector V i and the clustering center C j ; α and β represent the weighting coefficients.

[0039] Preferably, the generating feature groups and extracting the central features of each feature group include:

[0040] According to the principle of maximum similarity, attribute the feature vectors to the most similar clustering centers to form feature groups:

[0041]

[0042] In the formula, represents the feature vector V iThe final clustering center; arg max j W i,j represents the clustering center with the maximum similarity; C j represents the feature vector V i The most similar clustering center; G j represents the feature vector V i The feature grouping composed of...

[0043] Preferably, the central feature of the feature grouping is calculated as follows, and the center point of each feature grouping is calculated based on the mean vector:

[0044]

[0045] In the formula, M j represents the central feature of the feature grouping G j ...

[0046] Preferably, based on the central feature of the feature grouping, combined with the attack behavior data, the source of the attacker is located, the attack behavior path is determined, and a traceability report is generated; including:

[0047] Match the attack pattern based on the central feature of the feature grouping, and calculate the similarity between the central feature and the attack behavior data:

[0048]

[0049] In the formula, S(M j , A m ) represents the similarity between the central feature M j and the attack behavior data sample; A m represents the attack behavior data sample; ||M j || represents the norm of the central feature M j ; ‖A m ‖ represents the norm of the attack behavior data sample;

[0050] Construct an attack behavior path, and based on the timestamp T(A m ) and the target IP address IP(A m ) of the attack data, associate the attack events in chronological order:

[0051] Path = {A1 → A2 → … → A n}

[0052] In the formula, Path represents the constructed attack behavior path; A1 → A2 → … → A n represents the 1st to the nth attack behavior samples in the attack path, arranged in the order of the time of attack occurrence;

[0053] Extract the source IP address in the attack path, and combine it with the historical attack database to calculate the probability of the attacker's appearance on the path to identify the attacker's source:

[0054]

[0055] In the formula, Source represents the source IP address of the attacker; arg max IP P(IP∣Path) represents selecting the IP address with the highest probability on the attack path; P(IP∣Path) represents the probability of a certain IP address being the attacker under the condition of a given attack path;

[0056] Record the attack time, attack path, and attacker source to form structured traceability data and output a traceability report.

[0057] The second aspect of the present invention provides a network attack traceability device in an offensive and defensive confrontation scenario, including: a honeypot deployment module, a feature parsing module, a clustering analysis module, and a traceability report generation module;

[0058] The honeypot deployment module is used to deploy honeypots in the offensive and defensive confrontation scenario to simulate various attacked targets and obtain attack behavior data, where the attack behavior data includes an attack-side data set and an attacked-side data set;

[0059] The feature parsing module is used to parse the time distribution feature and access frequency feature of the attack-side data set, as well as the field interaction feature of the attacked-side data set, and construct a feature vector set based on the parsing results;

[0060] The clustering analysis module is used to classify the feature vector set by a clustering algorithm based on a dynamic adjustment mechanism, dynamically initialize the clustering center position by calculating the mean of the time distribution feature, the highest frequency of the access frequency feature, and the correlation of the field interaction feature, and calculate the similarity of each vector in the feature vector set to the clustering center, generate feature groups, and extract the central features of each feature group;

[0061] The traceability report generation module is used to locate the attacker's source according to the central features of the feature groups, determine the attack behavior path, and generate a traceability report in combination with the attack behavior data.

[0062] The third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the network attack traceability method in the offensive and defensive confrontation scenario.

[0063] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the network attack tracing method in the attack and defense confrontation scenario as described above.

[0064] Compared with the prior art, the beneficial effects of the present invention at least include: avoiding the deviation of the characteristics in a certain dimension from the overall effect in the traditional technology; by combining the honeypot technology and the dynamic clustering analysis algorithm, the present invention significantly improves the real-time performance and accuracy of network attack tracing; by obtaining attack behavior data in real time and dynamically adjusting the position of the clustering center, the source and attack path of the attacker can be quickly and accurately located, solving the problems of slow response and inaccurate positioning in the traditional tracing method; the generated tracing report provides a detailed attack analysis for network security protection, effectively supporting subsequent defense and countermeasures. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] Figure 1 is a flowchart of the network attack tracing method in the attack and defense confrontation scenario provided according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0066] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. The embodiments described in this application are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0067] As Figure 1 shown, Example 1 of the present invention provides a network attack tracing method in the attack and defense confrontation scenario, including the following steps:

[0068] Deploy a honeypot in the attack and defense confrontation scenario to simulate various attacked targets and obtain attack behavior data, where the attack behavior data includes an attack end data set and an attacked end data set;

[0069] Preferably, the step of deploying a honeypot in the attack and defense confrontation scenario to simulate various attacked targets includes:

[0070] Select a honeypot node in the attack and defense confrontation scenario, build a honeypot environment using a virtual server, and allocate independent computing resources;

[0071] Pre-configure different operating environments in the honeypot environment to simulate different types of attacked targets; configure network protocols and set communication protocol parameters to simulate network interaction behaviors;

[0072] Configure the open ports of the honeypot environment, including database services and remote login services, to provide different attack targets;

[0073] Collect and monitor attack behavior data. The attacker-side dataset includes source IP addresses, time distribution characteristics, access frequency characteristics, communication traffic characteristics, and protocol fields; the attacked-side dataset includes target IP addresses, service ports, field interaction characteristics, and protocol fields.

[0074] Analyze the time distribution characteristics and access frequency characteristics of the attacker-side dataset, as well as the field interaction characteristics of the attacked-side dataset, and construct a feature vector set;

[0075] Preferably, the analyzing the time distribution characteristics and access frequency characteristics of the attacker-side dataset, as well as the field interaction characteristics of the attacked-side dataset, and constructing a feature vector set includes:

[0076] Calculate the time interval between adjacent attack behaviors in the attacker-side dataset, segment the time series using a sliding window method, and perform numerical mapping on the time interval data based on adaptive scale transformation to obtain time distribution characteristics;

[0077] Count the number of accesses of the same source IP in the attacker-side dataset within a set time window, divide the access frequency into intervals using a hierarchical bucketing method, and obtain access frequency characteristics;

[0078] Analyze the communication protocol fields of the attacked-side dataset, construct a co-occurrence matrix, and calculate the semantic distance between fields as follows to obtain field interaction characteristics:

[0079]

[0080] In the formula, D(f i ,f j ) represents the semantic distance between field f i and field f j ; f i and f j represent two different fields in the communication protocol; f k represents an intermediate field affecting the relationship between f i and f j ; P(f k |f i ) represents the probability that f k and f i appear simultaneously; P(f k |f j ) represents the probability that f k and f j appear simultaneously;

[0081] The normalized time distribution features, access frequency features, and field interaction features are combined to generate a feature vector set.

[0082] The feature vector set is classified by a clustering algorithm based on a dynamic adjustment mechanism. The clustering center position is dynamically initialized by calculating the mean of the time distribution features, the highest frequency of the access frequency features, and the correlation of the field interaction features. The similarity between each vector in the feature vector set and the clustering center is calculated to generate feature groups, and the central features of each feature group are extracted.

[0083] Preferably, the dynamic initialization of the clustering center position by calculating the mean of the time distribution features, the highest frequency of the access frequency features, and the correlation of the field interaction features includes:

[0084] Outliers in the feature vector set are screened based on the highest frequency of the access frequency features, the distribution density of the access frequency features is calculated, and a threshold F th is set to remove outliers outside the set range:

[0085] F th = μ F + λσ F

[0086] In the formula, F th represents the set threshold; μ F represents the mean of the access frequency features; λ is a constant coefficient with a value range greater than 1; σ F represents the degree of dispersion of the eigenvalues;

[0087] In the screened feature vector set, the mean μ T of the time distribution features is calculated, and samples with time intervals falling within the set interval [μ T - σ T , μ T + σ T are used as candidate center points; where μ T represents the mean of the time distribution features; σ T is the standard deviation of the time distribution features, representing the fluctuation range of the time features;

[0088] The distribution of the center points is optimized based on the correlation of the field interaction features. The similarity between the field interaction feature vectors of the candidate center points is calculated, and the selection of the center points is optimized based on the spectral clustering method:

[0089]

[0090] In the formula, S i,j represents the similarity between the field interaction feature vector of the i-th candidate center point and the field interaction feature vector of the j-th candidate center point; The field interaction feature vector representing the i-th candidate center point; The field interaction feature vector representing the j-th candidate center point; The norm length of the field interaction feature vector representing the i-th candidate center point; The norm length of the field interaction feature vector representing the j-th candidate center point;

[0091] Determine the final clustering centers, calculate the density values of each candidate center point, and select N c points whose density values meet the set criteria as the final clustering centers to complete the dynamic initialization:

[0092]

[0093] In the formula, D i represents the density value of the i-th candidate center point; σ represents the scale for controlling the similarity square term; represents the weight of the similarity between two candidate center points.

[0094] Preferably, the highest access frequency for calculating the distribution density of the access frequency feature is calculated by the following formula:

[0095]

[0096] In the formula, P(F) represents the probability density of the access frequency; K(x) represents the kernel function; h represents the smoothing parameter; N represents the total number of attack behavior samples; F represents a specific value of the access frequency feature; F i represents the access frequency feature of the i-th sample.

[0097] Preferably, the similarity between each vector in the feature vector set and the clustering center is calculated by the following formula:

[0098]

[0099] In the formula, W i,j represents the similarity between the feature vector V i and the clustering center C j ; d i,j represents the distance metric between the feature vector V i and the clustering center C j ; α and β represent the weighting coefficients.

[0100] Preferably, the generating feature groups and extracting the central features of each feature group include:

[0101] According to the principle of maximum similarity, attribute the feature vectors to the most similar clustering centers to form feature groups:

[0102]

[0103] In the formula, represents the final clustering center of the feature vector V i arg max j W i,j represents the clustering center with the maximum similarity; C j represents the clustering center most similar to the feature vector V i ; G j represents the feature grouping composed of the feature vectors V i .

[0104] Preferably, the central feature of the feature grouping is calculated by the following formula, and the center point of each feature grouping is calculated based on the mean vector:

[0105]

[0106] In the formula, M j represents the central feature of the feature grouping G j .

[0107] Preferably, the original data set is extracted as (x1, x2,..., x n ), and each x i is a d-dimensional vector. The purpose of K-means clustering is to divide the original data into k clusters under the condition of a given number of classification groups k (k ≤ n); it includes:

[0108] Randomly select k vectors from the original data set as the initial cluster centers;

[0109] Calculate the similarities of the remaining elements to the k cluster centers respectively, and assign these elements to the clusters with the highest similarities;

[0110] The similarity calculation uses the adjusted cosine algorithm to calculate the cosine similarity;

[0111] For example, there is a two-dimensional vector set, and for two vectors a(x1, y1) and b(x2, y2) on the two-dimensional plane:

[0112] In the original data set, there is a mean value

[0113] in the y dimension. Similarly, there is also such an x c ;

[0114] X1 = x1 - x c ; Y1 = y1 - y c ;

[0115] The adjusted vectors are a,(X1, Y1) and b,(X2, Y2);

[0116] Then, use the cosine similarity algorithm for calculation:

[0117]

[0118] According to the clustering results, recalculate the centers of the k clusters respectively. The calculation method is to take the root mean square average of each dimension of all elements in the cluster.

[0119] Re - cluster all elements in the original dataset according to the new centers.

[0120] Repeat the previous step until the clustering results no longer change.

[0121] Output the results; obtain the center points of each distance - based cluster; complete the classification.

[0122] According to the central features of the feature groups, combined with the attack behavior data, locate the source of the attacker, determine the attack behavior path, and generate a traceability report.

[0123] Preferably, the step of according to the central features of the feature groups, combined with the attack behavior data, locating the source of the attacker, determining the attack behavior path, and generating a traceability report includes:

[0124] Match the attack pattern based on the central features of the feature groups, and calculate the similarity between the central features and the attack behavior data:

[0125]

[0126] In the formula, S(M j , A m ) represents the similarity between the central feature M j and the attack behavior data sample; A m represents the attack behavior data sample; ||M j || represents the norm of the central feature M j ; ‖A m ‖ represents the norm of the attack behavior data sample;

[0127] Construct the attack behavior path. Based on the timestamp T(A m ) and the target IP address IP(A m ) of the attack data, associate the attack events in time series:

[0128] Path = {A1 → A2 → … → A n}

[0129] In the formula, Path represents the constructed attack behavior path; A1 → A2 → … → A n represents the 1st to the nth attack behavior samples in the attack path respectively, arranged in the order of the time of attack occurrence;

[0130] Extract the source IP address in the attack path, and combine it with the historical attack database to calculate the probability of the attacker's appearance on the path to identify the attacker's source:

[0131]

[0132] In the formula, Source represents the source IP address of the attacker; arg max IP P(IP∣Path) represents selecting the IP address with the highest probability on the attack path; P(IP∣Path) represents the probability of a certain IP address being the attacker under the condition of a given attack path;

[0133] Record the attack time, attack path, and attacker source to form structured traceability data and output a traceability report.

[0134] Example 2 of the present invention provides a network attack traceability device in an attack and defense confrontation scenario, including: a honeypot deployment module, a feature parsing module, a clustering analysis module, and a traceability report generation module;

[0135] The honeypot deployment module is used to deploy honeypots in the attack and defense confrontation scenario to simulate various attacked targets and obtain attack behavior data. The attack behavior data includes an attack-side data set and an attacked-side data set;

[0136] The feature parsing module is used to parse the time distribution feature and access frequency feature of the attack-side data set, as well as the field interaction feature of the attacked-side data set, and construct a feature vector set based on the parsing results;

[0137] The clustering analysis module is used to classify the feature vector set by a clustering algorithm based on a dynamic adjustment mechanism, dynamically initialize the clustering center position by calculating the mean of the time distribution feature, the highest frequency of the access frequency feature, and the correlation of the field interaction feature, and calculate the similarity of each vector in the feature vector set to the clustering center, generate feature groups, and extract the central features of each feature group;

[0138] The traceability report generation module is used to locate the attacker's source according to the central features of the feature groups, determine the attack behavior path, and generate a traceability report in combination with the attack behavior data.

[0139] Example 3 of the present invention provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the network attack traceability method in the attack and defense confrontation scenario.

[0140] Example 4 of the present invention provides a computer-readable storage medium storing a computer program which, when executed by a processor, implements the network attack tracing method in the offensive and defensive confrontation scenario.

[0141] The beneficial effects of the present invention at least include: avoiding the deviation of the characteristics in a certain dimension from the overall effect in the traditional technology; by combining the honeypot technology and the dynamic clustering analysis algorithm, the present invention significantly improves the real-time performance and accuracy of network attack tracing; by obtaining attack behavior data in real time and dynamically adjusting the position of the clustering center, the source of the attacker and the attack path can be quickly and accurately located, solving the problems of slow response and inaccurate positioning of the traditional tracing method; the generated tracing report provides a detailed attack analysis for network security protection, effectively supporting subsequent defense and countermeasures.

[0142] The present disclosure may be a system, a method or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0143] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the claims of the present invention.

Claims

1. A network attack source tracing method in an attack-defense confrontation scenario, characterized in that: The following steps are involved: Deploy honeypots in attack-defense confrontation scenarios to simulate multiple attacked targets and obtain attack behavior data, which includes attack end data sets and attacked end data sets; Analyze the time distribution characteristics and access frequency characteristics of the attacking data set, as well as the field interaction characteristics of the attacked data set, and construct a feature vector set; The clustering algorithm based on the dynamic adjustment mechanism classifies the feature vector set, dynamically initializes the cluster center position by calculating the mean of the time distribution feature, the highest frequency of the access frequency feature and the correlation of the field interaction feature, and calculates the similarity between each vector in the feature vector set and the cluster center, generates feature groups, and extracts the central features of each feature group; Based on the central features of the feature grouping and combined with the attack behavior data, the attacker's source is located, the attack behavior path is determined, and a tracing report is generated.

2. According to a network attack tracing method in an attack and defense confrontation scenario according to claim 1, it is characterized by: The honeypot is deployed in the attack and defense confrontation scenario to simulate multiple attack targets, including: Select honeypot nodes in the attack and defense confrontation scenario, use virtual servers to build a honeypot environment, and allocate independent computing resources; Pre-configure different operating environments in the honeypot environment to simulate different types of attack targets; configure network protocols and set communication protocol parameters to simulate network interaction behaviors; Configure the open ports of the honeypot environment, including database services and remote login services, to provide different attack targets; Collect and monitor attack behavior data. The attacking data set includes source IP address, time distribution characteristics, access frequency characteristics, communication traffic characteristics, and protocol fields; the attacked data set includes target IP address, service port, field interaction characteristics, and protocol fields.

3. The method for tracing the source of a network attack in an attack-defense confrontation scenario according to claim 1 is characterized in that: The method of analyzing the time distribution characteristics and access frequency characteristics of the attacking end data set and the field interaction characteristics of the attacked end data set to construct a feature vector set includes: Calculate the time intervals between adjacent attack behaviors in the attack end data set, segment the time series using the sliding window method, and perform numerical mapping on the time interval data based on adaptive scaling to obtain time distribution characteristics; Count the number of visits from the same source IP in the attack end data set within the set time window, use the layered bucketing method to divide the access frequency into intervals, and obtain the access frequency characteristics; Parse the communication protocol fields of the attacked dataset, construct a co-occurrence matrix, and calculate the semantic distance between fields to obtain field interaction features using the following formula: In the formula, D(f i ,f j ) represents the field f i and field f j The semantic distance between i and f j Indicates two different fields in the communication protocol; f k Indicates the impact f i and f j The middle field of the relation; P(f k |f i ) indicates f k and f i The probability of simultaneous occurrence; P(f k |f j ) indicates f k and f j Probability of simultaneous occurrence; Normalize the time distribution features, access frequency features, and field interaction features, and combine them to generate a feature vector set.

4. The method for tracing the source of a network attack in an attack-defense confrontation scenario according to claim 1 is characterized in that: The method of dynamically initializing the cluster center position by calculating the mean of the time distribution feature, the highest frequency of the access frequency feature, and the correlation of the field interaction feature includes: Based on the highest frequency of access frequency features, filter outliers in the feature vector set, calculate the distribution density of access frequency features, and set the threshold F th To remove outliers that exceed the set range: F th =μ F +λσ F In the formula, F th Indicates the set threshold; μ F represents the mean of the access frequency feature; λ is a constant coefficient with a value range greater than 1; σ F Indicates the discreteness of the eigenvalue; In the filtered feature vector set, calculate the mean μ of the time distribution feature T , the screening time interval falls into the set interval [μ T -σ T ,μ T +σ T ] as candidate center points; where μ T Represents the mean of the time distribution characteristics; σ T is the standard deviation of the time distribution characteristics, indicating the fluctuation range of the time characteristics; The distribution of center points is optimized based on the correlation of field interaction features, the similarity between the field interaction feature vectors of candidate center points is calculated, and the center point selection is optimized based on spectral clustering: In the formula, S i,j Represents the similarity between the field interaction feature vector of the i-th candidate center point and the field interaction feature vector of the j-th candidate center point; The field interaction feature vector representing the i-th candidate center point; The field interaction feature vector representing the jth candidate center point; The modulus of the field interaction feature vector representing the i-th candidate center point; The modulus of the field interaction feature vector representing the jth candidate center point; Determine the final cluster center, calculate the density value of each candidate center point, and select N clusters whose density values ​​meet the set standards. c points as the final cluster centers to complete dynamic initialization: Where D i represents the density value of the i-th candidate center point; σ represents the scale of the control similarity square term; Represents the weight of the similarity between two candidate center points.

5. The method for tracing the source of a network attack in an attack-defense confrontation scenario according to claim 4 is characterized in that: The highest access frequency of the distribution density of the access frequency feature is calculated as follows: Where P(F) represents the probability density of access frequency; K(x) represents the kernel function; h represents the smoothing parameter; N represents the total number of attack behavior samples; F represents a specific value of the access frequency feature; F i Represents the access frequency characteristics of the i-th sample.

6. A network attack source tracing method in an attack-defense confrontation scenario according to claim 1 or 4, characterized in that: The similarity of each vector in the feature vector set to the cluster center is calculated as follows: Where W i,j Denotes the feature vector V i With cluster center C j Similarity; d i,j Denotes the feature vector V i With cluster center C j The distance metric of ; α and β represent weighting coefficients.

7. A network attack source tracing method in an attack-defense confrontation scenario according to claim 1 or 6, characterized in that: The generating of feature groups and extracting the central features of each feature group includes: According to the maximum similarity principle, the feature vectors are assigned to the most similar cluster centers to form feature groups: In the formula, Denotes the feature vector V i The final cluster center of j W i,j Represents the cluster center with the maximum similarity; C j Denotes the feature vector V i The most similar cluster center; G j Denotes the feature vector V i The features grouped together.

8. The method for tracing the source of a network attack in an attack-defense confrontation scenario according to claim 7 is characterized in that: The central feature of the feature grouping is calculated as follows, and the center point of each feature grouping is calculated based on the mean vector: Where M j Represents the feature group G j central feature.

9. A network attack source tracing method in an attack-defense confrontation scenario according to claim 1 or 8, characterized in that: The central feature of the feature grouping is combined with the attack behavior data to locate the attacker's source, determine the attack behavior path, and generate a tracing report; including: Based on the central feature of feature grouping, the attack pattern is matched and the similarity between the central feature and the attack behavior data is calculated: In the formula, S(M j ,A m ) represents the central feature M j Similarity with attack behavior data samples; A m represents the attack behavior data sample; ||M j || represents the central feature M j The modulus length of m ‖ represents the modulus length of the attack behavior data sample; Construct the attack behavior path based on the timestamp T(A m ) and the target IP address IP(A m ), and correlate attack events in time series: Path={A1→A2→…→A n } In the formula, Path represents the constructed attack behavior path; A1→A2→…→A n represents the 1st to nth attack behavior samples in the attack path, arranged in the order of the time when the attack occurred; Extract the source IP address in the attack path, and combine it with the historical attack database to calculate the probability of the attacker appearing on the path to identify the attacker's source: Where Source represents the attacker's source IP address; arg max IP P(IP|Path) means selecting the IP address with the highest probability on the attack path. P(IP|Path) means the probability of a certain IP address being an attacker under the condition of a given attack path. Record the attack time, attack path, and attacker source to form structured traceability data and output a traceability report.

10. A network attack source tracing device in an attack and defense confrontation scenario, comprising: Honeypot deployment module, feature analysis module, cluster analysis module, and traceability report generation module; A network attack source tracing method in an attack and defense confrontation scenario according to any one of claims 1 to 9 is run, characterized in that: Honeypot deployment module, used to deploy honeypots in attack and defense confrontation scenarios to simulate multiple attacked targets and obtain attack behavior data, which includes attack end data sets and attacked end data sets; The feature parsing module is used to parse the time distribution characteristics and access frequency characteristics of the attacking data set, as well as the field interaction characteristics of the attacked data set, and construct a feature vector set based on the parsing results; The clustering analysis module is used to classify the feature vector set based on the clustering algorithm of the dynamic adjustment mechanism. The cluster center position is dynamically initialized by calculating the mean of the time distribution feature, the highest frequency of the access frequency feature and the correlation of the field interaction feature, and the similarity of each vector in the feature vector set to the cluster center is calculated to generate feature groups and extract the central features of each feature group. The traceability report generation module is used to locate the source of the attacker, determine the attack behavior path, and generate a traceability report based on the central features of the feature grouping and combined with the attack behavior data.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is loaded into the processor, the network attack tracing method in the attack and defense confrontation scenario according to any one of claims 1 to 9 is implemented.

12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the network attack tracing method in the attack and defense confrontation scenario described in any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Data link traceability detection method and system based on abnormal network behaviors

    CN121125290A