Problem line determination method based on IP address classification and identification

By constructing IP correlation diagrams and dynamic threshold adjustment methods, the problem of high detection blind spots and false alarm rates for network attacks in the prior art is solved, and accurate identification and rapid response to low-frequency hidden and coordinated attacks are achieved.

CN120238372AActive Publication Date: 2025-07-01BEIJING ZHIXUN TIANCHENG TECH CO LTD

Patent Information

Application Number
CN202510712722.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
2045-05-30

AI Technical Summary

Technical Problem

When identifying network attacks, it is difficult to effectively detect low-frequency covert attacks or protocol-mixed attacks, and the lack of correlation mining of collaborative attack groups, resulting in high rates of missed detection and false alarms, especially under dynamic network loads, which is difficult to balance detection accuracy and real-time performance.

Method used

By collecting network traffic logs, extracting multi-dimensional features to build an IP association diagram, combining community division and PID feedback control mechanisms, dynamically adjusting detection thresholds, identifying abnormal IP groups and generating alarm information.

Benefits of technology

It realizes comprehensive identification of hidden attacks, reduces false alarm rates, improves detection sensitivity and real-time performance, and supports the rapid positioning of attack sources and the formulation of defense strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238372A_ABST
    Figure CN120238372A_ABST
Patent Text Reader

Abstract

The invention relates to the field of network security, and discloses a problem line determination method based on IP address classification and identification, which comprises the following steps of: collecting public network IP behavior data in a network flow log, extracting multi-dimensional characteristics such as a time regularity entropy value, a protocol diversity proportion, target port dispersity and a request rate change rate, and determining a problem line according to the extracted multi-dimensional characteristics; constructing an IP association graph fusing feature similarity, time synchronism and physical position constraints; an abnormal IP group is identified by adopting a community division algorithm, and the detection sensitivity is adaptively optimized according to a real-time network load and a historical false alarm rate in combination with a dynamic threshold adjustment mechanism based on a PID feedback control model; the attack type is judged through protocol-port mapping and Mahalanobis distance statistics, alarm information containing an abnormal IP list, a time window and a classification result is generated, and the alarm information is linked with safety equipment to execute a defense strategy. According to the method, the defects of a traditional method in the aspects of multi-dimensional attack recognition, dynamic environment adaptation and cooperative attack detection are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network security technology, and in particular to a method for determining problematic lines based on IP address classification and recognition. Background Art

[0002] With the complexity and scale of network attack means, traditional security detection methods based on traffic thresholds or static rules have gradually revealed limitations. Existing technologies usually use single-dimensional features (such as request frequency, port access times) for anomaly determination, and it is difficult to effectively identify low-frequency covert attacks or protocol hybrid attack behaviors. For example, in the face of the decentralization and behavior disguise of attack source IPs in distributed denial of service (DDoS) attacks, traditional methods are prone to missed detections due to the single feature dimension, or a large number of false positives due to fixed thresholds under dynamic network loads.

[0003] In addition, existing solutions mostly focus on the behavior analysis of individual IP addresses, lack the association mining of collaborative attack groups, and it is difficult to cope with new threats such as cross-IP collaborative port scanning and cross-protocol attacks. In terms of detection efficiency, the contradiction between the real-time processing requirements of massive IP data and the consumption of computing resources is becoming increasingly prominent. Especially in cloud environments or large enterprise networks, traditional centralized detection architectures are difficult to balance the requirements of detection accuracy and real-time performance.

[0004] The above defects make the existing technologies unable to meet the requirements of accurately identifying problematic lines and adaptively defending against collaborative attacks in complex network environments, and there is an urgent need for a comprehensive solution that integrates multi-dimensional behavior analysis, dynamic threshold adjustment, and efficient group detection. Summary of the Invention

[0005] The purpose of the present invention is to provide a method for determining problematic lines based on IP address classification and recognition, which overcomes the deficiencies of existing methods in multi-dimensional attack behavior recognition, dynamic threshold adaptation, and collaborative attack group detection.

[0006] To achieve the above purpose, the present invention is realized through the following technical solutions: A method for determining problematic lines based on IP address classification and recognition includes the following steps: Collect IP address behavior data in network traffic logs, where the behavior data includes time stamps, protocol types, target ports, and physical locations; Extract multi-dimensional features of each IP address according to the behavior data, and the multi-dimensional features at least include time regularity entropy value, protocol diversity ratio, target port dispersion degree, and request rate change rate; Construct an IP association graph based on the multi-dimensional features, where nodes represent IP addresses, and edge weights are calculated by fusing feature similarity, time synchronization, and physical location distance; Perform community partitioning on the IP association graph to detect abnormal IP groups with similar behavior patterns; Dynamically adjust the detection threshold according to the historical false alarm rate and real-time network load, and the dynamic adjustment is achieved through a feedback control mechanism; Based on the detection threshold and the statistical characteristics of the abnormal IP group, determine the problematic line and generate an alarm message.

[0007] Preferably, when collecting the IP address behavior data in the network traffic log, filter the internal network IP addresses and only retain the public network IP addresses for analysis.

[0008] Preferably, in the step of extracting multi-dimensional features of each IP address according to the behavior data: Temporal regularity entropy value: Calculate the entropy value of the time interval distribution based on the binned statistical results of the IP request timestamps; Protocol diversity ratio: Statistically calculate the ratio of the number of HTTP and HTTPS protocol requests to the total number of IP requests; Target port dispersion: Calculate its variance according to the target port numbers accessed by the IP to measure the concentration degree of port access; Request rate change rate: Quantify the mutation degree of the request rate through the second derivative of the request volume within a sliding window.

[0009] Preferably, the calculation formula for the edge weight in the IP association graph is: Among them, represents the feature vector of IP address and ; represents the cosine similarity between the feature vectors and ; represents the request time series data of IP address and within the time window; is the time synchronization coefficient, indicating the time correlation of the two IP request behaviors; represents the longitude and latitude coordinates of the physical locations of IP address and ; represents the spherical geographical distance calculated based on the longitude and latitude coordinates; is a non-negative weight parameter.

[0010] Preferably, the step of performing community partitioning on the IP association graph to detect abnormal IP groups with similar behavior patterns includes: Use the Louvain algorithm to perform community discovery on the IP association graph and output a set of IP groups with strong relevance; Calculate the protocol entropy of the IP addresses within each community, where the protocol entropy is used to measure the distribution randomness of network protocol types within the community; If the protocol entropy is lower than the preset threshold, shorten the current time window length to improve the detection sensitivity; Perform secondary clustering on the IP addresses within each community, and combine time synchronization and behavioral feature similarity to identify sudden abnormal subgroups.

[0011] Preferably, the step of dynamically adjusting the detection threshold according to the historical false alarm rate and real-time network load includes: Establish a feedback control model, whose input is the real-time false alarm rate and the scale of the abnormal IP group, and the output is the detection threshold adjustment amount; Calculate the threshold adjustment amount based on the PID control mechanism, and the formula is: Among them, is the detection threshold for the next moment; is the detection threshold for the current moment; , is the initial proportionality coefficient, is the current scale of the abnormal IP group, is the total number of active IPs; is the integral coefficient; is the differential coefficient; The real-time false alarm rate is defined as the ratio of the number of false alarms to the total number of alarms; The historical false alarm rate is the integral from the initial moment to the current moment; is the change rate of the false alarm rate over time; Compensate the threshold according to the real-time network load. When the load exceeds the preset safety value, reduce the detection threshold proportionally: Among them, is the final detection threshold after load compensation; is the real-time network load; is the preset network load safety threshold; is the maximum load threshold allowed by the network.

[0012] Preferably, the initial proportionality coefficient , the integral coefficient , the differential coefficient are obtained by training with historical attack data; When the network topology changes, retrain the initial proportionality coefficient , the integral coefficient , the differential coefficient based on the false alarm rate and attack data within a preset time period .

[0013] Preferably, when any of the following conditions is satisfied by the statistical features, the line is determined as a problematic line: The scale of the abnormal IP group exceeds the current detection threshold; The Mahalanobis distance between the abnormal group behavior and the normal mode exceeds a preset multiple of the standard deviation.

[0014] Preferably, the alarm information includes a list of abnormal IPs, an associated time window, and an attack type classification.

[0015] The present invention also provides a problematic line determination device based on IP address classification and recognition, including: A data acquisition module for filtering and partitioning IP behavior data in network traffic logs; A feature extraction module for calculating multi-dimensional features and generating a feature matrix; An association graph construction module for calculating the edge weights between IP nodes and generating an association graph; A community detection module for partitioning abnormal IP groups and dynamically adjusting the time window; A dynamic parameter adjustment module for optimizing the detection threshold through a PID feedback mechanism; An alarm output module for generating an alarm for a problematic line based on the threshold and statistical features.

[0016] In summary, the present invention includes at least one of the following beneficial technical effects: 1. By integrating heterogeneous features such as time regularity entropy value and protocol diversity ratio, the present invention constructs an IP association graph for group behavior analysis, overcoming the detection blind spots caused by traditional methods relying on a single feature. Compared with the detection methods based only on traffic thresholds or protocol statistics, the collaborative analysis of multi-dimensional features can more comprehensively identify concealed attacks (such as low-frequency slow attacks and mixed protocol scans), significantly improving the detection rate of abnormal behaviors.

[0017] 2. The present invention introduces a threshold dynamic adjustment mechanism based on PID feedback control and network load compensation to solve the contradiction between false alarms and missed detections of fixed thresholds in scenarios of traffic fluctuations. Through real-time false alarm rate and attack data feedback, the threshold can be intelligently adjusted according to the attack intensity and network load, maintaining detection sensitivity in high-concurrency service traffic and avoiding excessive alarms in low-load situations.

[0018] 3. Through protocol-port mapping table and Mahalanobis distance statistics determination, the present invention realizes the automatic classification of attack types (such as DDoS, port scanning, etc.). Combining the abnormal IP list with the associated time window information, it supports rapid location of the attack source and traceability analysis of the attack chain, providing data support for formulating subsequent defense strategies and shortening the security incident response cycle.

[0019] 4. The present invention adopts a detection architecture that combines community division and secondary clustering, and conducts fine-grained sub-group analysis on the basis of coarse-grained community division. This hierarchical detection strategy effectively reduces the computational complexity, and at the same time, through the dynamic time window adjustment driven by protocol entropy, improves the real-time capture ability of sudden attacks, and is applicable to efficient detection in large-scale network environments.

[0020] 5. The alarm information of the present invention is integrated with the API interfaces of third-party security devices (such as firewalls, traffic cleaning systems) to implement disposal operations such as automatic blocking of attack IPs and abnormal traffic traction. Through the closed-loop linkage of detection and disposal, the manual intervention delay is reduced, and the active defense ability of the network security protection system is improved. Brief Description of the Drawings

[0021] Figure 1 is a schematic flowchart of the method of the present invention; Figure 2 is a schematic structural diagram of the device of the present invention.

[0022] Among them, 10 is a data acquisition module; 20 is a feature extraction module; 30 is an association graph construction module; 40 is a community detection module; 50 is a dynamic parameter adjustment module; 60 is an alarm output module. Detailed Embodiment

[0023] The following will Figure 1 - Figure 2 , make a further detailed description of the present invention.

[0024] The present invention provides a method for determining a problematic line based on IP address classification and recognition, which realizes accurate detection and alarm of network abnormal lines through multi-dimensional feature fusion, dynamic threshold adjustment, and group behavior analysis.

[0025] As Figure 1 shown, the method for determining a problematic line based on IP address classification and recognition may include the following steps: S1. Collect IP address behavior data in network traffic logs and perform preprocessing; S2. Extract multi-dimensional features of each IP address according to the behavior data; S3. Construct an IP association graph based on the multi-dimensional features; S4. Perform community division on the IP association graph to detect abnormal IP groups with similar behavior patterns; S5. Dynamically adjust the detection threshold according to the historical false alarm rate and the real-time network load, and the dynamic adjustment is realized through a feedback control mechanism; S6. Based on the detection threshold and the statistical features of the abnormal IP groups, determine the problematic line and generate alarm information.

[0026] The following is a detailed description of each step in the method of the present invention, comprehensively elaborating on the specific implementation principles, technical details, and processes for each step.

[0027] For step S1, in this embodiment, the process of collecting and preprocessing the IP address behavior data in the network traffic log is implemented in the following manner: Specifically, the original data of the network traffic log is sourced from the mirror port of the core switch in the network infrastructure and the session record module of the border firewall, and is obtained in real time through a preset data collection interface.

[0028] The behavior data fields include but are not limited to: timestamp (recording accuracy is at the millisecond level), protocol type (including TCP, UDP, HTTP, HTTPS, ICMP, etc.), target port number (value range 1 - 65535), and IP address physical location information.

[0029] Among them, the physical location information is obtained by calling a third - party IP address location query interface, mapping the IP address to longitude and latitude coordinates, and storing it as geotag data.

[0030] Furthermore, in the data preprocessing stage, the original log is filtered using the internal network IP address range defined by the RFC 1918 standard. Preferably, the internal network IP addresses include reserved address segments such as 10.0.0.0 / 8, 172.16.0.0 / 12, and 192.168.0.0 / 16, and efficient filtering is achieved through regular expression matching. After this processing, only public network IP addresses are retained for subsequent analysis to avoid interference from the normal communication behavior of internal network devices on anomaly detection.

[0031] Preferably, to adapt to the dynamic network traffic characteristics, the collected log data is divided into time windows. The time window adopts a sliding window mechanism, with an initial window length of 1 hour and a sliding step of 5 minutes. Through this mechanism, the log data of the continuous time series is segmented into overlapping time - slice data blocks, thereby balancing the detection real - time performance and the integrity of the behavior pattern. For example, for the traffic data with a timestamp of t, the corresponding time window coverage interval is [t - 1 hour, t], and the next window slides to [t - 55 minutes, t + 5 minutes].

[0032] In addition, the data blocks within the time window are further stored in a standardized manner. Preferably, a column - based database is used to establish indexes for fields such as timestamps and protocol types to support efficient feature extraction queries. For a large - scale network environment, a distributed storage architecture is used to handle high - concurrent traffic logs, ensuring data collection throughput and system scalability.

[0033] It should be noted that the acquisition of physical location information needs to meet the requirements of data privacy compliance. Preferably, in the mapping process between IP addresses and longitude and latitude coordinates, a desensitization processing technology is adopted, and only the positioning accuracy at the city level is retained (for example, rounding the longitude and latitude coordinates to two decimal places) to avoid the risk of leakage of precise geographical location information.

[0034] Furthermore, the preprocessed data will be used as the input for multi-dimensional feature calculation. Through the time window division and sliding mechanism, the temporal variation characteristics of IP address behavior can be effectively captured, such as periodic access patterns or sudden request peaks, providing basic data support for the subsequent construction of the IP association graph.

[0035] For step S2, in this embodiment, the extraction of multi-dimensional features of IP addresses is achieved through the following methods: In this embodiment, the multi-dimensional features at least include the time regularity entropy value, the protocol diversity ratio, the target port dispersion degree, and the request rate change rate.

[0036] The calculation of the time regularity entropy value is based on the distribution characteristics of the request behavior of the IP address within the time window. Preferably, the time window is divided into multiple equally spaced bins, the number of requests in each bin is counted, and its probability distribution is calculated. Furthermore, the information entropy is used to quantify the randomness of the time distribution, and the calculation formula is: Among them, is the total number of bins, represents the proportion of the number of requests in the th bin, is the number of requests in this bin, is the total number of requests within the time window. Through this entropy value, the periodicity and regularity of the IP address access behavior can be effectively characterized. For example, a high entropy value reflects a random access pattern, while a low entropy value implies a regular behavior.

[0037] The protocol diversity ratio is used to measure the tendency of the IP address in the selection of application layer protocols. Preferably, the number of requests for the HTTP protocol (port 80) and the HTTPS protocol (port 443) is counted, and the proportion of the total number of requests is calculated: Among them, and are the number of requests for the corresponding protocols respectively. Through this ratio, abnormal protocol usage behaviors can be identified. For example, malicious scans are often accompanied by high-frequency access to non-standard protocols or ports.

[0038] The target port dispersion degree quantifies the concentration degree of the IP address access ports through variance. Specifically, the sequence of all target port numbers accessed by the IP address within the time window is counted , and calculate its variance: Wherein, is the mean of the port numbers. A high variance value indicates scattered port access, which may correspond to multi-service access by normal users; a low variance implies concentrated port access, which may be related to port scanning or brute-force attacks.

[0039] The rate of change of the request rate is used to capture the bursty characteristics of the IP address request behavior. Preferably, a sliding window mechanism is adopted to obtain the time series of the request volume , and calculate its second derivative: Wherein, is the sliding window step size. The second derivative can reflect the acceleration change of the request rate, such as a sudden increase or decrease, so as to identify instantaneous abnormal behaviors such as DDoS attacks.

[0040] It should be noted that the above multi-dimensional features need to be normalized after calculation. Preferably, the Z-score normalization method is adopted to map each feature value to the same dimension to avoid model bias caused by differences in numerical ranges. The normalization formula is: Wherein, , are the mean and standard deviation of this feature in the historical data, respectively.

[0041] Through the fusion of the above multi-dimensional features, the behavior pattern of the IP address can be comprehensively characterized, providing heterogeneous feature inputs for the subsequent construction of the IP association graph. For example, the time regularity entropy value and the rate of change of the request rate jointly reflect the time series behavior characteristics, and the protocol diversity ratio and the target port dispersion degree characterize the application layer behavior differences, thus supporting the multi-dimensional fusion calculation of the edge weights in the association graph.

[0042] For step S3, in this embodiment, the construction of the IP association graph is achieved by fusing multi-dimensional behavior features and spatio-temporal correlation, as follows: The IP association graph uses the set of public IP addresses after preprocessing as nodes, and the association strength between nodes is quantitatively characterized by edge weights. Preferably, if there is at least one communication behavior (such as successful TCP handshake or UDP packet interaction) between two IP addresses within the same time window, a corresponding edge connection is established to avoid the computational complexity of a fully connected graph.

[0043] Furthermore, the calculation of the edge weights is based on a comprehensive evaluation of feature similarity, time synchronization, and physical location distance. Specifically, the edge weight calculation formula is defined as: Among them, the meanings and calculation methods of each component are as follows: Represents the IP address and The cosine similarity of the multi-dimensional feature vectors of. Preferably, the feature vectors The time regularity entropy value extracted by step S2 , the protocol diversity ratio , the target port dispersion and the request rate change rate are composed and are normalized. The calculation formula is: Through this component, the similarity degree of the behavior patterns of two IPs can be effectively measured. For example, the traffic generated by the same attack tool usually has a similar feature distribution.

[0044] Is used to quantify the coordination of the request behaviors of two IP addresses within a time window. Preferably, the time window is divided into multiple sub-windows (for example, each 5 minutes is a sub-window), and the request count sequences of the two IPs in each sub-window are statistically counted and , and its Pearson correlation coefficient is calculated: Among them, , is the average request count. This component can identify coordinated attack behaviors. For example, the burst traffic of multiple IPs in a distributed denial of service (DDoS) attack often shows time synchronization.

[0045] Represents the spherical distance between the physical locations of two IP addresses. Preferably, based on the longitude and latitude coordinates and , the Haversine formula is used for calculation: Among them, is the radius of the earth, , . This component is used to suppress the mis-association caused by network topology of IPs with similar geographical locations, such as a normal server cluster in the same data center.

[0046] The weight parameter is used to adjust the contribution ratio of each component. Preferably, based on the known abnormal IP association samples in the historical attack data, the parameter combination is optimized by the gradient descent method, so that the edge weights of the abnormal group are significantly higher than those of the normal group. In addition, the in the denominatorTo avoid the abnormal amplification of weights caused by the denominator approaching zero when the geographical distance is too close.

[0047] It should be noted that the fusion calculation of edge weights not only covers behavioral characteristics and time dimensions, but also introduces physical location constraints, thereby overcoming the limitations of single-feature association. For example, traditional methods that only rely on protocol or port similarity may lead to misjudgments, while the proposed solution can more accurately distinguish collaborative attacks from the accidental behavioral similarities of normal user groups through multi-dimensional fusion.

[0048] Furthermore, the IP association graph is stored in an adjacency list structure, which is suitable for efficient graph traversal and community partitioning algorithms in large-scale network environments. Preferably, a sparse matrix compression format is used to reduce memory occupancy, and horizontal expansion is achieved through a distributed graph computing framework to support real-time association analysis of a large number of IP addresses.

[0049] For step S4, in this embodiment, the community partitioning and anomaly detection of the IP association graph are achieved through multi-stage collaborative analysis, as follows: Specifically, the community partitioning uses the Louvain algorithm based on modularity optimization. Preferably, by iteratively merging nodes to maximize the modularity metric, the IP association graph is divided into multiple communities, and the edge weights between nodes within each community are significantly higher than the connection weights between communities. The modularity calculation formula is: Among them, is the sum of all edge weights in the graph, , are the weighted degrees of nodes and respectively, , are the communities to which the nodes belong, takes 1 when the communities are the same, otherwise takes 0. Through this algorithm, it is possible to effectively identify IP groups with strong correlations, such as collaborative attack gangs or normal users in the same service cluster.

[0050] Furthermore, protocol entropy analysis is performed for each community. Protocol entropy is used to quantify the diversity of network protocol usage within a community, and the calculation formula is: Among them, is the probability of occurrence of protocol type within the community. Preferably, when When it is lower than the preset threshold, it is determined as a low protocol diversity scenario (for example, a large number of single protocols are used in a DDoS attack). At this time, the current time window length is dynamically shortened to improve the detection sensitivity. It should be noted that the shortening ratio of the time window is positively correlated with the deviation degree of the protocol entropy. For example, when When, the window length is adjusted to 50% of the original value to capture more fine-grained burst behaviors.

[0051] Regarding the local anomalies that may be missed by the coarse-grained community division, the present invention introduces secondary clustering analysis. Preferably, for the IP addresses within each community, based on time synchronization And behavioral feature similarity Construct a two-dimensional feature space, and use a density clustering algorithm (such as DBSCAN) to identify dense sub-populations. Specifically, if the sub-population simultaneously meets the following conditions: 1. Time synchronization (Preferably, ); 2. Feature similarity (Preferably, ); Then it is determined as a sudden abnormal sub-population. Through this step, local abnormal behaviors within a normal community can be effectively distinguished, such as highly coordinated port scanning attacks within a short period of time.

[0052] It should be noted that the collaborative mechanism of dynamic time window adjustment and secondary clustering can solve the problem of insufficient adaptability of traditional static detection methods in complex network environments. For example, in a low protocol diversity scenario, shortening the time window can quickly respond to attack behaviors; while secondary clustering avoids the risk of missed detection caused by overly coarse community division through multi-feature constraints.

[0053] Furthermore, the output results of the abnormal IP group include community identification, sub-clustering labels, and the associated time window range, providing a data basis for subsequent threshold adjustment and alarm generation. Preferably, a graph database is used to store community structure and attribute information, supporting traceability query and visualization display based on time windows.

[0054] For step S5, in this embodiment, the dynamic adjustment of the detection threshold is achieved through a feedback control mechanism and network load adaptive compensation, as follows: The feedback control model takes the real-time false alarm rate And the scale of the abnormal IP group As input variables, and outputs the detection threshold adjustment amount . Preferably, a proportional-integral-derivative (PID) control mechanism is used, and its mathematical expression is: Among them, the definitions of each component are as follows: Proportional term: Used to quickly respond to real-time false alarm rate deviations. Preferably, the proportional coefficient Dynamically correlates with the abnormal group size: Wherein, is the initial proportional coefficient, is the total number of currently active IPs. When the abnormal group size expands, the proportional coefficient increases proportionally, thereby enhancing the sensitivity to large-scale attacks.

[0055] Integral term: Used to eliminate the cumulative error of the historical false alarm rate. For example, the long-term steady-state deviation of the false alarm rate can be gradually corrected through the integral term, avoiding the threshold continuously deviating from the reasonable range.

[0056] Differential term: Used to suppress the instantaneous mutation interference of the false alarm rate, such as the false alarm rate spike caused by short-term network fluctuations.

[0057] It should be noted that the PID control parameters , , The initial values are obtained through training with historical attack data. Preferably, a supervised learning method is adopted, with minimizing the historical false alarm rate and attack undetected rate as the optimization objective, and the optimal parameter combination is iteratively solved through the gradient descent method.

[0058] Furthermore, the real-time network load is introduced as a threshold compensation factor. Preferably, the network load is measured by bandwidth utilization or the number of concurrent connections. When the load exceeds the preset security threshold , the detection threshold is reduced according to the overrun ratio: Wherein, is the maximum load threshold allowed by the network. For example, when , the threshold remains unchanged; when approaches , the threshold is reduced linearly to prevent the undetected risk in high-load scenarios.

[0059] It should be emphasized that the PID parameters are continuously updated during the system operation. Preferably, when the network topology changes (such as adding a server cluster or adjusting the firewall policy), the parameters are retrained based on the false alarm rate and attack data within a preset time period. For example, a sliding time window mechanism is adopted, and the data in the most recent 24 hours is used as the training set to ensure that the model adapts to the network environment changes.

[0060] Through the above dynamic adjustment mechanism, the detection threshold can adapt to changes in the network state and attack situation. For example, when the load is low and the false alarm rate is stable, the threshold is maintained at a relatively high level to reduce false alarms; while in high-load or sudden attack scenarios, the threshold is automatically lowered to improve detection sensitivity.

[0061] For step S6, in this embodiment, the determination of the problem line and the generation of the alarm information are realized through a multi-condition joint decision-making mechanism, which is specifically as follows: The determination conditions are jointly analyzed based on the detection threshold dynamically adjusted in step S5 and the statistical characteristics of the abnormal IP group. When any of the following conditions is met, the current line is determined to be a problem line: 1. The scale of the abnormal IP group exceeds the limit: If the scale of the abnormal IP group exceeds the final detection threshold after load compensation , that is This condition directly reflects whether the attack scale reaches the preset risk level. For example, in a large-scale DDoS attack, the number of abnormal IPs significantly exceeds the normal service traffic threshold.

[0062] 2. The group behavior deviates from the normal mode: If the Mahalanobis distance between the abnormal group behavior and the normal mode exceeds the dynamic determination threshold, that is Among them, is the Mahalanobis distance, is the feature vector of the current abnormal group (including the time regularity entropy value, protocol diversity ratio, etc.), , are respectively the mean vector and covariance matrix of the historical normal group features, is the standard deviation of the normal mode, is the multiple factor dynamically calculated based on the historical attack frequency. Preferably, the multiple factor is adaptively adjusted by statistically counting the attack frequency through a sliding window. For example, it is automatically increased during high attack frequency periods to reduce false alarms.

[0063] It should be noted that the introduction of the Mahalanobis distance can effectively eliminate the dimensional difference and correlation interference between features. For example, there may be a strong correlation between the time regularity entropy value and the change rate of the request rate. Directly using the Euclidean distance will lead to misjudgment, while the Mahalanobis distance realizes feature decoupling through the inverse operation of the covariance matrix, thereby improving the accuracy of statistical determination.

[0064] Furthermore, the generation of the alarm information includes the following fields: Abnormal IP List: Output all IP addresses within the abnormal sub-group detected in step S4. Preferably, they are sorted in descending order of activity (number of requests). Associated Time Window: A time interval accurate to the minute level (e.g., 2023-05-22 14:00~14:05), used to locate the time period when the attack occurred. Attack Type Classification: Mapped to a predefined attack type library based on the combination of protocol type and target port. For example, high-frequency access to the SSH port (22) is classified as a brute-force cracking attack, while multi-source IP high-frequency access to the same HTTP port (80) is classified as a CC attack.

[0065] Preferably, the attack type library is implemented through a preset protocol-port-threat type mapping table. Specifically, the mapping table contains the following typical entries: TCP / 22 (SSH) → Brute-force cracking UDP / 53 (DNS) → DNS amplification attack TCP / 80 (HTTP) → CC attack or web vulnerability exploitation

[0066] It should be emphasized that the alarm information is stored in a structured format (such as JSON or XML) and integrated into the alarm management module of the security operation and maintenance platform. Preferably, it is pushed to third-party security devices (such as firewalls or traffic cleaning devices) through API interfaces to trigger an automated handling process, such as IP blocking or traffic diversion.

[0067] Through the above determination and alarm mechanism, it is possible to quickly locate abnormal lines and accurately identify attack types, providing actionable security event data support for network defense. For example, by combining the dual determination conditions of dynamic thresholds and statistical distances, it not only avoids the limitations of single indicators but also ensures adaptability to new attack patterns.

[0068] Generally speaking, the present invention collects the public IP behavior data in the network traffic logs, extracts multi-dimensional features including time regularity entropy value, protocol diversity ratio, target port dispersion degree, and request rate change rate, constructs an IP association graph that integrates feature similarity, time synchronization, and physical location distance, uses the community division algorithm to detect abnormal IP groups, and dynamically adjusts the detection threshold in combination with the historical false alarm rate and real-time network load. Finally, it determines the problematic line according to the statistical features of the group size exceeding the limit or the behavior pattern deviating, generates alarm information including the abnormal IP list, associated time window, and attack type classification, and realizes the accurate identification and adaptive defense of network abnormal behaviors.

[0069] The problematic line determination device based on IP address classification recognition described below can be correspondingly referred to the problematic line determination method based on IP address classification recognition described above.

[0070] Please refer to the appendix Figure 2 , the present invention also provides a problem line determination device based on IP address classification and recognition, including: A data acquisition module 10 for filtering and partitioning IP behavior data in network traffic logs; A feature extraction module 20 for calculating multi-dimensional features and generating a feature matrix; An association graph construction module 30 for calculating edge weights between IP nodes and generating an association graph; A community detection module 40 for partitioning abnormal IP groups and dynamically adjusting the time window; A dynamic parameter adjustment module 50 for optimizing the detection threshold through a PID feedback mechanism; An alarm output module 60 for generating a problem line alarm based on the threshold and statistical features.

[0071] The device in this embodiment can be used to execute the above method embodiment, and its principle and technical effects are similar, which will not be elaborated here.

[0072] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for determining a problematic line based on IP address classification and recognition, characterized in that It includes the following steps: Collect IP address behavior data in network traffic logs, where the behavior data includes timestamp, protocol type, target port, and physical location; Extract multi-dimensional features of each IP address according to the behavior data, and the multi-dimensional features at least include time regularity entropy value, protocol diversity ratio, target port dispersion degree, and request rate change rate; Construct an IP association graph based on the multi-dimensional features, where nodes represent IP addresses, and edge weights are calculated by fusing feature similarity, time synchronization, and physical location distance; Perform community partitioning on the IP association graph to detect abnormal IP groups with similar behavior patterns; Dynamically adjust the detection threshold according to the historical false alarm rate and real-time network load, and the dynamic adjustment is achieved through a feedback control mechanism; Based on the detection threshold and statistical features of abnormal IP groups, determine the problem line and generate an alarm message.

2. The problem line determination method based on IP address classification and recognition according to claim 1, characterized in that, When collecting IP address behavior data in network traffic logs, filter internal network IP addresses and only retain public network IP addresses for analysis.

3. The method for determining a problematic line based on IP address classification and recognition according to claim 1, characterized in that, In the step of extracting multi-dimensional features of each IP address according to the behavior data: Time regularity entropy value: Calculate the entropy value of the time interval distribution based on the binned statistical results of IP request timestamps; Protocol diversity ratio: Statistically calculate the ratio of HTTP and HTTPS protocol requests to the total number of IP requests; Target port dispersion degree: Calculate its variance according to the target port numbers accessed by the IP to measure the concentration degree of port access; Request rate change rate: Quantify the mutation degree of the request rate through the second derivative of the request volume within a sliding window.

4. The problem line determination method based on IP address classification and recognition according to claim 1, wherein, The calculation formula for the edge weight in the IP association graph is: Among them, represents the IP address and the eigenvector; represents the eigenvector and the cosine similarity; represents the IP address and the request time series data within the time window; is the time synchronization coefficient, representing the time correlation of the IP request behaviors of two IP addresses; represents the IP address and the latitude and longitude coordinates of the physical location; represents the spherical geographical distance calculated based on the latitude and longitude coordinates; is a non - negative weight parameter.

5. The method for determining a problematic line based on IP address classification and recognition according to claim 1, characterized in that The step of performing community partitioning on the IP association graph to detect abnormal IP groups with similar behavior patterns includes: Use the Louvain algorithm to perform community discovery on the IP association graph and output a set of IP groups with strong relevance; Calculate the protocol entropy of the IP addresses within each community, and the protocol entropy is used to measure the randomness of the distribution of network protocol types within the community; If the protocol entropy is lower than the preset threshold, shorten the current time window length to improve the detection sensitivity; Perform secondary clustering on the IP addresses within each community, and combine time synchronization and behavior feature similarity to identify sudden abnormal sub-groups.

6. The method for determining a problematic line based on IP address classification and recognition according to claim 1, characterized in that, The step of dynamically adjusting the detection threshold according to the historical false alarm rate and real-time network load includes: Establish a feedback control model, whose input is the real-time false alarm rate and the scale of abnormal IP groups, and the output is the detection threshold adjustment amount; Calculate the threshold adjustment amount based on the PID control mechanism, and the formula is: Among them, is the detection threshold for the next moment; is the detection threshold for the current moment; , is the initial proportionality coefficient, is the scale of the current abnormal IP group, is the total number of active IPs; is the integral coefficient; is the differential coefficient; The real-time false alarm rate is defined as the ratio of the number of false alarms to the total number of alarms; The historical false alarm rate is the integral from the initial moment to the current moment; is the change rate of the false alarm rate over time; Compensate the threshold according to the real-time network load. When the load exceeds the preset safety value, reduce the detection threshold proportionally: Among them, is the final detection threshold after load compensation; is the real-time network load; is the preset network load security threshold; is the maximum load threshold allowed by the network.

7. The method for determining a problem line based on IP address classification and recognition according to claim 6, characterized in that, The initial proportionality coefficient , integral coefficient , differential coefficient are obtained by training with historical attack data; When the network topology changes, retrain the initial proportionality coefficient based on the false alarm rate and attack data within a preset time period , integral coefficient , differential coefficient .

8. The problem line determination method based on IP address classification and recognition according to claim 1, characterized in that, When any of the following conditions is met for the statistical features, it is determined as a problem line: The scale of the abnormal IP group exceeds the current detection threshold; The Mahalanobis distance between the behavior of the abnormal group and the normal mode exceeds a preset multiple of the standard deviation.

9. The method for determining a problematic line based on IP address classification and recognition according to claim 1, wherein The alarm message includes a list of abnormal IPs, the associated time window, and the attack type classification.

10. A problem line determination device based on IP address classification recognition, which is used to execute the method described in any one of claims 1-9, characterized in that It includes: A data collection module for filtering and partitioning IP behavior data in network traffic logs; Feature extraction module, which is used to calculate multi-dimensional features and generate a feature matrix; Association graph construction module, which is used to calculate the edge weights between IP nodes and generate an association graph; Community detection module, which is used to divide abnormal IP groups and dynamically adjust the time window; Dynamic parameter tuning module, which is used to optimize the detection threshold through a PID feedback mechanism; Alarm output module, which generates alarms for problem lines based on thresholds and statistical features.

Citation Information

Patent Citations

  • Network attack fingerprint extraction method and device, equipment and storage medium

    CN117240556A

  • Method and system for detecting network attack of telecommunication network signaling system

    CN118264473A

  • Anomaly detection to identify security threats

    US10673880B1

  • Network request processing method and apparatus, electronic device, and storage medium

    WO2019052469A1

Cited By

  • Network security sensing method and device based on artificial intelligence and network information entropy

    CN121077743A

  • Large model cue word attack detection method and system, terminal and medium

    CN121217372A

  • Database user abnormal behavior detection system

    CN121561898A

  • A database user abnormal behavior detection system

    CN121561898B

  • Abnormal attack identification method and system based on flow data analysis

    CN122027370A