Terminal network information processing and detecting method based on aggregation analysis
By adopting a three-tiered network information processing method, the problems of low efficiency in processing massive amounts of terminal data, insufficient accuracy in anomaly detection, and poor adaptability to dynamic network environments in traditional methods are solved. This method enables efficient, accurate, and real-time detection of abnormal terminal network behavior, thereby improving network security protection capabilities.
Patent Information
- Application Number
- CN202511076343.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-11-14
AI Technical Summary
Traditional network information processing methods suffer from low efficiency in processing massive amounts of terminal data, insufficient accuracy in anomaly detection, and poor adaptability to dynamic network environments, making it difficult to achieve efficient, accurate, and real-time detection of abnormal terminal network behavior.
A three-tier architecture is adopted: the data acquisition layer deploys lightweight probes to standardize data formats and perform dynamic sampling; the aggregation analysis layer performs three-dimensional aggregation analysis of temporal, spatial, and multi-protocol association entropy; and the dynamic detection layer combines linear regression and random forest to construct a hybrid detector and uses adaptive parameters and dynamic threshold mechanisms for anomaly scoring.
It enables efficient, accurate, and real-time detection of abnormal behavior on terminal networks, improves network security protection capabilities, reduces false alarm rates, and meets millisecond-level response requirements.
Smart Images

Figure CN120956465A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network security technology, and in particular to a network external connection analysis-driven attack clue linkage forensics method, which is used to dynamically perceive network external connection behavior, associate attack clues, fuse multi-source evidence chains and construct a knowledge graph to achieve accurate tracing and evidence collection of network attacks. Background Technology
[0002] With the rapid development of IoT and 5G technology, the number of terminal devices is growing exponentially, and traditional network information processing methods face three major challenges: (1) massive terminal data leads to low processing efficiency of single nodes; (2) insufficient accuracy of abnormal behavior detection in distributed network environments; and (3) limited real-time response capability under dynamic network topology.
[0003] Existing technologies mainly employ threshold alarms based on traffic statistics (RFC 7011) or isolated detection models based on machine learning (such as LSTM time-series prediction), but they have significant drawbacks: ① They only focus on single-dimensional traffic characteristics and lack multi-protocol layer correlation analysis; ② Static thresholds cannot adapt to dynamic network environments; ③ Processing latency is difficult to meet millisecond-level response requirements. Therefore, a novel method for terminal network information aggregation analysis and dynamic detection is urgently needed.
[0004] Full names of terms and abbreviations
[0005] • RFC 7011: Request for Comments document for the Traffic Mirroring Protocol (NetFlow v9), which defines the format and transmission specifications for traffic statistics.
[0006] •LSTM: Long Short-Term Memory, a recurrent neural network used for time-series data prediction.
[0007] IEC 61850: A standard developed by the International Electrotechnical Commission for communication networks and systems in substations, used for information exchange between equipment in power systems.
[0008] Modbus: An industrial communication protocol widely used for data transmission in devices such as smart meters.
[0009] • DNP3: Distributed Network Protocol, commonly used in remote monitoring systems in industries such as power and water conservancy.
[0010] • RTU: Remote Terminal Unit, a device used for real-time acquisition and transmission of field data. Summary of the Invention
[0011] This invention proposes a terminal network information processing and detection method based on aggregation analysis, which achieves efficient processing and accurate detection of terminal network information through a three-level architecture:
[0012] Data acquisition layer: Deploy lightweight probe agents to collect terminal network data, store it in a standardized format, and dynamically adjust the sampling period according to the number of connections to reduce the impact on terminal performance;
[0013] Aggregation Analysis Layer: Performs three-dimensional aggregation analysis from the time dimension (sliding window to calculate feature vectors), spatial dimension (constructing terminal relationship graphs), and multi-protocol association entropy (cross-protocol layer feature extraction) to mine multi-dimensional features;
[0014] Dynamic detection layer: It uses a hybrid model of linear regression and random forest to calculate anomaly scores, and combines a dynamic threshold mechanism with adaptive parameters to adjust the detection threshold in real time, so as to achieve accurate identification of abnormal network behavior.
[0015] Technical solution
[0016] The purpose of this invention is to provide a terminal network information processing and detection method based on aggregation analysis. This method aims to address the problems of low efficiency in processing massive amounts of terminal data, insufficient accuracy in anomaly detection, and poor adaptability to dynamic network environments in traditional network information processing methods. It achieves efficient, accurate, and real-time detection of abnormal terminal network behavior, thereby improving network security protection capabilities. The specific implementation process is as follows:
[0017] 1. Data Acquisition Layer
[0018] 1.1 Deployment of a lightweight probe agent
[0019] To achieve effective collection of network data from terminal devices, this invention employs a lightweight probe agent deployed on the terminal devices. This lightweight design aims to minimize the impact on terminal device performance, ensuring that it does not significantly increase device resource consumption, such as CPU utilization or memory usage, during operation.
[0020] The lightweight probe agent boasts high compatibility, enabling stable operation on a variety of terminal devices, including but not limited to personal computers, servers, and mobile devices. Its modular architecture allows for independent yet collaborative operation of each functional module, facilitating maintenance and expansion.
[0021] During deployment, the lightweight probe agent automatically detects the operating system type and network environment of the terminal device and adaptively configures itself based on the detection results. For example, for different operating systems (such as Windows, Linux, macOS, etc.), it calls the corresponding system interfaces to obtain network data; for different network environments (such as wired networks, wireless networks, etc.), it adjusts its collection strategy to ensure the accuracy and completeness of data collection.
[0022] 1.2 Data Format Standardization
[0023] The raw network data collected is usually diverse and complex, and needs to be standardized to facilitate subsequent processing and analysis. This invention standardizes the data format according to formula (1):
[0024] D i ={t,src_ip,dst_ip,proto,payload_len,flag_bits} (1)
[0025] in:
[0026] 't' represents the timestamp of data collection, accurate to the millisecond level. Recording timestamps is crucial for subsequent data analysis, helping us understand the order and temporal distribution of network data generation, thereby uncovering potential network anomalies.
[0027] `src_ip` is the source IP address, which is the IP address of the device sending the network data. The source IP address helps us determine the source of the data, thus enabling us to classify and manage different data sources.
[0028] dst_ip is the destination IP address, which is the IP address of the device receiving the network data. The destination IP address helps us understand the flow of data, thereby analyzing network communication patterns.
[0029] `.proto` represents a network protocol type, such as TCP, UDP, ICMP, etc. Different network protocols have different characteristics and uses, and understanding protocol types can help us better understand the meaning of network data.
[0030] payload_len represents the length of the data payload, that is, the number of bytes of data actually carried in a network data packet. Changes in the data payload length can reflect network communication traffic patterns, thereby detecting potential network attacks.
[0031] `flag_bits` are flag bits used to indicate special states of network packets, such as SYN, ACK, and FIN. Changes in these flag bits can help us understand the establishment, maintenance, and closure of network connections, thereby identifying potential network security threats.
[0032] This standardized data format allows us to process network data from different sources and of different types in a unified manner, improving the efficiency and accuracy of data processing.
[0033] 1.3 Dynamic adjustment of sampling period
[0034] To minimize the workload and network bandwidth usage while ensuring data acquisition quality, this invention employs a method of dynamically adjusting the sampling period Δt. The sampling period Δt is dynamically adjusted according to formula (2):
[0035] (N is the current number of connections)(2)
[0036] When the current number of connections N is small When the value is large, the sampling period Δt will take a larger value. This is to ensure that enough data can be collected. As the current number of connections N increases, The value will gradually decrease, when When the sampling period is less than 50ms, the sampling period Δt will be set to 50ms to avoid sampling too frequently and increasing the system load.
[0037] This method of dynamically adjusting the sampling period has the following advantages:
[0038] It can automatically adjust the sampling frequency according to the actual network usage, improve sampling accuracy when the number of network connections is small, and reduce the amount of data collected when the number of network connections is large, thereby achieving high efficiency and flexibility in data collection.
[0039] It can effectively reduce the use of network bandwidth and avoid interference with normal network communication caused by data collection.
[0040] It can adapt to different network environments and application scenarios, improving the stability and reliability of the system.
[0041] Through the design and implementation of the above data acquisition layer, this invention can efficiently and accurately collect network data from terminal devices, providing a solid foundation for subsequent data analysis and processing.
[0042] 2. Aggregation Analysis Layer
[0043] 2.1 Time-dimensional aggregation
[0044] In the analysis of network data, feature extraction along the time dimension is crucial for discovering patterns and anomalies in network behavior. This invention employs a sliding window technique for aggregation operations along the time dimension.
[0045] (1) Setting up a sliding window
[0046] A sliding window of T = 1 second was set, which is a suitable time interval determined through extensive experiments and practical applications. On the one hand, this time window can capture the changing characteristics of network behavior in a short period of time, avoiding the omission of some rapidly changing abnormal behaviors due to an excessively large window; on the other hand, it is not too small, which would lead to overly fragmented data and difficulty in extracting effective feature information.
[0047] The sliding window slides forward with a fixed step size, which can be adjusted according to specific application scenarios and analysis needs. In this invention, to ensure data continuity and integrity, the step size is set to 0.1s. This reduces data overlap to some extent while ensuring sufficient correlation information between adjacent windows.
[0048] (2) Calculation of eigenvectors
[0049] Within each sliding window, the following feature vector is calculated according to formula (3):
[0050] F T =[∑len,σ(len),H(proto),H(port)] (3)
[0051] ∑len: Represents the sum of the data payload lengths of all network packets within the window. This metric reflects the amount of network traffic within a time window. By analyzing it, one can understand the network's busyness and traffic trends. For example, if this value suddenly increases within a certain period, it may indicate a large amount of data transmission, such as file downloads or video streaming, or it may indicate a network attack, such as a DDoS attack.
[0052] σ(len): This is the standard deviation of the data payload length within the window. The standard deviation reflects the dispersion of the data; in network data, it reflects the fluctuation in network packet size. A larger standard deviation indicates greater variation in packet size, potentially suggesting the simultaneous operation of multiple different types of network applications. Conversely, a smaller standard deviation indicates more uniform packet size, possibly indicating that a specific network application is dominating network communication.
[0053] H(proto): This represents the information entropy of the network protocol type. Information entropy is an indicator used to measure the degree of uncertainty or disorder in information. In a network environment, different protocols have different functions and uses. By calculating the information entropy of protocol types, we can understand the distribution of network protocol usage. If the information entropy is high, it indicates that a wide variety of protocols are used in the network and their distribution is relatively even, indicating strong network diversity. Conversely, if the information entropy is low, it indicates that the network is mainly dominated by a few protocols.
[0054] H(port): This is the information entropy of the port number. Port numbers are used in network communication to identify different applications or services. Calculating the information entropy of a port number can help us understand the usage of different applications on the network. For example, in a corporate network, if the information entropy of a certain port number suddenly increases, it may mean that a new application has been introduced into the network, or that there is abnormal network activity.
[0055] Information entropy is calculated using formula (4):
[0056]
[0057] Where P(x) i ) is event x i The probability of occurrence. When calculating H(proto) and H(port), x i These represent different protocol types and port numbers, P(x) i The frequency of the protocol type or port number appearing in the window.
[0058] 2.2 Spatial Dimension Aggregation
[0059] The aggregation of spatial dimensions is mainly for analyzing the relationships and interaction patterns between terminal devices. This invention achieves this goal by constructing a terminal relationship graph.
[0060] (1) Construction of the terminal relationship diagram
[0061] Construct a terminal relationship graph G = (V, E), where V represents the set of terminal devices, and each node represents a terminal device; E represents the set of connection relationships between terminal devices, and each edge represents network communication between two terminal devices.
[0062] In practical applications, the connection relationship between terminal devices is determined by analyzing the source and destination IP addresses of network packets. When one terminal device sends a network packet to another, an edge is added between their corresponding nodes.
[0063] (2) Calculation of edge weights
[0064] To more accurately describe the strength of the relationships between terminal devices, each edge in the terminal relationship graph is assigned a weight value. The edge weights are calculated according to formula (5):
[0065]
[0066] IP i and IP j These represent the sets of IP addresses used for communication between the two terminal devices. |IP i ∩IP j | represents the number of elements in the intersection of two sets, i.e., the number of IP addresses shared by the two terminal devices for communication; | IP i ∪IP j | represents the number of elements in the union of two sets, which is the total number of IP addresses used for communication between the two terminal devices.
[0067] edge weight w ij The value of w ranges from [0, 1]. ij When w = 0, it means that the two terminal devices have the same IP address and their relationship is very close; when w ij When = 1, it means that the IP addresses of the two terminal devices do not overlap and the relationship between them is relatively distant.
[0068] By calculating edge weights, the strength of relationships between terminal devices can be quantified, thereby revealing the network topology and interaction patterns between terminal devices. For example, in an enterprise network, edge weights can be used to identify terminal devices closely related to the core server, as well as isolated terminal devices that interact less with other devices.
[0069] 2.3 Multi-protocol Association Entropy (Key Algorithm)
[0070] In complex network environments, different protocols often have interrelationships and dependencies. To extract these cross-protocol layer (L2-L7) correlation features, this invention proposes the concept of multi-protocol correlation entropy.
[0071] (1) Definition of multi-protocol association entropy
[0072] Multi-protocol association entropy H m Calculate according to formula (6):
[0073]
[0074] in:
[0075] K represents the number of protocol types, i.e., the number of different protocols used in the network. In a real network environment, multiple protocols such as TCP, UDP, and ICMP may coexist, and K reflects the diversity of network protocols.
[0076] n k This represents the number of times protocol k appears during the observation period. By counting the occurrences of different protocols, we can understand the frequency of their use in the network.
[0077] N is the total number of times all protocols appear, i.e.
[0078] T k The duration of protocol k is the length of time that protocol k continuously operates in the network. Different protocols may have different runtime characteristics. For example, the TCP protocol is typically used to establish reliable connections, and its duration may be relatively long; while the ICMP protocol is mainly used for network diagnostics and control, and its duration may be relatively short.
[0079] T total It is the total length of the observation period.
[0080] (2) The role of multi-protocol association entropy
[0081] Multi-protocol association entropy H m By comprehensively considering the frequency and duration of protocol occurrences, the correlation between different protocols in the network can be more fully reflected. Calculating multi-protocol correlation entropy can reveal the collaborative working patterns and potential abnormal behaviors among different protocols in the network. For example, a significant change in the frequency and duration of a particular protocol may lead to a change in the multi-protocol correlation entropy, thus indicating the presence of anomalies in the network, such as network attacks or protocol vulnerability exploitation.
[0082] Through the above-mentioned time-dimensional aggregation, spatial-dimensional aggregation, and multi-protocol association entropy calculation, this invention can perform aggregation analysis on network data from multiple perspectives, extract valuable feature information, and provide strong support for subsequent anomaly detection and security analysis.
[0083] 3. Dynamic Detection Layer
[0084] 3.1 Anomaly Scoring Model (Key Algorithm)
[0085] In the field of network security detection, single detection models often struggle to cope with complex and ever-changing network environments and diverse attack methods. To improve the accuracy and reliability of detection, this invention combines linear regression and random forest to construct a hybrid detector.
[0086] (1) Principle of hybrid detectors
[0087] Linear regression is a classic statistical model that predicts output values by linearly combining input features. In this invention, the linear regression component captures the linear relationship between features and anomaly scores, providing a basic predicted value for the anomaly scores. Its expression is β0 + β1F1 + ... + β n F n Where F1, F2, ..., F n These are the feature vectors extracted from the previous aggregation analysis layer, β0, β1, ..., β n These are the coefficients of the linear regression model, which are obtained by training on historical data.
[0088] Random forest is an ensemble learning method that consists of multiple decision trees. It improves the accuracy and stability of predictions by combining the results of these multiple decision trees. In this invention, the random forest model RF(F) can handle complex nonlinear relationships between features, supplementing and correcting the prediction results of linear regression models.
[0089] The results of linear regression and random forest are combined to obtain the final anomaly score, which is calculated using formula (7):
[0090] Score = α·(β0 + β1F1 + ... + β) n F n )+(1-α)·RF(F) (7)
[0091] (2) The role of dynamic weighting factors
[0092] The parameter α∈[0,1] is a dynamic weighting factor, which combines the detection capabilities of the linear regression and random forest models. When α is close to 1, it indicates that the linear regression model dominates in anomaly scoring; when α is close to 0, it indicates that the random forest model has a greater influence.
[0093] The dynamic weighting factor α can be adjusted in real time based on changes in the network environment and detection performance. For example, when the network environment is relatively stable and the linear relationship between features and anomaly scores is obvious, the value of α can be appropriately increased to fully leverage the advantages of the linear regression model; while when the network environment is complex and variable with many nonlinear relationships, the value of α can be decreased to allow the random forest model to play a greater role. In this way, the hybrid detector can adaptively adjust the weights of the two models, improving the performance of anomaly detection.
[0094] The real-time adjustment rules are based on the following network environment indicators:
[0095] 1) Flow fluctuation coefficient (C) f ):
[0096] Defined as the ratio of the current window flow standard deviation to the historical mean, i.e.:
[0097]
[0098] When C f When the threshold is >1.5 (preset traffic mutation threshold), the network is determined to have entered a complex environment, triggering a weight adjustment.
[0099] 2) Characteristic nonlinearity (N) f ):
[0100] The mutual information entropy of eigenvectors is used as a measure; if C f A threshold (such as 0.8) indicates the existence of a strong nonlinear relationship.
[0101] Adjustment logic:
[0102] When C f ≤1.5 and N f When ≤0.8 (stable environment): α=α base +Δα·sigmoid(t), where α base =0.7 (linear model dominant), Δα = 0.2 is the fine-tuning step size.
[0103] When C f >1.5 or N f >0.8 (complex environment): Where λ = 0.5 is the attenuation coefficient, and α prev The weights are set to the weights of the previous time step, and the weights decay exponentially to enhance the nonlinear fitting ability of the random forest model.
[0104] 3.2 Adaptive Threshold Mechanism
[0105] (1) Dynamic threshold update algorithm
[0106] In anomaly detection, setting the threshold is crucial for accurately identifying abnormal behavior. Traditional fixed threshold methods often fail to adapt to dynamic changes in the network environment, easily leading to false positives or false negatives. To address this issue, this invention employs a dynamic threshold update algorithm.
[0107] The core idea of this algorithm is to adjust the threshold in real time based on the abnormal ratings of the current window. Specifically, the threshold update formula (8) is:
[0108] θ t+1 =θ t +η·(μ t -θ t (8)
[0109] in:
[0110] θt It is the threshold at the current moment.
[0111] μ t This is an exponentially weighted moving average of the current window's score. The exponentially weighted moving average places greater emphasis on the impact of recent data, allowing the threshold to respond more quickly to changes in the network environment. Its calculation formula is μ. t =ω·Score t +(1-ω)·μ t-1 Score t ω is the abnormal rating of the current window, and ω is the weighting coefficient, which ranges from [0,1]. The closer ω is to 1, the more importance is attached to the rating of the current window.
[0112] η is the learning rate, which controls the speed at which the threshold is updated. If the learning rate is too large, the threshold may become overly sensitive, leading to frequent fluctuations; if the learning rate is too small, the threshold will update too slowly, failing to adapt to changes in the network environment in a timely manner. In practical applications, the value of the learning rate η needs to be adjusted according to the specific network environment and detection requirements.
[0113] By using a dynamic threshold update algorithm, the threshold can be dynamically adjusted as the network environment changes, thereby improving the accuracy and adaptability of anomaly detection.
[0114] (2) Dynamic threshold function (key algorithm)
[0115] To further improve the adaptability of the threshold to dynamic changes in network topology, this invention proposes a dynamic threshold function. This function comprehensively considers anomaly scoring, time factors, and adaptive parameters, and can dynamically adjust the threshold according to the real-time state of the network.
[0116] The formula (9) for calculating the dynamic threshold function is:
[0117]
[0118] in:
[0119] S(t) is the anomaly score at the current time.
[0120] S0 is a baseline score, defined as the median of historical anomaly scores, used to determine the center point of the threshold function. By statistically analyzing score data within historical normal business cycles, the median is taken as the baseline value under stable conditions (e.g., in a power grid scenario, the median of one week of historical normal data, S0 = 60).
[0121] `k` is a slope parameter that controls how steeply the threshold function changes with anomaly scores. When `k` is large, the threshold function is more sensitive to changes in anomaly scores; when `k` is small, the threshold function changes relatively smoothly. (For scenarios with high real-time requirements, `k=2` is used.)
[0122] θ min and θ max These are the minimum and maximum values of the threshold, respectively. They limit the range of the threshold value and prevent unreasonable fluctuations. Example: In a power grid scenario, based on the "Regulations for Security Protection of Power Monitoring Systems," we set: θ min =40 (corresponds to the lower limit of normal business fluctuations, allowing for a small number of non-critical anomalies); θ max =80 (corresponds to a high-risk threshold, triggering an emergency alarm).
[0123] λ and γ are adaptive parameters used to control the trend of the threshold changing over time. As time t increases, λe... -γt The value of gradually decreases, causing the threshold to gradually approach . For example, during the power grid's power supply guarantee period, λ = 10 and γ = 0.5 can be set to accelerate the stability of the threshold.
[0124] The dynamic threshold function incorporates an online learning mechanism for adaptive parameters k, λ, and γ, which can be adjusted in real time according to dynamic changes in the network topology. For example, when significant changes occur in the network topology, adjusting the values of k, λ, and γ allows the threshold function to adapt to the new network environment more quickly, thereby improving the accuracy and reliability of anomaly detection.
[0125] By combining an anomaly scoring model and an adaptive threshold mechanism, the dynamic detection layer of this invention can effectively detect abnormal behavior in the network and can adaptively adjust the detection strategy to adapt to the ever-changing network environment. Attached Figure Description
[0126] Figure 1 This is a flowchart of the method described in this invention. The flowchart illustrates the process from lightweight probe deployment and dynamic sampling at the data acquisition layer, to temporal, spatial, and multi-protocol correlation entropy analysis at the aggregation analysis layer, to anomaly scoring and adaptive threshold judgment at the dynamic detection layer, ultimately achieving accurate detection and alarm of abnormal network behavior. Detailed Implementation
[0127] The present invention will be further described below with reference to specific embodiments:
[0128] 1. Example 1
[0129] 1.1 Application Scenarios and Equipment Deployment
[0130] In smart grid scenarios, lightweight probe agents are deployed on each of the massive number of terminal devices (such as smart meters, relay protection devices, and remote terminal units) distributed in substations, transmission lines, and distribution substations.
[0131] Equipment types: Covering IEC 61850 protocol equipment (such as protection and control devices), Modbus protocol smart meters, TCP / IP-based video surveillance equipment, etc. from different manufacturers.
[0132] Deployment method: The probe agent is distributed and managed in batches through an edge computing gateway (such as Huawei NetEco 6000), and it supports cross-operating system compatibility (such as embedded Linux, VxWorks).
[0133] 1.2 Implementation Details of the Data Acquisition Layer
[0134] (1) Standardized data format
[0135] Based on the communication characteristics of power grid equipment, the data format of extended formula (1) is as follows:
[0136] D i ={t,src_ip,dst_ip,proto,payload_len,flag_bits,device_type,area_code
[0137] Added field:
[0138] device_type: Identifies the device type (e.g., "protection device", "smart meter", "RTU");
[0139] area_code: Labels the area code to which the device belongs (e.g., "Guangxi-01-003"), used for spatial dimension area aggregation analysis.
[0140] (2) Dynamic sampling strategy
[0141] Based on the periodic characteristics of power grid operations, the sampling period is set as follows:
[0142]
[0143] New parameter: T cycle For power grid business cycles (such as a meter reading cycle of 60s), the sampling cycle is automatically shortened to 50ms during peak business hours (such as 9:00-11:00 and 15:00-17:00 daily), and dynamically adjusted according to the number of connections during off-peak hours.
[0144] 1.3 Application of Aggregation Analysis Layer in Power Grids
[0145] (1) Time-dimensional aggregation optimization
[0146] To address the heartbeat message detection requirements of power grid equipment, anomaly features are added within a sliding window of T=1s:
[0147] F T=[∑len,σ(len),H(proto),H(port),heartbeat_loss]
[0148] New feature: heartbeat_loss is the number of missing heartbeat messages within the window, calculated based on the heartbeat interval (e.g., 30s) specified in the IEC 61850 protocol to determine the real-time missing rate.
[0149] (2) Spatial dimension aggregation of power grid topology mapping
[0150] Construct a power grid terminal relationship graph G = (V, E), and adjust the edge weight formula based on the topological characteristics of the power communication network:
[0151]
[0152] Parameter adjustment:
[0153] Introducing a physical distance factor (physical distance is the length of the optical cable between devices, D) max (To maximize the communication distance of the regional power grid), strengthen the interconnectivity of equipment within the same substation;
[0154] The correlation between IP communication and physical topology is balanced by a weighting coefficient (0.6:0.4).
[0155] (3) Power grid protocol adaptation of multi-protocol association entropy
[0156] For a hybrid protocol environment for power grids (IEC 61850, Modbus, DNP3), the number of protocol types K in formula (6) is 3, and the protocol duration T is defined. k :
[0157] IEC 61850 protocol: duration of real-time data transmission for corresponding protection devices;
[0158] Modbus protocol: corresponds to the duration of periodic meter reading tasks;
[0159] Through correlation entropy H m Anomalies in the detection protocol interaction (such as a sudden increase in Modbus traffic outside of meter reading periods).
[0160] 1.4 Power Grid Safety Detection Scenarios in Dynamic Detection Layer
[0161] (1) Training of attack features for anomaly scoring models
[0162] To address typical power grid attacks (such as IEC 61850 protocol flooding attacks and Modbus command tampering), protocol feature weights are added to the hybrid detector.
[0163] Linear regression part: β proto=1.2 (Protocol feature weight is higher than regular traffic feature);
[0164] Random Forest Model: Focuses on training the "protocol type-port anomalous combination" feature (e.g., the IEC 61850 protocol uses a port other than 8001).
[0165] (2) Adaptive threshold mechanism for power grid service adaptation
[0166] 1) Dynamic threshold update algorithm:
[0167] The learning rate η = 0.3 (higher than in normal scenarios) ensures rapid convergence to the threshold during brief flow fluctuations caused by power grid dispatching operations (such as remote switching).
[0168] 2) Dynamic threshold function:
[0169] Set the baseline score S0 = 60 (average score for normal business operations), and the threshold range θ. min =40, θ max =80, during special periods of the power grid, the threshold is accelerated to θ by using parameter γ=0.5. max Convergence improves detection sensitivity.
[0170] 1.5 Implementation Results and Patent Innovativeness
[0171] 1.5.1 Detection capability verification
[0172] (1) Test environment configuration:
[0173] Terminal equipment scale: In a pilot deployment in a provincial power grid, it covers 2000+ terminal devices, including:
[0174] Protection devices conforming to the IEC 61850 protocol (distributed in 50 substations);
[0175] 1200+ Modbus smart meters (covering 3 distribution zones);
[0176] 200+ TCP / IP-based video surveillance devices (along power transmission lines).
[0177] Network topology: It adopts a three-level architecture of "substation-aggregation layer-core layer" and the maximum physical distance of the communication link is 200 kilometers (from the transmission line to the main station).
[0178] (2) Data source:
[0179] Normal business data: Real-time communication logs of the power grid were collected from October to December 2024, including 50GB of raw traffic data, and marked with characteristics such as normal protocol interactions and heartbeat messages.
[0180] Attack dataset:
[0181] Inject IEC 61850 protocol flood attack (simulating abnormal traffic of 1000 packets / second);
[0182] Tampering with Modbus commands (changing meter readings to abnormal values);
[0183] Forged unauthorized IP communications (simulated device hijacking attack).
[0184] Evaluation indicators:
[0185] Accuracy: The proportion of correctly detected samples out of the total number of samples;
[0186] F1-score: The harmonic mean of precision and recall, calculated as follows:
[0187] False Positive Rate (FPR): The proportion of normal samples that are mistakenly identified as abnormal out of the total number of normal samples.
[0188] 1.5.2 Comparative Experimental Design
[0189] (1) Comparison method:
[0190] Traditional Method 1: SCADA system based on static thresholds (baseline method, widely used in power grids);
[0191] Traditional Method 2: Isolated LSTM time series prediction model (analyzing only single-dimensional traffic characteristics).
[0192] (2) Comparison of key indicators:
[0193]
[0194] (3) Experimental conclusions:
[0195] Accuracy Improvement: This approach improves accuracy by 16.7% compared to the static thresholding method and by 10.2% compared to the LSTM model, thanks to the ability of the 3D aggregation analysis and hybrid detection model to capture multi-dimensional features.
[0196] Reduced false alarm rate: The false alarm rate is reduced by 58% compared to the static threshold method (from 12.3% to 2.1%). The dynamic threshold mechanism effectively filters false alarms caused by business fluctuations.
[0197] Real-time advantage: The average detection latency is only 120ms, which meets the millisecond-level response requirements of the power grid (the latency of traditional methods is all >300ms), thanks to the lightweight probe and adaptive parameter update algorithm.
[0198] 1.5.3 Innovation Points in Power Grid Scenarios
[0199]
[0200] The embodiments of the present invention are not limited to those described above. The baseline value of the dynamic sampling period (e.g., 50ms) and the service period parameter T in the data acquisition layer can be adjusted according to the actual network security scenario. cycle The values of the following, including the sliding window duration T in the aggregation analysis layer, the coefficient ratio (e.g., 0.6:0.4) for calculating the edge weights of the terminal relationship graph, and the number of protocol types K for the multi-protocol association entropy, the values of the dynamic weight factor α of the hybrid model in the dynamic detection layer, the parameters k, λ, and γ of the adaptive threshold function, and the ratio of linear regression coefficients and random forest feature weights in the anomaly scoring model, all fall within the protection scope of this invention.
Claims
1. A terminal network information processing and detection method based on aggregation analysis, characterized in that, include: The data acquisition layer collects network data from terminal devices by deploying lightweight probe agents, and processes the data in a standardized D format. i ={t,src_ip,dst_ip,proto,payload_len,flag_bits} is stored, and the sampling period is dynamically adjusted according to the number of connections. Where N is the current number of connections; The aggregation analysis layer calculates the feature vector F from the time dimension using a 1-second sliding window. T =[∑len,σ(len),H(proto),H(port)], constructing a terminal relationship graph from a spatial dimension and according to... Calculate edge weights from the multi-protocol association entropy dimension. Extract cross-protocol layer features; The dynamic detection layer utilizes a hybrid model of linear regression and random forest, with the formula Score = α·(β0 + β1F1 + ... + β). n F n Anomaly scores are calculated using the formula )+(1-ɑ)·RF(F), and then passed through a dynamic threshold function. Real-time adjustment of detection thresholds enables the identification of abnormal network behavior.
2. The method according to claim 1, characterized in that, The lightweight probe agent adopts a modular architecture, has cross-operating system compatibility, can automatically detect the operating system type and network environment of the terminal device and perform adaptive configuration, and has a resource utilization rate of less than 5%.
3. The method according to claim 1, characterized in that, In the time-dimensional feature vector, ∑len is the total data load length within the window, σ(len) is the standard deviation of the data load length, and H(proto) and H(port) are the information entropy of the protocol type and port number, respectively. calculate.
4. The method according to claim 1, characterized in that, In the terminal relationship graph of the spatial dimension, nodes are terminal devices, edges represent network communication between devices, and edge weights are w. ij The value ranges from [0,1]. The closer the value is to 0, the higher the overlap of the IP address sets of the two devices.
5. The method according to claim 1, characterized in that, In the multi-protocol association entropy, K is the number of protocol types, and n k Let T be the number of times protocol k appears, N be the total number of times all protocols appear, and T be the number of times protocol k appears. k For the duration of protocol k, T total This represents the total length of the observation period.
6. The method according to claim 1, characterized in that, In the hybrid model of the dynamic detection layer, α∈[0,1] is the dynamic weighting factor, based on the flow fluctuation coefficient C. f and characteristic nonlinearity N f Adjust in real time, when C f ≤1.5 and N f When ≤0.8, α=α base +Δα·sigmoid(t); when C f >1.5 or N f When >0.8, 7. The method according to claim 1, characterized in that, In the dynamic threshold function, S(t) is the current anomaly score, S0 is the median of historical anomaly scores, k is the slope parameter, and θ is the slope parameter. min and θ max These are the minimum and maximum values of the threshold, respectively. λ and γ are adaptive parameters, and the parameters k, λ, and γ can be dynamically adjusted online according to changes in the network topology.
8. The method according to claim 1, characterized in that, In smart grid scenarios, the standardized format of the data acquisition layer is extended to D. i ={t,src_ip,dst_ip,proto,payload_len,flag_bits,device_type,area_code}, sampling period adjusted to Where ${T}_{cycle}$ represents the power grid operating cycle.
9. The method according to claim 8, characterized in that, In smart grid scenarios, the time dimension feature vector of the aggregation analysis layer is augmented with a heartbeat anomaly feature (heartbeat_loss), and the spatial dimension edge weights are adjusted to... In the multi-protocol association entropy, the number of protocol types K = 3, corresponding to IEC61850, Modbus, and DNP3 protocols.
Citation Information
Cited By
Communication data analysis system and method for network security
CN121333798A