Power system network asset dynamic sensing method based on full flow analysis

By using full-flow analysis and machine learning, a multi-dimensional communication behavior feature vector is constructed, which solves the problem of dynamic perception and anomaly detection of power system network assets, and realizes real-time, full-coverage, interference-free dynamic perception and high-precision anomaly detection of power system network assets.

CN122053170APending Publication Date: 2026-05-15ZHANGZHOU POWER SUPPLY COMPANY STATE GRID FUJIANELECTRIC POWER +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHANGZHOU POWER SUPPLY COMPANY STATE GRID FUJIANELECTRIC POWER
Filing Date
2026-02-13
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve real-time, full-coverage, and interference-free dynamic perception of power system network assets, and traditional methods cannot adapt to the dynamic changes and abnormal behavior detection of network assets.

Method used

By conducting full traffic analysis, deploying network traffic collection devices, performing session reconstruction and multi-protocol parsing, constructing multi-dimensional communication behavior feature vectors, and combining machine learning to establish network asset behavior baselines for continuous monitoring and anomaly detection.

Benefits of technology

It achieves non-interference, full coverage, and dynamic perception of power system network assets, reduces management complexity, improves the accuracy and real-time performance of anomaly detection, and enhances the level of intelligence in security monitoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053170A_ABST
    Figure CN122053170A_ABST
Patent Text Reader

Abstract

The invention discloses a power system network asset dynamic sensing method based on full-flow analysis, and aims to solve the problems of large scale and frequent change of power system network assets, high maintenance cost of a traditional ledger and incomplete sensing. The method comprises the following steps: deploying a flow collection device at a network-related side of a power system, and carrying out mirror image collection on total network flow; performing session reconstruction and multi-protocol analysis on the flow, and extracting quintuple information and application layer communication behavior characteristics; identifying network assets based on communication behaviors, and constructing asset feature vectors; establishing an asset behavior baseline through continuous statistical analysis; and performing deviation calculation on the real-time flow behavior and the baseline to realize dynamic perception and alarm of the abnormal behavior. The method supports multi-source data association analysis, and an advanced machine learning model can be adopted to improve the detection precision and the adaptive ability. The method does not need to depend on a static asset list, can achieve the continuous discovery and abnormality monitoring of the network assets of the power system, and improves the active defense capability of network security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power system network security and traffic analysis technology, and in particular to a dynamic perception method for power system network assets based on full traffic analysis. Background Technology

[0002] With the rapid development of new power systems, the reliance of power systems on communication networks in areas such as dispatch automation, substation integrated automation, and new energy grid integration is constantly deepening. The number of network assets continues to grow, the types of assets are complex, and their operating status and communication behavior exhibit highly dynamic characteristics.

[0003] Currently, power system network asset management mainly relies on the following two methods:

[0004] (1) Manual maintenance of asset ledgers: Information such as IP address, device type, and region of network assets is recorded manually. This method is slow to update, prone to omissions, and cannot reflect the online status and communication behavior changes of assets in real time. In large power networks, the number of assets is huge, manual maintenance is costly, and it is difficult to guarantee accuracy.

[0005] (2) Periodic Active Scanning: Network scanning tools (such as Nmap) are used to periodically probe the network to discover online assets. However, active scanning will generate additional traffic load on the network, which may interfere with normal business communication. Especially in power production control areas, frequent scanning may cause network latency or even equipment malfunction, which does not meet the power system's requirements for high availability and real-time performance. In addition, active scanning can only obtain basic information about assets (such as open ports, operating systems, etc.) and cannot obtain the continuous communication behavior characteristics of assets, so it is difficult to build a behavioral profile of assets.

[0006] Meanwhile, traditional security monitoring technologies often focus on single-point alarms or matching based on fixed rules, lacking continuous modeling and analysis of the overall situation of asset communication behavior. For example, traditional intrusion detection systems (IDS) are typically based on rule bases of known attack characteristics and cannot detect new attacks or abnormal behaviors. In addition, these technologies often require pre-configuration of asset information and cannot adapt to dynamic changes in assets. Summary of the Invention

[0007] In view of this, the purpose of this invention is to provide a dynamic perception method for power system network assets based on full flow analysis, which is based on real network traffic, non-interference, adaptive, and able to distinguish between normal evolution and malicious behavior, so as to improve the automation and intelligence level of network security monitoring.

[0008] To achieve the above objectives, the present invention adopts the following technical solution: a dynamic sensing method for power system network assets based on full flow analysis, comprising the following steps:

[0009] S1: Deploy network traffic acquisition devices on the grid-connected side of the power system to mirror and collect network traffic behavior data from the dispatching master station, substations, power plants, new energy power stations and their associated communication networks;

[0010] S2: Perform session reconstruction and multi-protocol parsing on the collected network traffic behavior data, extract the five-tuple information including source address, destination address, source port, destination port and transport layer protocol, and obtain the corresponding application layer protocol type and communication behavior characteristics.

[0011] S3: Based on the five-tuple information and communication behavior characteristics, identify the actual communication endpoints in the network, construct a network asset set, and establish a corresponding communication behavior feature vector for each asset to realize the dynamic discovery and identification of assets.

[0012] Identifying actual communication endpoints in the network specifically includes: First, asset discovery does not rely solely on traditional IP address scanning or port activity assessment. Instead, it intelligently identifies all active endpoints participating in communication by performing session reconstruction and deep multi-protocol analysis on the full traffic data collected from mirroring. Each unique IP address or device identifier that generates or receives network traffic is defined as a network asset, thus constructing an initial set of network assets. This step ensures comprehensive and real-time asset discovery coverage.

[0013] More importantly, there is the subsequent feature vector construction. This invention dynamically constructs a multi-dimensional communication behavior feature vector Fi=[Pi,Di,Ti,Ai] for each identified asset. This vectorized representation is one of the core innovations of this invention. It does not simply list asset attributes, but rather creates a unique, computable "behavioral DNA" for each asset by quantifying its historical and real-time communication behavior.

[0014] This feature vector integrates multi-dimensional information such as protocol preferences, social contacts, daily routines, and application access patterns, enabling precise characterization of an asset's role (e.g., server, workstation, industrial control equipment) and behavioral patterns within the network. This asset identification and classification method based on behavioral profiling overcomes the limitations of traditional ledgers and proactive scanning, which only provide static information. It achieves dynamic perception and understanding of the asset's intrinsic functions and business status, laying a solid foundation for subsequent abnormal behavior detection based on machine learning. Therefore, this process as a whole constitutes a significantly innovative technical solution.

[0015] Specifically, the dynamic discovery and identification of assets includes:

[0016] (1) Initial discovery and asset registration

[0017] The system automatically identifies all unique endpoints (primarily IP addresses) involved in communication by reconstructing the entire traffic session. Each discovered endpoint is initialized as a potential "network asset," and a unique profile is created for it within the system. This enables a zero-based start-up from "unknown" to "known," without requiring any pre-configured asset inventory.

[0018] (2) Behavioral feature extraction and vectorized modeling (core innovation)

[0019] This is the core of claim 3 and the key to achieving "dynamic identification." The system does not statically record asset information, but rather continuously and dynamically constructs and updates the communication behavior feature vector Fi = [Pi, Di, Ti, Ai] for each asset. This process is deep and multi-dimensional:

[0020] Protocol distribution characteristics (Pi): Quantifies the asset's preference for different protocols during communication. For example, a web server may exhibit a very high concentration of HTTP / HTTPS protocols.

[0021] Communication Object Distribution Characteristics (Di): Characterizes the "social" scope of an asset through information entropy or maximum concentration. A gateway that communicates with a large number of internal devices has a high entropy value; while an industrial control device that only communicates with a specific server has a high concentration.

[0022] Communication Period Characteristics (Ti): Establish the asset's activity pattern within 24 hours, clearly distinguishing between office terminals that are only active during working hours and monitoring servers that need to run 24 hours a day.

[0023] Application Access Characteristics (Ai): Using Embedding technology, discrete sequences such as URLs, APIs, and control point tables for asset access are transformed into continuous vectors, deeply representing their business behavior patterns.

[0024] (3) Asset classification and role identification

[0025] Based on the rich feature vectors constructed in step two, the system can use unsupervised clustering (such as K-means, DBSCAN) or supervised classification algorithms to automatically classify assets into different types or roles (e.g., database servers, engineer workstations, PLC controllers, video surveillance equipment, etc.). This behavior-based classification is far more accurate and dynamic than static classification based on IP address ranges, and can automatically identify newly added or unknown types of devices in the network.

[0026] (4) Dynamic tracking and baseline self-learning

[0027] An asset's identity is not static. The system continuously monitors changes in the feature vector of each asset.

[0028] Discovering new assets: When a previously unseen IP address begins communication, the system immediately creates a profile for it and begins building a feature vector to achieve dynamic discovery.

[0029] Identifying changes in asset status: If the communication period (Ti) of an asset suddenly changes from working hours to 24 / 7, or if an unknown protocol appears in its protocol distribution (Pi), these changes will be immediately reflected in the deviation of the feature vector, triggering subsequent anomaly detection.

[0030] Baseline Evolution: The asset's behavioral baseline Bi is also dynamic, and it will learn and adjust itself according to the long-term trend of the feature vector to adapt to normal business evolution.

[0031] In summary, the "Dynamic Asset Discovery and Identification" of this invention is an intelligent process that automatically learns from communication data and constructs an asset behavior profile. Through an innovative multi-dimensional feature vector model, it transforms the dynamic communication behavior of assets into a calculable and comparable digital representation. This not only enables the discovery of the "existence" of assets but also achieves deep identification of their "business role" and "normal behavior," laying a solid foundation for subsequent high-precision anomaly detection.

[0032] S4: Continuously perform statistical analysis and machine learning modeling on the aforementioned communication behavior feature vectors to construct a network asset behavior baseline. This baseline includes at least the asset online time baseline, network session baseline, and application access baseline. The core of the continuous statistical analysis lies in using a multi-layered, intelligent analytical framework to deeply mine historical behavior feature vectors. This process first uses descriptive statistics (such as calculating the mean, standard deviation, and probability distribution of each feature dimension) to characterize the normal level and normal fluctuation range of asset behavior, laying the data foundation for the baseline. Based on this, time series analysis (such as ARIMA and LSTM models) is introduced to capture and predict the periodicity and trend of asset activities, enabling the baseline to predictively adapt to the normal evolution of behavior. Furthermore, multivariate correlation modeling (such as calculating the covariance matrix and Mahalanobis distance) is used to analyze the intrinsic correlation between different behavioral characteristics (such as protocol type and communication time), thereby identifying anomalies that are not obvious in a single dimension but significant under multi-dimensional combined analysis. Ultimately, by combining unsupervised learning (such as clustering and density estimation), "normal behavior clusters" formed in historical data are automatically discovered, intelligently defining the boundary between normal and abnormal. The comprehensive application of this set of statistical machine learning methods ensures that the constructed behavioral baseline is not a static threshold, but an intelligent model that can comprehensively and dynamically represent the normal behavior patterns of assets and learn itself as the network evolves, providing a core basis for high-precision, low-false-positive anomaly detection.

[0033] S5: The deviation between the real-time collected network traffic behavior and the baseline of the network asset behavior is calculated. When the deviation exceeds a preset threshold, it is determined to be abnormal behavior and an alarm message is generated, so as to realize the dynamic perception and monitoring of abnormal behavior of power system network assets.

[0034] The deviation calculation in step S5 is a quantitative comparison process based on statistical learning. The system first analyzes the real-time network traffic to generate the asset's current behavior feature vector. The behavioral baseline model built with its historical data (usually characterized as a multivariate normal distribution) A comparison is performed. The core calculation uses the standardized Euclidean distance formula: This formula is obtained by dividing by the standard deviation of each dimension of features. Data standardization is achieved to eliminate differences in the dimensions of different features and to calculate the overall deviation. Subsequently, this deviation value will be compared with a dynamic threshold. (such as the 95th percentile of historical bias or The system compares the deviations with other data. To improve accuracy, the system also performs multi-level verification: it checks the persistence of deviations in the time dimension, analyzes the coordination anomalies of related assets in the spatial dimension, and only significant deviations that pass multiple verifications will trigger alarms with clear business semantics, thus achieving accurate perception.

[0035] In summary, the system collects and analyzes the current behavior feature vectors of network assets in real time and compares them quantitatively with a dynamic behavior baseline model built through learning from historical data. This comparison is not a simple numerical subtraction, but uses statistical metrics such as standardized Euclidean distance to calculate the comprehensive deviation of the current behavior from historical normal patterns. This deviation value is then compared with a dynamically set threshold that is adaptively set based on the asset's own historical fluctuation patterns. When the deviation significantly exceeds the threshold, an anomaly detection is triggered, and an alarm message is generated. This method can effectively detect subtle changes in behavior patterns, thereby achieving accurate and dynamic monitoring of abnormal states of power system network assets.

[0036] In a preferred embodiment, the multi-protocol resolution in step S2 includes at least HTTP, DNS, FTP, RDP, Telnet protocols, and general transport layer protocols based on TCP or UDP.

[0037] In a preferred embodiment, the communication behavior feature vector in step S3 is represented as:

[0038]

[0039] in, This represents the behavioral feature vector of the i-th network asset. Indicates the characteristics of protocol distribution. Indicates the distribution characteristics of communication objects. Indicates the characteristics of communication time periods. This indicates the application's access characteristics.

[0040] In a preferred embodiment, each dimension of the feature vector is further quantized using the following formula:

[0041] (1) Protocol distribution characteristics:

[0042]

[0043] Where m represents the total number of protocol types that can be parsed within the monitoring range. , 1≤j≤m;

[0044] (2) Distribution characteristics of communication objects:

[0045]

[0046] or

[0047]

[0048] in For asset i, the set of communication objects. The percentage of communication with object o;

[0049] (3) Characteristics of communication time periods:

[0050]

[0051] in, .

[0052] (4) Application access characteristics:

[0053]

[0054] in For accessing resource identifiers, embedding is implemented through Word2Vec, BERT pre-trained models, or domain-specific sequence encoders, transforming discrete resource identifier sequences into continuous vector representations.

[0055] In a preferred embodiment, the network asset behavior baseline described in step S4 is constructed in the following manner:

[0056]

[0057] in, This represents the baseline model of the behavior of the i-th network asset. This represents the behavioral feature vector of the asset during the k-th statistical period. This represents the total number of statistical periods used to construct the baseline.

[0058] i represents the i-th network asset, which is consistent with the entire document. The convention is consistent, used to uniquely identify a specific asset in the network (such as a specific server, workstation, or industrial control equipment); k represents the k-th statistical period. When building the baseline, the system does not use data from only one point in time, but continuously collects data and divides it into consecutive, equally long statistical time windows (e.g., each window is 1 hour or 1 day long). k is the sequence number of these time windows. Overall :therefore, Specifically, it refers to the communication behavior feature vector exhibited by the i-th network asset within the k-th statistical period. It is a "snapshot" of the asset's behavior obtained after statistically analyzing all network traffic within that time period (e.g., within the k-th hour).

[0059] Therefore, set This represents all historical data used to build the baseline model. It contains a series of behavioral feature vectors of the i-th asset, arranged chronologically from the 1st to the Nth statistical period. Baseline construction function. The task is precisely to extract data from this time series dataset. Learn from and extract the normal behavior patterns of this asset to form a baseline. This representation clearly demonstrates that the construction of the behavioral baseline is based on continuous observation data over a historical period, rather than the state at a single moment, which is the foundation for achieving dynamic and adaptive perception.

[0060] In a preferred embodiment, the network asset behavior baseline has self-learning and self-revision capabilities, triggering baseline revision when any of the following conditions are met:

[0061] The self-learning and self-revision mechanism of the network asset behavior baseline is a key intelligent process that ensures the model continuously adapts to dynamic changes in the network. Its revision is not a simple reset, but rather a gradual optimization and conditional reconstruction of the original baseline model based on new observation data. The specific revision method varies depending on the triggering conditions:

[0062] For newly added assets or assets that have been offline for a long time and are now back online: the system will initiate an initial learning period. During this period, the system will collect behavioral data of the asset over several statistical periods, directly build its initial behavioral baseline based on this new data, or quickly integrate its behavioral patterns into the existing baseline model through a high learning rate.

[0063] When significant changes occur in business application types or long-term network traffic trends, the behavior patterns of assets have fundamentally shifted. The system employs a time-decay weighting strategy, assigning higher weights to recent data while gradually reducing the influence of older historical data in the baseline model. This can be achieved by adjusting the forgetting factor (λ) in the incremental learning algorithm, for example, in the EWMA (Exponentially Weighted Moving Average) model. In the middle, the λ value is temporarily increased to enable the baseline to track new behavioral trends more quickly.

[0064] Model Reconstruction and Validation: When changes are very drastic, the system can automatically activate a sliding window mechanism to retrain the baseline model based only on the data of the most recent N periods (i.e., recalculate the mean vector μ_i and covariance matrix Σ_i), and use the validation set to evaluate the performance of the new model, ensuring that the revised baseline reflects the new patterns while maintaining stability.

[0065] The core of this self-revision mechanism lies in automatically judging the nature and extent of behavioral changes through algorithms and selecting the most appropriate revision strategy, so that the baseline model can evolve smoothly, adapting to normal business development and effectively avoiding model degradation caused by short-term noise or malicious traffic pollution.

[0066] (1) New network assets have been detected;

[0067] (2) Network assets have been offline for an extended period or need to be re-enabled;

[0068] (3) The type of business application has changed;

[0069] (4) The long-term trend of network traffic has changed significantly.

[0070] The present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory, wherein when the program is executed by the processor, it implements the dynamic sensing method for power system network assets based on full flow analysis as described above.

[0071] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the power system network asset dynamic sensing method based on full flow analysis as described above.

[0072] Compared with the prior art, the present invention has the following beneficial effects:

[0073] 1. Full traffic coverage and in-depth analysis: By deploying traffic acquisition devices for mirrored collection and combining multi-protocol in-depth analysis, it achieves non-interference and full-coverage discovery of power system network assets, overcoming the interference of active scanning on the production network and the blind spots of traditional ledgers.

[0074] 2. Dynamic Behavior Modeling: Based on communication behavior characteristics, an asset behavior baseline is constructed, which can adapt to network changes, realize continuous perception of asset status and anomaly detection, and improve the real-time performance and accuracy of security monitoring.

[0075] 3. Dependency-free automated discovery: It eliminates the need for static asset lists or manual maintenance, significantly reducing the complexity and cost of asset management, and is suitable for large-scale, highly dynamic power system network environments.

[0076] 4. Intelligent anomaly detection: Through machine learning and behavior analysis technology, early detection and intelligent alarm of abnormal communication behavior are achieved, enhancing the proactive defense and emergency response capabilities of the power system.

[0077] 5. Scalability and Foresight: It supports in-depth analysis of emerging protocols and power-specific protocols, and can integrate multi-source data for correlation verification, improving the applicability and detection reliability of the method in new power systems. Attached Figure Description

[0078] Fig. 1 A flowchart of a dynamic perception method for power system network assets based on full flow analysis provided in an embodiment of the present invention;

[0079] Fig. 2 This is a schematic diagram of the network asset communication behavior feature vector in an embodiment of the present invention;

[0080] Fig. 3 This is a schematic diagram illustrating the construction of network asset behavior baselines and anomaly detection in an embodiment of the present invention. Detailed Implementation

[0081] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0082] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.

[0083] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations according to this application; as used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise; furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof.

[0084] A dynamic sensing method for power system network assets based on full flow analysis, referencing Figs. 1-3 This includes the following steps:

[0085] Step S1: Full Traffic Acquisition

[0086] Traffic acquisition devices are deployed at key network nodes in dispatching master stations, substations, power plants, and new energy power stations. Through port mirroring or network splitting, all network traffic entering and leaving these nodes is collected without disturbance, ensuring the integrity and timeliness of traffic coverage.

[0087] Step S2: Multiprotocol parsing and session reconstruction

[0088] The system performs protocol identification and decoding on the collected raw traffic, supporting common application protocols including HTTP, DNS, FTP, RDP, and Telnet, as well as industrial control protocols such as IEC 104 and Modbus, IoT protocols such as MQTT and CoAP, and core protocols for smart substations such as the IEC 61850 series. Through session reconstruction technology, discrete data packets are reassembled into complete communication sessions, extracting five-tuple information and communication behavior characteristics, such as request-response patterns, data packet size distribution, and communication frequency. The multi-protocol parsing not only identifies protocol types but also performs deep decoding of application layer payloads to extract rich communication behavior characteristics. For example, for the HTTP protocol, fields such as request method (GET / POST / PUT, etc.), URI path, User-Agent, and Host can be extracted; for the DNS protocol, query domain name, record type, and response code can be extracted; for industrial control protocols (such as IEC 104 and Modbus), specific semantic information such as function codes, data object addresses, and control command values ​​can be extracted; for the IEC 61850 protocol, key semantic information such as Manufacturing Message Specification (MMS) service, General Object-Oriented Substation Event (GOOSE) message status, and Sampled Value (SV) data stream can be further parsed and extracted. These fine-grained features collectively form the basis for subsequent behavioral modeling.

[0089] Step S3: Dynamic Asset Identification and Feature Vector Construction

[0090] Based on session data, active communication endpoints (IP addresses, device identifiers, etc.) in the network are identified and defined as network assets. For each asset, its communication behavior features are extracted, and a multi-dimensional feature vector is constructed, formally represented as follows:

[0091]

[0092] in, This represents the behavioral feature vector of the i-th network asset. Indicates the characteristics of protocol distribution. Indicates the distribution characteristics of communication objects. Indicates the characteristics of communication time periods. This represents application access characteristics. To further characterize the roles and relationships of assets within the network topology, graph embedding features can be optionally introduced. This is achieved by learning a network graph with assets as nodes and communication relationships as edges using a graph neural network (GNN). Through clustering or classification algorithms, asset types (such as servers, workstations, industrial control equipment, etc.) can be identified.

[0093] To more specifically characterize asset behavior, each dimension of the feature vector can be further quantified using the following formula:

[0094] (1) Protocol distribution characteristics (discrete probability distribution):

[0095]

[0096] Where m represents the total number of protocol types that can be parsed within the monitoring range. .

[0097] (2) Distribution characteristics of communication objects (based on information entropy or concentration):

[0098]

[0099] or

[0100]

[0101] in For asset i, the set of communication objects. This represents the percentage of communication with object o.

[0102] (3) Communication time period characteristics (24-hour active mode):

[0103]

[0104] in, .

[0105] (4) Apply access characteristics (behavioral sequence vectors):

[0106]

[0107] in For accessed resource identifiers (such as URLs, API endpoints, and industrial control point tables), embedding can be implemented using pre-trained models such as Word2Vec and BERT, or sequence encoders trained for specific domains, to transform discrete resource identifier sequences into continuous vector representations.

[0108] Step S4: Behavioral Baseline Construction

[0109] A time-sliding window mechanism is used to statistically analyze the historical behavioral characteristics of each asset and establish its behavioral baseline model. The baseline includes:

[0110] 1) Online time baseline: the time pattern of asset activity;

[0111] 2) Network session baseline: Statistical distribution of communication objects, ports, and protocol combinations;

[0112] 3) Application access baseline: Typical characteristics of resource access and request patterns.

[0113] The baseline model supports self-learning updates. When changes in asset status are detected (such as addition, offline, or business changes), baseline revision is automatically triggered to ensure the timeliness of the model.

[0114] The behavioral baseline The construction can employ one or more of the following machine learning and statistical methods:

[0115] Statistical model: Calculate the historical mean of each dimension of the feature vector. With covariance matrix The baseline is defined as a multivariate normal distribution N( , Anomalies are measured using Mahalanobis distance.

[0116] Time series model: Establish an ARIMA or LSTM prediction model for time series features (such as communication frequency), with the baseline being the predicted value and its confidence interval.

[0117] Unsupervised clustering: Clustering historical feature vectors (such as DBSCAN), with the baseline defined by the core clusters, and anything falling outside the cluster is considered an anomaly.

[0118] Deep learning models: For complex nonlinear behavior patterns, deep autoencoders (DAEs) or generative adversarial networks (GANs) can be used to learn the data distribution of normal behavior and reconstruct the error or discriminator output as anomaly scores.

[0119] Graph learning models: To capture interaction patterns between assets, dynamic communication graphs can be constructed, and graph neural networks (GNNs) or graph autoencoders can be used to learn normal representations of the graph. Abnormal changes in graph structure or node features indicate threats.

[0120] To enhance the robustness of the behavioral baseline model and prevent it from being contaminated by deliberate low-rate spoofed traffic, an adversarial training mechanism can be introduced during the model training process.

[0121] The model supports incremental learning, where new data is incorporated into historical statistics with certain weights to achieve a smooth evolution of the baseline. A specific dynamic baseline construction method is based on Exponentially Weighted Moving Average (EWMA), as shown in the following formula:

[0122]

[0123]

[0124] in, and Let represent the mean and standard deviation estimates of the behavioral characteristics of asset i at the t-th update time; λ is the forgetting factor, with a value range of (0,1), preferably 0.85 to 0.99, used to control the weight decay rate of historical data. The larger λ is, the more sensitive it is to new data.

[0125] The baseline behavior at time t is:

[0126]

[0127] Step S5: Abnormal Behavior Detection and Alarm

[0128] Real-time collected network traffic is parsed to generate current behavioral characteristics, which are then compared with the behavioral baseline of the corresponding asset (e.g., cosine similarity, Mahalanobis distance, etc.). If the deviation exceeds a preset threshold, it is determined to be abnormal behavior, an alarm is generated, and abnormal context information is recorded for security personnel to analyze and handle.

[0129] Anomaly detection employs a multi-level threshold and correlation analysis strategy:

[0130] (a) Initial determination: Calculate the real-time feature vector Compared with baseline deviation value .like ( If the threshold is dynamic, a primary alarm will be triggered.

[0131] (b) Correlation Verification: Examine the context of the asset in terms of time (e.g., whether similar deviations have occurred recently) and space (e.g., whether related assets also show anomalies). A momentary exceedance in a single dimension may be caused by business fluctuations, while multi-dimensional and continuous correlation anomalies are highly likely to be a real threat. Furthermore, network traffic anomalies can be matched with external threat intelligence databases (e.g., CVE, NVD), or correlated with device system logs, security events, and other information to build a cross-data source evidence chain, thereby significantly improving the confidence of alerts.

[0132] (c) Alarm Generation: The final alarm information shall include at least: asset identifier (IP / hostname), anomaly type (protocol / object / time period / access anomaly), deviation value, related evidence, time window, and recommended handling level. The alarm information may include a feature contribution analysis report from the model interpretation module.

[0133] The deviation value The calculation can be performed using the standardized Euclidean distance:

[0134]

[0135] in, Let be the abnormal deviation value of asset i at time t; and let n be the behavioral feature vector. The total number of dimensions; Let be the value of the j-th dimension of the feature vector at time t; and denoted as the mean and standard deviation of the j-th feature in the baseline model, respectively.

[0136] Dynamic threshold It can be adaptively set according to the historical deviation distribution:

[0137]

[0138] or

[0139]

[0140] in, The dynamic alarm threshold for asset i; This is the average of all historical deviations of the asset during the baseline learning phase. The sample standard deviation of the historical deviation values; This is an adjustable sensitivity coefficient, typically set to 2 or 3, representing the tolerance for deviations exceeding the mean by several times the standard deviation. The dynamic threshold... The optimization process can employ automated hyperparameter tuning methods such as Bayesian optimization to adaptively find the optimal balance between false positive rate and false negative rate.

[0141] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A dynamic sensing method for power system network assets based on full flow analysis, characterized in that, Includes the following steps: S1: Deploy network traffic acquisition devices on the grid-connected side of the power system to mirror and collect network traffic behavior data from the dispatching master station, substations, power plants, new energy power stations and their associated communication networks; S2: Perform session reconstruction and multi-protocol parsing on the collected network traffic behavior data, extract the five-tuple information including source address, destination address, source port, destination port and transport layer protocol, and obtain the corresponding application layer protocol type and communication behavior characteristics. S3: Based on the five-tuple information and communication behavior characteristics, identify the actual communication endpoints in the network, construct a network asset set, and establish a corresponding communication behavior feature vector for each asset to realize the dynamic discovery and identification of assets. S4: Perform continuous statistical analysis and machine learning modeling on the communication behavior feature vector to construct a network asset behavior baseline, which includes at least an asset online time baseline, a network session baseline, and an application access baseline. S5: The deviation between the real-time collected network traffic behavior and the baseline of the network asset behavior is calculated. When the deviation exceeds a preset threshold, it is determined to be abnormal behavior and an alarm message is generated, so as to realize the dynamic perception and monitoring of abnormal behavior of power system network assets.

2. The method for dynamic sensing of power system network assets based on full flow analysis according to claim 1, characterized in that, The multi-protocol resolution in step S2 includes at least HTTP, DNS, FTP, RDP, Telnet protocols, and general transport layer protocols based on TCP or UDP.

3. The method for dynamic sensing of power system network assets based on full flow analysis according to claim 1, characterized in that, The communication behavior feature vector described in step S3 is represented as follows: in, This represents the behavioral feature vector of the i-th network asset. Indicates the characteristics of protocol distribution. Indicates the distribution characteristics of communication objects. Indicates the characteristics of communication time periods. This indicates the application's access characteristics.

4. The method for dynamic sensing of power system network assets based on full flow analysis according to claim 3, characterized in that, The feature vectors are further quantized into the following formulas: (1) Protocol distribution characteristics: Where m represents the total number of protocol types that can be parsed within the monitoring range. , 1≤j≤m; (2) Distribution characteristics of communication objects: or in For asset i, the set of communication objects. The percentage of communication with object o; (3) Characteristics of communication time periods: in, . (4) Application access characteristics: in For accessing resource identifiers, embedding is implemented through Word2Vec, BERT pre-trained models, or domain-specific sequence encoders, transforming discrete resource identifier sequences into continuous vector representations.

5. The method for dynamic sensing of power system network assets based on full flow analysis according to claim 1, characterized in that, The network asset behavior baseline described in step S4 is constructed in the following manner: in, This represents the baseline model of the behavior of the i-th network asset. This represents the behavioral feature vector of the asset during the k-th statistical period. This represents the total number of statistical periods used to construct the baseline.

6. The method for dynamic sensing of power system network assets based on full flow analysis according to claim 4, characterized in that, The network asset behavior baseline has self-learning and self-revision capabilities, and baseline revision is triggered when any of the following conditions are met: (1) New network assets have been detected; (2) Network assets have been offline for an extended period or need to be re-enabled; (3) The type of business application has changed; (4) The long-term trend of network traffic has changed significantly.

7. An electronic device comprising a processor, a memory, and a computer program stored in the memory, characterized in that, When the program is executed by the processor, it implements a dynamic perception method for power system network assets based on full flow analysis as described in any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of a dynamic perception method for power system network assets based on full flow analysis as described in any one of claims 1 to 6.