Power network intrusion analysis method and system based on data fusion

By deeply integrating multi-source monitoring data from the power network, a multi-dimensional fusion feature set is generated to identify complex attacks and respond in real time. This solves the problems of low detection accuracy and recovery efficiency in traditional power network intrusion analysis, and achieves efficient attack suppression and rapid system security reconstruction.

CN120567472BActive Publication Date: 2025-12-26STATE GRID SHANDONG ELECTRIC POWER CO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510688701.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-12-26
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Traditional power grid intrusion analysis methods struggle to identify complex attack methods, cannot simultaneously eliminate threats from multiple nodes along the attack chain, and lack mechanisms to repair the operational status of damaged equipment, resulting in long system recovery cycles and high secondary risks. Furthermore, the fixed-frequency data acquisition mode leads to incomplete feature extraction and data redundancy.

Method used

By deeply integrating and monitoring the operational data streams of power equipment nodes with network communication data streams, a multi-dimensional fusion feature set is generated. An abnormal behavior identification model is used for intrusion analysis, and the monitoring data acquisition strategy is dynamically adjusted in conjunction with real-time blocking and node status repair commands.

Benefits of technology

It achieves accurate detection of physical attack characteristics and network attack behaviors at the device operation level, simultaneously blocks attack paths and restores device status, improving intrusion detection accuracy and reducing system resource consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120567472B_ABST
    Figure CN120567472B_ABST
Patent Text Reader

Abstract

The application provides a power network intrusion analysis method and system based on data fusion, which comprises the following steps: acquiring a multi-source monitoring data set in a power network, performing dynamic feature fusion processing on the multi-source monitoring data set, generating a multi-dimensional fusion feature set associated with a power equipment node, inputting the multi-dimensional fusion feature set into a preset abnormal behavior recognition model for intrusion analysis processing, generating an abnormal behavior data set of the power equipment node, generating a network security defense strategy set according to the abnormal behavior data set, the network security defense strategy set comprising real-time blocking instructions and node state repair instructions for different abnormal behavior types, feeding back the network security defense strategy set to a power network control center to trigger a defense response operation, and adjusting the collection strategy of the multi-source monitoring data set according to the execution result of the defense response operation. Through the application, the intrusion detection accuracy is improved while reducing the system resource consumption.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of data processing, in particular to a power network intrusion analysis method and system based on data fusion. BACKGROUND

[0002] With the improvement of the intelligence degree of the power system, the power network intrusion analysis technology has become the core link to ensure energy security. The traditional power safety monitoring technology mainly adopts an independent monitoring mode: on the one hand, the sensors deployed at the power device nodes collect device operating parameters such as voltage fluctuation and current harmonic, and identify device overload or line fault by using threshold comparison; on the other hand, network traffic monitoring devices capture network behavior data such as communication protocol type and port access record, and detect conventional network attack behaviors by matching a rule library. However, such isolated analysis method is difficult to identify new composite attack means: the attacker triggers device operating parameter abnormalities by tampering with the communication protocol, or uses device overload state to cover malicious port scanning behavior, so that the correlation characteristics of device physical layer abnormalities and network layer attack behaviors are fragmented and detected. The response strategy for the detection results in the prior art is limited to single network connection blocking or device restart operation, which cannot simultaneously eliminate the multi-node threat on the attack link, and lacks a repair mechanism for the operating state of the damaged device, resulting in a long recovery period and high secondary risk after the system is attacked. In addition, the fixed frequency device data collection and network traffic packet capture mode cannot dynamically adapt to the feature extraction needs in complex attack scenarios, causing the analysis dilemma of missing key attack features and coexistence of redundant data. SUMMARY

[0003] The present application provides a power network intrusion analysis method and system based on data fusion.

[0004] In a first aspect, the embodiments of the present application provide a power network intrusion analysis method based on data fusion, comprising the following steps:

[0005] Obtain a multi-source monitoring data set in a power network, the multi-source monitoring data set comprising node operating data streams and network communication data streams of a plurality of power device nodes;

[0006] Perform dynamic feature fusion processing on the multi-source monitoring data set to generate a multi-dimensional fusion feature set associated with the power device nodes;

[0007] Input the multi-dimensional fusion feature set into a preset abnormal behavior recognition model for intrusion analysis processing to generate an abnormal behavior data set of the power device nodes;

[0008] generate a network security defense strategy set according to the abnormal behavior data set, the network security defense strategy set including real-time blocking instructions and node state repair instructions for different abnormal behavior types;

[0009] feedback the network security defense strategy set to a power network control center to trigger a defense response operation, and adjust the collection strategy of the multi-source monitoring data set according to the execution result of the defense response operation.

[0010] In a second aspect, an embodiment of the present application provides a computer system, comprising:

[0011] a memory, wherein the memory stores a computer program;

[0012] a processor, configured to load the computer program to implement the power network intrusion analysis method based on data fusion as described above.

[0013] The power network intrusion analysis method based on data fusion provided by the present application can realize dynamic feature fusion processing of device physical layer state features and network protocol layer behavior features by obtaining a deep fusion monitoring data set of the running data stream and the network communication data stream of the power device node, generate a multi-dimensional fusion feature set capable of representing the coupling relationship between device running abnormalities and network attack behaviors, so that when an abnormal behavior recognition model is used for intrusion analysis, not only physical attack features such as voltage fluctuation and power abnormality at the device running level can be identified, but also network attack behavior features such as port scanning and protocol tampering can be captured synchronously, and the accurate detection of composite attack modes can be realized through the spatio-temporal correlation analysis of device running state and network communication behavior. The dynamic generation mechanism of the defense strategy set is adopted, the attack path is traced and the key blocking point is located, the real-time blocking instructions and the node state repair instructions are executed in combination, the defects of the traditional defense strategy that only blocks attacks and ignores the recovery of damaged device states are effectively solved, and the attack behavior suppression and the rapid reconstruction of system running safety are simultaneously realized. Through dynamic optimization and adjustment of the multi-source monitoring data collection strategy, the collection granularity of device running parameters and network communication features is adaptively matched according to the defense response effect, the contradiction between incomplete feature coverage and coexistence of data redundancy under a fixed monitoring mode is broken through, the intrusion detection accuracy is improved, and the system resource consumption is reduced. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 is a flowchart of a power network intrusion analysis method based on data fusion provided by an embodiment of the present application.

[0015] Figure 2 is a composition schematic diagram of a computer system provided by an embodiment of the present application. DETAILED DESCRIPTION

[0016] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described, obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0017] Please refer to Figure 1 , Figure 1 A flowchart of a power network intrusion analysis method based on data fusion provided by an embodiment of the present application, the method can be executed by a computer system, comprising the following steps:

[0018] Step S100: acquiring a multi-source monitoring data set in the power network, the multi-source monitoring data set comprising node operation data flow and network communication data flow of a plurality of power device nodes.

[0019] The multi-source monitoring data set refers to a data set obtained from multiple different sources and containing multiple types of information, in the power network scenario, mainly composed of node operation data flow and network communication data flow. The node operation data flow is the data describing the running state of the power device node itself, covering various physical quantities and state information of the device in the running process, reflecting the real-time working condition of the power device. The network communication data flow is the data generated when the devices in the power network communicate with each other, containing information such as communication protocol and data packet transmission, embodying the information interaction between devices.

[0020] The multi-source monitoring data set in the power network can be acquired by using various data acquisition devices and technologies. For the node operation data flow, sensors can be used to acquire the relevant running data of the device. For example, voltage sensors are used to monitor the voltage value of the device in real time, current data is acquired through current transformers, and power sensors are used to measure the power of the device, etc. These sensors can be installed at the key parts of the power device, continuously acquire data according to the preset sampling frequency, and convert the acquired analog signals into digital signals, which are sent to the data storage and processing center through the data transmission line. For the network communication data flow, network flow monitoring devices can be used to capture and record. For example, network probes or intrusion detection systems (IDS) are deployed on the key nodes or communication links of the power network to listen to and analyze the data packets in the network. These devices can capture the data link layer frame structure and network layer packet structure, extract the source port number, target port number, protocol type identifier and other information, and organize these information into network communication data flow for storage and subsequent processing.

[0021] Step S200: Perform dynamic feature fusion processing on the multi-source monitoring data set to generate a multi-dimensional fusion feature set associated with the power equipment node.

[0022] Dynamic feature fusion processing is a process of comprehensively considering the characteristics of different types of data in the multi-source monitoring data set over time and organically combining these characteristics. This processing method fully considers the dynamics of the data and can more accurately reflect the running state and communication behavior of the power equipment node at different times. The multi-dimensional fusion feature set is a set of features containing multiple dimensions, each dimension describing the running state and network communication behavior of the power equipment node from a different perspective. By fusing these features, a more comprehensive and in-depth understanding of the working condition of the power equipment node can be obtained.

[0023] As an implementation, step S200, performing dynamic feature fusion processing on the multi-source monitoring data set to generate a multi-dimensional fusion feature set associated with the power equipment node, can specifically include steps S210-S240:

[0024] Step S210: Perform data preprocessing on the node running data stream to obtain a standardized node running data set, which includes device voltage fluctuation features, current phase offset features, and power abnormal fluctuation features.

[0025] Data preprocessing is a process of cleaning, transforming, and feature extraction on raw data to improve data quality and usability. When processing the node running data stream, the original data may contain noise, outliers, and other interference information, which need to be removed through data preprocessing to extract valuable features. The standardized node running data set is a data set with unified format and specifications after preprocessing, which contains device voltage fluctuation features, current phase offset features, and power abnormal fluctuation features, and can more accurately reflect the running state of the power equipment node.

[0026] Device voltage fluctuation features are features that describe the fluctuation of the voltage of the power equipment node during operation, reflecting the stability of the device voltage. Current phase offset features reflect the degree of phase offset of the current relative to the normal situation, which may be related to the load characteristics of the device, circuit faults, and other factors. Power abnormal fluctuation features represent abnormal changes in device power during operation, which may indicate that the device has a fault or is under external attack.

[0027] As an implementation, step S210, performing data preprocessing on the node running data stream to obtain a standardized node running data set, can specifically include steps S211-S216:

[0028] Step S211: Perform noise filtering on the original node running data stream to remove electromagnetic interference signals and sensor collection outliers, generating preliminary purified data stream.

[0029] The original node running data stream is the unprocessed data collected directly from the power equipment node, which may contain noise information such as electromagnetic interference signals and sensor collection outliers. Electromagnetic interference signals are interference signals generated by the complex electromagnetic environment in the power network, such as surrounding electrical equipment and transmission lines, which may affect the accuracy of the data. Sensor collection outliers refer to data collected by the sensor during the data collection process that does not conform to the actual situation due to its own failure, external interference, etc.

[0030] Noise filtering can be achieved using various filtering algorithms. For example, the mean filtering algorithm is used to calculate the average value of adjacent data points for each data point in the original node running data stream, and the average value is taken as the filtering result of the data point. Specifically, for a time series data {x1, x2,..., xN} of length N, select a window size M (M is an odd number), and for each data point xi (i = (M+1) / 2,..., N-(M-1) / 2), calculate its filtering value yi as the average value of the M data points in the window, i.e. yi = (xi-(M-1) / 2+xi-(M-1) / 2+1+...+xi+(M-1) / 2) / M. In this way, the data can be effectively smoothed and noise interference can be removed.

[0031] Step S212: Perform time series segmentation on the preliminary purified data stream to generate a set of device running state snapshots within equal-interval time windows.

[0032] Time series segmentation is to divide continuous time series data according to a pre-set time interval, dividing it into multiple equal-interval time windows, and the data in each time window constitutes a device running state snapshot. The device running state snapshot is a set of data reflecting the running state of the power equipment node within the pre-set time window, containing various running parameters of the device within the time window.

[0033] When performing time series segmentation, a suitable time window size can be selected according to actual needs. For example, a time window size of 1 minute is selected, and the preliminary purified data stream is divided by every minute. For each time window, extract the device running data such as voltage, current, power, etc., to form a device running state snapshot. In this way, continuous time series data can be converted into a discrete set of device running state snapshots, facilitating subsequent analysis and processing.

[0034] Step S213: Extracting the voltage effective value, current fundamental component and power factor feature parameters in each time window to generate the original operation feature parameter set.

[0035] The voltage effective value refers to the root mean square value of alternating voltage in a cycle, which can reflect the actual doing function of voltage. The current fundamental component refers to the component with the same frequency as the power frequency in alternating current, which is the main component of current. The power factor is an important indicator to measure the utilization efficiency of power equipment, which represents the ratio of active power consumed by the equipment to apparent power.

[0036] The voltage effective value, current fundamental component and power factor feature parameters in each time window can be extracted by using corresponding calculation methods. For the voltage effective value, it can be calculated by squaring, averaging and square root operation on the voltage data in the time window. For the current fundamental component, Fourier transform method can be used to convert the current data in the time window to frequency domain and extract the component with the same frequency as the power frequency. The power factor can be obtained by calculating the ratio of active power to apparent power in the time window. The active power can be calculated by voltage, current and cosine value of power factor angle, and the apparent power is the product of voltage effective value and current effective value. The voltage effective value, current fundamental component and power factor feature parameters extracted in each time window are combined to generate the original operation feature parameter set.

[0037] Step S214: Dynamic comparison processing of the voltage effective value in the original operation feature parameter set with the preset voltage safety range to generate the device voltage fluctuation feature in the standardized node operation data set.

[0038] The preset voltage safety range is the voltage fluctuation range allowed in advance according to the design requirements and operation standards of power system. The dynamic comparison processing is to compare the voltage effective value in the original operation feature parameter set with the preset voltage safety range in real time to analyze whether the voltage effective value is within the safety range and the degree of deviation from the safety range.

[0039] In the dynamic comparison processing, first, the preset voltage safety range is determined, for example, the voltage safety range is [Umin, Umax]. Then, for each voltage effective value Ui in the original operation feature parameter set, it is judged whether it is within the safety range. If Ui is within [Umin, Umax], it is considered that the voltage is normal; if Ui

[0040] Step S215: Perform harmonic component decomposition processing on the current fundamental component to extract the proportion of each harmonic content and generate the current phase shift feature in the standardized node operation data set.

[0041] Harmonic component decomposition processing is the process of decomposing the current fundamental component into harmonic components of different frequencies. In power systems, due to the existence of nonlinear loads, harmonic components may be contained in the current, which will cause the current phase to shift. The proportion of each harmonic content refers to the proportion of harmonic components of different frequencies in the total current.

[0042] Harmonic component decomposition processing on the current fundamental component can use the method of Fourier series expansion. The current fundamental component is represented as the superposition of a series of sine and cosine waves of different frequencies, and the proportion of each harmonic content is obtained by calculating the amplitude and phase of each harmonic component. Specifically, for a periodic current signal i(t), it can be expanded as a Fourier series: i(t) = a0 + ∑(an*cos(nωt) + bn*sin(nωt)), where a0 is the direct current component, an and bn are the coefficients of each harmonic, ω is the fundamental angular frequency, and n is the harmonic number. By calculating the coefficients of each harmonic, the amplitude and phase of each harmonic can be obtained, and the proportion of each harmonic content can be calculated. These proportions of each harmonic content are taken as the current phase shift feature and stored in the standardized node operation data set.

[0043] Step S216: Based on the correlation analysis processing of the power factor feature parameter and the device load rate data, generate the power abnormal fluctuation feature in the standardized node operation data set.

[0044] Device load rate data refers to the ratio of actual load to rated load of power equipment, reflecting the load condition of the device. Correlation analysis processing is to study the relationship between power factor feature parameters and device load rate data, analyze the change of power factor under different load rates, and judge whether the power appears abnormal fluctuation.

[0045] In the correlation analysis processing, a relationship model between power factor and device load rate can be established first. For example, by collecting a large amount of historical data, a linear regression model between power factor and device load rate is established by using regression analysis method: PF = a*LR + b, where PF is the power factor, LR is the device load rate, and a and b are the regression coefficients. Then, for the current power factor feature parameter and device load rate data, the device load rate is substituted into the regression model to calculate the expected power factor. Compare the expected power factor with the actual power factor, if the difference between them exceeds the preset threshold, it is considered that the power appears abnormal fluctuation. The quantification results of these power abnormal fluctuations are taken as the power abnormal fluctuation feature and stored in the standardized node operation data set.

[0046] Step S220: Protocol analysis processing is performed on the network communication data stream to extract a communication protocol feature field set, which includes a data packet transmission interval feature, a protocol type identification feature, and a port access frequency feature.

[0047] Protocol analysis processing is a process of analyzing data packets in a network communication data stream, identifying the communication protocol, and extracting key information in the protocol. The communication protocol feature field set is a set of fields containing key features of the network communication protocol, which can reflect the state and behavior of network communication.

[0048] The data packet transmission interval feature refers to the time interval between adjacent data packets, reflecting the timeliness and stability of data packet transmission. The protocol type identification feature is used to identify the protocol type used in network communication, to determine whether unauthorized protocols are used. The port access frequency feature refers to the number of access requests to a certain port per unit time, reflecting the frequency and activity level of the port.

[0049] As an embodiment, step S220, protocol analysis processing is performed on the network communication data stream to extract a communication protocol feature field set, which can specifically include the following steps S221-S226:

[0050] Step S221: Capture the data link layer frame structure and network layer data packet structure in the power communication network, and generate a set of original communication data packets.

[0051] The data link layer frame structure refers to the format of the data unit transmitted at the data link layer, which contains the control information of the data link layer and the upper layer data. The network layer data packet structure refers to the format of the data unit transmitted at the network layer, which contains the address information of the network layer and the upper layer data. The original communication data packet set is composed of the captured data link layer frame structure and network layer data packet structure.

[0052] Capturing the data link layer frame structure and network layer data packet structure in the power communication network can use network packet capture tools such as Wireshark. Deploy the network packet capture tool on the key nodes or communication links of the power communication network to listen to the data packet transmission in the network. When data packets pass through, the network packet capture tool will capture these data packets and store them as a set of original communication data packets.

[0053] Step S222: Analyze the transmission layer protocol header information in the original communication data packet set to extract the source port number, target port number, and protocol type identifier, and generate a basic protocol feature set.

[0054] The transport layer protocol header information refers to the header part in the transport layer protocol data packet, and contains key information such as source port number, target port number and protocol type identifier. The source port number is the port number used by the sender application, the target port number is the port number used by the receiver application, and the protocol type identifier is used to identify the protocol type used by the transport layer, such as TCP, UDP, etc.

[0055] The transport layer protocol header information in the original communication data packet set can be parsed according to the format specification of different transport layer protocols. For example, for the TCP protocol, its header information contains a fixed length of 20 bytes, of which the source port number and the target port number each occupy 2 bytes, and the protocol type identifier can be determined by the protocol field in the IP header. By parsing these fields, the source port number, the target port number and the protocol type identifier can be extracted, and these information can be combined to generate the basic protocol feature set.

[0056] Step S223: The access request frequency of the target port number in a unit of time is counted to generate the port access frequency feature in the communication protocol feature field set.

[0057] The access request frequency of the target port number in a unit of time refers to the number of access requests to a certain target port number in a preset time interval. By counting the access request frequency of the target port number in a unit of time, the use and activity level of the port can be understood.

[0058] The access request frequency of the target port number in a unit of time can be counted. A time window is set, for example, 1 minute, and the access requests to each target port number are counted in this time window. When the time window ends, the number of access requests to each target port number is recorded, and these numbers are taken as the port access frequency feature and stored in the communication protocol feature field set.

[0059] Step S224: The application layer protocol handshake message timing feature in the original communication data packet set is analyzed, and the adjacent data packet arrival time interval is calculated to generate the data packet transmission interval feature in the communication protocol feature field set.

[0060] The application layer protocol handshake message timing feature refers to the sending and receiving order and time interval of the handshake message in the connection establishment process of the application layer protocol. The adjacent data packet arrival time interval refers to the time difference between the arrival of two adjacent data packets, reflecting the timeliness and stability of data packet transmission.

[0061] When analyzing the timing characteristics of the application layer protocol handshake messages in the original communication data packet set, the handshake messages of the application layer protocol can be parsed, and the sending and receiving times of each message are recorded. Then, the time interval between the arrival of two adjacent data packets is calculated. For example, for the handshake process of the HTTP protocol, the time when the client sends a request message and the time when the server returns a response message are recorded, and the time interval between them is calculated. These adjacent data packet arrival time intervals are taken as data packet transmission interval characteristics and stored in the communication protocol characteristic field set.

[0062] Step S225: Identify non-standard protocol field formats in the network layer data packet structure, detect protocol field length anomalies and checksum error events to generate abnormal protocol handshake behavior identifiers in the communication protocol characteristic field set.

[0063] Non-standard protocol field format refers to the field format in the network layer data packet structure that does not conform to the standard protocol specification. Protocol field length anomaly refers to the length of the protocol field that does not conform to the length range specified in the standard protocol. Checksum error event refers to the case where the checksum calculation result in the protocol data packet does not match the actual value. Abnormal protocol handshake behavior identifier is used to identify abnormal protocol handshake behavior in network communication.

[0064] Identifying non-standard protocol field formats in the network layer data packet structure can be done by comparing with the standard protocol specification. For example, for the IP protocol, the header length field should be within the specified range, and if it is detected that the header length field exceeds this range, it is considered to be a non-standard protocol field format. Detecting protocol field length anomalies and checksum error events can be calculated and compared according to the provisions in the protocol specification. For example, for the TCP protocol, its header contains a checksum field, by recalculating the checksum and comparing it with the checksum in the data packet, if they are inconsistent, it is considered that there is a checksum error event. These abnormal cases are taken as abnormal protocol handshake behavior identifiers and stored in the communication protocol characteristic field set.

[0065] Step S226: Establish a compliance mapping relationship between protocol types and standard power communication protocols, and mark communication sessions using unauthorized protocol types to generate protocol type identification features in the communication protocol characteristic field set.

[0066] Standard power communication protocol refers to the communication protocol specified in the power industry that meets the safety and specification requirements. Compliance mapping relationship refers to associating protocol types with standard power communication protocols to determine whether the protocol type meets the standard requirements. Unauthorized protocol type refers to a protocol type that does not meet the requirements of the standard power communication protocol.

[0067] The mapping relationship between the protocol type and the compliance of the standard power communication protocol can be established by establishing a protocol type database, and storing the type and related information of the standard power communication protocol in the database. For each protocol type in the network communication data stream, a lookup and comparison are performed in the database. If the protocol type is not in the database, it is considered as an unauthorized protocol type. The communication session using the unauthorized protocol type is marked. An identification bit can be set in the communication protocol feature field set, and the identification bit of the communication session using the unauthorized protocol type is set to a specific value. The identification information is stored in the communication protocol feature field set as a protocol type identification feature.

[0068] Step S230: Aligning the standardized node operation data set and the communication protocol feature field set in a time window to determine the associated mapping relationship between the operation state features and the communication behavior features of the power equipment node in the same time sequence window.

[0069] The time window alignment processing is to match and align the standardized node operation data set and the communication protocol feature field set according to the same time window, so that the operation state features and the communication behavior features in the same time window can be corresponded. The associated mapping relationship is to describe the associated relationship between the operation state features and the communication behavior features of the power equipment node in the same time sequence window, and reflects the mutual influence between the device operation state and the network communication behavior.

[0070] As an implementation manner, in step S230, the standardized node operation data set and the communication protocol feature field set are aligned in a time window to determine the associated mapping relationship between the operation state features and the communication behavior features of the power equipment node in the same time sequence window, which can specifically include the following steps S231-S237:

[0071] Step S231: Extracting a device data timestamp sequence according to the time window identifier in the device operation state snapshot set, and extracting a communication data timestamp sequence from the communication protocol feature field set.

[0072] The time window identifier in the device operation state snapshot set is information for identifying the time window to which each device operation state snapshot belongs. The device data timestamp sequence is a sequence composed of the timestamps corresponding to each snapshot in the device operation state snapshot set. The communication data timestamp sequence is a sequence composed of the timestamps corresponding to each communication data in the communication protocol feature field set.

[0073] According to the time window identifier in the device running state snapshot set, the device data timestamp sequence is extracted. The device running state snapshot set can be traversed, the time window identifier of each snapshot is converted into the corresponding timestamp, and these timestamps are arranged in order to form the device data timestamp sequence. According to the communication protocol feature field set, the communication data timestamp sequence is extracted. The communication protocol feature field set can also be traversed to extract the timestamp of each communication data to form the communication data timestamp sequence.

[0074] Step S232: Perform time overlap interval matching processing on the device data timestamp sequence and the communication data timestamp sequence to generate a timestamp matching result containing device running state snapshots and communication protocol feature fields with the same time window identifier.

[0075] The time overlap interval matching processing is to compare the device data timestamp sequence and the communication data timestamp sequence, find the time overlap interval in the two sequences, and match the device running state snapshot and the communication protocol feature field in the overlapping interval. The timestamp matching result is the combination of the device running state snapshot and the communication protocol feature field with the same time window identifier.

[0076] When performing the time overlap interval matching processing, a double-pointer method can be used. Two pointers are set to point to the starting positions of the device data timestamp sequence and the communication data timestamp sequence, respectively. The timestamps pointed by the two pointers are compared. If the two timestamps are in the same time window, the corresponding device running state snapshot and communication protocol feature field are matched and recorded. Then, the pointers are moved to continue comparison until the two sequences are traversed. The matching result is stored as the timestamp matching result.

[0077] Step S233: Establish a first event association index between the device voltage fluctuation feature and the data packet transmission interval feature in the timestamp matching result to generate an association pair set of voltage fluctuation events and communication delay events.

[0078] The first event association index is an index for establishing the association relationship between the device voltage fluctuation feature and the data packet transmission interval feature. The voltage fluctuation event refers to the event of device voltage fluctuation, and the communication delay event refers to the event of long data packet transmission interval. The association pair set is a set of association pairs composed of voltage fluctuation events and communication delay events.

[0079] The first event correlation index between the device voltage fluctuation feature and the data packet transmission interval feature in the timestamp matching result can be established by traversing the timestamp matching result, judging whether the device voltage fluctuation feature exceeds a set threshold for each matched device running state snapshot and communication protocol feature field, and if the threshold is exceeded, considering that a voltage fluctuation event occurs; at the same time, judging whether the data packet transmission interval feature exceeds a set threshold, and if the threshold is exceeded, considering that a communication delay event occurs. The combination of the voltage fluctuation event and the communication delay event is taken as a correlation pair, which is stored in a correlation pair set.

[0080] Step S234: Spatial distribution matching processing is performed on the current phase offset feature and the port access frequency feature, the network path topological relationship between the device node position of the current abnormal area and the high-frequency port access node is analyzed, and a spatial correlation parameter set is generated.

[0081] The spatial distribution matching processing is a process of matching and analyzing the current phase offset feature and the port access frequency feature in space. The current abnormal area refers to the area where the current phase offset feature exceeds the set threshold. The high-frequency port access node refers to the node where the port access frequency feature exceeds the set threshold. The network path topological relationship refers to the connection relationship and path information between nodes in the power network. The spatial correlation parameter set is a set of parameters describing the network path topological relationship between the device node position of the current abnormal area and the high-frequency port access node.

[0082] The spatial distribution matching processing on the current phase offset feature and the port access frequency feature can first determine the device node position of the current abnormal area according to the current phase offset feature, and determine the position of the high-frequency port access node according to the port access frequency feature. Then, through the topological structure information of the power network, the network path topological relationship between these nodes is analyzed. For example, the shortest path length between two nodes, the number of nodes passed through, etc. These analysis results are taken as spatial correlation parameters and stored in the spatial correlation parameter set.

[0083] Step S235: Based on the synchronicity analysis processing of the power abnormal fluctuation feature and the protocol type identification feature, the distribution density parameter of the illegal protocol usage behavior in the power mutation period is calculated, and a protocol violation density sequence is generated.

[0084] The synchronicity analysis processing is a process of studying the synchronization relationship in time between the power abnormal fluctuation feature and the protocol type identification feature. The power mutation period refers to the time period in which the power abnormal fluctuation feature exceeds the set threshold. The illegal protocol usage behavior refers to the communication behavior using unauthorized protocol types. The distribution density parameter is a parameter describing the distribution of the illegal protocol usage behavior in the power mutation period. The protocol violation density sequence is a sequence composed of protocol violation density parameters in different power mutation periods.

[0085] Based on the synchronization analysis and processing of the power abnormal fluctuation feature and the protocol type identification feature, first, the power mutation period is determined according to the power abnormal fluctuation feature, and then the occurrence frequency and distribution of the illegal protocol usage behavior in these periods are counted. The power mutation period can be divided into several small time windows, and the ratio of the occurrence frequency of the illegal protocol usage behavior in each time window to the length of the time window is calculated as the protocol violation density parameter in the time window. The protocol violation density parameters are arranged in time sequence to generate a protocol violation density sequence.

[0086] Step S236: Perform multi-dimensional fusion processing on the association pair set, the spatial association parameter set, and the protocol violation density sequence to construct a multi-dimensional association feature vector reflecting the spatio-temporal association of the device running state feature and the communication behavior feature.

[0087] The multi-dimensional fusion processing is a process of integrating and fusing the association pair set, the spatial association parameter set, and the protocol violation density sequence from multiple dimensions. The multi-dimensional association feature vector is a vector containing multiple dimensional information, which can comprehensively and accurately reflect the spatio-temporal association between the device running state feature and the communication behavior feature.

[0088] The multi-dimensional fusion processing on the association pair set, the spatial association parameter set, and the protocol violation density sequence can be performed by feature splicing. Each element in the association pair set, the spatial association parameter set, and the protocol violation density sequence is taken as a dimension, and these dimensions are spliced in a set order to form a multi-dimensional vector. For example, the number of association pairs in the association pair set is taken as a dimension, the shortest path length in the spatial association parameter set is taken as a dimension, and the protocol violation density parameter in a certain time window in the protocol violation density sequence is taken as a dimension. In this way, the multi-dimensional association feature vector reflecting the spatio-temporal association of the device running state feature and the communication behavior feature is constructed.

[0089] Step S237: Generate an association mapping relationship between the running state feature and the communication behavior feature according to the weight distribution of each dimension in the multi-dimensional association feature vector, and the association mapping relationship is used to describe the spatio-temporal coupling strength of the device abnormal index and the protocol violation behavior.

[0090] The weight distribution of each dimension refers to the weight value corresponding to each dimension in the multi-dimensional association feature vector. These weight values reflect the importance of each dimension in describing the association relationship between the device running state feature and the communication behavior feature. The association mapping relationship is a mathematical model describing the association relationship between the device running state feature and the communication behavior feature, which describes the spatio-temporal coupling strength of the device abnormal index and the protocol violation behavior through the weighted combination of each dimension in the multi-dimensional association feature vector.

[0091] According to the weight distribution of each dimension in the multi-dimensional correlation feature vector, the correlation mapping relationship between the running state feature and the communication behavior feature can be generated by using a linear weighting method. First, the weight value of each dimension is determined, which can be determined by expert experience, machine learning algorithm, etc. Then, multiply each dimension value in the multi-dimensional correlation feature vector by the corresponding weight value, and then add these weighted dimension values to obtain a comprehensive correlation strength value. The correlation strength value is associated with the device anomaly index and the protocol violation behavior to form a correlation mapping relationship. For example, a linear regression model can be established: R = w1*x1 + w2*x2 +... + wn*xn, where R is the correlation strength value, wi is the weight value of the i-th dimension, and xi is the feature value of the i-th dimension. Through this model, the spatio-temporal coupling strength of the device anomaly index and the protocol violation behavior can be calculated according to the multi-dimensional correlation feature vector.

[0092] Step S240: performing feature dimension matching processing on the device voltage fluctuation feature, the current phase offset feature, the power abnormal fluctuation feature, the data packet transmission interval feature, the protocol type identification feature and the port access frequency feature based on the correlation mapping relationship, to generate a fusion feature vector in the multi-dimensional fusion feature set. Each feature dimension in the fusion feature vector corresponds to the correlation strength parameter of the device running state and the network communication behavior of the power device node within the preset time window.

[0093] The feature dimension matching processing is to combine and match the device voltage fluctuation feature, the current phase offset feature, the power abnormal fluctuation feature, the data packet transmission interval feature, the protocol type identification feature and the port access frequency feature according to the correlation mapping relationship, so that each feature dimension can accurately reflect the correlation strength of the device running state and the network communication behavior of the power device node within the preset time window. The fusion feature vector is a vector containing multiple feature dimensions, and each feature dimension represents a correlation strength parameter of the device running state and the network communication behavior.

[0094] As an implementation manner, in step S240, the feature dimension matching processing is performed on the device voltage fluctuation feature, the current phase offset feature, the power abnormal fluctuation feature, the data packet transmission interval feature, the protocol type identification feature and the port access frequency feature based on the correlation mapping relationship, to generate a fusion feature vector in the multi-dimensional fusion feature set, which can specifically include the following steps S241-S247:

[0095] Step S241: converting the time series data of the device voltage fluctuation feature into a standardized fluctuation amplitude index.

[0096] The time series data of the equipment voltage fluctuation feature refers to the sequence of data of the equipment voltage fluctuation feature changing over time. The standardized fluctuation amplitude index is an index obtained by standardizing the time series data of the equipment voltage fluctuation feature, which can more accurately reflect the amplitude of the equipment voltage fluctuation.

[0097] The time series data of the equipment voltage fluctuation feature can be converted into the standardized fluctuation amplitude index by using the normalization method. First, the mean and standard deviation of the time series data of the equipment voltage fluctuation feature are calculated. Then, for each value in the time series data, the mean is subtracted and divided by the standard deviation to obtain the standardized value. Finally, these standardized values are taken as the standardized fluctuation amplitude index. For example, for the time series data {x1, x2,..., xn} of the equipment voltage fluctuation feature, the mean μ = (x1 + x2 +... + xn) / n and the standard deviation σ = sqrt(((x1 - μ) 2 +(x2-μ) 2 +...+(xn-μ) 2 ) / n) are calculated. For each value xi, the standardized value yi = (xi - μ) / σ is calculated, and these yi are taken as the standardized fluctuation amplitude index.

[0098] Step S242: Fourier transform processing is performed on the current phase shift feature to extract the main harmonic frequency components.

[0099] The Fourier transform processing is the process of converting the time domain signal into the frequency domain signal. Through the Fourier transform, the current phase shift feature can be converted from the time domain to the frequency domain, so as to extract the main harmonic frequency components therein. The main harmonic frequency components refer to the harmonic frequency components with larger energy in the current phase shift feature, which have a greater impact on the current phase shift.

[0100] The Fourier transform processing can be performed on the current phase shift feature by using the fast Fourier transform (FFT) algorithm. First, the time series data of the current phase shift feature is taken as the input, and the FFT algorithm is used to convert it into a frequency domain signal. Then, the amplitude of each frequency component in the frequency domain signal is calculated, and the frequency components with larger amplitudes are selected as the main harmonic frequency components. For example, for a current phase shift feature time series data {x1, x2,..., xN} with a length of N, the FFT algorithm is used to convert it into a frequency domain signal {X1, X2,..., XN}. The amplitude |Xi| of each frequency component is calculated, and the first K frequency components with larger amplitudes are selected as the main harmonic frequency components.

[0101] Step S243: Map the power abnormal fluctuation feature into a discrete state transition probability matrix.

[0102] The discrete state transition probability matrix is a matrix used to describe the transition probabilities of the system between different discrete states. Mapping the power abnormal fluctuation feature into a discrete state transition probability matrix is to divide the power abnormal fluctuation feature into several discrete states, and calculate the transition probabilities of the system between these states.

[0103] Firstly, according to the value range of the power abnormal fluctuation feature, it is divided into several discrete states, such as low fluctuation state, medium fluctuation state and high fluctuation state. Then, the number of times of transition between different states in the historical data is counted. Finally, the transition probability of each state is calculated, and these probabilities are combined to form the discrete state transition probability matrix. For example, for a system containing three discrete states S1, S2, S3, the number of times of transition from state S1 to state S2 is n12, the number of times of transition from state S1 to state S3 is n13, and so on. The transition probabilities P12 = n12 / (n12+n13), P13 = n13 / (n12+n13) are calculated, and these probabilities are combined to form the discrete state transition probability matrix.

[0104] Step S244: Probability distribution fitting processing is performed on the data packet transmission interval feature to generate a transmission interval abnormal probability parameter.

[0105] The probability distribution fitting processing is to find a suitable probability distribution model that can better fit the actual distribution of the data packet transmission interval feature. The transmission interval abnormal probability parameter refers to the probability of the data packet transmission interval exceeding the normal range calculated according to the fitted probability distribution model.

[0106] The probability distribution fitting processing on the data packet transmission interval feature can use the maximum likelihood estimation method. First, a suitable probability distribution model is selected, such as normal distribution, exponential distribution, etc. Then, using the maximum likelihood estimation method, the parameters of the probability distribution model are estimated according to the historical data of the data packet transmission interval feature. Finally, according to the estimated parameters, the probability of the data packet transmission interval exceeding the normal range is calculated as the transmission interval abnormal probability parameter. For example, assuming that the normal distribution model N(μ,σ 2 ) is selected, and the mean μ and standard deviation σ are estimated using the maximum likelihood estimation method according to the historical data. For a given normal range [a, b], the probability P(X b), as the transmission interval anomaly probability parameter.

[0107] Step S245: encode the protocol type identification feature into a multi-dimensional one-hot vector form.

[0108] The multi-dimensional one-hot vector is a method of encoding discrete variables into vector form, where only one dimension has a value of 1 and the rest have a value of 0. Encoding the protocol type identification feature into a multi-dimensional one-hot vector form is to convert different protocol type identifications into corresponding multi-dimensional one-hot vectors to facilitate subsequent feature processing and analysis.

[0109] First, determine the set of all possible protocol type identifications. Then, for each protocol type identification, create a vector with a length equal to the size of the protocol type identification set, and set the corresponding dimension of the protocol type identification to 1 and the rest to 0. For example, given the protocol type identification set {TCP, UDP, HTTP, FTP}, for the protocol type identification TCP, create a four-dimensional vector [1, 0, 0, 0]; for the protocol type identification UDP, create a four-dimensional vector [0, 1, 0, 0], and so on. These multi-dimensional one-hot vectors are used as the encoding results of the protocol type identification feature.

[0110] Step S246: use the attention mechanism to perform feature weight allocation processing on the standardized fluctuation amplitude indicator, the main harmonic frequency component, the discrete state transition probability matrix, the transmission interval anomaly probability parameter, and the multi-dimensional one-hot vector.

[0111] Through the attention mechanism, different weights can be assigned to the standardized fluctuation amplitude indicator, the main harmonic frequency component, the discrete state transition probability matrix, the transmission interval anomaly probability parameter, and the multi-dimensional one-hot vector, so that important features are given more attention in subsequent processing.

[0112] Using the attention mechanism to perform feature weight allocation processing on these features can use a neural network-based attention model. First, input the standardized fluctuation amplitude indicator, the main harmonic frequency component, the discrete state transition probability matrix, the transmission interval anomaly probability parameter, and the multi-dimensional one-hot vector into the attention model. The attention model learns the relationship between the input features and calculates the attention weight of each feature. Then, multiply each feature by the corresponding attention weight to get the weighted feature. For example, for an input containing three features x1, x2, x3, the attention model calculates the attention weights w1, w2, w3, respectively, and the weighted features are w1*x1, w2*x2, w3*x3, respectively.

[0113] Step S247: combine the weighted dimensional features into a unified dimensional fusion feature vector through a feature concatenation layer.

[0114] The feature concatenation layer is a layer that combines multiple features of different dimensions into a unified dimension feature vector. Through the feature concatenation layer, the weighted standardized fluctuation amplitude index, the main harmonic frequency component, the discrete state transition probability matrix, the transmission interval abnormal probability parameter, and the multi-dimensional one-hot vector can be combined into a unified dimension fusion feature vector.

[0115] In the feature concatenation layer, the weighted features of different dimensions are arranged in a set order, and then they are concatenated together to form a unified dimension vector. For example, suppose the weighted standardized fluctuation amplitude index is a one-dimensional vector, the main harmonic frequency component is a two-dimensional vector, the discrete state transition probability matrix is a three-dimensional matrix, the transmission interval abnormal probability parameter is a one-dimensional vector, and the multi-dimensional one-hot vector is a four-dimensional vector. These features are concatenated together in order to form a one-dimensional vector, and the dimension of the vector is 1+2+3+1+4=11. The 11-dimensional vector is used as the fusion feature vector.

[0116] Step S300: input the multi-dimensional fusion feature set into a preset abnormal behavior recognition model for intrusion analysis processing to generate an abnormal behavior data set of the power equipment node.

[0117] The preset abnormal behavior recognition model is a pre-trained model for identifying abnormal behaviors of power equipment nodes. The intrusion analysis processing is a process of analyzing the multi-dimensional fusion feature set to determine whether the power equipment node has abnormal behaviors. The abnormal behavior data set is a set containing information related to abnormal behaviors of power equipment nodes, such as the type, occurrence time, and severity of abnormal behaviors.

[0118] As an implementation, in step S300, the multi-dimensional fusion feature set is input into the preset abnormal behavior recognition model for intrusion analysis processing to generate an abnormal behavior data set of the power equipment node, which can specifically include the following steps S310-S350:

[0119] Step S310: construct a feature analysis layer and a behavior judgment layer in the abnormal behavior recognition model, the feature analysis layer is used to extract the equipment operation abnormality index in the fusion feature vector, and the behavior judgment layer is used to identify protocol violation behaviors in the network communication data stream.

[0120] The feature analysis layer is an important component of the abnormal behavior recognition model, and its main function is to analyze the fused feature vector and extract indicators that can reflect device operation abnormalities. Device operation abnormality indicators are parameters that can reflect the abnormal state of power device node operation, such as voltage out-of-limit, current harmonic distortion, power mutation, etc. The behavior judgment layer is responsible for identifying protocol violation behaviors in network communication data streams. Protocol violation behaviors are behaviors that violate standard power communication protocols, such as illegal port scanning, abnormal protocol handshake, and communication frequency out-of-limit, etc.

[0121] As an implementation, step S310, constructing a feature analysis layer and a behavior judgment layer in an abnormal behavior recognition model, can specifically include the following steps S311-S315:

[0122] Step S311: deploying a bidirectional long short-term memory network unit in the feature analysis layer to capture the time sequence dependency of device operation features.

[0123] The bidirectional long short-term memory network (Bi-LSTM) unit is a neural network unit that can process sequence data. It can consider both past and future information of sequence data, thereby better capturing the time sequence dependency in sequence data. Deploying a bidirectional long short-term memory network unit in the feature analysis layer can process device operation features in the fused feature vector and extract time sequence dependency information.

[0124] The bidirectional long short-term memory network unit is composed of two direction LSTM units, namely forward LSTM unit and backward LSTM unit. The forward LSTM unit processes data from the starting position of the sequence, and the backward LSTM unit processes data from the end position of the sequence. The output results of the two direction LSTM units are merged to obtain the final output. For example, for a fused feature vector sequence {x1, x2,..., xT} of length T, the forward LSTM unit processes from x1 to calculate the hidden state h1f, h2f,..., hTf of each time step, and the backward LSTM unit processes from xT to calculate the hidden state h1b, h2b,..., hTb of each time step. Then, the forward and backward hidden states are merged to obtain the final hidden state ht = [htf; htb], where [;] represents the vector concatenation operation.

[0125] Step S312: constructing a convolutional neural network module in the behavior judgment layer to extract spatial local patterns of communication protocol features.

[0126] A convolutional neural network (CNN) module is a neural network module used to process data with spatial structure. It can extract spatial local patterns in the data through convolution operations. In the behavior judgment layer, a convolutional neural network module is constructed to process the communication protocol features in the network communication data stream and extract spatial local patterns, thereby better identifying protocol violation behaviors. A convolutional neural network module is usually composed of a convolutional layer, a pooling layer, and a fully connected layer. The convolutional layer slides over the input data with a convolution kernel to perform convolution operations and extract local features in the data. The pooling layer is used to downsample the output of the convolutional layer, reducing the dimension of the data while preserving important feature information. The fully connected layer connects the output of the pooling layer to perform classification or regression tasks. For example, for a two-dimensional communication protocol feature matrix, the convolutional layer can extract different local patterns with different convolution kernels, the pooling layer can perform max-pooling or average-pooling operations on the output of the convolutional layer, and the fully connected layer can classify protocol violation behaviors based on the output of the pooling layer.

[0127] Step S313: Set a cross-layer feature interaction channel to enable information fusion between the output features of the bidirectional long short-term memory network unit and the intermediate layer features of the convolutional neural network module.

[0128] The cross-layer feature interaction channel is a mechanism for realizing feature information interaction between different layers. By setting the cross-layer feature interaction channel, the output features of the bidirectional long short-term memory network unit and the intermediate layer features of the convolutional neural network module can be fused, thereby comprehensively utilizing the time sequence dependent information of device running features and the spatial local pattern information of communication protocol features to improve the accuracy of abnormal behavior recognition.

[0129] The cross-layer feature interaction channel can be set using the splicing or weighted sum method. For example, the output features of the bidirectional long short-term memory network unit and the intermediate layer features of the convolutional neural network module are spliced to obtain a new feature vector. Then, the new feature vector is input to the subsequent processing layer for further analysis and processing. Alternatively, the output features of the bidirectional long short-term memory network unit and the intermediate layer features of the convolutional neural network module are weighted and summed to obtain a fused feature vector.

[0130] Step S314: Integrate a random forest-based classifier in the model output layer to comprehensively evaluate the joint threat level of device running anomalies and communication behavior anomalies.

[0131] The random forest-based classifier is an ensemble learning-based classifier that combines multiple decision trees to perform classification tasks. By integrating a random forest-based classifier in the model output layer, information about device running anomalies and communication behavior anomalies can be considered comprehensively to classify the abnormal behavior of power device nodes and evaluate their joint threat level.

[0132] The training process of the random forest classifier is as follows: first, a plurality of subsets are randomly extracted from the training data, and each subset is used to train a decision tree. Then, each decision tree is trained according to its own training data to generate its own classification rule. Finally, the classification results of all decision trees are voted to obtain the final classification result. In the abnormal behavior recognition model, the output features of the bidirectional long short-term memory network unit and the convolutional neural network module are taken as the input of the random forest classifier, and the joint threat level of device operation abnormality and communication behavior abnormality is classified by the random forest classifier, such as high risk, medium risk, low risk, etc.

[0133] Step S315: The robustness parameters of the abnormal behavior recognition model are optimized by the adversarial training method, so that it can identify disguised hidden attack patterns.

[0134] The adversarial training method is a method of training the model by introducing adversarial samples, which can improve the robustness of the model and make it resistant to adversarial attacks and identify disguised hidden attack patterns. In the abnormal behavior recognition model, the robustness parameters of the model are optimized by the adversarial training method, which can enhance the model's ability to identify complex attack patterns.

[0135] The process of adversarial training is as follows: first, generate adversarial samples. Adversarial samples are generated by adding small perturbations to the original samples, so that the model makes mistakes in classifying them. Then, the adversarial samples and the original samples are input into the abnormal behavior recognition model for training. During the training process, the model will continuously adjust its parameters to improve the classification accuracy of the original samples and the adversarial samples. Through multiple iterations of training, the robustness parameters of the model will be optimized, making it better able to identify disguised hidden attack patterns.

[0136] Step S320: Perform device state anomaly detection processing on the fusion feature vector through the feature analysis layer to generate a first abnormal detection result set of the power device node, and the first abnormal detection result set includes voltage out-of-limit event identifiers, current harmonic distortion event identifiers, and power mutation event identifiers.

[0137] Device state anomaly detection processing is a process of analyzing the fusion feature vector to determine whether the running state of the power device node is abnormal. The first abnormal detection result set is a set containing information related to the abnormal running state of the power device node. The voltage out-of-limit event identifier in the set is used to identify events where the device voltage exceeds the safe range, the current harmonic distortion event identifier is used to identify events where the harmonic component of the current is abnormal, and the power mutation event identifier is used to identify events where the device power suddenly changes.

[0138] The equipment state anomaly detection processing on the fusion feature vector is performed by the feature analysis layer, and the time sequence dependency information of the equipment operation features extracted by the bidirectional long short-term memory network unit can be utilized. For the voltage overrun event, the voltage related features in the fusion feature vector are compared with the preset voltage safety range, and if it exceeds the safety range, it is marked as a voltage overrun event. For the current harmonic distortion event, the current harmonic component features in the fusion feature vector are analyzed, and if the harmonic content exceeds the set threshold, it is marked as a current harmonic distortion event. For the power mutation event, the change rate of the power features in the fusion feature vector is calculated, and if the change rate exceeds the preset threshold, it is marked as a power mutation event. The identifications of these abnormal events form the first anomaly detection result set.

[0139] Step S330: Perform communication protocol compliance verification processing on the communication protocol feature field set by the behavior judgment layer to generate a second anomaly detection result set, and the second anomaly detection result set includes an illegal port scanning behavior identification, an abnormal protocol handshake behavior identification, and a communication frequency overrun behavior identification.

[0140] The communication protocol compliance verification processing is a process of checking the communication protocol feature field set to determine whether the network communication conforms to the standard power communication protocol. The second anomaly detection result set is a set containing information related to network communication behavior anomalies, in which the illegal port scanning behavior identification is used to identify the behavior of scanning unauthorized ports, the abnormal protocol handshake behavior identification is used to identify the behavior that does not conform to the standard protocol handshake process, and the communication frequency overrun behavior identification is used to identify the behavior that the communication frequency exceeds the normal range.

[0141] The communication protocol compliance verification processing on the communication protocol feature field set by the behavior judgment layer can utilize the spatial local pattern information of the communication protocol features extracted by the convolutional neural network module. For the illegal port scanning behavior, the port access information in the communication protocol feature field set is checked, and if the access request frequency to unauthorized ports is found to be too high, it is marked as an illegal port scanning behavior. For the abnormal protocol handshake behavior, the protocol handshake message time sequence features in the communication protocol feature field set are analyzed, and if the situation that does not conform to the standard protocol handshake process is found, it is marked as an abnormal protocol handshake behavior. For the communication frequency overrun behavior, the data packet transmission frequency in the communication protocol feature field set is counted, and if it exceeds the normal range, it is marked as a communication frequency overrun behavior. The identifications of these abnormal behaviors form the second anomaly detection result set.

[0142] Step S340: Perform spatio-temporal correlation analysis processing on the first anomaly detection result set and the second anomaly detection result set to determine the causal relationship chain between the equipment operation anomaly indicators and the protocol violation behaviors.

[0143] Spatiotemporal correlation analysis is a process that comprehensively considers the temporal and spatial relationships between the first and second sets of anomaly detection results to analyze the causal relationship between abnormal equipment operation indicators and protocol violations. A causal chain is a chain-like structure describing the causal relationship between abnormal equipment operation indicators and protocol violations, which can help to better understand the occurrence mechanism and propagation path of abnormal behavior.

[0144] As one implementation method, step S340 involves performing spatiotemporal correlation analysis on the first and second anomaly detection result sets to determine the causal chain between abnormal equipment operation indicators and protocol violations. This may specifically include the following steps S341 to S349:

[0145] Step S341: Extract the timestamps of device anomaly events from the first anomaly detection result set and the timestamps of protocol violation events from the second anomaly detection result set.

[0146] The device anomaly event timestamp refers to the time information of each device anomaly event in the first anomaly detection result set. The protocol violation event timestamp refers to the time information of each protocol violation event in the second anomaly detection result set. By extracting these timestamps, time-dimensional information can be provided for subsequent spatiotemporal correlation analysis.

[0147] Extracting the timestamps of device anomaly events from the first anomaly detection result set and the timestamps of protocol violation events from the second anomaly detection result set can be achieved by traversing the first and second anomaly detection result sets, extracting the timestamp information of each anomaly event, and storing these timestamp information as two timestamp sequences.

[0148] Step S342: Generate a set of associated event pairs based on the time interval between the timestamp of the device abnormal event and the timestamp of the protocol violation event. Each event pair in the set of associated event pairs contains the device abnormal event and the protocol violation event that occurred within the same time window.

[0149] The time interval refers to the time difference between the timestamp of a device malfunction event and the timestamp of a protocol violation event. A set of related event pairs is a collection of event pairs consisting of device malfunction events and protocol violation events occurring within the same time window. By generating a set of related event pairs, device malfunction events and protocol violation events that may have a causal relationship can be initially identified.

[0150] The set of associated event pairs is generated according to the time interval between the device anomaly event timestamp and the protocol violation event timestamp. A time window size can be set, for example, 1 minute. For each device anomaly event timestamp, find the protocol violation events that occur within the time window in the sequence of protocol violation event timestamps, and group them into an event pair. Collect all such event pairs to form the set of associated event pairs.

[0151] Step S343: Perform network topology path analysis on each event pair in the set of associated event pairs to determine the physical connection path between the power device node corresponding to the device anomaly event and the network communication node associated with the protocol violation event.

[0152] Network topology path analysis is a process of studying the connection relationship and path information between nodes in the power network. By performing network topology path analysis on each event pair in the set of associated event pairs, the physical connection path between the power device node corresponding to the device anomaly event and the network communication node associated with the protocol violation event can be determined, providing spatial dimension information for subsequent causal relationship analysis.

[0153] Performing network topology path analysis on each event pair in the set of associated event pairs can utilize the topology structure information of the power network and use graph search algorithms such as breadth-first search (BFS) or depth-first search (DFS) algorithm to find the shortest path or all possible paths between the power device node corresponding to the device anomaly event and the network communication node associated with the protocol violation event. Store this path information for subsequent analysis.

[0154] Step S344: Generate a set of potential attack links based on the physical connection path, which contains the propagation direction and path node sequence between the device anomaly event and the protocol violation event.

[0155] The set of potential attack links is a set composed of possible attack propagation paths between the device anomaly event and the protocol violation event, which contains the propagation direction and path node sequence of the attack. By generating a set of potential attack links based on the physical connection path, the causal relationship and propagation mechanism between the device anomaly event and the protocol violation event can be further analyzed.

[0156] Based on the physical connection path, the propagation direction and path node sequence of the attack can be determined according to the direction and node sequence of the physical connection path. For example, if the physical connection path is from node A to node B and then to node C, the potential attack link can be represented as A->B->C, where A is the starting node of the attack and B and C are the intermediate nodes through which the attack passes. Collect all such potential attack links to form the set of potential attack links.

[0157] Step S345: Extract the device anomaly index change trend and the protocol violation feature evolution sequence of each attack link in the potential attack link set.

[0158] The device anomaly index change trend refers to the change of the device anomaly index over time in the attack link. The protocol violation feature evolution sequence refers to the change sequence of the protocol violation feature over time in the attack link. By extracting these information, the dynamic relationship between the device anomaly event and the protocol violation event can be further analyzed.

[0159] Extracting the device anomaly index change trend and the protocol violation feature evolution sequence of each attack link in the potential attack link set can find the device anomaly index and the protocol violation feature of the corresponding node in the first anomaly detection result set and the second anomaly detection result set according to the path node sequence of each attack link in the potential attack link set. Then, arrange these indexes and features in chronological order to obtain the device anomaly index change trend and the protocol violation feature evolution sequence.

[0160] Step S346: Perform time sequence alignment processing on the device anomaly index change trend and the protocol violation feature evolution sequence to generate the correlation strength parameter of each attack link.

[0161] Time sequence alignment processing is the process of aligning the device anomaly index change trend and the protocol violation feature evolution sequence in time, so that they can be compared and analyzed on the same time scale. The correlation strength parameter is a parameter that describes the degree of correlation between the device anomaly index change trend and the protocol violation feature evolution sequence.

[0162] The time sequence alignment processing of the device anomaly index change trend and the protocol violation feature evolution sequence can use interpolation or resampling methods. For example, for the device anomaly index change trend and the protocol violation feature evolution sequence, a uniform time interval is selected to resample them so that their time sequence lengths are the same. Then, calculate the correlation coefficient between the device anomaly index change trend and the protocol violation feature evolution sequence, such as the Pearson correlation coefficient, and use the correlation coefficient as the correlation strength parameter.

[0163] Step S347: Screen the attack links with correlation strength parameters exceeding the dynamically adjusted threshold to construct the initial causal relationship chain.

[0164] The dynamically adjusted threshold is a threshold dynamically adjusted according to actual conditions and data analysis results, which is used to screen attack links with high correlation strength. The initial causal relationship chain is a causal relationship chain composed of attack links with correlation strength parameters exceeding the dynamically adjusted threshold, which is a preliminary determination of the causal relationship between the device operation anomaly index and the protocol violation behavior.

[0165] The initial causal relationship chain is constructed by screening attack links whose correlation strength parameters exceed the dynamically adjusted threshold. The potential attack link set is traversed, and for each attack link, the correlation strength parameter thereof is compared with the dynamically adjusted threshold. If the correlation strength parameter exceeds the threshold, the attack link is added to the initial causal relationship chain.

[0166] Step S348: Multi-dimensional verification processing is performed on the initial causal relationship chain, and the verification dimensions include event time sequence continuity, path node matching degree, and data abnormal association.

[0167] The multi-dimensional verification processing is a process of verifying the initial causal relationship chain from multiple dimensions to ensure its rationality and reliability. The event time sequence continuity refers to whether the time sequence of event occurrence in the initial causal relationship chain conforms to logic. The path node matching degree refers to the matching degree of the path nodes in the initial causal relationship chain with the actual network topology structure. The data abnormal association refers to whether the data association between the device abnormal indicators and the protocol violation features in the initial causal relationship chain is reasonable.

[0168] The multi-dimensional verification processing is performed on the initial causal relationship chain, which can be verified from three dimensions of event time sequence continuity, path node matching degree, and data abnormal association. For event time sequence continuity, it is checked whether the time sequence of event occurrence in the initial causal relationship chain conforms to causal logic, such as whether the attack event should occur before the device abnormal event. For path node matching degree, the nodes in the initial causal relationship chain are compared with the actual network topology structure, and it is checked whether there is a mismatch. For data abnormal association, it is analyzed whether the data association between the device abnormal indicators and the protocol violation features in the initial causal relationship chain conforms to the actual situation, such as whether the change of the device abnormal indicators is related to the change of the protocol violation features.

[0169] Step S349: According to the verification result, the initial causal relationship chain that meets the multi-dimensional conditions is marked as a causal relationship chain between the device operation abnormal indicator and the protocol violation behavior.

[0170] According to the verification result, the initial causal relationship chain that meets the multi-dimensional conditions is marked as a causal relationship chain between the device operation abnormal indicator and the protocol violation behavior. The initial causal relationship chain is traversed, and for each attack link, it is checked whether it meets the conditions of the three dimensions of event time sequence continuity, path node matching degree, and data abnormal association. If all conditions are met, the attack link is marked as a causal relationship chain between the device operation abnormal indicator and the protocol violation behavior.

[0171] Step S350: Based on the causal relationship chain, threat level classification processing is performed on the first abnormal detection result set and the second abnormal detection result set, and high-risk behavior records and medium-risk behavior records in the abnormal behavior data set are generated.

[0172] The threat level classification processing is a process of classifying abnormal behaviors in the first abnormal detection result set and the second abnormal detection result set according to the causal relationship chain and determining the threat level thereof. The high-risk behavior record refers to an abnormal behavior record with a higher threat level, and the medium-risk behavior record refers to an abnormal behavior record with a moderate threat level.

[0173] Based on the causal relationship chain, the threat level classification processing is performed on the first abnormal detection result set and the second abnormal detection result set. The threat level of each abnormal behavior can be comprehensively evaluated according to factors such as the correlation strength of the attack link, the complexity of the attack path, and the severity of the abnormal behavior in the causal relationship chain. For example, for an abnormal behavior with high correlation strength, complex attack path, and severe abnormal behavior, it is classified as a high-risk behavior; for an abnormal behavior with moderate correlation strength, relatively simple attack path, and less severe abnormal behavior, it is classified as a medium-risk behavior. These high-risk behavior records and medium-risk behavior records are stored in the abnormal behavior data set.

[0174] Step S400: generating a network security defense policy set according to the abnormal behavior data set, the network security defense policy set including real-time blocking instructions and node state repair instructions for different abnormal behavior types.

[0175] The network security defense policy set is a set of strategies for ensuring the security of the power network. The real-time blocking instructions are used to prevent the occurrence and spread of abnormal behaviors in a timely manner, and the node state repair instructions are used to restore the normal operating state of the affected power equipment nodes. According to the abnormal behavior data set, the network security defense policy set can be generated according to the type and severity of different abnormal behaviors.

[0176] As an implementation, step S400, generating a network security defense policy set according to the abnormal behavior data set, can specifically include the following steps S410-S450:

[0177] Step S410: performing attack path tracing processing on the high-risk behavior record to determine the attacked entry position and attack propagation path of the power equipment node.

[0178] The attack path tracing processing is a process of tracking the source and propagation path of the attack by analyzing the high-risk behavior record. The attacked entry position refers to the starting node position of the attack entering the power network, and the attack propagation path refers to the path information of the attack propagating in the power network.

[0179] The attack path tracing processing on the high-risk behavior record can utilize the causal relationship chain and network topology structure information. Starting from the abnormal event in the high-risk behavior record, the propagation direction of the attack is traced in reverse according to the attack link in the causal relationship chain, and the starting node of the attack, i.e., the attack entry position, is found. Meanwhile, all the nodes and path information passed by the attack are recorded to form the attack propagation path.

[0180] Step S420: generating a first defense strategy subset according to the attack entry position, the first defense strategy subset including a non-authorized communication port closing instruction, an infected device node isolation instruction, and an encrypted communication key resetting instruction.

[0181] The attack entry position is the starting point of the attack on the power network, and the first defense strategy subset generated based on this is aimed at blocking the further penetration and spread of the attack from the source. The non-authorized communication port closing instruction is an important measure to prevent the attacker from continuing the attack by using the unauthorized open port. In the power network, there are a large number of communication ports for data interaction between devices, but some ports may become a security hazard due to improper configuration or being exploited by attackers. By closing these non-authorized communication ports, the attack surface can be effectively reduced. For example, through the network management system, the device corresponding to the attack entry position is port scanned to identify the unauthorized open port, and then a closing instruction is sent to prohibit external access to these ports.

[0182] The infected device node isolation instruction is to isolate the device node that has been infected by the attack from other normal devices in the network to avoid the spread of the attack. The access control list (ACL) can be modified through devices such as network switches or firewalls to block the communication between the infected device node and other nodes. For example, the IP address of the infected device node is added to the blacklist of the firewall to prohibit it from transmitting data with other devices in the network.

[0183] The encrypted communication key resetting instruction is to prevent the attacker from continuing the attack by using the cracked encryption key. In the power network, many communications use encryption technology to ensure the security of data, but if the encryption key is obtained by the attacker, the confidentiality and integrity of the data will be threatened. By resetting the encrypted communication key, the encryption mechanism can be updated to enhance the security of communication. For example, a new encryption key is generated by using a key management system and distributed to the relevant communication devices to ensure that the communication between devices uses the new encryption key for encryption and decryption.

[0184] Step S430: performing path blocking point analysis processing on the attack propagation path to determine a set of key path blocking device nodes in the power network.

[0185] The attack propagation path is a path passed by an attacker when attacking the power network. The path blocking point analysis processing is performed to find a key position that can effectively stop the attack from continuing to propagate. The key path blocking device node set refers to a set of device nodes on the attack propagation path that can block the attack propagation by taking corresponding measures.

[0186] The path blocking point analysis processing can be performed using a graph theory method. The power network is abstracted as a graph, where the nodes represent power devices and the edges represent communication connections between devices. According to the attack propagation path, the properties of each node and edge on the path are analyzed, such as the importance of the node, the bandwidth of the edge, etc. For example, for some nodes located in the core of the network, bearing a large amount of data transmission tasks, or key nodes connecting multiple subnets, these nodes are often important blocking points on the attack propagation path. At the same time, the security and controllability of the nodes can also be considered, and those nodes that are easy to operate and manage are selected as key path blocking device nodes. Through detailed analysis and evaluation of the attack propagation path, the key path blocking device node set is determined, providing a basis for subsequent defense strategy formulation.

[0187] Step S440: generating a second defense strategy subset according to the key path blocking device node set, the second defense strategy subset including deploying traffic cleaning instructions, enabling protocol whitelist filtering instructions, and strengthening identity authentication strength instructions.

[0188] The key path blocking device node set determines the key nodes that can effectively block the attack propagation in the power network. Based on this, the second defense strategy subset generated aims to further enhance the security of the network by operating these key nodes. Deploying traffic cleaning instructions is to clean the network traffic passing through the key path blocking device nodes, removing malicious traffic therefrom. Traffic cleaning devices such as Intrusion Detection System (IDS) or Intrusion Prevention System (IPS) can be deployed on the key nodes. These devices can monitor network traffic in real time, identify malicious traffic such as DDoS attack traffic, virus infection traffic, etc., and filter and clean them. For example, when a large amount of abnormal traffic is detected flowing to a certain key node, the traffic cleaning device can automatically direct these traffic to a predefined cleaning center, where the traffic is analyzed and processed to remove malicious components before normal traffic is returned to the network.

[0189] The enable protocol whitelist filtering instruction is to allow data transmission only using predefined and safe communication protocols, and prohibit the use of unauthorized protocols. In the power network, there are various communication protocols, and some protocols may have security vulnerabilities and be easily exploited by attackers. By enabling the protocol whitelist filtering instruction, the protocol whitelist is configured on the critical path blocking device node, and only protocols that meet the whitelist requirements are allowed to pass. For example, only standard power communication protocols such as IEC 61850, Modbus, etc. are allowed for communication, and some non-standard or security-risk protocols are prohibited.

[0190] The strengthen identity authentication strength instruction is to improve the identity authentication requirements of devices and users in the power network, ensuring that only legitimate devices and users can access network resources. Multi-factor authentication methods can be used, such as combining username / password, digital certificates, biometric identification, etc. On the critical path blocking device node, strict identity authentication is performed on all access requests. For example, when a device tries to access a critical node, in addition to providing a username and password, it also needs to provide a digital certificate for verification, and may need to undergo biometric verification such as fingerprint recognition or facial recognition. By strengthening the identity authentication strength, it can effectively prevent attackers from impersonating legitimate devices or users to launch attacks.

[0191] Step S450: The first defense strategy subset and the second defense strategy subset are combined according to the preset defense priority to generate real-time blocking instructions in the network security defense strategy set, and node state repair instructions are generated according to the medium-risk behavior record.

[0192] The preset defense priority is a strategy execution order preset according to the security requirements of the power network and the severity of attacks. The first defense strategy subset and the second defense strategy subset are combined according to the preset defense priority to ensure that the defense strategies are executed in the most effective order when responding to attacks. For example, for some urgent attack situations, the close unauthorized communication port instruction and the isolate infected device node instruction are executed first to quickly cut off the source and transmission path of the attack; for some relatively slow attacks, the deployment traffic cleaning instruction and the enable protocol whitelist filtering instruction can be executed on the basis of the above instructions.

[0193] By reasonably combining the first defense strategy subset and the second defense strategy subset, real-time blocking instructions in the network security defense strategy set are generated. These real-time blocking instructions can be quickly issued to the corresponding device nodes through the network management system to ensure that the defense strategies take effect in time.

[0194] The medium-risk behavior record refers to an abnormal behavior record with a medium threat level in the abnormal behavior data set. The node state repair instruction is generated according to the medium-risk behavior record, so as to restore the normal operation state of the power equipment node affected by the medium-risk abnormal behavior. For example, if the medium-risk behavior record shows that the communication protocol of a certain device node is abnormal, but the device itself is not seriously damaged, an instruction can be generated to restore the normal communication protocol configuration; if part of the data of the device is wrong, an instruction can be generated to repair and correct the data. Through these node state repair instructions, the affected device nodes can be repaired in time, and the influence of abnormal behavior on the power network can be reduced.

[0195] Step S500: feeding the network security defense strategy set to the power network control center to trigger the defense response operation, and adjusting the collection strategy of the multi-source monitoring data set according to the execution result of the defense response operation.

[0196] The power network control center is the core management and control mechanism of the power network. Feeding the network security defense strategy set to the power network control center can enable the control center to timely understand the security status of the network and take corresponding defense measures. The defense response operation refers to a series of security protection operations performed by the power network control center on the power network according to the received network security defense strategy set, such as executing real-time blocking instructions, repairing affected device nodes, etc.

[0197] As an implementation manner, in step S500, the network security defense strategy set is fed to the power network control center to trigger the defense response operation, and the collection strategy of the multi-source monitoring data set is adjusted according to the execution result of the defense response operation. Specifically, the following steps S510-S550 can be included:

[0198] Step S510: issuing a real-time blocking instruction to a target device node in the key path blocking device node set for execution, and monitoring the instruction execution state data of the target device node.

[0199] The real-time blocking instruction is an important part of the network security defense strategy set, and its purpose is to timely stop the continuous propagation of attacks. The real-time blocking instruction can be issued to the target device node in the key path blocking device node set, which can be achieved through a network management system or a special instruction transmission protocol. For example, the Simple Network Management Protocol (SNMP) is used to send the instruction to the target device node. After receiving the instruction, the target device node will perform the corresponding operation according to the requirements of the instruction, such as closing unauthorized communication ports, enabling protocol whitelist filtering, etc.

[0200] The instruction execution state data of the target device node is monitored to ensure that the instructions are executed correctly and to discover problems in the execution process in a timely manner. The instruction execution state data can be collected by installing monitoring software on the target device node or using the device's own log recording function. For example, information such as the time of receiving the instruction, the time of starting execution, the time of completing execution, and whether errors or abnormal situations occur during execution is recorded. This instruction execution state data is fed back to the power network control center for further analysis and processing.

[0201] Step S520: Evaluate the effectiveness parameter of the real-time blocking instruction according to the instruction execution state data. If the effectiveness parameter is below the preset threshold, start the backup defense strategy and update the path blocking point of the attack propagation path.

[0202] The instruction execution state data reflects the execution of the real-time blocking instruction. According to these data, the effectiveness parameter of the real-time blocking instruction can be evaluated. The effectiveness parameter can be measured by various indicators, such as the success rate of instruction execution and the degree of reduction of attack traffic. For example, if the attack traffic is significantly reduced after the unauthorized communication port closing instruction is executed, it indicates that the effectiveness of the instruction is high. If the attack traffic does not change significantly or even increases after the instruction is executed, it indicates that the effectiveness of the instruction is low.

[0203] The preset threshold is a judgment standard set in advance according to the security requirements of the power network and historical experience. If the effectiveness parameter is below the preset threshold, it indicates that the current real-time blocking instruction has failed to effectively block the propagation of attacks, and the backup defense strategy needs to be started. The backup defense strategy is a strategy prepared in advance for use when the main defense strategy fails, such as increasing firewall rules, strengthening traffic monitoring, etc. At the same time, the path blocking point of the attack propagation path needs to be updated. By reanalyzing the attack propagation path and combining the new attack situation, a new set of key path blocking device nodes is determined to take more effective defense measures.

[0204] Step S530: Perform running state recovery verification processing on the power device node after executing the node state repair instruction, and generate a device state recovery report and a remaining risk indicator.

[0205] The node state repair instruction is executed to restore the normal running state of the power device node affected by abnormal behavior. The running state recovery verification processing is performed on the power device node after executing the node state repair instruction to ensure that the device node has been restored to a normal running state and there is no residual security risk.

[0206] The running state recovery verification process can be performed in various ways, such as detecting various running parameters of the device, performing function tests, etc. For example, for a device node that has communication failure due to an attack, after executing the node state repair instruction, it is detected whether the communication port is working normally, whether the communication protocol is configured correctly, whether the data transmission is stable, etc. Through the detection and test of these aspects, it is judged whether the running state of the device node has been restored to normal.

[0207] According to the result of the running state recovery verification process, a device state recovery report is generated. The device state recovery report records in detail the recovery situation of the device node, including the comparison of running parameters before and after recovery, function test results, etc. At the same time, the remaining risk indicators of the device node are generated. The remaining risk indicators can be expressed in the form of risk level, risk probability, etc., providing a reference for subsequent security management.

[0208] Step S540: Adjust the collection strategy of the multi-source monitoring data set according to the device state recovery report and the remaining risk indicators.

[0209] The device state recovery report and the remaining risk indicators reflect the current state and security risk situation of the power device node. According to these information, the collection strategy of the multi-source monitoring data set is adjusted, which can make the monitoring data more targetedly reflect the security situation of the network.

[0210] If the device state recovery report shows that the device node has completely recovered to normal and the remaining risk indicators are low, the monitoring frequency of the device node can be appropriately reduced to reduce the workload of data collection. For example, the original device running data collected every minute is adjusted to be collected every five minutes. On the contrary, if the device state recovery report shows that the device node still has some problems or the remaining risk indicators are high, the monitoring frequency of the device node needs to be increased and the dimension of data collection needs to be increased. For example, in addition to collecting the basic running parameters of the device, the network traffic, communication protocol, etc. of the device are also monitored. By adjusting the collection strategy of the multi-source monitoring data set, potential security problems can be more effectively found and the security of the power network can be improved.

[0211] Step S550: Deploy a policy execution monitoring module in the power network control center to continuously monitor the implementation effect of the network security defense policy set and dynamically optimize the detection threshold parameters of the abnormal behavior recognition model.

[0212] The strategy execution monitoring module is a module for monitoring the implementation effect of the network security defense strategy set in real time. By deploying the strategy execution monitoring module in the power network control center, the execution of the network security defense strategy can be comprehensively and real-time monitored. The module can collect various related data such as the change of attack traffic, the state change of device nodes, the success rate of instruction execution, and analyze and process these data.

[0213] By continuously monitoring the implementation effect of the network security defense strategy set, problems that occur in the strategy execution process can be discovered in a timely manner, such as some defense strategies failing to achieve the expected effect, new attack types being unable to be effectively identified, etc. According to the monitoring results, the network security defense strategy can be adjusted and optimized to ensure that it can continuously and effectively protect the security of the power network.

[0214] At the same time, the detection threshold parameter of the abnormal behavior identification model is dynamically optimized. The detection threshold parameter of the abnormal behavior identification model is a standard for judging whether the network behavior is abnormal, and appropriate detection threshold parameter can improve the identification accuracy of the model. With the changes of the network security environment and the continuous updating of attack means, the original detection threshold parameter may no longer be applicable. Through monitoring and analysis of the implementation effect of the network security defense strategy, the identification performance of the model under different conditions can be understood, and the detection threshold parameter can be dynamically adjusted according to the actual situation. For example, if it is found that the model has a low recognition rate for some types of attacks, the detection threshold can be appropriately lowered to improve the sensitivity of the model; if it is found that the model has a high false positive rate, the detection threshold can be appropriately increased to reduce the occurrence of false positives. By dynamically optimizing the detection threshold parameter of the abnormal behavior identification model, the performance and adaptability of the model can be improved, and the security of the power network can be better protected.

[0215] In summary, the power network intrusion analysis method based on data fusion provided by the embodiment of the application can accurately identify abnormal behaviors in the power network by processing and analyzing multi-source monitoring data, and generate corresponding network security defense strategies. At the same time, by monitoring and adjusting the execution effect of the defense strategy and optimizing the abnormal behavior identification model, the security and reliability of the power network are continuously improved, and various network attacks and security threats can be effectively dealt with.

[0216] It can be understood that, in the above introduction of the embodiments of the present application, various algorithms involved, such as the FFT algorithm, the mean filtering algorithm and the like, can be known from the related content in the prior art, and in order to save space, they are not expanded too much in the embodiments of the present application. In addition, those skilled in the art can supplement the details according to the common knowledge in the art when implementing the scheme of the present application, for example, according to the common knowledge in the art, the normalization can be used to eliminate the dimensional conflict before feature fusion, the interpolation can be used to eliminate the dimensional difference, the threshold can be reasonably set in combination with historical data, experience or business scene requirements, the model can be trained based on a general model training manner, and the like, and the present application will not introduce the redundant implementation process too much in details.

[0217] Please refer to Figure 2 , Figure 2A structural schematic diagram of a computer system provided by the embodiment of the present application is shown in FIG. 1. The computer system includes at least a processor 101, a communication interface 102 and a memory 103. The processor 101, the communication interface 102 and the memory 103 can be connected through a bus or other means. The processor 101 (also referred to as a central processing unit (CPU)) is the computing core and control core of the computer system, which can parse various instructions in the computer system and process various data of the computer system. The communication interface 102 can optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and can be used for transmitting and receiving data under the control of the processor 101; the communication interface 102 can also be used for transmitting and interacting data within the computer system. The memory 103 is a memory device in the computer system, which is used for storing programs and data. It can be understood that the memory 103 can include a built-in memory of the computer system, and of course can also include an extended memory supported by the computer system. The memory 103 provides a storage space, which stores an operating system of the computer system, and the present application does not limit this.

[0218] In one embodiment, the processor 101 executes the computer program in the memory 103 to perform the power network intrusion analysis method based on data fusion provided by the embodiment of the present application.

Claims

1. A data fusion based power network intrusion analysis method, characterized in that, The method comprises the following steps: obtaining a multi-source monitoring data set in a power network, the multi-source monitoring data set comprising node operation data streams and network communication data streams of a plurality of power equipment nodes; performing dynamic feature fusion processing on the multi-source monitoring data set to generate a multi-dimensional fusion feature set associated with the power equipment nodes; specifically comprising: performing data preprocessing on the node operation data streams to obtain a standardized node operation data set, the standardized node operation data set comprising equipment voltage fluctuation features, current phase offset features and power abnormal fluctuation features; performing protocol analysis processing on the network communication data streams to extract a communication protocol feature field set, the communication protocol feature field set comprising data packet transmission interval features, protocol type identification features and port access frequency features; performing time window alignment processing on the standardized node operation data set and the communication protocol feature field set to determine the associated mapping relationship between the operation state features and the communication behavior features of the power equipment nodes within the same time sequence window; performing feature dimension matching processing on the equipment voltage fluctuation features, the current phase offset features, the power abnormal fluctuation features, the data packet transmission interval features, the protocol type identification features and the port access frequency features based on the associated mapping relationship to generate a fusion feature vector in the multi-dimensional fusion feature set; wherein each feature dimension in the fusion feature vector corresponds to an associated strength parameter of the equipment operation state and the network communication behavior of the power equipment nodes within a preset time window; inputting the multi-dimensional fusion feature set into a preset abnormal behavior recognition model for intrusion analysis processing to generate an abnormal behavior data set of the power equipment nodes; specifically comprising: constructing a feature analysis layer and a behavior judgment layer in the abnormal behavior recognition model, the feature analysis layer being used to extract equipment operation abnormality indicators in the fusion feature vector, and the behavior judgment layer being used to identify protocol violation behaviors in the network communication data streams; performing equipment state abnormality detection processing on the fusion feature vector through the feature analysis layer to generate a first abnormal detection result set of the power equipment nodes, the first abnormal detection result set comprising voltage out-of-limit event identification, current harmonic distortion event identification and power mutation event identification; performing communication protocol compliance verification processing on the communication protocol feature field set through the behavior judgment layer to generate a second abnormal detection result set, the second abnormal detection result set comprising illegal port scanning behavior identification, abnormal protocol handshake behavior identification and communication frequency out-of-limit behavior identification; performing spatio-temporal correlation analysis processing on the first abnormal detection result set and the second abnormal detection result set to determine a causal relationship chain between the equipment operation abnormality indicators and the protocol violation behaviors; performing threat level classification processing on the first abnormal detection result set and the second abnormal detection result set based on the causal relationship chain to generate high-risk behavior records and medium-risk behavior records in the abnormal behavior data set; The network security defense strategy set is generated based on the abnormal behavior data set, specifically including: performing attack path tracing processing on the high-risk behavior records to determine the attack entry point location and attack propagation path of the power equipment nodes; generating a first defense strategy subset based on the attack entry point location, the first defense strategy subset including instructions to close unauthorized communication ports, isolate infected device nodes, and reset encrypted communication keys; performing path blocking point analysis processing on the attack propagation path to determine a set of critical path blocking device nodes in the power network; generating a second defense strategy subset based on the set of critical path blocking device nodes, the second defense strategy subset including instructions to deploy traffic scrubbing, enable protocol whitelist filtering, and strengthen identity authentication; combining the first defense strategy subset and the second defense strategy subset according to a preset defense priority to generate real-time blocking instructions in the network security defense strategy set, and generating node status repair instructions based on the medium-risk behavior records; the network security defense strategy set includes real-time blocking instructions and node status repair instructions for different abnormal behavior types. The network security defense strategy set is fed back to the power network control center to trigger a defense response operation, and the collection strategy of the multi-source monitoring data set is adjusted according to the execution result of the defense response operation.

2. The method of claim 1, wherein, The step of feeding back the network security defense strategy set to the power network control center to trigger a defense response operation, and adjusting the acquisition strategy of the multi-source monitoring data set based on the execution result of the defense response operation, includes: The real-time blocking command is sent to the target device node in the set of critical path blocking device nodes for execution, and the command execution status data of the target device node is monitored; The effectiveness parameters of the real-time blocking instruction are evaluated based on the instruction execution status data. If the effectiveness parameters are lower than a preset threshold, a backup defense strategy is activated and the path blocking points of the attack propagation path are updated. After executing the node status repair command, the power equipment node is subjected to operation status recovery verification processing to generate an equipment status recovery report and remaining risk indicators. The collection frequency and feature dimensions of the multi-source monitoring data set are adjusted based on the equipment status recovery report and the remaining risk indicators; A policy execution monitoring module is deployed in the power network control center to continuously monitor the implementation effect of the network security defense policy set and dynamically optimize the detection threshold parameters of the abnormal behavior identification model.

3. The method of claim 1, wherein, The process of preprocessing the node operation data stream to obtain a standardized node operation dataset includes: The original node running data stream is subjected to noise filtering to remove electromagnetic interference signals and abnormal values ​​collected by sensors, generating a preliminary purified data stream. The preliminary purified data stream is processed by time series segmentation to generate a set of device operating status snapshots within equally spaced time windows; Extract the RMS voltage, fundamental current component, and power factor characteristic parameters within each time window to generate the original set of operating characteristic parameters; The effective value of voltage in the original operating characteristic parameter set is dynamically compared with a preset voltage safety range to generate a device voltage fluctuation characteristic in the standardized node operating data set; The current fundamental component is subjected to harmonic component decomposition processing to extract a harmonic content proportion to generate a current phase shift characteristic in the standardized node operating data set; Based on the association analysis processing of the power factor characteristic parameter and the device load rate data, a power abnormal fluctuation characteristic is generated in the standardized node operating data set.

4. The method of claim 1, wherein, The protocol analysis processing of the network communication data stream extracts a communication protocol characteristic field set, including: The data link layer frame structure and network layer packet structure in the power communication network are captured to generate an original communication packet set; The transport layer protocol header information in the original communication packet set is analyzed to extract a source port number, a target port number and a protocol type identifier to generate a basic protocol characteristic set; The access request frequency of the target port number within a unit time is counted to generate a port access frequency characteristic in the communication protocol characteristic field set; The application layer protocol handshake message timing characteristics in the original communication packet set are analyzed to calculate the adjacent packet arrival time interval to generate a data packet transmission interval characteristic in the communication protocol characteristic field set; Non-standard protocol field formats in the network layer packet structure are identified to detect protocol field length abnormalities and check code error events to generate an abnormal protocol handshake behavior identifier in the communication protocol characteristic field set; A compliance mapping relationship between the protocol type and the standard power communication protocol is established to mark a communication session using an unauthorized protocol type to generate a protocol type identification characteristic in the communication protocol characteristic field set.

5. The method of claim 1, wherein, The time window alignment processing of the standardized node operating data set and the communication protocol characteristic field set determines the association mapping relationship between the operating state characteristics and the communication behavior characteristics of the power device node within the same timing window, including: The device data timestamp sequence is extracted according to the time window identifier in the device operating state snapshot set, and the communication data timestamp sequence is extracted from the communication protocol characteristic field set; The device data timestamp sequence and the communication data timestamp sequence are subjected to time overlap interval matching processing to generate a timestamp matching result containing device operating state snapshots and communication protocol characteristic fields with the same time window identifier; A first event association index between the device voltage fluctuation characteristic and the data packet transmission interval characteristic is established in the timestamp matching result to generate an association pair set of voltage fluctuation events and communication delay events; The current phase shift characteristic and the port access frequency characteristic are subjected to spatial distribution matching processing to analyze the network path topology relationship between the device node position of the current abnormal area and the high-frequency port access node to generate a spatial association parameter set; Based on the synchronicity analysis processing of the power abnormal fluctuation characteristic and the protocol type identification characteristic, the distribution density parameter of the protocol violation usage behavior within the power mutation period is calculated to generate a protocol violation density sequence; The multi-dimensional fusion processing is performed on the association pair set, the spatial association parameter set and the protocol violation density sequence, and a multi-dimensional association feature vector reflecting the spatial association of the device running state feature and the communication behavior feature is constructed; An association mapping relationship of the running state feature and the communication behavior feature is generated according to the weight distribution of each dimension in the multi-dimensional association feature vector, and the association mapping relationship is used to describe the spatial coupling strength of the device abnormal index and the protocol violation behavior.

6. The method of claim 1, wherein, The feature dimension matching processing is performed on the device voltage fluctuation feature, the current phase offset feature, the power abnormal fluctuation feature, the data packet transmission interval feature, the protocol type identification feature and the port access frequency feature based on the association mapping relationship, and a fusion feature vector in the multi-dimensional fusion feature set is generated, including: The time series data of the device voltage fluctuation feature is converted into a standardized fluctuation amplitude index; The current phase offset feature is processed by Fourier transform to extract main harmonic frequency components; The power abnormal fluctuation feature is mapped into a discrete state transition probability matrix; The data packet transmission interval feature is processed by probability distribution fitting to generate a transmission interval abnormal probability parameter; The protocol type identification feature is encoded into a multi-dimensional one-hot vector form; The attention mechanism is used to perform feature weight distribution processing on the standardized fluctuation amplitude index, the main harmonic frequency component, the discrete state transition probability matrix, the transmission interval abnormal probability parameter and the multi-dimensional one-hot vector; The weighted dimensional features are combined into the fusion feature vector of a unified dimension through a feature splicing layer.

7. A computer system, characterized by It includes: A memory in which a computer program is stored; A processor for loading the computer program to implement the power network intrusion analysis method based on data fusion according to any one of claims 1-6.

Citation Information

Patent Citations

  • Micro-grid industrial control network-oriented intrusion detection method and system

    CN115314301A

  • Network security protection method and system

    CN117879970A