Multi-level network risk isolation method and system for full-automatic transformer substation anti-error detection
By employing multi-level network risk isolation methods and dynamic feature masking technology, the problems of single protection level, weak flood attack capability, and insufficient protocol identification in the anti-misoperation logic detection of intelligent substations have been solved. This has enabled full-chain protection and efficient and accurate protocol identification, thereby improving the security and real-time communication capabilities of substations.
Patent Information
- Application Number
- CN202511165074.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-10-31
AI Technical Summary
Existing intelligent substation anti-misoperation logic detection technologies have shortcomings in network risk isolation, including a single protection level, weak ability to cope with flood attacks, poor coordination of protection strategies, reliance on external security equipment, and deficiencies in protocol identification in terms of scope, real-time performance, dynamic adaptability, and encrypted message identification. They cannot achieve full-chain protection and efficient and accurate protocol identification, and cannot guarantee the safe and reliable operation of substation secondary systems.
A multi-level network risk isolation method is adopted. By constructing a layered parallel collaborative architecture that integrates fast filtering at the network/transport layer, dynamic port-load joint pre-matching, and deep parsing of application layer feature masks, and combining information entropy to quantify the importance of feature bits and feature extraction operations, efficient processing of packets from fast coarse screening to accurate parsing is achieved. The dynamic feature masking technology breaks through the limitations of encryption technology, and an intelligent dynamic learning system is constructed to adapt to dynamic changes in protocols and optimize the consumption of computing resources.
It significantly improves matching efficiency and system response speed, effectively solves the problem of encrypted message identification, enhances the system's adaptability to dynamic protocol changes, reduces computing resource consumption, and ensures the real-time communication needs and safe and stable operation of substations.
Smart Images

Figure CN120880957A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart substation safety protection and false detection prevention, and in particular to a multi-level network risk isolation method and system for false detection prevention in fully automated substations. Background Technology
[0002] In the operation of smart substations, anti-malfunction logic detection technology is crucial. Currently, this technology primarily focuses on the verification methods for anti-malfunction logic, such as manual setting and primary equipment status adjustment. However, in terms of network security protection, its protection methods have significant limitations, relying heavily on external security devices such as firewalls and isolation gateways, rather than the isolation capabilities integrated into the detection device itself. This is specifically reflected in the following aspects: From a protection layer perspective, existing protection measures are relatively simple, often concentrated at a single layer, such as IP filtering at the network layer or port restrictions at the transport layer. They lack a comprehensive, coordinated defense system spanning from the application layer to the physical layer, making it difficult to cope with complex attacks across multiple layers. Regarding protection dependency, the detection devices themselves lack hardware-level isolation mechanisms, relying excessively on external devices. This not only leads to delayed response times but also increases integration complexity. Furthermore, existing protection measures are mostly passive, relying solely on static rules to block attacks such as flooding attacks, lacking dynamic policy linkage, such as tiered escalation defenses from packet dropping to physical network port disabling, resulting in severely insufficient protection flexibility.
[0003] Specifically, existing technologies have several shortcomings. First, their network risk isolation capabilities are insufficient. Existing detection devices lack a built-in multi-level isolation architecture, failing to achieve full-chain protection from application-layer data verification and transport / network layer access control to physical-layer hardware blocking. This makes them vulnerable to cross-layer attacks, such as attackers using forged application-layer packets combined with unauthorized network-layer IP access. Second, their ability to handle flooding attacks is weak. For physical-layer traffic flooding attacks, such as UDP Flood and SYN Flood, existing devices lack hardware-level threshold control and network port disabling mechanisms. Relying solely on software-level traffic filtering is insufficient to prevent attacks from consuming system resources. Third, the coordination of protection strategies is poor. Each layer of protection measures, such as application-layer data verification and network-layer IP whitelisting, operates independently without a unified security policy engine for coordination. This allows lightweight attacks to escalate into high-risk threats if not handled promptly, resulting in low protection efficiency. Fourth, relying on external security equipment, some devices do not have risk isolation capabilities themselves and require additional deployment of firewalls and other equipment. This not only increases the complexity and cost of the system, but also creates inter-device coordination vulnerabilities, such as protection gaps caused by rule synchronization delays.
[0004] A comparison with existing technical literature and patents also reveals the shortcomings of current technologies. For example, patent CN115941564A does not adequately consider dynamic port learning and feature mask optimization in protocol identification, resulting in relatively weak adaptability to custom protocols. Patent CN114520838A uses a K-nearest neighbor model for protocol identification, which has high computational complexity and low efficiency when handling large-scale traffic, failing to meet the rapid matching requirements of multi-level network risk isolation in fully automated substation anti-misoperation logic detection. While the industrial control protocol identification method proposed in patent CN112788015A has some application value in the Industrial Internet of Things (IIoT), its identification scope is limited to industrial control protocols, and its processing efficiency is low in high-traffic scenarios.
[0005] In summary, existing intelligent substation anti-malfunction logic detection technologies have significant shortcomings in areas such as hierarchical coordination of network security protection, attack response, strategy linkage, and protocol identification. A better technical solution is urgently needed to address these issues and ensure the safe and stable operation of intelligent substations. Summary of the Invention
[0006] Therefore, the technical problem to be solved by the present invention is that the existing intelligent substation anti-misoperation logic detection technology has shortcomings in network risk isolation, such as single protection level, weak ability to deal with flood attacks, poor coordination of protection strategies, reliance on external security equipment, and protocol identification in terms of scope, real-time performance, dynamic adaptability, and encrypted message identification. It cannot achieve full-chain protection and efficient and accurate protocol identification to ensure the safe and reliable operation of the substation secondary system.
[0007] To address the aforementioned technical problems, this invention provides a multi-level network risk isolation method and system for preventing false detection in fully automated substations. The method includes the following steps: S1: Receive message data from the substation network in real time, and synchronously extract five-tuple information including source IP address, destination IP address, source port number, destination port number, transport layer protocol type, network layer protocol identifier, and transport layer status information from the message data; and construct a historical database storing successfully matched link information, and match the extracted information with the historical database: If the match is successful, the protocol type is obtained; If the match fails, proceed to step S2; S2: Deploy a port mapping baseline library and pre-configure static associations between fixed ports and protocols; collect port traffic feature data in real time based on a stream computing engine, dynamically optimize the port-protocol mapping model through an online clustering algorithm, and obtain dynamic port association results; at the same time, perform lightweight feature extraction on the load portion of unmatched packets to obtain load feature vectors; S3: Combine the port dynamic association results with the load feature vector to generate a pre-matching list containing multiple candidate protocols; S4: Based on the candidate protocols in the pre-matching list, a deep bitwise AND operation is performed on the message data to be parsed in the application layer and the feature mask in the template library through a dynamic feature mask generation mechanism and feature extraction operation to identify the protocol type.
[0008] In one embodiment of the present invention, in step S4, the method for identifying the protocol type includes: Based on the candidate protocols in the pre-matched protocol list, the corresponding protocol sample set is retrieved from the template library. The protocol sample set It consists of a fixed-length L-bit binary sequence, where the bit sequence of each sample is B = {b0, b1, ..., b}. L-1}, where b j Represents the Boolean value at position j; For each bit position j, calculate the probability that the bit is 1. : , This represents the number of samples where position j is 1, and N represents the total number of samples. According to the probability Calculate the entropy value H(j) of the corresponding bit: ; Select positions j where the entropy value exceeds the threshold θ to form the key feature bit set. The set of key feature bits Each bit position in the table represents a key feature bit of the protocol; A dynamic mask M is generated by aggregating all bit positions in the key feature bit set S through bitwise OR operation; Obtain the message data D to be parsed in the application layer, and perform standardization processing on the message data D: convert it into an L-bit integer, padding the high bits with zeros if it is less than L bits; Perform a bitwise AND operation on the standardized message data D and the dynamic mask M. The result bit is 1 if and only if both corresponding bits are 1, otherwise it is 0, thus obtaining the mask bit result. Extract the minimum bit position from the key feature bit set S. The mask bit result is then right-shifted by a number of bits. The aligned eigenvalues F are obtained. Perform a deep bitwise AND operation between the aligned feature value F and the feature mask corresponding to the pre-matched protocol in the template library, and determine the protocol type of the message by comparing the matching degree.
[0009] In one embodiment of the present invention, a method for generating a dynamic mask by aggregating all bit positions in the key feature bit set S through bitwise OR operation includes: For each bit position The operation shifts the binary number "1" to the left by j bits, and then combines all the results through a bitwise OR operation to generate a dynamic mask M. , in, This indicates a bitwise OR operation. This represents the operation of shifting the binary number "1" to the left by j bits.
[0010] In one embodiment of the present invention, the template library is used to store various protocol feature masks obtained by a dynamic feature mask generation algorithm. The protocol feature masks are generated based on bit entropy analysis of the protocol sample set and contain key feature bit information of the protocol. The template library dynamically updates the feature mask using an incremental learning algorithm. When a new protocol feature or a protocol version change is detected, a weighted merging method is used to fuse the old and new feature masks, and the feature masks of different protocol versions are stored in a version mapping table.
[0011] In one embodiment of the present invention, the method of fusing new and old feature masks using a weighted merging approach is as follows: , in, This represents the new feature mask generated after weighted merging at time t; This indicates the old feature mask that has been stored in the template library, i.e., the historical feature information that the current protocol already has; This represents the new feature mask generated when a new protocol feature is detected; These are the weighting coefficients.
[0012] In one embodiment of the present invention, the message data D is an L-bit integer, and if it is less than L bits, it is padded with zeros in the high bits.
[0013] In one embodiment of the present invention, S2, the method for obtaining the dynamic port association result includes: The initial data layer of the port mapping benchmark library is constructed, and the standard protocols of the substation are bound and stored with known fixed ports to form a static association table; For dynamic ports not recorded in the port mapping benchmark library, multiple cluster centers are initialized, each cluster center including... It includes the initial mean of port numbers and their corresponding traffic characteristics, and each cluster corresponds to a potential port-protocol association pattern; Set a sliding time window to continuously monitor the real-time traffic data of the port, and acquire the port traffic data points collected within the sliding time window. Calculate the port traffic data points With all historical cluster centers The distance, through Determine the closest cluster category; If the port traffic data point If the minimum distance to all historical cluster centers is less than a set distance threshold, then the port traffic data point will be... Classified into this category, and through the formula Dynamically updating cluster centers allows for real-time optimization based on changes in traffic characteristics. This represents the number of historical samples for the corresponding cluster. If the port traffic data point If the minimum distance to all historical cluster centers exceeds the aforementioned distance threshold, it is determined to be a new port pattern, a new cluster is created, and the data point is set as the initial cluster center. .
[0014] In one embodiment of the present invention, the method further includes lifecycle management of each cluster in the port mapping benchmark library, including: Count the number of samples in each cluster. If the number of samples If the cluster size is below a set threshold, it is marked as a low-frequency cluster. Extract the last update timestamp of each cluster. If the current time is the same as the stated time If the difference exceeds the timeout threshold, it is marked as a timeout cluster; Perform a deletion operation on clusters that simultaneously meet the conditions of low frequency and timeout, thereby releasing the storage space of the port mapping benchmark library.
[0015] In one embodiment of the present invention, in S2, the method for obtaining the load feature vector includes: Extract the payload data of unmatched messages, using a fixed length of N bytes as a baseline. If the payload length is less than N bytes, take the complete data as the feature extraction source. Load feature vectors are generated through multi-dimensional feature parsing, including: Perform a hash operation on the extracted payload fragment data to generate a hash value of fixed length; Count the frequency and percentage of each byte value in the first N bytes, and generate a byte frequency distribution histogram; The location of the protocol identifier bit in the payload data is identified, and the frequency of its value is counted to form a distinguishing feature.
[0016] In one embodiment of the present invention, S3, the method for generating a pre-matching list containing multiple candidate protocols includes: A two-dimensional matching model of port features and load features is established. Weight coefficients are assigned to the dynamic association results and the load feature vector. A comprehensive matching score is calculated for all potential protocol types. The comprehensive matching scores are sorted, and the top N configurable protocol types are selected to form a pre-matching list.
[0017] Based on the same inventive concept, the present invention also provides a multi-level network risk isolation system for preventing false detection in fully automatic substations. The multi-level network risk isolation system for preventing false detection in fully automatic substations includes: a message feature extraction and historical database matching module, a dynamic port analysis and load feature extraction module, a joint pre-matching list generation module, and an application layer feature mask deep recognition module. The message feature extraction and historical database matching module is configured to: receive message data from the substation network in real time; synchronously extract five-tuple information including source IP address, destination IP address, source port number, destination port number, transport layer protocol type, network layer protocol identifier, and transport layer status information from the message data; and construct a historical database storing successfully matched link information, and match the extracted information with the historical database. If the match is successful, the protocol type is obtained; If no match is found, the message feature extraction and historical database matching module is triggered. The dynamic port analysis and load feature extraction module is configured to: deploy a port mapping benchmark library and pre-configure static associations between fixed ports and protocols; collect port traffic feature data in real time based on a stream computing engine, dynamically optimize the port-protocol mapping model through an online clustering algorithm, and obtain dynamic port association results; and simultaneously perform lightweight feature extraction on the load portion of unmatched packets to obtain load feature vectors. The joint pre-matching list generation module is configured to: fuse the port dynamic association result with the load feature vector to generate a pre-matching list containing multiple candidate protocols; The application layer feature mask deep recognition module is configured to: based on the candidate protocols in the pre-matching list, perform a deep bitwise AND operation on the message data to be parsed in the application layer and the feature mask in the template library through a dynamic feature mask generation mechanism and feature extraction operation to identify the protocol type.
[0018] The technical solution of the present invention has the following advantages compared with the prior art: This invention achieves efficient processing of packets from rapid coarse screening to precise parsing by constructing a layered parallel collaborative architecture that integrates fast filtering at the network / transport layer, dynamic port-load joint pre-matching, and deep parsing of application layer feature masks. This significantly improves matching efficiency and system response speed, meeting the real-time communication needs of substations. The innovative dynamic feature masking technology leverages information entropy to quantify the importance of feature bits and adaptively extracts protocol features through feature extraction operations, overcoming the limitations of encryption technology and effectively solving the problem of encrypted packet recognition, thus improving feature utilization and recognition accuracy. The constructed intelligent dynamic learning system, based on multi-dimensional triggering conditions and distributed collaboration, achieves fully automatic and efficient updates to the rule base, reducing manual intervention and enhancing the system's adaptability to dynamic protocol changes. Through optimization techniques such as dual-hash linked list storage, multi-level caching architecture, and feature mask compression, the invention significantly reduces computational resource consumption, ensuring stable and efficient operation in resource-constrained scenarios such as edge computing devices, and expanding the application scope of the technology. Attached Figure Description
[0019] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein... Figure 1 This is a flowchart illustrating a multi-level network risk isolation method for preventing false detection in fully automated substations, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of a multi-level network risk isolation system for preventing false detection in fully automatic substations, provided in an embodiment of the present invention. Explanation of reference numerals in the accompanying drawings: 100, Message feature extraction and historical database matching module; 200, Message feature extraction and historical database matching module; 300, Joint pre-matching list generation module; 400, Application layer feature mask deep recognition module. Detailed Implementation
[0020] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.
[0021] Example 1: like Figure 1 As shown, this invention provides a multi-level network risk isolation method for preventing false detection in fully automated substations, comprising: S1: At the network / transport layer, receive message data from the substation network in real time, and synchronously extract five-tuple information including source IP address, destination IP address, source port number, destination port number, transport layer protocol type, network layer protocol identifier (such as IPv4, IPv6), and transport layer status information (such as TCP ESTABLISHED, SYN, etc.) from the message data; and build a historical database storing successfully matched link information, and match the extracted information with the historical database: If the match is successful, the protocol type is obtained, and the extracted information is determined to be trusted traffic. It is allowed to be transmitted through the normal path to avoid excessive isolation from affecting business operations. If no match is found, the extracted information will be marked as potential risk traffic and proceed to the next level of detection in step S2. S2: Deploy a port mapping baseline library and pre-configure static associations between fixed ports and protocols; collect port traffic feature data in real time based on a stream computing engine, dynamically optimize the port-protocol mapping model through an online clustering algorithm, and obtain dynamic port association results; at the same time, perform lightweight feature extraction on the load portion of unmatched packets to obtain load feature vectors; S3: Combine the port dynamic association results with the load feature vector to generate a pre-matching list containing multiple candidate protocols; S4: Based on the candidate protocols in the pre-matching list, a deep bitwise AND operation is performed on the message data to be parsed in the application layer and the feature mask in the template library through a dynamic feature mask generation mechanism and feature extraction operation to identify the protocol type.
[0022] As can be seen from the layered identification process of steps S1 to S4 in the above technical solution, the accuracy of protocol identification is gradually improved from rapid filtering to deep analysis. Multi-level network risk isolation corresponds to this process, achieving a progressive defense from coarse to fine: Level 1 isolation blocks unknown links, Level 2 isolation filters abnormal ports and loads, Level 3 isolation focuses on suspicious protocol ranges, and Level 4 isolation precisely blocks malicious packets. These two aspects form a closed loop of "identification-isolation-feedback," ensuring efficient transmission of compliant protocols while eliminating various risks at the protocol level (such as forgery, tampering, and variant attacks) through targeted isolation, ultimately achieving full-level security protection for the substation network.
[0023] Furthermore, in S1, a double-hash linked list structure is used to construct the historical database to store the link information (IP + port + protocol) of successfully matched links, ensuring fast query efficiency and efficient updates, and periodically eliminating expired records through a timestamp management mechanism.
[0024] Meanwhile, the template library serves as a core feature storage unit, used to store various protocol feature masks obtained through a dynamic feature mask generation algorithm. These protocol feature masks are generated based on bit entropy analysis of the protocol sample set: for a protocol sample sequence of fixed length L bits, the key feature bit set is filtered by calculating the information entropy of each bit position, and then aggregated through bitwise OR operations to accurately contain the core distinguishing feature bit information of the protocol.
[0025] The template library dynamically updates the feature mask using an incremental learning algorithm. When a new protocol feature is detected (e.g., an unmatched message is confirmed to belong to a new protocol through deep parsing) or a protocol version change is detected (e.g., a significant adjustment in the feature bit distribution), a weighted merging method is used to fuse the old and new feature masks, including: When the new feature mask and old feature mask The number of "1"s in the bitwise AND result (i.e., The number of "1"s in the old feature mask is less than the number of "1"s in the old feature mask. When the value is α times that of the target feature, it is determined that a significant new feature has appeared, triggering the generation of a new feature mask.
[0026] During incremental learning, a weighted merging formula is used at time t. The new and old feature masks are combined, where the weight coefficients are... The value range of β is 0.6≤β≤0.9. This setting retains the verified stable features (with higher weights) in the old feature mask and incorporates the newly added features (with lower weights) in the new feature mask, thus achieving a smooth transition and accumulation of features.
[0027] At the same time, by establishing a version mapping table Mask the basic features corresponding to different protocol versions Incremental feature mask (ΔM) between versions n The associated storage can accurately identify the current version of the protocol and backtrack to match the historical version of the protocol, ensuring the compatibility and identification of multiple version protocols and improving the template library's adaptability to protocol evolution.
[0028] Furthermore, this embodiment also provides a feature mask compression technique for template libraries, which uses a hybrid compression architecture that combines bit-plane coding and Huffman coding to optimize feature mask storage: The feature mask (binary sequence) is parsed hierarchically by bit dimension, decomposing it into multiple bit planes (e.g., an 8-bit mask is decomposed into 8 independent bit planes, each bit plane consisting of all the values of the corresponding bits in the original mask). This operation lays the foundation for subsequent targeted compression by stripping the feature distribution of different bit positions, while preserving the positional correlation of key bits in the feature mask. For each bit plane, a Huffman tree is constructed based on the probability of its bit value (0 or 1) to generate a variable-length coding table. Short codes are assigned to high-frequency bit patterns, and long codes are assigned to low-frequency patterns, achieving lossless compression of the bit plane data. This step utilizes the non-uniformity of bit distribution within the bit plane to maximize the compression ratio.
[0029] The feature mask compression technology achieves structured decomposition of the feature mask through bit-plane encoding, solving the problem of low compression efficiency caused by cross-interference of features in different bit planes in traditional overall encoding. Huffman encoding performs precise compression on the decomposed bit planes. The two work together to significantly reduce the storage resource overhead of the template library while ensuring that the key bit information of the feature mask is completely preserved (without affecting the matching accuracy of subsequent bitwise AND operations), and at the same time improve the retrieval response speed of the feature mask (the time consumed by reading and writing compressed data blocks is reduced).
[0030] This scheme employs bit-plane encoding to decompose the feature mask into bit-plane components, followed by Huffman coding for each sub-bit-plane for secondary compression. This approach significantly reduces storage resource overhead and effectively improves the storage density and retrieval response speed of the template library without compromising matching accuracy or efficiency.
[0031] Furthermore, in step S2, the method for obtaining the dynamic port association result includes: The initial data layer of the port mapping benchmark library is constructed, and the standard protocols of the substation are bound and stored with known fixed ports to form a static association table; For dynamic ports not recorded in the port mapping benchmark library, multiple cluster centers are initialized, each cluster center including... It includes the initial mean of port numbers and their corresponding traffic characteristics, and each cluster corresponds to a potential port-protocol association pattern; A 3-minute sliding time window is set to continuously monitor real-time port traffic data, and the port traffic data points collected within the sliding time window are obtained. Calculate the port traffic data points With all historical cluster centers The distance, through Determine the closest cluster category; If the port traffic data point If the minimum distance to all historical cluster centers is less than a set distance threshold, then the port traffic data point will be... It is classified into this cluster and determined by the formula. Dynamically updating cluster centers allows for real-time optimization based on changes in traffic characteristics. This represents the number of historical samples for the corresponding cluster. If the port traffic data point If the minimum distance to all historical cluster centers exceeds the aforementioned distance threshold, it is determined to be a new port pattern, a new cluster is created, and the data point is set as the initial cluster center. .
[0032] Meanwhile, to ensure the efficiency and timeliness of the port mapping benchmark library, lifecycle management of each cluster in the port mapping benchmark library is also required, including: Periodically count the number of samples in each cluster. If the number of samples If the frequency is below the set threshold, it is marked as a low-frequency cluster, indicating that the actual occurrence frequency of the port-protocol association pattern corresponding to the cluster is insufficient; Extract the last update timestamp of each cluster. If the current time is the same as the stated time If the difference exceeds the set timeout threshold, it is marked as a timeout cluster, indicating that the port-protocol association mode corresponding to the cluster has not been activated for a long time and may have become invalid. Redundant clusters that simultaneously meet the conditions of low frequency and timeout are subject to targeted deletion. By releasing invalid storage resources, the interference of outdated patterns on port-protocol association analysis is avoided. This ensures that the clusters remaining in the port mapping benchmark library are all effective association patterns with high activity and high timeliness, thereby maintaining the accuracy of dynamic port association results and the system's real-time adaptability to changes in the network environment.
[0033] Meanwhile, in step S2, lightweight feature extraction is performed on the payload portion of unmatched packets to obtain the payload feature vector. The methods include: The payload data fragments of unmatched messages are truncated as feature extraction sources, and a fixed length of N bytes is adopted for truncation (if the total length of the payload is less than N bytes, the complete payload data is directly taken) to ensure the consistency of the input dimensions for feature extraction. Load feature vectors are generated through multi-dimensional feature parsing, including: Perform hash operations (such as SHA-1 or MD5) on the preprocessed payload fragments to generate fixed-length hash values, and map variable-length payload data to standardized features for quickly comparing the similarity of payload content. The frequency and proportion of each byte value in the first N bytes of the preprocessed load segment are statistically analyzed to generate a byte frequency distribution histogram, which is then quantified into a feature vector dimension to reflect the byte composition pattern of the load data. By using preset protocol identifier rules (such as fields with specific offsets), the protocol identifier bits in the payload data are located, and the frequency and distribution patterns of their values are statistically analyzed to form distinctive identifier features, thereby enhancing the distinguishability of different protocol payloads.
[0034] By fusing the above multi-dimensional features, the generated load feature vector can not only retain the key distinguishing information of the load data, but also control the computational complexity through "lightweight" design (such as fixed-length truncation and statistical features), adapting to the real-time requirements of substation networks and providing efficient input for subsequent joint matching with port dynamic association results.
[0035] Furthermore, in this embodiment, in S3, the method for generating a pre-matching list containing multiple candidate protocols includes: A two-dimensional matching model of port features and load features is established. Based on the feature distribution characteristics of the substation network protocol, weight coefficients are assigned to the dynamic association results and the load feature vector (e.g., the port feature weight is set to ω1, the load feature weight is set to ω2, and ω1+ω2=1). For all potential protocol types, a comprehensive matching score is calculated by weighted summation: for the port feature dimension, the matching score s1 is calculated based on the dynamic association results of the ports; for the load feature dimension, the score s2 is generated based on the hash value comparison of the load feature vector, the similarity of the byte distribution, and the matching degree of the protocol identifier bit. Then the comprehensive matching score S = ω1・s1 + ω2・s2. Based on the comprehensive matching score S, the potential protocol types are sorted in descending order, and the top N protocol types (N is a configurable parameter, such as N=5) are selected to form a pre-matching list, providing an accurate candidate range for subsequent deep identification at the application layer, thus improving matching efficiency while ensuring identification coverage.
[0036] Further, in step S4, the method for identifying the protocol type includes: Based on the candidate protocols in the pre-matched protocol list, the corresponding protocol sample set is retrieved from the template library. The protocol sample set It consists of a fixed-length L-bit binary sequence, where the bit sequence of each sample is represented as B = {b0, b1, ..., b}. L-1}, where b j This represents the Boolean value for bit position j. For each bit position j, calculate the probability that the bit is 1. : , This represents the number of samples where position j is 1, and N represents the total number of samples. According to the probability The entropy value H(j) at position j is calculated using the Shannon entropy formula: ; The bits j whose entropy values exceed the threshold θ are selected to form a set of key feature bits. The set of key feature bits Each bit position in the table represents a key feature bit of the protocol; A dynamic mask M is generated by aggregating all bit positions in the key feature bit set S using a bitwise OR operation, that is: for each bit position By left shift operation Shift the binary number "1" left by j bits, and then perform a bitwise OR operation. By merging the results of operations on all bit positions, a dynamic mask M is generated: For example, when the set of key feature bits is S={2, 5, 7}, that is, the key bit positions are 2, 5, and 7, then: For each Perform a left shift operation: (Binary, the second bit is 1) (Binary, the 5th bit is 1) (Binary, the 7th bit is 1).
[0037] Perform bitwise OR aggregation: The final generated dynamic mask M has 1s in bits 2, 5, and 7, and 0s in the remaining bits, which accurately marks the key feature bits of the protocol; Obtain the message data D to be parsed in the application layer, and perform standardization processing on the message data D: convert it into an L-bit integer, padding the high bits with zeros if it is less than L bits; Perform a bitwise AND operation on the standardized message data D and the dynamic mask M. The result bit is 1 if and only if both corresponding bits are 1; otherwise, it is 0. Only the values of the key feature bits are retained to obtain the mask bit result Y. , Indicates a bitwise AND operation; Extract the minimum bit position from the key feature bit set S. The mask bit result is then right-shifted by a number of bits. The aligned eigenvalues F are obtained: , where min(S) is the position of the minimum bit in set S; The aligned feature value F is subjected to a deep bitwise AND operation with the feature mask corresponding to the pre-matched protocol in the template library. The matching degree is evaluated by calculating indicators such as Hamming distance or matching bit ratio, and the protocol type to which the message belongs is determined, so as to achieve accurate identification.
[0038] The above process uses bit entropy analysis to filter high-discrimination feature bits, and combines dynamic masking and normalized feature extraction to reduce unnecessary computational overhead while ensuring recognition accuracy, thus adapting to the real-time and accuracy requirements of substation networks for protocol recognition.
[0039] To further improve system performance, this embodiment incorporates the following innovative optimization strategies to meet the high real-time and high reliability requirements of fully automated substations: 1) Intelligent learning triggering mechanism: Construct a multi-dimensional adaptive triggering condition system to achieve dynamic iteration and accurate updates of the rule base. The specific mechanism is as follows: Based on the initial matching hit rate (with a set threshold ≤ 0.85) as the triggering condition, a weighted triggering model is established by combining key indicators such as system load rate (CPU / memory utilization), the frequency of new protocol feature emergence (the number of times unidentified protocols appear per unit time), and network traffic volatility (the ratio of the standard deviation to the mean of traffic within a sliding window). A hierarchical threshold judgment is used (e.g., increasing trigger sensitivity when the load rate is < 60% and the frequency of new protocols is > 5 times / minute) to dynamically adjust the timing and intensity of the learning strategy's activation, thus avoiding performance fluctuations caused by the learning task under high load conditions.
[0040] When the triggering conditions are met, a distributed learning task based on edge-center collaboration is initiated: edge nodes are responsible for the initial clustering and feature extraction of local traffic features, while the central server undertakes the fusion training of global protocol features and the generation of the rule base. Through lightweight data interaction (transmitting only feature vectors instead of original packets) and parallel task scheduling, incremental updates to the rule base are achieved, improving update efficiency by more than 30% compared to centralized learning, without affecting the real-time performance of the main business process.
[0041] 2) Multi-level cache optimization mechanism: A three-level cache queue architecture is constructed using the LRU-2Q (Least Recently Used-2 Queue) hybrid cache replacement algorithm to achieve efficient access to feature data and conserve computational resources. Set up a high-frequency cache queue (Q1), a medium-frequency cache queue (Q2), and a low-frequency cache queue (Q3) to store characteristic data items with high access frequency (e.g., ≥10 times per second), medium access frequency (3-10 times per second), and low access frequency (<3 times per second), respectively. Each data item x contains an access timestamp (t). x (accurate to the millisecond level) and access frequency (f x (Updated based on an exponential smoothing model).
[0042] Data item x contains an access timestamp (t) x ) and access frequency (f x The substitution rule is defined as follows: if data item x∈Q3 and f x If ≥th1, then migrate to Q2; if data item x∈Q2 and f x If the value is greater than or equal to th2, then migrate to Q1. New data is prioritized for Q3. If the upgrade condition is triggered, migrate according to the level to ensure that frequently accessed data is always in the low-latency cache queue.
[0043] The data item access frequency follows an exponential smoothing model: fx(t) = α・fx(t-1) + (1-α)・1 (α is the smoothing coefficient, with a value of 0.6 ≤ α ≤ 0.8). This model preserves historical access characteristics while also reflecting recent access trends. When the queue reaches its capacity threshold, the least recently accessed item is evicted according to the LRU criterion, and its current access frequency is determined based on f(t). x Perform cross-queue reallocation (e.g., if Q1 overflow item f) x If the value is greater than or equal to th2, it will be downgraded to Q2; otherwise, it will proceed directly to Q3, thus avoiding frequent replacement of valid data.
[0044] The above optimization strategy, through intelligent triggering and hierarchical caching design, significantly improves the real-time performance and resource utilization of the system while ensuring the accuracy of protocol identification, and adapts to the high-concurrency and low-latency operation requirements of substation networks.
[0045] Taking the IEC104 remote communication protocol commonly used in substations as an example, the first byte 0x68 in its message structure is a typical start identifier (fixed feature), but it is difficult to distinguish protocol variants or abnormal messages by relying solely on this byte.
[0046] The above technical solution, through bit entropy analysis, not only identifies the highly distinguishable feature bit 0x68 in the first byte (its entropy value is significantly higher than the threshold because this bit appears with a stable frequency in compliant messages), but also further mines the logical relationships of feature bits related to protocol functions in subsequent bytes. Examples include the "start / stop bit" in the control field and the "data length encoding bit" in the length field. These bits exhibit a regular distribution in normal communication (e.g., the start bit is only set to 1 when a connection is established), and their entropy values also satisfy the key feature bit conditions, thus being included in the feature set S.
[0047] The generated mask aggregates the key bits of the first byte 0x68 with the function bits of subsequent bytes through bitwise OR operations, forming a mask M containing multi-positional correlation information. For example, if bits 0-7 (binary 01101000) of the first byte 0x68 have a logical correlation with the "test frame identifier bit" (bit 6) of the third byte (this bit must be 1 in the test frame), then mask M will simultaneously mark these bit positions, ensuring that when performing bitwise AND operations on application layer messages, both the start identifier and the logical consistency of the function bits can be captured. This design breaks through the limitations of traditional fixed offset feature extraction, upgrading isolated features into feature chains, and significantly improving feature utilization.
[0048] The feature value F extracted from this mask can accurately reflect the compliance characteristics of the IEC104 protocol: if the message is a normal telemetry frame, F will match the combination pattern of "0x68 start bit + length bit valid value + data field check bit"; if it is an abnormal message (such as a forged start bit but with disordered function bits), the matching degree of F will be significantly reduced. This provides a highly recognizable feature input for subsequent protocol analysis (such as anomaly detection and traffic classification), ensuring that in the complex network environment of substations, it can efficiently filter compliant messages and accurately identify protocol variants or attack messages, thereby supporting the real-time performance and reliability of multi-level network risk isolation.
[0049] In summary, this application example verifies the effectiveness of the dynamic feature mask generation mechanism in the technical solution—by associating the fixed identifier of the comprehensive capture protocol with the dynamic function bits, it achieves an upgrade from single-point feature matching to multi-dimensional feature verification, providing core technical support for accurate protocol identification in substation anti-misdetection.
[0050] Example 2:
[0051] Based on the same inventive concept as Embodiment 1, this invention also provides a multi-level network risk isolation system for preventing false detection in fully automated substations, used to implement the steps of the multi-level network risk isolation method for preventing false detection in fully automated substations described in Embodiment 1. Figure 2 As shown, the multi-level network risk isolation system for preventing false detection in fully automatic substations includes: a message feature extraction and historical database matching module 100, a dynamic port analysis and load feature extraction module 200, a joint pre-matching list generation module 300, and an application layer feature mask deep recognition module 400. The message feature extraction and historical database matching module 100 is configured to: receive message data from the substation network in real time; synchronously extract five-tuple information including source IP address, destination IP address, source port number, destination port number, transport layer protocol type, network layer protocol identifier, and transport layer status information from the message data; and construct a historical database storing successfully matched link information, and match the extracted information with the historical database. If the match is successful, the protocol type is obtained; If no match is found, the message feature extraction and historical database matching module is triggered. The dynamic port analysis and load feature extraction module 200 is configured to: deploy a port mapping benchmark library and pre-configure the static association between fixed ports and protocols; collect port traffic feature data in real time based on the stream computing engine, dynamically optimize the port-protocol mapping model through an online clustering algorithm, and obtain the dynamic port association results; and simultaneously perform lightweight feature extraction on the load portion of unmatched packets to obtain load feature vectors. The joint pre-matching list generation module 300 is configured to: fuse the port dynamic association result with the load feature vector to generate a pre-matching list containing multiple candidate protocols; The application layer feature mask deep recognition module 400 is configured to: based on the candidate protocols in the pre-matching list, perform a deep bitwise AND operation on the message data to be parsed in the application layer and the feature mask in the template library through a dynamic feature mask generation mechanism and feature extraction operation to identify the protocol type.
[0052] This embodiment proposes a multi-level network risk isolation system for preventing false detection in fully automated substations. This system implements the aforementioned multi-level network risk isolation method for preventing false detection in fully automated substations. Therefore, the specific implementation of the multi-level network risk isolation system for preventing false detection in fully automated substations can be found in the embodiment section of the aforementioned method. For example, the message feature extraction and historical database matching module 100, the dynamic port analysis and load feature extraction module 200, the joint pre-matching list generation module 300, and the application layer feature mask deep recognition module 400 are respectively used to implement steps S1, S2, S3, and S4 in the multi-level network risk isolation method for preventing false detection in fully automated substations described in Embodiment 1. Therefore, its specific implementation can be referred to the descriptions of the corresponding embodiments. To avoid redundancy, further details are omitted here.
[0053] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0054] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0055] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0056] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0057] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.
Claims
1. A multi-level network risk isolation method for preventing false detection in fully automated substations, characterized in that, Includes the following steps: S1: Receive message data from the substation network in real time, and synchronously extract from the message data a five-tuple information including source IP address, destination IP address, source port number, destination port number, transport layer protocol type, network layer protocol identifier, and transport layer status information; And to build a historical database storing successfully matched link information, and match the extracted information with the historical database: If the match is successful, the protocol type is obtained; If the match fails, proceed to step S2; S2: Deploy a port mapping baseline library and pre-configure static associations between fixed ports and protocols; Port traffic feature data is collected in real time based on the stream computing engine, and the port-protocol mapping model is dynamically optimized through online clustering algorithm to obtain dynamic port association results; at the same time, lightweight feature extraction is performed on the load part of unmatched packets to obtain load feature vectors. S3: Combine the port dynamic association results with the load feature vector to generate a pre-matching list containing multiple candidate protocols; S4: Based on the candidate protocols in the pre-matching list, a deep bitwise AND operation is performed on the message data to be parsed in the application layer and the feature mask in the template library through a dynamic feature mask generation mechanism and feature extraction operation to identify the protocol type.
2. The multi-level network risk isolation method for preventing false detection in fully automated substations according to claim 1, characterized in that, In S4, the methods for identifying protocol types include: Based on the candidate protocols in the pre-matched protocol list, the corresponding protocol sample set is retrieved from the template library. The protocol sample set It consists of a fixed-length L-bit binary sequence, where the bit sequence of each sample is B = {b0, b1, ..., b}. L-1 }, where b j Represents the Boolean value at position j; For each bit position j, calculate the probability that the bit is 1. : , This represents the number of samples where position j is 1, and N represents the total number of samples. According to the probability Calculate the entropy value H(j) of the corresponding bit: ; Select positions j where the entropy value exceeds the threshold θ to form the key feature bit set. The set of key feature bits Each bit position in the table represents a key feature bit of the protocol; A dynamic mask M is generated by aggregating all bit positions in the key feature bit set S through bitwise OR operation; Obtain the message data D to be parsed in the application layer, and perform standardization processing on the message data D: convert it into an L-bit integer, padding the high bits with zeros if it is less than L bits; Perform a bitwise AND operation on the standardized message data D and the dynamic mask M. The result bit is 1 if and only if both corresponding bits are 1, otherwise it is 0, thus obtaining the mask bit result. Extract the minimum bit position from the key feature bit set S. The mask bit result is then right-shifted by a number of bits. The aligned eigenvalues F are obtained. Perform a deep bitwise AND operation between the aligned feature value F and the feature mask corresponding to the pre-matched protocol in the template library, and determine the protocol type of the message by comparing the matching degree.
3. The multi-level network risk isolation method for preventing erroneous detection in fully automated substations according to claim 2, characterized in that, The method for generating a dynamic mask by aggregating all bit positions in the key feature bit set S using bitwise OR operations includes: For each bit position The operation shifts the binary number "1" to the left by j bits, and then combines all the results through a bitwise OR operation to generate a dynamic mask M. , in, This indicates a bitwise OR operation. This represents the operation of shifting the binary number "1" to the left by j bits.
4. The multi-level network risk isolation method for preventing false detection in fully automated substations according to claim 2, characterized in that, The template library is used to store various protocol feature masks obtained by the dynamic feature mask generation algorithm. The protocol feature masks are generated based on bit entropy analysis of the protocol sample set and contain key feature bit information of the protocol. The template library dynamically updates the feature mask using an incremental learning algorithm. When a new protocol feature or a protocol version change is detected, a weighted merging method is used to fuse the old and new feature masks, and the feature masks of different protocol versions are stored in a version mapping table.
5. The multi-level network risk isolation method for preventing false detection in fully automated substations according to claim 4, characterized in that, The method for merging the old and new feature masks using a weighted merging approach is as follows: in, This represents the new feature mask generated after weighted merging at time t; This indicates the old feature mask that has been stored in the template library, i.e., the historical feature information that the current protocol already has; This represents the new feature mask generated when a new protocol feature is detected; These are the weighting coefficients.
6. The multi-level network risk isolation method for preventing false detection in fully automated substations according to claim 1, characterized in that, In S2, the methods for obtaining the dynamic port association results include: The initial data layer of the port mapping benchmark library is constructed, and the standard protocols of the substation are bound and stored with known fixed ports to form a static association table; For dynamic ports not recorded in the port mapping benchmark library, multiple cluster centers are initialized, each cluster center including... It includes the initial mean of port numbers and their corresponding traffic characteristics, and each cluster corresponds to a potential port-protocol association pattern; Set a sliding time window to continuously monitor the real-time traffic data of the port, and acquire the port traffic data points collected within the sliding time window. Calculate the port traffic data points With all historical cluster centers The distance, through Determine the closest cluster category; If the port traffic data point If the minimum distance to all historical cluster centers is less than a set distance threshold, then the port traffic data point will be... Classified into this category, and through the formula Dynamically updating cluster centers allows for real-time optimization based on changes in traffic characteristics. This represents the number of historical samples for the corresponding cluster. If the port traffic data point If the minimum distance to all historical cluster centers exceeds the aforementioned distance threshold, it is determined to be a new port pattern, a new cluster is created, and the data point is set as the initial cluster center. .
7. The multi-level network risk isolation method for preventing false detection in fully automatic substations according to claim 6, characterized in that, It also includes lifecycle management for each cluster in the port mapping benchmark library, including: Count the number of samples in each cluster. If the number of samples If the cluster size is below a set threshold, it is marked as a low-frequency cluster. Extract the last update timestamp of each cluster. If the current time is the same as the stated time If the difference exceeds the timeout threshold, it is marked as a timeout cluster; Perform a deletion operation on clusters that simultaneously meet the conditions of low frequency and timeout, thereby releasing the storage space of the port mapping benchmark library.
8. The multi-level network risk isolation method for preventing false detection in fully automated substations according to claim 1, characterized in that, In S2, the methods for obtaining the load feature vector include: Extract the payload data of unmatched messages, using a fixed length of N bytes as a baseline. If the payload length is less than N bytes, take the complete data as the feature extraction source. Load feature vectors are generated through multi-dimensional feature parsing, including: Perform a hash operation on the extracted payload fragment data to generate a hash value of fixed length; Count the frequency and percentage of each byte value in the first N bytes, and generate a byte frequency distribution histogram; The location of the protocol identifier bit in the payload data is identified, and the frequency of its value is counted to form a distinguishing feature.
9. The multi-level network risk isolation method for preventing false detection in fully automatic substations according to claim 1, characterized in that, In S3, methods for generating a pre-matching list containing multiple candidate protocols include: A two-dimensional matching model of port features and load features is established. Weight coefficients are assigned to the dynamic association results and the load feature vector. A comprehensive matching score is calculated for all potential protocol types. The comprehensive matching scores are sorted, and the top N configurable protocol types are selected to form a pre-matching list.
10. A multi-level network risk isolation system for preventing false detection in fully automated substations, characterized in that, The multi-level network risk isolation system for preventing misdetection in fully automated substations includes: a message feature extraction and historical database matching module, a dynamic port analysis and load feature extraction module, a joint pre-matching list generation module, and an application layer feature mask deep recognition module. The message feature extraction and historical database matching module is configured to: receive message data from the substation network in real time; synchronously extract five-tuple information including source IP address, destination IP address, source port number, destination port number, transport layer protocol type, network layer protocol identifier, and transport layer status information from the message data; and construct a historical database storing successfully matched link information, and match the extracted information with the historical database. If the match is successful, the protocol type is obtained; If no match is found, the message feature extraction and historical database matching module is triggered. The dynamic port analysis and load feature extraction module is configured to: deploy a port mapping benchmark library and pre-configure static associations between fixed ports and protocols; collect port traffic feature data in real time based on a stream computing engine, dynamically optimize the port-protocol mapping model through an online clustering algorithm, and obtain dynamic port association results; and simultaneously perform lightweight feature extraction on the load portion of unmatched packets to obtain load feature vectors. The joint pre-matching list generation module is configured to: fuse the port dynamic association result with the load feature vector to generate a pre-matching list containing multiple candidate protocols; The application layer feature mask deep recognition module is configured to: based on the candidate protocols in the pre-matching list, perform a deep bitwise AND operation on the message data to be parsed in the application layer and the feature mask in the template library through a dynamic feature mask generation mechanism and feature extraction operation to identify the protocol type.
Citation Information
Patent Citations
Industrial control protocol identification and analysis method based on industrial gateway
CN112788015A
Network message matching method of custom protocol application layer based on K-nearest neighbor
CN114520838A