Industrial network anomaly detection method and system based on machine learning
By conducting protocol features and timing feature analysis on real-time traffic data of industrial networks, the graph neural network and multimodal feature fusion generate device identity feature vectors, combined with watermark features and timing consistency scores, the problem of insufficient identification of camouflage attacks in the existing technology is solved, and efficient abnormal detection and continuous security protection of industrial networks are achieved.
Patent Information
- Application Number
- CN202510706147.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-08-05
AI Technical Summary
When facing complex and changing network attacks, existing industrial network anomaly detection methods are prone to false alarms or missed reports, especially lack of recognition capabilities for camouflage attacks and man-in-the-middle attacks, making it difficult to adapt to the diversified protocol environment and high hidden threats of industrial communications.
By extracting protocol features and recording communication timing feature of real-time traffic data of industrial networks, extracting device behavior patterns using graph neural networks, combining multimodal feature fusion and self-supervised learning to generate unique identity feature vectors of the device, analyzing watermark features and dynamic traffic timing consistency, and using random forest algorithms and long-term memory networks to classify abnormal traffic.
It realizes accurate identification of abnormal traffic in industrial networks, has adaptive learning ability, can dynamically update identity feature sets and feature extraction models, adapt to dynamic network environment, and improves the protection level of industrial network security.
Smart Images

Figure CN120434022A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of intelligent detection technology, and in particular relates to an industrial network anomaly detection method and system based on machine learning. Background Art
[0002] Industrial network security is a critical area that cannot be ignored in modern industrial systems. With the rapid development of the Industrial Internet, the threat posed by cyberattacks to production systems is becoming increasingly severe. Anomaly detection technology has become a crucial means of ensuring the security of industrial communications. Industrial networks not only carry real-time data exchange between devices but also affect production efficiency and economic lifeline. Therefore, building an efficient and accurate anomaly detection system is of irreplaceable strategic significance. However, traditional anomaly detection methods often rely on rule matching or static feature analysis, which often proves inadequate in the face of increasingly complex and dynamic cyberattacks. Existing solutions are prone to false positives or false negatives in dynamic environments, and are particularly incapable of identifying spoofing attacks and man-in-the-middle attacks. This directly limits the security level of industrial networks.
[0003] Existing technologies demonstrate that the limitations of industrial network anomaly detection stem primarily from insufficient in-depth analysis of network traffic characteristics and a lack of effective identification of disguised attack behaviors. These shortcomings make it difficult for systems to adapt to the diverse protocol environments and highly concealed threats found in industrial communications. Therefore, this paper proposes a machine learning-based industrial network anomaly detection method and system. Summary of the Invention
[0004] To solve the above technical problems, the present invention proposes an industrial network anomaly detection method and system based on machine learning to solve the problems existing in the above-mentioned prior art.
[0005] To achieve the above objectives, the present invention provides an industrial network anomaly detection method based on machine learning, comprising:
[0006] The protocol features of real-time traffic data in the industrial network are extracted and the communication timing features are recorded to obtain a comprehensive traffic feature set;
[0007] Determine a unique identity feature vector for each device based on the comprehensive traffic feature set to obtain a digital identity of the device;
[0008] Based on the digital identity, watermark features in the real-time traffic data are judged to obtain a preliminary abnormality mark;
[0009] The dynamic traffic in real-time traffic data is analyzed based on the long short-term memory network to obtain the temporal consistency score;
[0010] A classification label of abnormal traffic is obtained based on the preliminary abnormality mark and the temporal consistency score.
[0011] Optionally, the process of extracting protocol features from real-time traffic data in the industrial network and recording communication timing features to obtain a comprehensive traffic feature set includes:
[0012] Parsing the data packet header of the real-time traffic data to obtain a protocol feature set;
[0013] According to the protocol feature set, the protocol type is determined in combination with the data packet header information to obtain a protocol type set;
[0014] Performing communication time series feature annotation on the real-time traffic data by using a timestamp to obtain a time series feature set;
[0015] A feature extraction method is used to fuse the protocol feature set and the timing feature set to obtain a comprehensive traffic feature set.
[0016] Optionally, the process of determining the unique identity feature vector of each device based on the comprehensive traffic feature set to obtain the digital identity of the device includes:
[0017] Using a graph neural network to extract implicit features of the comprehensive traffic feature set;
[0018] Separating the device behavior pattern from the protocol features and communication timing features of the comprehensive traffic feature set based on the implicit features to obtain a device fingerprint representation;
[0019] Based on the device fingerprint representation, a unique identity feature vector for each device is obtained by using multimodal feature fusion and self-supervised learning;
[0020] A digital identity of the device is determined based on the unique identity feature vector of each device.
[0021] Optionally, the process of obtaining the device fingerprint representation includes:
[0022] Based on the implicit features, separate independent patterns from the protocol features to obtain a protocol behavior set;
[0023] When the protocol behavior set matches the communication timing, the timing regularity is marked by the timestamp to obtain the behavior pattern sequence;
[0024] Separating device behaviors from the behavior pattern sequence to obtain device-specific patterns;
[0025] Using feature separation technology to perform device fingerprint determination on the specific mode of the device to obtain a preliminary fingerprint determination result;
[0026] A device fingerprint representation is determined based on the fingerprint preliminary judgment result and the timestamp annotation information.
[0027] Optionally, based on the device fingerprint representation, a process of obtaining a unique identity feature vector for each device using multimodal feature fusion and self-supervised learning includes:
[0028] Processing the device fingerprint representation and device communication data using a multimodal technique to obtain an initial feature representation;
[0029] Optimizing the initial feature representation based on a self-supervised method to obtain an optimized identity representation;
[0030] Adjust the optimized identity representation based on feature fusion technology to obtain an adjusted vector;
[0031] For the adjusted vector, verify the consistency through the preset set and obtain the verification result;
[0032] If the consistency of the verification results exceeds a preset threshold, a unique identity feature vector for each device is determined.
[0033] Optionally, the process of determining the watermark feature in the real-time traffic data based on the digital identity to obtain a preliminary abnormality mark includes:
[0034] Obtain the watermark features of real-time traffic data to obtain independent watermark features;
[0035] Extracting normal watermark features of the digital identity recognition and generating standard feature representations from the normal watermark features using vector mapping technology;
[0036] Calculating the distance between the independent watermark feature and the standard feature representation, and when the distance exceeds a preset threshold, the independent watermark feature is a feature deviation and obtaining a deviation flag;
[0037] Extracting a traffic segment associated with the deviation identifier from the real-time traffic data based on the deviation identifier;
[0038] A preliminary anomaly mark is obtained based on the traffic segment associated with the deviation identifier.
[0039] Optionally, the process of analyzing the dynamic traffic in the real-time traffic data based on the long short-term memory network to obtain the temporal consistency score includes:
[0040] Long short-term memory network is used to process the communication timing features and device fingerprint representation in real-time traffic data to determine the timing consistency score.
[0041] Optionally, the process of obtaining a classification label for abnormal traffic based on the preliminary abnormality mark and the temporal consistency score includes:
[0042] A random forest algorithm is used to fuse the preliminary anomaly mark and the temporal consistency score to obtain a traffic classification result;
[0043] Extract abnormal traffic data from traffic classification results and determine abnormal traffic areas through weighted calculation;
[0044] According to the abnormal traffic area, a clustering algorithm is used to divide the abnormal mark set;
[0045] Get the updated classification label value through the abnormal label set;
[0046] Update the classification label value, integrate the abnormal traffic data, and determine the final abnormal traffic sequence;
[0047] The associated temporal consistency features are obtained from the final abnormal traffic sequence to determine the abnormality identification result.
[0048] The present invention also provides an industrial network anomaly detection system based on machine learning, which is used to implement an industrial network anomaly detection method based on machine learning. The system includes:
[0049] A traffic feature extraction module, configured to obtain real-time traffic data in an industrial network and obtain a comprehensive traffic feature set based on the real-time traffic data;
[0050] A device fingerprint generation module is used to extract implicit features based on the comprehensive traffic feature set using a graph neural network, separate device behavior patterns from protocol features and communication timing features, and obtain a device fingerprint representation;
[0051] The identity feature fusion module is used to obtain the unique identity feature vector of each device through multimodal feature fusion and self-supervised learning based on the device fingerprint representation and the pre-established normal communication data set;
[0052] A watermark detection module, configured to determine watermark features in real-time traffic data based on the digital identity identifier to obtain a preliminary abnormality mark;
[0053] The timing consistency analysis module is used to analyze the communication timing characteristics and device fingerprint representation in dynamic traffic through the long short-term memory network to obtain the timing consistency score;
[0054] The abnormal traffic classification module is used to fuse the preliminary abnormality mark and the time series consistency score using the random forest algorithm, output the abnormal traffic identification result through weighted calculation, and obtain the abnormal traffic classification label;
[0055] The watermark feature optimization module is used to extract the characteristic distribution of the forged watermark from the abnormal traffic identification results, and adjust the normal watermark features in the digital identity through an adaptive update mechanism to obtain the optimized identity feature set;
[0056] The feature extraction model update module is used to adjust the parameters of the graph neural network using online learning methods based on the optimized identity feature set, enhance the robustness of feature learning, and obtain a feature extraction model that adapts to dynamic environments;
[0057] The network security status update module is used to process subsequent traffic data through a feature extraction model, and output abnormal traffic identification results in real time by combining an autoencoder and self-attention mechanism to obtain a continuously updated industrial network security status.
[0058] Compared with the prior art, the present invention has the following advantages and technical effects:
[0059] The present invention discloses a method for identifying abnormal traffic in industrial networks. The method analyzes the protocol features and communication timing features of real-time traffic data packets, uses graph neural networks to extract device behavior patterns, and generates a unique device identity feature vector in combination with a pre-established normal communication data set. On this basis, the present invention achieves accurate identification of abnormal traffic by comparing watermark features, analyzing dynamic traffic timing features, and integrating multiple algorithms. At the same time, the present invention also has adaptive learning capabilities, and can dynamically update the identity feature set and feature extraction model based on the recognition results, thereby continuously optimizing the industrial network security status assessment. This method can not only effectively identify abnormal traffic such as forged watermarks, but also adapt to dynamically changing network environments, providing comprehensive and continuous protection for industrial network security. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0061] Figure 1 This is a flow chart of an industrial network anomaly detection method based on machine learning according to an embodiment of the present invention;
[0062] Figure 2 2 is a system structure diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0063] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0064] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0065] Example 1
[0066] like Figure 1 As shown, this embodiment provides an industrial network anomaly detection method based on machine learning, including the following steps:
[0067] S101. Extract protocol features from real-time traffic data in the industrial network and record communication timing features to obtain a comprehensive traffic feature set.
[0068] Furthermore, the process of extracting protocol features from real-time traffic data in an industrial network and recording communication timing features to obtain a comprehensive traffic feature set includes: parsing the data packet header of the real-time traffic data to obtain a protocol feature set; determining the protocol type based on the protocol feature set in combination with the data packet header information to obtain a protocol type set; annotating the real-time traffic data with communication timing features through a timestamp to obtain a timing feature set; and using a feature extraction method to fuse the protocol feature set and the timing feature set to obtain a comprehensive traffic feature set.
[0069] Furthermore, as a specific implementation of this embodiment, a network capture tool, such as Wireshark, is deployed in a factory's industrial control system to monitor inter-device communication traffic in real time. For example, in an automated production line scenario, the tool can capture data interactions between PLCs and sensors, generating a raw traffic dataset. This dataset typically contains numerous data packets, recording all the details of device communication.
[0070] In one possible implementation, parsing the packet header in the original traffic data set requires focusing on key fields such as the IP address, port number, and protocol identifier.
[0071] For example, extracting information from the packet header showing source IP address 192.168.1.10, destination IP address 192.168.1.20, and port number 502 may indicate typical characteristics of the Modbus protocol. By analyzing these fields and combining them with the protocol's message structure, protocol features can be extracted and formed into a protocol signature set. This signature set reflects the specific protocol type used in the communication and lays the foundation for subsequent analysis.
[0072] Specifically, when determining the protocol type, a match can be performed based on the packet header information and the protocol feature set.
[0073] In one embodiment, if the protocol signature set shows a fixed packet length of a fixed number of bytes and the port number is 1883, then it is determined to be the MQTT protocol; if the port number is 502 and the packet has a function code field, then it is confirmed to be the Modbus protocol. This will result in a set of protocol types that clearly mark the active protocols on the network.
[0074] As a specific implementation of this embodiment, in the captured traffic, a sensor sends a data request to the PLC every 500 milliseconds. This regular time interval is a time series feature. By annotating it with timestamps, a time series feature set is generated to reflect the rhythm and frequency of communication.
[0075] Preferably, a feature extraction method is used when fusing the protocol feature set and the timing feature set.
[0076] In one embodiment, the port number and protocol type in the protocol feature set are integrated with the time interval and frequency in the timing feature set to form a comprehensive traffic feature set.
[0077] For example, the communication of Modbus protocol shows a request every 200 milliseconds. Combining the protocol type and timing regularity, the traffic pattern can be fully described.
[0078] In industrial networks, analyzing a comprehensive set of signatures can reveal an abnormally high frequency of device communication, potentially indicating a potential fault or attack. Compared to single-signature analysis, this approach can verify the problem from multiple perspectives and reduce false positives.
[0079] In one embodiment, if a sudden surge in MQTT protocol traffic is found with disordered timing, it may indicate that the device is maliciously controlled, and timely warning can effectively improve network security.
[0080] For example, from another perspective, feature fusion can also optimize bandwidth management. For example, if a factory network runs both Modbus and MQTT protocols, comprehensive traffic feature set analysis reveals that Modbus occupies a large amount of bandwidth and has frequent communications. Adjusting its transmission frequency can improve overall network efficiency. This multi-faceted analysis ensures the practicality and reliability of the solution while enhancing the stability and security of the industrial network.
[0081] S102. For the comprehensive traffic feature set, a graph neural network is used to extract implicit features, and the device behavior pattern is separated from the protocol features and communication timing features to obtain the device fingerprint representation.
[0082] Furthermore, the process of obtaining the device fingerprint representation includes: based on the implicit features, separating independent patterns from the protocol features to obtain a protocol behavior set; when the protocol behavior set matches the communication timing, the timing rules are marked by timestamps to obtain a behavior pattern sequence; separating the device behavior from the behavior pattern sequence to obtain a device-specific pattern; using feature separation technology to perform device fingerprint judgment on the device-specific pattern to obtain a preliminary fingerprint judgment result; and determining the device fingerprint representation based on the preliminary fingerprint judgment result and timestamp marking information.
[0083] Furthermore, as a specific implementation of this example, in an industrial control network, traffic data is captured continuously for 5 minutes through a network interface, resulting in a collection of thousands of data packets. These packets carry information such as source address, destination address, and protocol identifier. When processing data using a graph neural network, the relationships between data packets can be modeled as a graph structure, with nodes representing devices and edges representing communication relationships.
[0084] In one possible implementation, assuming there are 10 devices in the network, a graph neural network uses three layers of convolution to compress each node's feature vector from its original 64-dimensional representation to a 32-dimensional latent feature representation. This representation preserves the communication patterns between devices, facilitating subsequent analysis. The latent feature representation is then used to separate independent patterns.
[0085] On an automated production line, one device sends status updates every 10 seconds, while another sends alerts only when an anomaly occurs. By analyzing the time intervals and message types in the sequence, we can isolate the two device-specific patterns.
[0086] Preferably, in combination with feature separation technology, when determining whether the device fingerprint exists, the unique identification field in the pattern can be checked.
[0087] In one embodiment, if a device's message header consistently contains the fixed sequence number 0x1234, its fingerprint is initially determined to exist. The initial fingerprint determination is then combined with the timestamp to further confirm that the device is active between 8:00 AM and 5:00 PM daily, generating a device fingerprint. A final behavioral identifier is then derived from the device fingerprint.
[0088] It should be noted that this technology may construct an identity by counting the frequency of device fingerprints at different time periods.
[0089] For example, a device might communicate five times per minute during peak production hours, but only once per minute during off-peak hours. The resulting behavioral signature not only reflects the device's identity but also its operating patterns. This signature can be used for device status monitoring or anomaly detection in industrial networks, possessing significant practical value.
[0090] In one possible implementation, if it is expanded to a multi-device scenario, a single business logic can still be maintained.
[0091] S103. Based on the device fingerprint representation and in combination with a pre-established normal communication data set, a unique identity feature vector for each device is obtained through multimodal feature fusion and self-supervised learning.
[0092] Furthermore, based on the device fingerprint representation, the process of using multimodal feature fusion and self-supervised learning to obtain a unique identity feature vector for each device includes: using multimodal technology to process the device fingerprint representation and device communication data to obtain an initial feature representation; optimizing the initial feature representation based on the self-supervised method to obtain an optimized identity representation; adjusting the optimized identity representation based on feature fusion technology to obtain an adjusted vector; for the adjusted vector, verifying consistency through a preset set to obtain a verification result; if the consistency of the verification result exceeds a preset threshold, determining the unique identity feature vector of each device.
[0093] Furthermore, as a specific implementation method of this embodiment, the IP address, port number and data packet sending frequency of the device are used as input, and text features and time series features are extracted respectively through multimodal technology, and then mapped to a common vector space. Assuming that the communication data of a device contains 10 data packets sent per second, and the IP address is a fixed value, then the initial feature representation is a multidimensional vector containing this information. The advantage of this method is that it can fully capture the communication characteristics of the device and provide rich basic information for subsequent analysis. For the initial feature representation, the self-supervision method is integrated for optimization to obtain the optimized identity representation. The principle of the self-supervision method is to learn through the laws of the data itself without the need for additional labeling.
[0094] Using the time intervals and packet sizes in the communication data, a model is trained to predict the arrival time of the next packet. If the predicted result is consistent with the actual result, the feature representation is optimized to a more discriminative identity representation.
[0095] Specifically, assuming a device's transmission interval is stable at 0.1 seconds, the optimized identity representation, through self-supervised learning, can highlight this regularity, thereby enhancing the uniqueness of the identity. The key to this step is improving the robustness of the feature, making it more representative of the device's core behavior. If the identity representation is consistent with the unique vector, data processing is used to determine the vector's integrity and confirm the device's identity.
[0096] In one embodiment, a vector is considered complete if the identity representation includes the device's MAC address, transmission frequency, and protocol type, and these information are not missing.
[0097] For example, a device's MAC address is AA:BB:CC:DD:EE:FF, and it transmits five times per second. By comparing this data against a pre-defined unique vector library, if a match occurs, the device can be identified as a specific device. This approach effectively reduces false positives and ensures accurate identification. By comparing the communication data with a normal dataset, feature fusion technology is used to adjust the identity representation, resulting in an adjusted vector.
[0098] Preferably, the actual communication data of the device can be compared with the historical normal data to find out the abnormal points.
[0099] For example, if a normal data set shows a device's average transmission interval of 0.2 seconds, while the current data shows 0.05 seconds, feature fusion technology can incorporate this abnormal feature into the identity representation. The resulting adjusted vector more accurately reflects the device's status. This adjustment allows for dynamic adaptation to changes in device behavior, improving the reliability of subsequent verification. The adjusted vector is then verified for consistency using a pre-set set to obtain the verification result.
[0100] For example, if the typical behavior of a device in the template library is to send 100 packets per minute, and the adjusted vector shows that the current device sends 98 packets per minute, the verification result shows a 98% consistency. The advantage of this verification method is that it uses quantitative indicators to clearly determine whether the device behavior meets expectations, providing a basis for the final judgment. If the verification results meet the requirements, the optimization method and multimodal technology are integrated to determine the final unique identity feature vector.
[0101] In one possible implementation, the features after self-supervised optimization can be further integrated with the temporal features extracted by multimodality.
[0102] For example, if the verification result consistency exceeds 95%, the device's transmission frequency of 0.1 seconds, protocol type TCP, and other information are integrated into a final vector. The advantage of this final feature vector is that it retains the static identity of the device while reflecting its dynamic behavior, significantly improving the accuracy and stability of device identification.
[0103] S104: Acquire watermark features in real-time traffic. If the watermark features deviate from a preset threshold value from the normal watermark features in the digital identity identifier, it is determined to be a forged watermark and a preliminary abnormality mark is obtained.
[0104] Furthermore, the process of judging the watermark features in the real-time traffic data based on the digital identity identification to obtain a preliminary abnormality mark includes: obtaining the watermark features of the real-time traffic data to obtain an independent watermark feature; extracting the normal watermark features of the digital identity identification, and using vector mapping technology to generate a standard feature representation of the normal watermark features; calculating the distance between the independent watermark feature and the standard feature representation, when the distance exceeds a preset threshold, the independent watermark feature is a feature deviation and a deviation mark is obtained; based on the deviation mark, a traffic segment associated with the deviation mark is extracted from the real-time traffic data; and a preliminary abnormality mark is obtained based on the traffic segment associated with the deviation mark.
[0105] Furthermore, as a specific implementation of this embodiment, in a real-world scenario, assuming real-time traffic of 1,000 packets per second, the extraction method targets 10 packets with specific identifiers, forming an independent set. This approach facilitates subsequent analysis and quickly focuses on key data. For independent watermark feature sets, extracting normal watermark features from digital identity identifiers is the basis for verifying authenticity.
[0106] Specifically, the normal watermark feature is a pre-stored device signature, such as a standard pattern represented by a 32-bit vector. When the vector mapping technology is used to generate the standard feature representation, the extracted watermark feature can be projected into the same vector space.
[0107] As a further implementation of this embodiment, the features of an independent watermark set are [0.8, 0.3, 0.5], while the normal features are [0.9, 0.2, 0.6]. Through mapping, a standardized comparison benchmark is obtained. This method helps unify the format and improves comparison efficiency. If the distance between the standard feature representation and the independent watermark feature set exceeds a preset threshold, a feature deviation is determined.
[0108] In one possible implementation, assuming a threshold of 0.2, the Euclidean distance between two vectors is calculated to be 0.25, exceeding the threshold and flagged as a deviation. The significance of the deviation flag is to indicate possible forgery or anomaly, requiring further investigation. This simple and straightforward decision logic effectively screens out anomalies.
[0109] Preferably, the extracted segments include 100 data packets before and after the deviation occurs. When using a clustering algorithm to divide the abnormal traffic area, such as using a K-means method, the data packets can be grouped by time or characteristics to identify the area where the abnormality is concentrated.
[0110] For example, clustering results showing an abnormally high packet frequency over a certain period of time could indicate a counterfeit watermark. This division helps narrow the scope of investigation. For areas with abnormal traffic, comparing watermark features with normal watermarks using an identity database is a key verification step.
[0111] In one embodiment, the match evaluation value is expressed as a percentage. While the match for a normal watermark is 90%, the match for the abnormal region is only 40%, below the preset threshold of 60%. This indicates the possibility of forgery. By integrating the tagging judgment logic with the threshold comparison results, the abnormality flags for forged watermarks are determined, which can more accurately locate the source of the problem and improve system security. Using abnormality flags, the data segments corresponding to the forged watermarks are separated from the real-time traffic as the final step.
[0112] For example, if the separated data segment consists of 50 packets within 10 seconds, analysis of its characteristics reveals an unusually concentrated distribution of timestamps, confirming it to be a forged watermark. The final abnormal traffic characteristics obtained can be used to optimize traffic monitoring. This separation method not only isolates problematic data but also provides a basis for system improvements.
[0113] S105. Analyze the communication timing characteristics and device fingerprint representation in dynamic traffic through the long short-term memory network to determine whether the traffic sequence conforms to the normal behavior pattern and obtain a timing consistency score.
[0114] The communication timing and device fingerprints in dynamic traffic are processed using a long short-term memory network to generate a time series feature sequence. Based on this time series feature sequence, a comparative calculation method is used to match it with the preset behavior pattern to obtain a deviation value. If the deviation value exceeds the preset threshold, the deviation calculation is used to determine the abnormal time series segment. For this abnormal time series segment, the associated communication timing data is extracted from the dynamic traffic to generate an abnormal feature set. The abnormal feature set is processed using a clustering algorithm to delineate abnormal behavior regions. Based on the abnormal behavior region, the device fingerprint is compared with the normal behavior pattern to obtain a behavior consistency assessment value. If the behavior consistency assessment value is lower than the preset threshold, the deviation value and the assessment value are combined to determine the abnormal traffic sequence.
[0115] Exemplarily, the communication timing and device fingerprints in dynamic traffic are processed by a long short-term memory network to generate a timing feature sequence.
[0116] For the time series feature sequence, a comparative calculation method is used to match it with the preset behavior pattern to obtain the deviation value. Specifically, the generated time series feature sequence is compared with the standard behavior pattern of the device stored in the database. The standard behavior pattern is "sending frequency 6-10 times / second, interval 0.1-0.3 seconds", while the real-time sequence is displayed as "15 times / second, interval 0.5 seconds". Through comparative calculation, the deviation value may reach 0.7, exceeding the preset threshold of 0.5. This indicates that the traffic behavior may be abnormal, which helps to detect potential problems in a timely manner. If the deviation value exceeds the preset threshold, the abnormal time series segment is determined through deviation calculation.
[0117] In one embodiment, a specific time window can be extracted from the complete time series. For example, in the example above, the 5-second segment corresponding to "15 times / second, 0.5-second interval" is marked as an anomaly. This approach has the advantage of quickly locating the problem area and improving the efficiency of subsequent analysis.
[0118] Extract related communication time series data from dynamic traffic to generate an abnormal feature set. It should be noted that this step also includes the context data of the abnormal fragment in the analysis.
[0119] For example, the traffic rate and packet size before and after the abnormal segment are combined into a single set, such as "rate 20 Mbps, packet size 1500 bytes." This provides a more comprehensive description of the abnormal characteristics and facilitates subsequent processing. Clustering algorithms are used to process the abnormal feature set and identify areas of abnormal behavior.
[0120] Preferably, an unsupervised clustering method is used to identify abnormal behavior areas. For example, assuming the abnormal feature set includes traffic data from multiple devices, clustering may result in two categories: "high-frequency transmission areas" and "low-frequency interruption areas." This helps distinguish different types of abnormal behavior and improves identification accuracy.
[0121] In a possible implementation, when comparing a device fingerprint with a normal behavior pattern based on an abnormal behavior area to obtain a behavior consistency evaluation value, the device fingerprint is a MAC address or a communication protocol feature.
[0122] For example, if a device uses protocol A in normal mode but displays protocol B in the abnormal region, the consistency assessment value may drop to 0.3, which is below the threshold of 0.6. This indicates that the device behavior deviates from expectations and may be counterfeit or faulty.
[0123] If the behavior consistency assessment value falls below the preset threshold, the deviation value is combined with the assessment value to determine an abnormal traffic sequence. For example, a deviation value of 0.7 combined with an assessment value of 0.3 might yield a comprehensive anomaly score of 0.85, confirming that the traffic sequence is abnormal. This multi-dimensional verification reduces the risk of misjudgment.
[0124] Specifically, it can be separated from the overall traffic flow and analyzed separately. For example, the data stream corresponding to the "high-frequency sending area" can be extracted to check for malicious injections. The advantage of this method is that it can effectively isolate problematic data and protect the stability of the overall system operation.
[0125] S106. Use the random forest algorithm to fuse the preliminary anomaly mark and the temporal consistency score, output the abnormal traffic identification result through weighted calculation, and obtain the classification label of the abnormal traffic.
[0126] Furthermore, the process of obtaining a classification label for abnormal traffic based on the preliminary abnormal mark and the temporal consistency score includes: using a random forest algorithm to fuse the preliminary abnormal mark and the temporal consistency score to obtain a traffic classification result; extracting abnormal traffic data from the traffic classification result, and determining the abnormal traffic area through weighted calculation; using a clustering algorithm to divide the abnormal mark set according to the abnormal traffic area; obtaining a classification label update value through the abnormal mark set; based on the classification label update value, fusing the abnormal traffic data to determine the final abnormal traffic sequence; obtaining associated temporal consistency features from the final abnormal traffic sequence to determine the abnormal identification result.
[0127] Furthermore, as a specific implementation of this example, if a region experiences significant traffic fluctuations, with a variance of 0.5, while normal traffic variance is typically less than 0.2, this region can be identified as an abnormal region after comprehensive scoring. This classification can improve the targeted nature of subsequent analysis. Based on the abnormal traffic regions, a clustering algorithm is used to create a set of abnormal markers. This process aims to further refine and categorize abnormal behavior.
[0128] Preferably, the K-means algorithm is used to cluster abnormal traffic into three categories according to time series characteristics: high-frequency short packets, low-frequency large packets, and mixed types.
[0129] For example, high-frequency short packets correspond to scanning behavior, while low-frequency large packets are associated with data leakage. The classification results provide a basis for subsequent processing. By comparing the set of abnormal tags with the preset behavior pattern, the updated classification label value is obtained. This step is a dynamic adjustment of the abnormal type.
[0130] For example, if the default behavior pattern specifies that a normal device sends no more than five packets per second, while a specific anomaly tag set indicates 10 packets per second, the label is updated to "High-Frequency Anomaly." This update allows the system to more accurately identify threat types. The classification label is updated, anomaly traffic data is integrated, and the final anomaly traffic sequence is determined. This process emphasizes the integration of multi-source data.
[0131] It's understandable that combining the updated "High-Frequency Anomaly" tag with traffic data reveals a 10-minute period of abnormal traffic flow, suggesting a possible malicious attack and ultimately confirming it as an abnormal sequence. This integration improves the reliability of the judgment. The associated temporal consistency features are extracted from the final abnormal traffic sequence to determine the anomaly identification result.
[0132] S107. Extract the characteristic distribution of the forged watermark from the abnormal traffic identification result, adjust the normal watermark characteristics in the digital identity through an adaptive update mechanism, and obtain an optimized identity feature set.
[0133] The characteristic distribution of counterfeit watermarks is extracted through abnormal traffic analysis to determine the initial abnormal feature set. An adaptive method is used to adjust the initial abnormal feature set to obtain a revised feature distribution. The identification information of the digital identity is obtained from the revised feature distribution to determine the baseline features of the normal watermark. If there is a deviation between the baseline features of the normal watermark and the characteristic distribution of the counterfeit watermark, the identity features are updated through feature adjustment. Based on the optimized results of the updated identity features, the boundary range of the abnormal traffic is determined. Based on the boundary range of the abnormal traffic, a clustering algorithm is used to divide the distribution area of the counterfeit watermark to obtain the final feature set. The digital identity after feature adjustment is obtained from the final feature set to determine the optimized identity feature set.
[0134] S108. For the optimized identity feature set, the online learning method is used to adjust the parameters of the graph neural network to enhance the robustness of feature learning and obtain a feature extraction model that adapts to dynamic environments.
[0135] Through online learning, real-time changes in identity features are captured, and the graph neural network parameters are adjusted to obtain an updated network structure. Intermediate results of feature learning are extracted from the updated network structure to determine a feature distribution with enhanced robustness. Feature distributions are used to determine changing trends in dynamic environments, resulting in a more adaptable feature extraction method. Network parameters are adjusted based on the feature extraction method to obtain an optimized graph neural network model. The impact of environmental changes on the optimized model is analyzed to determine the stable range of feature learning. The final representation of identity features is obtained from this stable range, resulting in a feature extraction result that adapts to dynamic environments.
[0136] S109. Process subsequent traffic data through the feature extraction model, combine the autoencoder and self-attention mechanism to output abnormal traffic identification results in real time, and obtain a continuously updated industrial network security status.
[0137] Traffic data is initially processed using a feature extraction model, and combined with an autoencoder to compress data dimensions, a streamlined feature representation is obtained. A self-attention mechanism is used to perform weighted calculations on this streamlined feature representation to obtain a preliminary distribution of abnormal traffic. Based on this preliminary distribution, if abnormal traffic exceeds a preset threshold, the model parameters are dynamically adjusted to obtain an updated recognition result. Based on this updated recognition result, the changing trend of abnormal traffic in the industrial network is determined to determine the current security status. Based on this changing security status, the traffic data is re-encoded using an autoencoder to obtain a feature set adapted to dynamic environments. Based on this feature set adapted to dynamic environments, the latest distribution of abnormal traffic is output in real time to obtain a continuously updated security status. Key indicators are extracted from this continuously updated security status to assess the operational stability of the industrial network.
[0138] Example 2
[0139] like Figure 2 As shown, this embodiment provides an industrial network anomaly detection system based on machine learning, including:
[0140] A traffic feature extraction module, configured to obtain real-time traffic data in an industrial network and obtain a comprehensive traffic feature set based on the real-time traffic data;
[0141] A device fingerprint generation module is used to extract implicit features based on the comprehensive traffic feature set using a graph neural network, separate device behavior patterns from protocol features and communication timing features, and obtain a device fingerprint representation;
[0142] The identity feature fusion module is used to obtain the unique identity feature vector of each device through multimodal feature fusion and self-supervised learning based on the device fingerprint representation and the pre-established normal communication data set;
[0143] A watermark detection module, configured to determine watermark features in real-time traffic data based on the digital identity identifier to obtain a preliminary abnormality mark;
[0144] The timing consistency analysis module is used to analyze the communication timing characteristics and device fingerprint representation in dynamic traffic through the long short-term memory network to obtain the timing consistency score;
[0145] The abnormal traffic classification module is used to fuse the preliminary abnormality mark and the time series consistency score using the random forest algorithm, output the abnormal traffic identification result through weighted calculation, and obtain the abnormal traffic classification label;
[0146] The watermark feature optimization module is used to extract the characteristic distribution of the forged watermark from the abnormal traffic identification results, and adjust the normal watermark features in the digital identity through an adaptive update mechanism to obtain the optimized identity feature set;
[0147] The feature extraction model update module is used to adjust the parameters of the graph neural network using online learning methods based on the optimized identity feature set, enhance the robustness of feature learning, and obtain a feature extraction model that adapts to dynamic environments;
[0148] The network security status update module is used to process subsequent traffic data through a feature extraction model, and output abnormal traffic identification results in real time by combining an autoencoder and self-attention mechanism to obtain a continuously updated industrial network security status.
[0149] The above are merely preferred embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A machine learning-based industrial network anomaly detection method, characterized in that: The following steps are involved: The protocol features of real-time traffic data in the industrial network are extracted and the communication timing features are recorded to obtain a comprehensive traffic feature set; Determine a unique identity feature vector for each device based on the comprehensive traffic feature set to obtain a digital identity of the device; Based on the digital identity, watermark features in the real-time traffic data are judged to obtain a preliminary abnormality mark; The dynamic traffic in real-time traffic data is analyzed based on the long short-term memory network to obtain the temporal consistency score; A classification label of abnormal traffic is obtained based on the preliminary abnormality mark and the temporal consistency score.
2. The method for detecting anomalies in an industrial network based on machine learning according to claim 1, wherein: The process of extracting protocol features from real-time traffic data in industrial networks and recording communication timing features to obtain a comprehensive traffic feature set includes: Parsing the data packet header of the real-time traffic data to obtain a protocol feature set; According to the protocol feature set, the protocol type is determined in combination with the data packet header information to obtain a protocol type set; Performing communication time series feature annotation on the real-time traffic data by using a timestamp to obtain a time series feature set; A feature extraction method is used to fuse the protocol feature set and the timing feature set to obtain a comprehensive traffic feature set.
3. The method for detecting anomalies in an industrial network based on machine learning according to claim 1, wherein: The process of determining the unique identity feature vector of each device based on the comprehensive traffic feature set to obtain the digital identity of the device includes: Using a graph neural network to extract implicit features of the comprehensive traffic feature set; Separating the device behavior pattern from the protocol features and communication timing features of the comprehensive traffic feature set based on the implicit features to obtain a device fingerprint representation; Based on the device fingerprint representation, a unique identity feature vector for each device is obtained by using multimodal feature fusion and self-supervised learning; A digital identity of the device is determined based on the unique identity feature vector of each device.
4. The method for detecting anomalies in an industrial network based on machine learning according to claim 3, wherein: The process of obtaining the device fingerprint representation includes: Based on the implicit features, separate independent patterns from the protocol features to obtain a protocol behavior set; When the protocol behavior set matches the communication timing, the timing regularity is marked by the timestamp to obtain the behavior pattern sequence; Separating device behaviors from the behavior pattern sequence to obtain device-specific patterns; Using feature separation technology to perform device fingerprint determination on the specific mode of the device to obtain a preliminary fingerprint determination result; A device fingerprint representation is determined based on the fingerprint preliminary judgment result and the timestamp annotation information.
5. The method for detecting anomalies in an industrial network based on machine learning according to claim 4, wherein: Based on the device fingerprint representation, the process of obtaining a unique identity feature vector for each device using multimodal feature fusion and self-supervised learning includes: Processing the device fingerprint representation and device communication data using a multimodal technique to obtain an initial feature representation; Optimizing the initial feature representation based on a self-supervised method to obtain an optimized identity representation; Adjust the optimized identity representation based on feature fusion technology to obtain an adjusted vector; For the adjusted vector, verify the consistency through the preset set and obtain the verification result; If the consistency of the verification results exceeds a preset threshold, a unique identity feature vector for each device is determined.
6. The method for detecting anomalies in an industrial network based on machine learning according to claim 5, characterized in that: The process of determining the watermark features in the real-time traffic data based on the digital identity identifier to obtain a preliminary abnormality mark includes: Obtain the watermark features of real-time traffic data to obtain independent watermark features; Extracting normal watermark features of the digital identity recognition and generating standard feature representations from the normal watermark features using vector mapping technology; Calculating the distance between the independent watermark feature and the standard feature representation, and when the distance exceeds a preset threshold, the independent watermark feature is a feature deviation and obtaining a deviation flag; Extracting a traffic segment associated with the deviation identifier from the real-time traffic data based on the deviation identifier; A preliminary anomaly mark is obtained based on the traffic segment associated with the deviation identifier.
7. The method for detecting anomalies in an industrial network based on machine learning according to claim 6, wherein: The process of analyzing dynamic traffic in real-time traffic data based on the long short-term memory network to obtain temporal consistency scores includes: Long short-term memory network is used to process the communication timing features and device fingerprint representation in real-time traffic data to determine the timing consistency score.
8. The method for detecting anomalies in an industrial network based on machine learning according to claim 7, wherein: The process of obtaining a classification label for abnormal traffic based on the preliminary abnormality mark and the temporal consistency score includes: A random forest algorithm is used to fuse the preliminary anomaly mark and the temporal consistency score to obtain a traffic classification result; Extract abnormal traffic data from traffic classification results and determine abnormal traffic areas through weighted calculation; According to the abnormal traffic area, a clustering algorithm is used to divide the abnormal mark set; Get the updated classification label value through the abnormal label set; Update the classification label value, integrate the abnormal traffic data, and determine the final abnormal traffic sequence; The associated temporal consistency features are obtained from the final abnormal traffic sequence to determine the abnormality identification result.
9. An industrial network anomaly detection system based on machine learning, characterized in that: For implementing the industrial network anomaly detection method based on machine learning as claimed in claim 1, the system comprises: A traffic feature extraction module, configured to obtain real-time traffic data in an industrial network and obtain a comprehensive traffic feature set based on the real-time traffic data; A device fingerprint generation module is used to extract implicit features based on the comprehensive traffic feature set using a graph neural network, separate device behavior patterns from protocol features and communication timing features, and obtain a device fingerprint representation; The identity feature fusion module is used to obtain the unique identity feature vector of each device through multimodal feature fusion and self-supervised learning based on the device fingerprint representation and the pre-established normal communication data set; A watermark detection module, configured to determine watermark features in real-time traffic data based on the digital identity identifier to obtain a preliminary abnormality mark; The timing consistency analysis module is used to analyze the communication timing characteristics and device fingerprint representation in dynamic traffic through the long short-term memory network to obtain the timing consistency score; The abnormal traffic classification module is used to fuse the preliminary abnormality mark and the time series consistency score using the random forest algorithm, output the abnormal traffic identification result through weighted calculation, and obtain the abnormal traffic classification label; The watermark feature optimization module is used to extract the characteristic distribution of the forged watermark from the abnormal traffic identification results, and adjust the normal watermark features in the digital identity through an adaptive update mechanism to obtain the optimized identity feature set; The feature extraction model update module is used to adjust the parameters of the graph neural network using online learning methods based on the optimized identity feature set, enhance the robustness of feature learning, and obtain a feature extraction model that adapts to dynamic environments; The network security status update module is used to process subsequent traffic data through a feature extraction model, and output abnormal traffic identification results in real time by combining an autoencoder and self-attention mechanism to obtain a continuously updated industrial network security status.