Method for Mining and Analyzing Abnormal Behaviors in Industrial Internet Security
By collecting and preprocessing multi-dimensional data in real time in the industrial Internet and establishing a baseline model using unsupervised learning, the problems of inefficient detection efficiency and high false alarm rate in the existing technology are solved, and more accurate and efficient detection of safe abnormal behaviors is achieved.
Patent Information
- Application Number
- CN202411510513.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-28
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2044-10-28
AI Technical Summary
When the prior art detects security abnormal behavior in the industrial Internet, it is inefficient, prone to false alarms and missed reports, and cannot effectively detect unknown or mutated attack behaviors.
By installing data acquisition devices on various nodes of the industrial Internet, multi-dimensional data is collected in real time, and the data is preprocessed, feature extraction and unsupervised learning analysis are carried out to establish a baseline model to detect abnormal behavior.
It improves the detection accuracy and response speed of abnormal safety behaviors in industrial Internet, can effectively detect unknown or mutated attack behaviors, and reduces the rate of false alarms and missed reports.
Smart Images

Figure CN119030799B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial Internet security, and particularly to a method for mining and analyzing abnormal behaviors in industrial Internet security. Background Art
[0002] In the prior art, methods for detecting abnormal behaviors in industrial Internet security mainly rely on traditional rule-based detection and signature-based detection. These methods identify abnormal behaviors through predefined rules or known attack signatures. In the industrial Internet environment, data is usually collected through network traffic analysis, log analysis, and device status monitoring, and analyzed through data mining and machine learning techniques to detect potential security threats and abnormal behaviors.
[0003] However, there are some main problems in the prior art. First, rule-based detection and signature-based detection methods need to frequently update rule and signature libraries and cannot effectively detect unknown or mutated attack behaviors. Second, traditional methods are often inefficient in processing massive multi-dimensional data in the industrial Internet and are prone to false positives and false negatives. In addition, the prior art has limitations in feature extraction and data preprocessing and cannot comprehensively reflect the complexity of device operating states and network behaviors.
[0004] Therefore, there is an urgent need for a method that can efficiently and accurately mine and analyze abnormal behaviors in the industrial Internet to improve the security and reliability of the system. Summary of the Invention
[0005] This application provides a method for mining and analyzing abnormal behaviors in industrial Internet security to improve the detection accuracy and response speed of abnormal behaviors in industrial Internet security.
[0006] The method for mining and analyzing abnormal behaviors in industrial Internet security provided by this application includes:
[0007] Collecting multi-dimensional data in the industrial Internet in real time through data collection devices installed on each node of the industrial Internet, where the multi-dimensional data includes device operating state data, network traffic data, user operation logs, and sensor data;
[0008] Preprocessing the multi-dimensional data to generate a preprocessed data set, where the preprocessing includes data cleaning, missing value filling, noise filtering, and data format standardization;
[0009] Extracting features from the preprocessed data set to obtain key features reflecting device operating states and network behaviors, where the key features include time series features and frequency features;
[0010] Analyze the extracted key features through unsupervised learning algorithms to establish a baseline model for the normal operation of the industrial Internet; during the actual operation of the industrial Internet, collect new data in real time, perform preprocessing and feature extraction, and input the extracted features into the baseline model for comparison to detect abnormal behaviors that deviate from the baseline model.
[0011] Further, the analysis of the extracted key features through unsupervised learning algorithms to establish a baseline model for the normal operation of the industrial Internet includes:
[0012] Represent the extracted key features as a set of feature vectors , where each key feature vector includes temporal features and frequency features, is the number of key features;
[0013] Use the k-means clustering algorithm to perform clustering analysis on the set of feature vectors , set the number of clusters , calculate the distance from each feature vector in the set of feature vectors to the cluster center, and iteratively update the cluster center , until convergence; where the th cluster center is calculated according to the following formula (1):
[0014]
[0015] where, represents the set of feature vectors belonging to the th cluster; is the number of feature vectors in the set of feature vectors ; is the mean vector of all feature vectors; is an adjustment parameter used to balance the offset and dispersion of the cluster center;
[0016] For each feature vector , calculate its distance to the nearest cluster center as the anomaly score according to the following formula (2):
[0017]
[0018] where, is the anomaly score of the feature vector ; is the Euclidean distance between the feature vector and the cluster center ; is a parameter for adjustment, used to adjust the weights of Euclidean distance and absolute difference; represents a feature vector at the -th dimension; represents a cluster center at the -th dimension; is the number of dimensions of the feature vector;
[0019] Perform statistical analysis on all anomaly scores and fit their probability distribution; according to the fitted probability distribution, calculate the mean value and standard deviation ; calculate the threshold according to the following formula (3):
[0020]
[0021] where is the threshold for detecting abnormal behavior; is the number of anomaly scores; is the -th anomaly score;
[0022] Take the set of feature vectors with anomaly scores less than as the baseline model for the normal operation of the industrial Internet.
[0023] Furthermore, inputting the extracted feature vectors into the baseline model for comparison to detect abnormal behaviors deviating from the baseline model includes:
[0024] Calculate a new anomaly score ,
[0025]
[0026] where is the feature vector corresponding to the extracted feature;
[0027] If , it is determined as an abnormal behavior; otherwise, it is determined as a normal behavior.
[0028] Furthermore, performing feature extraction on the preprocessed data set to obtain key features reflecting the device operation status and network behavior includes:
[0029] Divide the preprocessed data set into multiple subsets according to a fixed time window, and calculate the mean value and standard deviation respectively within each time window according to the following formulas (5) and (6):
[0030] ,
[0031] wherein, is the th data point within the time window; is the weight of the th data point; is the timestamp of the th data point, is the time decay coefficient; is the total number of data points within the time window;
[0032]
[0033] wherein, is the th data point within the time window; is the weight of the th data point; is the th data point's timestamp; is the time decay coefficient; is the total number of data points within the time window;
[0034] The autocorrelation coefficient is calculated according to the following formula (7) :
[0035]
[0036] wherein, is the th data point within the time window; is the lag step; is the total number of data points within the time window;
[0037] The calculated mean , standard deviation and autocorrelation coefficient are determined as time series features;
[0038] Apply the discrete Fourier transform to the data within each time window to transform the time domain data to the frequency domain and obtain the frequency domain representation ; Calculate the main frequency according to the following formula (8):
[0039]
[0040] The calculated is determined as the frequency feature;
[0041] Concatenate the determined timing features and frequency features to generate key features reflecting the device operating status and network behavior.
[0042] Furthermore, through the data acquisition devices installed on each node of the industrial Internet, multi-dimensional data in the industrial Internet is collected in real time, including:
[0043] Real-time collect device operating status data through sensors installed on industrial devices, and transmit the collected device operating status data to the data processing center through the industrial Internet for centralized storage and processing;
[0044] Through network traffic monitoring devices, capture network traffic data in the industrial Internet; perform packet splitting on the network traffic data, extract the source address, destination address, protocol type, and data volume in the data packets; sort the network traffic data according to the time stamp and transmit it to the data processing center for analysis;
[0045] Deploy logging tools on user terminals and servers to record user operation behaviors; transmit user operation log data to the data processing center through a secure encrypted channel.
[0046] Furthermore, the sensors include temperature sensors, pressure sensors, and vibration sensors.
[0047] Furthermore, user operation behaviors include user login, logout, command execution, and file access.
[0048] Furthermore, preprocess the multi-dimensional data to generate a preprocessed data set, including:
[0049] Perform a preliminary screening on the collected data to eliminate obviously incorrect and invalid data;
[0050] Check the integrity and consistency of the data, and merge and deduplicate duplicate data;
[0051] Identify and process outliers and extreme values to ensure the accuracy of the data.
[0052] Furthermore, preprocessing the multi-dimensional data to generate a preprocessed data set also includes:
[0053] Analyze the parts of the data set with missing values to determine the distribution and influence range of the missing values;
[0054] Use linear interpolation to fill in the missing values in time series data;
[0055] For non-time series data, use the mean filling method or multiple imputation method for processing.
[0056] Further, the preprocessing of the multi-dimensional data to generate a preprocessed data set further includes:
[0057] Using a moving average filter to smooth the data and reduce the influence of random noise;
[0058] Adopting a method combining a high-pass filter and a low-pass filter to perform frequency-domain filtering on the data to remove high-frequency and low-frequency noises;
[0059] Performing format standardization processing on the data to convert the data into a unified format and dimension for subsequent analysis and processing.
[0060] The beneficial effects of the technical solution provided by this application include:
[0061] (1) Through the data acquisition devices installed on each node of the industrial Internet, multi-dimensional data is collected in real time, including device operation status data, network traffic data, user operation logs, and sensor data. This can comprehensively monitor various activities in the industrial Internet and ensure the real-time and comprehensiveness of the data. (2) The multi-dimensional data is preprocessed to generate a preprocessed data set. The preprocessing steps include data cleaning, missing value filling, noise filtering, and data format standardization. These steps greatly improve the quality of the data, reduce errors caused by data missing and noise, and improve the accuracy of subsequent analysis. (3) Feature extraction is performed on the preprocessed data set to obtain key features reflecting the device operation status and network behavior, including time series features and frequency features. Accurate feature extraction can better reflect the operation status of the device and the network, providing high-quality input for subsequent anomaly detection. (4) By analyzing the extracted key features through unsupervised learning algorithms, a baseline model for the normal operation of the industrial Internet is established. Unsupervised learning can automatically adapt to data changes, establish a dynamic baseline model, and effectively cope with the complexity and variability of the industrial Internet environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a flowchart of a method for mining and analyzing security abnormal behaviors in the industrial Internet provided by the first embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0063] Many specific details are set forth in the following description in order to provide a thorough understanding of this application. However, this application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this application. Therefore, this application is not limited by the specific implementations disclosed below.
[0064] The first embodiment of this application provides a method for mining and analyzing security abnormal behaviors in the industrial Internet. Please refer to Figure 1, this figure is a schematic diagram of the first embodiment of this application. The following will be combined with Figure 1 to provide a method for mining and analyzing industrial Internet security abnormal behaviors in the first embodiment of this application in detail.
[0065] Step S101: Through the data collection devices installed on each node of the industrial Internet, collect multi-dimensional data in the industrial Internet in real time, where the multi-dimensional data includes device operation status data, network traffic data, user operation logs, and sensor data.
[0066] In step S101, through the data collection devices installed on each node of the industrial Internet, collect multi-dimensional data in the industrial Internet in real time. First, the collection of device operation status data requires installing corresponding sensors on each industrial device, and these sensors include but are not limited to temperature sensors, pressure sensors, vibration sensors, and humidity sensors, etc. These sensors will monitor the operation status of the device in real time and transmit the data to the centralized data processing center through the industrial Internet. Secondly, the collection of network traffic data requires deploying network traffic monitoring devices on network nodes. These devices can capture network traffic data packets, extract information such as source address, destination address, protocol type, data packet size, and transmission time, and store these data in the network traffic database.
[0067] At the same time, the collection of user operation logs requires installing log recording tools on user terminals and servers. These tools record every operation of users, including operation logs such as login, logout, command execution, file access, and system changes. These log data are transmitted to the data processing center through a secure encryption channel to ensure the integrity and confidentiality of the data during transmission. In addition, the data collection devices for environmental sensors also include gas sensors, light sensors, etc. These sensors are used to monitor changes in the industrial environment to ensure that the environmental conditions are within a safe range.
[0068] To ensure the comprehensiveness and accuracy of data collection, each data collection device needs to be calibrated and maintained regularly. The sensor data is preliminarily processed through edge computing devices, such as data compression and format conversion, to reduce transmission latency and bandwidth occupancy. All the collected data will be sent to the data center at regular intervals. At the data center, the data will be further stored and processed. Encryption technologies, such as SSL / TLS, are used during the data transmission process to ensure the security and anti-tampering of data transmission.
[0069] In summary, step S101 ensures the real-time, comprehensive, and accurate collection of multi-dimensional data in the industrial Internet through precise arrangement and use of various data collection devices, combined with edge computing and encrypted transmission technologies, providing a solid data foundation for subsequent preprocessing, feature extraction, and abnormal behavior detection.
[0070] Furthermore, the data acquisition devices installed on each node of the industrial Internet collect multi-dimensional data in the industrial Internet in real time, including:
[0071] The operating status data of industrial equipment is collected in real time by sensors installed on the industrial equipment, and the collected operating status data of the equipment is transmitted to the data processing center through the industrial Internet for centralized storage and processing;
[0072] The network traffic data in the industrial Internet is captured by network traffic monitoring devices; the network traffic data is packetized, and the source address, destination address, protocol type, and data volume in the data packets are extracted; the network traffic data is sorted according to the timestamp and transmitted to the data processing center for analysis;
[0073] Logging tools are deployed on user terminals and servers to record users' operation behaviors; the user operation log data is transmitted to the data processing center through a secure encrypted channel.
[0074] The data acquisition devices installed on each node of the industrial Internet collect multi-dimensional data in the industrial Internet in real time, specifically including the following aspects:
[0075] First, various types of sensors are installed on industrial equipment to collect the operating status data of the equipment in real time. These sensors may include temperature sensors, pressure sensors, vibration sensors, humidity sensors, etc. These sensors need to be properly connected and calibrated to the industrial equipment to ensure the accuracy and reliability of the collected data. The collected data is transmitted to the data processing center through the industrial Internet for centralized storage and processing. During the data transmission process, advanced communication protocols and encryption technologies are used to ensure the security and integrity of data transmission.
[0076] Second, the network traffic data in the industrial Internet is captured by network traffic monitoring devices. These devices are usually deployed at key nodes of the network and can capture and record the detailed information of network data packets in real time. In order to process and analyze these data packets, they need to be packetized. Specifically, packetization includes decomposing the data packets into multiple fields and extracting key information such as source address, destination address, protocol type, and data volume. These field information helps to understand the specific situation and traffic characteristics of network communication. Next, the extracted network traffic data is sorted according to the timestamp to ensure the temporal consistency of the data, and the sorted data is transmitted to the data processing center for further analysis. The data processing center uses specialized network analysis tools and algorithms to deeply analyze these data to identify potential network anomalies and security threats.
[0077] Deploying logging tools on user terminals and servers to record users' operation behaviors is another important step. These logging tools can record various operations of users in the system in detail, including logins, logouts, file accesses, command executions, and system configuration changes. To protect the security of this sensitive data, secure encryption channels such as the SSL / TLS protocol are used during the transmission of log data to ensure that the data will not be stolen or tampered with during transmission. After being transmitted to the data processing center, these log data will be stored in a dedicated log database and classified and indexed for subsequent query and analysis. By analyzing users' operation logs, abnormal user behaviors can be detected, such as frequent login failures, abnormal file accesses, and unauthorized system configuration changes.
[0078] The entire data collection and transmission process requires the comprehensive application of various technologies and devices to ensure the comprehensiveness, accuracy, and security of the data. Through efficient data processing and analysis methods, various security abnormal behaviors in the industrial Internet can be identified and addressed in a timely manner to ensure the stable operation of the system and information security.
[0079] Furthermore, the sensors include a temperature sensor, a pressure sensor, and a vibration sensor.
[0080] First, the temperature sensor is used to monitor the temperature changes of industrial equipment. The temperature sensor is installed at key parts of the equipment, such as the engine, heat exchanger, or other heat-generating components, to ensure that the actual operating temperature of the equipment can be accurately measured. The temperature sensor needs to be connected to the control system of the equipment, and usually transmits temperature data to the data acquisition device through cables or wireless transmission methods. To ensure the accuracy of the data, the temperature sensor should be calibrated before installation to eliminate measurement errors.
[0081] Second, the pressure sensor is used to monitor the pressure changes generated during the operation of the equipment. The pressure sensor is usually installed on the hydraulic system, pneumatic system, or other key components of the equipment to obtain real-time system pressure data. By using appropriate interfaces and pipelines, the pressure sensor is connected to the pressure point of the equipment. The collected pressure data is transmitted to the industrial Internet network through the data acquisition device. During the installation and use of the pressure sensor, attention needs to be paid to its measurement range and response speed to ensure that it can accurately reflect the operating state of the equipment.
[0082] Vibration sensors are used to monitor the vibration conditions of equipment, helping to identify whether the equipment is in a normal operating state. Vibration sensors are usually installed on the bearings, transmission devices or other parts of the equipment that are prone to vibration. By measuring the vibration frequency and amplitude of the equipment, vibration sensors can detect the running smoothness of the equipment and potential mechanical failures. Install the vibration sensor on the equipment through appropriate fixing devices and ensure its stability to obtain accurate vibration data. The vibration data is transmitted to the industrial Internet through a data acquisition device for further analysis.
[0083] To achieve effective data transmission, all the data collected by sensors will be transmitted to the data processing center through the industrial Internet. During the transmission process, advanced communication protocols and encryption technologies are adopted to ensure the integrity and security of the data. After receiving these data, the data processing center will centrally store and process them. Through the comprehensive analysis of temperature, pressure and vibration data, the running state of the equipment can be comprehensively understood, and potential abnormal behaviors and failures can be discovered and warned in a timely manner.
[0084] Furthermore, the user's operation behaviors include user login, logout, command execution and file access.
[0085] First, for the monitoring of user login and logout behaviors, the log recording tool will record relevant information every time a user attempts to log in to the system. The recorded content includes the username, login time, login IP address and login result (success or failure). When the user successfully logs in, the system will generate a log record containing the above information and the unique identifier of the user session. Similarly, when the user logs out of the system, the log recording tool will record the time when the user logs out and the session identifier, ensuring that every login and logout behavior has a detailed record.
[0086] Second, for the monitoring of command execution, the log recording tool will capture every command executed by the user in the system. The recorded content includes the time when the command is executed, the command content, the executing user and the command execution result. This can detail the operation path and specific behaviors of the user in the system. Especially when an exception or security event occurs, the specific operation steps and responsible persons can be traced through the command log. For example, when the user executes important system configuration commands or database operation commands on the server, the log recording tool will detail these operations for subsequent auditing and analysis.
[0087] Monitoring file access behavior is equally crucial, especially in cases involving sensitive data or critical business files. Logging tools record all user access behaviors to files, including file reading, modification, deletion, and creation operations. The recorded content includes the file access path, access time, operation type, executing user, and operation result. For example, when a user reads a file containing sensitive information, the logging tool will detail the behavior, including the time of reading the file, the user identity, and the reading result.
[0088] To ensure data security, all log data is transmitted to the data processing center through a secure encrypted channel after being recorded. The transmission of this data uses encryption technologies such as the SSL / TLS protocol to ensure that the logs cannot be intercepted or tampered with by unauthorized third parties during transmission. The log data transmitted to the data processing center is stored in a dedicated log database and classified and indexed for subsequent query and analysis.
[0089] At the data processing center, specialized analysis tools analyze the collected log data to identify potential security anomalies. For example, by analyzing login and logout logs, abnormally frequent login attempts or unauthorized login behaviors can be detected. By analyzing command execution logs, abnormal system configuration changes or database operations can be identified. By analyzing file access logs, unauthorized file reading or data leakage behaviors can be discovered.
[0090] Step S102: Preprocess the multi-dimensional data to generate a preprocessed data set, where the preprocessing includes data cleaning, missing value filling, noise filtering, and data format standardization.
[0091] In step S102, the multi-dimensional data is preprocessed to generate a preprocessed data set, and this process involves multiple key steps to ensure that the collected data can be accurately and effectively used for subsequent analysis and processing.
[0092] First, data cleaning is performed. The purpose of data cleaning is to identify and correct or delete errors, duplicates, and invalid data in the data. For example, for device operation status data, there may be abnormal readings due to sensor failures or communication problems. By setting reasonable threshold ranges, obviously unreasonable temperature, pressure, or vibration readings can be identified and deleted. For network traffic data, the cleaning process can include deleting packets with incorrect formats, identifying and deleting duplicate traffic records. User operation logs also need to be cleaned to remove invalid or duplicate log records and ensure that the format of each log record is consistent.
[0093] Next, handle the missing values in the data. In the industrial Internet environment, data missing is a common problem, which may be caused by reasons such as sensor failures and network delays. To fill in the missing values, various methods can be adopted. For time series data, linear interpolation is a commonly used method, which calculates the missing values through the linear relationship of known data points. For more complex situations, Lagrange interpolation or multiple imputation methods can be used. These methods take into account the overall trend and pattern of the data and provide more accurate filling results. For non-time series data, mean filling is a simple and effective method, which replaces the missing values with the mean of the variable, thereby reducing the impact of data missing on the analysis results.
[0094] Then, perform noise filtering. Noise is the random error or irrelevant information in the data, which will affect the accuracy of the analysis results. To filter out the noise, various filtering techniques can be used. For example, the moving average filter reduces random fluctuations by smoothing the data and is applicable to stationary time series data. For time series data with periodicity or trend, the method of combining high-pass filters and low-pass filters can be used to remove low-frequency and high-frequency noise respectively, thereby retaining the main information in the data. In addition, the Kalman filter is also an effective noise filtering method, especially suitable for data processing in dynamic systems.
[0095] Finally, perform data format standardization. The multi-dimensional data in the industrial Internet comes from different devices and systems, with different formats and dimensions. To ensure the consistency and comparability of the data, it is necessary to standardize the data. Commonly used standardization methods include normalizing the data according to its mean and standard deviation, so that it is converted into a standard normal distribution with a mean of zero and a standard deviation of one. Another method is to scale the data to the range of [0,1], which is achieved by subtracting the minimum value and dividing by the range (the maximum value minus the minimum value). The standardized data is not only convenient for comparison and analysis, but also can improve the efficiency and effect of model training.
[0096] Through the above detailed and clear preprocessing steps, the quality and consistency of the data can be effectively improved, laying a solid foundation for subsequent feature extraction and abnormal behavior detection.
[0097] Furthermore, the preprocessing of the multi-dimensional data to generate a preprocessed data set includes:
[0098] Conduct a preliminary screening of the collected data to eliminate obvious errors and invalid data;
[0099] Check the integrity and consistency of the data, and merge and deduplicate the duplicate data;
[0100] Identify and process outliers and extreme values to ensure the accuracy of the data.
[0101] First, conduct a preliminary screening of the collected data to eliminate obviously incorrect and invalid data. The purpose of this step is to ensure the basic quality and credibility of the dataset. The specific operations include checking whether the data format is correct, whether the numerical range is reasonable, and whether there are obvious logical errors. For example, for the data of the device operating status, if the temperature value recorded by the sensor far exceeds the normal operating range of the device (such as -100°C or 1000°C), then these data are obviously unreasonable and need to be eliminated. Similarly, for the network traffic data, if the source address or destination address format of the data packet is incorrect, these data should also be marked as invalid data and eliminated.
[0102] Secondly, check the integrity and consistency of the data, and merge and deduplicate the duplicate data. The data integrity check includes confirming that all necessary data fields have been filled and there is no missing key data. For example, in the user operation log, each record should include information such as the username, operation type, and timestamp. Data records lacking this information will be regarded as incomplete. The data consistency check ensures that the data remains consistent across different sources or different time periods. For example, the data of the same event recorded by multiple sensors should be consistent. If inconsistencies are found, further investigation and processing are required. For duplicate data, it can be identified by comparing the content and timestamp of the records. If two records are exactly the same within a short period of time, they can be merged into one record to reduce data redundancy and improve data processing efficiency.
[0103] Finally, identify and process outliers and extreme values to ensure the accuracy of the data. Outliers refer to those values that deviate significantly from other data points and may be caused by sensor failures, data transmission errors, etc. The methods for processing outliers include using statistical analysis techniques, such as the standard deviation method, to identify data points that exceed a certain standard deviation range, and then selecting to delete, correct, or mark them as outliers according to the specific situation. For example, for the vibration sensor data, if the vibration amplitude of some data points is significantly higher than the vibration range under normal operating conditions, these data points can be marked as outliers and further analyzed. Extreme values refer to those data points that are significantly different from other data points and may reflect potential system failures or abnormal behaviors. The methods for processing extreme values include using clustering analysis techniques to divide the data points into normal groups and extreme value groups, and then conducting detailed analysis and processing on the extreme value groups.
[0104] Through the above detailed steps, the multi-dimensional data collected is preprocessed to generate a high-quality and reliable preprocessed dataset. This dataset not only eliminates obviously incorrect and invalid data, ensures the integrity and consistency of the data, but also improves the accuracy and reliability of the data by identifying and processing outliers and extreme values.
[0105] Further, the preprocessing of the multi-dimensional data to generate a preprocessed data set further includes:
[0106] Analyze the part of the data set with missing values to determine the distribution and scope of influence of the missing values;
[0107] Use linear interpolation to fill in the missing values in the time series data;
[0108] For non-time series data, use the mean filling method or multiple imputation method for processing.
[0109] First, conduct a detailed analysis of the part of the data set with missing values. The purpose of this step is to determine the distribution and scope of influence of the missing values in the data set. When analyzing the missing values, it is necessary to count the number and location of the missing values in each variable to understand the overall distribution of the missing values. For example, some sensors may fail to collect data during a specific period, resulting in missing values concentrated in that period. Through such analysis, it is possible to clarify which variables and time periods of data need to be focused on for processing.
[0110] For the missing values in the time series data, use linear interpolation to fill them. Linear interpolation is a commonly used filling method that calculates the missing values according to a linear relationship by using the values of adjacent data points. Specifically, for a missing value in the time series data, the two nearest known values before and after the missing value can be found, and a linear interpolation can be calculated based on the time and values of these two known values. For example, if there are known values and at time points and , and there is a missing value at time point , then the missing value can be calculated by the following formula:
[0111] ,
[0112] This method is simple and has high computational efficiency, and is suitable for situations where the data changes relatively smoothly.
[0113] For non-time series data, the mean imputation method or multiple imputation method is used for processing. The mean imputation method is the simplest imputation method, which replaces the missing values with the mean of the variable. The advantage of this method is simple calculation, but the disadvantage is that it may underestimate the variability of the data. The specific operation is to calculate the mean of each variable, and then replace the missing values with the mean of the corresponding variable. The multiple imputation method is a more complex and accurate method, which estimates the missing values by generating multiple alternative values and takes into account the inherent variability of the data. The multiple imputation method generally includes the following steps: First, generate multiple data sets containing alternative values; then, analyze each data set to obtain multiple analysis results; finally, synthesize these results to obtain the final imputation result. The multiple imputation method can better maintain the statistical characteristics and variability of the data and is applicable to cases where the data is relatively complex and uneven.
[0114] Through the above steps, the missing values existing in the data set can be effectively processed, and the integrity of the data set and the accuracy of the analysis can be improved.
[0115] Furthermore, the preprocessing of the multi-dimensional data to generate a preprocessed data set further includes:
[0116] Use a moving average filter to smooth the data and reduce the influence of random noise;
[0117] Adopt a method combining a high-pass filter and a low-pass filter to perform frequency domain filtering on the data to remove high-frequency and low-frequency noise;
[0118] Perform format standardization processing on the data to convert the data into a unified format and dimension for subsequent analysis and processing.
[0119] The preprocessing of the multi-dimensional data to generate a preprocessed data set further includes several important steps, which include using a moving average filter for smoothing, adopting a method combining a high-pass filter and a low-pass filter for frequency domain filtering, and performing format standardization processing on the data.
[0120] First, use a moving average filter to smooth the data to reduce the influence of random noise. The moving average filter smooths the data by calculating the average of a series of data points. This process involves sliding a fixed-size window over the data sequence, calculating the average of the data points within the window at each position, and then replacing the data point at the center of the window with this average. This can effectively weaken the random fluctuations in the data and make the data smoother and more stable. For example, for time series data, the moving average filter can help eliminate short-term fluctuations in sensor data and highlight long-term trends.
[0121] Next, a method combining a high-pass filter and a low-pass filter is used to perform frequency-domain filtering on the data to remove high-frequency and low-frequency noises. Frequency-domain filtering analyzes the frequency components of the data and selectively retains or removes signals within a specific frequency range. The high-pass filter is used to remove low-frequency noises, such as slow-changing trends or DC components, while the low-pass filter is used to remove high-frequency noises, such as fast-changing interference signals. During specific operations, the data can first be subjected to a fast Fourier transform (FFT) to convert the time-domain data to the frequency domain, then the high-pass and low-pass filters are applied to remove the unwanted frequency components, and finally, the data is converted back to the time domain through an inverse Fourier transform (IFFT). The processed data is not only smoother but also more accurately reflects the actual operating state.
[0122] Finally, format standardization processing is performed on the data to convert the data into a unified format and dimension, facilitating subsequent analysis and processing. The purpose of format standardization processing is to eliminate the differences between different data sources and different types of data, making the data numerically comparable. Standardization methods include normalizing the data according to its mean and standard deviation to convert it into a standard normal distribution with a mean of zero and a standard deviation of one, or scaling the data to the range of [0,1]. For data with different dimensions, such as temperature (in degrees Celsius), pressure (in Pascals), and vibration (in Hertz), standardization processing can convert them to the same dimension range through linear transformation or logarithmic transformation. In this way, the standardized data is not only convenient for comparison and analysis but also can improve the efficiency and effectiveness of subsequent model training and algorithm processing.
[0123] Through the above detailed steps, the multi-dimensional data collected is processed through smoothing, frequency-domain filtering, and format standardization to generate a preprocessed dataset with high quality and a unified format. This dataset not only reduces the influence of noise, ensuring the smoothness and stability of the data, but also eliminates the differences between different data sources, improving the comparability and processing efficiency of the data.
[0124] Step S103: Extract features from the preprocessed dataset to obtain key features reflecting the device operating state and network behavior, where the key features include temporal features and frequency features.
[0125] In step S103, features are extracted from the preprocessed dataset to obtain key features reflecting the device operating state and network behavior, and these features include temporal features and frequency features.
[0126] First, for time series feature extraction, the preprocessed data needs to be divided according to a fixed time window. The data within each time window is treated as a unit. For example, if one minute is selected as the time window, the data for each minute is regarded as a separate unit for analysis. Within each time window, multiple time series statistics are calculated, including the mean, standard deviation, autocorrelation coefficient, etc. The mean can reflect the overall level of the data within that time window, while the standard deviation is used to measure the degree of data fluctuation. The autocorrelation coefficient is used to describe the correlation between time series data at different time points, thereby identifying periodic and trend changes in the data.
[0127] Next, for frequency feature extraction, frequency domain analysis needs to be performed on the data within each time window. A commonly used method is to apply the discrete Fourier transform (DFT) to convert the time domain data to the frequency domain. Through the Fourier transform, the original time series data can be decomposed into the superposition of different frequency components, thereby obtaining a frequency domain representation. For each frequency component, its amplitude and phase are calculated, with particular attention paid to the main frequency and its corresponding amplitude, because the main frequency reflects the most significant periodic component in the data. By analyzing the spectrogram, the main frequency components in the data and their energy distribution can be identified, thereby extracting frequency features that reflect the system operating state.
[0128] In addition, to further improve the effect of feature extraction, wavelet transform can be combined to perform multi-scale analysis on the data. Wavelet transform can analyze the local features of the data in both the time domain and the frequency domain by decomposing the data into components of different scales. This method is particularly suitable for processing non-stationary time series data because it can capture instantaneous changes and abrupt features in the data.
[0129] Finally, the extracted time series features and frequency features are fused to form a comprehensive feature vector. To ensure the unity and comparability of the feature vector, the feature vector usually needs to be standardized. Standardization methods include subtracting the mean of each feature from it and dividing by its standard deviation, or scaling the feature to a specific range (such as [0,1]). The standardized feature vector can be better used for subsequent model training and anomaly detection.
[0130] Through the above steps, high-quality key features can be extracted from the preprocessed dataset. These features can not only accurately reflect the device operating state and network behavior, but also provide a reliable data basis for subsequent unsupervised learning algorithms, thereby realizing the effective mining and analysis of industrial Internet security abnormal behaviors.
[0131] Furthermore, the feature extraction of the preprocessed dataset to obtain key features reflecting the device operating state and network behavior includes:
[0132] The preprocessed dataset is divided into multiple subsets according to a fixed time window, and the mean value is calculated respectively according to the following formulas (5) and (6) within each time window and the standard deviation :
[0133] ,
[0134] where is the th data point within the time window; is the weight of the th data point; is the timestamp of the th data point, is the time decay coefficient; is the total number of data points within the time window;
[0135]
[0136] where is the th data point within the time window; is the weight of the th data point; is the timestamp of the th data point; is the time decay coefficient; is the total number of data points within the time window;
[0137] The autocorrelation coefficient is calculated according to the following formula (7) :
[0138]
[0139] where is the th data point within the time window; is the lag step; is the total number of data points within the time window;
[0140] The calculated mean value , standard deviation and autocorrelation coefficient are determined as time series features;
[0141] Apply the discrete Fourier transform to the data within each time window to transform the time-domain data into the frequency domain and obtain the frequency-domain representation ; Calculate the dominant frequency according to the following formula (8) :
[0142]
[0143] The calculated is determined as the frequency feature;
[0144] The determined time series feature and frequency feature are concatenated to generate the key feature reflecting the device operating state and network behavior.
[0145] First, the preprocessed dataset is divided into multiple subsets according to a fixed time window. This division process is for detailed statistical analysis within each time window to ensure real-time monitoring and analysis of the device operating state and network behavior. The size of the time window can be selected according to the specific application scenario, such as one minute or five minutes.
[0146] Within each time window, first calculate the mean value of the data . The calculation formula for the mean value is:
[0147]
[0148] where, where, is the th data point within the time window; is the weight of the th data point; is the timestamp of the th data point; is the time decay coefficient; is the total number of data points within the time window; The mean value reflects the central tendency of the data within this time window and is a basic description of the data distribution.
[0149] Next, calculate the standard deviation of the data , and the calculation formula for the standard deviation is:
[0150]
[0151] where, is the th data point within the time window; is the weight of the th data point; is the timestamp of the th data point; is the time decay coefficient; is the total number of data points within the time window;
[0152] The standard deviation is used to measure the degree of dispersion of the data and describes the distribution of the data points around the mean value. A high standard deviation indicates that the data points are more dispersed, while a low standard deviation indicates that the data points are concentrated near the mean value.
[0153] In addition, the autocorrelation coefficient also needs to be calculated , the calculation formula of the autocorrelation coefficient is:
[0154]
[0155] Among them, is the th data point within the time window; is the lag step. The autocorrelation coefficient is used to measure the correlation between time series data at different time points and can reveal the periodic and trend changes in the data.
[0156] The calculated mean value , standard deviation and autocorrelation coefficient are determined as time series features; these time series features can reflect the operating state of the device within a specific time window.
[0157] Next, apply the discrete Fourier transform (DFT) to the data within each time window to transform the time-domain data into the frequency domain and obtain the frequency-domain representation . Through frequency-domain analysis, the frequency characteristics of the data can be revealed, and the periodic components in the data can be captured.
[0158] Apply the discrete Fourier transform to the data within each time window to transform the time-domain data into the frequency domain and obtain the frequency-domain representation ; calculate the main frequency according to the following formula (8) :
[0159]
[0160] The main frequency reflects the most significant periodic component in the data and is an important result of frequency-domain analysis. The calculated main frequency is determined as the frequency feature.
[0161] Finally, splice the determined time series features and frequency features to generate key features that reflect the operating state of the device and network behavior. The spliced feature vector contains information in both the time series and frequency dimensions and can comprehensively describe the operating state of the device and network behavior. These key features will serve as the basis for the analysis of subsequent unsupervised learning algorithms and provide reliable data support for the detection of abnormal behaviors in industrial Internet security.
[0162] Step S104: Analyze the extracted key features through an unsupervised learning algorithm to establish a baseline model for the normal operation of the industrial Internet; during the actual operation of the industrial Internet, collect new data in real time, perform preprocessing and feature extraction, input the extracted features into the baseline model for comparison, and detect abnormal behaviors that deviate from the baseline model.
[0163] In step S104, the extracted key features are analyzed through an unsupervised learning algorithm to establish a baseline model for the normal operation of the industrial Internet. In the actual implementation process, first, a training data set containing a large amount of normal operation data needs to be prepared, and this data should have undergone the aforementioned preprocessing and feature extraction steps.
[0164] To establish the baseline model, a suitable unsupervised learning algorithm is selected, such as K-means clustering. Before starting the clustering analysis, the number of clusters K is set, and this parameter can be determined according to the specific application scenario and data characteristics. Usually, an empirical rule or a data-based analysis method is used to select the value of K to ensure that the clustering results can accurately represent the internal structure of the data.
[0165] The set of feature vectors is input into the K-means algorithm. First, K initial cluster centers are randomly selected. Then, each feature vector is assigned to the nearest cluster center, which is achieved by calculating the Euclidean distance from the feature vector to each cluster center. Each feature vector is assigned to the corresponding cluster according to the principle of the minimum distance.
[0166] Next, calculate the new center positions of each cluster, that is, calculate the mean of all feature vectors belonging to the same cluster. After updating the cluster centers, reassign the feature vectors to the new cluster centers, and repeat the above process until the positions of the cluster centers no longer change significantly or reach the preset number of iterations. The finally obtained cluster centers represent the baseline model in the normal operation state.
[0167] During the actual operation of the industrial Internet, after the newly collected data is preprocessed and feature extracted, new feature vectors will be obtained. To detect abnormal behaviors, these new feature vectors are input into the baseline model, and the distances from them to each cluster center are calculated. Specifically, for each new feature vector, calculate its Euclidean distance to all cluster centers, and select the minimum distance as the anomaly score of this feature vector.
[0168] Based on the normal behavior data in the aforementioned training data set, calculate the anomaly scores of all feature vectors to obtain the distribution of anomaly scores. Through statistical analysis, calculate the mean and standard deviation of the anomaly scores, and set the detection threshold. Usually, the threshold can be set as the mean plus twice the standard deviation, or the multiple can be adjusted according to specific requirements to control the sensitivity and false alarm rate of anomaly detection.
[0169] In practical applications, when the anomaly score of a new feature vector exceeds the preset threshold, it is determined as an abnormal behavior. This anomaly detection method based on unsupervised learning can adapt to the dynamic changes of data and improve the accuracy and robustness of detection.
[0170] Further, analyzing the extracted key features through an unsupervised learning algorithm to establish a baseline model for the normal operation of the industrial Internet, including:
[0171] Represent the extracted key features as a set of feature vectors
[0172] , where each key feature vector includes temporal features and frequency features, is the number of key features;
[0173] Use the k-means clustering algorithm to perform clustering analysis on the set of feature vectors , set the number of clusters , calculate the distance from each feature vector in the set of feature vectors to the cluster center, and iteratively update the cluster center until convergence; where the k-th cluster center is calculated according to the following formula (1):
[0174]
[0175] where, represents the set of feature vectors belonging to the k-th cluster; is the number of feature vectors in the set of feature vectors ; is the mean vector of all feature vectors; is an adjustment parameter used to balance the offset and dispersion of the cluster center;
[0176] For each feature vector , calculate its distance to the nearest cluster center as the anomaly score according to the following formula (2):
[0177]
[0178] where, is the anomaly score of the feature vector ; is the Euclidean distance between the feature vector and the cluster center ; is an adjustment parameter used to adjust the weights of the Euclidean distance and the absolute difference; represents the value of the feature vector in the -th dimension; represents the value of the cluster center in the Values on each dimension; Is the number of dimensions of the feature vector;
[0179] For all anomaly scores Perform statistical analysis, fit its probability distribution; According to the fitted probability distribution, calculate the mean value of the anomaly scores And standard deviation ; Calculate the threshold according to the following formula (3):
[0180]
[0181] Where, Is the threshold for detecting abnormal behavior; Is the number of anomaly scores; Is the th anomaly score;
[0182] The set of feature vectors with anomaly scores less than Is used as the baseline model for the normal operation of the industrial Internet. As the baseline model for the normal operation of the industrial Internet.
[0183] First, represent the extracted key features as a set of feature vectors , where each feature vector Includes time series features and frequency features; Is the number of feature vectors. This set contains all the feature vectors extracted from the dataset and is ready for clustering analysis.
[0184] Next, use The k-means clustering algorithm to perform clustering analysis on the set of feature vectors First, set the number of clusters , this value can be selected according to the characteristics and requirements of the data. Then, through an iterative process, calculate the distance from each feature vector to the cluster center and update the cluster center. After initializing cluster centers, assign each feature vector to the nearest cluster center. Specifically, the th cluster center Is calculated according to the following formula:
[0185]
[0186] Where, Represents the set of feature vectors belonging to the th cluster; Is the set of feature vectors The number of feature vectors in; Is the mean vector of all feature vectors; It is a parameter for adjustment, used to balance the offset and dispersion of the cluster centers; this formula takes into account the mean and dispersion of the feature vectors, making the cluster centers more representative.
[0187] Then, for each feature vector , calculate its distance to the nearest cluster center as the anomaly score. The calculation formula is as follows:
[0188]
[0189] Where, is the anomaly score of the feature vector ; is the Euclidean distance between the feature vector and the cluster center ; is a parameter for adjustment, used to adjust the weights of the Euclidean distance and the absolute difference; represents the value of the feature vector in the th dimension; represents the value of the cluster center in the th dimension; is the number of dimensions of the feature vector;
[0190] Perform statistical analysis on all the anomaly scores and fit their probability distribution. By calculating the mean and standard deviation of the anomaly scores, the distribution characteristics of the anomaly scores can be determined. The calculation formula is as follows:
[0191]
[0192] Where, is the threshold for detecting abnormal behaviors; is the number of anomaly scores; is the th anomaly score; this formula not only considers the mean and standard deviation, but also introduces a skewness adjustment term, making the threshold calculation more accurate and sensitive.
[0193] Finally, take the set of feature vectors with anomaly scores less than as the baseline model for the normal operation of the industrial Internet. Through this method, abnormal behaviors can be effectively identified and detected, ensuring the normal operation and security of the system.
[0194] Furthermore, inputting the extracted feature vectors into the baseline model for comparison to detect abnormal behaviors deviating from the baseline model includes:
[0195] Calculate the new anomaly score according to the following formula (4). :
[0196]
[0197] where is the eigenvector corresponding to the extracted feature;
[0198] If , it is determined as an abnormal behavior; otherwise, it is determined as a normal behavior.
[0199] First, input the newly extracted eigenvector into the established baseline model. The baseline model is obtained by analyzing and clustering the data in the normal operating state and contains multiple cluster centers .
[0200] To calculate the anomaly score of the new eigenvector, it is necessary to calculate its distances to all cluster centers. Specifically, for the new eigenvector , calculate its Euclidean distance to each cluster center and select the minimum distance as the anomaly score. The calculation formula is as follows:
[0201] ,
[0202] In this formula, represents the Euclidean distance between the new eigenvector and the th cluster center . By comparing the distances between the new eigenvector and all cluster centers, the closest cluster center can be found, and this distance is used as the anomaly score of the eigenvector.
[0203] Next, by comparing the calculated anomaly score with the pre-set threshold , it is determined whether the new eigenvector belongs to an abnormal behavior. If , it is determined that the behavior corresponding to the eigenvector is an abnormal behavior; otherwise, it is determined as a normal behavior.
[0204] Specifically, the threshold is calculated based on the statistical analysis of historical data and usually takes into account the average value, standard deviation, and other statistical characteristics of the data to ensure that the threshold can effectively distinguish normal and abnormal behaviors. In practical applications, the setting of the threshold can be adjusted according to specific application scenarios and requirements to balance the risks of false positives and false negatives.
[0205] The second embodiment of the present application provides an electronic device, and the electronic device includes:
[0206] Processor;
[0207] A memory for storing a program which, when read and executed by the processor, executes the method for mining and analyzing industrial Internet security abnormal behaviors provided in the first embodiment of the present application.
[0208] The third embodiment of the present application provides a computer-readable storage medium with a computer program stored thereon. When the program is executed by a processor, it executes the method for mining and analyzing industrial Internet security abnormal behaviors provided in the first embodiment of the present application.
[0209] Although the present application is disclosed above in preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application shall be subject to the scope defined by the claims of the present application.
Claims
1. A method for mining and analyzing abnormal security behaviors in the industrial Internet, characterized in that: include: The multi-dimensional data in the industrial Internet is collected in real time through the data collection device installed on each node of the industrial Internet, wherein the multi-dimensional data includes equipment operation status data, network traffic data, user operation logs and sensor data; Preprocess multi-dimensional data to generate preprocessed data sets, where preprocessing includes data cleaning, missing value filling, noise filtering and data format standardization; Perform feature extraction on the preprocessed data set to obtain key features that reflect the device operating status and network behavior, where the key features include timing features and frequency features; The extracted key features are analyzed through unsupervised learning algorithms to establish a baseline model for the normal operation of the Industrial Internet. In the actual operation of the Industrial Internet, new data is collected in real time and preprocessed and feature extracted. The extracted features are input into the baseline model for comparison to detect abnormal behaviors that deviate from the baseline model. The extracted key features are analyzed by unsupervised learning algorithms to establish a baseline model for the normal operation of the industrial Internet, including: The extracted key features are represented as a feature vector set X = {x1, x2, ..., x n }, where each key feature vector x i It includes time series features and frequency features, n is the number of key features; Use the K-means clustering algorithm to perform cluster analysis on the feature vector set X, set the number of clusters k, calculate the distance from each feature vector in the feature vector set X to the cluster center, and iteratively update the cluster center C = {c1, c2, ..., c k }, until convergence; among them, the jth cluster center c j Calculate according to the following formula (1): Among them, S j represents the set of feature vectors belonging to the jth cluster; |S j | is the feature vector set S j The number of eigenvectors in; m is the mean vector of all eigenvectors; λ1 is an adjustment parameter used to balance the offset and dispersion of cluster centers; For each feature vector x i , according to the following formula (2), calculate its distance to the nearest cluster center as the anomaly score: Among them, d i is the eigenvector x i The anomaly score of ||x i -c j || is the eigenvector x i With cluster center c j The Euclidean distance between them; α is an adjustment parameter used to adjust the weights of the Euclidean distance and the absolute difference; x i,t represents the feature vector x i The value in the tth dimension; c j,t Represents the cluster center c j The value in the tth dimension; T is the dimension of the feature vector; For all anomaly scores {d1, d2, ..., d n } Perform statistical analysis and fit its probability distribution; according to the fitted probability distribution, calculate the average value μ of the abnormal score d and standard deviation σ d ; Calculate the threshold value according to the following formula (3): Where θ is the threshold for detecting abnormal behavior; n is the number of anomaly scores; d i is the i-th anomaly score; The set of feature vectors {x i |d i <θ} serves as a baseline model for the normal operation of the Industrial Internet; Perform feature extraction on the preprocessed data set to obtain key features that reflect the device operating status and network behavior, including: The preprocessed data set is divided into multiple subsets according to fixed time windows, and the mean μ is calculated in each time window according to the following formulas (5) and (6): t and standard deviation σ t : Among them, x i is the i-th data point in the time window; w i is the weight of the i-th data point; t i is the timestamp of the i-th data point; λ is the time decay coefficient; n is the total number of data points in the time window; Among them, x i is the i-th data point in the time window; w i is the weight of the i-th data point; t i is the timestamp of the i-th data point; λ is the time decay coefficient; n is the total number of data points in the time window; The autocorrelation coefficient ρ is calculated according to the following formula (7): t : Among them, x i is the i-th data point in the time window; k is the number of lag steps; n is the total number of data points in the time window; The calculated mean μ t , standard deviation σ t And the autocorrelation coefficient ρ t Determined as a time series feature; Apply discrete Fourier transform to the data in each time window to convert the time domain data into the frequency domain and obtain the frequency domain representation X f ; Calculate the main frequency f according to the following formula (8) peak : The calculated f peak Determined as frequency feature; The determined timing features and frequency features are spliced together to generate key features that reflect the device operating status and network behavior.
2. The method for mining and analyzing abnormal security behaviors of the industrial Internet according to claim 1 is characterized in that: The extracted feature vectors are input into the baseline model for comparison to detect abnormal behaviors that deviate from the baseline model, including: According to the following formula (4), calculate the new anomaly score d new : Among them, x new is the feature vector corresponding to the extracted feature; If d new >θ, it is judged as abnormal behavior; otherwise, it is judged as normal behavior.
3. The method for mining and analyzing abnormal security behaviors of the industrial Internet according to claim 2 is characterized in that: Through the data collection devices installed on each node of the Industrial Internet, multi-dimensional data in the Industrial Internet is collected in real time, including: The sensors installed on industrial equipment collect the equipment operation status data in real time, and transmit the collected equipment operation status data to the data processing center through the industrial Internet for centralized storage and processing; Capture network traffic data in the industrial Internet through network traffic monitoring equipment; perform packet processing on network traffic data, extract source address, destination address, protocol type and data volume in the data packet; sort network traffic data by timestamp and transmit it to the data processing center for analysis; Deploy logging tools on user terminals and servers to record user operation behaviors; transmit user operation log data to the data processing center through a secure encrypted channel.
4. The method for mining and analyzing abnormal security behaviors of the industrial Internet according to claim 3 is characterized in that: The sensors include temperature sensors, pressure sensors and vibration sensors.
5. The method for mining and analyzing abnormal security behaviors of the industrial Internet according to claim 4 is characterized in that: User operations include login, logout, command execution, and file access.
6. The method for mining and analyzing abnormal security behaviors of the industrial Internet according to claim 5 is characterized in that: Preprocess the multi-dimensional data to generate a preprocessed data set, including: Conduct a preliminary screening of the collected data to eliminate obviously erroneous and invalid data; Check the integrity and consistency of data, merge and remove duplicate data; Identify and process abnormal values and outliers to ensure data accuracy.
7. The method for mining and analyzing abnormal security behaviors of the industrial Internet according to claim 6 is characterized in that: Preprocessing multi-dimensional data to generate preprocessed data sets also includes: Analyze the missing values in the data set to determine the distribution and impact of the missing values; Use linear interpolation to fill missing values in time series data; For non-time series data, mean imputation or multiple imputation methods are used.
8. The method for mining and analyzing abnormal security behaviors of the industrial Internet according to claim 7 is characterized in that: Preprocessing multi-dimensional data to generate preprocessed data sets also includes: Use a moving average filter to smooth the data and reduce the impact of random noise; The data is filtered in the frequency domain by combining high-pass filter and low-pass filter to remove high-frequency and low-frequency noise; The data format is standardized and converted into a unified format and dimension to facilitate subsequent analysis and processing.
Citation Information
Patent Citations
Network information security analysis method and system based on data analysis
CN118784348A