A method for constructing a data anomaly monitoring model based on machine learning

CN110851422BActive Publication Date: 2025-05-09NAT COMPUTER NETWORK & INFORMATION SECURITY MANAGEMENT CENT SHANXI BRANCH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201911078822.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-11-06
Publication Date
2025-05-09
Estimated Expiration
2039-11-06

Smart Images

  • Figure CN110851422B_ABST
    Figure CN110851422B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for constructing a data anomaly monitoring model based on machine learning. Considering the one-sidedness and insufficiency of a single feature analysis result, in the analysis of platform business data, the Pearson correlation coefficient and variance expansion factor are introduced to extract features from different features, and suitable feature data is extracted to enter a clustering model, which greatly improves the accuracy of the model. In addition, for the clustering model, an improved algorithm I-K-means algorithm model of the K-means algorithm is selected. Since the main purpose of the design is to do anomaly processing, the algorithm does not need to perform clustering to the end before finding anomalies, and is faster than other original algorithms. In summary, the model established based on machine learning has the characteristics of self-learning and self-evolution, can adapt to complex and changeable network environments, can detect unknown anomalies, and meet real-time and accurate requirements.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for constructing a data anomaly monitoring model based on machine learning, and belongs to the technical field of data anomaly detection. Background Art

[0002] At present, the coal, metallurgy, chemical industry, electromechanical and other industries are facing a critical period of automation, digitalization and intelligent transformation. A large number of traditional industrial control systems are transforming and upgrading towards the industrial Internet, and the information security situation is becoming increasingly severe. In recent years, network security incidents have occurred frequently, network attacks have intensified, and network security defense technology has lagged behind. With network security issues becoming increasingly prominent today, it is particularly important to detect abnormal network communication behaviors in a timely and effective manner. Abnormal communication behavior refers to situations that deviate from normal data in a network environment. Of course, normal behavior is not fixed. It changes according to changes in user operations, business processes, and network management. The traditional method is to detect abnormal network behavior through static rule matching. It is difficult to detect unknown anomalies and attack types in a dynamic and complex network environment, and cannot meet the requirements of network security detection. Summary of the invention

[0003] The technical problem to be solved by the present invention is to provide a method for constructing a data anomaly monitoring model based on machine learning. The model established based on machine learning has the characteristics of self-learning and self-evolution, can adapt to complex and changeable network environments, can detect unknown anomalies, and meet real-time and accurate requirements.

[0004] In order to solve the above technical problems, the present invention adopts the following technical solution: The present invention designs a method for constructing a data anomaly monitoring model based on machine learning, which is used to obtain a data anomaly monitoring model corresponding to a target application in the industrial Internet, comprising the following steps:

[0005] Step A. Extract the identification fingerprint features corresponding to the communication traffic of the target application in the target historical time period, construct a feature set, and then proceed to step B;

[0006] Step B. Preprocess each identification fingerprint feature in the feature set, update the feature set, and then proceed to step C;

[0007] Step C. Screening each identification fingerprint feature in the feature set, updating the feature set, and then proceeding to step D;

[0008] Step D: Use the feature set to perform model training on a preset specified clustering model to obtain a trained model as a data anomaly detection model corresponding to the target application.

[0009] As a preferred technical solution of the present invention: it also includes step E as follows, after executing step D, entering step E;

[0010] Step E: For the data anomaly detection model corresponding to the target application, use the cross-validation method or the grid search method to adjust parameters and update the data anomaly detection model.

[0011] As a preferred technical solution of the present invention: it also includes step F as follows, after executing step E, entering step F;

[0012] Step F: Apply the Rand Index algorithm and the Silhouette Coefficient algorithm to evaluate the data anomaly detection model.

[0013] As a preferred technical solution of the present invention: in the step B, for each identification fingerprint feature in the feature set, data cleaning, data standardization, and data normalization operations are performed in sequence to achieve preprocessing of the feature set, and then enter step C;

[0014] Among them, data cleaning is used to analyze each identification fingerprint feature in the feature set, find out the missing data in each identification fingerprint feature, fill it with data, and delete the identification fingerprint features whose missing data accounts for more than a preset ratio threshold, thereby realizing the update of the feature set;

[0015] Data standardization is used to standardize each identification fingerprint feature in the feature set, obtain a unified format between each identification fingerprint feature, and realize the update of the feature set;

[0016] Data normalization is used to apply the Logistic function method to normalize each fingerprint feature in the feature set to update the feature set.

[0017] As a preferred technical solution of the present invention: Step C. for each identification fingerprint feature in the feature set, feature deletion, feature selection, and feature derivation operations are performed in sequence to achieve screening of the feature set, and then enter step D;

[0018] Feature deletion, which is used to apply the variance expansion factor method to judge each identification fingerprint feature in the feature set, and delete the identification fingerprint features whose variance expansion factors exceed the preset threshold range, so as to realize the screening of the feature set;

[0019] Feature selection, which is used to apply the Pierreson correlation coefficient method to process each identification fingerprint feature in the feature set, obtain each identification fingerprint feature whose correlation is greater than a preset correlation threshold, and replace all the identification fingerprint features in the feature set with the individual identification fingerprint features to achieve screening of the feature set;

[0020] Feature derivation: for the first-level special diagnosis of each identification fingerprint feature in the feature set, the corresponding second-level features are derived and added to the feature set as each identification fingerprint feature to realize the screening of the feature set.

[0021] As a preferred technical solution of the present invention, in step D, a feature set is used to perform model training for the IK-means algorithm according to the following steps to obtain a trained model as a data anomaly detection model corresponding to the target application;

[0022] Step D1. For each identification fingerprint feature in the feature set M, a preset number of identification fingerprint features are arbitrarily selected as each cluster center, and the remaining identification fingerprint features are clustered based on each cluster center to obtain each cluster, and then enter step D2;

[0023] Step D2. For each cluster, obtain the new cluster center of the cluster, and calculate the average cluster radius of the cluster. Then, take the identification fingerprint features in the cluster that exceed the average cluster radius as the abnormal identification fingerprint features, and classify them into the abnormal candidate feature set N. Then, obtain the intersection of the feature set M and the abnormal candidate feature set N, use the intersection to update the feature set M, and enter step D3.

[0024] Step D3. Determine whether the feature set M converges, that is, obtain the trained model as the data anomaly detection model corresponding to the target application; otherwise, return to step D1.

[0025] The method for constructing a data anomaly monitoring model based on machine learning described in the present invention adopts the above technical solution and has the following technical effects compared with the prior art:

[0026] The method for constructing a data anomaly monitoring model based on machine learning designed by the present invention takes into account the one-sidedness and shortcomings of single feature analysis results. In the analysis of platform business data, the Pearson correlation coefficient and variance expansion factor are introduced to extract features from different features, and suitable feature data is extracted to enter the clustering model, which greatly improves the accuracy of the model; and for the clustering model, an improved algorithm IK-means algorithm model of the K-means algorithm is selected. Since the main purpose of the design is to do exception processing, the algorithm does not need to carry out clustering to the end before finding anomalies, and is faster than other original algorithms; in summary, the model established based on machine learning has the characteristics of self-learning and self-evolution, can adapt to complex and changeable network environments, can detect unknown anomalies, and meet real-time and accurate needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic diagram of the architecture of a method for constructing a data anomaly monitoring model based on machine learning designed by the present invention. DETAILED DESCRIPTION

[0028] The specific implementation modes of the present invention will be further described in detail below in conjunction with the accompanying drawings.

[0029] With network security issues becoming increasingly prominent today, it is particularly important to detect abnormal network communication behaviors in a timely and effective manner. The data communication anomaly detection model based on machine learning proposed in this patent studies the communication rules of industrial Internet devices, cloud platforms, and industrial APPs in key industrial Internet application scenarios, extracts communication traffic identification fingerprint features; cleans and filters invalid and abnormal data in the original data, and supplements and marks missing information data. Then, feature engineering is used to extract features from the processed data, and appropriate features are selected for model training using IK-means, and the model is evaluated to determine the parameters of the model. By collecting a large amount of communication traffic from industrial Internet devices, cloud platforms, and industrial APPs in key industrial Internet application scenarios, based on the model established by model building module training, abnormal communication behaviors are detected, and alarms are issued for possible abnormal communication behaviors.

[0030] In practical applications, the present invention designs a method for constructing a data anomaly monitoring model based on machine learning, as shown in Question 1, which specifically includes the following steps.

[0031] Step A. Study the communication patterns of industrial Internet devices, cloud platforms, and industrial apps in key industrial Internet application scenarios, extract the identification fingerprint features corresponding to the communication traffic of the target application in the target historical time period, build a feature set, and then proceed to step B.

[0032] Step B. For each identification fingerprint feature in the feature set, data cleaning, data standardization, and data normalization operations are performed in sequence to pre-process the feature set, and then step C is entered;

[0033] Among them, data cleaning is used to analyze each identification fingerprint feature in the feature set, find out the missing data in each identification fingerprint feature, fill it with data, and delete the identification fingerprint features whose missing data accounts for more than a preset ratio threshold, thereby realizing the update of the feature set.

[0034] Data standardization is used to standardize each identification fingerprint feature in the feature set to obtain a unified format between each identification fingerprint feature. For example, if there is text in the communication message feature, use Word2vec to vectorize it and unify the data format into a unified format, so that the format type of the data becomes a unified data format; thereby realizing the update of the feature set.

[0035] Data normalization is used to apply the Logistic function method to normalize each fingerprint feature in the feature set so that the data range of all features is compressed between 0 and 1 to ensure that the data will not affect the weight of the model due to the value of the feature, thereby achieving the update of the feature set.

[0036] Step C. For each fingerprint feature in the feature set, perform feature deletion, feature selection, and feature derivation operations in sequence to screen the feature set, and then proceed to step D;

[0037] Feature removal, used to apply the Variance Inflation Factor method:

[0038]

[0039] Get the VIF corresponding to each identification fingerprint feature, make a judgment on each identification fingerprint feature in the feature set, delete the identification fingerprint features whose variance expansion factor exceeds the preset threshold range, and implement the screening of the feature set. For example, for the identification fingerprint feature "communication message header length", if 0<VIF<10, it is normal data, and if VIF>10, the collinearity of the identification fingerprint feature is relatively serious and needs to be deleted.

[0040] Feature selection is used to apply the Pearson correlation coefficient method (Pearson correlation Coefficient), according to the following formula:

[0041]

[0042] Each identification fingerprint feature in the feature set is processed to obtain each identification fingerprint feature whose correlation is greater than a preset correlation threshold, and all identification fingerprint features in the feature set are replaced with the individual identification fingerprint features to achieve screening of the feature set.

[0043] Feature derivation: for the first-level special diagnosis of each identification fingerprint feature in the feature set, the corresponding second-level features are derived and added to the feature set as each identification fingerprint feature to realize the screening of the feature set.

[0044] Step D. Use the feature set to train the model according to the following steps for the IK-means algorithm to obtain the trained model as the data anomaly detection model corresponding to the target application, and then proceed to step E.

[0045] Step D1. For each identification fingerprint feature in the feature set M, a preset number of identification fingerprint features are arbitrarily selected as each cluster center, and the remaining identification fingerprint features are clustered based on each cluster center to obtain each cluster, and then enter step D2.

[0046] Step D2. For each cluster, obtain the new cluster center of the cluster and calculate the average cluster radius of the cluster. Then, take the identification fingerprint features in the cluster that exceed the average cluster radius as the abnormal identification fingerprint features and classify them into the abnormal candidate feature set N. Then, obtain the intersection of the feature set M and the abnormal candidate feature set N, use the intersection to update the feature set M, and enter step D3.

[0047] Step D3. Determine whether the feature set M converges, that is, obtain the trained model as the data anomaly detection model corresponding to the target application; otherwise, return to step D1.

[0048] Step E: For the data anomaly detection model corresponding to the target application, apply the cross-validation method or the grid search method, adjust the parameters, update the data anomaly detection model, and then proceed to step F.

[0049] Step F: Apply the Adjusted rand index algorithm and the Silhouette Coefficient algorithm to evaluate the data anomaly detection model.

[0050] Among them, in the Rand Index algorithm (Adjusted rand index):

[0051]

[0052]

[0053] In RI, C is the category information, a represents the number of element logs of the same category in both C and K, and b represents the number of element logs of different categories in both C and K. The value range of ARI is [-1, 1]. The larger the value, the more consistent the clustering result is with the actual situation. In a broad sense, ARI measures the degree of consistency between the distributions of two data.

[0054]

[0055] In the above Silhouette Coefficient algorithm, for a single sample, let a be the average distance to other samples in the same category, and b be the average distance to the closest samples in different categories. For a sample set, its silhouette coefficient is the average of all sample silhouette coefficients. The range of SS is [-1, 1]. The closer the distance between samples in the same category and the farther the distance between samples in different categories, the higher the score.

[0056] After obtaining the data anomaly detection model corresponding to the target application through steps A to F above, by collecting a large amount of communication traffic from industrial Internet devices, cloud platforms, and industrial APPs in key industrial Internet application scenarios, the model established through model building module training is used to detect abnormal communication behaviors and issue alarms for possible abnormal communication behaviors.

[0057] The data anomaly monitoring model construction method based on machine learning designed by the above technical solution takes into account the one-sidedness and shortcomings of the single feature analysis results. In the analysis of the platform business data, the Pearson correlation coefficient and variance expansion factor are introduced to extract different features, and the appropriate feature data is extracted into the clustering model, which greatly improves the accuracy of the model; and for the clustering model, the improved algorithm IK-means algorithm model of the K-means algorithm is selected. Since the main purpose of the design is to do anomaly processing, the algorithm does not need to carry out clustering to the end before finding anomalies, and is faster than other original algorithms; in summary, the model established based on machine learning has the characteristics of self-learning and self-evolution, can adapt to complex and changeable network environments, can detect unknown anomalies, and meet real-time and accurate needs.

[0058] The embodiments of the present invention are described in detail above with reference to the accompanying drawings, but the present invention is not limited to the above embodiments, and various changes can be made within the knowledge scope of ordinary technicians in this field without departing from the purpose of the present invention.

Claims

1. A method for constructing a data anomaly monitoring model based on machine learning, characterized in that: The method is used to obtain a data anomaly monitoring model corresponding to a target application in the industrial Internet, including the following steps: Step A. Extract the identification fingerprint features corresponding to the communication traffic of the target application in the target historical time period, construct a feature set, and then proceed to step B; Step B. Preprocess each identification fingerprint feature in the feature set, update the feature set, and then proceed to step C; Step C. Screening each identification fingerprint feature in the feature set, updating the feature set, and then proceeding to step D; In the above step C, for each identification fingerprint feature in the feature set, feature deletion, feature selection, and feature derivation operations are performed in sequence to screen the feature set, and then step D is entered; Feature deletion, which is used to apply the variance expansion factor method to judge each identification fingerprint feature in the feature set, and delete the identification fingerprint features whose variance expansion factors exceed the preset threshold range, so as to realize the screening of the feature set; Feature selection, which is used to apply the Pierreson correlation coefficient method to process each identification fingerprint feature in the feature set, obtain each identification fingerprint feature whose correlation is greater than a preset correlation threshold, and replace all the identification fingerprint features in the feature set with the individual identification fingerprint features to achieve screening of the feature set; Feature derivation: for the first-level special diagnosis of each identification fingerprint feature in the feature set, derive the corresponding second-level features, add them to the feature set as each identification fingerprint feature, and realize the screening of the feature set; Step D. Using the feature set, perform model training on the preset specified clustering model to obtain the trained model as the data anomaly detection model corresponding to the target application; In the above step D, the feature set is used to train the model according to the following steps for the IK-means algorithm to obtain the trained model as the data anomaly detection model corresponding to the target application; Step D1. For each identification fingerprint feature in the feature set M, a preset number of identification fingerprint features are arbitrarily selected as each cluster center, and the remaining identification fingerprint features are clustered based on each cluster center to obtain each cluster, and then enter step D2; Step D2. For each cluster, obtain the new cluster center of the cluster, and calculate the average cluster radius of the cluster. Then, take the identification fingerprint features in the cluster that exceed the average cluster radius as the abnormal identification fingerprint features, and classify them into the abnormal candidate feature set N. Then, obtain the intersection of the feature set M and the abnormal candidate feature set N, use the intersection to update the feature set M, and enter step D3. Step D3. Determine whether the feature set M converges, that is, obtain the trained model as the data anomaly detection model corresponding to the target application; otherwise, return to step D1.

2. According to the method for constructing a data anomaly monitoring model based on machine learning as claimed in claim 1, it is characterized by: It also includes step E as follows, after executing step D, entering step E; Step E: For the data anomaly detection model corresponding to the target application, use the cross-validation method or the grid search method to adjust parameters and update the data anomaly detection model.

3. The method for constructing a data anomaly monitoring model based on machine learning according to claim 2, characterized in that: The step F is as follows, after executing step E, the process proceeds to step F; Step F: Apply the Rand Index algorithm and the Silhouette Coefficient algorithm to evaluate the data anomaly detection model.

4. A method for constructing a data anomaly monitoring model based on machine learning according to any one of claims 1 to 3, characterized in that: In the step B, for each identification fingerprint feature in the feature set, data cleaning, data standardization, and data normalization operations are performed in sequence to achieve preprocessing of the feature set, and then enter step C; Among them, data cleaning is used to analyze each identification fingerprint feature in the feature set, find out the missing data in each identification fingerprint feature, fill it with data, and delete the identification fingerprint features whose missing data accounts for more than a preset ratio threshold, thereby realizing the update of the feature set; Data standardization is used to standardize each identification fingerprint feature in the feature set, obtain a unified format between each identification fingerprint feature, and realize the update of the feature set; Data normalization is used to apply the Logistic function method to normalize each fingerprint feature in the feature set to update the feature set.

Citation Information

Patent Citations

  • Industrial control abnormal behavior detection method based on multiple machine learning algorithms

    CN110324316A

  • System for detecting attack suspected anomal event

    KR101623071B1