Industrial internet abnormal data detection method and system based on knowledge graph
By constructing a knowledge graph in the industrial Internet and using graph convolution networks and hierarchical representation learning technology, multivariate time series data of industrial equipment is analyzed, and the problem of insufficient detection accuracy and real-time in the existing technology is solved, efficient and accurate abnormal detection is achieved, and the interpretability of the detection results is enhanced.
Patent Information
- Application Number
- CN202510050381.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-06
AI Technical Summary
The existing industrial Internet anomaly data detection methods are insufficient in the detection accuracy and real-timeness and lack explanatory when dealing with complex data relationships and multivariate data.
Using a knowledge graph-based method, a knowledge graph of industrial equipment and its operating state is constructed through graph convolution networks and hierarchical representation learning technology that can distinguish hierarchical pooling, the relationship between various features in multivariable time series data is analyzed, and outlier classification is performed.
It realizes efficient and accurate abnormal detection, reduces false alarms and missed reports, improves data quality and consistency, and enhances the interpretability of detection results.
Smart Images

Figure CN119938937A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial Internet data governance, and specifically relates to an industrial Internet abnormal data detection method and system based on knowledge graph. Background Art
[0002] With the advent of the Industrial 4.0 era, the Industrial Internet of Things (IIOT), as its core component, is profoundly changing the operation mode of traditional manufacturing and industrial production. The Industrial Internet combines physical devices, sensors, control systems and information technology to achieve the interconnection of equipment, real-time data collection and analysis, and intelligent management of production processes. This change has improved production efficiency and product quality, and promoted the flexibility and sustainability of industrial systems. However, with the expansion and increasing complexity of the Industrial Internet system, data management and analysis face unprecedented challenges, especially in abnormal data detection.
[0003] During the long-term operation of industrial equipment, various abnormal situations may occur, such as equipment failure, abnormal operating parameters, data transmission errors, etc. These abnormal situations may not only cause production interruptions and economic losses, but may also pose a threat to production safety and environmental protection. Timely and accurate detection and response to abnormal data, therefore, paying attention to abnormal data monitoring is crucial to ensuring the stability and reliability of industrial systems.
[0004] At present, abnormal data detection in the industrial Internet mainly relies on traditional statistical methods and methods based on machine learning and deep learning. These methods include but are not limited to:
[0005] 1) Statistical methods: such as mean variance analysis and control charts, which monitor the statistical characteristics of data and identify abnormal points that deviate from the normal range. These methods are simple to implement and have high computational efficiency, but they often show low detection accuracy when dealing with complex data relationships and multivariate data.
[0006] 2) Machine learning methods: including supervised learning and unsupervised learning algorithms, such as support vector machines (SVM), decision trees, cluster analysis, neural networks, etc. These methods can capture complex patterns and nonlinear relationships in data and improve the accuracy of anomaly detection. However, machine learning methods rely on a large amount of labeled data for training, and when faced with high-dimensional data and dynamically changing industrial environments, they often have problems with insufficient model generalization and poor real-time performance.
[0007] 3) Deep learning methods: such as Convolutional Neural Network (CNN), Recurrent Neural Network (RNN), Autoencoder, etc., can automatically extract high-level features from large-scale data and further improve the performance of anomaly detection. However, its "black box" characteristics also limit its widespread application in the industrial field, because industrial systems usually require interpretability of detection results.
[0008] Although the above methods have improved the anomaly detection capabilities in the Industrial Internet to a certain extent, there are still many challenges:
[0009] 1) Insufficient modeling of complex relationships: There are complex dependencies and interactions between devices and data in industrial systems. Traditional methods find it difficult to fully capture and model these relationships, resulting in limited accuracy and comprehensiveness of anomaly detection.
[0010] 2) Balance between real-time performance and efficiency: Industrial production has high requirements for the real-time performance of anomaly detection, but complex models often find it difficult to meet the needs of real-time processing while ensuring high detection accuracy.
[0011] 3) Data heterogeneity and dynamism: The Industrial Internet system involves various types of equipment and data, and the data formats are diverse and dynamically changing. Traditional methods are difficult to effectively process heterogeneous data and adapt to dynamic environments.
[0012] 4) Insufficient explainability: In the industrial field, detection results need to be not only accurate but also explainable so that engineers can understand the cause of the anomaly and take corresponding measures. Many detection methods often lack sufficient explainability, which limits their trust and adoption in practical applications. Summary of the invention
[0013] The purpose of the present invention is to provide an industrial Internet abnormal data detection method and system based on knowledge graph, which is used to solve the technical problem that the existing anomaly detection methods have low accuracy.
[0014] In order to achieve the above object, the present invention adopts the following technical solutions:
[0015] The present invention discloses an industrial Internet abnormal data detection method based on knowledge graph, comprising the following steps:
[0016] Collect data from industrial equipment;
[0017] Preprocess the data of industrial equipment, then perform outlier calibration on the preprocessed data to obtain multivariate time series data;
[0018] A classification model is constructed using graph convolutional networks and hierarchical representation learning technology with distinguishable hierarchical pooling. Multivariate time series data (label column (normal / abnormal)) are used to train and optimize the classification model. The relationship between each feature in the multivariate time series data is analyzed based on the classification model, and outliers are classified according to the relationship between each feature.
[0019] Furthermore, the data of industrial equipment is collected according to the device data transmission protocol in the actual industrial Internet scenario.
[0020] Furthermore, the specific steps of collecting data of industrial equipment and transmitting it to the database are as follows:
[0021] Based on industrial communication protocols, design and develop corresponding data collection programs for various types of industrial equipment in specific industrial scenarios; then formulate data collection strategies based on the characteristics of communication protocols of different industrial equipment, collect data according to data collection programs and data collection strategies, and obtain data values at each point;
[0022] The collected data values of each point and the corresponding equipment point information are stored in the historical database in an orderly manner.
[0023] Furthermore, the type of the industrial communication protocol is Modbus, oBIX or IEC104.
[0024] Furthermore, the specific steps for preprocessing the data in the database are as follows:
[0025] The data stored in the database is preprocessed and converted into time series data in which each feature changes over time.
[0026] Furthermore, after the data stored in the database is preprocessed, the method further includes uniformly formatting the data of different devices of the original data according to the timestamp.
[0027] Furthermore, the specific steps for outlier calibration of the preprocessed data are as follows:
[0028] By analyzing the historical operating data and fault records of industrial equipment as well as expert knowledge, the abnormal characteristics and thresholds of problems occurring in industrial equipment under various working conditions are determined. Based on the abnormal characteristics and thresholds, the abnormal data points in the preprocessed data are preliminarily screened out. The preliminarily screened abnormal data points are further verified and accurately calibrated by combining statistical analysis and machine learning algorithms. The calibrated abnormal points are used to train and optimize the classification model based on graph convolutional networks and hierarchical representation learning technology with distinguishable hierarchical pooling.
[0029] Furthermore, the specific steps of analyzing the relationship between the features in the preprocessed data based on the classification model and classifying outliers based on the learned relationship are as follows:
[0030] Construct a knowledge graph that reflects industrial equipment and its operating status, design a graph convolutional network architecture suitable for the knowledge graph, perform graph convolution operations on the node features in the knowledge graph according to the graph convolutional network architecture, and extract high-dimensional feature representations; then adopt a hierarchical representation learning method to abstract and aggregate high-dimensional feature representations layer by layer through a hierarchical pooling mechanism, output the normal / abnormal probability of time series sample data through a linear layer, and mark and respond to detected anomalies based on the classification results.
[0031] Furthermore, the knowledge graph includes equipment nodes, feature nodes of industrial equipment and the association relationships therebetween.
[0032] The present invention also discloses an industrial Internet abnormal data detection system based on a knowledge graph for implementing the above-mentioned detection method.
[0033] Compared with the prior art, the present invention has the following beneficial effects:
[0034] The present invention discloses a method for detecting abnormal data of industrial Internet based on knowledge graph. By constructing and dynamically maintaining a knowledge graph reflecting industrial equipment and its operating status, combined with a graph convolutional network (GCN) and a hierarchical representation learning technology with distinguishable hierarchical pooling, the complex relationship between the features in the time series data is deeply analyzed, thereby realizing efficient and accurate anomaly detection; by constructing a knowledge graph, the association relationship between industrial equipment and its operating status can be clearly displayed, thereby more accurately identifying abnormal data. The information such as entities, relationships and attributes in the knowledge graph provides rich contextual information for anomaly detection, which helps to reduce false positives and false negatives. The data preprocessing and outlier calibration steps can remove redundancy and noise in the data and improve the quality and consistency of the data. This helps to reduce the computational burden of subsequent processing steps and improve the efficiency of the entire anomaly detection process; the above method solves the technical problem that the existing anomaly detection method has low accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 It is a flow chart of the method for detecting abnormal data in the industrial Internet based on the knowledge graph of the present invention;
[0036] Figure 2 A schematic diagram of the structure of a classification model constructed by the present invention using a graph convolutional network and a hierarchical representation learning technique with distinguishable hierarchical pooling. DETAILED DESCRIPTION
[0037] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0038] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0039] The present invention discloses an industrial Internet abnormal data detection method based on knowledge graph, which specifically includes the following steps:
[0040] Step 1: Data collection and transmission: According to the data transmission protocol of the specific equipment used in the Industrial Internet, collect equipment data and transmit it to the database. Through various industrial communication protocols, ensure that the data of different equipment can be accurately collected and stored;
[0041] Step 2: Data preprocessing: format and preprocess the collected time series data;
[0042] Step 3: Outlier calibration: Based on historical operation data and fault records, use statistical methods or machine learning techniques to calibrate outliers on the preprocessed data;
[0043] Step 4: Build the model: First, build a knowledge graph of industrial equipment and its operating status, use graph convolutional networks and hierarchical representation learning technology with distinguishable hierarchical pooling to analyze the relationship between features in time series data, and classify outliers based on the learned relationship.
[0044] Preferably, according to the data transmission protocol of the specific equipment used in the industrial Internet, the specific implementation method of collecting equipment data and transmitting it to the database is as follows:
[0045] Based on common industrial communication protocols such as Modbus, oBIX, and IEC104, we design and develop corresponding data acquisition programs for various types of equipment in specific industrial scenarios. According to the characteristics of the communication protocols of different devices, we formulate personalized data acquisition strategies to ensure the integrity and accuracy of the data, set the time interval for data collection to balance the real-time nature of the data and the load of the system; match the collected data values of each point with the pre-defined static table to ensure the consistency and traceability of the data. The information of the equipment points and their corresponding values are stored in the historical database in an orderly manner, and a detailed historical data record is established to support subsequent data analysis and anomaly detection.
[0046] Preferably, the specific implementation method of formatting and preprocessing the collected time series data is as follows:
[0047] Preprocess the data stored in the database and convert it into time series data in which each feature changes over time to ensure the uniformity of the data format and facilitate subsequent analysis and processing; then format the data of different devices of the original data uniformly according to the timestamp to ensure data consistency.
[0048] Preferably, based on historical operation data and fault records, statistical methods or machine learning techniques are used to calibrate outliers on preprocessed data, and the specific implementation of identifying and recording potential outliers is as follows:
[0049] By analyzing the historical operating data and fault records of industrial equipment and expert knowledge, the abnormal characteristics and thresholds of various operating problems are determined; a rule-based detection method is applied to preliminarily screen out potential abnormal data points; statistical analysis and machine learning algorithms are combined to further verify and accurately calibrate the preliminarily screened abnormal data; the calibrated outliers are used to train and optimize the classification model based on graph convolutional networks and hierarchical representation learning technology with distinguishable hierarchical pooling, thereby improving the detection accuracy and robustness of the model.
[0050] Preferably, the relationship between features in time series data is analyzed by using graph convolutional networks and hierarchical representation learning technology with distinguishable hierarchical pooling, and the specific method for outlier classification based on the learned relationship is as follows:
[0051] Construct a knowledge graph that reflects industrial equipment and its operating status, including equipment nodes, feature nodes and the relationships between them; define the edge types and weights between nodes to represent the mutual influence and dependency between different features; design a graph convolution network architecture suitable for the knowledge graph that can effectively capture local and global feature information between nodes; perform graph convolution operations on node features in the knowledge graph to extract high-dimensional feature representations; ensure the distinguishability during the hierarchical pooling process so that features at different levels can be effectively distinguished and combined, and use a hierarchical representation learning method to abstract and aggregate feature representations layer by layer through a hierarchical pooling mechanism to capture multi-level feature information. After GCN and hierarchical pooling, the normal / abnormal probability of the time series sample data is output through the linear layer, and the detected abnormal data is marked and responded to based on the classification results.
[0052] The present invention is further described in detail below in conjunction with the accompanying drawings:
[0053] The model adopted by the present invention is specifically described as follows:
[0054] The anomaly detection method consists of an offline training module and an online testing module, such as Figure 1 As shown. Throughout the training and testing process, the processed data are all multivariate time series, including multiple feature columns and a label column. In the offline training phase, the system will learn the feature change rules that can distinguish normal and abnormal states based on the input multivariate time series data, and build a model that adapts to the target scenario; in the online testing phase, the model will analyze the real-time data stream, capture the dynamic changes of the dependencies between sensors, identify potential abnormal behaviors, thereby realizing anomaly detection and fault diagnosis of the integrated energy system, and providing support for the stable operation of the system.
[0055] A. Model Framework
[0056] The input original time series is passed through the linear layer to extract the key time features, and then the attention weight matrix between each feature column is obtained through the multi-head attention mechanism, and then the adjacency matrix is generated through GumbelSoftMax. The adjacency matrix and the original time series are input into two L-layer GCNs, the nodes are clustered or pooled together, and the node embedding matrix and the allocation matrix are generated respectively. Then the embedding matrix and the allocation matrix are used as the input of the next layer of GCN, and the input is continuously coarsened. Finally, the output cluster node embedding matrix is output as a normal / abnormal probability matrix through the linear layer.
[0057] B. Building a knowledge graph
[0058] Input device sensor sample data S at the same time t ∈R 1×N (N physical points).
[0059] First, a linear layer is used to extract key time features H from the original time series input data. t .H t =S t W l , W l represents learnable parameters.
[0060] At the same time, in order to achieve more efficient feature capture and reasoning of complex dependencies, a self-attention mechanism model is adopted. The self-attention mechanism can dynamically adjust the weight distribution between time steps and highlight important time series features, thereby better representing the overall performance of the model.
[0061] The goal of the encoder is to learn the latent dependencies of time series features A1|S t Following the distribution, you can use A1 represents the dependency between features, which is obtained through the attention mechanism:
[0062] Q=H t W Q
[0063] K=H t W K
[0064] Θ=QK T ∈R N×N ; (1)
[0065] Among them, H t is the extracted key time feature, W Q ,W K It is a trainable parameter matrix, Q represents Query, K is key, and Θ is the similarity matrix, which represents the relationship weight between nodes.
[0066] In order to improve the model's fitting ability and focus on different aspects of information, we further use the multi-head attention mechanism to define multiple different sets of W for the same input X. Q ,W K , and finally different parameters are obtained by calculating respectively.
[0067] Q i =H t W i Q
[0068] K i =H t W i K
[0069]
[0070] Among them, W i Q ,W i K represents the training parameter matrix of the i-th attention head, Q i ,K i They represent the Query and key of the i-th head respectively.
[0071] head i represents the i-th attention head, 1≤i≤h, h is the number of heads, W 0 is the weight parameter of linear projection,
[0072]
[0073] MultiHead(S)∈R h×N×N The attention weight matrix between each feature, that is, the relationship matrix between each node, is obtained through the multi-head attention mechanism. There are h types of dependencies in total, and the SoftMax function is used to ensure that the sum of the probabilities of all these types is equal to 1.
[0074] The obtained attention weight matrix can be passed through Gumbel SoftMax to get the adjacency matrix of the original data. Gumbel SoftMax allows the process of sampling from discrete distributions in the model to become differentiable, thereby allowing the gradient to be used to update the model parameters during back propagation, that is, the following process:
[0075] Unnormalized probability for the categorical distribution Add Gumbel noise and use the SoftMax function to Converted to probability distribution, and introduced temperature parameter τ to control the smoothness of the distribution.
[0076]
[0077] y soft Convert to one-hot encoding and you get:
[0078]
[0079] In the forward propagation, one-hot encoding is used to obtain the adjacency matrix between the feature columns of the original time series, that is, the one-hot hard-coded y=y output during the forward propagation hard . However, during the back propagation, due to the gradient back propagation, hard Unable to perform derivation, resulting in the inability to backpropagate the gradient, the gradient is still based on y soft , retaining the gradient information of the sample, allowing the optimizer to update the parameters. That is, the adjacency matrix is obtained
[0080] C. Graph Convolutional Network GCN and Hierarchical Representation Learning Technology with Distinguishable Hierarchical Pooling
[0081] The input is a multivariate time series sample S t ∈R M×N , and the adjacency matrix A ∈ {0, 1} obtained after processing the attention weight matrix by Gumbel SoftMax N×N .
[0082] Based on the graph convolutional neural network GCN, it can be trained in a supervised manner within an end-to-end learning framework.
[0083]
[0084] Among them, W (k) is the weight matrix of the Kth layer of the neural network, and ReLU is the activation function. H (k) is the node embedding calculated after K steps of GCN. The input node embedding at the initial message passing iteration (k = 1) uses the original time series H (0) = S for initialization. The complete GCN module will run the above formula iteratively K times to generate the final output node embedding Z = H (k) , where k is usually between 2 and 6. The GCN module obtained by K message passing iterations using the adjacency matrix A and the original input time series S is denoted as Z = GCN(S, A).
[0085] C.1 Graph Convolutional Network GCN
[0086] Traditional GCN models are usually flat, only performing information transfer at the edges of the graph. The output of each layer depends on the features of its neighbor nodes, lacking deep graph structure information. By stacking multiple GCN modules in a hierarchical manner, an attempt is made to enable the model to learn higher-order graph features. Through a pooling operation, a new coarsened graph containing m nodes is generated, where m < n, for dimensionality reduction and aggregation of the input graph. This entire process is repeated L times, with a coarser graph output each time. By stacking multiple GCN modules, the model can learn increasingly abstract graph structure information at each layer, as Figure 2 shown.
[0087] Use two independent GCNs to generate the embedding matrix Z (l) and the assignment matrix S (l) , with the input being the cluster node feature matrix X (l) of the Lth layer and the coarsened adjacency matrix A (l) . The embedding GCN of the Lth layer is a standard GCN module applied to the following input:
[0088] Z(l) =GNN l,embed (A (l) ,X (l) ) (7)
[0089] The cluster assignment matrix learned at layer L is expressed as The row corresponds to n of layer L L nodes, the columns correspond to n nodes in the L+1th layer L+1 nodes, and softly assigns each node in layer L to a cluster in the next coarsening layer L+1. The pooling GCN in layer L uses the input feature matrix X (l) and the adjacency matrix A (L) Generate the allocation matrix S (l) :
[0090] S (L) =softmax(GNN l,pool (A (l) ,X (l) )) (8)
[0091] Among them, the SoftMax function is applied row by row. l,pool The output dimension of corresponds to the predefined maximum number of clusters in layer l and is a hyperparameter of the model.
[0092] When L=0, input X (l) is the original time series, A (l) is an adjacency matrix generated from a time series.
[0093] C.2 Hierarchical Representation Learning with Differentiable Hierarchical Pooling
[0094] Use the cluster node embedding matrix Z extracted from layer L that is useful for graph classification (l) And the learned cluster assignment matrix S (l) The input pooling layer is used as the input of another L-layer GCN model. Given the input diffpoolLayer(A (l+1) ,X (l+1) )=DIFFPOOL(A (l) ,Z (l) ) coarsens the input graph and generates a new adjacency matrix A for the coarsened graph (l+1) and the new node embedding matrix X (l+1) , the equation is as follows:
[0095]
[0096] Through the above formula (9) and formula (10), combined with the clustering allocation matrix S (l) and the aggregate node embedding matrix Z (l) n l+1nodes to generate a new embedding matrix X (l+1) , and generate a coarsened adjacency matrix A (l+1) . A (l+1) is a fully connected edge-weighted graph, where is the connection strength between node i and node j, X (l+1) Denotes the embedding of the i-th row corresponding to node i. The coarsened adjacency matrix A (l+1) and the new embedding matrix X (l+1) Together they serve as the input of the next GCN layer, that is, they are input into formulas (7) and (8).
[0097] After multiple GCN layers and layered pooling, the final GCN structure output is expressed as the probability matrix output∈R of the node normal / abnormal through the linear layer output n×2 .
[0098] D. Combined application of multiple loss functions
[0099] In the multivariate time series outlier detection task in industrial scenarios, the combination of link prediction loss, cross entropy loss and Focal Loss loss functions can effectively improve the robustness and accuracy of the model when processing complex time series data. Link prediction loss helps the model better understand the correlation between data, cross entropy loss is used to optimize classification accuracy, and Focal Loss effectively addresses the problem of data imbalance, ensuring that the model maintains high performance when identifying minority class anomalies.
[0100] 1. Link prediction loss: Make the predicted adjacency matrix consistent with the actual adjacency matrix.
[0101] 2. The cross entropy loss function can minimize the difference between the predicted probability and the true label, helping the model to effectively distinguish normal data from abnormal data.
[0102]
[0103] Where N is the total number of samples, y i is the true label (0 or 1) of the i-th sample, p i is the predicted probability of the ith sample, w i represents the weight of the i-th sample.
[0104] 3. Focal loss function
[0105] In outlier detection, most of the data is usually normal data (i.e., negative class), while the proportion of abnormal data (i.e., positive class) is very small. Traditional loss functions may cause the model to over-focus on normal data. Focal loss is based on binary cross entropy CE. Through a dynamic scaling factor, it can dynamically reduce the weight of easily distinguishable samples during training, thereby focusing on difficult-to-distinguish samples.
[0106] FL(p t )=-α t (1-p t ) γ log(p t ); (12)
[0107] Among them, (1-p t ) γ The loss contribution of easy-to-distinguish samples can be reduced, that is, the imbalance of the number of simple / difficult-to-distinguish samples can be controlled by γ, and α t The imbalance in the number of positive and negative samples can be suppressed.
[0108] E. Evaluation Method
[0109] In the present invention, in order to comprehensively evaluate the performance of the industrial Internet abnormal data detection method based on knowledge graph, two evaluation methods are adopted: label evaluation and degree evaluation method.
[0110] E.1 Label evaluation method
[0111] The label evaluation method aims to evaluate the classification performance of the model by comparing the consistency between the model prediction results and the actual labels. The specific steps are as follows:
[0112] 1. Model testing and prediction: The test set divided by prediction is input into the optimal parameter model obtained during the training process. The model analyzes each test sample and outputs the corresponding probability score of normal state or abnormal state.
[0113] 2. Label generation and comparison: According to the probability score, a threshold is set, and samples with scores higher than the threshold are marked as "abnormal", and samples with scores lower than the threshold are marked as "normal", thereby generating abnormal labels. The generated predicted labels are compared with the actual labels marked in the test set to form a confusion matrix (Confusion Matrix), including true positives (True Positive, TP), true negatives (True Negative, TN), false positives (False Positive, FP) and false negatives (False Negative, FN).
[0114] Note:
[0115] Positive Class: abnormal state, label is 1.
[0116] Negative Class: Normal state, label is 0.
[0117] 3. Confusion Matrix Construction
[0118] Prediction Normal(0) Predicting anomalies (1) Actual normal (0) TN FP Actual exception (1) FN TP
[0119] 4. Performance index calculation:
[0120] Accuracy: measures the proportion of samples that the model predicts correctly. The calculation formula is:
[0121]
[0122] Precision: The proportion of samples that are actually abnormal among all samples predicted to be abnormal. The calculation formula is:
[0123]
[0124] Recall: The proportion of samples that are correctly predicted to be abnormal among all samples that are actually abnormal. The calculation formula is:
[0125]
[0126] F1 score (F1-score): The harmonic mean of precision and recall, which comprehensively evaluates the performance of the model. The calculation formula is:
[0127]
[0128] AUC: By drawing the ROC curve, the area under the curve is calculated. The closer the AUC value is to 1, the better the classification performance of the model.
[0129] 5. Result analysis and optimization:
[0130] The classification performance of the model is comprehensively evaluated by calculating indicators such as accuracy, precision, recall, F1 score, AUC, TPR, and FPR. Based on the evaluation results, the model structure is optimized, hyperparameters are adjusted, or more features are introduced to improve the overall detection performance.
[0131] E.2 Degree Evaluation Method
[0132] The degree evaluation method uses the interdependence between sensors to evaluate the anomalies in multivariate time series by calculating the sum of the in-degree and out-degree of each sensor as the anomaly score (Degree). The specific steps are as follows:
[0133] 1. Sensor dependency analysis: In multivariate time series data, there are complex dependencies between sensors. The knowledge graph is used to build a relationship network between sensors to reflect the mutual influence and dependency between sensors.
[0134] 2. Degree calculation and anomaly scoring
[0135]
[0136] in, They represent the mobile filtered in / out degree of node i in the learning graph, respectively, taking into account the bidirectional dependency changes of the time series and providing a comprehensive indicator that can reveal the changes in the system state.
[0137] 3. Abnormal definition and filtering:
[0138] Anomalies are usually based on sudden changes rather than slow gradual changes, and it is necessary to exclude slow-changing, normal trend fluctuations. Use moving average filtering to smooth time series data, reduce noise interference, and suppress gradually changing noise signals. The specific steps include:
[0139] Moving average filtering: Select an appropriate time window, calculate the average value of the data in the window, and replace the current data point to smooth the data sequence, that is, smooth the local random fluctuations.
[0140] Outlier identification: In the smoothed data, outliers are identified based on the mutation of the degree value.
[0141] 4. Local information capture and abnormal pattern discovery: Through the degree evaluation method, more local information can be captured and fine-grained abnormal patterns can be identified. The relationship between sensors in the comprehensive knowledge graph can be used to further analyze the causes of abnormal points and improve the accuracy and explainability of anomaly detection.
[0142] In summary, with the rapid development of the Industrial Internet, traditional abnormal data detection methods can no longer meet the needs of efficient and accurate detection in complex industrial environments. By constructing a knowledge graph of each device point in industrial equipment, combining the graph convolutional network GCN with the hierarchical representation learning technology with distinguishable hierarchical pooling, and analyzing the correlation and dependency between the features of the time series, this paper provides an innovative solution for data monitoring and analysis in the Industrial Internet, and promotes the safety, efficiency and sustainability of industrial production.
[0143] The above contents are only for explaining the technical idea of the present invention and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the claims of the present invention.
Claims
1. A method for detecting abnormal data in industrial Internet based on knowledge graph, characterized in that: The following steps are involved: Collect data from industrial equipment; Preprocess the data of industrial equipment, then perform outlier calibration on the preprocessed data to obtain multivariate time series data; A classification model is constructed using graph convolutional networks and hierarchical representation learning technology with distinguishable hierarchical pooling. Multivariate time series data is used to train and optimize the classification model. The relationship between each feature in the multivariate time series data is analyzed based on the classification model, and outliers are classified according to the relationship between each feature.
2. According to the method for detecting abnormal data in industrial Internet based on knowledge graph in claim 1, it is characterized in that: The data collected from industrial equipment is carried out according to the equipment data transmission protocol in the actual industrial Internet scenario; The multivariate time series data includes abnormal points and normal data obtained by performing abnormal value calibration on the preprocessed data.
3. According to the method for detecting abnormal data in industrial Internet based on knowledge graph according to claim 2, it is characterized in that: The specific steps of collecting data of industrial equipment and transmitting it to the database are as follows: Based on industrial communication protocols, design and develop corresponding data collection programs for various types of industrial equipment in specific industrial scenarios; then formulate data collection strategies based on the characteristics of communication protocols of different industrial equipment, collect data according to data collection programs and data collection strategies, and obtain data values at each point; The collected data values of each point and the corresponding equipment point information are stored in the historical database in an orderly manner.
4. According to the method for detecting abnormal data in industrial Internet based on knowledge graph according to claim 3, it is characterized in that: The type of the industrial communication protocol is Modbus, oBIX or IEC104.
5. The method for detecting abnormal data in industrial Internet based on knowledge graph according to claim 1 is characterized in that: The specific steps for preprocessing the data in the database are as follows: The data stored in the database is preprocessed and converted into time series data in which each feature changes over time.
6. According to the method for detecting abnormal data in industrial Internet based on knowledge graph according to claim 5, it is characterized in that: After the data stored in the database is preprocessed, the method further includes uniformly formatting the data of different devices of the original data according to the timestamp.
7. The method for detecting abnormal data in industrial Internet based on knowledge graph according to claim 6 is characterized in that: The specific steps for outlier calibration of preprocessed data are as follows: By analyzing the historical operation data and fault records of industrial equipment and expert knowledge, the abnormal characteristics and thresholds of problems in industrial equipment under various working conditions are determined; based on the abnormal characteristics and thresholds, the abnormal data points in the preprocessed data are preliminarily screened out; combined with statistical analysis and machine learning algorithms, the preliminarily screened abnormal data points are further verified and accurately calibrated; The calibrated outliers are used to train and optimize a classification model based on graph convolutional networks and hierarchical representation learning techniques with differentiable hierarchical pooling.
8. The method for detecting abnormal data in industrial Internet based on knowledge graph according to claim 7 is characterized in that: The specific steps for analyzing the relationship between the features in the preprocessed data based on the classification model and classifying outliers based on the learned relationship are as follows: Construct a knowledge graph that reflects industrial equipment and its operating status, design a graph convolutional network architecture suitable for the knowledge graph, perform graph convolution operations on the node features in the knowledge graph according to the graph convolutional network architecture, and extract high-dimensional feature representations; then adopt a hierarchical representation learning method to abstract and aggregate high-dimensional feature representations layer by layer through a hierarchical pooling mechanism, output the normal / abnormal probability of time series sample data through a linear layer, and mark and respond to detected anomalies based on the classification results.
9. The method for detecting abnormal data in industrial Internet based on knowledge graph according to claim 8 is characterized in that: The knowledge graph includes equipment nodes, feature nodes of industrial equipment and the association relationships between them.
10. An industrial Internet abnormal data detection system based on knowledge graph, characterized in that: Used to implement the detection method described in any one of claims 1 to 9.