Production abnormity early warning method based on abnormity detection
By conducting preliminary abnormality detection and repair of industrial production data, an automatic codec was built for reconstruction and training, which solved the problem of poor ability of automatic codec to handle abnormal data, and achieved efficient and automated abnormal warning effect, which was suitable for industrial production environments.
Patent Information
- Application Number
- CN202510448509.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-08-01
AI Technical Summary
In the existing abnormal point detection methods, the automatic codec has poor ability to handle abnormal data, resulting in errors in the built model and affecting the abnormal detection effect. In addition, the traditional evaluation system relies on manual annotation, making it difficult to effectively evaluate the abnormal detection effect in industrial production.
By collecting multi-dimensional time series data of the equipment production process, performing preliminary anomaly detection and repair, building an automatic codec for reconstruction training, and using isolated forest algorithms and linear interpolation to repair abnormal points, calculating the reconstruction error threshold for abnormal judgment, and evaluating model performance with dynamic early warning indicators.
It improves the robustness and accuracy of the model, reduces the cost of manual labeling, increases the effective warning rate, adapts to the complexity and high real-time requirements of industrial scenarios, and supports the expansion of diverse scenarios.
Smart Images

Figure CN120408188A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of digital data processing, and in particular relates to a production anomaly early warning method based on anomaly detection. Background Art
[0002] Generally, classic outlier detection methods are usually divided into four categories: statistics-based, clustering-based, classification-based, and proximity-based. However, with the development of cloud computing and big data technologies, the traditional method of mining abnormal node information through single-node computing can no longer meet the growing data computing needs. Artificial intelligence technology represented by deep learning technology has provided a new research direction for outlier detection. The more common deep learning-based anomaly detection method uses automatic codecs to perform dimensionality reduction modeling on normal high-dimensional data, and predicts the test data based on the constructed model. If the reconstruction deviation is too large, it will be identified as abnormal data.
[0003] Automatic codecs often place high demands on the quality of normal data during model training. However, real-world industrial data often contains abnormal data, which can lead to errors in the constructed model and affect the effectiveness of anomaly detection. Furthermore, traditional anomaly detection evaluation systems are typically based on metrics such as accuracy, recall, and F1 score. The calculation of these metrics requires accurate labeling of abnormal data. However, in actual industrial production processes, only the nodes where anomalies occur can be clearly labeled. Unless a large amount of manpower is invested in manual labeling, it is impossible to effectively evaluate the system's early warning effectiveness. Summary of the Invention
[0004] In order to solve the problem that the automatic codec in the existing anomaly point detection method has poor ability to process abnormal data, resulting in errors in the constructed model, thereby affecting the anomaly detection effect, the present invention provides a production anomaly early warning method based on anomaly detection, which repairs the detected anomalies and then inputs them into the automatic codec for training, thereby improving the anomaly detection accuracy of the model.
[0005] The technical solutions of the present invention are as follows:
[0006] A production anomaly early warning method based on anomaly detection, the method comprising the following steps:
[0007] Step 1: Collect parameters from each node of the equipment production process to form a multi-dimensional time series, which includes normal production data and abnormal production data with corresponding annotations;
[0008] Step 2: Divide the historical data into a training set and a validation set, and remove abnormal production data;
[0009] Step 3: Perform preliminary anomaly detection on the training set and repair abnormal points;
[0010] Step 4: Construct an autoencoder, input the repaired data into the network, and perform reconstruction training;
[0011] Step 5: Input the validation set into the trained autoencoder, calculate the reconstruction error threshold, and determine whether the data is abnormal according to the score.
[0012] Step 6: Calculate the anomaly warning evaluation index based on the production anomaly annotation points to evaluate the model performance.
[0013] Furthermore, the operating parameters of each node of the production equipment are collected in real time through sensors to form a multi-dimensional time series.
[0014] Furthermore, step 3 is specifically: using the Isolation Forest algorithm to perform preliminary anomaly detection on the training set, the input data in the preliminary anomaly detection includes normal production data, and linear interpolation is used to repair the anomaly points.
[0015] Furthermore, the dimension of the input layer of the autoencoder is consistent with the number of data features. The autoencoder is composed of two fully connected layers. Except for the final output using the Sigmoid function, the activation function uses the ReLU function for the rest.
[0016] Furthermore, the training input data of the autoencoder is the preliminarily repaired normal production data, while the input data in the verification stage is the unrepaired normal production data.
[0017] Furthermore, the formula for calculating the reconstruction error threshold in step 5 is specifically as follows:
[0018]
[0019] where: μ represents the mean of the reconstruction error, σ represents the standard deviation, the reconstruction error threshold is μ + 3σ, e j represents the reconstruction error of the j-th sample, and m is the number of samples in the validation set;
[0020] For all data points, judge whether they exceed the reconstruction error threshold. If they exceed, they are determined to be abnormal.
[0021] Furthermore, the anomaly warning evaluation index in step 6 includes the effective warning rate and the missed warning rate within a preset time range, and the calculation formulas are as follows:
[0022]
[0023] The higher the effective warning rate, the stronger the model's ability to predict actual anomalies, and the more timely and accurate it can provide anomaly warning information for production; the lower the missed warning rate, the stronger the model's ability to capture anomalies.
[0024] Compared with the prior art, the present invention has the following beneficial effects:
[0025] (1) Improve model robustness: The repaired training data significantly reduces the model's dependence on "clean data" and adapts to the complexity of data mixing in industrial scenarios.
[0026] (2) Reduce labor costs: No need to label a large amount of abnormal data, and automatic evaluation can be achieved through dynamic early warning indicators.
[0027] (3) Efficient early warning capability: Experiments show that compared with direct anomaly detection without taking any repair or remediation measures, the 12-hour effective early warning rate is increased by 16% under the same missed warning rate, which is suitable for industrial production environments with high real-time requirements.
[0028] (4) Strong scalability: It supports the replacement of different anomaly detection algorithms (such as LOF, One-Class SVM) or encoder structures (such as variational autoencoder) to adapt to diverse scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 Algorithm steps for multi-level prediction;
[0030] Figure 2 Flowchart of the production anomaly early warning system based on multi-level prediction;
[0031] Figure 3 It is the automatic codec network structure. DETAILED DESCRIPTION
[0032] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0033] A production anomaly early warning method based on anomaly detection, the method comprising the following steps:
[0034] Step 1: Collect parameters from each node of the equipment production process to form a multi-dimensional time series, which includes normal production data and abnormal production data with corresponding annotations;
[0035] Step 2: Divide the historical data into a training set and a validation set, and remove abnormal production data;
[0036] Step 3: Perform preliminary anomaly detection on the training set and repair abnormal points;
[0037] Step 4: Construct an automatic encoder-decoder and input the repaired data into the network for reconstruction training;
[0038] Step 5: Input the validation set into the trained automatic encoder / decoder, calculate the reconstruction error threshold, and determine whether it is abnormal data based on the score.
[0039] Step 6: Calculate the anomaly warning evaluation index based on the production anomaly marking points and evaluate the model performance.
[0040] The present invention will be further described below in conjunction with a specific embodiment:
[0041] See Figure 1 , a production anomaly warning method based on anomaly detection, the method is as follows:
[0042] Step 1: Multidimensional time series data collection, that is, collect the parameters of each node in the equipment production process to form a multi-dimensional time series, including two types of data: normal production and abnormal production, and there are obvious markings;
[0043] The operating parameters (such as temperature, pressure, speed, etc.) of each node of the production equipment are collected in real time through sensors to form 64-dimensional time series data, the data interval is 10 seconds, the signal data of production anomalies are collected, and the multi-dimensional data are marked.
[0044] Step 2: Data preprocessing and divide the data into a training set and a validation set;
[0045] Data preprocessing: Aggregate the historical data at intervals of 10 seconds, use mean smoothing for multiple data falling within the same time interval, and use forward interpolation to fill in the missing data; after removing the data marked as production anomalies, in order to eliminate the differences in dimension and numerical range of different features and improve the model training effect, use max-min normalization, and the formula is: where x min and x max are the minimum and maximum values of the data respectively;
[0046] Dataset division: Divide the preprocessed data into a training set and a validation set according to a ratio of 7:3.
[0047] Step 3: Primary anomaly detection and repair; Use the isolation forest algorithm to perform preliminary anomaly detection on the training set and repair the anomaly points;
[0048] First-level anomaly detection (isolation forest algorithm): Use the isolation forest algorithm to perform preliminary anomaly detection on the training set, set the number of trees to 100, the subsample size to 256, and detect potential anomaly points.
[0049] Data repair: For the detected anomaly points, use linear interpolation to repair, and fill the anomaly values based on the mean of the 10 normal
[0050] data points before and after, and generate the repaired training set, and its calculation formula is: where x t-1 and x t+1respectively represent the i-th normal data point before and after the outlier, x repaired is the data value after repair, used to replace the outlier.
[0051] Step 4: Construct an autoencoder, input the repaired training set data into the network, and perform reconstruction training;
[0052] Network structure: Build a deep autoencoder. The dimension of the input layer is the same as the number of data features. Both the encoder and decoder consist of two fully connected layers. The activation function uses ReLU for all except the final output which uses Sigmod. Its network structure is as Figure 3 shown. The encoder compresses the feature dimension from 64 dimensions to 16 dimensions through two fully connected layers, and the decoder restores the feature dimension from 16 dimensions to 64 dimensions through two fully connected layers. The loss function during the training process is the mean squared error (MSE), and the calculation formula is: where y i is the input data, is the reconstructed output, and n is the number of samples.
[0053] Training process: Input the repaired training set into the autoencoder for unsupervised training. The number of iterations is 500 rounds, the batch size is 64, and the optimizer is Adam (learning rate 0.001). With the minimization of the loss function as the training goal, construct an encoding and decoding model.
[0054] Step 5: Input the validation set data into the trained autoencoder, calculate the reconstruction error score, and determine whether it is abnormal data based on the score;
[0055] Validation set evaluation: Input the validation set data into the trained autoencoder to calculate the reconstruction error threshold of the data;
[0056] Abnormality determination: The specific formula for calculating the reconstruction error threshold is as follows:
[0057]
[0058] where: μ represents the mean of the reconstruction error, σ represents the standard deviation, and the reconstruction error threshold is μ + 3σ (the reconstruction error threshold for determining abnormal data is equal to the mean of the reconstruction error plus three times the standard deviation), e j represents the reconstruction error of the j-th sample, and m is the number of samples in the validation set;
[0059] For all data points, determine whether they exceed the reconstruction error threshold. If they exceed, they are determined to be abnormal.
[0060] Step 6: Calculate the abnormal warning evaluation index based on the production abnormal annotation points of the original data.
[0061] To measure the effectiveness of the model's early warning, in this embodiment, the effective early warning rate is calculated every 12 hours (), and its calculation formula is as follows:
[0062]
[0063] The higher this ratio, the stronger the model's ability to predict actual anomalies, and it can provide early warning information for production more timely and accurately, helping enterprises take measures in advance to reduce production losses.
[0064] The above-mentioned anomaly early warning evaluation indicators include the effective early warning rate and the missed early warning rate within a preset time range, and the calculation formulas are as follows:
[0065]
[0066] The lower the missed early warning rate, the stronger the model's ability to capture anomalies.
[0067] The above are only the embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied to other related technical fields, shall be similarly included in the patent protection scope of the present invention.
Claims
1. A production anomaly warning method based on anomaly detection, characterized in that, The method comprises the following steps: Step 1: Collect parameters from each node of the equipment production process to form a multi-dimensional time series, which includes normal production data and abnormal production data with corresponding annotations; Step 2: Divide the historical data into a training set and a validation set, and remove abnormal production data; Step 3: Perform preliminary anomaly detection on the training set and repair abnormal points; Step 4: Construct an automatic encoder-decoder and input the repaired data into the network for reconstruction training; Step 5: Input the validation set into the trained automatic encoder / decoder, calculate the reconstruction error threshold, and determine whether it is abnormal data based on the score. Step 6: Calculate the abnormal warning evaluation index based on the production abnormality marked points and evaluate the model performance.
2. The production anomaly warning method based on anomaly detection according to claim 1, wherein In step 1: the operating parameters of each node of the production equipment are collected in real time through sensors to form a multi-dimensional time series.
3. The production anomaly warning method based on anomaly detection according to claim 1, characterized in that, Step 3 is specifically as follows: using the isolation forest algorithm to perform preliminary anomaly detection on the training set, the input data in the preliminary anomaly detection includes normal production data, and linear interpolation is used to repair abnormal points.
4. A production anomaly warning method based on anomaly detection according to claim 1, characterized in that, The input layer dimension of the automatic codec is consistent with the number of data features. The automatic codec consists of two fully connected layers. The activation function uses the ReLU function except for the last output which uses the Sigmoid function.
5. The production anomaly warning method based on anomaly detection according to claim 1, characterized in that, The training input data of the automatic codec is the normal production data after preliminary repair, while the input data in the verification stage is the unrepaired data of the normal production data.
6. The production anomaly warning method based on anomaly detection according to claim 1, wherein, The formula for calculating the reconstruction error threshold in step 5 is as follows: Among them: μ represents the mean of the reconstruction error, σ represents the standard deviation, the reconstruction error threshold is μ + 3σ, and e j represents the reconstruction error of the j-th sample, and m is the number of samples in the validation set; For all data points, determine whether they exceed the reconstruction error threshold. If they exceed, they are considered abnormal.
7. The production anomaly warning method based on anomaly detection according to claim 1, wherein The abnormal warning evaluation indicators described in step 6 include the effective warning rate and missed warning rate within the preset time range, and the calculation formula is as follows: The higher the effective warning rate, the stronger the model's ability to predict the occurrence of actual abnormalities, and the more timely and accurate it can provide abnormal warning information for production; the lower the missed warning rate, the stronger the model's ability to capture abnormalities.
Citation Information
Cited By
Insulated cable production quality detection method and system
CN120612017A