Production line fault identification method based on BiLSTM-Logistic regression
By using the parallel hybrid model of BiLSTM-Logistic regression in production line fault recognition, the problem that traditional methods cannot accurately identify fault types and high costs is solved, and efficient and accurate fault recognition effect is achieved.
Patent Information
- Application Number
- CN202510673249.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional production line fault identification methods cannot accurately identify the fault type, and rely on a large number of sensor equipment and manual inspections, resulting in high cost and low efficiency.
Using a parallel hybrid fault identification model based on BiLSTM-Logistic regression, the long-term dependence of timing data is mined through the BiLSTM network, the logistic regression model quickly and explicitly modeled non-timing features, and the final fault identification is performed through the SVM classifier.
It realizes efficient identification of production line failures, excellent recall effect and extremely high accuracy, avoiding excessive smoothing of structured data by the depth model.
Smart Images

Figure CN120180253A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of industrial fault diagnosis, and particularly to a production line fault identification method based on BiLSTM-Logistic regression. Background Art
[0002] Traditional production line fault identification mainly relies on deploying sensor devices at various positions on the production line. This method can detect the occurrence of faults in a timely manner, but it cannot directly feedback the types of faults occurring on the production line, and manual troubleshooting of fault types is required. Deploying sensors on each production line undoubtedly requires significant capital investment, and troubleshooting the cause of faults also requires labor costs. The application of intelligent control technology can bring higher flexibility, reliability and self-adaptability to the automated production line. For example, integrating fault intelligent alarm technology into the automated production line can effectively detect the occurrence of faults and accurately identify the types of faults, realizing the efficient utilization of personnel, equipment, raw materials, etc., so as to maximize production efficiency and reduce resource waste. Therefore, it is crucial and of great significance to explore how to make full use of the advanced functions and algorithms of intelligent control technology, combined with the characteristics and requirements of the automated production line, to achieve fault identification of the production line.
[0003] Currently, there are already various fault identification models for production lines. For example, directly using a logistic regression model, such models use a single model for identification, often with low accuracy, resulting in limited scope of application. Later, fault identification models integrating multiple models emerged, such as the ResNet-BiLSTM model disclosed in CN115392333A, whose accuracy has been significantly improved compared to the former. The current research focus is on how to integrate various models to further improve the accuracy. Summary of the Invention
[0004] In order to solve the above problems, the present invention provides a production line fault identification method based on BiLSTM-Logistic regression by considering the time series data and fault judgment logic on the production line to accurately and efficiently identify the faults of the production line.
[0005] In order to achieve the above object, the solution provided by the present invention is as follows: A production line fault identification method based on BiLSTM-Logistic regression, comprising the following steps: S1. Determine the characteristic parameters representing the operation of the production line, and collect the data of each characteristic parameter at different times under normal working conditions and fault conditions respectively; S2. Construct a sample set and a time series data set. The samples in the sample set are independent data points of feature parameters at a certain moment, and the samples in the time series data set are composed of multiple data points of the samples in the sample set within a continuous time window. The samples in the sample set are aligned with the samples in the time series data set so that the samples in each sample set have the same time series ranking in the corresponding samples of the time series data set. S3. Construct a fault identification model. The fault identification model includes a bidirectional long short-term memory network unit (BiLSTM network unit) and a logistic regression model (Logistic regression model) connected in parallel. The features extracted by the BiLSTM network unit and the Logistic regression model are concatenated together to form the final features, and an SVM classifier is used to identify faults in the final features. S4. Use the sample set and the time series data set to train the fault identification model to obtain the final fault identification model. The samples in the sample set are input into the Logistic regression model, and the samples corresponding to the samples in the sample set in the time series data set are input into the BiLSTM network unit. S5. Extract the feature parameters of the production line to be predicted as the samples to be predicted and input them into the Logistic regression model of the final fault identification model. Take the data points of the samples to be predicted within a continuous time window to construct the samples to be predicted in the time series and input them into the BiLSTM network unit of the final fault identification model, and use the final fault identification model to identify the faults of the production line to be predicted.
[0006] In the present invention, there may be some problems with the data collected in step S1. For example, there may be missing values, the data volatility is relatively large, and the data under normal working conditions is much more than that under fault conditions, resulting in imbalance of various types of data. The data can be preprocessed specifically. For example, traverse all the data and supplement the missing values by linear interpolation method. For example, traverse all the data and delete the samples with missing values. For example, use min-max normalization (Min-Max Scaling / MAX-MIN) to process the data. For example, perform undersampling on the data under normal working conditions to balance various types of data.
[0007] As a specific implementation manner of the present invention, the ratio of the data under fault conditions to the data under normal working conditions in the data of step S1 is 1:2 to 1:3 to prevent overfitting during the deep learning process. During specific operation, undersampling or oversampling can be performed on a certain type of data so that the ratio of the data under fault conditions to the data under normal working conditions falls within the target range.
[0008] In step S2, the size of each sample in the time series data set can be set as needed. For example, take 10 consecutive time series-related data points, or take 21 consecutive time series-related data points, but it is necessary to ensure that the sampling interval time of each sample within the continuous time window is the same.
[0009] As a specific implementation manner of the present invention, in step S3, the BiLSTM network unit includes two BiLSTM networks connected in series. The output of the first BiLSTM network is processed by a Relu activation function and used as the input of the second BiLSTM network. The output of the second BiLSTM network is processed by a Relu activation function, an average pooling layer, and a fully connected layer in sequence and used as the output of the BiLSTM network unit. The Logistic regression model processes the input data and generates a fitted value. The Sigmoid function is used to transform the fitted value generated by the Logistic regression model into [0,1] as the output of the Logistic regression model. The output of the BiLSTM network structure and the output of the Logistic regression model are concatenated together, processed by a Relu activation function, and then an SVM is used as a classifier to output the final classification result.
[0010] Beneficial effects: The present invention constructs a parallel hybrid fault recognition model based on BiLSTM-Logistic regression, and uses the BiLSTM network structure and the logistic regression model to identify faults in the data respectively. The BiLSTM network structure mines the long-term and short-term time series dependencies of the data, and the logistic regression model performs fast and explicit modeling on non-time series features, avoiding excessive smoothing of structured data by deep models. The parallel outputs of the two are fused or cascaded through weighted decision-making, which is more efficient than the simple connection of dual-channel models. It has been verified that the recall effect of this model is excellent and the precision rate is extremely high, realizing the efficient recognition of production line faults. Description of the Drawings
[0011] Figure 1 It is a flowchart of the fault recognition model in the embodiment of the present invention. Specific Embodiment
[0012] The following combines embodiments and drawings to further elaborate on the present invention in detail, but the implementation manners of the present invention are not limited thereto.
[0013] Please refer to Figure 1 , the production line fault recognition method based on BiLSTM-Logistic regression in this embodiment includes the following steps: A production line fault recognition method based on BiLSTM-Logistic regression includes the following steps: S1. Determine the characteristic parameters representing the operation of the production line, and collect the data of each characteristic parameter at different times under normal working conditions and fault conditions respectively; In this embodiment, the production line data comes from the "Teddy Cup" Data Mining Challenge 2024 (the 12th) - Problem A: Automatic Fault Identification of Production Lines. When the production line is running, operation and fault information such as material pushing, container uploading, container positioning, material filling, product capping, product screwing, and product quality inspection are collected in real time.
[0014] By analyzing the data characteristics of each device fault in the operation process of the production line, the main characteristics affecting various device faults are found. The main characteristics of various faults are shown in Tables 1 - 9.
[0015] Table 1 Fault 1001 of Material Pushing Device
[0016] Table 2 Fault 2001 of Material Detection Device
[0017] Table 3 Detection Fault 4001 of Filling Device
[0018] Table 4 Positioning Fault 4002 of Filling Device
[0019] Table 5 Filling Fault 4003 of Filling Device
[0020] Table 6 Positioning Fault 5001 of Capping Device
[0021] Table 7 Capping Fault 5002 of Capping Device
[0022] Table 8 Positioning Fault 6001 of Screwing Device
[0023] Table 9 Screwing Fault 6002 of Screwing Device
[0024] By analyzing the causes of faults and their corresponding data characteristics, it can be seen that the occurrence of each fault is not related to all characteristics, and after screening, it is found that there is no situation where two faults occur simultaneously. Therefore, the characteristic data corresponding to each fault is separately extracted for training, so as to achieve the purpose of feature dimensionality reduction.
[0025] Data preprocessing: After inspection, there is no missing data at each time point in the production line time series data, and the time intervals between sampling points are all fixed; since the number of normal samples is 75,032,405 and the number of faulty samples is 256,692, there is a problem of various data imbalances. Therefore, undersampling is performed on the data under normal working conditions. Finally, in the obtained data, there are 256,692 faulty samples and 513,384 normal samples, and the ratio between the two is 1:2. The MAX-MIN method is used to normalize the above data.
[0026] S2. Construct a sample set and a time series data set. Among them, the samples in the sample set are independent data points of characteristic parameters at a certain moment (t21). The Kafka system is used to extract the data points of the samples in the sample set within a continuous time window (t1 - t21) to construct the time series data set. Therefore, the samples in the time series data set are composed of 21 data points of the samples in the sample set within the continuous time window (t1 - t21), that is, the data points of the samples in the sample set at the previous 20 moments + the samples in the sample set. Each sample in the time series data set is a group of time series data with a fixed length. The samples in the sample set are aligned with the samples in the time series data set, so that the time series rankings of the samples in each sample set in the corresponding samples in the time series data set are the same and at the end.
[0027] Both the sample set and the time series data set are divided into a training set and a validation set according to a ratio of 9:1. Among them, the samples in the time series data set need to be assigned to the same category as the corresponding samples in the sample set. That is, if sample 1 in the sample set belongs to the training set, then the samples constructed by the data points of sample 1 at different moments within the continuous time window in the time series data set should also be assigned to the training set.
[0028] S3. Construct a fault identification model. For the fault identification model, see Figure 1 , which includes a parallel BiLSTM network unit and a Logistic regression model. The features extracted by the BiLSTM network unit and the Logistic regression model are concatenated together to form the final features, and an SVM classifier is used to identify faults for the final features; Specifically, the BiLSTM network unit includes two BiLSTM networks connected in series. Each BiLSTM network is composed of two long short-term memory networks (LSTM). One LSTM processes the input time series data forward, and the other LSTM processes the time series data backward. After the processing is completed, the outputs of the two LSTMs are concatenated as the output of this BiLSTM network. The output of the first BiLSTM network is processed by the Relu activation function and then used as the input of the second BiLSTM network. The output of the second BiLSTM network is processed by the Relu activation function, average pooling layer, and fully connected layer in sequence as the output of the BiLSTM network unit; the loss function of the Logistic regression model is the cross-entropy loss. The Logistic regression model processes the input data and generates a fitted value. The Sigmoid function is used to transform the fitted value generated by the Logistic regression model to [0,1] as the output of the Logistic regression model; the output of the BiLSTM network structure and the output of the Logistic regression model are concatenated together, processed by the Relu activation function, and then the SVM is used as the classifier to output the final classification result.
[0029] S4. Use the sample set and time series data set to train the fault identification model to obtain the final fault identification model. Among them, the samples in the sample set are input into the Logistic regression model, and the samples corresponding to the samples in the sample set in the time series data set are input into the BiLSTM network unit; S5. Extract the characteristic parameters of the production line to be predicted as the sample to be predicted. After normalizing it, input it into the Logistic regression model of the final fault identification model; take the data points of the sample to be predicted at different times to construct the time series sample to be predicted, input it into the BiLSTM network unit of the final fault identification model, and use the final fault identification model to identify the faults in the data of the production line to be predicted.
[0030] In addition, in order to verify the effects of different fault identification models, multiple control groups are set in this embodiment, which are respectively: Control group 1, Logistic regression model; Control group 2, BiLSTM model; Control group 3, ResNet-BiLSTM dual-channel model (from CN115392333A); Control group 4, dual-channel feature fusion model based on attention mechanism (from CN118193935A); The set effect indicators are as follows: Precision: Precision is for the prediction result. It indicates how many of the samples predicted as faults are real faults. Its calculation formula is as follows: ; Wherein, TP (True positive) represents the number of faults correctly detected; FP (False positive) represents the number of faults wrongly detected; Recall rate: The recall rate is for the original samples. It represents how many faults in the samples are correctly identified. In the problem of fault identification, the recall rate is the most important indicator when examining the results. Its calculation formula is as follows: ; Wherein, FN (False negative) represents the number of non-faults wrongly detected; The harmonic mean of the precision rate and the recall rate (F1-score), and its calculation formula is as follows: ; The preprocessed data is respectively used to train and test the fault identification models in this embodiment and the comparative examples. The specific results are shown in Table 10 - Table 14.
[0031] Table 10 Classification results of the model of Comparative Example 1
[0032] Note: support represents the number of samples with actual faults in the test set.
[0033] Table 11 Classification results of the model of Comparative Example 2
[0034] Table 12 Classification results of the model of Comparative Example 3
[0035] Table 13 Classification results of the model of Comparative Example 4
[0036] Table 14 Classification results of the model of the embodiment of the present invention
[0037] As can be seen from Table 10 - 14, the models of the present invention perform significantly better than other models in terms of recall rate and precision rate. In addition, during the edge deployment process, the parallel model of BiLSTM-Logistic regression of the present invention only takes 8.1 ms to learn a time window, the ResNet-BiLSTM dual-channel model of Comparative Example 3 takes 17 ms, and the dual-channel feature fusion model based on the attention mechanism of Comparative Example 4 takes 23 ms.
[0038] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the embodiments of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for identifying production line faults based on BiLSTM-Logistic regression, characterized in that, It includes the following steps: S1. Determine the characteristic parameters representing the operation of the production line, and collect the data of each characteristic parameter at different times under normal working conditions and fault conditions respectively; S2. Construct a sample set and a time series data set. Among them, the samples in the sample set are independent data points of the characteristic parameters at a certain moment, and the samples in the time series data set are composed of multiple data points of the samples in the sample set within a continuous time window; the samples in the sample set are aligned with the samples in the time series data set, so that the samples in each sample set have the same time series ranking in the corresponding samples in the time series data set; S3. Construct a fault identification model. The fault identification model includes a parallel BiLSTM network unit and a Logistic regression model. The features extracted by the BiLSTM network unit and the Logistic regression model are spliced together to form the final feature, and an SVM classifier is used to identify faults for the final feature; S4. Use the sample set and the time series data set to train the fault identification model to obtain the final fault identification model. Among them, the samples in the sample set are input into the Logistic regression model, and the samples corresponding to the samples in the sample set in the time series data set are input into the BiLSTM network unit; S5. Extract the characteristic parameters of the production line to be predicted as the sample to be predicted, and input them into the Logistic regression model of the final fault identification model; take the data points of the sample to be predicted within a continuous time window to construct the time series sample to be predicted, and input it into the BiLSTM network unit of the final fault identification model, and use the final fault identification model to identify the faults of the production line to be predicted.
2. The method for identifying production line faults based on BiLSTM-Logistic regression according to claim 1, characterized in that, In step S1, the ratio of the data of the fault condition to the data of the normal condition is 1:2 to 1:
3.
3. The method for identifying production line faults based on BiLSTM-Logistic regression according to claim 1, characterized in that, In step S3, the BiLSTM network unit includes two BiLSTM networks connected in series. The output of the first BiLSTM network is processed by the Relu activation function and used as the input of the second BiLSTM network. The output of the second BiLSTM network is processed by the Relu activation function, the average pooling layer, and the fully connected layer in sequence and used as the output of the BiLSTM network unit; the Sigmoid function is used to convert the fitting value generated by the Logistic regression model to [0,1] as the output of the Logistic regression model; The output of the BiLSTM network structure and the output of the Logistic regression model are spliced together, processed by the Relu activation function, and then the SVM is used as the classifier to output the final classification result.
4. The method for identifying production line faults based on BiLSTM-Logistic regression according to claim 3, characterized in that, The loss function of the Logistic regression model is the cross-entropy loss.
Citation Information
Patent Citations
Image classification model and method based on improved convolutional neural network and application thereof
CN110969171A
Fault detection method and device based on multi-model fusion, equipment and medium
CN113592019A
Pantograph fault diagnosis method and system, storage medium and equipment
CN113887440A
Thermodynamic data anomaly detection and restoration method based on mechanism and data cooperative driving
CN117574290A
Clustering method and device based on two-dimensional multivariable time sequence feature fusion
CN118035778A
Cited By
Automobile production line abnormity identification and root cause positioning method and system
CN122490378A