An intelligent manufacturing monitoring method and system based on big data

By collecting and processing multi-dimensional data from intelligent manufacturing production lines in real time, combining multi-modal fusion and feature extraction technology, an abnormality detection model is built, which solves the real-time processing and extreme abnormality detection problems in intelligent manufacturing production lines data monitoring, and realizes efficient abnormality detection and early warning.

CN118859800BActive Publication Date: 2025-07-18WUHAN WISTRON WEIZUN SOFTWARE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410913729.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2025-07-18
Estimated Expiration
2044-07-09

AI Technical Summary

Technical Problem

The existing intelligent manufacturing production line data monitoring methods have shortcomings in real-time processing capabilities and extreme outliers processing, resulting in a degradation in model performance, unable to meet real-time processing requirements, and insensitive to extreme outliers.

Method used

Multidimensional monitoring data is collected in real time through the sensor network, cleaned and standardized, and multimodal data fusion is carried out. Time and frequency domain features are extracted using sliding window technology and Fourier transform, real-time abnormality monitoring and extreme abnormality detection models are constructed, combined with random forest and decision tree algorithms for identification and rating, and finally multi-level early warning is performed.

Benefits of technology

It improves the real-time processing capability of intelligent manufacturing production line data and the accuracy of extreme anomaly detection, ensures the robustness and sensitivity of the model, can detect and handle abnormal situations in a timely manner, and provides a multi-level early warning mechanism to deal with abnormalities of varying degrees.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118859800B_ABST
    Figure CN118859800B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of intelligent manufacturing technology, and provides an intelligent manufacturing monitoring method and system based on big data, including: collecting an original multi-dimensional monitoring data set of an intelligent manufacturing production line in real time through a sensor network; performing real-time cleaning, denoising and standardization processing on the original multi-dimensional monitoring data set to obtain a first multi-dimensional monitoring data set; performing multi-modal data fusion on the first multi-dimensional monitoring data set to obtain a second multi-dimensional monitoring data set; using a sliding window technique and a fast Fourier transform on the second multi-dimensional monitoring data set to extract the monitoring time domain features and monitoring frequency domain features of the second multi-dimensional monitoring data set; constructing a real-time anomaly monitoring model and an extreme anomaly detection model, and identifying the monitoring time domain features and monitoring frequency domain features based on the real-time anomaly monitoring model to obtain the anomaly degree rating and extreme anomaly rating of the second multi-dimensional monitoring data set; outputting the anomaly degree rating and extreme anomaly rating in real time, and giving an early warning according to the anomaly degree rating and extreme anomaly rating. Through hierarchical anomaly detection, the intelligent manufacturing production line data is given an early warning rating in real time by comprehensively considering the anomaly degree rating and extreme anomaly rating.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent manufacturing, and in particular, to an intelligent manufacturing monitoring method and system based on big data. Background Art

[0002] With the rapid development of Industry 4.0 and intelligent manufacturing, enterprises have begun to collect and utilize a large amount of data from intelligent devices on the production line on a large scale to achieve predictive maintenance and intelligent management. However, the current intelligent manufacturing big data processing methods face many challenges, especially in the aspects of real-time processing ability and extreme outlier processing, with significant deficiencies. First, the complexity and quality problems of industrial data make it difficult to extract valuable information from a large amount of monitoring data. Second, the existing data processing methods mostly adopt batch processing mode, which is difficult to meet the real-time processing requirements and cannot detect and process production anomalies in a timely manner. Third, in the aspect of extreme outlier processing, the existing methods have problems such as insufficient sensitivity, high false alarm rate, lack of context understanding, and poor adaptability. Finally, there is a coordination dilemma between real-time processing and extreme outlier detection, and few methods can effectively solve these two problems at the same time. These limitations seriously affect the performance of intelligent diagnosis and prediction models and hinder the further optimization and development of intelligent manufacturing systems.

[0003] However, for scenarios that require real-time processing, the models adopted by the existing intelligent manufacturing production line data monitoring methods have delays, which affect the real-time performance, and no special processing method for extreme outliers is given, which will affect the performance of the model on data with extreme outliers. Summary of the Invention

[0004] In view of this, the present invention proposes an intelligent manufacturing monitoring method and system based on big data to solve the problems of low real-time processing ability and insensitivity to extreme outliers in the existing intelligent manufacturing production line data monitoring.

[0005] The technical solution of the present invention is realized as follows: On the one hand, the present invention provides an intelligent manufacturing monitoring method based on big data, including the following steps:

[0006] S1, real-time collect the original multi-dimensional monitoring data set of the intelligent manufacturing production line through a sensor network;

[0007] S2, perform real-time cleaning, denoising and standardization processing on the original multi-dimensional monitoring data set to obtain a first multi-dimensional monitoring data set;

[0008] S3, perform multi-modal data fusion on the first multi-dimensional monitoring data set to obtain a second multi-dimensional monitoring data set;

[0009] S4. Use the sliding window technique and fast Fourier transform on the second multi-dimensional monitoring dataset to extract the monitoring time-domain features and monitoring frequency-domain features of the second multi-dimensional monitoring dataset;

[0010] S5. Construct a real-time anomaly monitoring model and an extreme anomaly detection model. Based on the real-time anomaly monitoring model, identify the monitoring time-domain features and monitoring frequency-domain features to obtain the anomaly degree rating of the second multi-dimensional monitoring dataset. When the anomaly degree rating exceeds the anomaly detection threshold, use the extreme anomaly detection model to identify the monitoring time-domain features and monitoring frequency-domain features to obtain the extreme anomaly rating of the second multi-dimensional monitoring dataset;

[0011] S6. Output the anomaly degree rating and the extreme anomaly rating in real time, and issue a warning according to the anomaly degree rating and the extreme anomaly rating.

[0012] Based on the above technical solutions, preferably, step S2 includes:

[0013] Mark the missing data in the original multi-dimensional monitoring dataset. The missing data is the data for which the sensor network in the original multi-dimensional monitoring dataset did not collect data values;

[0014] According to the missing dataset, complete dataset, and the relationship between the missing dataset and the complete dataset in the original multi-dimensional monitoring dataset, perform data pre-interpolation on the missing data. The calculation formula for the missing data y q is:

[0015]

[0016]

[0017] where α i is the Pearson correlation coefficient between the i-th type of data in the complete dataset and the i-th type of data in the missing dataset, δ(A i , B) is the co-correlation value of the same dimension between the i-th type of data in the complete dataset and the missing dataset, Ci is the mean value of the i-th type of data in the complete dataset, δ(A i , B j ) is the co-correlation value of the j-th dimension between the i-th type of data in the complete dataset and the missing dataset, n is the total number of data categories in the complete dataset and the missing dataset, m is the dimension of the original multi-dimensional monitoring dataset, D qk is the k-th data in the data category where y q is located, and l is the total number of data in the data category where y q is located;

[0018] P is the co-correlation value matrix, δ(A n , B m) is the m - dimensional correlation value between the n - th type of data in the complete dataset and the missing dataset, a n is any data of the n - th type of data in the complete dataset, b m is any dimension of the missing dataset, p(a n ) is the marginal probability distribution of selecting any data of the n - th type of data in the complete dataset, p(b m ) is the marginal probability distribution of selecting any dimension of the missing dataset, p(a n , b m ) is the joint probability distribution of simultaneously selecting any data of the n - th type of data in the complete dataset and any dimension of the missing dataset.

[0019] Based on the above technical solutions, preferably, step S3 includes:

[0020] Perform time alignment and feature selection on the first multi - dimensional monitoring dataset according to the timestamp mark to obtain a multi - source monitoring data feature group {E1, E2,..., E N}, perform multi - modal data fusion on the multi - source monitoring data feature group to obtain a multi - dimensional monitoring fusion dataset, perform normalization processing on the multi - dimensional monitoring fusion dataset to obtain a second multi - dimensional monitoring dataset, and the calculation formula of the second multi - dimensional monitoring dataset is:

[0021]

[0022] where EE is the multi - dimensional monitoring fusion dataset, E1 is the monitoring data feature vector of the first sensor, E2 is the monitoring data feature vector of the second sensor, E N is the monitoring data feature vector of the N - th sensor, ee a,b is the element in the a - th row and b - th column of the second multi - dimensional monitoring dataset, e a,b is the element in the a - th row and b - th column of the multi - dimensional monitoring fusion dataset.

[0023] Based on the above technical solutions, preferably, step S4 includes:

[0024] Perform time - domain analysis on the second multi - dimensional monitoring dataset using the sliding window technique to extract the monitoring time - domain features of the second multi - dimensional monitoring dataset, and the monitoring time - domain features include mean, standard deviation, peak value, root mean square, skewness, kurtosis; perform Fourier transform on the second multi - dimensional monitoring dataset to extract the monitoring frequency - domain features of the second multi - dimensional monitoring dataset, and the monitoring frequency - domain features include spectral energy, main frequency, spectral entropy, spectral centroid, frequency band energy ratio.

[0025] Based on the above technical solutions, preferably, step S5 includes:

[0026] S51. Obtain historical monitoring data, and divide the historical monitoring data into a training set and a test set according to a ratio of 7:3;

[0027] S52. Construct an initial real-time anomaly monitoring model based on a random forest model, perform iterative training on the initial real-time anomaly monitoring model through the training set to obtain a trained initial real-time anomaly monitoring model, evaluate the trained initial real-time anomaly monitoring model through the test set. When the evaluation passes, obtain the real-time anomaly monitoring model; when the evaluation fails, perform iterative training on the model again until the evaluation passes;

[0028] S53. Construct an initial extreme anomaly detection model based on a decision tree algorithm, attach class weights to extreme anomaly samples in the training set to obtain an extreme sample training set, perform iterative training on the initial extreme anomaly detection model through the extreme sample training set to obtain a trained initial extreme anomaly detection model, evaluate the trained initial extreme anomaly detection model through the test set. When the evaluation passes, obtain the extreme anomaly detection model; when the evaluation fails, perform iterative training on the model again until the evaluation passes;

[0029] S54. Identify the monitoring time-domain features and monitoring frequency-domain features based on the real-time anomaly monitoring model to obtain the anomaly degree rating of the second multi-dimensional monitoring data set. When the anomaly degree rating exceeds the anomaly detection threshold, identify the monitoring time-domain features and monitoring frequency-domain features through the extreme anomaly detection model to obtain the extreme anomaly rating of the second multi-dimensional monitoring data set.

[0030] Based on the above technical solutions, preferably, step S52 includes:

[0031] Construct K decision trees based on the training set, and construct an initial real-time anomaly monitoring model based on the random forest classifier RF0 = {T1, T2,..., T K}. The iteration stop conditions for the iterative training include: the number of decision trees reaches a preset value, and the first loss function converges;

[0032] The first loss function L is:

[0033]

[0034] where M is the number of samples, x h is the anomaly degree of the h-th sample, and p h is the probability of predicting the h-th sample as normal;

[0035] Evaluate the trained initial real-time anomaly monitoring model through the test set, calculate the mAP value, and when the mAP value is not less than the preset value, the evaluation passes, and the real-time anomaly monitoring model is obtained:

[0036]

[0037] Among them, YC u (X s , X p ) is the probability that the u-th decision tree predicts that the dataset X is abnormal. X s is the time-domain feature of the dataset X. X p is the frequency-domain feature of the dataset X. Z(r) is the outlier of the r-th class data of the dataset X. R is the number of data classes of the dataset X. YC(X) is the real-time anomaly monitoring probability of the dataset X.

[0038] Based on the above technical solutions, preferably, step S53 includes:

[0039] Construct a decision tree weak classifier based on the decision tree algorithm. An initial extreme anomaly detection model is composed of the decision tree weak classifier. Attach class weights to the extreme anomaly samples in the training set to obtain an extreme sample training set. Train the decision tree weak classifier through the extreme sample training set until the weight of the decision tree weak classifier approaches zero to obtain a decision tree final classifier. Construct a trained initial extreme anomaly detection model based on the decision tree final classifier. Evaluate the trained initial extreme anomaly detection model through the test set, using the accuracy as the evaluation index of the model. When the accuracy of the model is not less than the preset value, the evaluation passes to obtain an extreme anomaly detection model. The calculation formula is:

[0040]

[0041] Among them, J(c) is the decision tree final classifier, sign(·) is the sign function, and j g (c) is the g-th decision tree weak classifier, and λ g is the weight of the g-th decision tree weak classifier. G is the number of decision tree weak classifiers. P(vc) is the conditional probability. is the optimization function. P(v = 1c) is the probability that the sample c is not an extreme anomaly value. exp(·) is the natural exponential function.

[0042] Based on the above technical solutions, preferably, step S54 includes:

[0043] Identify the monitored time-domain feature and monitored frequency-domain feature based on the real-time anomaly monitoring model to obtain the anomaly degree rating Q of the second multi-dimensional monitoring dataset. When Q exceeds the anomaly detection threshold Q y , identify the monitored time-domain feature and monitored frequency-domain feature through the extreme anomaly detection model to obtain the extreme anomaly rating V of the second multi-dimensional monitoring dataset:

[0044]

[0045] Among them, X1 = {X 1s , X 1p} is the monitoring time-domain feature and monitoring frequency-domain feature of the second multi-dimensional monitoring data set, P(v = 1c d ) is the probability that the d-th sample of X1 is not an extreme outlier, and D is the number of samples of X1;

[0046] Based on the extreme anomaly rating V, the anomaly detection threshold Q is updated in real time y .

[0047] On the basis of the above technical solutions, preferably, step S6 includes:

[0048] Give an early warning according to the anomaly degree rating and the extreme anomaly rating:

[0049]

[0050] Among them, W is the early warning rating, Q is the anomaly degree rating, V is the extreme anomaly rating, D1 is the first anomaly degree rating threshold, D2 is the second anomaly degree rating threshold, Q y is the anomaly detection threshold, I1 is the first extreme anomaly rating threshold, I2 is the second extreme anomaly rating threshold, is the first early warning rating, is the second early warning rating, is the third early warning rating, is the fourth early warning rating, is the fifth early warning rating, is the sixth early warning rating;

[0051] Give an early warning to the intelligent manufacturing supervision personnel based on different early warning ratings.

[0052] On the other hand, the present invention also provides an intelligent manufacturing monitoring system based on big data, and the system includes:

[0053] A data acquisition module for real-time collecting the original multi-dimensional monitoring data set of the intelligent manufacturing production line through a sensor network;

[0054] A preprocessing module for performing real-time cleaning, denoising, and standardization processing on the original multi-dimensional monitoring data set to obtain a first multi-dimensional monitoring data set;

[0055] A data fusion module for performing multi-modal data fusion on the first multi-dimensional monitoring data set to obtain a second multi-dimensional monitoring data set;

[0056] A feature extraction module for using the sliding window technology and the fast Fourier transform on the second multi-dimensional monitoring data set to extract the monitoring time-domain feature and monitoring frequency-domain feature of the second multi-dimensional monitoring data set;

[0057] The model construction and feature recognition module is used to construct a real-time anomaly monitoring model and an extreme anomaly detection model, identify the monitoring time-domain features and monitoring frequency-domain features based on the real-time anomaly monitoring model, obtain the anomaly degree rating of the second multi-dimensional monitoring dataset, and when the anomaly degree rating exceeds the anomaly detection threshold, identify the monitoring time-domain features and monitoring frequency-domain features through the extreme anomaly detection model to obtain the extreme anomaly rating of the second multi-dimensional monitoring dataset;

[0058] The evaluation and warning module is used to output the anomaly degree rating and the extreme anomaly rating in real time, and issue a warning based on the anomaly degree rating and the extreme anomaly rating.

[0059] The intelligent manufacturing monitoring method and system based on big data of the present invention has the following beneficial effects compared with the prior art:

[0060] (1) Through hierarchical anomaly detection, a lightweight real-time anomaly monitoring model is used to preliminarily identify the monitoring time-domain features and monitoring frequency-domain features of the intelligent manufacturing production line data. According to the identified anomaly degree rating, it is judged whether there is an extreme anomaly in the input intelligent manufacturing production line data. When an extreme anomaly is detected in the monitoring, the monitoring time-domain features and monitoring frequency-domain features of the intelligent manufacturing production line data are further identified through the extreme anomaly detection model, so as to judge the extreme anomaly rating of the intelligent manufacturing production line data. Finally, the anomaly degree rating and the extreme anomaly rating are comprehensively used to perform a warning rating on the intelligent manufacturing production line data in real time, and different corresponding solutions are adopted for different levels of warning ratings;

[0061] (2) By marking and pre-imputing the missing data in the original multi-dimensional monitoring dataset, the data missing problem is effectively solved. Considering the missing dataset, the complete dataset and the relationship between the two, the integrity and reliability of the data are improved, providing a more accurate data basis for subsequent analysis;

[0062] (3) Through time alignment and feature selection of the first multi-dimensional monitoring dataset, a multi-source monitoring data feature group is obtained, and then multi-modal data fusion and normalization processing are performed, ensuring the synchronization of different sensor data in the time dimension, improving the consistency and comparability of the data, and reducing the data volume and noise impact through feature selection and fusion, and improving the model calculation efficiency;

[0063] (4) Based on the random forest model, a real-time anomaly monitoring model is constructed. The integration method using multiple decision trees reduces the bias of a single decision tree, improves the robustness and generalization ability of the model, and ensures the efficiency and effect of model training by setting iteration stop conditions;

[0064] (5) An extreme anomaly detection model is constructed based on the decision tree algorithm, and class weights are attached to the extreme anomaly samples in the training set, improving the recognition ability for extreme anomaly situations. The accuracy rate is used as the evaluation index to ensure the high performance of the model in the extreme anomaly detection task. BRIEF DESCRIPTION OF THE DRAWINGS

[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0066] Figure 1 It is a flowchart of the intelligent manufacturing monitoring method based on big data of the present invention;

[0067] Figure 2 It is a structure diagram of the intelligent manufacturing monitoring system based on big data of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0068] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0069] Please refer to Figure 1 , the present invention provides an intelligent manufacturing monitoring method based on big data, including the following steps:

[0070] S1. Real-time collect the original multi-dimensional monitoring data set of the intelligent manufacturing production line through the sensor network;

[0071] S2. Perform real-time cleaning, denoising and standardization processing on the original multi-dimensional monitoring data set to obtain the first multi-dimensional monitoring data set;

[0072] S3. Perform multi-modal data fusion on the first multi-dimensional monitoring data set to obtain the second multi-dimensional monitoring data set;

[0073] S4. Use the sliding window technology and the fast Fourier transform on the second multi-dimensional monitoring data set to extract the monitoring time domain features and monitoring frequency domain features of the second multi-dimensional monitoring data set;

[0074] S5. Construct a real-time anomaly monitoring model and an extreme anomaly detection model. Based on the real-time anomaly monitoring model, identify the monitored time-domain features and monitored frequency-domain features to obtain the anomaly degree rating of the second multi-dimensional monitoring dataset. When the anomaly degree rating exceeds the anomaly detection threshold, use the extreme anomaly detection model to identify the monitored time-domain features and monitored frequency-domain features to obtain the extreme anomaly rating of the second multi-dimensional monitoring dataset;

[0075] S6. Output the anomaly degree rating and the extreme anomaly rating in real time, and issue a warning based on the anomaly degree rating and the extreme anomaly rating.

[0076] Specifically, in the intelligent manufacturing environment, extreme outliers may usually occur in the case of sudden equipment failures or sensor malfunctions. The intelligent manufacturing monitoring method based on big data in this embodiment performs hierarchical anomaly detection. First, a lightweight real-time anomaly monitoring model is used to preliminarily identify the monitored time-domain features and monitored frequency-domain features of the intelligent manufacturing production line data. According to the identified anomaly degree rating (actually the anomaly probability), it is determined whether there is an extreme anomaly in the input intelligent manufacturing production line data. When an extreme anomaly is detected, the extreme anomaly detection model is further used to identify the monitored time-domain features and monitored frequency-domain features of the intelligent manufacturing production line data, so as to judge the extreme anomaly rating of the intelligent manufacturing production line data. Finally, the anomaly degree rating and the extreme anomaly rating are comprehensively used to perform a real-time warning rating on the intelligent manufacturing production line data, and different corresponding solutions are adopted for different levels of warning ratings.

[0077] Specifically, the original multi-dimensional monitoring dataset of the intelligent manufacturing production line is collected in real time through a sensor network. There are various different data in different dimensions in the intelligent manufacturing production line, and these data together represent the comprehensive status of the intelligent manufacturing production line;

[0078] In a specific embodiment, the original multi-dimensional monitoring dataset includes: high-frequency sampling data, such as vibration, noise, current, and rotational speed; continuous time-series data, such as temperature, humidity, air pressure, and light intensity; multi-dimensional data, such as production speed, production quantity, and product quality.

[0079] Step S2 includes:

[0080] Mark the missing data in the original multi-dimensional monitoring dataset. The missing data is the data for which the sensor network in the original multi-dimensional monitoring dataset fails to collect data values;

[0081] According to the missing dataset, complete dataset in the original multi-dimensional monitoring dataset, and the relationship between the missing dataset and the complete dataset, perform data pre-interpolation on the missing data. The missing data y q The calculation formula is:

[0082]

[0083]

[0084] where α i is the Pearson correlation coefficient between the data of the i-th class in the complete data set and the data of the i-th class in the missing data set, and δ(A i , B) is the co-correlation value of the same dimension between the data of the i-th class in the complete data set and the missing data set, and C i is the mean value of the data of the i-th class in the complete data set, and δ(A i , B j ) is the co-correlation value of the j-th dimension between the data of the i-th class in the complete data set and the missing data set, n is the total number of data categories in the complete data set and the missing data set, m is the dimension of the original multi-dimensional monitoring data set, D qk is the k-th data in the data category where y q is located, and l is the total number of data in the data category where y q is located;

[0085] P is the co-correlation value matrix, and δ(A n , B m ) is the co-correlation value of the m-th dimension between the data of the n-th class in the complete data set and the missing data set, a n is any data of the n-th class in the complete data set, b m is any dimension of the missing data set, p(a n ) is the marginal probability distribution of selecting any data of the n-th class in the complete data set, p(b m ) is the marginal probability distribution of selecting any dimension of the missing data set, and p(a n , b m ) is the joint probability distribution of simultaneously selecting any data of the n-th class in the complete data set and any dimension of the missing data set.

[0086] Specifically, step S2 significantly improves the integrity of the data by marking and pre-imputing the missing data in the original multi-dimensional monitoring data set. The existence of missing data will lead to a decrease in the accuracy of subsequent data analysis and model training. Through the pre-imputation technology, these data gaps can be effectively filled, thus ensuring the continuity and integrity of the data set.

[0087] The pre-imputation of missing data not only fills the data gaps, but also removes the noise and outliers in the data through cleaning and denoising processing.

[0088] After the cleaning is completed, there is also a standardization process for the data, enabling data in different dimensions to be compared and analyzed on the same scale. The standardization process eliminates the dimensional differences between different data dimensions, making data fusion and feature extraction more effective.

[0089] The missing data calculation formula in this embodiment takes into account the reference effect of the existing complete data set on the missing data set, and uses the marginal probability distribution of a single variable and the joint probability distribution of multiple variables to introduce the correlation effect between multiple variables, thereby improving the accuracy of data pre-imputation.

[0090] Missing data y q The calculation formula estimates the missing value by calculating the marginal probability distribution of the missing data set and the joint probability distribution of the complete data set and the missing data set. This method can capture the complex dependence relationships between data, thereby improving the accuracy and reliability of the imputed data.

[0091] In a specific embodiment, p(a n ) can represent the marginal probability distribution of temperature, used to reflect the overall distribution of temperature, and p(b m ) can represent the marginal probability distributions of temperatures in different dimensions, such as temperatures in different regions and temperatures at different times, used to reflect the overall distributions of temperatures in different dimensions. At this time, p(a n , b m ) represents the joint probability distribution of temperature and temperature dimension, and can predict the probability that the temperature value and the temperature dimension are both specific values. If the missing data y q is the temperature data during the daytime in the production workshop area, then D qk is the k-th data during the daytime in the production workshop area. To calculate the mean of this category of data, first assign this mean to the missing data y q , and then introduce a bias correction value, that is, in the formula. This bias correction value takes into account the mean of the complete data set corresponding to this missing data, the Pearson correlation coefficient between the complete data set and the missing data set, and the multi-dimensional mutual correlation value between the complete data set and the missing data set. The mean of the complete data set plays a reference role, and the Pearson correlation coefficient is used to quantify the linear correlation degree between the missing data set and the complete data set. In fact, the missing data can already be pre-imputed through the mean of the complete data set and the Pearson correlation coefficient between the complete data set and the missing data set. However, in order to reduce the influence of the missing data dimension, consider the weights of the data corresponding to the missing data in the missing data set and the complete data set in all dimensions, that is, in the formula, thus obtaining the calculation formula for the missing data in this embodiment.

[0092] Step S3 includes:

[0093] Align the first multi-dimensional monitoring data set in time according to the time stamp and perform feature selection to obtain a multi-source monitoring data feature group {E1, E2,..., E N}, perform multi-modal data fusion on the multi-source monitoring data feature group to obtain a multi-dimensional monitoring fusion data set, and perform normalization processing on the multi-dimensional monitoring fusion data set to obtain a second multi-dimensional monitoring data set. The calculation formula of the second multi-dimensional monitoring data set is:

[0094]

[0095] where EE is the multi-dimensional monitoring fusion data set, E1 is the monitoring data feature vector of the first sensor, E2 is the monitoring data feature vector of the second sensor, and E N is the monitoring data feature vector of the Nth sensor, ee a,b is the element in the ath row and bth column of the second multi-dimensional monitoring data set, and e a,b is the element in the ath row and bth column of the multi-dimensional monitoring fusion data set.

[0096] Specifically, step S3 ensures the synchronization of data from different sensors in the time dimension in the first multi-dimensional monitoring data set through time alignment, improving the consistency and comparability of the data;

[0097] Then, the most relevant and informative features are selected through feature selection to reduce the amount of data, reduce the influence of noise, and improve the calculation efficiency of the model;

[0098] Furthermore, integrate the data of different sensors into a multi-source monitoring data feature group {E1, E2,..., E N}, where N is the number of sensors, and then perform feature fusion on the features of each sensor through multi-modal data fusion to obtain a multi-dimensional monitoring fusion data set. Considering that the data types and scales of each sensor are different, normalization processing is also used to unify the scales of different features and reduce the influence of dimensions.

[0099] In this embodiment, N 1×H vectors are fused into a complete H×H matrix through matrix operations, which is convenient for subsequent data processing.

[0100] Step S4 includes:

[0101] Perform time-domain analysis on the second multi-dimensional monitoring data set using the sliding window technique to extract the monitoring time-domain features of the second multi-dimensional monitoring data set. The monitoring time-domain features include mean, standard deviation, peak value, root mean square, skewness, and kurtosis; perform Fourier transform on the second multi-dimensional monitoring data set to extract the monitoring frequency-domain features of the second multi-dimensional monitoring data set. The monitoring frequency-domain features include spectral energy, main frequency, spectral entropy, spectral centroid, and frequency band energy ratio.

[0102] Specifically, step S4 performs time-domain analysis on the second multi-dimensional monitoring dataset through the sliding window technique, thereby extracting the time-domain features of the second multi-dimensional monitoring dataset, including mean, standard deviation, peak value, root mean square, skewness, and kurtosis. The sliding window technique has the characteristic of dynamic monitoring, which can perform dynamic real-time monitoring on data, capture the time-varying trend of data at any time, and can capture the local features of data in different time periods, improving the sensitivity of anomaly detection. Through the smoothing process of the data within the window, the influence of noise on feature extraction can be reduced.

[0103] After time-domain analysis, the second multi-dimensional monitoring dataset is subjected to frequency-domain conversion through Fourier transform, converting the time-domain data to the frequency domain, and extracting the frequency-domain features of the second multi-dimensional monitoring dataset, including spectral energy, main frequency, spectral entropy, spectral centroid, and frequency band energy ratio, so as to analyze the periodicity and frequency components of the data in the second multi-dimensional monitoring dataset. It can also filter out high-frequency noise and extract useful frequency components. Frequency-domain features can analyze data at different frequency scales and provide richer information.

[0104] In this embodiment, through the sliding window technique and Fourier transform, the time-domain and frequency-domain features of the second multi-dimensional monitoring dataset are extracted. These features provide a multi-dimensional description of the data, improving the comprehensiveness and accuracy of data analysis, and providing a basis for subsequent real-time anomaly monitoring and extreme anomaly detection. Through the dynamic monitoring and analysis of the sliding window technique, abnormal changes can be captured in a timely manner, improving the response speed and sensitivity of the system, thereby enhancing the overall performance and reliability of the intelligent manufacturing monitoring method.

[0105] Step S5 includes:

[0106] S51, obtaining historical monitoring data, and dividing the historical monitoring data into a training set and a test set according to 7:3;

[0107] S52, constructing an initial real-time anomaly monitoring model based on the random forest model, iteratively training the initial real-time anomaly monitoring model through the training set to obtain the trained initial real-time anomaly monitoring model, evaluating the trained initial real-time anomaly monitoring model through the test set. When the evaluation passes, a real-time anomaly monitoring model is obtained. When the evaluation fails, the model is iteratively trained again until the evaluation passes;

[0108] S53, constructing an initial extreme anomaly detection model based on the decision tree algorithm, attaching class weights to the extreme anomaly samples in the training set to obtain an extreme sample training set, iteratively training the initial extreme anomaly detection model through the extreme sample training set to obtain the trained initial extreme anomaly detection model, evaluating the trained initial extreme anomaly detection model through the test set. When the evaluation passes, an extreme anomaly detection model is obtained. When the evaluation fails, the model is iteratively trained again until the evaluation passes;

[0109] S54. Based on the real-time anomaly monitoring model, identify the monitored time-domain features and monitored frequency-domain features to obtain the anomaly degree rating of the second multi-dimensional monitoring data set. When the anomaly degree rating exceeds the anomaly detection threshold, identify the monitored time-domain features and monitored frequency-domain features through the extreme anomaly detection model to obtain the extreme anomaly rating of the second multi-dimensional monitoring data set.

[0110] Specifically, for step S51, obtain the historical monitoring data with the confirmed anomaly degree from the historical database. The various attribute indicators of these data are determined and can be fully applied to the training and testing of the model.

[0111] By dividing the training set and the test set, overfitting of the model can be effectively avoided, and the performance of the model on unseen data can be improved.

[0112] The ratio of 7:3 can achieve a balance between training and testing, which can not only ensure the training effect of the model but also effectively evaluate the performance of the model.

[0113] In this embodiment, by dividing the historical monitoring data into a training set and a test set according to the ratio of 7:3, the rationality and effectiveness of model training and evaluation are ensured, the generalization ability and reliability of the model are improved, the performance of the model in actual applications is ensured. At the same time, the reasonable data division ratio also ensures the full utilization and representativeness of the data.

[0114] Step S52 includes:

[0115] Construct K decision trees based on the training set, and construct an initial real-time anomaly monitoring model based on the random forest classifier RF0 = {T1, T2,..., T K}. The iteration stop conditions for the iterative training include: The iteration stop conditions for the iterative training include: the number of decision trees reaches the preset value, and the first loss function converges;

[0116] The first loss function L is:

[0117]

[0118] where M is the number of samples, x h is the anomaly degree of the h-th sample, and p h is the probability of predicting the h-th sample as normal;

[0119] Evaluate the trained initial real-time anomaly monitoring model through the test set, calculate the mAP value. When the mAP value is not less than the preset value, the evaluation passes, and a real-time anomaly monitoring model is obtained:

[0120]

[0121] Among them, YC u (X s , X p ) is the probability that the u-th decision tree predicts that the dataset X is abnormal. X s is the time-domain feature of the dataset X, and X p is the frequency-domain feature of the dataset X. Z(r) is the outlier of the r-th type of data in the dataset X, R is the number of data categories in the dataset X, and YC(X) is the real-time anomaly monitoring probability of the dataset X.

[0122] Specifically, in step S52, by constructing K decision trees, a random forest classifier is formed. The random forest improves the stability and accuracy of the model by integrating the results of multiple decision trees. The decision tree is constructed based on the training set and updated in real time and dynamically. Each training will modify the decision tree based on the previous training. Integrating multiple decision trees reduces the bias of a single decision tree, thereby improving the robustness and generalization ability of the model.

[0123] The first loss function L is improved based on the cross-entropy loss function and is used to measure the gap between the model prediction and the actual label. The model parameters are optimized by minimizing this loss function.

[0124] When the change rate of the loss function value for several consecutive iterations of the model is less than the preset threshold, it is determined that the loss function converges at this time, that is, the iterative training is stopped to obtain the trained model. If the number of decision trees reaches the preset value before the loss function converges, the iterative training will also be stopped in advance.

[0125] The initial real-time anomaly monitoring model in step S52 is constructed based on the random forest model. The initial real-time anomaly monitoring model is trained through the training set, mainly by inputting the time-domain feature and frequency-domain feature in the training set, that is, X s and X p in the formula, so as to perform the ability training of the anomaly prediction of the model.

[0126] It should be noted that YC(X) is the real-time anomaly monitoring probability of the dataset X. In this embodiment, by using the excellent classification ability of the random forest model, the anomaly detection is converted into a binary classification problem, so as to perform binary classification on the dataset, that is, the classification of normal and abnormal, and the size of the probability represents the degree of anomaly. Therefore, a lightweight random forest classifier is used to realize the preliminary real-time anomaly detection of the data.

[0127] In a specific embodiment, the decision tree is initially set to 100, and 5 are added for each iterative training. The preset value is 1000, that is, the model will stop iterative training after at most 1800 rounds of training.

[0128] It should be noted that the mAP value is used to evaluate the model, and its calculation method is as follows:

[0129] premax(w) = max(precision(w, w + 1));

[0130]

[0131] Among them, mAP is the evaluation index, AP is the average precision, W is the number of samples in the training set, w is the current sample, U(w) is the recall rate of the current training set sample, U(w + 1) is the recall rate of the next training set sample, premax(w) is the maximum value of the accuracy between the current training set sample and the next training set sample, and precision(w, w + 1) is the accuracy of the current training set sample and the accuracy of the next training set sample.

[0132] The following parameters are specified:

[0133] TP: The number of samples correctly classified as positive samples; actually positive samples, and also classified as positive samples by the model;

[0134] FP: The number of samples misclassified as positive samples; actually negative samples, but classified as positive samples by the model;

[0135] TN: The number of samples correctly classified as negative samples; actually negative samples, and also classified as negative samples by the model;

[0136] FN: The number of samples misclassified as negative samples; actually positive samples, but classified as negative samples by the model;

[0137] Accuracy Precision = TP / (TP + FP), Recall = TP / (TP + FN);

[0138] Taking the Precision and Recall of each type as the horizontal and vertical axes respectively, finding the AP value is actually equivalent to finding the area enclosed by Precision and Recall. Since the composed curve region jitters very severely, interpolation is used here to smooth the curve region. If the recall value of the current sample is U(w), the premax used for interpolation is the maximum value of Precision between the recall of the next sample being U(w + 1). Where w represents the current sample. The mAP value is the average of the AP values of all types, and W here represents the number of samples in the training set.

[0139] Step S53 includes:

[0140] Construct a decision tree weak classifier based on the decision tree algorithm. An initial extreme anomaly detection model is composed of decision tree weak classifiers. Attach class weights to the extreme anomaly samples in the training set to obtain an extreme sample training set. Train the decision tree weak classifier with the extreme sample training set until the weight of the decision tree weak classifier approaches zero, obtaining the final decision tree classifier. Construct a trained initial extreme anomaly detection model based on the final decision tree classifier. Evaluate the trained initial extreme anomaly detection model with a test set, using the accuracy as the evaluation metric for the model. When the accuracy of the model is not less than the preset value, the evaluation passes, obtaining the extreme anomaly detection model. The calculation formula is as follows:

[0141]

[0142] Among them, J(c) is the final decision tree classifier, sign(·) is the sign function, and j g (c) is the g-th decision tree weak classifier, and λ g is the weight of the g-th decision tree weak classifier, G is the number of decision tree weak classifiers, P(v|c) is the conditional probability, is the optimization function, P(v = 1|c) is the probability that sample c is not an extreme anomaly value, and exp(·) is the natural exponential function.

[0143] Specifically, in step S53, a decision tree weak classifier is constructed based on the decision tree algorithm to construct an initial extreme anomaly detection model. The decision tree algorithm is a supervised learning method that forms a tree-like structure by splitting the data set and can handle classification problems well. A weak classifier refers to a classifier with relatively weak classification ability, but a strong classifier can be formed by integrating multiple weak classifiers.

[0144] Considering that the extreme anomaly detection model is used to identify extreme anomaly values in the data, in this embodiment, class weights are attached to the extreme anomaly samples in the training set, which can make the model pay more attention to these extreme anomaly samples, thereby improving the model's detection ability for extreme anomalies. Attaching class weights is to increase the influence of extreme anomaly samples during the training process, enabling the model to better learn the characteristics of these samples during training, similar to the effect of introducing an attention mechanism.

[0145] In this embodiment, the decision tree weak classifier is trained. The decision tree weak classifier is trained with the extreme sample training set until the weight of the decision tree weak classifier approaches zero, which can ensure that each weak classifier has fully learned the characteristics of the extreme anomaly samples. The training process is an iterative process, and the weight of the weak classifier is adjusted in each iteration until the weight approaches zero, indicating that the weak classifier has fully learned the sample characteristics.

[0146] After training enough weak classifiers, an initial extreme anomaly detection model is constructed based on the decision tree final classifier, which can form a strong classifier to effectively detect extreme anomaly samples. The final classifier is composed of multiple weak classifiers integrated together. By integrating the results of multiple weak classifiers, the stability and accuracy of the model can be improved.

[0147] After training the decision tree final classifier, an initial extreme anomaly detection model is constructed based on the decision tree final classifier. Then, the trained initial extreme anomaly detection model is evaluated using the test set, with the accuracy rate as the evaluation metric for the model. When the accuracy rate of the model is not less than the preset value, the evaluation passes, and the extreme anomaly detection model is obtained. The evaluation process is to verify the performance of the model to ensure that the model can effectively detect extreme anomaly samples in practical applications. Using the accuracy rate as the evaluation metric can intuitively reflect the detection ability of the model.

[0148] In this embodiment, by constructing multiple decision tree weak classifiers j g (c), and then integrating them to obtain the decision tree final classifier, the final classification result is obtained by weighted summing the results of all weak classifiers. Then, the sign of the classification result is confirmed based on the sign function to classify the data as abnormal. The conditional probability P(vc) is used to determine whether a sample is an outlier given the sample.

[0149] Step S54 includes:

[0150] Based on the real-time anomaly monitoring model, the monitoring time-domain features and monitoring frequency-domain features are identified to obtain the anomaly degree rating Q of the second multi-dimensional monitoring data set. When Q exceeds the anomaly detection threshold Q y , the monitoring time-domain features and monitoring frequency-domain features are identified by the extreme anomaly detection model to obtain the extreme anomaly rating V of the second multi-dimensional monitoring data set:

[0151]

[0152] where X1 = {X 1s , X 1p} are the monitoring time-domain features and monitoring frequency-domain features of the second multi-dimensional monitoring data set, P(v = 1|c d ) is the probability that the d-th sample of X1 is not an extreme outlier, and D is the number of samples of X1;

[0153] Based on the extreme anomaly rating V, the anomaly detection threshold Q y is updated in real time.

[0154] Specifically, in step S54, the trained real-time anomaly monitoring model and extreme anomaly detection model are deployed to the intelligent manufacturing production line. An embedded system can be used as the carrier, and the intelligent manufacturing production line data is used as the input for actual application.

[0155] The real-time anomaly monitoring model identifies the monitoring time-domain features and monitoring frequency-domain features, and obtains the real-time anomaly monitoring probability of the second multi-dimensional monitoring data set, which can quickly detect abnormal situations in the data set. The real-time anomaly monitoring model uses time-domain and frequency-domain features to identify anomalies, and these features can capture mutations and abnormal patterns in the data, thus enabling rapid response.

[0156] In a specific embodiment, the real-time anomaly monitoring probability output by the real-time anomaly monitoring model is in the range of [0, 1]. The greater the probability, the higher the degree of anomaly. When there is only one set of time-domain features and frequency-domain features in the second multi-dimensional monitoring data set, the anomaly degree rating Q at this time is the real-time anomaly monitoring probability. When there are multiple sets of time-domain features and frequency-domain features in the second multi-dimensional monitoring data set, the mean value of the real-time anomaly monitoring probabilities of multiple sets of features needs to be calculated as the anomaly degree rating.

[0157] In a specific embodiment, the predicted anomaly detection threshold Q y is 0.8. When the real-time anomaly monitoring probability of the second multi-dimensional monitoring data set exceeds 0.8, the second multi-dimensional monitoring data set is further subjected to extreme anomaly detection. The second multi-dimensional monitoring data set is identified by the extreme anomaly detection model to obtain the probability that the data in the second multi-dimensional monitoring data set is not an extreme anomaly value. It should be noted that the extreme anomaly detection model identifies each sample data in the second multi-dimensional monitoring data set to obtain the probability that each sample is an extreme anomaly value.

[0158] In a specific embodiment, the probability that each sample is an extreme anomaly value is in the range of [0, 1]. The greater the probability, the more the number of samples of extreme anomaly values. The extreme anomaly detection model needs to detect each sample in the data set one by one, so the calculation amount is large, but the accuracy will be higher. Finally, the extreme anomaly rating obtained is the mean value of the probabilities of each sample, and this mean value is used as the number of samples of extreme anomaly values.

[0159] By combining the real-time anomaly monitoring model and the extreme anomaly detection model, the accuracy and timeliness of anomaly detection can be improved, ensuring the timely discovery and handling of abnormal situations. The real-time anomaly monitoring model can quickly detect abnormal situations, while the extreme anomaly detection model can further identify extreme abnormal situations. The combination of the two can improve the accuracy and timeliness of anomaly detection, ensuring the timely discovery and handling of abnormal situations.

[0160] Step S6 includes:

[0161] Give an early warning according to the abnormal degree rating and the extreme abnormal rating:

[0162]

[0163] where W y is the early warning rating, Q is the abnormal degree rating, V is the extreme abnormal rating, D1 is the first abnormal degree rating threshold, D2 is the second abnormal degree rating threshold, Q y is the abnormal detection threshold, I1 is the first extreme abnormal rating threshold, I2 is the second extreme abnormal rating threshold, is the first early warning rating, is the second early warning rating, is the third early warning rating, is the fourth early warning rating, is the fifth early warning rating, is the sixth early warning rating;

[0164] Give an early warning to the intelligent manufacturing supervisors based on different early warning ratings.

[0165] Specifically, step S6 comprehensively evaluates the abnormal degree rating Q and the extreme abnormal rating V to comprehensively understand the abnormal conditions in the data set. This comprehensive evaluation method can more accurately reflect the abnormal state of the data set.

[0166] According to different abnormal degree ratings and extreme abnormal ratings, step S6 sets up a multi-level early warning mechanism. Specifically, it includes six early warning ratings (the first early warning rating, the second early warning rating, the third early warning rating, the fourth early warning rating, the fifth early warning rating, the sixth early warning rating), and each early warning rating corresponds to different abnormal and extreme abnormal threshold intervals (D1, D2, Q y , I1, I2). This multi-level early warning mechanism can give a graded early warning according to the severity of the abnormality and provide more detailed early warning information.

[0167] In a specific embodiment, D1 can be set to 0.3, D2 can be set to 0.6, Q y can be set to 0.8, I1 can be set to 0.4, and I2 can be set to 0.7.

[0168] By combining the real-time abnormal monitoring model and the extreme abnormal detection model, step S6 can improve the accuracy and timeliness of abnormal detection. The real-time abnormal monitoring model is used to quickly detect abnormal conditions, while the extreme abnormal detection model can further identify extreme abnormal conditions. The combination of the two can more accurately and timely discover and handle abnormal conditions.

[0169] Based on different warning ratings, step S6 can issue warnings to intelligent manufacturing supervisors. This warning mechanism can help supervisors promptly understand and handle abnormal situations, ensuring the stable operation of the intelligent manufacturing system.

[0170] In a specific embodiment, different levels of response measures are required for different warning ratings:

[0171] First warning rating - Slight abnormality (no extreme outliers, relatively low abnormality rating):

[0172] Record the abnormal situation, notify the relevant departments for routine inspections, and increase the monitoring frequency of relevant parameters.

[0173] Second warning rating - Mild abnormality (no extreme outliers, medium abnormality rating):

[0174] Execute all measures at the first level, conduct a preliminary fault diagnosis, prepare spare equipment or parts that may be needed, and formulate a preliminary response plan.

[0175] Third warning rating - Moderate abnormality (no extreme outliers, relatively high abnormality rating):

[0176] Execute all measures at the previous two levels, start a preventive maintenance program, adjust the production plan, reduce the load on the affected equipment, and notify technical experts for in-depth analysis.

[0177] Fourth warning rating - Severe abnormality (there are extreme outliers, and the number of extreme outliers is small):

[0178] Execute all measures at the previous three levels, start an emergency response program, consider temporarily stopping the operation of the affected equipment, mobilize additional manpower and resources for handling, evaluate the impact on the overall production, and formulate a response strategy.

[0179] Fifth warning rating - High abnormality (there are extreme outliers, and the number of extreme outliers is medium):

[0180] Execute all measures at the previous four levels, immediately stop the operation of the affected equipment, evacuate non-essential personnel, start the backup system or emergency plan, notify senior management, and prepare a public relations statement.

[0181] Sixth warning rating - Extreme abnormality (there are extreme outliers, and the number of extreme outliers is large):

[0182] Implement all measures of the first five levels, immediately stop the operation of the entire production line or factory, activate the comprehensive emergency plan, hold an emergency crisis management meeting, consider seeking assistance from external experts, evaluate the long-term impact and formulate a recovery plan, and prepare a detailed accident report and public statement.

[0183] Please refer to Figure 2 , based on the above method embodiments, the present invention further provides an intelligent manufacturing monitoring system based on big data, and the system includes:

[0184] A data acquisition module, configured to collect the original multi-dimensional monitoring data set of the intelligent manufacturing production line in real time through a sensor network;

[0185] A preprocessing module, configured to perform real-time cleaning, denoising, and standardization processing on the original multi-dimensional monitoring data set to obtain a first multi-dimensional monitoring data set;

[0186] A data fusion module, configured to perform multi-modal data fusion on the first multi-dimensional monitoring data set to obtain a second multi-dimensional monitoring data set;

[0187] A feature extraction module, configured to use a sliding window technique and a fast Fourier transform on the second multi-dimensional monitoring data set to extract the monitoring time-domain features and monitoring frequency-domain features of the second multi-dimensional monitoring data set;

[0188] A model construction and feature recognition module, configured to construct a real-time anomaly monitoring model and an extreme anomaly detection model, identify the monitoring time-domain features and monitoring frequency-domain features based on the real-time anomaly monitoring model to obtain the anomaly degree rating of the second multi-dimensional monitoring data set, and when the anomaly degree rating exceeds the anomaly detection threshold, identify the monitoring time-domain features and monitoring frequency-domain features through the extreme anomaly detection model to obtain the extreme anomaly rating of the second multi-dimensional monitoring data set;

[0189] An evaluation and warning module, configured to output the anomaly degree rating and the extreme anomaly rating in real time, and issue a warning according to the anomaly degree rating and the extreme anomaly rating.

[0190] The above system embodiments and method embodiments correspond one by one. For the brief description of the system embodiments, please refer to the method embodiments.

[0191] The above is only the preferred embodiment of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An intelligent manufacturing monitoring method based on big data, characterized in that, Including the following steps: S1. Real-time collect the original multi-dimensional monitoring data set of the intelligent manufacturing production line through the sensor network; S2. Perform real-time cleaning, denoising and standardization processing on the original multi-dimensional monitoring data set to obtain the first multi-dimensional monitoring data set; S3. Perform multi-modal data fusion on the first multi-dimensional monitoring data set to obtain the second multi-dimensional monitoring data set; S4. Use the sliding window technique and fast Fourier transform on the second multi-dimensional monitoring data set to extract the monitoring time-domain features and monitoring frequency-domain features of the second multi-dimensional monitoring data set; S5. Construct a real-time anomaly monitoring model and an extreme anomaly detection model, identify the monitoring time-domain features and monitoring frequency-domain features based on the real-time anomaly monitoring model to obtain the anomaly degree rating of the second multi-dimensional monitoring data set. When the anomaly degree rating exceeds the anomaly detection threshold, identify the monitoring time-domain features and monitoring frequency-domain features through the extreme anomaly detection model to obtain the extreme anomaly rating of the second multi-dimensional monitoring data set; Step S5 includes: S51. Obtain historical monitoring data, and divide the historical monitoring data into a training set and a test set according to 7:3; S52. Construct an initial real-time anomaly monitoring model based on the random forest model, iteratively train the initial real-time anomaly monitoring model through the training set to obtain the trained initial real-time anomaly monitoring model, evaluate the trained initial real-time anomaly monitoring model through the test set. When the evaluation passes, obtain the real-time anomaly monitoring model. When the evaluation fails, iteratively train the model again until the evaluation passes; S53. Construct an initial extreme anomaly detection model based on the decision tree algorithm, attach class weights to the extreme anomaly samples in the training set to obtain an extreme sample training set, iteratively train the initial extreme anomaly detection model through the extreme sample training set to obtain the trained initial extreme anomaly detection model, evaluate the trained initial extreme anomaly detection model through the test set. When the evaluation passes, obtain the extreme anomaly detection model. When the evaluation fails, iteratively train the model again until the evaluation passes; S54. Identify the monitoring time-domain features and monitoring frequency-domain features based on the real-time anomaly monitoring model to obtain the anomaly degree rating of the second multi-dimensional monitoring data set. When the anomaly degree rating exceeds the anomaly detection threshold, identify the monitoring time-domain features and monitoring frequency-domain features through the extreme anomaly detection model to obtain the extreme anomaly rating of the second multi-dimensional monitoring data set; Step S54 includes: Based on the real-time anomaly monitoring model, identify the monitored time-domain features and monitored frequency-domain features to obtain the anomaly degree rating of the second multi-dimensional monitoring data set Q , when Q exceeds the anomaly detection threshold Q y , identify the monitored time-domain features and monitored frequency-domain features through the extreme anomaly detection model to obtain the extreme anomaly rating of the second multi-dimensional monitoring data set V : ; ; Among them, are the monitoring time domain features and monitoring frequency domain features of the second multi-dimensional monitoring data set, is X the probability that the d th sample of 1 is not an extreme outlier, D is X the number of samples of 1; Based on the extreme anomaly rating V Update the anomaly detection threshold in real time Q y ; S6. Real-time output the anomaly degree rating and the extreme anomaly rating, and give an early warning according to the anomaly degree rating and the extreme anomaly rating.

2. The intelligent manufacturing monitoring method based on big data according to claim 1, wherein Step S2 includes: Mark the missing data in the original multi-dimensional monitoring data set, where the missing data is the data for which the sensor network in the original multi-dimensional monitoring data set did not collect data values; Based on the missing data set, complete data set in the original multi-dimensional monitoring data set, and the relationship between the missing data set and the complete data set, perform data pre-imputation on the missing data, the missing data The calculation formula is: ; ; ; Among them, is the Pearson correlation coefficient between the i -th class data of the complete data set and the i -th class data of the missing data set, is the co - correlation value of the same dimension between the i -th class data of the complete data set and the missing data set, C i is the mean value of the i -th class data of the complete data set, is the i -th class data of the complete data set and the j co - correlation value of the n dimensions between the complete data set and the missing data set, m is the total number of data categories of the complete data set and the missing data set, is the k -th data in the data category where l is located, is the total number of data in the data category where P is the mutual correlation value matrix, is the n -dimensional mutual correlation value of the m th class data of the complete data set and the is the n th class data of the complete data set, is any one dimension of the missing data set, is the marginal probability distribution of any one data of the n th class data selected from the complete data set, is the marginal probability distribution of any one dimension selected from the missing data set, is the joint probability distribution of any one data of the n th class data selected from the complete data set and any one dimension of the missing data set.

3. The intelligent manufacturing monitoring method based on big data according to claim 2, wherein Step S3 includes: Time-align and feature-select the first multi-dimensional monitoring data set according to the timestamp markings to obtain a multi-source monitoring data feature group , perform multi-modal data fusion on the multi-source monitoring data feature group to obtain a multi-dimensional monitoring fusion data set, perform normalization processing on the multi-dimensional monitoring fusion data set to obtain a second multi-dimensional monitoring data set, and the calculation formula for the second multi-dimensional monitoring data set is: ; ; Among them, EE is a multi-dimensional monitoring fusion data set, E 1 is the monitoring data feature vector of the first sensor, E 2 is the monitoring data feature vector of the second sensor, E N is the N monitoring data feature vector of the is the a th row and b th column element of the second multi-dimensional monitoring data set, is the a th row and b th column element of the multi-dimensional monitoring fusion data set.

4. The intelligent manufacturing monitoring method based on big data according to claim 3, characterized in that, Step S4 includes: Perform time-domain analysis on the second multi-dimensional monitoring dataset using the sliding window technique to extract the monitoring time-domain features of the second multi-dimensional monitoring dataset, where the monitoring time-domain features include mean, standard deviation, peak value, root mean square, skewness, and kurtosis; perform Fourier transform on the second multi-dimensional monitoring dataset to extract the monitoring frequency-domain features of the second multi-dimensional monitoring dataset, where the monitoring frequency-domain features include spectral energy, main frequency, spectral entropy, spectral centroid, and frequency band energy ratio.

5. The intelligent manufacturing monitoring method based on big data according to claim 4, wherein, Step S52 includes: Constructed based on the training set K decision trees, and based on a random forest classifier Construct an initial real-time anomaly monitoring model. The iteration stop conditions for the iterative training include: the number of decision trees reaches a preset value, and the first loss function converges; The first loss function L is as follows: ; Among them, M is the number of samples, is the outlier degree of the h th sample, is the probability of predicting that the h th sample is normal; Evaluate the trained initial real-time anomaly monitoring model using the test set, calculate the mAP value, and if the mAP value is not less than the preset value, the evaluation passes to obtain the real-time anomaly monitoring model: ; ; ; Among them, is the u th decision tree prediction data set X is the probability of being abnormal, X s is the time-domain feature of the data set X ; X p is the frequency-domain feature of the data set X ; is the outlier of the X th r class data of the data set R is the number of data categories of the data set X ; YC ( X ) is the real-time anomaly monitoring probability of the data set X .

6. The intelligent manufacturing monitoring method based on big data according to claim 5, wherein Step S53 includes: Construct a decision tree weak classifier based on the decision tree algorithm, form an initial extreme anomaly detection model through the decision tree weak classifier, attach class weights to the extreme anomaly samples in the training set to obtain an extreme sample training set, train the decision tree weak classifier through the extreme sample training set until the weight of the decision tree weak classifier approaches zero to obtain the final decision tree classifier, construct the trained initial extreme anomaly detection model based on the final decision tree classifier, evaluate the trained initial extreme anomaly detection model using the test set, use the accuracy rate as the evaluation index of the model, and if the accuracy rate of the model is not less than the preset value, the evaluation passes to obtain the extreme anomaly detection model, and the calculation formula is: ; ; ; Among them, J ( c ) is the final classifier of the decision tree, is the sign function, j g ( c ) is the g-th weak classifier of the decision tree, is the g weight of the g-th weak classifier of the decision tree, G is the number of weak classifiers of the decision tree, is the conditional probability, is the optimization function, is the probability that the sample c is not an extreme outlier, is the natural exponential function.

7. The intelligent manufacturing monitoring method based on big data according to claim 6, wherein Step S6 includes: Give an early warning according to the anomaly degree rating and the extreme anomaly rating: ; Among them, W is the warning rating, Q is the abnormal degree rating, V is the extreme abnormal rating, D 1 is the first abnormal degree rating threshold, D 2 is the second abnormal degree rating threshold, Q y is the abnormal detection threshold, I 1 is the first extreme abnormal rating threshold, I 2 is the second extreme abnormal rating threshold, is the first warning rating, is the second warning rating, is the third warning rating, is the fourth warning rating, is the fifth warning rating, is the sixth warning rating; Give an early warning to the intelligent manufacturing supervision personnel based on different early warning ratings.

8. An intelligent manufacturing monitoring system based on big data, which is used to implement an intelligent manufacturing monitoring method based on big data as described in any one of claims 1-7, characterized in that, The system includes: A data acquisition module for collecting the original multi-dimensional monitoring dataset of the intelligent manufacturing production line in real time through a sensor network; A preprocessing module for performing real-time cleaning, denoising, and standardization processing on the original multi-dimensional monitoring dataset to obtain a first multi-dimensional monitoring dataset; A data fusion module for performing multi-modal data fusion on the first multi-dimensional monitoring dataset to obtain a second multi-dimensional monitoring dataset; A feature extraction module for using the sliding window technique and the fast Fourier transform on the second multi-dimensional monitoring dataset to extract the monitoring time-domain features and monitoring frequency-domain features of the second multi-dimensional monitoring dataset; A model construction and feature recognition module for constructing a real-time anomaly monitoring model and an extreme anomaly detection model, identifying the monitoring time-domain features and monitoring frequency-domain features based on the real-time anomaly monitoring model to obtain the anomaly degree rating of the second multi-dimensional monitoring dataset, and when the anomaly degree rating exceeds the anomaly detection threshold, identifying the monitoring time-domain features and monitoring frequency-domain features through the extreme anomaly detection model to obtain the extreme anomaly rating of the second multi-dimensional monitoring dataset; An evaluation and early warning module for outputting the anomaly degree rating and the extreme anomaly rating in real time, and giving an early warning according to the anomaly degree rating and the extreme anomaly rating.

Citation Information

Patent Citations

  • Method for diagnosing and interpolating abnormal water regimen data based on RF-Adaboost model

    CN114493023A

  • Intelligent manufacturing data real-time monitoring method and system based on Internet of Things technology

    CN117309042A

  • Operation performance anomaly detection method and system for fan unit

    CN118030409A

  • Abnormal data detection method for mechanical state monitoring

    CN118094421A