Industrial production Internet of Things data anomaly detection method, medium and system
By using principal component analysis, virtual sensors, and pseudo-statistical mechanics models to reduce the dimensionality of high-dimensional sensor data, and combining dynamic thresholds and adaptive parameter optimization, the problem of high computational complexity in traditional methods is solved, achieving efficient and accurate anomaly detection.
Patent Information
- Application Number
- CN202511323890.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-17
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-09-17
AI Technical Summary
Traditional anomaly detection algorithms have high computational complexity in high-dimensional sensor data and cannot effectively process multi-dimensional sensor data of industrial equipment, resulting in poor detection results.
Principal component analysis and mutual information algorithm are used for dimensionality reduction. A virtual sensor algorithm and a pseudo-statistical mechanics anomaly analysis model are constructed. Combined with dynamic threshold adjustment and parameter adaptive update algorithm, data transmission is optimized through virtual sensors and network topology. Statistical mechanics theory is used to map high-dimensional data to low-dimensional features and establish an anomaly pattern classification system.
It achieves efficient anomaly detection, reduces computational complexity, improves detection accuracy and robustness, adapts to different working conditions and time changes, and reduces false alarm rate and false negative rate.
Smart Images

Figure CN120822158A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of industrial production Internet of Things, and specifically relates to a method, medium and system for detecting anomaly in industrial production Internet of Things data. Background Art
[0002] Device anomaly detection in the Industrial Internet of Things (IIoT) environment is a key technology for ensuring industrial production safety and improving equipment operational efficiency. Traditional anomaly detection methods primarily rely on techniques such as statistical analysis, machine learning, and pattern recognition. These methods establish a baseline for normal equipment operation to identify abnormal behavior that deviates from normal conditions. These methods have been widely used in industrial scenarios such as power system monitoring, petrochemical plant safety protection, and manufacturing equipment maintenance. However, traditional anomaly detection techniques suffer from significant drawbacks when processing high-dimensional sensor data. As the dimensionality of sensor data increases, data points become sparsely distributed in the high-dimensional space, rendering traditional distance metrics ineffective. Furthermore, the computational complexity of high-dimensional data increases exponentially with the number of dimensions, making it difficult for traditional algorithms to complete anomaly detection tasks within a reasonable timeframe. Due to the high dimensionality, complex correlations, and large data volumes of multidimensional sensor data generated by industrial equipment, traditional anomaly detection algorithms are unable to effectively transform the massive amount of high-dimensional data into manageable low-dimensional feature representations. The lack of effective dimensionality compression and information aggregation mechanisms results in significant computational resource consumption and poor detection results. In other words, there is a technical problem in the existing technology that the sensor data generated by industrial equipment is extremely high in dimension, and traditional algorithms are prone to the curse of dimensionality in high-dimensional space, resulting in an exponential increase in computational complexity and the inability to achieve effective anomaly detection. Summary of the Invention
[0003] In view of this, the present invention provides a method, medium and system for detecting anomalies in industrial production Internet of Things data, which can solve the technical problems in the prior art that the sensor data generated by industrial equipment is extremely high in dimension, and traditional algorithms are prone to the curse of dimensionality in high-dimensional space, resulting in an exponential increase in computational complexity and the inability to achieve effective anomaly detection.
[0004] The present invention is implemented as follows: In a first aspect, the present invention provides an industrial production Internet of Things data anomaly detection method, comprising collecting sensor data generated by industrial equipment and preprocessing the data to establish a multidimensional sensor data set, performing dimensionality reduction processing on the multidimensional sensor data set using principal component analysis and performing feature selection using a mutual information algorithm to establish a dimensionality reduction feature data set, constructing a virtual sensor algorithm to generate virtual sensor data from the dimensionality reduction feature data set to form an extended sensor data set, establishing an anomaly analysis model based on statistical mechanics laws to map the extended sensor data set to a statistical mechanics space to calculate a statistical mechanics feature vector, analyzing the statistical mechanics feature vector using an industrial equipment status assessment model, calculating an anomaly detection threshold through a dynamic threshold adjustment function to identify an abnormal mode of equipment operation, establishing an abnormal mode classification system to classify and mark the abnormal mode of equipment operation, calculating an anomaly level through an abnormality severity assessment function to generate an anomaly detection report, constructing a feedback optimization mechanism to adjust detection parameters according to the anomaly detection report, optimizing the performance of the industrial equipment status assessment model through a parameter adaptive update algorithm, and realizing continuous improvement of the anomaly detection system.
[0005] Among them, the virtual sensor algorithm is specifically a calculation method based on existing sensor data to infer missing or faulty sensor data. By establishing a correlation model between sensors and using the data of normal sensors to infer the data value of the virtual sensor, the data missing problem in the sensor network is compensated.
[0006] Among them, the virtual sensor algorithm adopts a virtual collection router model, which is specifically a data path optimization model based on graph theory and network topology analysis. By constructing a connection relationship graph between sensor nodes and calculating the shortest path weights between nodes, it identifies weak links and data transmission bottlenecks in the network from the perspective of the sensor network topology structure.
[0007] Among them, when the virtual collection router model detects a sensor node failure or data transmission interruption, it automatically reconstructs the data transmission path and generates virtual sensor node data based on the redundant paths and backup node information of the network topology structure, and makes up for the missing data through the neighboring node data interpolation and time series data extrapolation algorithm.
[0008] Among them, the simulated statistical mechanics anomaly analysis model is specifically an equipment state analysis model constructed based on statistical mechanics theory and thermodynamic equilibrium principles. It maps the multi-dimensional operating parameters of industrial equipment to the microscopic particle state distribution and macroscopic thermodynamic parameters in the statistical mechanics space, and simulates the equipment operation state by establishing a virtual particle ensemble and thermodynamic system.
[0009] Among them, the statistical mechanics anomaly analysis model maps the random fluctuations of equipment operating parameters into Brownian motion and thermal fluctuations of virtual particles, maps the changes in equipment energy consumption into the internal energy and enthalpy changes of the virtual thermodynamic system, and maps the equipment operation stability into the entropy value and free energy gradient of the virtual system.
[0010] Among them, the statistical mechanics anomaly analysis model analyzes the regularity of equipment operation data through the Boltzmann distribution law and Maxwell statistical distribution, calculates the partition function, correlation function and fluctuation dissipation relationship of the virtual particle ensemble, and identifies abnormal patterns by detecting the entropy increase anomaly, phase change phenomenon and statistical distribution deviation of the virtual thermodynamic system.
[0011] Among them, the statistical mechanics characteristic vector is specifically a set of vectors that characterize the operating state of the equipment and is calculated by simulating the statistical mechanics anomaly analysis model, which includes the microscopic state distribution characteristics, macroscopic thermodynamic equilibrium state and statistical fluctuation law during the operation of the equipment.
[0012] The abnormal operation mode of the equipment is specifically a device behavior mode that deviates from the normal operation state and is identified by the dynamic threshold adjustment function, and includes information on the time of occurrence of the abnormality, the type of abnormality, and the degree of abnormality.
[0013] Among them, the abnormal severity assessment function is used to quantify the severity of the abnormal situation. The input includes the abnormal deviation degree, duration, impact range, and historical abnormal frequency. The output is the abnormal level value, which divides the equipment operation abnormal mode into four levels: minor, moderate, severe, and emergency.
[0014] The abnormality detection report is a comprehensive report document containing abnormality level values and detailed information on abnormal equipment operation modes, which is used to guide equipment maintenance and fault handling decisions.
[0015] Among them, the parameter adaptive update algorithm is specifically an optimization algorithm that automatically adjusts model parameters according to the detection effect, receives anomaly detection reports as input, monitors the accuracy and recall rate of anomaly detection, and automatically adjusts model weights and threshold parameters when the detection performance deteriorates.
[0016] The industrial equipment status assessment model is a computational model that comprehensively assesses the operating status of industrial equipment. It takes statistical mechanics eigenvectors as input, calculates the overall health status of the equipment through a composite assessment algorithm, and provides a status benchmark for anomaly detection. The industrial equipment status assessment model is optimized using a state fusion network based on an attention mechanism. The state fusion network dynamically assigns attention weights based on the importance and relevance of sensor data. The calculation of attention weights is determined based on three key parameters: sensor data dimension, data acquisition frequency, and sensor reliability coefficient. The structure of the state fusion network is a multi-layer attention encoder architecture, which includes a data embedding layer, a multi-head self-attention layer, a feedforward neural network layer, and an output mapping layer. The entire network uses residual connections and layer normalization technology to improve training stability and convergence speed.
[0017] The dynamic threshold adjustment function is specifically a function that dynamically adjusts the anomaly detection threshold according to the equipment operating environment and historical data, receives the output results of the industrial equipment status assessment model, and automatically adjusts the detection sensitivity to reduce false positives and missed positives.
[0018] A second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions are run in a computer, they are used to execute the above-mentioned method for detecting anomaly in industrial production Internet of Things data.
[0019] The third aspect of the present invention provides an industrial production Internet of Things data anomaly detection system, which includes the above-mentioned computer-readable storage medium. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.
[0020] The present invention constructs a statistical mechanics-like anomaly analysis model to convert the microscopic state distribution of massive high-dimensional sensor data into a few macroscopic thermodynamic parameters, achieving significant data dimensionality compression and effectively addressing the curse of dimensionality problem of high-dimensional sensor data. Drawing on the successful experience of statistical physics in processing large numbers of microscopic particle systems, the statistical mechanics-like model treats each sensor data point as the microscopic state of a virtual particle. Using statistical mechanics theory, the complex high-dimensional data associations are converted into a few macroscopic thermodynamic parameters such as temperature, entropy, internal energy, and free energy. This conversion process inherently possesses the ability to compress dimensions, capable of compressing thousands of dimensions of sensor data into a dozen or so thermodynamic parameters while preserving the data's core statistical characteristics and anomaly information. By establishing a virtual particle ensemble and thermodynamic system to simulate the device's operating state, and utilizing the Boltzmann distribution law and Maxwell statistical distribution to analyze the regularity of the device's operating data, anomaly detection is transformed into detecting entropy increase anomalies, phase transition phenomena, and statistical distribution shifts in the virtual thermodynamic system. This dimensionality reduction method, based on physical principles, has a sound theoretical foundation and mathematical interpretability. In summary, the present invention solves the technical problem mentioned in the background technology that the sensor data generated by industrial equipment is extremely high in dimension, and traditional algorithms are prone to the curse of dimensionality in high-dimensional space, resulting in an exponential increase in computational complexity and the inability to achieve effective anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 is a flow chart of the method of the present invention.
[0022] Figure 2 This is a schematic diagram of the network structure of the industrial equipment status assessment model involved in the present invention.
[0023] Figure 3 This is a topology diagram of the virtual sensor network in Example 2.
[0024] Figure 4 This is a comparison diagram of the statistical mechanics characteristic vector distribution in Example 2.
[0025] Figure 5 This is a timing analysis diagram of the abnormality detection results in Example 2. DETAILED DESCRIPTION
[0026] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0027] like Figure 1 FIG. 1 is a flowchart of a method for detecting anomaly in industrial production Internet of Things data provided by the first aspect of the present invention. The method comprises the following steps: S01. Collect sensor data generated by industrial equipment and pre-process the collected sensor data, including data cleaning and preliminary filtering of outliers, to establish a multi-dimensional sensor data set; S02, performing dimensionality reduction processing on the multidimensional sensor data set by principal component analysis, extracting main components by eigenvalue decomposition, and performing feature selection by using mutual information algorithm to establish a dimensionality reduction feature data set; S03, constructing a virtual sensor algorithm to generate virtual sensor data for the dimension-reduced feature dataset, and compensating for data loss from the perspective of the sensor network topology through a virtual acquisition router model to form an extended sensor dataset; S04. Establishing a statistical mechanics anomaly analysis model based on the laws of statistical mechanics, mapping the extended sensor data set into a statistical mechanics space, performing regularity analysis on the equipment operation data through microscopic state statistical distribution and macroscopic thermodynamic parameters, and calculating a statistical mechanics eigenvector; S05. Analyze the mechanical characteristic vector using an industrial equipment status assessment model, calculate an anomaly detection threshold using a dynamic threshold adjustment function, and identify abnormal equipment operation modes; S06. Establish an abnormal pattern classification system, classify and mark the detected abnormal patterns of equipment operation, calculate the abnormality level through the abnormality severity evaluation function, and generate an abnormality detection report; S07. Construct a feedback optimization mechanism, adjust detection parameters according to the anomaly detection report, optimize the performance of the industrial equipment status assessment model through a parameter adaptive update algorithm, and achieve continuous improvement of the anomaly detection system.
[0028] The virtual sensor algorithm is a calculation method that infers missing or faulty sensor data based on existing sensor data. By establishing an association model between sensors and using the data of normal sensors to infer the data value of the virtual sensor, the data missing problem in the sensor network is compensated.
[0029] The virtual collection router model is a data path optimization model based on graph theory and network topology analysis. By constructing a connection relationship graph between sensor nodes and calculating the shortest path weights between nodes, it identifies weak links and data transmission bottlenecks in the sensor network topology. When a sensor node failure or data transmission interruption is detected, the virtual collection router model automatically reconstructs the data transmission path and generates virtual sensor node data based on the redundant paths and backup node information of the network topology. It compensates for missing data through neighboring node data interpolation and time series data extrapolation algorithms, and dynamically adjusts network load balancing and data transmission priority to ensure the reliable transmission and integrity of critical sensor data.
[0030] The simulated statistical mechanics anomaly analysis model is an equipment state analysis model constructed based on statistical mechanics theory and thermodynamic equilibrium principles. It maps the multidimensional operating parameters of industrial equipment to the microscopic particle state distribution and macroscopic thermodynamic parameters in the statistical mechanics space, and simulates the equipment operation state by establishing a virtual particle ensemble and a thermodynamic system. The simulated statistical mechanics anomaly analysis model maps the random fluctuations of the equipment operation parameters to the Brownian motion and thermal fluctuations of the virtual particles, maps the changes in equipment energy consumption to the internal energy and enthalpy changes of the virtual thermodynamic system, and maps the equipment operation stability to the entropy value and free energy gradient of the virtual system. The equipment operation data is analyzed for regularity through the Boltzmann distribution law and Maxwell statistical distribution, and the partition function, correlation function and fluctuation-dissipation relationship of the virtual particle ensemble are calculated. When the equipment operation deviates from the normal operating conditions, the simulated statistical mechanics anomaly analysis model identifies the abnormal pattern by detecting the entropy increase anomaly, phase change phenomenon and statistical distribution offset of the virtual thermodynamic system, and establishes a mapping relationship between statistical mechanics parameters and the equipment health state.
[0031] The statistical mechanics characteristic vector is a set of vectors representing the operating state of the equipment obtained by calculation through the simulated statistical mechanics anomaly analysis model, which includes the microscopic state distribution characteristics, macroscopic thermodynamic equilibrium state and statistical fluctuation law during the operation of the equipment.
[0032] The industrial equipment status assessment model is a computational model for comprehensively assessing the operating status of industrial equipment. It takes the statistical mechanics eigenvector as input and calculates the overall health status of the equipment through a composite assessment algorithm, providing a status benchmark for anomaly detection.
[0033] The dynamic threshold adjustment function is a function that dynamically adjusts the anomaly detection threshold according to the equipment operating environment and historical data, receives the output results of the industrial equipment status assessment model, and automatically adjusts the detection sensitivity to reduce false positives and missed negatives.
[0034] The device operation abnormality pattern is a device behavior pattern that deviates from a normal operating state and is identified by the dynamic threshold adjustment function, and includes information on the abnormality occurrence time, abnormality type, and abnormality degree.
[0035] The abnormal severity assessment function is used to quantify the severity of the abnormal situation. The input includes the abnormal deviation degree, duration, impact range, and historical abnormal frequency. The output is the abnormal level value, which divides the equipment operation abnormal mode into four levels: minor, moderate, severe, and emergency.
[0036] The anomaly detection report is a comprehensive report document containing the anomaly level value and detailed information on the abnormal operation mode of the equipment, and is used to guide equipment maintenance and fault handling decisions.
[0037] The parameter adaptive update algorithm is an optimization algorithm that automatically adjusts model parameters according to the detection effect. It receives the anomaly detection report as input, monitors the accuracy and recall rate of anomaly detection, and automatically adjusts the model weights and threshold parameters when the detection performance decreases to ensure the long-term stability and accuracy of the detection system.
[0038] The industrial equipment status assessment model is optimized using a state fusion network based on the attention mechanism. The state fusion network dynamically allocates attention weights according to the importance and relevance of sensor data. The calculation of attention weights needs to be determined based on three key parameters: sensor data dimension, data acquisition frequency, and sensor reliability coefficient. The sensor data dimension affects the number of attention heads, the data acquisition frequency determines the length of the time window, and the sensor reliability coefficient adjusts the weight allocation ratio of each sensor data.
[0039] The structure of the state fusion network is a multi-layer attention encoder architecture, which includes a data embedding layer, a multi-head self-attention layer, a feedforward neural network layer and an output mapping layer. The data embedding layer converts the mechanical feature vector into a unified vector representation. The multi-head self-attention layer captures the complex correlation between sensor data by parallel calculation of multiple attention heads. The feedforward neural network layer performs nonlinear transformation and feature extraction on the attention output. The output mapping layer maps high-dimensional features into equipment state evaluation results. The entire network uses residual connection and layer normalization technology to improve training stability and convergence speed.
[0040] The steps for establishing a training data set for the state fusion network include collecting equipment operation data in different industrial scenarios, covering sensor data records of normal operating states, various abnormal states and failure modes, performing quality assessment and noise filtering on the collected raw data, eliminating data samples with obvious errors and serious omissions, classifying and labeling the data according to equipment type and operating conditions, establishing a supervised learning data set containing equipment state labels, dividing the data set into training set, validation set and test set according to time series, ensuring that data in different time periods are evenly distributed, normalizing and performing data enhancement operations on the training data, generating more diverse training samples and improving the model generalization ability.
[0041] The state fusion network training steps include initializing network parameters and optimizer configuration, setting learning rate scheduling strategy and regularization parameters, using batch gradient descent algorithm for model training, calculating model output and loss function value through forward propagation, using backpropagation algorithm to calculate parameter gradients and update network weights, regularly evaluating model performance on the validation set during training, monitoring the changing trends of training loss and validation accuracy, adopting early stopping strategy to avoid overfitting when validation performance no longer improves, using learning rate decay and gradient clipping technology to improve training stability, and performing final performance evaluation on the test set after training to ensure that the model's generalization ability on unseen data meets the expected requirements.
[0042] The state weight adjustment function is used to adjust the attention weight distribution of the state fusion network. It is calculated based on multiple data such as sensor data variance, data correlation coefficient, anomaly detection historical accuracy, and current device load rate to obtain an attention adjustment coefficient value. When the attention adjustment coefficient value belongs to different ranges, it is used to adjust the attention weight parameter of the state fusion network using different weight distribution strategies, including four adjustment ranges. The first interval value is calculated by taking the square root of the product of the sensor data variance and the data correlation coefficient. The second interval value is calculated by taking the geometric mean of the anomaly detection historical accuracy and the current device load rate. The third interval value is calculated by taking the geometric mean of the sensor data variance and the data correlation coefficient. According to the harmonic mean of the variance and the historical accuracy of the anomaly detection, when the attention adjustment coefficient value is less than the first interval value, a conservative weight distribution strategy is adopted to reduce the dynamic adjustment amplitude of the attention weight; when the attention adjustment coefficient value is between the first interval value and the second interval value, a balanced weight distribution strategy is adopted to maintain a moderate adjustment intensity of the attention weight; when the attention adjustment coefficient value is between the second interval value and the third interval value, an active weight distribution strategy is adopted to enhance the response sensitivity of the attention weight; when the attention adjustment coefficient value is greater than the third interval value, an aggressive weight distribution strategy is adopted to maximize the adaptive adjustment ability of the attention weight.
[0043] The specific implementation of the above steps is described in detail below.
[0044] Step S01 is implemented by acquiring raw data from various sensors on industrial equipment through a distributed data acquisition system. These sensors include devices that monitor various physical quantities, such as temperature, pressure, vibration, current, and speed. The data acquisition process utilizes a time synchronization mechanism to ensure temporal consistency across sensor data. The acquisition frequency is set to an appropriate value between 1Hz and 1000Hz, depending on the device type. The preprocessing phase begins with data cleaning, using a sliding window smoothing filter algorithm to remove high-frequency noise. The window length is set between 5 and 15 sampling points. Initial outlier filtering utilizes a combination of the 3σ criterion based on statistical distribution and the interquartile range method. Data points that deviate from the mean by more than three standard deviations or more than 1.5 times the interquartile range are marked as outliers. The multidimensional sensor dataset is constructed through data alignment and interpolation to ensure the integrity and consistency of the data across all dimensions along the time axis. This step aims to provide high-quality basic data for subsequent analysis. Preprocessing improves data quality and reduces the impact of noise on anomaly detection accuracy.
[0045] The implementation of step S02 is to perform dimensionality reduction and feature selection on the multidimensional sensor data set. The principal component analysis dimensionality reduction process uses the singular value decomposition algorithm to calculate the eigenvalues and eigenvectors of the covariance matrix, and retains the principal components with a cumulative contribution rate of 85% to 95% as the features after dimensionality reduction. The eigenvalue decomposition process is implemented by the Jacobi iterative algorithm or the QR decomposition algorithm, and the iterative convergence threshold is set to When the mutual information algorithm performs feature selection, it calculates the mutual information value between each feature dimension and the target variable, uses the maximum information coefficient method to quantify nonlinear correlations, and retains the feature variables with the top 70% to 90% mutual information values. The dimension of the dimensionality-reduced feature dataset is usually reduced from dozens of dimensions of the original data to 10 to 20 dimensions, which not only preserves the key information of the data but also reduces computational complexity. The purpose of this step is to eliminate data redundancy and reduce the curse of dimensionality, thereby improving the computational efficiency and model generalization ability of subsequent algorithms.
[0046] Step S03 involves constructing a virtual sensor algorithm to generate data from missing or faulty sensors. This virtual sensor algorithm uses machine learning methods such as multivariate linear regression, radial basis function networks, or support vector regression to establish a correlation model between sensors. The training process uses historical normal operation data and optimizes model parameters using the least squares method or gradient descent algorithm. The model training error threshold is set to a root mean square error of less than 5%. The virtual collection router model constructs a sensor network topology based on graph theory algorithms, with nodes representing sensor locations and edge weights representing the strength of correlation between sensors. The shortest path calculation uses the Dijkstra or Floyd-Warshall algorithms, with the path weight threshold set to sensor pairs with a correlation coefficient greater than 0.6. When a sensor failure is detected, the model automatically selects a backup data source based on redundant network paths and generates virtual sensor data using linear interpolation, spline interpolation, or time series extrapolation algorithms. The extended sensor dataset is formed by fusing the original data with the virtual data, achieving data integrity exceeding 95%. This step aims to address data loss caused by sensor failure or communication interruptions and ensure the continuity and integrity of the input data for the anomaly detection algorithm.
[0047] The implementation method of step S04 is to establish a statistical mechanics anomaly analysis model to map the sensor data into the statistical mechanics space for analysis. The model regards the equipment operating parameters as a virtual particle system. The random fluctuations of the parameters correspond to the Brownian motion of the particles, and the changes in the equipment energy consumption correspond to the changes in the internal energy of the system. The statistical distribution analysis uses the Boltzmann distribution and Maxwell distribution to describe the velocity and energy distribution of the virtual particles, and the distribution parameters are fitted by the maximum likelihood estimation method. The partition function calculation uses the Monte Carlo method for numerical integration, and the sampling number is set to to times to ensure the calculation accuracy. The correlation function calculation obtains the time and space correlation of the system through autocorrelation and cross-correlation analysis, and the correlation length threshold is set to 0.1 to 0.3. The fluctuation dissipation relationship describes the relationship between the system response and thermal fluctuations through the Einstein relationship. The statistical mechanics eigenvector contains physical quantities such as the system entropy value, free energy gradient, partition function value, correlation length, and the vector dimension is usually 8 to 15 dimensions. The purpose of this step is to understand the operation law of the equipment from a physical perspective, and to provide the physical basis and theoretical support for anomaly detection through statistical mechanics theory.
[0048] Step S05 is implemented by using an industrial equipment condition assessment model to analyze statistical mechanical eigenvectors and calculate anomaly detection thresholds. The condition assessment model utilizes a neural network architecture based on an attention mechanism, with statistical mechanical eigenvectors as input and an equipment health score as output. A dynamic threshold adjustment function dynamically adjusts the detection threshold based on factors such as the equipment's operating ambient temperature, load conditions, and maintenance history. The ambient temperature correction factor ranges from 0.8 to 1.2, and the load correction factor ranges from 0.9 to 1.1. The anomaly detection threshold is determined using statistical control charts, employing a combination of moving average and exponentially weighted moving average control charts, with control limits set at three standard deviations. Equipment operating anomaly pattern recognition utilizes pattern matching and cluster analysis algorithms, including K-means clustering, hierarchical clustering, or density clustering. The number of clusters is determined to be three to eight categories based on historical anomaly types. The purpose of this step is to establish a quantitative assessment standard for equipment condition, adapting to varying operating conditions through a dynamic threshold mechanism, and improving the accuracy and robustness of anomaly detection.
[0049] Step S06 involves establishing an abnormality pattern classification system to classify and assess the levels of detected anomalies. Abnormality pattern classification utilizes supervised learning to train a classifier, with input features such as equipment status parameters at the time of the anomaly, the duration of the anomaly, and the degree of anomaly deviation. The classification algorithm can be a support vector machine, random forest, or deep neural network, with a classification accuracy threshold set at or above 90%. The abnormality severity assessment function uses a weighted factor of 0.4 for the degree of anomaly deviation, 0.3 for the duration, 0.2 for the impact range, and 0.1 for the historical frequency. Anomalies are graded using a four-tier system: minor anomalies are scored from 0 to 25, moderate anomalies from 25 to 50, severe anomalies from 50 to 75, and urgent anomalies from 75 to 100. Anomaly detection reports are generated using a templated format, including information such as the anomaly time, location, type, level, impact range, and recommended actions. This step facilitates structured analysis and grading of abnormalities, providing a quantitative basis for maintenance decision-making and guidance for prioritization.
[0050] The implementation method of step S07 is to build a feedback optimization mechanism to adjust the model parameters according to the detection effect. The parameter adaptive update algorithm monitors the performance indicators such as accuracy, recall rate, false alarm rate, etc. of anomaly detection. The accuracy threshold is set to 85%, the recall threshold is set to 80%, and the false alarm threshold is set to 15%. When the detection performance is lower than the threshold, the algorithm automatically adjusts the model weight parameters, detection threshold parameters and feature selection parameters. The weight update is optimized using a gradient descent algorithm or a genetic algorithm, the learning rate is set to 0.001 to 0.01, and the number of iterations is set to 100 to 1000 times. The threshold parameter adjustment uses a grid search or Bayesian optimization method to find the optimal parameter combination, and the search space is set to a reasonable range based on historical experience. The feature selection parameters evaluate the detection effect of different feature combinations through a cross-validation method, and select the feature subset with the best performance. The purpose of this step is to achieve adaptive optimization and continuous improvement of the detection system, and to maintain the stable performance of the system under different working conditions and time conditions through a feedback mechanism.
[0051] The detailed structure of the industrial equipment condition assessment model adopts a state fusion network architecture based on an attention mechanism. It consists of four main components: a data embedding layer, a multi-head self-attention layer, a feedforward neural network layer, and an output mapping layer. The data embedding layer converts the input statistical mechanical feature vector into a unified high-dimensional vector representation through linear transformation and positional encoding. The embedding dimension is set to 128 to 512 dimensions. The multi-head self-attention layer uses a parallel computing architecture. The number of attention heads is determined by the input feature dimension, ranging from 4 to 16. Each attention head independently calculates the query, key, and value matrices, capturing complex inter-feature relationships through a scaled dot product attention mechanism. The feedforward neural network layer adopts a two-layer fully connected network structure. The number of neurons in the first layer is 2 to 4 times the embedding dimension, using ReLU or GELU activation functions. The number of neurons in the second layer is the same as the embedding dimension. The output mapping layer uses linear transformation to map the high-dimensional features into equipment condition assessment results. The output dimension is determined by the number of condition categories. The entire network uses residual connections and layer normalization techniques to improve training stability. The dropout probability is set to 0.1 to 0.3 to prevent overfitting.
[0052] The detailed steps for building a training dataset include data collection, quality assessment, classification and labeling, data partitioning, and preprocessing. The data collection phase involves collecting equipment operating data from various industrial scenarios, covering typical equipment in industries such as steel, chemical, electric power, and machinery manufacturing. The data spans at least six months and contains complete records of normal operation, various anomalies, and failure modes. Quality assessment is performed through data integrity checks, consistency verification, and noise level assessment. Data samples with missing data exceeding 20%, timestamp errors, and obvious anomalies are eliminated. Classification and labeling are performed manually by domain experts based on equipment maintenance records and fault reports to create a supervised learning dataset with four status labels: normal, minor anomaly, moderate anomaly, and severe anomaly. Data partitioning follows the time series principle, with the dataset divided into a training set (60%), a validation set (20%), and a test set (20%) to ensure even distribution of data across different time periods and cover a wide range of operating conditions. Preprocessing includes data normalization, standardization, and data augmentation. Normalization utilizes maximum-minimum scaling or Z-score normalization. Data augmentation generates more training samples through time warping, noise injection, and sliding window sampling.
[0053] The reason why the industrial equipment status assessment model is suitable for solving the technical problems of the present invention is that it can effectively integrate multi-dimensional sensor data and accurately assess the operating status of the equipment. Compared with traditional anomaly detection methods based on single threshold judgment, this model dynamically assigns the importance weight of each sensor data through an attention mechanism, and can adaptively identify key fault characteristics. Compared with existing anomaly detection technologies based on statistical analysis, this model does not rely on preset statistical distribution assumptions and automatically learns complex patterns of equipment status through deep learning methods. The advantage of the model is that it can handle high-dimensional, nonlinear sensor data associations, and simultaneously focus on local features and global trends of time series through a multi-head attention mechanism. Compared with traditional support vector machines or random forest methods, it has stronger feature expression capabilities and generalization performance.
[0054] Virtual sensor algorithms are suitable for addressing sensor failure and missing data. Their technical advantage lies in reconstructing missing data through multi-sensor information fusion. Compared to traditional data interpolation methods, virtual sensor algorithms establish prediction models based on the physical relationships between sensors, enabling more accurate estimation of missing sensor values. Compared to existing time series prediction methods, this algorithm considers the spatial correlation between sensors and establishes a more accurate sensor correlation model through multivariate regression or neural network methods. This maintains data continuity and reliability even in the event of sensor failure.
[0055] The technical advantage of the virtual collection router model lies in its ability to optimize data transmission paths and fault recovery mechanisms from a network topology perspective. Compared to traditional fixed-route transmission methods, this model dynamically calculates the optimal transmission path using graph theory algorithms, automatically switching to an alternate path in the event of a network node failure. Compared to existing network redundancy technologies, the virtual collection router model not only implements path backup but also dynamically adjusts data transmission priorities based on network load and transmission quality, ensuring the real-time and integrity of critical sensor data.
[0056] The technical advantage of the statistical mechanics-based anomaly analysis model lies in its ability to understand device operating patterns from a physical perspective while simultaneously achieving efficient data dimensionality reduction. Using statistical mechanics principles, the model transforms the microscopic fluctuations of massive sensor data points into a small number of macroscopic thermodynamic parameters, which are used to describe the device's overall operating state. This effectively maps high-dimensional, complex data to low-dimensional physical characteristics. Compared to traditional empirical threshold-based detection methods, this model not only provides a theoretical foundation but also compresses hundreds of sensor measurements into a dozen or so statistical mechanics feature vectors through statistical aggregation, significantly reducing the computational complexity of data processing. Compared to existing mathematical dimensionality reduction methods such as principal component analysis, the dimensionality reduction process of the statistical mechanics-based model has clear physical meaning. It describes device states through statistical laws such as the Boltzmann distribution and the Maxwell distribution, converting complex multidimensional sensor correlations into interpretable physical quantities such as entropy, partition function, and correlation length. This provides improved detection capabilities and interpretability for small sample sizes and unknown anomaly types.
[0057] The key technical concepts of this invention include virtual sensor data reconstruction, statistical mechanics-like state modeling, attention mechanism state fusion, and adaptive parameter optimization. Virtual sensor data reconstruction technology, compared to traditional interpolation methods, takes into account the physical correlation between sensors and reconstructs missing data through multivariate statistical models and machine learning methods, ensuring the physical consistency and logical rationality of the data. Statistical mechanics-like state modeling technology maps the device operating state into a statistical physical space, providing a solid theoretical foundation compared to empirical threshold methods. It identifies anomalies through changes in entropy and phase transitions, providing an explainable physical mechanism. Attention mechanism state fusion technology dynamically assigns importance weights to different sensor data. Compared to fixed-weight methods, it can adaptively identify key fault characteristics, improving the accuracy and robustness of anomaly detection. Adaptive parameter optimization technology automatically adjusts model parameters based on detection results. Compared to static parameter settings, it can adapt to different operating conditions and time changes, maintaining the long-term stability of the detection system. The synergistic effect of these four technical concepts forms a complete anomaly detection solution. Virtual sensor technology ensures data integrity, statistical mechanics modeling provides theoretical support, the attention mechanism enables intelligent feature fusion, and adaptive optimization ensures system stability. Compared to existing technologies, it has higher detection accuracy, stronger adaptability, and better engineering practicality.
[0058] It should be noted that the present invention solves the technical problem of reduced reliability of anomaly detection systems caused by missing data and sensor failures in sensor networks. In the industrial Internet of Things environment, sensors often experience data loss or complete failure due to factors such as equipment aging, environmental interference, and communication failures. Traditional anomaly detection methods are highly dependent on complete data. Once data is missing, the detection accuracy will be greatly reduced or even the system will fail. The present invention constructs a virtual sensor algorithm and a virtual acquisition router model, establishes an association model between sensors, and uses the data of normal sensors to infer the data values of virtual sensors. At the same time, it optimizes the data transmission path based on graph theory and network topology analysis. When a sensor node failure is detected, the data transmission path is automatically reconstructed and virtual sensor node data is generated. The missing data is supplemented by neighboring node data interpolation and time series data extrapolation algorithms to ensure that the anomaly detection system can maintain normal operation even when some sensors fail.
[0059] It should be noted that the present invention also solves the technical problem that the fixed anomaly detection threshold leads to excessively high false alarm and missed alarm rates. Traditional anomaly detection methods usually use fixed thresholds for anomaly judgment, but the operating status of industrial equipment will change dynamically with factors such as load changes, environmental conditions and equipment aging. Fixed thresholds cannot adapt to such dynamic changes, and are prone to false alarms when the equipment is lightly loaded, and missed alarms when the equipment is heavily loaded. The present invention dynamically adjusts the anomaly detection threshold according to the equipment operating environment and historical data by establishing a dynamic threshold adjustment function and a parameter adaptive update algorithm, automatically adjusts the detection sensitivity by receiving the output results of the industrial equipment status assessment model, and establishes a feedback optimization mechanism to monitor the accuracy and recall rate of anomaly detection. When the detection performance decreases, the model weights and threshold parameters are automatically adjusted to achieve adaptive optimization of the detection system, effectively reducing the false alarm and missed alarm rates.
[0060] A second aspect of the present invention provides a computer-readable storage medium, in which program instructions are stored. When the program instructions are run in a computer, they are used to execute the above-mentioned method for detecting anomaly in industrial production Internet of Things data.
[0061] The third aspect of the present invention provides an industrial production Internet of Things data anomaly detection system, which includes the above-mentioned computer-readable storage medium. The system is any one of a computer, a server, and a single-chip microcomputer. The computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor that executes the program instructions stored in the computer-readable storage medium.
[0062] Specifically, the present invention addresses the curse of dimensionality in anomaly detection of high-dimensional sensor data. Its core principle is to leverage successful statistical mechanics methods for processing massive microscopic particle systems, transforming the microscopic states of high-dimensional sensor data into macroscopic thermodynamic parameters for dimensionality reduction analysis. In statistical mechanics theory, complex systems containing microscopic particles at the Avogadro constant level can be fully described by a handful of macroscopic thermodynamic parameters, such as temperature, pressure, volume, and entropy. This micro-to-macroscopic mapping is essentially a highly efficient dimensionality reduction process. The present invention applies this principle to industrial sensor data processing, treating each sensor data point as the microscopic state of a virtual particle. Multidimensional sensor data constitutes a virtual particle ensemble. By establishing a virtual thermodynamic system, random fluctuations in sensor data are mapped to the Brownian motion and thermal fluctuations of virtual particles. Changes in equipment energy consumption are mapped to the internal energy and enthalpy changes of the virtual thermodynamic system. Furthermore, equipment operational stability is mapped to the entropy and free energy gradient of the virtual system. This mapping process utilizes the partition function, correlation function, and fluctuation-dissipation relations of statistical mechanics to calculate macroscopic thermodynamic parameters, compressing sensor data originally dimensionalized into a dozen or so statistical mechanics eigenvectors, achieving exponential dimensionality reduction. Because thermodynamic parameters can reflect the equilibrium and non-equilibrium characteristics of a system, when equipment operation deviates from normal operating conditions, the virtual thermodynamic system will exhibit observable macroscopic changes such as entropy increase anomalies, phase transitions, and statistical distribution shifts. These changes can serve as effective indicators for anomaly detection. Preprocessing with principal component analysis and mutual information algorithms further optimizes data input quality, while the virtual sensor algorithm ensures data integrity. The state fusion network dynamically adjusts weight distribution through an attention mechanism. The entire technical solution forms a complete closed loop from data preprocessing, dimensionality compression, anomaly analysis, and feedback optimization, significantly reducing computational complexity while maintaining the accuracy and reliability of anomaly detection.
[0063] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows.
[0064] The specific implementation of step S01 is to obtain industrial equipment sensor data through a distributed data acquisition system and pre-process it. The specific representation of the sliding window smoothing filter algorithm is as follows: ; Where, For the The filtered output value of the sampling points; is the original sensor data sequence; is the half-length of the window, ranging from 2 to 7; is the total length of the sliding window. The 3σ criterion is used for the initial filtering of outliers, and the judgment conditions are as follows: ; Where, is the data point to be detected; is the mean of the data series; is the standard deviation of the data series; when this condition is met, Mark as an outlier.
[0065] The specific implementation of step S02 is to perform dimensionality reduction processing on the multidimensional sensor data set. The covariance matrix eigenvalue decomposition of the principal component analysis is expressed as follows: ; ; Where, for dimensional covariance matrix; for dimensional normalized data matrix, is the sample size, is the feature dimension; is a matrix The transposed matrix of ; For the feature vectors with dimension ; For the corresponding The principal component selection is determined by the cumulative contribution rate: ; Where, is the number of principal components retained. The calculation of mutual information feature selection is expressed as follows: ; Where, Characterized by With the target variable The mutual information value of For the characteristic variables; is the characteristic variable The specific value of is the target variable The specific value of for and The joint probability density of and They are and The marginal probability density of .
[0066] The specific implementation of step S03 is to construct a virtual sensor algorithm to generate missing sensor data. The multivariate linear regression model is expressed as follows: ; Where, Estimate values for virtual sensors; For the Normal sensor data; is the intercept term; is the regression coefficient; is the number of sensors involved in the regression; is the estimated error term, ranging from 0.01 to 0.05. The path weight calculation of the virtual collection router model is expressed as follows: ; Where, For sensor nodes and The path weight between them; For sensors and Correlation coefficient of the data; is the physical distance between nodes, in meters.
[0067] The specific implementation of step S04 is to establish a statistical mechanics anomaly analysis model, which is composed of the following equations. The Boltzmann distribution of the equipment operating parameters is expressed as follows: ; Where, The device's operating status The probability distribution of Status The corresponding effective energy is obtained by mapping the normalized sensor data in joules; is the partition function; is the Boltzmann constant, which takes the value Joule per Kelvin; is the effective temperature parameter in Kelvin, obtained by fitting the equipment operation data. The calculation of the partition function is expressed as follows: ; Where, is the total number of states, which is determined by the degree of discretization of sensor data and typically ranges from 100 to 1000. The system entropy value is calculated as follows: ; Where, is the system entropy value, in joules per Kelvin. The average energy calculation is expressed as follows: ; Where, is the average energy of the system in joules. The energy fluctuation calculation is expressed as follows: ; Where, is the standard deviation of energy fluctuations, in joules; is the expected value of the square of energy, through Calculated; is the square of the average energy. The heat capacity calculation is expressed as follows: ; Where, is the constant volume heat capacity in joules per kelvin. The relevant length calculation is expressed as follows: ; Where, is the spatial correlation length, in meters; Status The corresponding spatial position vector, in m; Status The corresponding spatial position vector, in m; Status and The Euclidean distance between and Status and The probability distribution of . The relevant time calculation is expressed as follows: ; Where, is the system-related time, in seconds; is the time variable, the unit is s; is the normalized autocorrelation function; for The instantaneous energy of the system at the moment; is the system energy at the initial moment; is the time-dependent average value of energy at different moments. The construction of the statistical mechanics eigenvector is expressed as follows: ; Where, is a 6-dimensional statistical mechanics eigenvector.
[0068] The specific implementation of step S05 is to use the industrial equipment status assessment model to analyze the statistical mechanics characteristic vector. The dynamic threshold adjustment function is expressed as follows: ; Where, is the detection threshold after dynamic adjustment; is the basic threshold; is the ambient temperature correction factor, ranging from 0.8 to 1.2; is the load correction factor, ranging from 0.9 to 1.1; To maintain historical correction factors, the range is 0.85 to 1.15.
[0069] The specific implementation of step S06 is to establish an abnormal pattern classification system. The abnormal severity evaluation function is expressed as follows: ; Where, score the severity of the abnormality; is the degree of abnormal deviation; is the abnormal duration; For the scope of influence; is the historical abnormal frequency; , , , is the corresponding weight coefficient.
[0070] The specific implementation of step S07 is the same as above and will not be described in detail here.
[0071] The attention adjustment coefficient value of the state weight adjustment function is calculated as follows: ; Where, is the attention adjustment coefficient value; is the sensor data variance; is the data correlation coefficient, ranging from 0 to 1; is the historical accuracy of anomaly detection, ranging from 0 to 1; The current device load rate ranges from 0 to 1; and is the weighting coefficient, usually , ; To adjust the error term, the range is -0.05 to 0.05. The first interval value is calculated as follows: ; Where, is the first interval value. The second interval value is calculated as follows: ; Where, is the second interval value. The third interval value is calculated as follows: ; Where, is the third interval value, calculated using the harmonic mean method.
[0072] The principles and effects of each formula are explained as follows. Sliding window smoothing filter formula Based on the finite impulse response filter principle in digital signal processing, high-frequency noise is suppressed by performing weighted averaging on adjacent data points. Compared with the traditional single-point detection method, this formula can effectively reduce the impact of random noise on anomaly detection accuracy and improve the stability and reliability of data preprocessing. Criteria judgment formula Based on the statistical characteristics of the normal distribution, the standard deviation of the data from the mean is used to identify outliers. Compared with the fixed threshold method, this formula can adaptively adjust the judgment criteria according to the data distribution characteristics, reducing misjudgments caused by data distribution differences.
[0073] Principal component analysis eigenvalue decomposition formula Based on the matrix eigendecomposition theory in linear algebra, dimensionality reduction is achieved by finding the projection direction with the largest data variance. Compared with traditional feature selection methods, this formula can retain the main information of the data while reducing the dimension, thereby improving the computational efficiency of subsequent algorithms. Based on the principle of information measurement in information theory, feature importance is evaluated by calculating the statistical dependence between variables. Compared with the linear correlation coefficient method, this formula can capture nonlinear correlation relationships and improve the accuracy of feature selection.
[0074] Virtual sensor multiple regression formula Based on the least squares estimation theory, the missing data is predicted by establishing a linear relationship model between sensors. Compared with the simple interpolation method, this formula takes into account the synergy of multiple sensors and improves the accuracy and physical rationality of data reconstruction. Path weight calculation formula The correlation of sensor data and physical distance factors are comprehensively considered, and the connection strength between nodes is quantified through the distance-normalized correlation coefficient. Compared with the fixed topology structure, this formula can dynamically optimize the data transmission path and improve the network's adaptability.
[0075] Boltzmann distribution formula Derived from the canonical ensemble theory in statistical physics, it maps the operating state of the device to the energy distribution of the thermodynamic system and describes the state probability through an exponential decay function. Compared with the empirical probability model, this formula provides a solid physical theoretical foundation and improves the interpretability of anomaly detection. As a core concept of statistical mechanics, the probability distribution is normalized by the statistical summation of all possible states. Compared with direct probability calculation, this formula ensures the mathematical rigor and physical consistency of the probability distribution. Based on the second law of thermodynamics, the degree of disorder of the system is quantified by the weighted logarithm sum of the probability distribution. Compared with the traditional variance index, this formula can more accurately reflect the complexity and stability of the system state. Average energy formula The expected energy value of the system is calculated by probability weighted summation, which provides the basic thermodynamic parameters for equipment status evaluation. The statistical fluctuation of energy calculated based on the variance definition reflects the stability of the system. Compared with direct variance calculation, this formula has a clear physical meaning in the framework of statistical mechanics. Based on the fluctuation-dissipation theorem, the heat capacity is calculated through the relationship between energy fluctuation and temperature, providing a quantitative indicator for the system response characteristics.
[0076] Dynamic threshold adjustment formula The adaptive adjustment of the threshold is achieved through the product form of multi-factor correction, taking into account factors such as environmental conditions, operating load and maintenance history. Compared with the fixed threshold method, this formula can adapt to changes in different working conditions and significantly reduce the false alarm rate and missed alarm rate. A weighted linear combination is used to combine multiple abnormal characteristic indicators into a single scoring value. The importance of each factor is balanced by the weight coefficient determined by expert experience. Compared with single indicator evaluation, this formula can comprehensively quantify the severity of the abnormality and provide a scientific basis for maintenance decision-making.
[0077] State weight adjustment function The linear combination form of weighted geometric mean is adopted to comprehensively evaluate the system state through the weighted summation of two geometric mean terms. Compared with the simple linear combination, the geometric mean term in this formula can better handle the relationship between parameters of different dimensions. 、 The mathematical forms of geometric mean are used respectively. The geometric mean is suitable for representing the comprehensive level of two positive numbers. The harmonic mean formula More suitable for processing the average of ratio data, compared with simple arithmetic mean, these formulas can more accurately reflect the intrinsic relationship between different parameters, and improve the scientificity and effectiveness of the weight allocation strategy.
[0078] To better understand and implement the present invention, Example 2, a specific application scenario, is provided below: A technical team deployed an industrial production IoT data anomaly detection system on a production line, equipped with 276 sensors to monitor equipment operating status. The system operated in an ambient temperature range of 15°C to 45°C, with equipment load fluctuating between 60% and 95%. The sensor data acquisition frequency was set between 50Hz and 500Hz.
[0079] The technical team first established a distributed data acquisition system and deployed 42 temperature sensors, 18 pressure sensors and 12 The concentration sensor has a measurement range of 800℃ to 1200℃, 0.5MPa to 2.5MPa and 200ppm to 800ppm. The roughing mill is equipped with 36 vibration sensors, 24 current sensors and 15 speed sensors. The vibration acceleration measurement range is to m / s², current range is 100A to 2000A, and speed range is 50rpm to 800rpm. The finishing mill is equipped with 48 temperature sensors, 32 pressure sensors, 28 vibration sensors, and 20 displacement sensors, with displacement measurement accuracy reaching 10μm. The coiler area is equipped with 21 tension sensors and 18 temperature sensors, with a tension measurement range of 500N to N.
[0080] The sliding window smoothing filter algorithm was used in the sensor data preprocessing stage, and the window length was set to 9 sampling points, which effectively removed high-frequency noise with a frequency higher than 25Hz. During the initial filtering of outliers, it was found that the mean temperature data of the heating furnace area was 1050℃ and the standard deviation was 45℃. 218 outlier data points were identified and removed. The mean vibration data of the roughing mill was 0.8m / s² and the standard deviation was 0.15m / s², and 134 outliers were filtered out. The established multidimensional sensor data set contains 276 feature dimensions, with a time span of 72 consecutive hours and a total data volume of sampling points.
[0081] The principal component analysis (PCA) dimensionality reduction stage calculated the eigenvalue distribution of the covariance matrix. The cumulative contribution of the first 15 principal components reached 91.3%, so these 15 principal components were retained as features after dimensionality reduction. The mutual information feature selection algorithm calculated the correlation between each feature and equipment abnormality and selected the top 68 feature variables with the highest mutual information values. The final dimensionality reduction feature dataset was 68, a 75.4% reduction compared to the original 276-dimensional data.
[0082] During the construction of the virtual sensor algorithm, the technical team discovered that the No. 3 vibration sensor of the roughing mill had intermittent failures, with a data loss rate of 12.8%. A multivariate linear regression model was used to establish the correlation between this sensor and eight adjacent sensors. The regression coefficients were 0.73, -0.42, 0.56, 0.31, -0.28, 0.65, -0.39, and 0.44, respectively. The intercept term was 0.067, and the estimated error was controlled within 3.2%. The virtual acquisition router model established a node connection diagram based on the sensor network topology, such as Figure 3As shown in the figure, six critical path nodes and 12 redundant transmission paths were identified. When sensor No. 3 failed, the system automatically switched to the backup data source and generated virtual sensor data by interpolating data from neighboring sensors, improving data integrity to 98.7%.
[0083] The statistical mechanics anomaly analysis model maps the equipment operating parameters to the statistical physical space. The effective temperature parameter fitting value is 850K, which corresponds to the comprehensive operating temperature level of the equipment. The partition function calculation uses the Monte Carlo method for numerical integration, and the sampling number is set to times, the calculation accuracy reaches Level. The system entropy value is stable at Near J / K, the average energy is J, the standard deviation of energy fluctuation is J. The calculated constant volume heat capacity is J / K, the spatial correlation length is 2.3m, and the temporal correlation length is 42s. The constructed statistical mechanics eigenvector contains 6 components, which effectively compresses the complex information of high-dimensional sensor data.
[0084] The industrial equipment condition assessment model uses a state fusion network based on an attention mechanism. The network structure consists of a 128-dimensional data embedding layer, an 8-head self-attention layer, a 256-dimensional feedforward neural network layer, and an output mapping layer. The training dataset contains 65,000 samples of normal operation data, 8,500 samples of minor anomaly data, 3,200 samples of moderate anomaly data, and 1,800 samples of severe anomaly data. The data distribution is shown in Table 1.
[0085] Table 1. Statistics of equipment abnormality type distribution
[0086] The dynamic threshold adjustment function makes real-time adjustments based on environmental conditions. During the observation period, the ambient temperature correction factor varied from 0.85 to 1.15, the load correction factor varied from 0.92 to 1.08, and the maintenance history correction factor remained stable around 0.95. The base threshold was set at 0.75, and after dynamic adjustment, the detection threshold fluctuated between 0.68 and 0.82, effectively adapting to varying operating conditions.
[0087] The abnormality detection operation results show that the system detected 127 equipment operation abnormalities during the 72-hour monitoring period, including 72 minor abnormalities, 38 moderate abnormalities, 15 serious abnormalities, and 2 emergency abnormalities. Figure 4As shown, the statistical mechanics eigenvectors exhibit significant distribution differences under different abnormal conditions. Under normal conditions, the system entropy is concentrated in the low range, while under abnormal conditions, the entropy increases significantly, reflecting the increasing degree of disorder in the system. The calculation results of the anomaly severity assessment function show that the average score for minor anomalies is 18.5, the average score for moderate anomalies is 42.3, the average score for severe anomalies is 68.7, and the average score for emergency anomalies is 89.2. The score distribution is reasonable and the discrimination is clear.
[0088] During the calculation of the attention adjustment coefficient value of the state weight adjustment function, the sensor data variance is 0.34, the data correlation coefficient is 0.67, the anomaly detection historical accuracy is 0.91, and the current device load rate is 0.78. The calculated attention adjustment coefficient value is 0.72, which is between the second interval value of 0.83 and the third interval value of 0.59. The system adopts an active weight allocation strategy to enhance the response sensitivity of the attention weight. Figure 5 As shown, compared with the traditional fixed threshold detection method, the dynamic threshold adjustment mechanism of the present invention significantly reduces false alarms and improves the accuracy of anomaly detection.
[0089] A feedback optimization mechanism continuously adjusts model parameters based on detection results. During the monitoring period, the system maintained an accuracy of 92.4%, a recall rate of 88.7%, and a false alarm rate of less than 6.8%. The adaptive parameter update algorithm adjusted network weights 15 times, optimized detection threshold parameters eight times, and updated feature selection parameters five times, ensuring the continued stability of the detection system's performance.
[0090] The system established a sensor connectivity graph consisting of 276 nodes and 456 edges, identifying 18 key hub nodes and 32 redundant path branches. The network diameter was 7 hops, the average path length was 3.2 hops, and the clustering coefficient was 0.43, demonstrating good small-world network characteristics. In the event of a sensor failure, the network was able to reconstruct the path within an average of 1.8 seconds, keeping the increase in data transmission latency to less than 15%.
[0091] Statistical analysis of the test results showed that the accuracy of anomaly detection in the heating furnace area reached 94.2%, with over-temperature and uneven combustion being the primary anomaly types. The accuracy of anomaly detection in the roughing mill was 91.8%, with abnormal vibration and current fluctuation being the primary failure modes. The accuracy of detection in the finishing mill was 93.5%, with abnormal pressure and displacement being typical anomalies. The accuracy of detection in the coiler area was 90.7%, with tension fluctuation and temperature anomalies being the most common. The anomaly detection performance indicators for each area are shown in Table 2.
[0092] Table 2 Statistics of abnormal detection performance in each device area
[0093] This invention represents a significant technological advancement over traditional anomaly detection methods. Traditional methods primarily rely on single-sensor data and fixed thresholds, making them susceptible to environmental interference and changes in equipment operating conditions, resulting in high false alarm rates and poor adaptability. This invention addresses the problem of missing data caused by sensor failures through a virtual sensor algorithm. The data reconstruction method based on multi-sensor information fusion offers greater physical plausibility and estimation accuracy than simple interpolation. A statistical mechanics-inspired anomaly analysis model maps equipment operating states into a statistical physics space, leveraging the principles of entropy increase and phase transitions to identify anomaly patterns. Compared to empirical threshold methods, this model has a solid theoretical foundation and greater interpretability. The attention mechanism state fusion network dynamically assigns importance weights to different sensor data, adaptively identifying key fault signatures and significantly improving anomaly detection accuracy and robustness compared to fixed-weight methods. A dynamic threshold adjustment mechanism optimizes detection parameters in real time based on environmental conditions and equipment status. Compared to static threshold settings, this mechanism better adapts to changing operating conditions and effectively reduces false alarm and missed alarm rates. An adaptive parameter update algorithm continuously optimizes model performance through feedback from detection results, ensuring long-term system stability and accuracy. This approach is more intelligent and efficient than traditional manual parameter adjustment methods.
[0094] It should be noted that the detailed explanation of the variables involved in the present invention is shown in Table 3.
[0095] Table 3 Variable Explanation Table
[0096] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.
Claims
1. A method for detecting anomaly in industrial production Internet of Things data, characterized in that: It includes collecting sensor data generated by industrial equipment and preprocessing it to establish a multidimensional sensor data set, using principal component analysis to reduce the dimensionality of the multidimensional sensor data set and using the mutual information algorithm to perform feature selection to establish a reduced dimensionality feature data set, constructing a virtual sensor algorithm to generate virtual sensor data from the reduced dimensionality feature data set to form an extended sensor data set, establishing a statistical mechanics-like anomaly analysis model based on the laws of statistical mechanics to map the extended sensor data set to the statistical mechanics space to calculate the statistical mechanics feature vector, using the industrial equipment status assessment model to analyze the statistical mechanics feature vector, calculating the anomaly detection threshold through the dynamic threshold adjustment function to identify the abnormal operation mode of the equipment, establishing an abnormal mode classification system to classify and mark the abnormal operation mode of the equipment, calculating the anomaly level through the anomaly severity assessment function to generate an anomaly detection report, and constructing a feedback optimization mechanism to adjust the detection parameters according to the anomaly detection report, and optimizing the performance of the industrial equipment status assessment model through the parameter adaptive update algorithm to achieve continuous improvement of the anomaly detection system.
2. The method for detecting anomaly in industrial production Internet of Things data according to claim 1, characterized in that: The virtual sensor algorithm is specifically a calculation method based on existing sensor data to infer missing or faulty sensor data. By establishing an association model between sensors and using the data of normal sensors to infer the data value of the virtual sensor, the data missing problem in the sensor network is compensated.
3. The method for detecting anomaly in industrial production Internet of Things data according to claim 2, characterized in that: The virtual sensor algorithm adopts a virtual acquisition router model, specifically a data path optimization model based on graph theory and network topology analysis. By constructing a connection relationship graph between sensor nodes and calculating the shortest path weights between nodes, it identifies weak links and data transmission bottlenecks in the network from the perspective of the sensor network topology structure.
4. The method for detecting anomaly in industrial production Internet of Things data according to claim 3, characterized in that: When a sensor node failure or data transmission interruption is detected, the virtual collection router model automatically reconstructs the data transmission path and generates virtual sensor node data based on the redundant paths and backup node information of the network topology, and compensates for the missing data through neighboring node data interpolation and time series data extrapolation algorithms.
5. The method for detecting anomaly in industrial production Internet of Things data according to claim 4, characterized in that: The simulated statistical mechanics anomaly analysis model is specifically an equipment state analysis model constructed based on statistical mechanics theory and thermodynamic equilibrium principles. It maps the multi-dimensional operating parameters of industrial equipment to the microscopic particle state distribution and macroscopic thermodynamic parameters in the statistical mechanics space, and simulates the equipment operating state by establishing a virtual particle ensemble and thermodynamic system.
6. The method for detecting anomaly in industrial production Internet of Things data according to claim 5, characterized in that: The simulated statistical mechanics anomaly analysis model maps the random fluctuations of equipment operating parameters into Brownian motion and thermal fluctuations of virtual particles, maps the changes in equipment energy consumption into the internal energy and enthalpy changes of a virtual thermodynamic system, and maps the equipment operation stability into the entropy value and free energy gradient of the virtual system.
7. The method for detecting anomaly in industrial production Internet of Things data according to claim 6, characterized in that: The simulated statistical mechanics anomaly analysis model analyzes the regularity of equipment operation data through the Boltzmann distribution law and Maxwell statistical distribution, calculates the partition function, correlation function and fluctuation-dissipation relationship of the virtual particle ensemble, and identifies abnormal patterns by detecting the entropy increase anomaly, phase change phenomenon and statistical distribution deviation of the virtual thermodynamic system.
8. The method for detecting anomaly in industrial production Internet of Things data according to claim 7, characterized in that: The statistical mechanics characteristic vector is specifically a set of vectors representing the operating state of the equipment obtained by calculating the statistical mechanics anomaly analysis model, which includes the microscopic state distribution characteristics, macroscopic thermodynamic equilibrium state and statistical fluctuation law during the operation of the equipment.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program instructions, and when the program instructions are run in a computer, they are used to execute the method for detecting anomaly in industrial production Internet of Things data according to any one of claims 1 to 8.
10. An industrial production Internet of Things data anomaly detection system, characterized in that: The computer-readable storage medium according to claim 9 is included, the system is any one of a computer, a server, and a single-chip microcomputer, the computer-readable storage medium is arranged in the system, and the system is provided with a microprocessor for executing program instructions stored in the computer-readable storage medium.
Citation Information
Patent Citations
Spacecraft system anomaly detection method based on high-dimensional space mapping
CN111274543A
Photovoltaic power station performance monitoring and analyzing method
CN118032124A
Intelligent control method for chemical industry park based on digital twinning
CN119940005A
Artificial intelligence detection system and method for computer system fault
CN120508421A
Detection of abnormal behaviour of devices from associated unlabeled sensor observations
US20220092432A1
Cited By
Self-adaptive network decoy system and method based on dynamic honey spots
CN121547299A
Industrial supply chain intelligent regulation and control method, medium and system based on AI vertical domain model
CN121581773A
Data processing method and system for insulator inspection by unmanned aerial vehicle, electronic equipment and medium
CN122336610A