An online monitoring and fault early warning method for safe operation of industrial equipment
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SICHUAN ACAD OF SAFETY SCI & TECH
- Filing Date
- 2026-04-30
- Publication Date
- 2026-08-07
AI Technical Summary
[0011](1)解决工业现场强干扰、变工况环境下数据失真与特征漂移问题,通过工况分段归一化预处理与微点突变捕获机制,剔除噪声干扰,消除工况波动影响,同时捕捉故障萌芽阶段的微弱瞬态信号,为早期故障识别提供支撑;
[0079]本发明与现有技术相比,具有以下突出的实质性特点与显著的进步,具备极高的创造性、新颖性与实用性:
Smart Images

Figure CN122528022A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent operation and maintenance and fault diagnosis technology for industrial equipment. Specifically, it relates to an online monitoring and fault early warning method for the safe operation of industrial equipment. It is applicable to real-time safety monitoring, early fault identification, and graded early warning of rotating machinery (pumps, fans, machine tool spindles, gearboxes), stationary equipment (pressure vessels, pipelines, heat exchangers), and electrical equipment (motors, frequency converters) in various industrial production scenarios. It can be widely used in industrial fields such as petrochemicals, metallurgy and mining, intelligent manufacturing, power energy, and rail transportation, providing technical support for predictive maintenance of industrial equipment throughout its entire life cycle and ensuring the continuity, safety, and economy of the production process. Background Technology
[0002] With the deepening advancement of Industry 4.0, industrial production is rapidly developing towards intelligence, continuous operation, and high-load capacity. As the core carrier of the production process, the operating status of industrial equipment directly determines production efficiency, product quality, and production safety. Once equipment malfunctions, it can not only lead to production interruptions and huge economic losses, but also cause safety accidents such as equipment damage and personal injury. Therefore, realizing online real-time monitoring and early warning of faults for industrial equipment, and promoting the transformation of operation and maintenance models from "reactive maintenance" to "predictive maintenance," has become a key technical problem that urgently needs to be solved in the industrial sector.
[0003] Currently, existing technologies in the field of industrial equipment monitoring and fault early warning all have obvious defects and shortcomings, and cannot fully meet the high requirements of modern industrial production. The core pain points are as follows:
[0004] 1. Traditional monitoring modes have fundamental limitations: manual inspection mode is limited by personnel experience and inspection frequency, making it impossible to achieve real-time monitoring, resulting in serious delays in early warning and safety risks in high-risk scenarios; single-parameter offline monitoring mode cannot fully reflect the fault evolution law of equipment multi-physical field coupling, which is prone to false and missed early warnings, and offline analysis leads to untimely early warnings.
[0005] 2. Insufficient core capabilities of the multi-parameter online monitoring solution: First, weak data preprocessing capabilities. In industrial environments with strong interference, traditional filtering methods cannot effectively remove noise, resulting in a low signal-to-noise ratio. Second, weak early fault identification capabilities. Traditional feature extraction methods can only obtain shallow features and cannot capture weak "soft signs" in the early stages of faults. Early fault features are easily missed when they highly overlap with normal signals. Third, rigid threshold settings. Fixed thresholds cannot adapt to load fluctuations, environmental changes, and normal state drift caused by equipment aging, resulting in high false alarm and false alarm rates. Fourth, insufficient depth of multi-source data fusion. Various sensor parameters form information silos. Cross-domain correlation features are not explored by combining the physical topology of the equipment and coupling mechanisms, resulting in poor fault location and root cause tracing capabilities.
[0006] 3. Existing intelligent monitoring solutions have application shortcomings: Deep learning-based monitoring solutions are mostly designed for single types of equipment, resulting in poor versatility. The scarcity of fault samples in industrial settings leads to insufficient model training and high computational demands, making it difficult to meet real-time requirements. Existing dynamic threshold solutions mostly optimize single parameters and fail to achieve deep fusion of multi-source data, thus failing to solve the problem of identifying multiple coupled faults under complex operating conditions. Most models remain fixed after deployment, lacking online self-evolution capabilities. As equipment ages and operating conditions drift, model performance rapidly degrades, making it unable to adapt to the full lifecycle maintenance needs of equipment. Furthermore, the lack of an edge-cloud collaborative architecture makes it impossible to balance the computational power requirements of real-time inference on-site and deep analysis in the cloud, limiting industrial applicability.
[0007] 4. Disconnect between early warning and closed-loop operation and maintenance management: Existing solutions mostly rely on binary judgment of "normal / abnormal" without combining equipment degradation mechanism, remaining life prediction and maintenance decision-making for risk classification. The interpretability of early warning information is poor, and it cannot provide accurate fault tracing and maintenance guidance. The conversion efficiency of early warning information into maintenance actions is low, making it difficult to achieve closed-loop management of the entire process of monitoring, early warning, maintenance and optimization.
[0008] In summary, existing technologies all suffer from drawbacks such as low monitoring accuracy, difficulty in early fault identification, delayed early warning, poor adaptability to operating conditions, lack of self-evolutionary model capabilities, and insufficient practicality, failing to fully meet the high requirements of modern industrial production for safe equipment operation. Therefore, developing an online monitoring and fault early warning method that combines high accuracy, strong adaptability, a closed-loop process, and high industrial applicability has become an urgent need in this field. Summary of the Invention
[0009] In response to the core deficiencies of existing technologies, the present invention aims to provide an online monitoring and fault early warning method for the safe operation of industrial equipment, which realizes real-time and comprehensive monitoring of the operating status of industrial equipment, accurate identification of early faults, hierarchical adaptive early warning, fault root cause tracing and whole-system self-optimization, completely solving the core pain points of existing technologies and significantly improving the inventiveness, novelty and practicality of the patent.
[0010] The specific objectives of this invention include:
[0011] (1) Solve the problems of data distortion and feature drift under strong interference and variable working conditions in industrial sites. Through the working condition segmentation normalization preprocessing and micro point mutation capture mechanism, noise interference is eliminated, the influence of working condition fluctuation is eliminated, and weak transient signals in the fault initiation stage are captured to provide support for early fault identification.
[0012] (2) Achieve deep fusion of multi-source heterogeneous data, construct a three-level feature extraction system of “shallow features - deep features - spatiotemporal graph aggregation features”, combine equipment physical topology and coupling mechanism, break information silos, fully explore cross-domain fault correlation features, and improve the early fault identification accuracy to over 96%;
[0013] (3) Construct a digital twin-driven working condition adaptive dynamic baseline and a three-level progressive early warning system, combined with risk-driven dynamic threshold optimization, to completely solve the problem of rigid fixed thresholds, and control the false early warning rate to below 1.5% and the missed early warning rate to below 0.8%;
[0014] (4) Construct a full-process architecture of edge-cloud collaboration, balance the computing power requirements of real-time on-site early warning and cloud-based in-depth analysis, realize millisecond-level edge inference and continuous optimization of cloud-based models, and meet the real-time and long-term stability requirements of industrial sites;
[0015] (5) Achieve accurate fault tracing and closed-loop operation and maintenance management. By combining spatiotemporal diagram anomaly propagation path analysis and cause-effect diagram backtracking, accurately locate the root cause and propagation path of the fault, generate targeted maintenance suggestions, and improve maintenance efficiency by more than 40%.
[0016] (6) Construct an online self-evolution mechanism for the entire lifecycle model. Through incremental learning, adversarial training, and federated updates, solve the problems of static model solidification and performance degradation, extend the model performance stability period by more than 3 times, and improve the versatility and generalization ability of the solution to adapt to all types of industrial equipment and complex industrial scenarios.
[0017] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: an online monitoring and fault early warning method for the safe operation of industrial equipment, comprising the following eight core steps, constructing a closed-loop system covering the entire process of "data acquisition - preprocessing - feature extraction - status assessment - hierarchical early warning - fault tracing - self-optimization - visualization," as detailed below:
[0018] Step S1: Equipment monitoring point layout and edge-cloud collaborative data acquisition
[0019] Based on the type, structural characteristics, fault-prone parts, and operating conditions of industrial equipment, key monitoring parts and parameters are determined, multi-source sensor arrays are deployed, and an edge-cloud collaborative multi-source data acquisition system is constructed to achieve real-time synchronous acquisition and low-latency transmission of equipment operating parameters.
[0020] The specific implementation details are as follows:
[0021] (1) Monitoring parameters and sensor selection: The monitoring parameters cover four categories: mechanical parameters, electrical parameters, environmental parameters, and process parameters. Among them, mechanical parameters include vibration acceleration / velocity / displacement, rotational speed, noise, and ultrasonic emission signals; electrical parameters include current, voltage, power, power factor, winding temperature, and leakage flux signals; environmental parameters include ambient temperature, humidity, dust concentration, and corrosive gas concentration; and process parameters include medium pressure, flow rate, temperature, and concentration. The sensor selection follows the principle of "accuracy matching, environmental matching, and stable performance". The vibration sensor is a piezoelectric accelerometer (measurement range 0.1-100g, frequency range 10-10000Hz), the temperature sensor is a PT100 resistance temperature detector (measurement range -50℃~200℃, accuracy ±0.5℃), and the current sensor is a Hall current sensor (accuracy ±0.2%). An ultrasonic emission sensor is added to capture the transient impact signal of early faults, with a sampling frequency of not less than 25.6kHz.
[0022] (2) Monitoring point layout: Monitoring points are set up for the bearings, spindles and gearboxes of rotating machinery, the interfaces, welds and stress-bearing parts of stationary equipment, and the windings and terminals of electrical equipment, which are prone to failure. Multiple types of sensors are set up at the same key parts to realize the synchronous acquisition of multiple physical field parameters.
[0023] (3) Edge-Cloud Collaborative Acquisition Architecture: An industrial-grade edge computing gateway is deployed, equipped with a high-speed data acquisition card. The sampling frequency can be adaptively adjusted according to the equipment type and operating conditions (10Hz-1000Hz, with a maximum of 25.6kHz for high-frequency ultrasonic signals). After A / D conversion and preliminary filtering, the acquired data is transmitted to the edge computing gateway via wired (Ethernet, RS485) or wireless (5G, LoRa) methods, with transmission latency controlled within 100ms. The IEEE1588v2 precision time protocol is used to achieve millisecond-level synchronization of multi-channel data, with timestamp errors controlled within 10ms. The data acquisition module has a local caching capability of more than 72 hours, automatically caching data when the network is interrupted and retransmitting it after recovery to avoid data loss. The edge gateway is responsible for data preprocessing and lightweight real-time inference, while the cloud server is responsible for deep feature mining, model training, and global optimization. The two work together to complete the entire data processing process.
[0024] Step S2: Adaptive preprocessing of multi-source data and micro-point mutation capture
[0025] After receiving the raw data, the edge computing gateway performs a full-process preprocessing of "operating condition segmentation - three-level data cleaning - operating condition normalization - dual feature stream extraction" to remove noise and outliers, eliminate the influence of operating condition drift, and extract micro-point mutation event streams representing early faults and feature streams reflecting macroscopic states of gradual change, laying the foundation for subsequent feature extraction.
[0026] The specific implementation details are as follows:
[0027] Sub-step S21: Operating condition segmentation and three-level data cleaning
[0028] (1) Operating condition segmentation: Based on the equipment operating parameters and process instructions, the data is divided into steady state segment (operating condition parameter fluctuation <3% for more than 10 minutes), transition segment (operating condition switching period such as start-up, shutdown, load adjustment, etc.), and transient segment (instantaneous action period such as safety valve release). Differentiated preprocessing strategies are adopted for different segments.
[0029] (2) Three-level data cleaning: The three-level cleaning process of "outlier identification - noise removal - missing value completion" is adopted: ① Outlier identification adopts the combination of 3σ criterion and isolated forest algorithm. First, suspected outliers are marked by 3σ criterion, and then the isolated forest algorithm is used to distinguish between real outliers and extreme values of normal operating conditions, so as to avoid the deletion of valid data by mistake; ② Noise removal adopts the combination of Kalman filtering and wavelet packet transform. First, high-frequency electromagnetic interference is removed by Kalman filtering, and then the noise is deeply removed by db4 wavelet 3-level decomposition. After processing, the signal-to-noise ratio of the data is improved to more than 28dB; ③ Missing value completion adopts LSTM time series prediction algorithm, which combines historical parameter data and real-time data of related parameters to complete missing values. The completion accuracy is not less than 98%. If the missing value cannot be completed within 10 minutes, the sensor fault warning is triggered.
[0030] Sub-step S22: Operating condition normalization processing. For transition and variable operating condition data, a three-dimensional operating condition reference space of load-speed-ambient temperature is constructed. The local weighted regression method is used to map the original features to the equivalent values under rated operating conditions, eliminating the modulation effect of load fluctuation, speed change, and ambient temperature and humidity change on feature amplitude, realizing cross-operating condition feature comparability, and solving the problem of false early warning caused by feature drift under variable operating conditions.
[0031] Sub-step S23: Dual Feature Flow Extraction
[0032] (1) Micro-point mutation event stream extraction: For high-frequency vibration and ultrasonic emission signals, an impact event detector combining continuous wavelet transform and peak hold downsampling is used to locate and parameterize non-stationary, short-time energy mutation events (micro-point mutations) in the signal, and generate a four-dimensional feature vector containing the event occurrence time, energy intensity, duration and frequency centroid to form a micro-point mutation event stream, and capture weak signals in the early stage of faults such as early micro-stripping of bearings and partial discharge of insulation;
[0033] (2) Extraction of slowly changing trend feature flow: For slowly changing signals such as temperature, pressure, and current, extract the mean, standard deviation, and slope of change within a short window, and combine them with the Mann-Kendall trend test statistic to form a slowly changing trend feature flow that reflects the macroscopic state drift of the equipment.
[0034] Step S3: Multi-dimensional fault feature extraction and deep fusion of spatiotemporal graph
[0035] The preprocessed data is transmitted to the cloud server, and a three-level feature extraction process of "shallow feature extraction - deep feature mining - spatiotemporal graph aggregation and fusion" is adopted to fully explore the fault features in multi-source data and solve the problems of insufficient feature extraction and information silos in existing technologies.
[0036] The specific implementation details are as follows:
[0037] Sub-step S31: Shallow feature extraction
[0038] For different types of monitoring parameters and dual-feature flows, the corresponding shallow fault features are extracted: ① Mechanical parameters: extract the time-domain features (peak value, RMS value, kurtosis, waveform factor), frequency-domain features (dominant frequency, band energy, harmonic components), and statistical features of micro-point mutation events (event occurrence rate, average energy, frequency centroid distribution); ② Electrical parameters: extract the RMS values of current and voltage, harmonic distortion rate, three-phase imbalance, power fluctuation value, and winding temperature change rate; ③ Environmental and process parameters: extract parameter fluctuation values, deviation rate, and trend stability coefficient; ④ Statistical features and trend test results of slowly changing trend characteristic flows.
[0039] Sub-step S32: Deep Feature Mining
[0040] An improved ResNet50-1D deep learning model is used to perform deep feature mining on preprocessed one-dimensional time-series data. The original ResNet50 model is adapted to one-dimensional time-series data, and a temporal attention mechanism is added to strengthen the weight allocation of fault-related features. The model outputs a 128-dimensional deep feature vector. To address the scarcity of industrial field fault samples, a combination of transfer learning and data augmentation is adopted. The model is first pre-trained on a public fault dataset, and then fine-tuned with a small number of field samples. At the same time, time stretching and noise addition are used to expand the sample and improve the model's generalization ability.
[0041] Sub-step S33: Spatiotemporal graph feature fusion
[0042] A spatiotemporal graph of the device sensor topology is constructed, where the graph nodes are the monitoring points of key components of the device. The connection relationships of the edges are predefined based on the physical topology of the device, the transmission path and the coupling relationship of the physical field. The edge weights are dynamically updated through node feature mutual information and attention mechanism. Shallow and deep features are mapped to graph nodes, and spatiotemporal graph neural network (ST-GNN) is used for feature aggregation. In the spatial dimension, the propagation path and causal relationship of anomalies across components and sensors are captured, and in the temporal dimension, the state evolution trajectory is tracked. Finally, a comprehensive fault feature vector that integrates multi-source information and has both spatial correlation and temporal evolution characteristics is output, which completely breaks down the information silos of sensor data.
[0043] Step S4: Equipment operating status assessment and fault identification
[0044] Based on the integrated fault feature vector, the cloud server adopts the process of "digital twin dynamic baseline health assessment - fault type identification - degradation trend prediction" to comprehensively assess the equipment's operating status, accurately identify fault types and severity, and predict remaining useful life (RUL).
[0045] The specific implementation details are as follows:
[0046] Sub-step S41: Health assessment based on digital twin dynamic baseline
[0047] During the equipment health break-in period, normal operating data across the entire operating range is collected to train a normalized health representation model based on a deep variational autoencoder and a Gaussian mixture model, constructing a dynamic digital twin baseline for the equipment. This baseline can dynamically generate the normal probability distribution manifold and confidence boundary of features under the current operating condition based on real-time operating parameters, replacing the traditional fixed baseline. A fuzzy comprehensive evaluation method is used, combining the standardized residuals of real-time features and the dynamic baseline, and the comprehensive fault feature vector, to calculate the equipment health score (0-100 points), classifying the equipment operating status into four levels: normal state (85-100 points), sub-healthy state (70-84 points), abnormal state (40-69 points), and fault state (0-39 points).
[0048] Sub-step S42: Fault identification and classification
[0049] An improved LSTM neural network model is used for fault identification based on comprehensive fault feature vectors. The model incorporates dropout layers, batch normalization layers, and an attention mechanism to reduce overfitting and enhance attention to fault features. It can accurately identify various faults such as bearing wear, spindle imbalance, gear tooth breakage, winding insulation aging, pipeline leakage, and pressure vessel corrosion. The overall fault identification accuracy is no less than 99%, and the early fault identification accuracy is no less than 96%. Based on the severity, scope of impact, and health level of the fault, faults are classified into three levels: Level 1 faults (minor faults, corresponding to a sub-healthy state), Level 2 faults (general faults, corresponding to an abnormal state), and Level 3 faults (serious faults, corresponding to a faulty state).
[0050] Sub-step S43: Degradation Trend and Remaining Life Prediction
[0051] Based on the health sequence and comprehensive fault characteristics of continuous time windows, a probabilistic degradation prediction model is constructed that integrates physical degradation mechanism constraints and data-driven learning. The model outputs the evolution trend of equipment health index, remaining service life range and uncertainty boundary, providing data support for preventive maintenance decisions.
[0052] Step S5: Adaptive Hierarchical Early Warning and Collaborative Linkage
[0053] Based on the equipment operation status assessment results, fault identification results, and degradation trend prediction results, a "three-level progressive adaptive early warning" mechanism is constructed. Combined with dynamic threshold optimization, it can achieve accurate hierarchical early warning and realize multi-device collaboration and linkage with existing industrial systems.
[0054] The specific implementation details are as follows:
[0055] Sub-step S51: Three-level progressive adaptive early warning
[0056] Combining digital twin dynamic baselines and fault classification, a three-level progressive early warning system is established. The early warning thresholds at each level are dynamically optimized using a beta distribution self-learning algorithm and L1 trend filtering technology. This allows for adaptive adjustment based on changes in equipment operating conditions, aging levels, and lifecycle stages, completely resolving the problem of rigid fixed thresholds.
[0057] (1) Level 1 Deviation Warning (Yellow Warning): When the health of a certain sub-dimension continues to deviate from the dynamic baseline confidence boundary, but the overall health and fault characteristics are not significantly abnormal, it is triggered. It corresponds to Level 1 fault. The warning is triggered by the monitoring platform pop-up window and the mobile APP push. The processing flow is to strengthen monitoring and daily inspection.
[0058] (2) Level II event rate warning (orange warning): When the occurrence rate and energy intensity of micro-point mutation events at key monitoring points show a statistically significant upward trend, or when the overall health enters an abnormal range, it is triggered. Corresponding to the Level II fault, an on-site audible and visual alarm is added, requiring maintenance personnel to arrive on-site within 2 hours to handle the situation.
[0059] (3) Three-level cascaded coupling early warning (red warning): When the health of multiple dimensions across physical fields and components deteriorates synchronously, the spatiotemporal graph model identifies the fault cascade propagation mode, or the equipment enters a fault state, it is triggered. For the corresponding three-level fault, a new person in charge is notified via SMS, and the shutdown process is immediately triggered. The maintenance personnel arrive at the scene within 30 minutes to handle the situation.
[0060] Sub-step S52: Collaborative Linkage Mechanism
[0061] (1) Multi-equipment collaborative monitoring: Establish a correlation model of related equipment in the production line. When a certain equipment triggers an early warning, the operating status of upstream and downstream related equipment is automatically monitored, the risk of chain failure is identified, and a linkage early warning is triggered to avoid production interruption;
[0062] (2) Industrial system linkage and docking: It is connected with the production scheduling system, equipment operation and maintenance management system and safety management system. When the three-level warning is triggered, it automatically sends a shutdown request to the production scheduling system, pushes fault information and generates maintenance work orders to the operation and maintenance management system, and sends a safety warning to the safety management system to achieve coordinated response of the whole system.
[0063] Step S6: Fault tracing and maintenance suggestion generation
[0064] When the equipment triggers a level 2 or higher warning, the system performs fault source analysis based on the fault identification results, spatiotemporal map features, and multi-source historical data to determine the root cause and propagation path of the fault, and generates targeted maintenance suggestions to achieve a closed-loop connection between early warning and operation and maintenance.
[0065] The specific implementation details are as follows:
[0066] Sub-step S61: Fault source analysis
[0067] A three-level tracing mechanism is constructed, consisting of "spatiotemporal graph propagation path analysis + causal graph backtracking + data backtracking": ① By using the abnormal propagation weights of the spatiotemporal graph neural network, the propagation path and scope of the fault from the root component to related components are determined; ② A fault causal graph is established based on the fault type to clarify the causal relationship between the fault and monitoring parameters and components; ③ Multi-source monitoring data from 12 hours before the fault occurred are backtracked to locate the fault start time and initial abnormal parameters, and the root cause and cause of the fault are accurately determined by combining the causal graph and propagation path.
[0068] Sub-step S62: Maintenance suggestion generation
[0069] Based on the fault tracing results, combined with the equipment's historical maintenance records, equipment manuals, and spare parts inventory information, a full-process maintenance suggestion is generated, which includes maintenance priority, detailed steps, required spare parts and tools, personnel configuration, estimated maintenance time, and safety precautions. This suggestion is automatically pushed to the maintenance personnel's APP and maintenance management system to help maintenance personnel quickly complete fault handling. After the maintenance is completed, the maintenance results can be entered, providing data support for the system's self-optimization.
[0070] Step S7: Full System Adaptive Optimization and Online Self-Evolution of the Model
[0071] Based on equipment operation data, fault identification results, and maintenance feedback, a full-process adaptive optimization mechanism is constructed. At the same time, online incremental self-evolution of the model is achieved through edge-cloud collaboration to ensure long-term stable operation of the system and adapt to changes in the status of the equipment throughout its entire life cycle.
[0072] The specific implementation details are as follows:
[0073] Sub-step S71: Adaptive optimization of full-process parameters
[0074] Regularly perform adaptive optimization of data acquisition parameters, preprocessing parameters, feature extraction methods, and early warning thresholds: increase the sampling frequency for fault-prone areas and decrease the sampling frequency for stable operating areas; adjust filtering and outlier identification parameters according to on-site interference; optimize feature extraction dimensions and fusion weights based on fault identification results; and dynamically optimize early warning thresholds based on equipment aging, operating condition changes, and false alarm feedback to adapt to changes in equipment status.
[0075] Sub-step S72: Online incremental self-evolution of the model
[0076] A closed-loop self-evolution mechanism is constructed, encompassing "Operational Feedback - Sample Return - Incremental Update - Secure Deployment": ① The system periodically collects fault samples, false alarm samples, and samples of new normal operating conditions labeled by operations and maintenance personnel to construct an incremental training set; ② An incremental learning strategy combining knowledge distillation and elastic weight consolidation is adopted to perform lightweight updates on the dynamic baseline model, feature extraction model, and fault identification model without forgetting the original fault identification capabilities; ③ Adversarial training is conducted on false alarm samples to improve the model's anti-interference ability; federated learning is used to integrate optimization experience from similar devices to improve the model's generalization ability; ④ After the updated model passes the shadow mode verification, it is deployed to the edge gateway for hot switching, achieving online self-evolution throughout the model's lifecycle and solving the problem of model performance degradation. The adaptive optimization cycle can be set according to the device's operating conditions, typically one month, but can be shortened to 15 days for high-fault devices.
[0077] Step S8: Data Storage and Visualization
[0078] The cloud server uses a distributed database (Hadoop) to classify and store raw data, feature data, status assessment results, early warning information, maintenance records, etc., supporting at least one year of historical data storage. A multi-terminal visualization monitoring platform is built to display equipment operating status, monitoring parameters, health status, early warning information, and fault tracing results in the form of line charts, heat maps, and 3D scatter plots. It supports access from computers, mobile phones, and tablets, and has functions such as historical data playback, report generation, and fault tracing query. It realizes a transparent display from macro-health overview to micro-signal mechanism, making it convenient for operation and maintenance personnel and managers to grasp the equipment status in real time.
[0079] Compared with the prior art, this invention has the following outstanding substantive features and significant progress, and possesses extremely high inventiveness, novelty and practicality:
[0080] 1. A qualitative breakthrough in early fault identification capability and a significant improvement in monitoring accuracy: Through the micro-point mutation capture mechanism, weak transient signals in the fault initiation stage that cannot be identified by existing technologies are captured. Combined with three-level feature extraction and deep fusion of spatiotemporal maps, the accuracy of early fault identification is no less than 96%, and the overall fault identification accuracy is no less than 99%. Through adaptive preprocessing, the data signal-to-noise ratio is improved to over 28dB, completely solving the problem of data distortion under strong interference environment, and the monitoring accuracy far exceeds that of existing technologies.
[0081] 2. Completely solves the pain point of rigid thresholds, significantly improving the accuracy and timeliness of early warnings: Through the digital twin operating condition adaptive dynamic baseline and three-level progressive early warning system, combined with risk-driven dynamic threshold optimization, it perfectly adapts to the state drift caused by changing operating conditions and equipment aging. The false early warning rate is controlled below 1.5%, and the missed early warning rate is controlled below 0.8%, which is far superior to existing technologies; the hierarchical early warning mechanism and collaborative linkage process ensure timely handling of faults and avoid the escalation of faults.
[0082] 3. Advanced architecture with strong industrial applicability and practicality: The edge-cloud collaborative architecture balances the computing power requirements of real-time on-site early warning and deep cloud analysis, with edge inference latency of less than 50ms, meeting the real-time requirements of industrial sites; the full-process solution is compatible with all types of industrial equipment, including rotating machinery, stationary equipment, and electrical equipment, and can be widely used in various industrial fields without large-scale modification of existing systems; the fault tracing and maintenance suggestion generation function improves maintenance efficiency by more than 40%, significantly reducing operation and maintenance costs and production losses.
[0083] 4. Possesses full lifecycle self-evolution capability and excellent long-term stability: Through an online self-evolution mechanism of incremental learning, adversarial training, and federated updates, it solves the industry pain points of static solidification and rapid performance degradation of existing models, extending the model performance stability period by more than 3 times; the full-process adaptive optimization mechanism eliminates the need for manual periodic parameter adjustments, significantly reducing the workload of operation and maintenance, and ensuring the continuous and stable operation of the system throughout the entire lifecycle of the equipment.
[0084] 5. Complete closed-loop solution, broad protection scope, and high patent value: This invention constructs a complete closed-loop technology system from data acquisition to model self-optimization, covering all aspects of equipment monitoring and early warning. The core innovations are all combinations of solutions not disclosed in existing technologies, and are highly non-obvious. At the same time, the solution is highly versatile, adaptable to all types of industrial equipment and all industrial fields, with a broad patent protection scope and extremely strong commercial value and exclusivity. Attached Figure Description
[0085] Figure 1 This is a flowchart illustrating the overall process of the method of the present invention.
[0086] Figure 2 This is a schematic diagram of the edge-cloud collaborative data acquisition and processing architecture of the present invention;
[0087] Figure 3 This is a schematic diagram of the multi-source data preprocessing and micro-point mutation capture process of the present invention;
[0088] Figure 4 This is a schematic diagram illustrating the principle of spatiotemporal graph feature fusion and fault tracing in this invention. Detailed Implementation
[0089] The present invention will be further described in detail below with reference to specific embodiments. This embodiment uses the ISG100-160 centrifugal water pump from the petrochemical industry as the monitoring object. This pump is used in the crude oil transportation process and operates in a high-temperature, high-humidity, and dusty environment. Long-term continuous operation is prone to failures such as bearing wear, impeller corrosion, pump body leakage, and motor overload. The method of the present invention is used to achieve real-time monitoring and fault early warning. The specific implementation process is as follows:
[0090] I. Implementation Preparation
[0091] Hardware Deployment and Sensor Selection: Based on the pump structure and fault-prone areas, a multi-source sensor array is deployed, including 4 piezoelectric vibration sensors, 5 PT100 temperature sensors, 1 Hall current sensor, 1 voltage transformer, 2 pressure sensors, 1 flow sensor, 1 noise sensor, 1 ultrasonic emission sensor, 1 humidity sensor, and 1 dust concentration sensor. All sensors are designed to be waterproof, dustproof, and corrosion-resistant. An industrial-grade edge computing gateway and high-speed data acquisition card are deployed, and a cloud server (configured with an Intel Xeon E5-2690 CPU, 32GB of memory, and an NVIDIA RTX 3090 graphics card) is built to establish a multi-terminal visual monitoring platform, which is integrated with the plant's production scheduling, operation and maintenance management, and safety management systems.
[0092] Model pre-training and baseline construction: The improved ResNet50-1D model, LSTM fault recognition model, and spatiotemporal graph neural network were pre-trained using the CWRU motor fault dataset and the XJTU-SY gearbox fault dataset; normal operating data under all working conditions were collected during the one-month healthy break-in period after the water pump overhaul to train the digital twin dynamic baseline model and complete the construction of the normalized health representation model under working conditions; system parameters were initialized, including a high-frequency sampling frequency of 500Hz for bearings and main shaft, an ultrasonic signal sampling frequency of 25.6kHz, an adaptive optimization period of one month, and initial thresholds for each level of early warning.
[0093] II. Specific Implementation Steps
[0094] Edge-cloud collaborative data acquisition: Sensors collect multi-source data in real time, including pump vibration, temperature, current, pressure, ultrasonic emission, and process parameters. The data acquisition card performs A / D conversion and preliminary filtering. The IEEE1588v2 protocol is used to achieve multi-channel data time synchronization, and the transmission delay is controlled within 80ms. Data is cached locally when the network is interrupted and automatically retransmitted after recovery. After receiving the data, the edge gateway performs preprocessing and lightweight real-time inference, and synchronously transmits the preprocessed data to the cloud server.
[0095] Adaptive preprocessing and micro-point mutation capture: The edge gateway performs operating condition segmentation on the raw data, dividing it into steady-state, transition, and transient segments; outliers and noise are removed through a three-level data cleaning process, improving the data signal-to-noise ratio to 28dB; three-dimensional operating condition space normalization is performed on the variable operating condition data to eliminate the influence of load and speed fluctuations; micro-point mutation detection is performed on high-frequency vibration and ultrasonic signals to extract four-dimensional event feature vectors to form an event stream, and trend features are extracted from slowly changing signals to form a slowly changing feature stream.
[0096] Three-level feature extraction and spatiotemporal graph fusion: The cloud server receives the preprocessed data and first extracts shallow features and micro-point mutation statistical features of vibration, electrical, and process parameters; it then mines deep features using an improved ResNet50-1D model and outputs a 128-dimensional feature vector; a spatiotemporal graph of the pump sensor topology is constructed, with nodes including 10 monitoring points such as the motor drive end bearing, the pump front end bearing, the motor winding, and inlet and outlet pressures. Edges are defined based on the transmission path and heat conduction relationship. Feature aggregation is completed through a spatiotemporal graph neural network, and a comprehensive fault feature vector integrating multi-source information is output.
[0097] Condition assessment and fault identification: Based on the digital twin dynamic baseline, the real-time health score of the water pump is calculated to classify the operating status level; the fault type is identified by the improved LSTM model, and the identification accuracy of 99.3% is achieved for 8 common faults such as bearing wear, impeller corrosion, and pump body leakage, with an early fault identification accuracy of 96.2%; the remaining service life of the water pump is predicted based on the health sequence to provide support for maintenance decisions.
[0098] Tiered early warning and collaborative linkage: A three-tiered progressive early warning mechanism is adopted, with dynamic thresholds optimized monthly based on pump operating conditions and aging levels. When slight wear occurs in the pump's front-end bearing, the incidence of ultrasonic micro-point mutation events increases significantly, triggering a level-two orange alert. On-site audible and visual alarms are activated simultaneously, and maintenance personnel arrive within 2 hours to inspect and resolve the fault after replenishing lubricating oil. When a level-three bearing fracture fault is detected, a red alert is immediately triggered, automatically sending a shutdown request to the production scheduling system, starting the backup pump, generating a maintenance work order for the maintenance system, and requiring maintenance personnel to arrive within 30 minutes to handle the situation, thus preventing crude oil leaks and production interruptions.
[0099] Fault tracing and maintenance suggestion generation: For bearing wear faults, through spatiotemporal diagram propagation path analysis, combined with cause-effect diagram and 12-hour historical data backtracking, the root cause of the fault was located as insufficient lubrication oil supply due to lubrication system blockage. The propagation path was "lubrication system blockage → insufficient lubrication oil → bearing wear → increased vibration → motor current fluctuation". Based on the tracing results, maintenance suggestions including maintenance steps, required spare parts, and personnel configuration were generated and pushed to the maintenance personnel's APP, improving maintenance efficiency by 42%.
[0100] Adaptive optimization and model self-evolution: The system performs full-process parameter optimization once a month. For the front-end bearings that are prone to failure, the sampling frequency is increased from 500Hz to 600Hz, and the filtering parameters are optimized to improve anti-interference ability. Based on the fault samples and false alarm samples labeled by the operation and maintenance personnel, the incremental learning strategy is used to update the model. The robustness of false alarm scenarios is optimized through adversarial training. The updated model is then distributed to the edge gateway after being verified by shadow mode. The model performance stability period is extended by more than 3 times.
[0101] Visualization: The visualization platform displays pump operating parameters, health scores, and early warning information in real time. It shows the distribution of micro-point mutation events through a 3D scatter plot and the status of each monitoring point through a heat map. It supports historical data playback and fault tracing query. Maintenance personnel can view equipment status and early warning information at any time via mobile phone.
[0102] III. Implementation Results Verification
[0103] This embodiment, after six months of continuous operation and verification, has achieved the following significant results: the signal-to-noise ratio after data preprocessing is stable above 28dB, the overall fault identification accuracy is 99.3%, the early fault identification accuracy is 96.2%, and the early bearing wear faults are identified 7-10 days earlier than traditional solutions; the false alarm rate is 1.2%, and the missed alarm rate is 0.6%, which is far lower than the industry average; the average fault handling time is shortened by 60%, maintenance efficiency is improved by 42%, and production losses are reduced by approximately 580,000 yuan within six months; the online self-evolution mechanism of the model ensures that the system operates without performance degradation over a long period of time, eliminating the need for manual periodic parameter adjustments, reducing operation and maintenance costs by 35%, and fully achieving the expected design goals.
[0104] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for online monitoring and fault early warning of safe operation of industrial equipment, characterized in that: Includes the following steps: S1. Equipment monitoring point layout and edge-cloud collaborative data acquisition: Based on the structural characteristics and fault-prone parts of industrial equipment, multi-source sensor arrays are deployed to build an edge-cloud collaborative multi-source data acquisition system, which synchronously collects multi-source parameters of equipment operation and completes low-latency transmission. S2. Multi-source data adaptive preprocessing and micro-point mutation capture: Perform operating condition segmentation and three-level data cleaning on the collected raw data, complete the normalization processing of variable operating condition data, and extract the micro-point mutation event stream representing early faults and the feature stream reflecting the gradual change trend of macro state. S3. Multi-dimensional fault feature extraction and deep fusion of spatiotemporal graph: shallow feature extraction, deep feature mining and spatiotemporal graph aggregation and fusion are performed on the preprocessed data to obtain a comprehensive fault feature vector that integrates multi-source information; S4. Equipment operating status assessment and fault identification: Based on the digital twin-driven adaptive dynamic baseline, the equipment health assessment is completed, the fault type and severity level are accurately identified, and the equipment degradation trend and remaining service life are predicted. S5. Adaptive hierarchical early warning and collaborative linkage: Construct a three-level progressive adaptive early warning system, dynamically optimize the early warning thresholds at each level, and realize multi-device collaborative monitoring and linkage response with existing industrial systems; S6. Fault tracing and maintenance suggestion generation: When a level 2 or higher warning is triggered, the root cause and propagation path of the fault are located through the three-level tracing mechanism, and maintenance suggestions for the whole process are generated in combination with equipment operation and maintenance data. S7. Full-system adaptive optimization and online self-evolution of the model: Based on equipment operation data and maintenance feedback, complete the adaptive optimization of parameters throughout the entire process, and realize online self-evolution of the model throughout its entire life cycle through incremental learning; S8. Data storage and visualization: Classify and store data throughout the entire process, build a multi-terminal visualization monitoring platform, and realize transparent display of equipment operating status.
2. The method for online monitoring and fault early warning of safe operation of industrial equipment according to claim 1, characterized in that: In step S1, the collected monitoring parameters cover four major categories: mechanical parameters, electrical parameters, environmental parameters, and process parameters. Monitoring points are set up for equipment fault-prone areas, and multiple types of sensors are simultaneously deployed at the same key location. The IEEE 1588v2 precision time protocol is used to achieve millisecond-level synchronization of multi-channel data. The edge gateway is responsible for data preprocessing and lightweight real-time inference, while the cloud server is responsible for deep feature mining, model training, and global optimization.
3. The method for online monitoring and fault early warning of safe operation of industrial equipment according to claim 1, characterized in that: In step S2, the operating condition segmentation divides the original data into steady-state, transition, and transient segments, and adopts differentiated preprocessing strategies for different segments; the three-level data cleaning uses a combination of the 3σ criterion and the isolated forest algorithm to identify outliers, Kalman filtering and wavelet packet transform to remove noise, and the LSTM time series prediction algorithm to fill in missing values; a three-dimensional operating condition reference space of load-speed-ambient temperature is constructed to complete the operating condition normalization, and the micro-point mutation event flow of high-frequency signals and the slow-changing trend feature flow of slowly changing signals are extracted respectively.
4. The method for online monitoring and fault early warning of safe operation of industrial equipment according to claim 1, characterized in that: In step S3, shallow feature extraction extracts the corresponding time domain, frequency domain, statistical and trend features for different types of monitoring parameters and dual feature streams, respectively. Deep feature mining employs an improved ResNet50-1D model adapted to one-dimensional time-series data, adds a temporal attention mechanism, and combines transfer learning and data augmentation to enhance the model's generalization ability; spatiotemporal graph feature fusion constructs a spatiotemporal graph of the device sensor topology based on the device's physical topology and coupling relationship, and uses a spatiotemporal graph neural network to complete feature aggregation.
5. The method for online monitoring and fault early warning of safe operation of industrial equipment according to claim 1, characterized in that: In step S4, the digital twin dynamic baseline is constructed based on the normal operating data of the equipment during the health break-in period under all operating conditions. It adopts a deep variational autoencoder and a Gaussian mixture model to dynamically generate the normal probability distribution and confidence boundary of the features under the current operating conditions. Based on the health score, the equipment operating status is divided into four levels, and fault classification and remaining service life prediction are completed simultaneously.
6. The method for online monitoring and fault early warning of safe operation of industrial equipment according to claim 1, characterized in that: In step S5, the three-level progressive adaptive early warning system includes a first-level deviation early warning, a second-level event rate early warning, and a third-level cascaded coupling early warning. The early warning thresholds at each level are dynamically optimized using a beta distribution self-learning algorithm and L1 trend filtering technology. The collaborative linkage mechanism includes collaborative monitoring and early warning of upstream and downstream related equipment, as well as linkage and docking with production scheduling, operation and maintenance management, and safety management systems.
7. The method for online monitoring and fault early warning of safe operation of industrial equipment according to claim 1, characterized in that: In step S6, the three-level tracing mechanism includes spatiotemporal diagram anomaly propagation path analysis, fault cause-effect diagram backtracking, and historical data backtracking before the fault occurred, accurately locating the root cause, propagation path, and cause of the fault; maintenance suggestions are generated by combining equipment historical maintenance records, equipment manuals, and spare parts inventory information, covering the core elements of the entire maintenance process.
8. The method for online monitoring and fault early warning of safe operation of industrial equipment according to claim 1, characterized in that: In step S7, the full-process parameter adaptive optimization dynamically adjusts the parameters for data acquisition, preprocessing, feature extraction, and early warning thresholds based on the equipment operating status and fault identification effect. The model is built online through self-evolution, constructing a closed-loop mechanism of "operation and maintenance feedback - sample return - incremental update - secure deployment". It adopts an incremental learning strategy that combines knowledge distillation and elastic weight consolidation, and completes model updates by combining adversarial training and federated learning. After verification in shadow mode, it is deployed to the edge gateway.
9. The method for online monitoring and fault early warning of safe operation of industrial equipment according to claim 2, characterized in that: In step S1, a sampling frequency of no less than 25.6kHz is set for the ultrasonic emission signal, the edge computing gateway has a local data caching capability of no less than 72 hours, and the data transmission delay is controlled within 100ms.
10. The method for online monitoring and fault early warning of safe operation of industrial equipment according to claim 1, characterized in that: In step S8, a distributed database is used to complete the classification and storage of data throughout the entire process, supporting at least one year of historical data backtracking; the multi-terminal visualization platform supports access from multiple terminals and has functions for historical data playback, report generation, and fault tracing query.