A Method for Online Monitoring and Anomaly Early Warning of Particulate Matter in Industrial Processes Based on Electrostatic Induction and Deep Learning
By combining an electrostatic induction sensor array with temperature and humidity compensation and a deep learning algorithm, the real-time performance and environmental interference problems of traditional industrial particulate matter monitoring methods have been solved. This has enabled high-precision monitoring of particulate matter concentration and size distribution, as well as anomaly warning, thereby improving the level of intelligence in industrial processes.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANDONG UNIV OF SCI & TECH
- Filing Date
- 2026-05-12
- Publication Date
- 2026-07-17
Smart Images

Figure CN122171411B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of industrial process monitoring and deep learning technology, specifically to a method for online monitoring and early warning of anomalies in industrial processes based on electrostatic induction and deep learning. Background Technology
[0002] With the rapid development of industrial production, particulate matter emitted from coal-fired boilers, cement kilns, and metallurgical sintering processes has become a significant source of air pollution, posing a serious threat to the ecological environment and human health. Traditional industrial particulate matter monitoring methods mainly include offline gravimetric analysis, beta-ray diffraction, and optical scattering. These methods have significant limitations in practical applications. Specifically, while offline gravimetric analysis is the standard method, its sampling and analysis cycle is long, making it unsuitable for real-time monitoring. Beta-ray diffraction, although capable of continuous monitoring, suffers from high equipment costs, complex maintenance, and insufficient sensitivity to low concentrations of particulate matter. Optical scattering is easily affected by environmental humidity, particulate matter color, and refractive index, with measurement errors significantly increasing under high humidity and high concentration conditions.
[0003] Existing electrostatic induction monitoring methods utilize the induced current generated by a sensor after particulate matter is charged to achieve concentration detection. These methods offer advantages such as simple structure, fast response speed, and insensitivity to optical properties. However, single electrostatic induction signals are easily affected by factors such as ambient temperature and humidity, and uneven particle size distribution, and it is difficult to simultaneously achieve concentration monitoring and particle size distribution inversion. Furthermore, traditional electrostatic induction signal processing methods are mostly based on simple feature statistics and linear regression, lacking the ability to extract nonlinear features under complex operating conditions. This limits monitoring accuracy and robustness, and the lack of an effective anomaly warning mechanism makes it impossible to promptly identify abnormal particulate matter emissions in industrial processes.
[0004] Traditional industrial particulate matter monitoring methods suffer from poor real-time performance, weak environmental resistance, limitations of single-parameter monitoring, and a lack of anomaly warnings. These issues easily lead to delayed monitoring data and untimely detection of emissions exceeding standards, thus impacting environmental regulation and industrial process optimization. The industry urgently needs an intelligent monitoring method that integrates multi-feature information from electrostatic induction, possesses environmental adaptability, enables integrated concentration monitoring and particle size distribution inversion, and incorporates anomaly detection algorithms to improve the real-time performance, accuracy, and reliability of industrial process particulate matter emission monitoring.
[0005] Therefore, it is necessary to develop an online monitoring and anomaly early warning method for particulate matter in industrial processes based on electrostatic induction and deep learning. By combining the nonlinear feature extraction capability of deep learning algorithms with data fusion technology, it is possible not only to achieve high-precision real-time monitoring of particulate matter concentration and particle size distribution, but also to promptly identify emission anomalies. This will provide technical support for environmental supervision and industrial process optimization, and ensure the green and sustainable development of atmospheric environmental quality and industrial production. Summary of the Invention
[0006] The purpose of this invention is to provide an online monitoring and anomaly early warning method for industrial process particulate matter based on electrostatic induction and deep learning, so as to solve the technical problems of poor real-time performance, weak environmental anti-interference ability, large limitations of single parameter monitoring, and lack of anomaly early warning in the existing traditional industrial particulate matter monitoring methods.
[0007] This invention provides a method for online monitoring and early warning of particulate matter in industrial processes based on electrostatic induction and deep learning, comprising the following steps: The multimodal sensing system synchronously collects electrostatic induction current signals and temperature and humidity data of particulate matter in the industrial process. The multimodal sensing system includes an electrostatic induction sensor array and a temperature and humidity sensor. The electrostatic induced current signal of the particulate matter is denoised to obtain a denoised signal. The denoised signal is then subjected to temperature and humidity compensation processing based on temperature and humidity data to obtain a preprocessed signal. Extract the time-domain features, frequency-domain features, and time-frequency-domain features of the preprocessed signal; The extracted time-frequency domain features are used to form a time-frequency image, which is then input into a trained deep learning model. The deep learning model then outputs particulate matter concentration monitoring results and particle size distribution inversion results simultaneously. The system collects time-domain features, frequency-domain features, and time-frequency-domain features under normal operating conditions, processes them to obtain an early warning threshold, processes the real-time time-domain features, frequency-domain features, and time-frequency-domain features to calculate a real-time anomaly score, and compares the real-time anomaly score with the early warning threshold to identify abnormal particulate matter emission and trigger an early warning.
[0008] Furthermore, the deep learning model is a convolutional neural network combined with a bidirectional long short-term memory network (CNN-Bi-LSTM); the structure of the CNN-Bi-LSTM model consists of an input layer, a convolutional layer, a pooling layer, a Bi-LSTM layer, a fully connected layer, and an output layer connected in sequence.
[0009] Furthermore, when training the CNN-Bi-LSTM model, the loss function used is a weighted sum of the mean square error loss of particulate matter concentration prediction and the KL divergence loss of particle size distribution inversion. The Adam optimizer is used to minimize the loss function. The initial learning rate is 0.001, which decays to 0.5 every 50 training cycles. The batch size is 32. An early stopping mechanism is enabled during training to prevent overfitting.
[0010] Furthermore, the noise reduction process employs an improved wavelet thresholding method. Specifically, it involves performing multi-scale analysis of the electrostatic induced current signal of the particulate matter through wavelet decomposition and adaptive thresholding of the high-frequency noise coefficient. The threshold function used in the improved wavelet thresholding method is: ; in, Let be the wavelet coefficients after thresholding, denoted as the denoised signal, and j be the scaling parameter, corresponding to the frequency coarsening level in the wavelet transform. These are translation parameters, corresponding to spatial / temporal displacements at the same scale; These are the original wavelet coefficients. For adaptive threshold, As a regulating factor, The value range is 0.5-1.5. It is a symbolic function.
[0011] Furthermore, the temperature and humidity compensation processing is implemented through a radial basis function (RBF) neural network, specifically: the RBF neural network takes the collected temperature and humidity data as input and outputs compensation coefficients that act on the noise reduction signal.
[0012] Furthermore, the warning threshold is calculated based on the isolated forest algorithm; the warning threshold for: ; in, This represents the average score for abnormalities under normal operating conditions. The standard deviation of outlier scores. As the early warning coefficient, The value range is 2-3. When the real-time anomaly score exceeds... An alert is triggered at any time.
[0013] Furthermore, the electrostatic induction sensor array includes a plurality of electrostatic induction sensors distributed in space; the plurality of electrostatic induction sensors are arranged at a preset interval along the particulate flow direction, and / or arranged at different radial positions on a cross section perpendicular to the particulate flow direction, so that sensors located at different spatial positions generate distinguishable electrostatic induction signals to the same particulate, the distinguishability including signal strength difference or time phase difference. The step of extracting the time-frequency domain features of the preprocessed signal further includes: obtaining a spatial difference feature vector to characterize the particle size distribution based on the spatial response differences of multiple electrostatic induction sensors to the same particulate matter; the deep learning model further uses the spatial difference feature vector as input to output the particle size distribution inversion result.
[0014] Furthermore, the distinguishable electrostatic induction signal is transmitted through electrostatic induction current. and the current response vector of the electrostatic induction sensor array It is determined that the electrostatic induced current is [not specified]. , used to determine the current of a single electrostatic induction sensor; in, The total charge of particulate matter. This refers to the number of particulate matter per unit volume. The average charge per particle. The velocity of the particles. The effective sensing area of the sensor; The current response vector of the electrostatic induction sensor array for: ; The current response vector of the electrostatic induction sensor array This is used to combine the current intensities of different electrostatic induction sensors into a current response vector to characterize the spatial response differences between sensors. in, For the first The sensor for the first The sensitivity coefficient of a single particulate matter For the first The charge of each particulate matter For the number of sensors, This represents the number of particulate matter.
[0015] Furthermore, the extraction of the time-domain features includes the following features: Average current ; in, For the first Current values at each sampling point This represents the number of sampling points; Peak current ; Root mean square value ; The frequency domain features extracted include the following features: main frequency ; in, The power spectral density of the signal; Power spectral entropy ; in, For the first Normalized power spectrum values at each frequency point This represents the number of frequency points. The extraction of the time-frequency domain features includes the following features: ; in, For the first Energy of each frequency band Let j be the total energy and j be the number of decomposition levels. Let be the wavelet packet energy entropy.
[0016] Furthermore, the particulate matter concentration monitoring results and particle size distribution inversion results synchronously output by the deep learning model are calculated based on the correlation formula between particulate matter charge and particle size, wherein the correlation formula is: ; in, Particle size The particulate matter charge, The vacuum permittivity, Let be the relative permittivity of the particulate matter. denoted as electric field strength.
[0017] The beneficial effects of this invention are as follows: 1. It overcomes the limitations of traditional single-sensor monitoring, significantly improving the robustness of monitoring under complex working conditions: Traditional electrostatic induction monitoring methods often rely on a single sensor, whose signals are easily interfered with by factors such as ambient temperature and humidity, and uneven particle size distribution, making it difficult to guarantee monitoring accuracy and stability. This invention employs an electrostatic induction sensor array combined with a multimodal sensing mechanism with temperature and humidity compensation, forming cross-validation from both spatial sensing and environmental compensation dimensions, effectively solving the aforementioned problems.
[0018] On the one hand, the electrostatic induction sensor array acquires distinguishable sensing signals generated by the same particle at different spatial locations by using multiple sensors arranged at preset intervals along the particle flow direction, as well as sensors arranged at different radial positions on the same cross section. This spatial arrangement allows particles of different sizes to generate distinguishable responses in signal intensity and phase on each sensor in the array due to differences in charge and motion characteristics. This provides rich multidimensional spatial information for particle size distribution inversion, breaking through the technical bottleneck that a single sensor can only sense the total amount but cannot distinguish particle size.
[0019] On the other hand, addressing the strong interference of ambient temperature and humidity on electrostatic induction signals—especially under high humidity conditions (relative humidity > 70%), where particulate matter charge efficiency decreases by 30%-60% and a water film easily forms on the sensor surface leading to signal drift—this invention establishes a temperature and humidity compensation model based on a radial basis function (RBF) neural network. This model takes real-time temperature and humidity data as input and outputs dynamic compensation coefficients acting on the noise-reducing signal, adaptively eliminating interference from environmental factors. Implementation examples demonstrate that within a wide range of relative humidity variation (25%-85%), the baseline drift of the compensated signal is controlled within ±3%, significantly enhancing the environmental robustness of the preprocessed signal.
[0020] 2. An innovative CNN-Bi-LSTM deep learning model is introduced to achieve high-precision simultaneous monitoring of particulate matter concentration and particle size distribution: Traditional electrostatic induction signal processing methods are mostly based on simple feature statistics and linear regression. They are severely lacking in the ability to extract and model the complex nonlinear mapping relationship between particulate matter concentration and multidimensional electrostatic features, as well as the deep correlation between particle size distribution and signal spectral features. The improved CNN-Bi-LSTM deep learning model proposed in this invention organically combines the local feature extraction capability of convolutional neural networks (CNNs) with the temporal dependency capture capability of bidirectional long short-term memory networks (Bi-LSTMs), completely solving this technical challenge.
[0021] Specifically, the convolutional layer uses learnable convolutional kernels to extract local features from the time-frequency feature map of the electrostatic induction signal and the spatial difference feature vector formed by the spatial response differences of the sensor array, automatically mining potential feature patterns related to particulate matter size and concentration. Building on this, the Bi-LSTM layer fully utilizes the temporal continuity of particulate matter emission states in industrial processes, capturing long-term and short-term dependencies in the feature sequence through forward and backward hidden state updates, further enhancing the model's adaptability to dynamically changing operating conditions.
[0022] During model training, this invention employs an innovative joint loss function, weighted by combining the mean squared error loss of concentration prediction with the KL divergence loss of particle size distribution inversion (weights of 0.6 and 0.4 respectively). Through multi-task joint optimization, the model can share underlying feature representations during learning, mutually enhancing the learning performance of each task. Testing shows that this method achieves an average relative error of less than 4.2% for concentration monitoring and improves particle size distribution inversion accuracy by more than 31.1% compared to traditional single-sensor linear inversion methods. This truly realizes high-precision synchronous online monitoring of concentration and particle size distribution, filling a gap in existing technology in this area.
[0023] 3. An improved wavelet threshold denoising method is embedded to effectively suppress noise interference in industrial settings: Industrial sites are characterized by significant electromagnetic interference and mechanical vibration noise. Directly using raw electrostatic induction signals for analysis would severely reduce monitoring accuracy. This invention addresses the characteristics of electrostatic induction signals, such as pronounced particulate pulse features and noise primarily concentrated at high frequencies, by designing an improved wavelet thresholding denoising method. This method employs an adaptive thresholding strategy, where the threshold λ is determined by both the noise standard deviation estimate and the signal length, dynamically adjusting according to the actual noise level of each layer's coefficients. Simultaneously, the designed nonlinear threshold function achieves a smooth transition between hard and soft thresholds, preserving the amplitude of the main particulate pulse features while avoiding ringing and pseudo-Gibbs phenomena in the reconstructed signal. In the embodiment, by adjusting the factor a=1.0, high-frequency electromagnetic interference is effectively filtered out while perfectly preserving the key pulse signals caused by particulate matter passing through the sensor, providing a high-quality data foundation for subsequent feature extraction and model input.
[0024] 4. A multi-level anomaly early warning mechanism based on the Isolation Forest algorithm is established, which has high early warning accuracy and fast response speed: Traditional industrial process particulate matter monitoring generally lacks effective anomaly early warning mechanisms, making it difficult to identify emission anomalies in a timely manner, leading to environmental pollution incidents. This invention establishes an anomaly detection and early warning system independent of deep learning models based on the Isolation Forest algorithm. The Isolation Forest algorithm leverages the "few and distinct" characteristics of anomaly samples, constructing multiple isolated trees by randomly selecting features and segmentation values. This allows for the rapid isolation of anomalies in a high-dimensional feature space without the need for pre-labeling of anomaly samples, making it particularly suitable for real-world scenarios where industrial site anomaly data is scarce.
[0025] Building upon this foundation, the present invention further establishes a scientific three-tiered early warning threshold system: by statistically analyzing the distribution characteristics of anomaly scores under normal operating conditions, and using the mean and standard deviation as a basis, thresholds of 2.0, 2.5, and 3.0 times the standard deviation are respectively set as the thresholds for Level 1 yellow alert, Level 2 orange alarm, and Level 3 red alarm. This tiered early warning mechanism can sensitively identify anomalies of different degrees while effectively controlling the false alarm rate. Remote continuous 72-hour testing results in the embodiment show that the anomaly early warning accuracy reaches 95.5%, the false alarm rate is less than 4%, and the average early warning response delay is only 8.2 seconds, fully meeting the stringent requirements of industrial applications for real-time performance and accuracy.
[0026] 5. Form an integrated technological closed loop of "sensing-processing-monitoring-early warning" to promote intelligent control of particulate matter emissions from industrial processes: Another significant advantage of this invention lies in its integration of multimodal sensing, signal preprocessing, feature extraction, deep learning monitoring, and anomaly early warning into a complete technological closed loop. Through an embedded acquisition system, high-speed synchronous acquisition of electrostatic induced current at 10kHz and real-time reading of temperature and humidity data at 1Hz are achieved. After preprocessing and feature extraction, the CNN-Bi-LSTM model performs online synchronous inversion of concentration and particle size, while the isolated forest model performs parallel anomaly detection. When an anomaly is detected, the system automatically triggers an early warning and completes local data storage and cloud upload, providing real-time, complete, and traceable data support for environmental supervision and industrial process optimization.
[0027] In summary, this invention systematically solves a series of technical problems in traditional industrial particulate matter monitoring methods through the synergistic innovation of multiple technical means, including multi-sensor spatial array design, environmental adaptive compensation, deep learning multi-task modeling, improved signal denoising, and intelligent anomaly detection. These problems include poor real-time performance, weak environmental anti-interference ability, difficulty in simultaneously achieving concentration monitoring and particle size distribution inversion, and lack of effective anomaly early warning mechanisms. This invention significantly improves the accuracy, robustness, and intelligence level of particulate matter emission monitoring under complex operating conditions, providing core technical support for the control of particulate matter emissions in industrial processes and the optimization of green production. Attached Figure Description
[0028] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort, wherein: Figure 1 This is a block diagram of the overall structure of the industrial particulate matter online monitoring and anomaly early warning system in an embodiment of the present invention; Figure 2 This is a schematic diagram of the electrostatic induction signal processing and deep learning monitoring process in an embodiment of the present invention. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only for explaining the invention and are not intended to limit the invention; that is, the described embodiments are merely some embodiments of the invention, and not all embodiments. The components of the embodiments of the invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0030] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0031] It should be noted that relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0032] The features and performance of the present invention will be further described in detail below with reference to embodiments.
[0033] Example: This embodiment provides a method for online monitoring and anomaly early warning of particulate matter in industrial processes based on electrostatic induction and deep learning. The overall system architecture is as follows: Figure 1 As shown, the signal processing and monitoring process is as follows: Figure 2 As shown. The method includes the following steps: I. Establishing a multimodal sensing system and synchronous data acquisition: 1. Design and layout of electrostatic induction sensor array: The electrostatic induction sensor array in this embodiment employs four ring electrodes. The electrode substrate is 316L stainless steel, and the surface is coated with a 0.1mm thick aluminum oxide insulating layer by plasma spraying to prevent leakage caused by particulate matter deposition. The effective sensing area of each ring electrode (i.e., the projected area of the inner wall of the electrode facing the airflow direction) is A = 1.0 × 10⁻⁶. -3 m 2 This area value was calculated based on an electrode inner diameter of 10mm and an axial length of 30mm (projected area = inner diameter × length = 10mm × 30mm ≈ 3 × 10). -4 m 2 To further improve sensing sensitivity, this embodiment adopts a two-ring parallel structure, achieving an equivalent sensing area of 1.0 × 10⁻⁶. -3 m 2 ).
[0034] The spatial arrangement of the four electrodes on the exhaust duct (100mm inner diameter) is as follows: Along the direction of particulate flow (axial direction), electrodes are installed sequentially at equal intervals of 0.2 m to form an axial array, numbered E1, E2, E3, and E4. This axial spacing was determined through fluid dynamics simulation (COMSOL software): simulation results show that when the wind speed is 5 m / s, the flow field and charge distribution of 10 μm particles stabilize within approximately 0.15 m after passing one electrode. Using 0.2 m ensures the measurement independence between each electrode.
[0035] At the same axial position (E2 position), three radially distributed electrodes are added to the pipe cross-section perpendicular to the flow direction, located at the center of the pipe, at 1 / 2 radius, and near the wall (5 mm from the inner wall), respectively, to sense the radial distribution differences of particulate matter concentration and particle size at different locations on the same cross-section.
[0036] Therefore, the entire array contains 6 electrostatic induction sensors (4 axial + 2 additional radial, or replacing one of the original axial electrodes on the E2 section, for a total of 4 + 2 = 6), and this embodiment uses 6 signal channels. In this embodiment, it is more generally described as "arranged at preset intervals and / or arranged at different radial positions," and this embodiment is a combination of both.
[0037] 2. Selection and placement of temperature and humidity sensors: The Sensirion SHT35 digital temperature and humidity sensor was selected, with a temperature measurement accuracy of ±0.1°C and a humidity accuracy of ±1.5%RH. The sensor was placed at the end of the array (0.1m after E4) to measure the ambient temperature and humidity across the entire sensing segment. It was connected to the acquisition system via an I²C bus, with a sampling rate set to 1Hz. This sampling rate was chosen because the time constant of temperature and humidity changes in industrial processes is typically on the order of minutes, and 1Hz is sufficient to capture these environmental changes.
[0038] 3. Dust generation and exhaust system: The system employs an SAG-410L dust aerosol generator, capable of uniformly dispersing fine dust (particle size range 0.1-10μm) conforming to ISO12103-1A2 standards. By adjusting the feed belt speed and nozzle pressure, particulate matter concentrations from 0-1000 mg / m³ can be achieved. The exhaust system uses a variable frequency centrifugal fan, allowing for frequency adjustment to control the air velocity within the duct within the range of 0.5-15 m / s. A high-efficiency filter is installed at the fan outlet to protect the environment.
[0039] 4. Embedded data acquisition system design: The core of the data acquisition system uses an STM32F407 microcontroller with an ARM Cortex-M4 core and a main frequency of 168MHz. The electrostatic induction current signal conditioning circuit includes: I / V Conversion: An Analog Devices ADA4530-1 electrometer-level operational amplifier with a feedback resistor Rf=10MΩ is used to achieve current-to-voltage conversion with a gain of 10. 7 V / A.
[0040] Anti-aliasing filter: Second-order active low-pass filter with a cutoff frequency fcut=5kHz.
[0041] ADC Sampling: A 24-bit Σ-Δ analog-to-digital converter (ADS1256) is used with a sampling rate of 10kHz, and four channels are sampled synchronously (four along the axis). The other two axial signals are acquired through another ADS1256 chip to ensure strict synchronization of the six signals. The 10kHz sampling rate is chosen because the pulse spectrum generated by particulate matter passing through the sensor is mainly in the 0-3kHz range, and according to the Nyquist sampling theorem, 10kHz can completely preserve the signal characteristics.
[0042] Temperature and humidity data are read at a frequency of 1Hz via the STM32F407's I²C interface. All data is transferred to an external SRAM (2MB) cache via DMA and uploaded to the host computer in real time via TCP / IP protocol through the on-chip Ethernet controller (W5500). Simultaneously, the original data is stored on a local SD card as a backup.
[0043] 5. Reference Instruments and Truth Acquisition: To provide the labeled data required for model training, this embodiment uses two reference instruments: Particulate matter concentration reference value: The TEOM1405 oscillating balance particulate matter monitor from Thermo Fisher Scientific, USA, is used. This instrument has high measurement accuracy and can be used as a standard.
[0044] Particle size distribution reference values: The TSI APS3321 aerodynamic particle size spectrometer was used to measure the particle size range of 0.5-20μm. In this embodiment, only 32 logarithmic intervals in the 0.5-10μm range are used as labels.
[0045] 6. Simulated operating conditions and data acquisition experiments: By adjusting the feed rate and air velocity, data was collected under three environmental conditions: dry (RH=25%), moderate (RH=50%), and high humidity (RH=80%), combining different concentrations (0-1000 mg / m³, with intervals of approximately 50 mg / m³, for a total of 21 concentration points) and different particle size distributions (single-peak, double-peak, etc.). Data was collected continuously for 30 minutes under each condition, totaling approximately 1500 samples (each sample representing 1 minute of data). Simultaneously, abnormal states (such as sudden concentration exceeding limits, equipment malfunctions, etc.) were randomly inserted under different operating conditions to construct an anomaly detection dataset. A total of approximately 50,000 valid samples were obtained, each containing 6 electrostatic signals (10 kHz), temperature and humidity data (1 Hz), and the corresponding reference concentration and particle size distribution. This dataset comprehensively covers various industrial processes and disturbances, serving as the foundation for subsequent model training and validation.
[0046] II. Signal Preprocessing and Temperature and Humidity Compensation: 1. Improved wavelet thresholding for noise reduction: For each electrostatic induction signal, a frame with a length of N=1024 sampling points (approximately 0.1024 seconds) is extracted from the original 10kHz sampling sequence and denoised.
[0047] Wavelet basis and decomposition level selection: By comparing the mean square error of various wavelet bases such as db2, db4, and sym4 in the reconstructed signal and the original noiseless signal, the db4 wavelet showed the best preservation of particulate pulse signals. Therefore, the Daubechies4 (db4) wavelet was selected. The decomposition level J=5, which was determined based on the sampling rate and the highest frequency of the useful signal. At a sampling rate of 10kHz, the highest frequency is 5kHz. After 5 levels of decomposition, the lowest frequency approximation coefficient corresponds to a frequency band of 0-156Hz, which matches the signal frequency range of particulate induced pulses.
[0048] Adaptive threshold calculation: For each layer of high-frequency coefficients (j=1-5), the noise standard deviation σ is first estimated. The robust estimation method proposed by Donoho is adopted: the median of the absolute values of the first layer of high-frequency coefficients is taken and divided by 0.6745 as the estimated value of σ: σ=median(|w j,k |) / 0.6745; Then calculate the adaptive threshold λ: λ = σ·sqrt(2log(Nj)); Where Nj is the length of the coefficient of this layer.
[0049] Improved threshold function processing: The wavelet coefficients w are processed using the following improved threshold function: if|w|>=λ: ŵ=sign(w)*(|w|-a*λ / exp((|w|-λ) / λ)); else: ŵ=0.
[0050] In this embodiment, the adjustment factor a = 1.0. This value was determined experimentally by searching within the range of 0.8-1.2 with a step size of 0.1, after comprehensively balancing the signal-to-noise ratio and pulse amplitude preservation after denoising a known standard signal. When a = 1.0, it can effectively remove high-frequency noise and control the peak amplitude loss caused by particulate matter to within 5%.
[0051] Reconstruction: Perform inverse wavelet transform on the processed high-frequency coefficients of each layer and the low-frequency approximation coefficients of the 5th layer to obtain the denoised signal sequence s_denoised.
[0052] 2. Temperature and humidity compensation based on RBF neural network: Ambient temperature and humidity affect electrostatic signals in two main ways: increased humidity leads to a decrease in air insulation resistance, increasing the sensor's leakage current and resulting in a rise in the signal baseline; humidity also reduces the particulate matter charging efficiency, causing a decrease in the overall signal amplitude. To compensate for this effect, an RBF neural network compensation model is established.
[0053] Network structure: 2 nodes in the input layer (temperature T, relative humidity RH), 40 radial basis function neurons in the hidden layer, and 1 node in the output layer (compensation coefficient α). The number of hidden layer nodes (40) was determined during training using 10-fold cross-validation to minimize the mean squared error on both the training and validation sets.
[0054] Training data generation: Under controlled temperature and humidity conditions in the laboratory (temperature 20-45°C, humidity 25%-85%), the output baseline voltage Vb(T,RH) of each sensor was recorded in clean air free of particulate matter. Standard conditions were defined as T=25°C, RH=30%, and the standard baseline was Vb. For the denoised signal s_denoised under arbitrary temperature and humidity, the goal is to bring its baseline back to the standard baseline; therefore, the target compensation coefficient α_target=Vb / Vb(T,RH) (assuming the signal baseline is proportional to the leakage current and inversely proportional to the insulation resistance). A total of 200 sets of baseline data under different temperature and humidity conditions were collected as training samples for the RBF network.
[0055] Training method: K-means clustering algorithm is used to extract 40 cluster centers from 200 groups (T, RH) as the centers of the hidden layer basis functions. The width σ of the basis function is taken as the mean of the distances between each center and its nearest neighbor. Then, the pseudo-inverse method is used to calculate the output layer weights to obtain the RBF network model.
[0056] Online compensation: In actual use, the current temperature and humidity are input, and the network feedforward calculates α. Then the preprocessed signal after temperature and humidity compensation is: s_pre=α×s_denoised; In a 72-hour test of this compensation method, the signal deviation under high humidity conditions was reduced from ±18% before compensation to within ±3%.
[0057] III. Multi-dimensional Feature Extraction: The compensated signal sequence s_pre is processed in frames, with a frame length of N = 1024 points (0.1024 seconds) and a frame shift of 512 points (approximately 0.05 seconds) to ensure temporal continuity. The following features are extracted from each frame:
[0058] 1. Temporal characteristics: The mean current μ: μ = (1 / N)Σs_pre[n], reflects the average intensity of charge passing through within this frame. Typical value range: 10. -12 -10 -6 A (The voltage value after I / V conversion is divided by the feedback resistor to calculate the current).
[0059] Peak current I_peak: I_peak = max(|s_pre[n]|), reflecting the maximum charged pulse when a single particle or a group of particles passes through. Value range: 10. -11 -10 -5 A.
[0060] Root mean square value (RMS): RMS = sqrt((1 / N)Σ(s_pre[n]) 2 This reflects the total energy of the signal. Its value ranges approximately 10. -12 -10 -6 A.
[0061] 2. Frequency domain characteristics: After applying a Hanning window to the frame signal, perform a 1024-point FFT to obtain the amplitude spectrum of the signal, and then calculate the power spectral density PSD[k]=|X[k]|², k=0-511.
[0062] The dominant frequency f_dom is defined as follows: f_dom = argmax(PSD(f)), with a search range of 0-5000Hz. The Nyquist frequency at 10kHz sampling is 5kHz, hence the frequency range is 0-5000Hz. The dominant frequency of typical industrial electrostatic induction signals occurs between 50-2000Hz.
[0063] Power spectral entropy H_pse: First, normalize the power spectrum: p_k = PSD[k] / ΣPSD[k], then calculate H_pse = -Σp_klog2(p_k), k = 0-511. This entropy value reflects the complexity of the spectrum; the wider the particle size distribution, the higher the entropy value usually is. Typical values range from 2.0 to 8.0 bits.
[0064] 3. Time-frequency domain characteristics: A three-level wavelet packet decomposition was performed using the db2 wavelet basis, resulting in eight terminal frequency bands. The energy of each frequency band was calculated. E_j=Σ(w_packet,j[m])², j=1…8.
[0065] Total energy E_total = ΣE_j.
[0066] Wavelet packet energy entropy H_wpe = -Σ(E_j / E_total)log2(E_j / E_total). This feature effectively reflects the energy distribution of different frequency components and is closely related to the particle size distribution. Typical values range from 1.0 to 6.0 bits.
[0067] 4. Spatial differences: For 6 (or 4) signals, the aforementioned time-domain and time-frequency domain features are extracted respectively. To characterize the spatial response differences of the sensor array, the following approach is constructed: The time-domain features (mean, peak, RMS) of the four axial sensors and their respective wavelet packet energy entropy, totaling 4×4=16 dimensions, are combined into a single vector.
[0068] The same features of the three radial sensors on the same cross section are combined to form another set of vectors.
[0069] This embodiment actually uses the time-domain features, spectral features, and wavelet packet features of 6 signals to form a 24-dimensional initial feature vector. Then, principal component analysis (PCA) or direct concatenation is used to compress it into a 12-dimensional new feature vector as input to the deep learning model (this compression step is not necessary, but aims to reduce the amount of computation, and is obtained by retaining 95% of the variance in the training set through PCA).
[0070] To utilize sequence information, features from 10 consecutive frames (approximately 1 second) are taken to form a time-frequency feature map X (10×12 dimensions). At the same time, the spatial position response differences of each frame (such as the peak ratio between the radial sensor and the central sensor) are also used to form a spatial difference feature map (10×3). Both are input into the model.
[0071] IV. Construction, Training, and Online Inference of CNN-Bi-LSTM Deep Learning Models: 1. Model Structure: The model structure details are as follows: Input branch 1: Time-frequency feature map, size 10×12×1 (time step × number of features × channels).
[0072] Input branch 2: Spatial difference feature map, size 10×3×1.
[0073] Convolutional layers (each branch): Use one-dimensional convolution (Conv1D), with a kernel size of 3×1, a stride of 1, and zero padding to maintain the same size. The number of kernels is 32, and the activation function is ReLU. A max-pooling layer follows the convolution, with a pooling size of 2 and a stride of 2. Therefore, the output size of branch 1 is 5×32, and the output size of branch 2 is 5×32.
[0074] Fusion layer: Flatten and concatenate the outputs of the two branches to obtain a feature vector with a length of 5×32×2=320.
[0075] Reshaping into a sequence: In order to input Bi-LSTM, the fused features are reshaped into a time step of T=5, with a feature dimension of 64 per step (because the convolution extracts abstract features, the original time step number is used directly).
[0076] Bi-LSTM layer: The number of unidirectional hidden units dh=64, and the dimensions of the forward hidden state h_f and the backward hidden state h_b are each 64, so the output after concatenation at each time step is 128-dimensional. The output of the last time step (concatenation of forward and backward) is taken as the representation of the entire sequence, i.e., a 128-dimensional vector. The number of hidden units of 64 is chosen based on empirical trials and validation set results; a smaller number of units can prevent overfitting and maintain low computational latency.
[0077] Fully connected output layer: Concentration regression branch: Dense(1), activation function linear. Particle size distribution branch: Dense(32), activation function softmax, corresponding to 32 logarithmic particle size intervals. The total number of model parameters is approximately 150K, and the single inference time on STM32F407 is approximately 45ms, which meets the real-time requirements.
[0078] 2. Model Training: Loss function: Joint loss function is used. Loss = λ1 * L_MSE + λ2 * L_KL; Wherein, the mean square error of concentration prediction L_MSE=(1 / B)Σ(c_true,m–c_pred,m)², and B is the batch size.
[0079] Particle size distribution KL divergence L_KL=Σ_{i=1}^{32}p_true(i)*log(p_true(i) / (p_pred(i)+ε)),ε=1e-7 to prevent log(0).
[0080] The weights λ1=0.6 and λ2=0.4 were obtained through grid search on the independent validation set (λ1 ranges from 0.3 to 0.9, with a step size of 0.1). This combination enables both the relative error of concentration and the KL divergence of particle size to reach a relatively good level.
[0081] Training settings: Optimizer: Adam, initial learning rate η = 0.001.
[0082] Learning rate scheduling: Every 50 epochs, η decays to 0.5 times its original value.
[0083] Batch size B=32 (limited by GPU memory and training stability).
[0084] The total number of epochs is 200, and an early stopping mechanism is adopted: if the validation loss does not decrease for 15 consecutive epochs, training is stopped and the optimal model parameters are restored.
[0085] Data partitioning: 50,000 samples were randomly divided into a training set (70%, 35,000 samples), a validation set (15%, 7,500 samples), and a test set (15%, 7,500 samples).
[0086] Training results: On the test set, the mean relative error (MRE) of concentration prediction was 4.2%, calculated as MRE = (1 / M)Σ|c_pred–c_true| / c_true × 100%. The average KL divergence between the particle size distribution inversion and the true distribution measured by APS was 0.273, while the KL divergence of the traditional single sensor + linear regression method was 0.396, representing a 31.1% improvement in accuracy. These results fully demonstrate the effectiveness of this model.
[0087] 3. Online reasoning process: During online operation, the embedded system continuously collects data. A sliding window (10 frames) is used every second to construct input features, and a trained CNN-Bi-LSTM model is invoked for forward computation. The model outputs concentration values and a 32-dimensional particle size probability vector. To achieve smoothing, the concentration values from the five most recent outputs (within 2.5 seconds) are averaged, and the particle size probability vectors are averaged as well. The final output is a concentration value (mg / m³) and a particle size distribution (a 32-interval probability histogram). These results are displayed locally and also fed into an anomaly warning module.
[0088] V. Anomaly Detection and Three-Level Early Warning Based on Isolated Forest: 1. Anomaly detection model training: In the first month of equipment operation, only data under normal operating conditions were collected (based on annotations by process experts or judgment by reference instruments indicating no abnormalities). Approximately 10,000 normal samples were accumulated (each sample is a feature vector within a 10-second window). The feature vectors used the 24-dimensional features extracted in the above steps (including time domain, frequency domain, time-frequency domain, and spatial difference features).
[0089] The feature data is standardized using StandardScaler (mean = 0, variance = 1). Then, the Isolation Forest model is trained: Parameter settings: n_estimators = 100 (number of trees), max_samples = 256 (random sampling size for each tree), contamination = 'auto' (automatically estimated based on the data. In fact, for the model trained on normal data, contamination is set to 0.05, that is, 5% of the boundary points are expected). These parameters are determined based on experience and adjusted on the held-out validation set to ensure that the false positive rate for normal data is less than 3%.
[0090] After the model is trained, the anomaly scores are calculated using all normal samples. The calculation method of the anomaly score s is referred to the original paper, and its value range is (0, 1). The closer it is to 1, the more anomalous it indicates.
[0091] The mean μ_s and standard deviation σ_s of the anomaly scores of all normal samples are statistically calculated. In this embodiment, μ_s = 0.52 and σ_s = 0.12 (the data is statistically based on actual tests).
[0092] 2. Early warning thresholds and response mechanisms: Set three-level early warning thresholds (in this embodiment, the warning coefficients κ1 = 2.0, κ2 = 2.5, κ3 = 3.0):<( Th1 = μ_s + 2.0 * σ_s = 0.52 + 0.24 = 0.76; Th2 = μ_s + 2.5 * σ_s = 0.52 + 0.30 = 0.82; Th3 = μ_s + 3.0 * σ_s = 0.52 + 0.36 = 0.88; During online operation, the current feature vector is calculated every 10 seconds, and the anomaly score s_new is obtained through the Isolation Forest. Then the following logic is executed: If s_new < Th1: The status is normal, no action.
[0093] If Th1 ≤ s_new < Th2: Trigger a first-level yellow prompt. The yellow audible and visual alarm in the local control cabinet is activated, a prompt window pops up on the upper computer interface, and the abnormal data is recorded but production is not interrupted.
[0094] If Th2 ≤ s_new < Th3: Trigger a second-level orange alarm. In addition to the above actions, the system automatically sends a text message alarm to the mobile phones of the preset management personnel through the 4G module.
[0095] If s_new ≥ Th3: Trigger a Level 3 red alarm. In addition to audible and visual alarms and SMS messages, the system outputs dry contact signals via relays, which can interlock and control the induced draft fan speed or bypass valves to achieve automatic emergency emission reduction. At this time, the system automatically packages all data, including the original 6 channels of static electricity signals, temperature and humidity data, feature sets, and anomaly scores from the previous 10 minutes (total 20 minutes), stores them on a local industrial SD card (capacity 128GB, capable of cyclic recording for 1 year), and simultaneously uploads them to the environmental monitoring cloud platform via a 4G network.
[0096] 3. Anomaly detection and early warning performance: During a 72-hour continuous test, several abnormal operating conditions were simulated (such as a sudden doubling of concentration, sensor malfunction, and fan shutdown), resulting in a total of 21 abnormal events. The system detected 20 of these, achieving a warning accuracy rate of 95.5%, with only two false alarms (both at the humidity abrupt change boundary), a false alarm rate of 4%. The average delay from the occurrence of an abnormal event to the system issuing a warning was 8.2 seconds (including feature window calculation delay), significantly faster than manual inspection. This demonstrates that the warning mechanism is characterized by its fast response, accuracy, and practicality.
[0097] VI. Overall System Testing and Performance Verification: After integrating all the above modules, a continuous 72-hour test was conducted in a real industrial simulation environment. Test conditions: simulated fly ash from a coal-fired boiler, with concentrations randomly fluctuating between 50-800 mg / m³, temperature between 20-45°C, and relative humidity between 25%-85%. The system operated automatically throughout the entire process.
[0098] Test results: Average relative error of concentration monitoring: 3.8%-4.6% (different operating conditions), overall MRE 4.1%.
[0099] Particle size distribution inversion accuracy (KL divergence compared with APS3321): average 0.28, which is 31% better than traditional methods.
[0100] Signal stability after temperature and humidity compensation: baseline drift is less than 3% (18% drift without compensation in traditional methods).
[0101] Anomaly warning accuracy: 95.5%, false alarm rate: 4%, average response delay: 8.2 seconds.
[0102] Overall system reliability: No crashes or data loss within 72 hours; embedded system CPU utilization was 65% and memory usage was 40%.
[0103] The above test results demonstrate that the method and system of this invention can effectively solve the problems of poor real-time performance, large environmental interference, inability to synchronously invert particle size, and lack of abnormal early warning in traditional particulate matter monitoring technology, and have extremely high industrial application value.
[0104] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for online monitoring and anomaly early warning of particulate matter in industrial processes based on electrostatic induction and deep learning, characterized in that, Includes the following steps: The multimodal sensing system synchronously collects electrostatic induction current signals and temperature and humidity data of particulate matter in the industrial process. The multimodal sensing system includes an electrostatic induction sensor array and a temperature and humidity sensor. The electrostatic induced current signal of the particulate matter is denoised to obtain a denoised signal. The denoised signal is then subjected to temperature and humidity compensation processing based on temperature and humidity data to obtain a preprocessed signal. Extract the time-domain features, frequency-domain features, and time-frequency-domain features of the preprocessed signal; The extracted time-frequency domain features are used to form a time-frequency image, which is then input into a trained deep learning model. The deep learning model then outputs particulate matter concentration monitoring results and particle size distribution inversion results simultaneously. The system collects time-domain features, frequency-domain features, and time-frequency-domain features under normal operating conditions, processes them to obtain an early warning threshold, processes the real-time time-domain features, frequency-domain features, and time-frequency-domain features to calculate a real-time anomaly score, and compares the real-time anomaly score with the early warning threshold to identify abnormal particulate matter emission and trigger an early warning. The electrostatic induction sensor array includes multiple electrostatic induction sensors distributed in space; the multiple electrostatic induction sensors are arranged at a preset interval along the particulate flow direction, and / or arranged at different radial positions on a cross section perpendicular to the particulate flow direction, so that sensors located at different spatial positions generate distinguishable electrostatic induction signals to the same particulate matter, the distinguishability including signal strength difference or time phase difference. The step of extracting the time-frequency domain features of the preprocessed signal further includes: obtaining a spatial difference feature vector to characterize the particle size distribution based on the spatial response differences of multiple electrostatic induction sensors to the same particulate matter; the deep learning model further uses the spatial difference feature vector as input to output the particle size distribution inversion result. The distinguishable electrostatic induction signal is transmitted through electrostatic induction current. and the current response vector of the electrostatic induction sensor array It is determined that the electrostatic induced current is [not specified]. , used to determine the current of a single electrostatic induction sensor; in, The total charge of particulate matter. This refers to the number of particulate matter per unit volume. The average charge per particle. The velocity of the particles. The effective sensing area of the sensor; The current response vector of the electrostatic induction sensor array for: ; The current response vector of the electrostatic induction sensor array This is used to combine the current intensities of different electrostatic induction sensors into a current response vector to characterize the spatial response differences between sensors. in, For the first The sensor for the first The sensitivity coefficient of a single particulate matter For the first The charge of each particulate matter For the number of sensors, This refers to the number of particulate matter. The particulate matter concentration monitoring results and particle size distribution inversion results synchronously output by the deep learning model are calculated based on the correlation formula between particulate matter charge and particle size. The correlation formula is as follows: ; in, Particle size The particulate matter charge, The vacuum permittivity, Let be the relative permittivity of the particulate matter. denoted as electric field strength.
2. The method according to claim 1, characterized in that, The deep learning model is a convolutional neural network combined with a bidirectional long short-term memory network (CNN-Bi-LSTM). The structure of the CNN-Bi-LSTM model consists of an input layer, a convolutional layer, a pooling layer, a Bi-LSTM layer, a fully connected layer, and an output layer connected in sequence.
3. The method according to claim 2, characterized in that, When training the CNN-Bi-LSTM model, the loss function used is the weighted sum of the mean squared error loss of particulate matter concentration prediction and the KL divergence loss of particle size distribution inversion. The Adam optimizer is used to minimize the loss function. The initial learning rate is 0.001, which decays to 0.5 every 50 training epochs. The batch size is 32. An early stopping mechanism is enabled during training to prevent overfitting.
4. The method according to claim 1, characterized in that, The noise reduction process employs an improved wavelet threshold noise reduction method. Specifically, it involves multi-scale analysis of the electrostatic induced current signal of the particulate matter through wavelet decomposition and adaptive threshold processing of the high-frequency noise coefficient. The threshold function used in the improved wavelet threshold noise reduction method is as follows: ; in, The wavelet coefficients after thresholding are denoted as the denoised signal, and j is the scale parameter, corresponding to the frequency coarsening in the wavelet transform; These are translation parameters, corresponding to spatial / temporal displacements at the same scale; These are the original wavelet coefficients. For adaptive threshold, As a regulating factor, The value range is 0.5-1.
5. It is a symbolic function.
5. The method according to claim 1, characterized in that, The temperature and humidity compensation process is implemented through a radial basis function (RBF) neural network. Specifically, the RBF neural network takes the collected temperature and humidity data as input and outputs compensation coefficients that act on the noise reduction signal.
6. The method according to claim 1, characterized in that, The warning threshold is calculated based on the isolated forest algorithm; the warning threshold for: ; in, This represents the average score for abnormalities under normal operating conditions. The standard deviation of outlier scores. As the early warning coefficient, The value range is 2-3. When the real-time anomaly score exceeds... An alert is triggered at any time.
7. The method according to claim 4, characterized in that, The extraction of the time-domain features includes the following features: Average current ; in, For the first Current values at each sampling point This represents the number of sampling points; Peak current ; Root mean square value ; The frequency domain features extracted include the following features: main frequency ; in, The power spectral density of the signal; Power spectral entropy ; in, For the first Normalized power spectrum values at each frequency point This represents the number of frequency points. The extraction of the time-frequency domain features includes the following features: ; in, For the first Energy of each frequency band For total energy, The number of decomposition layers, Let be the wavelet packet energy entropy.