Industrial sensor missing data filling method and system based on EEMD (ensemble empirical mode decomposition)

Through adaptive dynamic EEMD decomposition and adversarial generation network calibration, the modal aliasing and reconstruction error problems of the EEMD method in industrial sensor data loss filling are solved, and high-precision and real-time filling of missing data is achieved.

CN120524104AInactive Publication Date: 2025-08-22SICHUAN DIEMENG TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510659093.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-08-22
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing EEMD methods have problems such as modal aliasing, poor physical interpretability, reconstruction errors caused by static reconstruction strategies and computational redundancy and insufficient prediction caused by single missing mode processing methods in the filling of industrial sensor data.

Method used

Adaptive dynamic EEMD decomposition, dynamic adjustment of white noise amplitude based on Bayesian noise modeling, combined with missing mode classification and adversarial generation network (GAN) calibration signal, and optimized filling accuracy and real-timeness through dynamic weights.

Benefits of technology

Effectively suppress modal aliasing, improve data decomposition accuracy and physical interpretability, balance real-time and accuracy, ensure the reliability of the signal under time domain, frequency domain and physical constraints, and avoid error accumulation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120524104A_ABST
    Figure CN120524104A_ABST
Patent Text Reader

Abstract

The invention discloses an EEMD (ensemble empirical mode decomposition)-based industrial sensor missing data filling method and system, and relates to the technical field of sensor data processing.The method comprises the steps that firstly, self-adaptive dynamic EEMD is carried out, the amplitude of white noise is dynamically adjusted through Bayesian noise modeling, modal aliasing is inhibited, and a physically interpretable IMF component is generated; thirdly, based on missing mode classification, a divide-and-conquer strategy is adopted, dynamic weight fusion IMF components are combined, and filling precision and real-time performance are optimized; and finally, evaluating the authenticity of the reconstructed signal by using a GAN framework, dynamically adjusting a smoothing factor, and verifying the multi-scale consistency. According to the method, the accuracy and the physical interpretability of data are improved through adaptive dynamic EEMD decomposition, the accuracy and the real-time performance of missing data filling are optimized through a divide-and-conquer strategy, the reliability and the multi-scale consistency of signals are enhanced through the generative adversarial network, and therefore the limitation of a traditional method in a high-noise and non-stationary industrial environment is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sensor data processing technology, and in particular to a method and system for filling missing data of industrial sensors based on EEMD. Background Art

[0002] Industrial sensor data often suffers from data loss during the data collection process due to equipment failures, signal interference, or network transmission issues, directly impacting real-time monitoring and decision-making in industrial systems. Among traditional data infill methods, ensemble empirical mode decomposition (EEMD)-based techniques are widely used due to their ability to handle non-stationary signals. However, existing EEMD methods have significant drawbacks in industrial scenarios: First, the amplitude of white noise is fixed during the decomposition process and cannot adapt to the dynamically changing noise intensity, resulting in severe modal aliasing and reducing the physical interpretability of the IMF components. Second, a static weight distribution strategy is used during signal reconstruction, ignoring the dynamic contribution differences of each IMF component under different working conditions, which is prone to accumulating reconstruction errors under sudden or complex working conditions.

[0003] These problems limit the practical application of traditional methods in high-noise and non-stationary industrial environments.

[0004] Further research shows that the limitations of existing methods are specifically reflected in the following aspects: 1. Noise sensitivity: EEMD decomposition with a fixed noise amplitude is prone to over-decomposition or under-decomposition when the signal-to-noise ratio fluctuates, causing the IMF component to contain redundant noise or lose key features, affecting the subsequent filling accuracy; 2. Static reconstruction strategy: Fixed weights cannot capture local signal features (such as instantaneous energy changes in sudden segments), resulting in significant deviations between the reconstructed signal and the real data. This leads to serious error accumulation, especially in scenarios with long-term continuous loss. 3. Single treatment of missing patterns: Existing methods do not distinguish between random missing (short intervals) and continuous missing (long intervals), and uniformly adopt the same filling strategy. This leads to computational redundancy in short missing scenarios and insufficient prediction ability in long missing scenarios.

[0005] In order to solve the above-mentioned defects, a technical solution is now provided. Summary of the Invention

[0006] The purpose of the present invention is to solve the problems of modal aliasing and poor physical interpretability caused by the fixed white noise amplitude when filling missing industrial sensor data in the EEMD method in the prior art, as well as the reconstruction error and computational redundancy caused by the static weight distribution strategy and single missing mode processing method, and to propose a method and system for filling missing industrial sensor data based on EEMD.

[0007] The purpose of the present invention can be achieved through the following technical solutions: A missing data filling method for industrial sensors based on EEMD includes the following steps: S1: Adaptive dynamic EEMD decomposition, dynamically adjusting the white noise amplitude through Bayesian noise modeling, suppressing modal aliasing and generating physically interpretable IMF components; S2: IMF prediction and reconstruction, using a divide-and-conquer strategy based on missing pattern classification, combining dynamic weights to fuse IMF components, optimizing filling accuracy and real-time performance; S3: Real-time calibration of the generative adversarial network, using the GAN framework to evaluate the authenticity of the reconstructed signal, dynamically adjust the smoothing factor and verify multi-scale consistency to ensure the reliability of the signal under time domain, frequency domain and physical constraints.

[0008] Furthermore, the execution process of S1 is as follows: S1.1, Noise intensity estimation: Data volume segmentation, input sensor signal At a fixed sampling rate Receive, every The time buffer is a data block, and the sliding window length Dynamic selection: Based on calculate ,satisfy , each sliding step of the window ,guaranteed 50% overlap to cover the signal mutation; S1.2, Bayesian noise modeling: for the current window Internal data, calculate the local mean : ; Assume that noise obeys , the prior distribution is the inverse Gamma distribution , calculate the posterior probability: Where is the noise standard deviation, which obeys the inverse Gamma prior distribution ; is the data block of the current sliding window; For the window The original signal value of the sampling point; For the window The local mean of the sampling points; is the hyperparameter of the inverse Gamma distribution; is the sliding window length; Solving the maximum a posteriori estimate by gradient ascent ; S1.3, dynamic noise amplitude injection: according to Adjust the EEMD white noise amplitude: ,in is the amplitude of white noise injected in EEMD decomposition, is the optimal noise standard deviation obtained by Bayesian posterior estimation; is the empirical coefficient; right Perform EEMD decomposition: Add 100 groups of amplitudes White noise, generating IMF set ; Calculate the instantaneous frequency for each IMF Variance ,like , judged as aliasing mode, discarded and trigger local re-decomposition, only for the current window ;in For the The instantaneous frequency of each IMF component; is the variance of the instantaneous frequency, which is used to detect modal aliasing; is the sampling rate of the sensor signal.

[0009] Furthermore, the specific operation steps of S2 are as follows: S2.1. Missing area positioning: missing interval detection, real-time monitoring of input signals , when continuous When the data within the time is missing or exceeds the range, the missing interval is marked , extract the missing interval before and after Known data as context ;in Represent the start and end time of the missing interval respectively, is the length of the known data extracted before and after the missing interval, Represents the context dataset before and after the extracted missing interval; S2.2. Missing pattern classification: obtaining missing intervals , when the missing duration , it is judged as random missing, otherwise it is judged as continuous missing, where threshold for missing pattern classification; S2.3. A divide-and-conquer strategy based on missing pattern classification is adopted, including: When it is determined to be a random missing point, that is, a short interval missing point, it is initialized by neighboring interpolation: for each missing point in the missing interval , using bilinear interpolation: ,in It is the most recent valid time point before and after the missing point; Calculate the local variance of the interpolated signal. If it exceeds the normal fluctuation range, which is 1.5 times the standard deviation of the historical data, trigger the LSTM prediction correction; When it is determined to be missing continuously, a temporal attention bidirectional LSTM prediction is performed; S2.4. Combine the dynamic weights to fuse the IMF components.

[0010] Furthermore, the specific steps for performing temporal attention bidirectional LSTM prediction in S2.3 are as follows: Model loading and input construction: pre-trained bidirectional LSTM model, input layer dimension 64, hidden layer 128, attention layer 32, embedding, for each IMF component ,from Extract subsequences from , is the length of the extracted subsequence, Points, normalized and input into the model, Index for time points; Attention Mechanism Calculation: LSTM Hidden State With input Splicing, generating attention weights through the fully connected layer: ,in is the time step The attention weight, For bidirectional LSTM at time step The hidden state of is the time step Input features of is the weight matrix, and its dimensions are , ; is the trainable parameter vector and bias term, , ; is the activation function, is an activation function used to introduce nonlinearity; Weighted context vector , input the fully connected layer and output the predicted value ; Online loss calculation and optimization, loss function updated in real time: ; After predicting 10 missing intervals, the model parameters are fine-tuned using the data in the sliding window with a learning rate of 10 −4 ;in is the total loss function; is the mean square error; MF component and The correlation between them.

[0011] Furthermore, the specific operation steps of S2.4 are as follows: Contribution calculation and signal reconstruction: in the neighborhood of missing intervals , Seconds, the time-varying weight of each IMF is calculated in segments at 0.1 second intervals: ,in For the IMF at the time Contribution weight; For the The variance of the IMF; is the signal-to-noise ratio of the IMF, which is obtained by calculating the instantaneous energy ratio through Hilbert transform: ,in is the integration time interval, used to calculate the signal-to-noise ratio, ; For the filled IMF set according to Weighted summation is performed to obtain the final reconstructed signal: , the weight is updated every 0.1 seconds; To finally reconstruct the signal; is the residual term, which represents the signal part that cannot be explained by the IMF component; N is the total number of IMF components.

[0012] Furthermore, the missing pattern classification threshold is determined in S2.2 The specific steps are as follows: Initial threshold setting: Maintain a length of = 24-hour sliding window, recording the duration of all missing intervals Every time = 1 hour, calculate the quantile of missing duration in the window: , that is, choose a value that makes 95% of the missing intervals shorter than the current threshold; Dynamic adjustment: for each missing interval , extract the following features: Missing duration ; Volatility of data before and after: Calculate the standard deviation of the data 10 seconds before and after the missing interval ; Correlation decay: Calculate the Pearson correlation coefficient of the data 10 seconds before the missing interval ; Initial threshold The corrected calculation logic is as follows: ,in ; is the corrected missing pattern classification threshold; Online feedback optimization: For each filled missing interval, its physical constraints and multi-scale consistency are verified. If the filling fails and the original classification is a short missing, the threshold is lowered: If the filling fails and the original classification is a long deletion, the trigger threshold is raised: .

[0013] Furthermore, the specific operation steps of S3 are as follows: S3.1. GAN online inference framework: Generator G and Discriminator D are deployed. Generator structure: 3-layer convolutional network, input 1D signal, output calibration signal, convolution kernel size 5, stride 2, number of channels [16, 32, 1] per layer; Discriminator structure: 4-layer convolutional network, output authenticity probability, convolution kernel size 3, number of channels [32, 64, 128, 1]; pre-trained model is embedded in the real-time system, and calibration is triggered every 10 seconds of data received; S3.2. Adversarial calibration process: signal optimization and parameter adjustment, input reconstructed signal To generator G, output calibration signal ; Discriminator D is With real data Calculating confidence ; S3.3, Dynamic Smoothing Factor Adjustment: ,in ; Over-limit point , corrected to: ;in is the measuring range of the sensor signal; like , triggering generator online fine-tuning: freezing the discriminator parameters to Update G and iterate 5 times; is the loss function of the generator; S3.4. Multi-scale consistency verification: Calculation of original data and calibration signal Autocorrelation coefficients at three time scales: 1 second scale: calculate the autocorrelation with a lag of 10 points ; 10-second scale: calculate the autocorrelation of 100 points lag ; 1-minute scale: calculate the autocorrelation with a lag of 600 points ; like ,in, is the autocorrelation coefficient of the original data, If the autocorrelation coefficient of the calibration signal is not equal to the autocorrelation coefficient of the calibration signal, it is judged as inconsistent and the following operations are triggered: Adjust the IMF weights for the corresponding time period ,according to renew, is the difference ratio; re-execute the LSTM prediction and weight assignment in S2.

[0014] An industrial sensor missing data filling system based on EEMD, comprising: Dynamic noise adjustment and EEMD decomposition module dynamically adjusts the EEMD white noise amplitude based on Bayesian noise estimation to generate physically interpretable IMF components. It also suppresses aliasing through modal purity verification to improve decomposition robustness. The missing pattern classification and divide-and-conquer filling module is used to dynamically classify missing patterns based on missing duration, data volatility, and correlation, and combines neighbor interpolation with attention LSTM divide-and-conquer filling to balance real-time performance and accuracy. Dynamic weight fusion and signal reconstruction module, used to calculate the time-varying contribution of IMF components, weighted fusion fill results, update weights every 0.1 seconds, and optimize the physical consistency of the reconstructed signal; The adversarial generative network calibration module is used to optimize the reconstructed signal through the generator, evaluate the authenticity through the discriminator, dynamically adjust the smoothing factor and constrain the range, and iteratively optimize the signal fidelity by combining multi-scale autocorrelation tests; The real-time data stream management module is used to manage sensor data input, sliding window segmentation, historical data caching, and parallel computing scheduling to optimize processing delays.

[0015] Compared with the prior art, the present invention has the following beneficial effects: (1) The present invention uses adaptive dynamic EEMD decomposition and Bayesian noise modeling to dynamically adjust the white noise amplitude, effectively suppressing modal aliasing and generating physically interpretable IMF components. Compared with the traditional EEMD method with a fixed white noise amplitude, the present invention can better adapt to dynamically changing noise intensity and avoid over-decomposition or under-decomposition, thereby improving the accuracy of data decomposition and physical interpretability. (2) This invention adopts a divide-and-conquer strategy based on missing pattern classification and combines dynamic weights to fuse IMF components. For short-interval random missing, neighbor interpolation initialization is used; for long-interval continuous missing, temporal attention bidirectional LSTM prediction is used. The divide-and-conquer strategy can balance real-time performance and accuracy, avoiding the computational redundancy of traditional methods in short-interval missing scenarios and the prediction deficiency in long-interval missing scenarios, significantly improving the filling accuracy. (3) The present invention uses a generative adversarial network (GAN) framework to evaluate the authenticity of the reconstructed signal, dynamically adjusts the smoothing factor and verifies the multi-scale consistency; optimizes the reconstructed signal by the generator and the discriminator to evaluate the authenticity, combined with the multi-scale autocorrelation test to ensure the reliability of the signal in the time domain, frequency domain and physical constraints; the multi-scale consistency verification mechanism can effectively avoid the error accumulation of traditional methods in long-term continuous loss scenarios, and improve the overall fidelity of the signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to facilitate understanding by those skilled in the art, the present invention will be further described below with reference to the accompanying drawings; Figure 1 Flow chart of the method of the present invention. DETAILED DESCRIPTION

[0017] The following will clearly and completely describe the technical solutions of the present invention in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.

[0018] It should be understood that the terms “include” and “comprising” used in the specification and claims of the present disclosure indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0019] It should also be understood that the terminology used in this disclosure is for the purpose of describing specific embodiments only and is not intended to limit the disclosure. As used in this disclosure and the claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should be further understood that the term "and / or" as used in this disclosure and the claims refers to any and all possible combinations of one or more of the associated listed items, including and including these combinations.

[0020] like Figure 1 As shown, a method for filling missing data of industrial sensors based on EEMD includes the following steps: Step 1: Adaptive dynamic EEMD decomposition, dynamically adjusting the white noise amplitude through Bayesian noise modeling, suppressing modal aliasing and generating physically interpretable IMF components, improving decomposition robustness; Noise intensity estimation: Data volume segmentation: Input sensor signal At a fixed sampling rate Receive, every The time buffer is a data block, and the sliding window length Dynamic selection: Based on calculate ,satisfy , each sliding step of the window ,ensure 50% overlap to cover signal mutations; Bayesian noise modeling: for the current window Internal data, calculate the local mean : (5-point moving average); assuming the noise obeys , the prior distribution is the inverse Gamma distribution , calculate the posterior probability: Where is the noise standard deviation, which obeys the inverse Gamma prior distribution ; is the data block of the current sliding window; For the window The original signal value of the sampling point; For the window The local mean of the sampling points (5-point sliding average); are the hyperparameters of the inverse Gamma distribution, set to 1 and 0.5 respectively; is the sliding window length; the maximum a posteriori estimate is solved by the gradient ascent method , the calculation is completed within 10ms; Dynamic noise amplitude injection: According to Adjust the EEMD white noise amplitude: ,in is the amplitude of white noise injected in EEMD decomposition, is the optimal noise standard deviation obtained by Bayesian posterior estimation; is an empirical coefficient used to adjust the noise amplitude ratio, and ;right Perform EEMD decomposition: Add 100 groups of amplitudes White noise, generating IMF set ; Calculate the instantaneous frequency for each IMF Variance ,like , judged as aliasing mode, discarded and trigger local re-decomposition, only for the current window ;in For the The instantaneous frequency of each IMF component; is the variance of the instantaneous frequency, which is used to detect modal aliasing; is the sampling rate of the sensor signal.

[0021] Step 2: IMF prediction and reconstruction. Based on the missing pattern classification (short / long interval), a divide-and-conquer strategy (neighboring interpolation or LSTM prediction) is adopted. IMF components are fused with dynamic weights to optimize filling accuracy and real-time performance. Missing area positioning: missing interval detection, real-time monitoring of input signals , when continuous When the data within the time is NaN (missing data) or exceeds the range, the missing interval is marked , extract the missing interval before and after Known data as context ;in Represent the start and end time of the missing interval respectively, is the length of the known data extracted before and after the missing interval, Represents the context dataset before and after the extracted missing interval; Missing pattern classification: obtaining missing intervals , when the missing duration , it is judged as random missing, otherwise it is judged as continuous missing, where is the missing pattern classification threshold, and the specific determination process is as follows: Initial threshold setting: Maintain a length of = 24-hour sliding window, recording the duration of all missing intervals Every time = 1 hour, calculate the quantile of missing duration in the window: , that is, choose a value that makes 95% of the missing intervals shorter than the current threshold; Dynamic adjustment: for each missing interval , extract the following features: Missing duration ; Volatility of data before and after: Calculate the standard deviation of the data 10 seconds before and after the missing interval ; Correlation decay: Calculate the Pearson correlation coefficient of the data 10 seconds before the missing interval ; Initial threshold The corrected calculation logic is as follows: ,in ;when When the value is low (the correlation between the previous and next data is weak), increase the threshold appropriately to avoid misjudging long deletions as short deletions. is the corrected missing pattern classification threshold; Online feedback optimization: For each filled missing interval, verify its physical constraints (such as range) and multi-scale consistency. If filling fails and the original classification is short missing, the threshold is lowered: If the filling fails and the original classification is a long deletion, the trigger threshold is raised: .

[0022] When it is determined to be a random missing point, that is, a short interval missing point, it is initialized by neighboring interpolation: for each missing point in the missing interval , using bilinear interpolation: ,in The most recent valid time point before and after the missing point; calculate the local variance of the interpolated signal. If it exceeds the normal fluctuation range (1.5 times the standard deviation of historical data), trigger the LSTM prediction correction; When it is determined to be missing, the temporal attention bidirectional LSTM prediction is performed, and the process is as follows: Model loading and input construction: Pre-trained bidirectional LSTM model (input layer dimension 64, hidden layer 128, attention layer 32) is embedded, and each IMF component ,from Extract subsequences from , is the length of the extracted subsequence, Points, normalized and input into the model, Index for time points; Attention Mechanism Calculation: LSTM Hidden State With input Splicing, generating attention weights through the fully connected layer: ,in is the time step The attention weight, For bidirectional LSTM at time step The hidden state of is the time step Input features of is the weight matrix, and its dimensions are , ; is the trainable parameter vector and bias term, , ; is the activation function used to convert the attention score into a probability distribution, is an activation function used to introduce nonlinearity; weighted context vector , input the fully connected layer and output the predicted value ; Online loss calculation and optimization: Loss function is updated in real time: After predicting 10 missing intervals, the model parameters are fine-tuned using the data in the sliding window (learning rate 10 −4 );in is the total loss function; is the mean square error; MF component and the correlation between Dynamic Weight Fusion: Contribution Calculation and Signal Reconstruction: In the Neighborhood of Missing Intervals , Seconds, the time-varying weight of each IMF is calculated in segments at 0.1 second intervals: ,in For the IMF at the time Contribution weight; For the The variance of the IMF; is the signal-to-noise ratio of the IMF, which is obtained by calculating the instantaneous energy ratio through Hilbert transform: ,in is the integration time interval, used to calculate the signal-to-noise ratio, ; For the filled IMF set according to Weighted summation is performed to obtain the final reconstructed signal: , the weight is updated every 0.1 seconds; To finally reconstruct the signal; is the residual term, which represents the signal part that cannot be explained by the IMF component; N is the total number of IMF components.

[0023] Step 3: Real-time calibration of the generative adversarial network, using the GAN framework to evaluate the authenticity of the reconstructed signal, dynamically adjust the smoothing factor and verify multi-scale consistency to ensure the reliability of the signal under time domain, frequency domain and physical constraints; GAN online inference framework: Generator (G) and Discriminator (D) deployment. Generator architecture: 3-layer convolutional network (input 1D signal, output calibration signal), each layer has a convolution kernel size of 5, a stride of 2, and a number of channels in the range of [16, 32, 1]. Discriminator architecture: 4-layer convolutional network (output authenticity probability), a convolution kernel size of 3, and a number of channels in the range of [32, 64, 128, 1]. The pre-trained model is embedded in the real-time system, and calibration is triggered every 10 seconds of data received. Adversarial calibration process: signal optimization and parameter adjustment, input reconstructed signal To generator G, output calibration signal ; Discriminator D is With real data (Historical complete data cache) Calculate confidence ; Dynamic smoothing factor adjustment: ,in ; For over-limit points , corrected to: ;in is the measuring range of the sensor signal; if , triggering generator online fine-tuning: freezing the discriminator parameters to Update G and iterate 5 times; is the loss function of the generator; Multi-scale consistency verification: Calculation of raw data (historical cache) and calibration signals Autocorrelation coefficients at three time scales: 1 second scale: Calculate autocorrelation at 10 lags ; 10-second scale: calculate the autocorrelation of 100 points lag ; 1 minute scale: calculate the autocorrelation with a lag of 600 points .

[0024] like , ( is the autocorrelation coefficient of the original data, is the autocorrelation coefficient of the calibration signal) is determined to be inconsistent, triggering the following operations: Adjust the IMF weights for the corresponding time period ,according to renew, is the difference ratio; re-execute the LSTM prediction and weight assignment in step 2.

[0025] An EEMD-based missing data filling system for industrial sensors, including a dynamic noise adjustment and EEMD decomposition module, a missing pattern classification and divide-and-conquer filling module, a dynamic weight fusion and signal reconstruction module, a generative adversarial network (GAN) calibration module, and a real-time data stream management module; The dynamic noise adjustment and EEMD decomposition module dynamically adjusts the EEMD white noise amplitude based on Bayesian noise estimation to generate physically interpretable IMF components. It also suppresses aliasing through modal purity verification to improve decomposition robustness. The missing pattern classification and divide-and-conquer filling module dynamically classifies missing patterns (short / long intervals) based on missing duration, data volatility, and correlation, and combines neighbor interpolation with attention LSTM divide-and-conquer filling to balance real-time performance and accuracy. The dynamic weight fusion and signal reconstruction module calculates the time-varying contribution (variance and signal-to-noise ratio) of the IMF components, weights the fusion results, and updates the weights every 0.1 seconds to optimize the physical consistency of the reconstructed signal. The Generative Adversarial Network (GAN) calibration module optimizes the reconstructed signal through the generator, evaluates the authenticity through the discriminator, dynamically adjusts the smoothing factor and constrains the range, and iteratively optimizes the signal fidelity by combining multi-scale autocorrelation tests; The real-time data stream management module manages sensor data input, sliding window segmentation, historical data caching, and parallel computing scheduling, ensuring end-to-end processing latency within 50ms and supporting high-throughput real-time filling.

[0026] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. A missing data filling method for industrial sensors based on EEMD, characterized in that: The following steps are involved: S1: Adaptive dynamic EEMD decomposition, dynamically adjusting the white noise amplitude through Bayesian noise modeling, suppressing modal aliasing and generating physically interpretable IMF components; S2: IMF prediction and reconstruction, using a divide-and-conquer strategy based on missing pattern classification, combining dynamic weights to fuse IMF components, optimizing filling accuracy and real-time performance; S3: Real-time calibration of the generative adversarial network, using the GAN framework to evaluate the authenticity of the reconstructed signal, dynamically adjust the smoothing factor and verify multi-scale consistency to ensure the reliability of the signal under time domain, frequency domain and physical constraints.

2. The method for filling missing data of industrial sensors based on EEMD according to claim 1, characterized in that: The execution process of S1 is as follows: S1.1, Noise intensity estimation: Data volume segmentation, input sensor signal At a fixed sampling rate Receive, every The time buffer is a data block, and the sliding window length Dynamic selection: Based on calculate ,satisfy , each sliding step of the window ,ensure 50% overlap to cover signal mutations; S1.2, Bayesian noise modeling: for the current window Internal data, calculate the local mean : ; Assume that the noise obeys , the prior distribution is the inverse Gamma distribution , calculate the posterior probability: Where is the noise standard deviation, which obeys the inverse Gamma prior distribution ; is the data block of the current sliding window; For the window The original signal value of the sampling point; For the window The local mean of the sampling points; is the hyperparameter of the inverse Gamma distribution; is the sliding window length; Solving the maximum a posteriori estimate by gradient ascent ; S1.3, dynamic noise amplitude injection: according to Adjust the EEMD white noise amplitude: ,in is the amplitude of white noise injected in EEMD decomposition, is the optimal noise standard deviation obtained by Bayesian posterior estimation; is the empirical coefficient; right Perform EEMD decomposition: Add 100 groups of amplitudes White noise, generating IMF set ; Calculate the instantaneous frequency for each IMF Variance ,like , judged as aliasing mode, discarded and trigger local re-decomposition, only for the current window ;in For the The instantaneous frequency of each IMF component; is the variance of the instantaneous frequency, which is used to detect modal aliasing; is the sampling rate of the sensor signal.

3. The method for filling missing data of industrial sensors based on EEMD according to claim 1, characterized in that: The specific operation steps of S2 are as follows: S2.

1. Missing area positioning: missing interval detection, real-time monitoring of input signals , when continuous When the data within the time is missing or exceeds the range, the missing interval is marked , extract the missing interval before and after Known data as context ;in Represent the start and end time of the missing interval respectively, is the length of the known data extracted before and after the missing interval, Represents the context dataset before and after the extracted missing interval; S2.

2. Missing pattern classification: obtaining missing intervals , when the missing duration , it is judged as random missing, otherwise it is judged as continuous missing, where threshold for missing pattern classification; S2.

3. A divide-and-conquer strategy based on missing pattern classification is adopted, including: When it is determined to be a random missing point, that is, a short interval missing point, it is initialized by neighboring interpolation: for each missing point in the missing interval , using bilinear interpolation: ,in It is the most recent valid time point before and after the missing point; Calculate the local variance of the interpolated signal. If it exceeds the normal fluctuation range, which is 1.5 times the standard deviation of the historical data, trigger the LSTM prediction correction; When it is determined to be missing continuously, a temporal attention bidirectional LSTM prediction is performed; S2.

4. Combine the dynamic weights to fuse the IMF components.

4. The method for filling missing data of industrial sensors based on EEMD according to claim 3, characterized in that: The specific steps for performing temporal attention bidirectional LSTM prediction in S2.3 are as follows: Model loading and input construction: pre-trained bidirectional LSTM model, input layer dimension 64, hidden layer 128, attention layer 32, embedding, for each IMF component ,from Extract subsequences from , is the length of the extracted subsequence, Points, normalized and input into the model, Index for time points; Attention Mechanism Calculation: LSTM Hidden State With input Splicing, generating attention weights through the fully connected layer: ,in is the time step The attention weight, For bidirectional LSTM at time step The hidden state of is the time step Input features of is the weight matrix, and its dimensions are , ; is the trainable parameter vector and bias term, , ; is the activation function, is an activation function used to introduce nonlinearity; Weighted context vector , input the fully connected layer and output the predicted value ; Online loss calculation and optimization, loss function updated in real time: ; After predicting 10 missing intervals, the model parameters are fine-tuned using the data in the sliding window with a learning rate of 10 −4 ;in is the total loss function; is the mean square error; MF component and The correlation between them.

5. The method for filling missing data of industrial sensors based on EEMD according to claim 3, characterized in that: The specific operation steps of S2.4 are as follows: Contribution calculation and signal reconstruction: in the neighborhood of missing intervals , Seconds, the time-varying weight of each IMF is calculated in segments at 0.1 second intervals: ,in For the IMF at the time Contribution weight; For the The variance of the IMF; is the signal-to-noise ratio of the IMF, which is obtained by calculating the instantaneous energy ratio through Hilbert transform: ,in is the integration time interval, used to calculate the signal-to-noise ratio, ; For the filled IMF set according to Weighted summation is performed to obtain the final reconstructed signal: , the weight is updated every 0.1 seconds; To finally reconstruct the signal; is the residual term, which represents the signal part that cannot be explained by the IMF component; N is the total number of IMF components.

6. The method for filling missing data of industrial sensors based on EEMD according to claim 3, characterized in that: Determine the missing pattern classification threshold in S2.2 The specific steps are as follows: Initial threshold setting: Maintain a length of = 24-hour sliding window, recording the duration of all missing intervals Every time = 1 hour, calculate the quantile of missing duration in the window: , that is, choose a value that makes 95% of the missing intervals shorter than the current threshold; Dynamic adjustment: for each missing interval , extract the following features: Missing duration ; Volatility of data before and after: Calculate the standard deviation of the data 10 seconds before and after the missing interval ; Correlation decay: Calculate the Pearson correlation coefficient of the data 10 seconds before the missing interval ; Initial threshold The corrected calculation logic is as follows: ,in ; is the corrected missing pattern classification threshold; Online feedback optimization: For each filled missing interval, its physical constraints and multi-scale consistency are verified. If the filling fails and the original classification is a short missing, the threshold is lowered: If the filling fails and the original classification is a long deletion, the trigger threshold is raised: .

7. The method for filling missing data of industrial sensors based on EEMD according to claim 1, characterized in that: The specific operation steps of S3 are as follows: S3.

1. GAN online inference framework: Generator G and Discriminator D are deployed. Generator structure: 3-layer convolutional network, input 1D signal, output calibration signal, convolution kernel size 5, stride 2, number of channels [16, 32, 1] per layer; Discriminator structure: 4-layer convolutional network, output authenticity probability, convolution kernel size 3, number of channels [32, 64, 128, 1]; pre-trained model is embedded in the real-time system, and calibration is triggered every 10 seconds of data received; S3.

2. Adversarial calibration process: signal optimization and parameter adjustment, input reconstructed signal To generator G, output calibration signal ; Discriminator D is With real data Calculating confidence ; S3.3, Dynamic Smoothing Factor Adjustment: ,in ; Over-limit point , corrected to: ;in is the measuring range of the sensor signal; like , triggering generator online fine-tuning: freezing the discriminator parameters to Update G and iterate 5 times; is the loss function of the generator; S3.

4. Multi-scale consistency verification: Calculation of original data and calibration signal Autocorrelation coefficients at three time scales: 1 second scale: calculate the autocorrelation with a lag of 10 points ; 10-second scale: calculate the autocorrelation of 100 points lag ; 1-minute scale: calculate the autocorrelation with a lag of 600 points ; like ,in, is the autocorrelation coefficient of the original data, If the autocorrelation coefficient of the calibration signal is not equal to the autocorrelation coefficient of the calibration signal, it is judged as inconsistent and the following operations are triggered: Adjust the IMF weights for the corresponding time period ,according to renew, is the difference ratio; re-execute the LSTM prediction and weight assignment in S2.

8. A system for filling missing data of industrial sensors based on EEMD according to any one of claims 1 to 7, characterized in that: include: Dynamic noise adjustment and EEMD decomposition module dynamically adjusts the EEMD white noise amplitude based on Bayesian noise estimation to generate physically interpretable IMF components. It also suppresses aliasing through modal purity verification to improve decomposition robustness. The missing pattern classification and divide-and-conquer filling module is used to dynamically classify missing patterns based on missing duration, data volatility, and correlation, and combines neighbor interpolation with attention LSTM divide-and-conquer filling to balance real-time performance and accuracy. Dynamic weight fusion and signal reconstruction module, used to calculate the time-varying contribution of IMF components, weighted fusion fill results, update weights every 0.1 seconds, and optimize the physical consistency of the reconstructed signal; The adversarial generative network calibration module is used to optimize the reconstructed signal through the generator, evaluate the authenticity through the discriminator, dynamically adjust the smoothing factor and constrain the range, and iteratively optimize the signal fidelity by combining multi-scale autocorrelation tests; The real-time data stream management module is used to manage sensor data input, sliding window segmentation, historical data caching, and parallel computing scheduling to optimize processing delays.

Citation Information

Cited By

  • Self-adaptive pH value regulation and control method and system in collagen peptide enzymolysis process

    CN121538354A