Digital Twin Signal Virtualization via Ensemble Imputation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The inefficiency of data ingestion and storage in digital twin systems due to large volumes of data from multiple sources, which can lead to performance issues and increased bandwidth and storage requirements.
Innovation Solution
Implementing imputation methods to estimate and generate data points between checkpoints, reducing the need for actual data transmission and storage, using ensemble-based imputation techniques such as Last Observation Carried Forward, Next Observation Carried Backward, Rolling Moving Average, and forecasting models to minimize data flow and volume.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If data from multiple data sources and databases is ingested into the digital twin, then the digital twin can perform simulations using real data, but the volume of data becomes prohibitively large and impacts operation efficiency
Solution Approach 1:
The patent extracts only the essential features and patterns from the raw data rather than ingesting complete datasets. The digital twin system identifies and extracts key signal characteristics, statistical properties, and critical data points, discarding redundant information. This extraction approach maintains simulation reliability while dramatically reducing data volume to improve operational efficiency.
Solution Approach 2:
Instead of ingesting all raw data and then processing it, the patent inverts the approach by first defining the essential simulation requirements and then generating only the necessary synthetic data that meets those requirements. The system works backward from the simulation needs to generate minimal sufficient data, rather than forward from comprehensive data collection.
2Adaptability or versatility
If data from multiple sources is transmitted to the digital twin, then comprehensive simulations can be performed, but bandwidth consumption increases significantly
Solution Approach 1:
The patent creates simplified copies or representations of the original data sources rather than transmitting the actual complete datasets. Synthetic data generators produce copies that replicate the essential statistical properties, distributions, and relationships of the source data, enabling comprehensive simulations while consuming minimal bandwidth.
Solution Approach 2:
The system changes the parameters of data representation from complete raw data to condensed statistical parameters. Instead of transmitting full time-series data, the patent transmits key parameters such as mean, variance, autocorrelation coefficients, and other statistical descriptors that capture the essential characteristics needed for accurate simulations.
3Measurement precision
If complete data is stored in the digital twin, then accurate analysis can be performed, but storage requirements become prohibitive
Solution Approach 1:
The patent transforms complete data into condensed parameter representations for storage. Instead of storing full datasets, the system stores statistical parameters, model coefficients, and compressed representations that preserve analysis accuracy while occupying minimal storage space. The stored parameters can be expanded back into full synthetic datasets when needed for analysis.
Solution Approach 2:
The patent applies different levels of data representation to different parts of the system. Critical data that requires high fidelity for analysis is stored in detailed form, while less critical data is stored in compressed parameter form. This local differentiation of data quality maintains analysis accuracy for essential components while reducing overall storage requirements.
Data Source
AI summary
Signal virtualization using an ensemble of imputation operations is disclosed. In a digital twin, observations or data points are imputed using a combination of imputation operations. The imputed observations are generated between checkpoint operations, which correspond to observations from a data source. The number of imputed observations can be determined in advance.


