Synthetic Power Equipment Data Under Fidelity and Privacy Constraints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The challenge of generating high-fidelity synthetic data for power equipment monitoring is hindered by privacy and competitiveness concerns, limiting the availability of operational and historical data necessary for testing monitoring tools and creating machine learning models.
Innovation Solution
A method for generating synthetic datasets using a combination of real data, domain knowledge, and statistical tools, involving the generation and refinement of synthetic values based on predetermined constraints to ensure data fidelity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If real operational and historical data from power equipment is collected, then data fidelity and authenticity are improved, but data privacy and competitiveness concerns worsen
Solution Approach 1:
The patent creates synthetic copies of real power equipment data that replicate statistical properties, temporal patterns, and operational characteristics without containing actual sensitive information. The synthetic dataset mirrors the structure and behavior of real data while being mathematically generated, thus preserving data fidelity for testing and analysis while eliminating privacy and competitiveness risks associated with using real operational data
Solution Approach 2:
The patent introduces synthetic data as an intermediary between real data and testing/analysis applications. This intermediary layer allows researchers and developers to access data with realistic properties for model training and tool testing without direct exposure to confidential operational data, thus resolving the contradiction between data authenticity and privacy protection
2Productivity
If synthetic data is generated without rigorous constraint validation, then data generation speed is improved, but data fidelity and realism worsen
Solution Approach 1:
The patent performs preliminary actions by pre-defining all constraints, boundaries, and statistical properties before data generation begins. The synthetic data generation process is designed with predetermined rules for temporal correlations, operational ranges, and event patterns, allowing rapid generation while ensuring fidelity through built-in constraint validation rather than post-generation checking
Solution Approach 2:
The patent implements feedback mechanisms where generated synthetic values are continuously validated against predefined constraints and statistical properties. The system monitors whether generated data points satisfy temporal correlations, operational boundaries, and event patterns, and adjusts subsequent generation parameters accordingly to maintain data fidelity while preserving generation speed through efficient validation loops
Data Source
Figure 1
Figure 2
Figure 2
AI summary
The present disclosure relates to a method for generating a synthetic dataset of a power equipment. The method comprises 1a. obtaining a dataset of the power equipment having a plurality of data points, 1b. generating an initial synthetic value for at least one data point in the dataset obtained in 1a, 1c. determining whetherthe initial synthetic value satisfies at least one predetermined constraint, 1d. determining that the initial synthetic value is a final synthetic value when it is determined that the initial synthetic value satisfies the at least one predetermined constraint, or 1e. discarding the initial synthetic value and generating a new synthetic value when it is determined that the initial synthetic value does not satisfy the at least one predetermined constraint. Steps 1c. to 1e. are repeated using the new synthetic value generated in step 1e. until the new synthetic value satisfies the at least one predetermined constraint and is determined to be a final synthetic value. The method further comprises step 1f. performing steps 1b. to 1e. for all subsequent data points of the dataset, and step 1g. outputting a generated synthetic dataset comprising the final synthetic values.