Utility Dataset Anonymization That Preserves Grid Topology
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Utility datasets containing personal identifiable information (PII) pose challenges for utility companies as they need to maintain privacy while preserving valuable grid topology and consumption data, and conventional methods often strip away useful information.
Innovation Solution
A two-phase anonymization technique involving a one-way hashing algorithm for assigning anonymous identifiers and a data swapping method to maintain topology and consumption data integrity, ensuring PII removal without losing utility dataset value.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If conventional approaches are used to remove PII from utility datasets, then privacy protection is improved, but useful information such as grid topology information and equipment specifications is lost
Solution Approach 1:
The patent extracts and removes only the personal identifiable information (names, addresses, contact information) from the utility dataset while deliberately preserving the grid topology information, equipment specifications, and other useful technical data. This selective extraction approach resolves the contradiction by removing harmful PII without losing valuable grid information.
Solution Approach 2:
The patent applies different treatment to different parts of the dataset: PII fields are removed or anonymized while grid topology fields, equipment specification fields, and consumption data fields are preserved. This local differentiation approach ensures privacy protection in sensitive areas while maintaining data utility in technical areas.
2Object-affected harmful factors
If conventional approaches are used to remove PII from utility datasets, then privacy protection is improved, but equipment specifications and other useful data are stripped away
Solution Approach 1:
The patent selectively extracts and removes only PII elements while preserving equipment specifications and other technical data. The anonymization process targets specific PII fields rather than removing all identifiable information, thus protecting privacy while retaining equipment specification details needed for grid management.
Solution Approach 2:
The patent applies privacy protection measures specifically to PII-containing fields while leaving equipment specification fields intact. This localized approach ensures that privacy concerns are addressed without sacrificing the technical information needed for equipment maintenance and grid operations.
3Reliability
If anonymization techniques are applied to utility datasets, then data security is improved, but data complexity and processing requirements increase
Solution Approach 1:
The patent performs anonymization processing as a preliminary step before data analysis or sharing. By pre-anonymizing the utility dataset, the system establishes security foundations upfront, allowing subsequent data usage without requiring complex real-time anonymization during analysis operations.
Solution Approach 2:
The patent creates an anonymized copy of the original utility dataset that can be used for analysis, sharing, and processing without compromising the security of the original data containing PII. This copying approach allows multiple users to work with secure data while the original sensitive data remains protected.
4Object-affected harmful factors
If PII is removed from utility datasets, then privacy obligations are met, but the usefulness of the data for grid analysis and research is reduced
Solution Approach 1:
The patent extracts and removes only the PII components (customer names, addresses, contact information) while preserving all grid analysis-relevant data including topology structures, equipment specifications, and consumption patterns. This selective extraction maintains both privacy compliance and analytical usefulness.
Solution Approach 2:
The patent applies privacy protection selectively to PII fields while preserving the quality and completeness of grid analysis fields. The anonymized dataset maintains full topological relationships, equipment details, and consumption data needed for grid analysis, load forecasting, and research purposes.
Data Source
AI summary
A data anonymization technique for datasets including personal identifiable information (PII) such as utility datasets may primarily include assigning anonymous identifiers to nodes in the dataset and swapping or otherwise moving portions of information between nodes of the dataset. The methodology of the swapping or moving operation may vary optionally based on a number of parameters, and may include swapping endpoints under a single parent, swapping endpoints between similar parents, and/or swapping similar endpoints between parents. The anonymization technique may output an anonymized dataset which reflects a topology of the original dataset and may optionally be updatable and modifiable to include additional data about existing or new nodes.


