Anonymizing IoT Data Subsets via Correlation Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods fail to effectively anonymize datasets from network-connected devices, making it difficult to protect the identity of entities in IoT applications, particularly in connected car data analysis where non-essential identification data is collected.
Innovation Solution
A method and system that select anonymized subsets of parameters from datasets by calculating autocorrelation and cross-correlation, applying these metrics to a decision function to determine if the selected subset is insufficient for identifying network-connected devices, using a quotient threshold to assess anonymity.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If datasets containing parameters from network-connected devices are collected for analysis, then the quantity and quality of data available for processing is improved, but the ability to protect the identity of entities (anonymity) deteriorates
Solution Approach 1:
The patent extracts and removes identification-capable parameters from the dataset. The system automatically identifies parameters that can be used to identify entities and extracts only the anonymized subset that is insufficient for identification, thereby separating useful analytical data from identifying information.
Solution Approach 2:
The patent segments the original dataset into two distinct parts: identification parameters and anonymized parameters. By dividing the data into these segments and retaining only the anonymized portion for analysis, the system maintains data utility while protecting entity identities.
2Loss of information
If parameters are removed from datasets to achieve anonymization, then the anonymity of entities is improved, but the quality and usefulness of data for analysis deteriorates
Solution Approach 1:
The patent employs feedback mechanisms through automated algorithms that evaluate the anonymized dataset to ensure it remains sufficient for analysis. The system continuously assesses whether the remaining parameters maintain analytical value while achieving anonymization, adjusting the selection process based on this feedback.
Solution Approach 2:
The patent changes the parameters of the dataset by transforming identification-capable parameters into anonymized versions. The system modifies the data structure and parameter selection to maintain analytical quality while removing identifying characteristics, using automated criteria to determine which parameters to retain.
3Productivity
If automated algorithms are used to select anonymized subsets of parameters, then the productivity of data processing is improved, but the device complexity increases
Solution Approach 1:
The patent implements self-service through automated algorithms that autonomously perform the anonymization process. The system automatically identifies, selects, and processes parameters without requiring manual intervention, enabling the system to serve itself in terms of data preparation and anonymization tasks.
Solution Approach 2:
The patent creates a universal anonymization system that can process multiple types of datasets from various network-connected devices. The automated algorithm is designed to handle different data formats and structures, making the system multi-functional and adaptable to diverse data sources while maintaining consistent anonymization standards.
Data Source
AI summary
A method and a system for selecting an anonymized subset of parameters from datasets of network-connected devices are provided herein. The method may include: obtaining a plurality of datasets, comprising a set of parameters related to one of a plurality of network-connected devices; automatically selecting a subset of parameters from at least one of the datasets, wherein the selecting is based on specified selection criteria; calculating an autocorrelation of the selected subset of parameters; calculating a correlation of the selected subset of parameters and one or more subsets of parameters selected from the datasets relating to network-connected devices other than said one of the plurality of network-connected devices; and applying the correlation and the autocorrelation to a decision function to determine whether the selected subset of parameters is an anonymized subset that is insufficient for determining an identity of the one of the plurality of the network-connected devices.


