Privacy Preserving Data-Mining Protocol for Healthcare Research
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data privacy regulations, such as HIPAA, restrict the use of personal health information, limiting the ability to perform high-resolution analysis and data manipulation necessary for healthcare and pharmaceutical research, while also preventing the breach of individual privacy.
Innovation Solution
A Privacy Preserving Data-Mining Protocol that allows for the analysis of raw, high-resolution data across multiple sources while maintaining privacy restrictions, using an aggregator processor to collect and aggregate data from source-entity processors, filtering out sensitive information to prevent identity disclosure.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If data privacy regulations (HIPAA) are enforced to protect individual privacy, then individual privacy is preserved, but the ability to perform high-resolution data analysis and manipulation is limited
Solution Approach 1:
The patent segments data into multiple representations: original high-resolution data retained by source entities, de-identified data for analysis, and synthesized data for sharing. This segmentation allows different uses of the same data with appropriate privacy protections at each stage.
Solution Approach 2:
The patent introduces an intermediary de-identification process that acts as a mediator between raw data and analysis systems. This intermediary layer removes or generalizes identifiers while preserving analytical value, enabling analysis without direct access to sensitive personal information.
2Object-affected harmful factors
If data is de-identified and aggregated to protect privacy, then individual privacy is preserved, but the resolution and detail of data analysis are reduced
Solution Approach 1:
The patent applies different quality levels to different portions of data based on local needs. Source entities maintain high-resolution data locally for detailed analysis, while sharing lower-resolution de-identified data externally. This local quality approach preserves precision where needed while protecting privacy where data is shared.
Solution Approach 2:
The patent adds a new dimension to data representation by creating synthesized data that captures statistical properties and patterns without containing individual records. This dimensional transformation allows analysis of data characteristics without exposing sensitive details.
3Productivity
If full transparency of data is allowed to enable comprehensive analysis, then data utilization value is maximized, but individual privacy may be breached
Solution Approach 1:
The patent implements partial transparency by providing just enough data detail to enable meaningful analysis without revealing complete individual information. De-identified data provides sufficient granularity for research while removing identifiers that would enable re-identification, achieving partial action that balances utility and privacy.
Data Source
AI summary
Privacy Preserving Data-Mining Protocol, between a secure “aggregator” and “sources” having respective access to privacy-sensitive micro-data, the protocol including: the “aggregator” accepting a user query and transmitting a parameter list for that query to the “sources” (often including privacy-problematic identifiable specifics to be analyzed); the “sources” then forming files of privacy-sensitive data-items according to the parameter list and privacy filtering out details particular to less than a predetermined quantity of micro-data-specific data-items; and the “aggregator” merging the privacy-filtered files into a data-warehouse to formulate a privacy-safe response to the user—even though the user may have included privacy-problematic identifiable specifics.


