Genomic Data Indexing for Privacy-Controlled Visualization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
There is a need for secure and efficient distribution of genomic data while ensuring privacy and control over access, as existing systems lack effective mechanisms for secure data sharing and privacy protection in genomic visualization.
Innovation Solution
A method and system for correlating genome data with electronic medical record (EMR)/personal health record (PHR) data by indexing and encrypting genome files with unique identifiers, storing index files in a central depository, and allowing owners to control access through a brokering server, enabling privacy-controlled genomic visualization and secure data sharing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If genomic data is distributed through existing visualization systems, then data accessibility and visualization capability are improved, but data privacy and security control are compromised
Solution Approach 1:
The system segments genomic data into two distinct components: encrypted genomic data files stored locally at source sites, and unencrypted index files containing only search criteria stored centrally. This segmentation allows the index to be freely accessible for search operations while the actual genomic data remains protected and under local control, resolving the contradiction between data accessibility and privacy protection.
Solution Approach 2:
The system introduces an intermediary mechanism where index files act as mediators between users and encrypted genomic data. The index files contain searchable information without exposing the actual genomic data, enabling users to search and identify relevant data without direct access to the sensitive information. This intermediary layer maintains both accessibility and security.
2Productivity
If genomic data is centralized for easy search and access, then search efficiency is improved, but security risks and loss of owner control increase
Solution Approach 1:
The system separates searchable index information from sensitive genomic data, storing the index centrally for efficient searching while keeping the actual data distributed and encrypted at source sites. This resolves the contradiction by enabling centralized search functionality without centralizing the sensitive data itself.
Solution Approach 2:
The system creates simplified copies (index files) of genomic data that contain only search-relevant information without the sensitive details. These index copies are stored centrally for efficient searching, while the original encrypted data remains at source sites, eliminating security risks associated with centralizing full genomic datasets.
3Reliability
If genomic data is encrypted to protect privacy, then data security is improved, but data sharing and visualization capabilities are reduced
Solution Approach 1:
The system segments data into encrypted genomic data files and unencrypted index files. The index files can be freely shared and accessed for visualization purposes without encryption, while the actual genomic data remains encrypted for security. This segmentation maintains both data security and sharing capability.
Solution Approach 2:
The unencrypted index file serves as an intermediary that enables data sharing and visualization without exposing encrypted genomic data. Researchers can use the index information for visualization and analysis while the actual sensitive data remains protected, resolving the contradiction between security and adaptability.
4Productivity
If full genomic data is made accessible for research, then research capability is improved, but patient anonymity and privacy are compromised
Solution Approach 1:
The system segments information into index data that supports research functionality and encrypted genomic data that protects patient anonymity. The index files provide sufficient information for research and visualization without containing personally identifiable information or sensitive genomic details that would compromise anonymity.
Solution Approach 2:
The system creates index file copies that contain research-relevant information without including sensitive patient identifiers or full genomic sequences. These simplified copies enable research capability while preserving patient anonymity by omitting identifying information from the accessible copies.
Data Source
AI summary
A method and system for correlating genome data with EMR/PHR data is disclosed herein. The method includes identifying a plurality of sources of genome data. The method also includes generating an index file for each of the plurality of genome files. The method also includes transmitting each index file to a central depository. The method also includes identifying electronic medical record (EMR) and/or personal health record (PHR) data at each source of the plurality of sources of genome data. The method also includes correlating each genome file of the plurality of genome files with a corresponding EMR/PHR data.


