Derivative Data Visualization for Privacy-Preserving Genome Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems face challenges in effectively analyzing and sharing large data sets, particularly in genetic epidemiology, due to difficulties in distinguishing true genetic associations from statistical artifacts and ethical concerns related to patient privacy.

Innovation Solution

The development of systems and methods for automated analysis, visualization, and sharing of large data sets, including dynamic genome browsers and meta-analysis tools, which facilitate rapid gene association search, replication, and open data release while ensuring privacy protection through anonymization and derivative data sharing.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If open access data sharing of unabridged GWAS data is implemented, then data sharing and replication opportunities are improved, but patient privacy protection deteriorates

Engineering Contradiction:
Improvedata sharing capabilityVSAvoidprivacy risk
Core Design Contradiction:
Adaptability or versatilityVSObject-affected harmful factors

Solution Approach 1:

The patent creates derivative data products (manhattan plots, genome browsers, association tables) that copy and transform the original GWAS data into summarized formats. These derivatives retain analytical value while removing direct patient identifiers, enabling sharing without compromising privacy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system extracts and removes personally identifiable information from the raw GWAS data, separating the genetic association signals from patient-level identifiers. This extraction allows the core scientific data to be shared while the privacy-sensitive elements are excluded.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If statistical thresholds for genome-wide significance are applied, then false positive associations are reduced, but true functional associations below the threshold are discarded

Engineering Contradiction:
Improveassociation validityVSAvoidtrue associations
Core Design Contradiction:
ReliabilityVSLoss of information

Solution Approach 1:

The patent segments the GWAS results into multiple categories: genome-wide significant hits, sub-threshold associations, and candidate regions. This segmentation allows different analytical approaches for different result categories, preserving potentially important signals that would be discarded by a single threshold approach.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs more comprehensive analysis than traditional GWAS by examining all SNPs and variants regardless of p-value threshold, then applying multiple layers of validation. This excessive initial screening ensures no true associations are prematurely discarded, with subsequent filtering occurring only after thorough examination.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If genome-wide association studies interrogate millions of SNPs, then disease-gene associations are identified, but statistical artifacts increase

Engineering Contradiction:
Improvediscovery rateVSAvoidstatistical artifacts
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The patent implements multiple feedback loops including replication analysis in independent cohorts, functional annotation verification, and consistency checking across different statistical models. These feedback mechanisms allow continuous refinement of associations, filtering out statistical artifacts that arise from interrogating millions of SNPs.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary quality control and annotation of all SNPs and variants before formal association testing. This preliminary action identifies and removes problematic variants and populations at risk of generating statistical artifacts, reducing false positives before they enter the analysis pipeline.

Inventive Principle:
Principle #10Preliminary action

4Object-affected harmful factors

If patient anonymization is attempted with genetic data, then privacy protection is improved, but data sharing capability deteriorates

Engineering Contradiction:
Improveprivacy protectionVSAvoiddata sharing capability
Core Design Contradiction:
Object-affected harmful factorsVSAdaptability or versatility

Solution Approach 1:

The patent creates derivative data products (manhattan plots, genome browsers, association tables) that copy and transform the original GWAS data into summarized formats. These derivatives retain analytical value while removing direct patient identifiers, enabling sharing without compromising privacy.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system introduces intermediate data representations that serve as mediators between raw patient data and research applications. These intermediates (aggregated statistics, visualizations, annotated variants) bridge the gap between privacy requirements and sharing needs, allowing data to be shared in a privacy-preserving manner.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS9910957B2Visualization, sharing and analysis of large data sets
Publication Date: 2018.03.06 SAINT PETERSBURG STATE UNIVERSITY
  • US9910957B2 patent drawing
  • US9910957B2 patent drawing
  • US9910957B2 patent drawing

AI summary

Systems and methods for visualization, sharing and analysis of large data sets are described. Systems and methods may include receiving an input data set, wherein the input data set includes data that can be classified in classification dimensions wherein a first classification dimension is a linear ordering of data entries and a second classification dimension represents analysis criteria, traits of the data entries, or aspects of the data entries; obtaining an unabridged data table listing results for each combination of coordinates in the first classification dimension and the second classification dimension; and displaying contents of the unabridged data table as a visual array wherein two axes correspond to the coordinates and a third axis corresponds to a third classification dimension, wherein the third classification dimension represents an actual value of the respective data point for the coordinates. Methods may also assess the visual array, such as by identifying one or more regions of high density of signals.