Genetic Data Integration via Standardized Model and Validation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for integrating and analyzing genetic and phenotypic data in clinical settings are inefficient, as they fail to combine data from multiple sources into a standardized format, leading to difficulties in sharing information within the biomedical community and making effective predictions.

Innovation Solution

A system that integrates genetic and phenotypic data into a standardized data model, using expert and statistical relationships to validate and analyze the data, while ensuring data privacy and allowing for compensation of individuals and validators, and enabling secure access through biometric authentication.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If data from multiple sources is integrated into a standardized format, then information sharing and analysis capability is improved, but system complexity and implementation difficulty increase

Engineering Contradiction:
Improveinformation sharing capabilityVSAvoidsystem complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments the complex data integration task into distinct functional modules: a data intake service that receives data from multiple sources, a validation service that checks data quality, and an analysis service that processes validated data. This modular architecture reduces system complexity by making each component independently manageable while maintaining overall integration capability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a standardized data model as an intermediary layer between diverse data sources and analysis tools. This intermediate representation format acts as a universal interface that translates various source formats into a common structure, enabling information sharing without requiring direct integration between all source systems.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If comprehensive data validation is performed using expert and statistical relationships, then data reliability is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvedata reliabilityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary validation by establishing expert rules and statistical relationships in advance during system setup. These pre-configured validation criteria are stored and automatically applied to incoming data, eliminating the need to compute validation rules from scratch for each data set, thus reducing processing time while maintaining reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The validation mechanism operates autonomously by automatically applying pre-configured expert rules and statistical models to validate incoming data without requiring manual intervention. The system self-manages the validation process, checking data quality metrics and flagging anomalies automatically, which reduces processing overhead compared to manual validation methods.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If genetic and phenotypic data are aggregated from multiple subjects, then predictive accuracy is improved, but privacy protection challenges increase

Engineering Contradiction:
Improvepredictive accuracyVSAvoidprivacy risks
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The system extracts and separates personally identifiable information from genetic and phenotypic data during the intake process. By removing direct identifiers and storing them separately with enhanced security controls, the system enables aggregation of clinical data for improved predictive accuracy while mitigating privacy risks through the extraction of sensitive personal information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system creates anonymized copies of genetic and phenotypic data for aggregation and analysis purposes. These copied data sets retain the clinical information needed for predictive modeling while removing direct personal identifiers, allowing multiple subjects' data to be combined for improved accuracy without exposing individual privacy.

Inventive Principle:
Principle #26Copying

4Manufacturing precision

If manual data integration methods are used, then data quality control is improved, but resource intensity and cost increase

Engineering Contradiction:
Improvedata quality controlVSAvoidresource intensity
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The system implements automated validation that performs data quality control self-service by automatically applying pre-configured expert rules and statistical models to incoming data. This automated approach maintains data quality control comparable to manual methods while significantly reducing resource intensity by eliminating the need for manual review of each data entry.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual data integration processes with automated computational systems that use algorithmic validation rules. This substitution of mechanical human labor with automated computing resources maintains data quality control through systematic application of validation criteria while reducing overall resource intensity by scaling efficiently with data volume.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS8024128B2System and method for improving clinical decisions by aggregating, validating and analysing genetic and phenotypic data
Publication Date: 2011.09.20 NATERA INC
  • US8024128B2 patent drawing
  • US8024128B2 patent drawing
  • US8024128B2 patent drawing

AI summary

The information management system disclosed enables caregivers to make better decisions by using aggregated data. The system enables the integration, validation and analysis of genetic, phenotypic and clinical data from multiple subjects. A standardized data model stores a range of patient data in standardized data classes comprising patient profile, genetic, symptomatic, treatment and diagnostic information. Data is converted into standardized data classes using a data parser specifically tailored to the source system. Relationships exist between standardized data classes, based on expert rules and statistical models, and are used to validate new data and predict phenotypic outcomes. The prediction may comprise a clinical outcome in response to a proposed intervention. The statistical models and methods for training those models may be input according to a standardized template. Methods are described for selecting, creating and training the statistical models to operate on genetic, phenotypic, clinical and undetermined data sets.