Synthetic Genomic Datasets for Bioinformatics Validation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current genomic analysis tools lack effective methods for validating the performance and accuracy of mutation calling algorithms, particularly in whole genome sequencing, due to variations in sample analysis and underlying assumptions that impact computing platforms.

Innovation Solution

The development of systems and methods using synthetic digital patient data sets with defined mutations to simulate tumor and normal tissue genomes, allowing for the evaluation and validation of genomic analysis algorithms by assessing accuracy, sensitivity, and specificity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If synthetic digital patient data sets with defined mutations are used to validate genomic analysis algorithms, then measurement precision of mutation calling is improved, but device complexity increases due to the need for virtual genome generation and validation infrastructure

Engineering Contradiction:
Improvemeasurement precisionVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent creates synthetic digital copies of patient genomes with known defined mutations to serve as validation data. These virtual genomes replicate real genomic structures and variations, allowing algorithm validation without requiring additional physical biological samples. The copying principle enables precise measurement of algorithm performance against known ground truth while avoiding the complexity of generating new biological specimens.

Inventive Principle:
Principle #26Copying

2Reliability

If multiple synthetic patient datasets with various genetic alterations are generated, then reliability of algorithm validation is improved, but loss of time increases due to the computational effort required for virtual genome generation and analysis

Engineering Contradiction:
ImprovereliabilityVSAvoidloss of time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary generation of synthetic digital patient datasets with predefined mutations before algorithm validation. By preparing multiple virtual genomes with known genetic alterations in advance, the system establishes a ready-to-use validation framework that can be applied to test algorithms without repeated data generation. This preliminary action improves validation reliability across multiple testing scenarios while reducing cumulative time loss.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If synthetic genomes are used to evaluate algorithm performance across different sample variations, then adaptability of validation systems is improved, but manufacturing precision decreases due to the complexity of simulating diverse genetic alterations

Engineering Contradiction:
ImproveadaptabilityVSAvoidmanufacturing precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies local quality by introducing specific defined mutations at particular locations within synthetic digital genomes. Rather than uniformly randomizing all genetic variations, the system precisely controls the type, position, and frequency of specific genetic alterations (SNPs, insertions, deletions) in different genomic regions. This approach enables adaptable validation across diverse sample types while maintaining precision in the known ground truth of each local mutation.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10984890B2Synthetic WGS bioinformatics validation
Publication Date: 2021.04.20 NANTOMICS LLC

AI summary

Systems, methods, and devices for generating synthetic genomic datasets and validating bioinformatic pipelines for genomic analysis are disclosed. In preferred embodiments, synthetic maternal and paternal datasets with known variants are used with matched normal synthetic datasets to validate various bioinformatic pipelines. Bioinformatic pipelines are evaluated using the synthetic datasets to assess design changes and improvements. Accuracy, PPV, specificity, sensitivity, reproducibility, and limit of detection of the pipelines in calling variants in synthetic datasets is reported.