Linkage Disequilibrium Analysis for Low-Coverage Genomic Identity

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for determining whether two DNA samples originate from the same individual are challenging when DNA coverage is low, as they rely on identifying alleles at polymorphic sites, which is difficult due to limited sequence information and fragmented DNA, especially in ancient or degraded samples.

Innovation Solution

A method using linkage disequilibrium (LD) scores calculated from independent DNA sequences to determine the likelihood that sequences are from a single individual, employing publicly available SNP databases and next-generation sequencing to analyze polymorphic sites, even with low sequence coverage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If low-coverage shotgun sequencing is used to analyze ancient or degraded DNA samples, then the ability to sequence highly fragmented DNA is improved, but the ability to identify alleles at polymorphic sites deteriorates

Engineering Contradiction:
Improveability to sequence highly fragmented DNAVSAvoidability to identify alleles at polymorphic sites
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The invention transitions from analyzing individual SNP alleles in isolation to examining pairs of SNPs in linkage disequilibrium. By adding the dimension of pairwise relationships between polymorphic sites, the method extracts more information from low-coverage data. The LD score aggregation across many SNP pairs compensates for the low coverage at individual sites, enabling reliable individual identification even when most positions have no observations in a single library.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If direct comparison of alleles at polymorphic sites is performed between two low-coverage libraries, then individual identification can be attempted, but the number of informative sites for comparison deteriorates to a very small number

Engineering Contradiction:
Improveindividual identification capabilityVSAvoidnumber of informative sites
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The invention merges information from many SNP pairs by aggregating LD scores across the genome. Instead of relying on a small number of directly observed informative sites, the method combines weak signals from numerous SNP pairs in linkage disequilibrium. This aggregation approach accumulates sufficient statistical power to reliably distinguish between samples from the same individual versus different individuals, even when each individual SNP has very low coverage.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If SNP-based forensic analysis using a small set of curated SNPs is used, then individual identification can be performed, but the markers are not observed in low-coverage shotgun sequencing of the genome

Engineering Contradiction:
Improveindividual identification accuracyVSAvoidobservability of markers in low-coverage sequencing
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The invention creates a universal method that works with both targeted SNP panels and low-coverage whole-genome sequencing data. By using linkage disequilibrium patterns across many polymorphic sites rather than relying on a specific small set of curated SNPs, the approach is adaptable to different sequencing strategies. The method can utilize any polymorphic sites present in the low-coverage data, making it universally applicable regardless of which specific markers are observed.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentEP3158488B1Method for determining relatedness of genomic samples using partial sequence information
Publication Date: 2024.05.22 RGT UNIV OF CALIFORNIA
  • EP3158488B1 patent drawingFigure 1A
  • EP3158488B1 patent drawingFigure 1B
  • EP3158488B1 patent drawingFigure 1C

AI summary

Disclosed are methods for testing biological samples containing genomic nucleic acids obtained from an organism having a genome, such as a human genome. It is often desirable to analyze a DNA sample or more than one, different DNA samples, to determine whether the sample comes from one individual or two individuals. The present method requires very low amounts of DNA and can use partial sequences of DNA fragments. Partial sequences are analyzed for the presence of polymorphisms (e.g. SNP's) that can be mapped to a reference SNP map. The distance between similar SNPS, which are genetically linked, can be used to statistically determine a likelihood of identity of individuality in a sample.