Proximity Ligation Mapping for Structural Variant Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current sequencing technologies face challenges in producing high-quality, highly contiguous genome sequences, particularly from preserved samples like formalin-fixed, paraffin-embedded (FFPE) samples, due to a lack of efficient methods for analyzing and assembling genomic data.
Innovation Solution
Methods involving mapping read pairs onto a reference nucleic acid scaffold, assigning positions to bins, generating two-dimensional images, calculating z-scores, and identifying candidate hits for structural variations, along with systems for modeling allelic variations and restructuring sequence scaffolds to reduce density variations, are employed to detect and correct genomic rearrangements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If read pair mapping and binning methods are used to detect structural variants, then measurement precision is improved, but device complexity increases
Solution Approach 1:
The genome is divided into bins of varying sizes, with larger bins for regions expected to have few read pairs and smaller bins for regions expected to have many read pairs. This segmentation allows the system to apply appropriate detection thresholds and analysis methods to different genomic regions, improving structural variant detection precision while managing computational complexity through region-specific processing.
2Manufacturing precision
If complex computational methods are applied to analyze genomic data, then manufacturing precision is improved, but productivity decreases
Solution Approach 1:
Different bin sizes are applied to different genomic regions based on expected read pair density. Regions with low read pair density use larger bins to reduce computational burden, while regions with high read pair density use smaller bins to maintain detection sensitivity. This local adaptation of analysis parameters improves genome assembly quality in critical regions while maintaining overall analysis productivity.
3Loss of information
If traditional sequencing methods are used, then ease of operation is maintained, but loss of information increases
Solution Approach 1:
The method transforms one-dimensional linear genomic data into two-dimensional contact frequency data by analyzing spatial relationships between genomic loci. This dimensional transformation enables recovery of phasing information and detection of structural variants that are invisible in traditional linear sequencing, capturing additional biological information without requiring complex sample processing.
Data Source
AI summary
The disclosure provides methods, systems, and algorithms to identify and report genome or chromosome level structural information, such as the presence of structural variations. In some cases, structural variations include copy number variations, inversions, deletions, tandem duplications, or inverted duplications. Further provided herein are methods, systems and algorithms for assembling read-paired genomic data, including creating and optimizing scaffo


