Linking Sequence Reads to Anchor Positions for Structural Variant Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA sequencing methods struggle to accurately detect and determine structural variants due to the loss of connectivity information between sequence fragments during the sequencing process, leading to computationally intensive assembly processes and potential errors.
Innovation Solution
The method involves analyzing sequence reads located within clusters on a flowcell, determining the position of anchor sequence reads, and calculating a threshold distance to link unknown sequence reads to anchor reads, thereby determining their actual position in the genome with high confidence.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional shotgun sequencing approach is used to sequence large genomic DNA fragments, then the sequencing process can be performed on smaller fragmented pieces, but the connectivity information between sequence fragments is lost and computationally intensive assembly processes are required
Solution Approach 1:
The patent applies preliminary action by performing in-situ fragmentation directly on the flow cell surface before sequencing. This preserves the spatial relationships between DNA fragments and their original positions in the template, maintaining connectivity information that would otherwise be lost. The fragments remain associated with their location data, enabling later reconstruction of the original sequence without intensive assembly computations.
Solution Approach 2:
The patent introduces spatial dimensionality by performing fragmentation and sequencing on a flow cell surface rather than in solution. This adds a spatial coordinate system (x, y positions on the flow cell) that preserves connectivity information. The two-dimensional spatial arrangement on the flow cell surface encodes the relationship between fragments and their original positions, allowing reconstruction without losing connectivity data.
2Reliability
If in-situ fragmentation is performed on the flow cell, then connectivity information is preserved, but the device complexity increases
Solution Approach 1:
The patent merges multiple functions into a single integrated flow cell system. The flow cell simultaneously serves as the substrate for library preparation, the surface for in-situ fragmentation, and the platform for sequencing. This consolidation preserves connectivity information while avoiding the need for separate complex devices for each step, thereby managing device complexity through functional integration.
Solution Approach 2:
The flow cell surface performs self-service by providing built-in features that enable in-situ fragmentation and sequencing without requiring external complex equipment. The flow cell's surface chemistry and structure are designed to facilitate fragment attachment and sequencing reactions directly on its surface, reducing the need for additional specialized devices and simplifying the overall system.
3Measurement precision
If all sequence reads from the entire genome are analyzed to detect structural variants, then comprehensive detection is achieved, but the data processing volume and time increase significantly
Solution Approach 1:
The patent applies local quality by focusing analysis on specific local regions around anchor reads rather than uniformly processing all sequence reads. Anchor reads serve as localized reference points, and analysis is concentrated in their vicinity where structural variants are most likely to be detected. This localized approach maintains high detection accuracy while significantly reducing the total data processing volume and time required.
Solution Approach 2:
The patent extracts and utilizes spatial position information from sequence reads as a filtering criterion. By extracting the spatial coordinates of reads on the flow cell surface, the method identifies and analyzes only those reads located within threshold distances of anchor reads. This extraction of spatial data enables selective processing of relevant reads, reducing overall data processing requirements while maintaining comprehensive variant detection capability.
Data Source
AI summary
Described are DNA sequencing systems and methods. Systems and methods may establish a link between read sequences when the sequences are within a threshold distance. The read sequences that are mapped and aligned with high confidence may be used to determine the location of the nearby linked read sequences that would otherwise be difficult to place. The systems and methods may identify structural variants in the polynucleotide by analyzing sequence reads located within a threshold distance to the anchor sequence reads on the flowcell to determine sequence reads linked to the anchor sequence reads.


