Genomic Alignment via Interval Ratios for Structural Variant Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current DNA sequencing techniques, such as next-generation sequencing (NGS), are limited in reading lengths and struggle to accurately detect structural variants due to errors in measurement data, including apparent expansion and contraction of DNA fragments, which can lead to erroneous recognition of structural variants.
Innovation Solution
An information processing device with a processor and memory that calculates ratios of intervals between partial sequences in both reference and target nucleic acid sequences, constructs an index to handle apparent expansion and contraction, and aligns labeled positions by comparing these ratios to identify accurate positions on the reference genome.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If current DNA sequencing techniques (NGS) are used to read base sequences, then the reading process is fast and cost-effective, but the read length is limited to about hundreds of bases making it impossible to capture large genome regions for structural variant detection
Solution Approach 1:
The patent segments the genome analysis task into two parts: using NGS for high-throughput sequencing of short fragments, and using genome mapping with optical maps for long-range structural information. This segmentation allows each method to operate in its optimal range - NGS for base-level detail and optical mapping for large-scale structural variants.
Solution Approach 2:
The patent merges NGS data with genome mapping data to achieve comprehensive genome analysis. By combining the high-resolution base sequence information from NGS with the long-range structural information from optical mapping, the system overcomes the limitations of both individual methods.
2Length of moving object
If long-read sequencing techniques are used to obtain base sequences of tens of thousands of bases, then the read length increases sufficient for structural variant detection, but the ability to identify positions of repetitive sequences deteriorates
Solution Approach 1:
The patent segments the information extraction into two layers: long-read sequencing provides the structural framework and identifies large variants, while genome mapping with short labeled sequences provides precise position identification even in repetitive regions.
Solution Approach 2:
The patent introduces genome mapping as an intermediary method that bridges the gap between long-read sequencing and precise position identification. The optical map serves as a reference framework that helps anchor and verify the positions of long reads, particularly in repetitive regions where long reads alone are ambiguous.
3Measurement precision
If genome mapping is used to identify positions on the genome by labeling specific base sequences, then the ability to detect structural variants improves, but errors in measurement data including apparent expansion and contraction of DNA fragments can lead to erroneous recognition of structural variants
Solution Approach 1:
The patent implements a feedback mechanism where NGS data is used to verify and correct genome mapping results. When apparent expansions or contractions are detected in optical mapping, the system cross-references with NGS data to determine whether these represent true structural variants or measurement artifacts.
Solution Approach 2:
The patent prepares for potential measurement errors by using multiple independent measurement methods (NGS and genome mapping) that can compensate for each other's weaknesses. This redundant approach cushions against the reliability issues inherent in any single method.
4Productivity
If alignment is performed on measurement data with a large number of errors, then the alignment process completes, but abnormalities in labeled positions caused by errors may be erroneously recognized as structural variants
Solution Approach 1:
The patent uses a feedback loop where initial alignment results are evaluated for consistency, and suspicious regions are re-examined using the other data type. This feedback mechanism prevents erroneous recognition of artifacts as true variants while maintaining overall processing efficiency.
Solution Approach 2:
The patent implements a dynamic alignment strategy that adjusts the stringency and methods used based on the quality and characteristics of the input data. For regions with high error rates or ambiguous signals, the system dynamically switches to more conservative alignment criteria or uses the complementary data type for verification.
Data Source
AI summary
Labeled positions alignment capable of dealing with apparent expansion and contraction of a target nucleic acid sequence is performed. An information processing device calculates first ratios of intervals between partial sequences in a reference nucleic acid sequence, constructs an index indicating a combination of the first ratios and information indicating a position of a partial sequence in the nucleic acid sequence corresponding to the combination of the first ratios, calculates second ratios of intervals between partial sequences in a target nucleic acid sequence, extracts a combination of the first ratios corresponding to a combination of the second ratios based on a comparison result between the combination of the second ratios and the combination of the first ratios indicated by the index, and outputs information indicating a position of a partial sequence corresponding to the extracted combination of the first ratios in the reference nucleic acid sequence.


