Genomic Alignment via Interval Ratios for Structural Variant Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current DNA sequencing techniques, such as next-generation sequencing (NGS), are limited in reading lengths and struggle to accurately detect structural variants due to errors in measurement data, including apparent expansion and contraction of DNA fragments, which can lead to erroneous recognition of structural variants.

Innovation Solution

An information processing device with a processor and memory that calculates ratios of intervals between partial sequences in both reference and target nucleic acid sequences, constructs an index to handle apparent expansion and contraction, and aligns labeled positions by comparing these ratios to identify accurate positions on the reference genome.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If current DNA sequencing techniques (NGS) are used to read base sequences, then the reading process is fast and cost-effective, but the read length is limited to about hundreds of bases making it impossible to capture large genome regions for structural variant detection

Engineering Contradiction:
Improvesequencing throughputVSAvoidread length
Core Design Contradiction:
ProductivityVSLength of moving object

Solution Approach 1:

The patent segments the genome analysis task into two parts: using NGS for high-throughput sequencing of short fragments, and using genome mapping with optical maps for long-range structural information. This segmentation allows each method to operate in its optimal range - NGS for base-level detail and optical mapping for large-scale structural variants.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges NGS data with genome mapping data to achieve comprehensive genome analysis. By combining the high-resolution base sequence information from NGS with the long-range structural information from optical mapping, the system overcomes the limitations of both individual methods.

Inventive Principle:
Principle #5Merging (Combining)

2Length of moving object

If long-read sequencing techniques are used to obtain base sequences of tens of thousands of bases, then the read length increases sufficient for structural variant detection, but the ability to identify positions of repetitive sequences deteriorates

Engineering Contradiction:
Improveread lengthVSAvoidposition identification accuracy
Core Design Contradiction:
Length of moving objectVSMeasurement precision

Solution Approach 1:

The patent segments the information extraction into two layers: long-read sequencing provides the structural framework and identifies large variants, while genome mapping with short labeled sequences provides precise position identification even in repetitive regions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces genome mapping as an intermediary method that bridges the gap between long-read sequencing and precise position identification. The optical map serves as a reference framework that helps anchor and verify the positions of long reads, particularly in repetitive regions where long reads alone are ambiguous.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If genome mapping is used to identify positions on the genome by labeling specific base sequences, then the ability to detect structural variants improves, but errors in measurement data including apparent expansion and contraction of DNA fragments can lead to erroneous recognition of structural variants

Engineering Contradiction:
Improveposition measurement accuracyVSAvoidstructural variant detection reliability
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent implements a feedback mechanism where NGS data is used to verify and correct genome mapping results. When apparent expansions or contractions are detected in optical mapping, the system cross-references with NGS data to determine whether these represent true structural variants or measurement artifacts.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent prepares for potential measurement errors by using multiple independent measurement methods (NGS and genome mapping) that can compensate for each other's weaknesses. This redundant approach cushions against the reliability issues inherent in any single method.

Inventive Principle:
Principle #11Beforehand cushioning (Prior cushioning)

4Productivity

If alignment is performed on measurement data with a large number of errors, then the alignment process completes, but abnormalities in labeled positions caused by errors may be erroneously recognized as structural variants

Engineering Contradiction:
Improvealignment processing speedVSAvoidstructural variant detection accuracy
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent uses a feedback loop where initial alignment results are evaluated for consistency, and suspicious regions are re-examined using the other data type. This feedback mechanism prevents erroneous recognition of artifacts as true variants while maintaining overall processing efficiency.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent implements a dynamic alignment strategy that adjusts the stringency and methods used based on the quality and characteristics of the input data. For regions with high error rates or ambiguous signals, the system dynamically switches to more conservative alignment criteria or uses the complementary data type for verification.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240331804A1Information processing device and information processing method
Publication Date: 2024.10.03 HITACHI LTD
  • US20240331804A1 patent drawing
  • US20240331804A1 patent drawing
  • US20240331804A1 patent drawing

AI summary

Labeled positions alignment capable of dealing with apparent expansion and contraction of a target nucleic acid sequence is performed. An information processing device calculates first ratios of intervals between partial sequences in a reference nucleic acid sequence, constructs an index indicating a combination of the first ratios and information indicating a position of a partial sequence in the nucleic acid sequence corresponding to the combination of the first ratios, calculates second ratios of intervals between partial sequences in a target nucleic acid sequence, extracts a combination of the first ratios corresponding to a combination of the second ratios based on a comparison result between the combination of the second ratios and the combination of the first ratios indicated by the index, and outputs information indicating a position of a partial sequence corresponding to the extracted combination of the first ratios in the reference nucleic acid sequence.