Compressed Molecular Tagged Nucleic Acid Sequence Data Fusion Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for analyzing molecular tagged nucleic acid sequence data are inefficient in terms of storage and processing, particularly for detecting genetic fusions in cancer cells, as they require large amounts of memory and are not optimized for compressed data analysis.

Innovation Solution

A method and system for compressing molecular tagged nucleic acid sequence data by determining consensus sequence reads and alignments based on flow space signal measurements, generating a compressed data structure, and using this structure to detect fusions, thereby reducing memory requirements and enhancing detection efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If molecular tagged nucleic acid sequence data is stored and processed using conventional methods, then fusion detection can be performed, but large amounts of memory are required and processing efficiency is low

Engineering Contradiction:
Improvefusion detection capabilityVSAvoidmemory requirements
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts only the essential information from the original sequence data by determining consensus sequences and consensus alignments. This extraction process removes redundant data while preserving the critical fusion detection information, thereby reducing memory requirements while maintaining detection capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

Instead of processing all raw sequence data to detect fusions, the patent inverts the approach by first compressing the data into consensus representations and then performing fusion detection on the compressed data. This inversion enables efficient processing with reduced memory footprint.

Inventive Principle:
Principle #13The other way round (Inversion)

2Quantity of substance

If molecular tagged nucleic acid sequence data is compressed into consensus data structures, then memory requirements are reduced, but processing complexity increases

Engineering Contradiction:
Improvedata volumeVSAvoidprocessing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent performs preliminary compression of the sequence data into consensus sequences and consensus alignments before fusion detection. This preliminary action organizes the data in advance, making subsequent fusion detection operations simpler and more efficient despite the initial compression complexity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates simplified copies of the original data in the form of consensus sequences and alignments. These consensus copies capture the essential information in a condensed format, reducing data volume while enabling efficient analysis through the copied representation.

Inventive Principle:
Principle #26Copying

3Quantity of substance

If consensus sequence reads and alignments are determined from flow space signal measurements, then data compression is achieved, but processing time increases

Engineering Contradiction:
Improvedata compression ratioVSAvoidprocessing time
Core Design Contradiction:
Quantity of substanceVSLoss of time

Solution Approach 1:

The patent merges multiple sequence reads into a single consensus sequence and multiple alignments into a single consensus alignment for each molecular tag family. This merging process compresses the data by combining redundant information while maintaining the essential fusion detection signal, achieving compression without excessive time penalty.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20240203525A1Methods for detection of fusions using compressed molecular tagged nucleic acid sequence data
Publication Date: 2024.06.20 LIFE TECHNOLOGIES CORP
  • US20240203525A1 patent drawing
  • US20240203525A1 patent drawing
  • US20240203525A1 patent drawing

AI summary

A method for compressing nucleic acid sequence data wherein each sequence read is associated with a molecular tag sequence, wherein a portion of the sequence reads alignments correspond to sequence reads mapped to a targeted fusion reference sequence includes determining a consensus sequence read for each family of sequence reads based on flow space signal measurements corresponding to the family of sequence reads, determining a consensus sequence alignment for each family of sequence reads, wherein a portion of the consensus sequence alignments correspond to the consensus sequence reads aligned with the targeted fusion reference sequence, generating a compressed data structure comprising consensus compressed data, the consensus compressed data including the consensus sequence read and the consensus sequence alignment for each family, and detecting a fusion using the consensus sequence reads and the consensus sequence alignments from the compressed data structure.