Metagenome Strain Profiling via L1-L2 Index Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current metagenomic analysis techniques face challenges in accurately profiling microbial samples at the strain level due to high genomic similarity, limited read support, and the need for costly and complex pre-processing, especially when updating reference databases, which leads to estimation errors and inefficiencies.

Innovation Solution

A system and method utilizing L1-L2 indexing techniques to generate L1 and L2 indices from strain-level sequences, allowing for non-homology based, alignment-free profiling by segregating reference microbial sequences into pre-defined chunks and using k-merization to map query k-mers, enabling accurate abundance estimation at the strain level.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If homology based microbial profiling approach is used, then microbial sequences can be analyzed based on similarity, but complex computation and intensive pre-processing are required

Engineering Contradiction:
Improvemicrobial profiling accuracyVSAvoidcomputation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The reference genome collection is divided into multiple non-overlapping chunks, each indexed separately. This segmentation reduces the computational complexity of indexing and querying by breaking down the large-scale homology search into smaller, manageable sub-problems, while maintaining comprehensive coverage of the reference database

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The reference genome collection is pre-processed into chunks and indexed before actual metagenomic analysis. This preliminary action includes creating chunk indices that store pre-computed homology information, eliminating the need for intensive pre-processing during each new metagenomic sample analysis

Inventive Principle:
Principle #10Preliminary action

2Productivity

If marker sequences are used for metagenomic analysis, then taxonomic composition can be estimated, but accurate strain level profiling is difficult due to limited read support

Engineering Contradiction:
Improvetaxonomic composition estimationVSAvoidstrain level profiling accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The approach transitions from one-dimensional marker sequence analysis to multi-dimensional whole genome chunk analysis. By utilizing entire genome chunks instead of limited marker sequences, the method increases the dimensionality of information used for profiling, thereby improving strain level discrimination capability while maintaining taxonomic composition estimation

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Adaptability or versatility

If reference database is updated with new sequences, then comprehensive microbial coverage is improved, but significant re-indexing and pre-processing is required

Engineering Contradiction:
Improvereference database coverageVSAvoidre-indexing time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The reference database is segmented into independent chunks that can be updated individually. When new sequences are added, only the relevant chunks need to be re-indexed rather than the entire database, significantly reducing the time and computational resources required for database updates while maintaining comprehensive microbial coverage

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The reference database is pre-organized into chunks with pre-computed indices. This preliminary structuring allows for efficient incremental updates where only affected chunks require re-indexing, rather than requiring complete re-indexing of the entire database when updates occur

Inventive Principle:
Principle #10Preliminary action

4Device complexity

If homology approaches work on selected subset of markers, then computation is simplified, but estimation errors increase due to reduced read support

Engineering Contradiction:
Improvecomputation complexityVSAvoidabundance estimation accuracy
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

Instead of using a selected subset of markers, the genome is segmented into multiple comprehensive chunks that collectively cover the entire reference collection. This segmentation approach maintains computational efficiency by dividing the search space while ensuring that abundance estimation utilizes reads mapping to any chunk, thereby increasing total read support and reducing estimation errors

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4086912B1A method and a system for profiling of a metagenome sample
Publication Date: 2024.07.10 TATA CONSULTANCY SERVICES LTD
  • EP4086912B1 patent drawingFigure 1
  • EP4086912B1 patent drawingFigure 2
  • EP4086912B1 patent drawingFigure 3A

AI summary

This disclosure relates generally to a method and a system for profiling of metagenome samples. Most state of-art techniques for metagenomic profiling use homology-based, curated database of identified marker sequences generated after complex and costly pre-processing. The disclosed method and system for profiling of metagenome samples are a non-homology based, a non-marker based and an alignment free strain level profiling tools for microbe profiling. The disclosure works with a several k-mer based indexing techniques for constructing a compact and comprehensive multi-level indexing, wherein the multi-level indexing includes a LI-Index and a L2-Index. The multi-level indexing is used for profiling metagenomics by abundance estimation, wherein the abundance estimation includes a relative abundance and an absolute abundance.