Pangenome Read Mapping with Graph Embedding and Winnowing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional genome analysis methods relying on linear references are time-consuming and memory-intensive, limiting accurate and efficient mapping of read sequences, especially when dealing with large-scale non-linear genome variation graphs.

Innovation Solution

A system and method utilizing graph embedding and winnowing techniques to generate an index for genome variation graphs, enabling efficient and accurate mapping of read sequences by constructing subgraphs for optimal alignment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional linear reference methods are used for read sequence mapping, then the mapping process is straightforward and simple, but the time consumption and memory requirements increase significantly

Engineering Contradiction:
Improvemapping simplicityVSAvoidmapping time
Core Design Contradiction:
Ease of operationVSLoss of time

Solution Approach 1:

The patent segments the genome variation graph into smaller subgraphs that can be processed independently. Each read sequence is mapped to relevant subgraphs rather than the entire graph, reducing computational time while maintaining mapping accuracy. This segmentation allows parallel processing and reduces the search space for each mapping operation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary indexing structure that mediates between the read sequences and the genome variation graph. This index pre-processes and organizes graph information, enabling faster lookup and reducing the time required for actual mapping operations without sacrificing the ability to handle complex variations.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If conventional linear reference methods are used for read sequence mapping, then the implementation is simple, but the memory consumption increases significantly

Engineering Contradiction:
Improvesystem complexityVSAvoidmemory usage
Core Design Contradiction:
Device complexityVSQuantity of substance

Solution Approach 1:

By dividing the genome variation graph into manageable subgraphs and creating targeted indexes for each, the patent reduces the amount of data that needs to be loaded into memory simultaneously. This segmentation strategy maintains system complexity at acceptable levels while dramatically reducing peak memory consumption during mapping operations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by pre-processing the genome variation graph to create compressed and optimized index structures before actual mapping occurs. This pre-computation phase organizes data in memory-efficient formats, reducing the memory footprint during the actual mapping process while keeping the overall system manageable.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If graph embedding and winnowing techniques are applied to generate index for genome variation graphs, then the mapping accuracy improves, but the index construction complexity increases

Engineering Contradiction:
Improvemapping accuracyVSAvoidindex construction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical graph traversal methods with graph embedding techniques that map the graph structure into a lower-dimensional space. This substitution maintains high mapping accuracy by preserving topological relationships while simplifying the index construction process and enabling more efficient search operations.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent applies parameter changes by transforming the genome variation graph through embedding parameters that capture essential structural information. This transformation changes the representation parameters from raw graph topology to embedded coordinates, improving mapping accuracy while making the index construction more tractable through dimensionality reduction.

Inventive Principle:
Principle #35Parameter changes

4Productivity

If conventional methods restrict gapped alignment to small regions, then the processing speed increases, but the mapping accuracy decreases

Engineering Contradiction:
Improveprocessing speedVSAvoidmapping accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary action by pre-identifying and indexing relevant regions of the genome variation graph that are likely to contain alignment targets. This pre-processing step enables the system to quickly navigate to appropriate regions without exhaustively searching the entire graph, maintaining high processing speed while ensuring accurate mapping by focusing computational resources on relevant areas.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12417820B2Method and system for mapping read sequences using a pangenome reference
Publication Date: 2025.09.16 TATA CONSULTANCY SERVICES LTD
  • US12417820B2 patent drawing
  • US12417820B2 patent drawing
  • US12417820B2 patent drawing

AI summary

There is a demand for low-cost efficient robust method for mapping read sequences with genome variation graph in genomic study. This disclosure herein relates to a method and system for mapping read sequences with genome variation graph by constructing a subgraph using a novel combination of graph embedding and graph winnowing techniques. The system processes the obtained plurality of read sequences and a genome variation graph for constructing the subgraph by computing an embedding for the genome variation graph utilizing a graph embedding technique. Further, graph index is generated for the genome variation graph based on the embedding and the genome variation graph utilizing the graph winnowing technique. Then computes gapped alignment score for read sequence (r) with its corresponding subgraph. Thus, enables a reliable method for read sequence with accurate, memory efficient and scalable system for mapping read sequences with genome variation graph.