Pangenome Read Mapping with Graph Embedding and Winnowing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional genome analysis methods relying on linear references are time-consuming and memory-intensive, limiting accurate and efficient mapping of read sequences, especially when dealing with large-scale non-linear genome variation graphs.
Innovation Solution
A system and method utilizing graph embedding and winnowing techniques to generate an index for genome variation graphs, enabling efficient and accurate mapping of read sequences by constructing subgraphs for optimal alignment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional linear reference methods are used for read sequence mapping, then the mapping process is straightforward and simple, but the time consumption and memory requirements increase significantly
Solution Approach 1:
The patent segments the genome variation graph into smaller subgraphs that can be processed independently. Each read sequence is mapped to relevant subgraphs rather than the entire graph, reducing computational time while maintaining mapping accuracy. This segmentation allows parallel processing and reduces the search space for each mapping operation.
Solution Approach 2:
The patent introduces an intermediary indexing structure that mediates between the read sequences and the genome variation graph. This index pre-processes and organizes graph information, enabling faster lookup and reducing the time required for actual mapping operations without sacrificing the ability to handle complex variations.
2Device complexity
If conventional linear reference methods are used for read sequence mapping, then the implementation is simple, but the memory consumption increases significantly
Solution Approach 1:
By dividing the genome variation graph into manageable subgraphs and creating targeted indexes for each, the patent reduces the amount of data that needs to be loaded into memory simultaneously. This segmentation strategy maintains system complexity at acceptable levels while dramatically reducing peak memory consumption during mapping operations.
Solution Approach 2:
The patent performs preliminary actions by pre-processing the genome variation graph to create compressed and optimized index structures before actual mapping occurs. This pre-computation phase organizes data in memory-efficient formats, reducing the memory footprint during the actual mapping process while keeping the overall system manageable.
3Measurement precision
If graph embedding and winnowing techniques are applied to generate index for genome variation graphs, then the mapping accuracy improves, but the index construction complexity increases
Solution Approach 1:
The patent replaces complex mechanical graph traversal methods with graph embedding techniques that map the graph structure into a lower-dimensional space. This substitution maintains high mapping accuracy by preserving topological relationships while simplifying the index construction process and enabling more efficient search operations.
Solution Approach 2:
The patent applies parameter changes by transforming the genome variation graph through embedding parameters that capture essential structural information. This transformation changes the representation parameters from raw graph topology to embedded coordinates, improving mapping accuracy while making the index construction more tractable through dimensionality reduction.
4Productivity
If conventional methods restrict gapped alignment to small regions, then the processing speed increases, but the mapping accuracy decreases
Solution Approach 1:
The patent performs preliminary action by pre-identifying and indexing relevant regions of the genome variation graph that are likely to contain alignment targets. This pre-processing step enables the system to quickly navigate to appropriate regions without exhaustively searching the entire graph, maintaining high processing speed while ensuring accurate mapping by focusing computational resources on relevant areas.
Data Source
AI summary
There is a demand for low-cost efficient robust method for mapping read sequences with genome variation graph in genomic study. This disclosure herein relates to a method and system for mapping read sequences with genome variation graph by constructing a subgraph using a novel combination of graph embedding and graph winnowing techniques. The system processes the obtained plurality of read sequences and a genome variation graph for constructing the subgraph by computing an embedding for the genome variation graph utilizing a graph embedding technique. Further, graph index is generated for the genome variation graph based on the embedding and the genome variation graph utilizing the graph winnowing technique. Then computes gapped alignment score for read sequence (r) with its corresponding subgraph. Thus, enables a reliable method for read sequence with accurate, memory efficient and scalable system for mapping read sequences with genome variation graph.


