RNA Alignment via Reduced Reference Genome for Amplicon Sequencing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional RNA alignment methods are inefficient and computationally demanding, particularly when dealing with amplicon sequencing data, due to issues like primer artifacts, false-positive amplifications, and non-uniform coverage, which require large amounts of RAM and are not well-suited for many computer systems, and they complicate workflows by necessitating additional analyses to identify transcript targets.
Innovation Solution
A computer-implemented method that aligns RNA by receiving primer and transcript sequences, generates a modified reference genome excluding non-amplifiable sequences, and aligns sequence reads to this genome, reducing computational demands and streamlining the process by using a microprocessor to determine target sequences and generate an alignment profile.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional RNA alignment methods are used to align reads to a complete reference genome, then comprehensive transcript identification is achieved, but computational resource requirements (RAM) increase significantly to 32 gigabytes
Solution Approach 1:
The patent extracts only the relevant portions of the reference genome that correspond to amplicon targets defined by primer sequences. Instead of using the complete reference genome, a reduced reference genome is constructed containing only the sequences that can be amplified by the specified primers, thereby reducing RAM requirements while maintaining alignment accuracy for relevant transcripts.
Solution Approach 2:
The reference genome is segmented into discrete amplicon regions based on primer binding sites. Each amplicon region is independently defined and indexed, allowing the alignment algorithm to process only these segmented regions rather than the entire genome, thus reducing computational resource requirements.
2Measurement precision
If conventional RNA alignment methods align reads to a complete reference genome, then all possible transcripts can be identified, but additional analysis steps are required to identify transcript targets
Solution Approach 1:
The patent performs preliminary action by pre-defining amplicon regions and creating a reduced reference genome that is specifically tailored for amplicon sequencing data. This preliminary preparation eliminates the need for additional analysis steps later, as the reference genome is already optimized to directly identify transcript targets from amplicon reads.
Solution Approach 2:
The reduced reference genome serves multiple functions simultaneously: it acts as the alignment reference, defines transcript boundaries, and directly identifies transcript targets. This multi-functionality eliminates the need for separate analysis steps that would otherwise be required to identify which transcripts the aligned reads correspond to.
3Adaptability or versatility
If conventional RNA alignment methods are used, then standard whole-genome alignment analyses can be applied, but primer artifacts and false-positive amplifications complicate the data interpretation
Solution Approach 1:
The patent applies local quality by creating a reference genome that is specifically optimized for amplicon sequencing rather than using a generic whole-genome reference. The reduced reference genome incorporates local characteristics of amplicon regions, including primer binding sites and expected amplicon structures, enabling the alignment algorithm to properly interpret amplicon-specific features and reduce false positives.
Solution Approach 2:
The reduced reference genome acts as an intermediary that mediates between the raw amplicon sequencing data and the biological interpretation. It translates the primer-artifact-containing amplicon reads into meaningful transcript information by providing a reference structure that accounts for primer binding and amplification characteristics, thereby filtering out false positives during alignment.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Provided is a computer-implemented method of aligning RNA including receiving onto a data storage unit primer sequences and transcript sequences transcribable from a reference genome based on a gene model, generating target sequences to be amplified from a combination of the primer sequences and the transcript sequences, generating a modified reference genome based on the plurality of target sequences, aligning sequence reads generated from a test sample comprising RNA amplicon molecules to the of target sequences, and generating an alignment profile for the test sample based on the aligning. Also provided is a computer system for performing the foregoing method.