Dynamic Reference Sequence for Nucleic Acid Alignment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional gene sequence analysis methods face reduced alignment accuracy when polymorphism, mutation, or methylation occurs in the target sequences, particularly when using a reference sequence without mutations, and require frequent updates to the analysis program due to changing mutation information.
Innovation Solution
A sequence analysis method that aligns read sequences with a single reference sequence comprising multiple rearrangement sequences, including those with known polymorphisms, mutations, and methylation, allowing for accurate mapping even with new mutations and eliminating the need for program modifications as mutation information evolves.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a reference sequence without mutations is used for alignment, then the alignment process is simple, but alignment accuracy is reduced when polymorphism or mutation occurs in the target sequences
Solution Approach 1:
The reference sequence is segmented into multiple rearrangement sequences, each representing different mutation patterns. This segmentation allows the system to handle complex genetic variations by breaking down the reference into manageable, mutation-specific segments that can be individually compared against read sequences.
Solution Approach 2:
The reference sequence is made dynamic by incorporating multiple rearrangement sequences that can be selectively applied based on detected mutations. The system dynamically updates the reference sequence composition during analysis, adding new rearrangement sequences as mutations are identified, thereby adapting to the specific sample being analyzed.
2Measurement precision
If multiple reference sequences are used to account for mutations, then alignment accuracy improves, but the number of reference sequences increases requiring frequent program updates
Solution Approach 1:
The system performs preliminary actions by pre-defining a template for creating rearrangement sequences and establishing a database structure that can accommodate new mutations. This preliminary setup enables rapid integration of new mutation data without requiring fundamental program changes, as the framework is already in place to receive and process new rearrangement sequences.
Solution Approach 2:
The reference sequence structure is designed with universal applicability through a standardized template that can represent any mutation pattern. This multi-functional design allows the same system architecture to handle diverse mutation types by simply populating the template with new sequence data, eliminating the need for separate processing pathways for different mutation kinds.
3Measurement precision
If the reference sequence is updated frequently to include new mutations, then alignment accuracy is maintained, but the time and resources required for updates increase
Solution Approach 1:
The system uses copying by creating new rearrangement sequences as copies or variations of existing templates rather than developing entirely new reference sequences. This copying approach allows rapid replication of proven sequence structures with minor modifications to accommodate new mutations, significantly reducing the time and computational resources required for updates compared to creating de novo reference sequences.
4Productivity
If a single reference sequence containing multiple rearrangement sequences is used, then the number of reference sequences is reduced, but the complexity of managing the reference sequence increases
Solution Approach 1:
Multiple rearrangement sequences are merged into a single integrated reference sequence structure. This merging consolidates what would otherwise be separate reference files into one unified resource, reducing the number of files to manage and streamlining the alignment process while maintaining all necessary mutation information within the single sequence.
Data Source
AI summary
Disclosed is a sequence analysis method for analyzing nucleic acid sequence, the sequence analysis method including: obtaining a plurality of read sequences read from the nucleic acid sequence; and determining each nucleic acid sequence by aligning each read sequence with reference to a single reference sequence, wherein the reference sequence includes at least a first rearrangement sequence and a second rearrangement sequence that is different from the first rearrangement sequence.


