Sequence Clustering Using Reference IDs for Faster Virus Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing databases for virus detection in bio-pharmaceuticals face challenges in efficiently clustering nucleotide sequences due to sequence duplications, leading to difficulty in identifying virus contamination.
Innovation Solution
A sequence clustering method that includes steps for read acquisition, ID group formation, high-frequency read identification, relevant reference information acquisition, and cluster formation, utilizing a computer system to measure similarity and reduce the number of references for rapid analysis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If databases contain many duplicate base sequences to ensure high comprehensiveness and low risk of failing to detect viruses, then detection reliability is improved, but device complexity and analysis difficulty increase due to multiple matches per read
Solution Approach 1:
The patent merges multiple duplicate reference sequences that represent the same virus into a single clustered reference sequence. By grouping references with identical or highly similar base sequences and assigning them a common cluster ID, the system maintains comprehensive virus detection coverage while eliminating redundant matches during read analysis, thus reducing database complexity without sacrificing detection reliability
Solution Approach 2:
The patent creates a simplified copy of the database structure by generating cluster IDs that represent groups of duplicate references. Instead of storing and processing all original duplicate references, the system uses these cluster ID copies to efficiently track and analyze virus-related sequences, reducing computational complexity while preserving detection capability
2Reliability
If databases include all sequenced virus base sequences to ensure comprehensive coverage, then detection sensitivity is improved, but processing time increases due to the enormous number of sequences to examine
Solution Approach 1:
The patent combines multiple duplicate virus reference sequences into single clustered entries with unique cluster IDs. This merging process maintains comprehensive virus detection coverage by preserving all unique virus types while dramatically reducing the total number of sequences that must be processed during read matching, thus decreasing search processing time without compromising detection sensitivity
Solution Approach 2:
The patent performs preliminary clustering of reference sequences before the actual virus detection process. By pre-grouping duplicate sequences and assigning cluster IDs in advance, the system prepares a streamlined database structure that enables faster processing during subsequent read analysis, reducing the time required for virus detection while maintaining comprehensive coverage
3Reliability
If databases store numerous duplicate base sequences from different specimens, then comprehensiveness is improved, but ease of operation deteriorates due to difficulty in aggregation and analysis
Solution Approach 1:
The patent merges duplicate base sequences from different specimens into clustered groups represented by unique cluster IDs. This approach maintains comprehensive virus detection coverage by preserving all unique virus types while simplifying data aggregation and analysis operations, as users can now easily group and analyze results by cluster ID rather than dealing with numerous individual duplicate entries
Solution Approach 2:
The patent creates cluster IDs that serve multiple functions simultaneously: they uniquely identify virus types, group duplicate sequences, and facilitate efficient data aggregation and analysis. This universal identifier system enables comprehensive database coverage while greatly improving ease of operation for various analytical tasks
Data Source
Figure 1
Figure 2A
Figure 2B
AI summary
[Problem to be Solved] To provide a sequence clustering method that can easily cluster sequences of sequences included in a database in relation to specimen-derived reads. [Solution] A sequence clustering method comprising: a read acquisition step of acquiring reads derived from specimens; a read ID acquisition step of acquiring identification information on the reads; an ID group acquisition step of checking reads against references included in a database and acquiring ID groups; a high-frequency read ID acquisition step of obtaining the number of pieces of read identification information included in ID groups and obtaining high-frequency read identification information; a high-frequency read identification information acquisition step of obtaining high-frequency reference identification information; a relevant reference information acquisition step of acquiring, using high-frequency reference identification information, information including organism species of high-frequency reference sequences stored in association with the high-frequency reference identification information from the database; a reference cluster step of allowing high-frequency reference identification information to be replaced with first cluster identification information; and a replaced ID group acquisition step of converting high-frequency reference identification information included in ID groups into first cluster identification information.