OTU Delineation via Abundance-Weighted Sequence Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current OTU delineation methods, such as Qiime, Mothur, and Usearch, overestimate the number of operational taxonomic units (OTUs) in microbial community studies, leading to the generation of spurious OTUs and distortion of microbial community composition profiles, which hinders the isolation and verification of functionally important bacteria.
Innovation Solution
A modified approach that involves obtaining relative abundance values of qualified sequences, ranking them from high to low, separating them into high and low abundance groups, and using only high abundance sequences for initial OTU delineation, with low abundance sequences remapped to OTUs only if they have at least 97% sequence similarity, to minimize pseudo OTUs and improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If conventional OTU delineation methods (Qiime, Mothur, Usearch) are used to process all sequences, then complete microbial community coverage is achieved, but the number of OTUs is overestimated and spurious OTUs are generated
Solution Approach 1:
The patent segments the sequence processing into two distinct phases: (1) initial OTU delineation using only high-abundance sequences, and (2) subsequent remapping of low-abundance sequences. This segmentation prevents low-abundance erroneous sequences from creating spurious OTUs while preserving the ability to assign them to existing OTUs, thereby resolving the contradiction between complete coverage and accuracy.
Solution Approach 2:
The patent applies different processing quality standards to different sequence groups. High-abundance sequences undergo standard OTU delineation, while low-abundance sequences undergo a more stringent remapping process with higher similarity thresholds. This local quality differentiation ensures that sequences most likely to be erroneous (low-abundance) receive more rigorous filtering, improving overall OTU accuracy without sacrificing coverage.
2Quantity of substance
If all sequences including low abundance sequences are used for OTU delineation, then comprehensive microbial diversity is captured, but pseudo OTUs are generated and composition profiles are distorted
Solution Approach 1:
The patent performs preliminary OTU delineation using only high-abundance sequences before incorporating low-abundance sequences. This preliminary action establishes a reliable foundation of true OTUs, which then serves as the reference for remapping low-abundance sequences. This preliminary filtering prevents pseudo OTUs from being created in the first place, while still allowing comprehensive diversity capture in the subsequent remapping phase.
Solution Approach 2:
The patent introduces high-abundance sequences as an intermediary layer between raw sequencing data and final OTU composition. These high-abundance sequences act as a filter and reference standard, mediating the incorporation of low-abundance sequences into the final OTU table. This intermediary step ensures that only low-abundance sequences matching established high-confidence OTUs are incorporated, preventing composition distortion.
3Measurement precision
If low abundance sequences are remapped with high similarity threshold (97%), then spurious OTU assignments are reduced, but some valid low abundance OTUs may be missed
Solution Approach 1:
The patent applies dynamic similarity thresholds in the remapping process. The 97% threshold is applied specifically during the remapping phase for low-abundance sequences, which is more stringent than the initial clustering threshold. This dynamic adjustment of stringency based on sequence abundance and processing phase optimizes both accuracy for common species and sensitivity for rare species, resolving the contradiction between precision and detection capability.
Data Source
AI summary
A method in which a microorganism operational taxonomic unit (OTU) in a sample is defined based on a DNA sequence of a system generation information gene of microorganism in the sample. In the method, qualified sequence segments are obtained by means of processing and reading of an original sequence; the segments are sorted according to a relative abundance value of each segment; and only the qualified sequences with the high abundance values are used to obtain the temporary OTU. The qualified sequences with the low abundance values are reallocated; and the qualified sequence can be distributed to the proper temporary OTU respectively when a sequence similarity degree between the qualified sequence and an OTU sequence reaches at least 97%. The present disclosure also provides a sequence-assisted microorganism separation method.


