A high-throughput sequencing data sequence clustering method and system based on cyclic self-blast

By using a cyclic self-blast method to remove redundancy from environmental DNA macrobarcode sequencing data, the problem of the same species being represented by multiple OTUs or ASVs in OTU and ASV methods is solved, generating a concise OTU clustering result set, which improves the accuracy of species annotation and the reliability of diversity analysis.

CN122067616BActive Publication Date: 2026-07-10NANJING NORMAL UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING NORMAL UNIVERSITY
Filing Date
2026-04-03
Publication Date
2026-07-10

AI Technical Summary

Technical Problem

In existing technologies, OTU and ASV methods result in the same species being represented by multiple OTUs or ASVs when processing environmental DNA macrobarcode sequencing data, leading to taxonomic redundancy, increased complexity of downstream analysis, and risk of species abundance overestimation.

Method used

A sequence clustering method based on cyclic self-blast is adopted. By deduplication, sorting, partitioning into subsets and parent sets, blast multiple alignment and cyclic alignment are performed, and sequences that meet the similarity and coverage criteria are merged to generate a concise OTU clustering result set.

Benefits of technology

It significantly reduces taxonomic redundancy, improves the accuracy of species annotation and the reliability of diversity analysis, and reduces species abundance overestimation and community pattern misinterpretation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122067616B_ABST
    Figure CN122067616B_ABST
Patent Text Reader

Abstract

The application discloses a high-throughput sequencing data sequence clustering method and system based on cyclic self-blast: the high-throughput sequencing data is de-duplicated; unique sequences with a number of repetitions lower than a threshold value a are filtered, and the filtered unique sequences are sorted according to the number of repetitions; the sorted unique sequences are divided into subsets and a parent set; the subsets are subjected to blast multiple alignment to obtain a de-redundant subset; the parent set and the de-redundant subset are subjected to cyclic blast alignment to obtain a de-redundant parent set; all representative sequences in a preliminary clustering set are combined, sorted according to the number of repetitions and distributed with identification tags to generate a final OTU clustering result set; in the prior art, when OTU clustering and ASV methods are used to process environmental DNA macro-barcode sequencing data, different OTU or ASV sequences are annotated to the same species, so that a large amount of redundancy still exists in the classified units after clustering, and the reliability of species analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Genome decontaminating method

    CN119068986A

  • Method and device for estimating repeating sequence content of genome

    WO2013097149A1