Genomic data redundancy elimination method
By combining genomic similarity and quality evaluation indicators, similar connection clusters are identified and ranked, and representative genomes are selected. This solves the problem of low efficiency in deduplication of genomic data in existing technologies, and achieves efficient and accurate genomic data processing and analysis.
Patent Information
- Application Number
- CN202511746991.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-27
AI Technical Summary
Existing methods for removing redundancy from genomic data are inefficient, cannot process tens of thousands of genomes quickly, require a large amount of memory, cannot combine similarity and coverage for accurate screening, and lack quality standards for optimizing the sorting and representative screening of genomes within clusters.
By calculating genome similarity, establishing connections, identifying similar connection clusters, and ranking them based on quality evaluation indicators, the genomes with the highest quality are selected as cluster centers to form a representative genome set.
It enables rapid and efficient deduplication of genomic data, reduces computational overhead, outputs high-quality representative genome sets, improves analytical accuracy and efficiency, and is suitable for building standard databases.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] This invention relates to the fields of bioinformatics, genomics and related data processing technologies, and in particular to a method for deduplicating genomic data. Background Technology
[0002] With the widespread application of high-throughput sequencing technology, a large amount of genomic data from various organisms, including microorganisms, plants, animals, and humans, is constantly being generated and accumulated. Especially in the field of microbiology, different research institutions or projects often perform repeated sequencing on the same or highly similar strains or populations, resulting in serious redundancy in genomic data.
[0003] Redundant genomic data not only increases storage costs but also interferes with subsequent analytical results. For example, in species annotation, phylogenetic analysis, or metagenomic research, redundant data reduces the representativeness and efficiency of the analysis and interferes with species identification. Therefore, effectively clustering and deduplicating massive amounts of genomic data to select representative reference genomes is of great significance for constructing standard databases and improving the efficiency of bioinformatics processing.
[0004] Existing methods such as genome nucleotide consistency or shared gene sets are used to remove redundancy and select representative genomes, but these methods often have the following problems: Redundancy removal algorithms are inefficient and cannot quickly process tens of thousands of genomes, which is time-consuming. The redundancy removal algorithm requires a large amount of memory, at the terabyte level, and cannot be performed on a regular server. It is not possible to combine similarity (such as ANI) and coverage (such as AF, Alignment Fraction) simultaneously for precise filtering; There is a lack of mechanisms to combine quality standards for optimizing the sequencing and representative screening of genomes within clusters.
[0005] Therefore, there is an urgent need for a fast, efficient, and standardized method to achieve genome redundancy removal and representative genome selection in order to adapt to the processing and utilization of large-scale genomic data. Summary of the Invention
[0006] This invention covers the following technical solutions: One aspect of the present invention relates to a method for deduplicating genomic data, comprising the following steps: 1) Calculate the similarity of each pair of genomes to be processed, and retain genome pairs that simultaneously meet the preset similarity threshold conditions; 2) Establish connections based on genome pairs that meet the similarity threshold, identify all directly or indirectly connected node sets, and form several non-overlapping similar connection clusters; 3) For each similar connection cluster, the genomes are sorted according to the preset quality evaluation index, and the genome with the highest quality is selected as the cluster center. Genomes with similarity not lower than the set threshold are assigned to the same cluster. The above steps are repeated for genomes that are not assigned to any cluster until all genomes in the cluster have been clustered. The quality evaluation index is used to comprehensively evaluate the integrity, continuity and contamination of genome assembly. 4) The cluster centers in each cluster are used as representative genome outputs to form a non-redundant representative genome set.
[0007] Another aspect of the invention relates to a computer-readable medium for storing computer instructions, programs, code sets, or instruction sets that, when run on a computer, cause the computer to perform the methods described above.
[0008] Another aspect of the present invention relates to an electronic device comprising: One or more processors; and A computer-readable storage medium for storing computer instructions, programs, code sets, or instruction sets that, when executed on a computer, cause the one or more processors to implement the methods described above.
[0009] The method described in this invention combines connection relationship construction based on similarity thresholds with recursive clustering based on quality indicators to achieve automatic grouping and redundancy elimination of large amounts of genomic data. This method can effectively reduce the amount of redundant data and computational overhead while ensuring optimal representative genome quality, and output a stable and structured representative genome set.
[0010] The representative genome set obtained by the above-mentioned redundancy-removing clustering method is of high quality and representativeness, and can be directly used to construct a standardized species reference database. It can also improve the accuracy and computational efficiency of analysis in bioinformatics applications such as metagenomic annotation, species identification, pangenome analysis and phylogenetic tree construction, thereby significantly expanding the application value of this method. Detailed Implementation
[0011] Reference will now be made to detailed embodiments of the present invention, one or more of which are described below. Each example is provided for explanation and not for limitation of the invention. In fact, it will be apparent to those skilled in the art that various modifications and variations can be made to the invention without departing from its scope or spirit. For example, features described or illustrated as part of one embodiment may be used in another embodiment to produce further embodiments.
[0012] Unless otherwise stated, all terms used to disclose this invention (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Further guidance is provided below for a better understanding of the teachings of this invention. The terminology used herein in the specification of this invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0013] In this invention, unless otherwise stated, the scientific and technical terms used herein have the meanings commonly understood by those skilled in the art. Furthermore, the genomics, molecular biology, and bioinformatics terms used herein are all widely used terms in their respective fields.
[0014] The terms “containing,” “comprising,” and “including” as used in this invention are synonyms and are inclusive or open-ended, not excluding additional, uncited members, elements, or method steps.
[0015] In this invention, the numerical range represented by endpoints includes all numerical values and fractions contained within that range, as well as the endpoints mentioned.
[0016] As used in this invention, the term "about" or "approximately" means within 20%, preferably within 10%, and more preferably within 5%, of a given value or range. It also includes specific numbers, such as about 20 including 20.
[0017] Furthermore, in describing representative embodiments of the invention, this specification may present the methods and / or processes of the invention as a specific sequence of steps. However, the method or process should not be limited to the specific order of the steps described herein, to the extent that the method or process does not depend on the specific order of the steps presented herein. As will be understood by those skilled in the art, other sequences of steps are also possible. Therefore, the specific order of steps presented in the specification should not be construed as a limitation of the claims. Additionally, the claims relating to the methods and / or processes of the invention should not be limited to the execution of their steps in the order they are written, and those skilled in the art will readily recognize that the sequence can be changed while still remaining within the spirit and scope of the invention.
[0018] As used in this invention, unless otherwise stated, the singular forms of the articles “a,” “an,” and “the” include plural referents.
[0019] In this invention, the terms "multiple" or "various" are used unless otherwise specified, referring to a quantity of 2 or more.
[0020] In this invention, the technical features described in an open-ended manner include both closed-ended technical solutions composed of the listed features and open-ended technical solutions that include the listed features.
[0021] In this invention, terms such as "preferred," "better," "more suitable," and "ideal" merely describe implementation methods or embodiments with better effects and should be understood not to limit the scope of protection of this invention. In this invention, terms such as "optionally," "optionally," and "optional" mean that something is optional, that is, selected from either "with" or "without" a parallel solution. If multiple "optional" statements appear in a technical solution, unless otherwise specified and without contradiction or mutual constraint, each "optional" statement is independent.
[0022] The first aspect of the present invention relates to a method for deduplicating genomic data, comprising the following steps: 1) Calculate the similarity of each pair of genomes to be processed, and retain genome pairs that simultaneously meet the preset similarity threshold conditions; 2) Establish connections based on genome pairs that meet the similarity threshold, identify all directly or indirectly connected node sets, and form several non-overlapping similar connection clusters; 3) For each similar connection cluster, the genomes are sorted according to the preset quality evaluation index, and the genome with the highest quality is selected as the cluster center. Genomes with similarity not lower than the set threshold are assigned to the same cluster. The above steps are repeated for genomes that are not assigned to any cluster until all genomes in the cluster have been clustered. The quality evaluation index is used to comprehensively evaluate the integrity, continuity and contamination of genome assembly. 4) The cluster centers in each cluster are used as representative genome outputs to form a non-redundant representative genome set.
[0023] This invention has at least the following technical effects: a) Grouping and classifying redundant genomes by screening with similarity thresholds and identifying connected components.
[0024] By calculating genomic similarity and retaining genomic pairs that meet the threshold conditions to form similar connection clusters, genomes with high similarity in the overall dataset are automatically assigned to the same candidate cluster, reducing irrelevant alignment calculations and improving clustering efficiency and accuracy.
[0025] b) By ranking and recursive clustering based on quality evaluation indicators, we ensure that the quality of the genome represented in each cluster is optimal.
[0026] Genomes were ranked according to quality evaluation indicators, and the highest-quality genomes were selected as cluster centers. This ensured that the representative genomes of each cluster were optimal in terms of completeness, continuity, and low contamination, guaranteeing the reliability and biological quality of the representative data. The above steps were repeated for genomes that did not belong to any cluster, forming a recursive clustering process. This avoided omissions and duplicate classifications, ensuring that the clustering results were complete and non-overlapping.
[0027] c) Achieve redundant data compression and structured output of results through the representative genome output step.
[0028] Ultimately, only the cluster centers of each cluster are output as representative genomes, significantly reducing the amount of data to be stored and analyzed, and achieving a structured set of results after redundancy removal. This result set has the following characteristics: high representativeness, no redundancy (only one representative is retained for each cluster), good traceability (each redundant item can be traced back to its corresponding representative), and suitability for building a reference database.
[0029] In a preferred embodiment, this method exhibits significant technical advantages in terms of efficiency and resource consumption: by avoiding the construction of fully connected matrices and complex hierarchical clustering, and establishing connections only on genome pairs that meet the threshold and performing intra-cluster recursive clustering, the task, which originally required TB-level memory and approximately one week of runtime, can be reduced to tens of GB of memory and several hours of runtime. In an exemplary embodiment test targeting 10,000 genomes, memory consumption was controlled at approximately 1 GB and runtime at approximately 1 minute. Furthermore, this method offers configurable similarity thresholds and quality evaluation metrics, adapting to different data types and heterogeneous backgrounds, and possesses good scalability and portability. The unified sorting and clustering rules ensure consistent output results and a high degree of automation, directly supporting subsequent large-scale analysis and database construction needs.
[0030] In some implementations, the similarity is defined as Average Nucleotide Identity (ANI) and / or Alignment Fraction (AF). ANI reflects the overall similarity between two genome sequences at the nucleotide level and is a commonly used indicator for determining genomic relationships. AF describes the proportion of effectively aligned regions between two genomes within a reference sequence, reflecting the coverage of similar regions. By combining ANI and AF, genomic similarity can be assessed more accurately, avoiding misjudgments based solely on sequence identity due to fragmentary or localized similarity.
[0031] The similarity value can be obtained using efficient comparison tools such as fastANI, skani, mash, and BLAST.
[0032] In large-scale scenarios, this step can be executed in parallel and combined with threshold judgment to filter out comparison results that do not meet the requirements in real time, avoiding the storage of a huge comparison matrix and thus saving memory resources.
[0033] The similarity threshold can be flexibly adjusted according to the specific species, genus, or research needs. In some implementations, the similarity threshold is set as follows: ANI of 80%–100% and AF of 20%–100%, preferably AF ≥ 50%. Genomes with an ANI below 80% are generally considered to belong to different species or distantly related groups, while an AF below 20% indicates too few aligned regions to support structural similarity between the two genomes. By using these dual threshold conditions, low-quality alignments or irrelevant sequences can be effectively filtered out, ensuring that connections are established only between genomes with sufficient similarity.
[0034] By adopting the similarity index and threshold settings defined above, we can significantly reduce the amount of invalid computation while ensuring the accuracy of the comparison, making the subsequent connection relationship construction and clustering steps more efficient and reliable, and ultimately improving the overall redundancy removal accuracy and speed.
[0035] In some implementations, the quality assessment metrics include one or more of the following: genome length, N50 value, completeness, contamination, number of scaffolds, number of contigs, or sequencing quality score. Genome length reflects the overall scale of the assembly; the N50 value represents sequence continuity and is a commonly used metric for evaluating genome assembly quality; completeness measures the proportion of the assembled genome containing the intended marker genes; and contamination reflects the degree of foreign sequence or assembly errors.
[0036] Furthermore, the number of scaffolds and contigs can be used to reflect the degree of fragmentation in the assembly; a lower number generally indicates a more complete and continuous assembly. The sequencing quality score is a statistical parameter assessing the accuracy of base recognition or the reliability of sequencing, and can be used as an auxiliary indicator in the overall evaluation. In practical applications, these indicators can be flexibly selected or combined according to the specific data type (such as bacterial genomes, eukaryotic genomes, or metagenomic assembly results) to more comprehensively reflect the quality of genome assembly.
[0037] By introducing the above-mentioned multi-dimensional quality evaluation indicators, the genomes within the cluster can be sorted using multiple parameters. This can effectively avoid the bias caused by relying on a single indicator (such as length or N50), thereby ensuring that the representative genomes selected as cluster centers perform optimally in terms of assembly integrity, continuity, and low contamination, thus providing quality assurance for the subsequent output of representative genomes.
[0038] In some implementations, the genomes within a cluster are sorted by a weighted average score calculated from various quality evaluation indicators, with the genome with the highest score serving as the cluster center.
[0039] Specifically, weighting coefficients can be assigned based on the importance of different indicators. For example, integrity and N50 can be given higher weights to highlight assembly quality factors; at the same time, contamination and contig quantity can be given negative weights to prevent low-quality genomes from being selected as representative. Through this weighted ranking, a balance can be achieved among multiple evaluation dimensions, ensuring that the genome with the highest score is representative in terms of overall quality.
[0040] In some implementations, to address the differentiated needs of specific research or database construction scenarios, the method further includes allowing the specification of cluster centers within certain clusters to add new cluster centers or replace those selected through sorting. This operation can be based on existing annotation information, species priority, sample source, or external quality assessment results, supplementing the results of automatic sorting or replacing automatically selected cluster centers. For example, when an authoritative version of a species' reference genome has been established in a public database, the user can directly set that genome as the cluster center, thereby improving the consistency and controllability of representative data. By combining automatic weighted sorting with a manually specified adjustable mechanism, a flexible human-machine collaborative selection method can be provided while ensuring the determinism of the algorithm. This ensures both the objectivity of representative genome selection and the practical application needs of specific projects, making the clustering results more aligned with the goals of dataset construction and downstream analysis. This feature is particularly useful in clinical tracing and standard strain analysis.
[0041] In some implementations, the sequencing of the genomes within the cluster is used to calculate a comprehensive score using the following weighted scoring function: Score = αR + βL + γN50 + δC – εP – ζn_contig, where R is genome priority, L is genome length, N50 is assembly continuity index, C is integrity, P is contamination level, and n_contig is the number of contigs; α, β, γ, δ, ε, and ζ are weighting coefficients between 0 and 1.
[0042] In all embodiments of the present invention, scoring is performed according to the scoring function described above.
[0043] The scoring function described above maintains the reproducibility of the calculation process while allowing for flexible adjustment of parameter weights based on different data characteristics. For example, the weight of the contamination factor can be increased in metagenomic samples to enhance low-contamination screening, while the weights of N50 and length can be increased in eukaryotic genome scenarios to enhance assembly continuity. The setting of the genome priority R allows the algorithm to retain controllable human preference adjustment capabilities on the basis of automation, enabling the priority selection of specific species or authoritative reference genomes when results are similar, thereby enhancing representativeness and biological significance.
[0044] The genome priority is used to prioritize specific genomes when their overall scores or quality indicators are similar across multiple genomes. Typically, this priority can be determined based on external information or manually defined rules, such as: genome species representativeness, known reference status, sample source reliability, annotation completeness, database inclusion status, or user-specified priority weights. By introducing this parameter, biological or data management preferences can be further reflected on top of automated sorting, making the selection of cluster centers more aligned with practical application needs.
[0045] The comprehensive score ranking calculated by this weighted scoring function can significantly reduce the impact of randomness or sequence dependence on the selection of cluster centers, ensuring that the selection of representative genomes is deterministic, stable and interpretable, thereby improving the reliability and reproducibility of clustering results.
[0046] In some implementations, when multiple genomes have the same overall score, cluster centers are selected according to a unique identifier or a preset sorting to ensure that the clustering results are stable and independent of the input order.
[0047] By adopting this decision rule, the selection results of cluster centers can be kept consistent under different operating environments and different input orders, thereby achieving order-independent input and deterministic output in the clustering process. Compared with existing greedy or randomly initialized clustering algorithms, this invention eliminates the uncertainty of ranking through an explicit tie-breaking rule, making the clustering results completely reproducible.
[0048] The introduction of this feature significantly improves the determinism and reproducibility of the method, generating completely consistent clustering results in multiple runs, distributed execution, or environments with different computing nodes. This not only reduces errors caused by the randomness of the algorithm but also facilitates result verification, database updates, and the engineering deployment of large-scale genome clustering tasks.
[0049] In some implementations, establishing the connection includes: A sparse connectivity graph is constructed based on genome pairs that meet the similarity threshold condition. The graph traversal algorithm is used to identify the set of nodes that are directly or indirectly connected and to determine the connected components, thus forming the similar connectivity cluster.
[0050] Specifically, each genome can be viewed as a node in a graph, and genome pairs that meet similarity threshold conditions (such as ANI and AF both not lower than a set threshold) are used as edge connections to form a sparse adjacency graph. On this sparse graph, depth-first search (DFS), breadth-first search (BFS), or Union-Find algorithms can be used to traverse the nodes to identify all connected components. Each connected component corresponds to a set of genomes with high similarity to each other, i.e., a similarity cluster. Since the sparse graph only retains highly similar genome pairs as edges, the density of the graph structure is much lower than that of a complete n² alignment, thus significantly reducing memory usage and computational complexity.
[0051] Establishing connections in this manner enables efficient cluster partitioning while maintaining clustering accuracy. This process not only quickly identifies potentially redundant genome sets but also provides clear computational boundaries for subsequent clustering and representative genome selection, thus achieving an efficient, stable, and scalable clustering infrastructure for large-scale datasets.
[0052] In some implementations, the clusters are formed using a greedy clustering algorithm, a k-medoids clustering algorithm, a heuristic clustering algorithm, or a center-based clustering algorithm. Preferably, the center-based clustering algorithm is used. Specifically, within each similarity cluster, all unclassified genomes are first calculated and sorted according to the quality evaluation index, and the genome with the highest current ranking is selected as the cluster center. Then, genomes with a similarity to this cluster center not lower than a set threshold (radius) are assigned to the same cluster and removed from the candidate set. Afterward, a new highest-quality genome is selected from the remaining unclassified genomes as the next round of cluster centers, and the above operation is repeated until all genomes within the similarity cluster are assigned to their corresponding clusters.
[0053] By employing the aforementioned centered clustering algorithm, computational efficiency, clustering stability, and result reproducibility can be balanced in large-scale genomic data environments. Compared to existing algorithms that rely on random initial points or heuristic selection, this method eliminates sequence dependencies and nondeterministic factors, making the clustering process linearly controllable. It is particularly suitable for execution in cloud computing or distributed computing environments, facilitating batch processing and database updates.
[0054] A second aspect of the invention relates to a computer-readable medium for storing computer instructions, programs, code sets, or instruction sets that, when run on a computer, cause the computer to perform the methods described above.
[0055] Any combination of one or more computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory, optical storage devices, magnetic storage devices, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.
[0056] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including—but not limited to—electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of transmitting, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0057] The program code contained on a computer-readable medium may be transmitted using any suitable medium, including—but not limited to—wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0058] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages—such as Java, Smalltalk, C++, Swift, and Python—as well as conventional procedural programming languages—such as the "C" language or similar programming languages. Additionally, scripting languages such as Python and R can be used. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0059] A third aspect of the present invention relates to an electronic device comprising: One or more processors; and A computer-readable storage medium for storing computer instructions, programs, code sets, or instruction sets that, when executed on a computer, cause the one or more processors to implement the methods described above.
[0060] In some embodiments, the electronic device may also include a transceiver. The processor and the transceiver are connected, such as via a bus. It should be noted that in practical applications, the transceiver is not limited to one unit, and the structure of the electronic device does not constitute a limitation on the embodiments of this application.
[0061] The processor can be a CPU, a general-purpose processor, a DSP, an ASIC, an FPGA, or other programmable logic device, transistor logic device, hardware component, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0062] A bus can include a pathway for transmitting information between the aforementioned components. The bus can be a PCI bus or an EISA bus, etc. Buses can be categorized as address buses, data buses, control buses, etc.
[0063] The embodiments of the present invention will be described in detail below with reference to examples. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. For experimental methods in the following embodiments where specific conditions are not specified, please refer to the guidelines given in this invention, or follow experimental manuals or conventional conditions in the art, or other experimental methods known in the art, or follow the conditions recommended by the manufacturer.
[0064] In the specific embodiments described below, the measurement parameters involving raw material components may have slight deviations within the weighing accuracy range unless otherwise specified. Temperature and time parameters are subject to acceptable deviations due to instrument testing accuracy or operational precision.
[0065] Example 1: Rapid Redundancy Removal Analysis Based on Large-Scale Microbial Genomes This embodiment focuses on the rapid redundancy removal of a microbial genome of a certain species. The total number of genomes is 10,000, with an average size of about 5Mb per genome. The data sources include self-tested sequencing data and data collected from public databases.
[0066] The specific steps are as follows: Step 1: Obtain genome similarity The skani tool was used to perform pairwise alignments on all 10,000 genomes, obtaining the average nucleotide identity (ANI) and alignment coverage (AF) between each pair. The parameters were: --fragLen 1000 --minFraction 0.2 --minANI 80. This resulted in approximately 80 million valid alignments (ANI ≥ 80%, AF ≥ 20%).
[0067] To save resources, the comparison process uses 32 threads to run in parallel, taking about 1 hour in actual operation, with a maximum memory consumption of about 20GB.
[0068] Step 2: Construct similar connection clusters The ANI threshold ani_threshold was set to 95%, and the AF threshold af_threshold was set to 80%. Based on the alignment results that met the above dual threshold conditions, a connectivity graph between genomes was constructed; all nodes were traversed, and connected subgraphs were identified, with each subgraph being an initial similar connectivity cluster; a total of 4200 initial clusters were finally identified, with the number of genomes in a single cluster ranging from 2 to 60.
[0069] Step 3: Forming clusters For each initial similarity cluster, the genomes within the cluster are sorted in descending order according to their genome length; the genome with the highest ranking is selected as the center; with this genome as a reference, all members with an ANI ≥ 95% are collected to form a fine cluster; the same operation is repeated for the remaining unclassified members until all genomes in the initial cluster have been clustered.
[0070] Step 4: Generate a representative genome Each fine cluster retains the top-ranked genome as a representative genome; after redundancy removal, the total number of representative genomes is approximately 5200, representing the data structure and diversity of all 10,000 genomes.
[0071] This representative genome set can serve as a standard reference database for subsequent metagenomic annotation, pathogen tracing analysis, genome evolution research, and other purposes.
[0072] Comparative analysis of implementation effects:
[0073] The results show that the method described in this invention significantly reduces memory consumption and running time while maintaining consistent redundancy removal, making it suitable for large-scale genome alignment analysis scenarios.
[0074] Example 2: Rapid redundancy removal based on 200,000 viral genomes To construct a non-redundant viral reference database, we plan to perform redundancy removal on approximately 200,000 viral genome sequences collected from public databases (such as NCBI RefSeq and GISAID). These sequences are from diverse sources, have a high repetition rate, and cover multiple categories including RNA viruses, DNA viruses, double-stranded viruses, segmented viruses, and retroviruses. Their lengths range from approximately 1kb to 300kb, exhibiting a wide species span, numerous recombination events, and significant differences in manual assembly. Therefore, it is necessary to efficiently identify redundancy and select representative sequences. The goal is to complete the clustering and redundancy removal of viral genomes in a resource-constrained server environment (16 threads, 128GB memory), ultimately outputting a representative viral genome set to provide a foundation for viral classification, origin tracing, and the construction of pan-genome databases.
[0075] Step 1: Obtain genome similarity The skani tool was used to perform rapid pairwise comparisons of viral genomes. The parameters were: t=16, preset='--small-genomes' (adapting to viral sequence length), min_AF=20; the data was saved as an ani record file skani.ani. Step 2: Construct similar connection clusters Highly similar genome pairs that meet the criteria are filtered according to thresholds: ANI ≥ 95% and AF ≥ 80%. Similarity clusters are extracted using traversal, and each connected cluster is an initial similar connection cluster. Finally, 26,317 similar connection clusters are obtained, of which the largest cluster contains as many as 2,112 viral strains and the smallest cluster contains only 1.
[0076] Step 3: Construct clusters Perform the following operations on the viral genome in each similarity cluster: All genomes within the cluster are sorted according to the following criteria: genome length (longest priority) and N50 value (highest priority). • Select the first element in the sorted sequence as the center sequence and calculate its ANI value relative to other members within the cluster; • Group genomes with an ANI ≥ 95% into the same cluster; • Repeat the operation on the remaining unclustered genomes until all members in the original similarity relationship cluster are assigned to a specific cluster.
[0077] For example: Cluster ID: Cluster123 initially contains 104 viral genomes. In the first round, 56 viruses are grouped around genome G1. Then, 28 viruses are grouped around G22, which is ranked first among the remaining viruses. Finally, 3 viruses are grouped into a subclass centered around G45. The remaining 3 viruses are isolated points that do not form clusters. This connection cluster is eventually divided into 3 fine clusters + 3 isolated viruses.
[0078] Step 4: Representative genome output • The center sequence of each cluster is selected as the representative genome, resulting in the final deredundant virus set: Quantity: 42,761 representative viral genomes; accounting for 21.4% of the original sequence quantity, with a redundancy compression rate of approximately 78.6%.
[0079] III. Performance Evaluation and Comparison
[0080] This representative genome set can be used for: constructing virus origin comparison databases (such as Kraken2 and Centrifuge); integrating metagenomic virus type annotation workflows; constructing virus evolutionary relationship profiles; distinguishing highly similar variants (such as SARS-CoV-2 and adeno-associated virus); and constructing a viral pangenome graph as the basis for structural variation analysis.
[0081] Verification has shown that when using this representative set for metagenomic annotation, compared to using all 200,000 original sequences, the alignment speed is increased by more than 120 times and memory consumption is reduced by 90%, while maintaining a basically consistent species recognition rate. It can be widely deployed on conventional servers and even high-performance laptop environments.
[0082] Example 3: Comparison with dRep tool To verify the efficiency and resource consumption of the method of this invention, the proposed fast redundancy removal method was compared with the commonly used genome redundancy removal tool dRep. Experimental data were selected from sample sets containing 5000, 10000, and 20000 metagenomic assembled genomes (MAGs), with each sample set repeated three times. The genome lengths ranged from 1 kbp to 2.1 Mbp. The experimental platform was a computing server with a 96-core CPU and 512 GB of memory.
[0083] 1. Experimental Methods dRep: Performs pre-clustering (Mash sketch) and ANI alignment (skani) using default parameters, and selects representative genomes based on the clustering results.
[0084] The method of this invention uses ANI ≥ 95% and AF ≥ 20% as similarity judgment thresholds to construct similar connectivity relationships, perform iterative clustering based on quality index ranking, and output a representative genome set.
[0085] 2. Result Comparison
[0086] The results show that the method of the present invention has significant advantages over dRep in terms of runtime and memory consumption, with the total runtime reduced from 106 times to 347.9 times and the memory usage reduced from 8.5 times to 56.5 times. Moreover, the improvement effect is more significant as the genome set increases, making it more suitable for redundancy removal scenarios with more than 100,000 genomes.
[0087] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims, and the specification can be used to interpret the content of the claims.
Claims
1. A method for deduplicating redundancy in genomic data, characterized in that, Includes the following steps: 1) Calculate the similarity of each pair of genomes to be processed, and retain genome pairs that simultaneously meet the preset similarity threshold conditions; 2) Establish connections based on genome pairs that meet the similarity threshold, identify all directly or indirectly connected node sets, and form several non-overlapping similar connection clusters; 3) For each similar connection cluster, the genomes are sorted according to the preset quality evaluation index, and the genome with the highest quality is selected as the cluster center. Genomes with similarity not lower than the set threshold are assigned to the same cluster. The above steps are repeated for genomes that are not assigned to any cluster until all genomes in the cluster have been clustered. The quality evaluation index is used to comprehensively evaluate the integrity, continuity and contamination of genome assembly. 4) The cluster centers in each cluster are used as representative genome outputs to form a non-redundant representative genome set.
2. The method according to claim 1, characterized in that, The similarity is defined as average nucleotide identity (ANI) and / or genome alignment coverage (AF). Optionally, the similarity threshold is set as follows: ANI is 80% to 100%, AF is 20% to 100%, and preferably AF ≥ 50%.
3. The method according to claim 1, characterized in that, The quality assessment metrics include one or more of the following: genome length, N50 value, integrity, contamination level, number of scaffolds, number of contigs, or sequencing quality score.
4. The method according to claim 3, characterized in that, The genomes within the cluster are sorted by a comprehensive score obtained by weighting various quality evaluation indicators, with the genome with the highest score being used as the cluster center. Optionally, the method further includes: allowing the specification of cluster centers in certain clusters to add new cluster centers or replace the cluster centers selected by sorting.
5. The method according to claim 4, characterized in that, The sequencing of the genomes within the cluster is used to calculate a comprehensive score using the following weighted scoring function: Score = αR + βL + γN50 + δC – εP – ζn_contig, where R is genome priority, L is genome length, N50 is assembly continuity index, C is integrity, P is contamination level, and n_contig is the number of contigs; α, β, γ, δ, ε, and ζ are weighting coefficients between 0 and 1.
6. The method according to claim 1, characterized in that, When multiple genomes have the same overall score, cluster centers are selected according to a unique identifier or a preset sorting to ensure that the clustering results are stable and independent of the input order.
7. The method according to any one of claims 1-6, characterized in that, The establishment of the connection relationship includes: A sparse connectivity graph is constructed based on genome pairs that meet the similarity threshold condition. The graph traversal algorithm is used to identify the set of nodes that are directly or indirectly connected and to determine the connected components, thus forming the similar connectivity cluster.
8. The method according to any one of claims 1-6, characterized in that, The clusters are formed using greedy clustering, k-medoids clustering, heuristic clustering, or circle-centered clustering.
9. A computer-readable medium for storing computer instructions, programs, code sets, or instruction sets, characterized in that, When it is run on a computer, it causes the computer to perform the method described in any one of claims 1-8.
10. An electronic device, comprising: One or more processors; as well as A computer-readable storage medium for storing computer instructions, programs, code sets, or instruction sets, characterized in that, when run on a computer, it causes the one or more processors to implement the method according to any one of claims 1-8.