Management method and system of algae eDNA database

By combining automated data collection and phylogenetic tree construction with a comparison strategy of BLAST algorithm and Hidden Markov Model, the problems of data dispersion and inconsistent annotation in the management of cyanobacteria eDNA database were solved. This achieved high-quality centralized data management and improved annotation accuracy, especially in the ability to identify low-similarity sequences and new taxonomic units.

CN121565271APending Publication Date: 2026-02-24HANGZHOU INST OF ECOLOGICAL & ENVIRONMENTAL SCI (HANGZHOU URBAN ECOLOGICAL ENVIRONMENT MONITORING STATION) +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511768219.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-24

AI Technical Summary

Technical Problem

Existing technologies for managing cyanobacterial eDNA databases suffer from data fragmentation, inconsistent annotation standards, and limited functionality, making it difficult to meet the needs of cyanobacterial ecological research and monitoring of harmful algal blooms. Furthermore, the reference sequence library has limited coverage, and the annotation methods lack the integration of multiple algorithms, making it difficult to achieve standardized data integration and dynamic maintenance.

Method used

By automatically collecting specific marker cpc gene sequences from multiple public databases and authoritative literature, a phylogenetic tree is constructed. The BLAST algorithm and Hidden Markov Model are then combined to perform similarity comparisons, screen candidate matching sequences, and perform joint annotation using the phylogenetic tree. This establishes a high-quality cyanobacteria-specific eDNA alignment and annotation database, enabling centralized data management and multi-dimensional verification.

Benefits of technology

It enables centralized management of cyanobacterial eDNA data, improves the accuracy of annotation results, especially the ability to identify low-similarity sequences and potential new taxonomic units, reduces the cost and error of manual screening, and provides a more comprehensive and reliable data support platform.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121565271A_ABST
    Figure CN121565271A_ABST
Patent Text Reader

Abstract

The invention discloses a management method and system for an algae eDNA database, and belongs to the technical field of data management, and the method comprises the steps: constructing a cyanophyta specific sequence data set and a phylogenetic tree by automatically collecting cyanophyta specific marker cpc gene sequences in a public database and authoritative literatures; after a blue-green algae eDNA query sequence is received, similarity comparison is performed by adopting a BLAST algorithm and a hidden Markov model double strategy, candidate matching sequences are screened, and joint annotation is performed in combination with a phylogenetic tree clustering position. According to the method, the problems that cyanobacteria eDNA data are dispersed, annotation standards are not uniform, new species are difficult to recognize and the like are solved, full-process automation from data collection, annotation to version management is achieved, and a high-accuracy and traceable eDNA database solution is provided for algae diversity monitoring and ecological research.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data management technology, and more specifically to a management method and system for an algal eDNA database. Background Technology

[0002] Environmental DNA (eDNA) technology, as a non-invasive biomonitoring method, has been widely used in aquatic ecosystem research and algal diversity analysis in recent years. Currently, eDNA-based species identification and database management methods are mostly concentrated in the bacterial field, with the strategy of using the variable region of the 16S rRNA gene as a universal primer for amplification and sequencing becoming the mainstream. However, this method has significant limitations in applications targeting specific phyla (such as cyanobacteria): First, the 16S rRNA gene has low variability in cyanobacteria, making it difficult to achieve high-resolution species annotation; second, existing eDNA databases are mostly geared towards broad microbial communities, lacking systematic integration of cyanobacteria-specific entries, resulting in scattered data, a low proportion of cyanobacteria-related sequences in the detection results (often below 40%), inconsistent annotation standards, and limited functionality, failing to meet the specific needs of cyanobacterial ecological research and harmful algal bloom monitoring.

[0003] In recent years, researchers have begun to explore the use of cyanobacteria-specific markers for cpc genes to improve annotation accuracy. For example, Jiang Yongguang and Xiao Peng et al. proposed using phycocyanin in the journal *Harmful Algae* in 2017. CPC The PC-IGS sequences of genes and their spacer regions were obtained, and PCβF and PCαR454 primers were designed. A phylogenetic tree was constructed to perform cluster analysis on cyanobacterial sequences in environmental samples. Although this method improves the specificity of cyanobacterial identification to a certain extent, the following problems still exist: (1) The coverage of the reference sequence library is limited, and it fails to make full use of the large amount of publicly available cyanobacterial genomes or newly uploaded gene sequence data in recent years; (2) The annotation method mainly relies on phylogenetic tree construction and lacks the integration of multiple algorithms such as sequence similarity comparison, resulting in insufficient annotation ability for unknown or rare sequences; (3) There is a lack of a complete and continuously updated database management system, making it difficult to achieve standardized integration, dynamic maintenance and phylogenetic tree association analysis of data.

[0004] Therefore, how to provide a specific eDNA database management method and system for cyanobacteria that can achieve standardization and integration of the entire process from data acquisition, sequence annotation, phylogenetic analysis to database maintenance, thereby providing a more comprehensive, accurate and reliable data support platform for cyanobacterial eDNA research, is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0005] In view of this, the present invention provides a management method and system for algal eDNA databases. By automating data collection, optimizing alignment and annotation strategies, and combining phylogenetic tree integration analysis with continuous data maintenance and user feedback mechanisms, a high-quality cyanobacteria-specific eDNA alignment and annotation database is constructed to achieve accurate annotation of cyanobacteria eDNA query sequences.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: On one hand, the present invention provides a method for managing an algal eDNA database, comprising: S1. Automatically collect nucleic acid sequence data of cpc genes, which are specific markers related to cyanobacteria, from multiple public databases and / or authoritative literature, and construct a specific sequence dataset for cyanobacteria. S2. Construct a phylogenetic tree based on a specific sequence dataset; S3. Receive cyanobacterial eDNA query sequences, and perform similarity comparisons between the query sequences and the specific sequence dataset using the BLAST algorithm or a hidden Markov model constructed based on the conserved structural domain of the cyanobacterial-specific phycocyanin cpc gene. Screen candidate matching sequences, and combine them with the phylogenetic tree to construct a local phylogenetic tree. Perform joint annotation based on cluster positions and similarity comparison results. Construct a cyanobacterial-specific eDNA alignment and annotation database based on the query sequences, annotation results, specific sequence dataset, and phylogenetic tree. S4. Regularly perform redundancy removal and quality control on the cyanobacteria-specific eDNA alignment and annotation database, collect new data and import it into the database after processing according to S1-S3, and establish a data version management and traceability mechanism.

[0007] Preferably, S3 includes: Receive the cyanobacterial eDNA query sequence input by the user; The BLAST algorithm is used to compare the similarity with a specific sequence dataset, and reference sequences with a similarity greater than a first preset value are selected as candidate matching sequences. If no reference sequence with similarity greater than the first preset value is found, a hidden Markov model based on the conserved structural domain of the cyanobacterial phycocyanin cpc gene is used for comparison to screen out candidate matching sequences with a structural domain matching degree greater than the second preset value. By combining the phylogenetic tree, the query sequence and the candidate matching sequence are jointly reconstructed into a local phylogenetic tree. The final annotation result is determined based on the cluster position of the query sequence in the phylogenetic tree. By integrating query sequences, annotation results, specific sequence datasets, and phylogenetic trees, and establishing associations through unique sequence IDs, a cyanobacteria-specific eDNA alignment and annotation database is formed.

[0008] Preferably, the BLAST algorithm is used to compare the similarity with a specific sequence dataset, and reference sequences with a similarity greater than a first preset value are selected as candidate matching sequences, including: Set the parameters for the BLAST algorithm; Perform a global alignment between the query sequence and the specific sequence dataset, and calculate global similarity; Based on global similarity ranking, reference sequences with global similarity greater than a first preset value are selected as preliminary candidate matching sequences; Perform local alignment on the preliminary candidate matching sequences and calculate local similarity and similarity coverage; Select preliminary candidate matching sequences with local similarity greater than the first preset value and similarity coverage greater than the third preset value as the second candidate matching sequences; The second candidate matching sequence is verified by reverse BLAST. The second candidate matching sequence with a similarity difference of less than the fourth preset value and a coverage difference of less than the fifth preset value in the bidirectional alignment results is retained as the candidate matching sequence.

[0009] Preferably, a hidden Markov model based on the conserved structural domain of the phycocyanin cpc gene of cyanobacteria is used for comparison to screen candidate matching sequences with a structural domain matching degree greater than a second preset value, including: The cyanobacterial specific phycocyanin cpc gene sequence was extracted from the specific sequence dataset. Sequence alignment was performed using a multiple sequence alignment tool to identify and extract the conserved structural domain regions of the cpc gene, which were then used as the model training dataset. Hidden Markov Model (HMM) models were trained using a model training dataset using a hidden Markov model building tool to generate a hidden Markov model targeting the conserved structural region of the cyanobacterial-specific phycocyanin cpc gene. Input the query sequence into the trained Hidden Markov Model to perform structural domain matching analysis, and calculate the matching degree and matching coverage between the query sequence and the conservative structural region. Select reference sequences that simultaneously satisfy a matching degree greater than the second preset value and a matching coverage greater than the sixth preset value as candidate matching sequences; The selected candidate matching sequences are subjected to structural region boundary verification. By comparing the positional distribution of the matching region with the known conserved structural regions of the cpc gene, candidate matching sequences whose matching regions deviate from the conserved structural domain region by more than the seventh preset value are eliminated, and the final candidate matching sequences are determined.

[0010] Preferably, by combining the phylogenetic tree, a local phylogenetic tree is reconstructed using the query sequence and candidate matching sequences. Based on the cluster position of the query sequence within the phylogenetic tree, the final annotation result is determined, including: The query sequence is compared with the candidate matching sequence through multiple sequence alignment, and the local phylogenetic tree is reconstructed using the maximum likelihood method. Obtain the branch cluster where the query sequence belongs, and calculate the branch distance between the query sequence and each candidate matching sequence within the branch cluster; If the branch distance between the query sequence and a candidate matching sequence is less than the eighth preset value, and the self-expansion support rate of the query sequence and a candidate matching sequence in the branch cluster is greater than the ninth preset value, then the query sequence and the candidate matching sequence are determined to belong to the same classification unit. If the query sequence forms a branch on its own and the branch distance to multiple candidate matching sequences is greater than the eighth preset value, then based on the known classification information of the candidate matching sequences, the query sequence is determined to be a potential new taxonomic unit or a closely related unrecorded group, and marked as pending verification, prompting the user to further confirm it by combining it with other molecular markers.

[0011] Preferably, the method further includes: Establish a user feedback verification mechanism to receive user verification information on annotation results, including confirmation of correctness, classification correction, and suggestions for adding new classifications; Correction information that has been verified by at least three independent users or at least two different experiments is marked as verified and updated to the specific sequence dataset and phylogenetic tree, triggering database version iteration; For query sequences marked as pending verification or potential new taxonomic units, the latest public databases and literature data are automatically retrieved periodically, automated re-annotation is performed, and the updated annotation results are pushed to the original users.

[0012] On the other hand, the present invention provides a management system for an algal eDNA database, comprising: The data acquisition unit is used to automatically collect nucleic acid sequence data of specific marker cpc genes related to cyanobacteria from multiple public databases and / or authoritative literature, and construct a specific sequence dataset for cyanobacteria. The phylogenetic tree building module is used to construct phylogenetic trees based on specific sequence datasets. The annotation module receives cyanobacterial eDNA query sequences, performs similarity comparisons with the specific sequence dataset using a BLAST algorithm or a hidden Markov model constructed based on the conserved structural domain of the cyanobacterial-specific phycocyanin cpc gene, filters candidate matching sequences, and constructs a local phylogenetic tree by combining the query sequence and candidate matching sequences with the phylogenetic tree. Joint annotation is performed based on cluster positions and similarity comparison results. A cyanobacterial-specific eDNA alignment and annotation database is constructed based on the query sequences, annotation results, specific sequence dataset, and phylogenetic tree. The quality control module is used to periodically remove redundancy and perform quality control on the cyanobacteria-specific eDNA alignment and annotation database, collect new data and import it into the database, and establish a data version management and traceability mechanism.

[0013] Preferably, the system further includes: The optimization module is used to establish a user feedback verification mechanism, receive user verification information on the annotation results, update the specific sequence dataset and phylogenetic tree, and trigger database version iteration.

[0014] As can be seen from the above technical solution, compared with the prior art, this invention discloses a management method and system for algal eDNA databases. By automatically collecting specific marker cpc gene sequences from multiple public databases and authoritative literature, it achieves centralized and systematic management of algal eDNA data, avoiding the problems of data dispersion and cumbersome acquisition. Simultaneously, the constructed specific sequence dataset provides high-quality basic data for subsequent analysis, reducing the cost and error of manual screening. Furthermore, this invention employs a dual alignment strategy of BLAST algorithm and Hidden Markov Model based on conserved structural domains of cpc genes, combined with phylogenetic tree clustering analysis, to achieve multi-dimensional verification of query sequences. This effectively reduces the limitations of single alignment methods and improves the accuracy of annotation results, especially for identifying low-similarity sequences or potential new taxonomic units. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the process provided by the present invention.

[0017] Figure 2 This is a structural schematic diagram provided for the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] This invention discloses a method for managing an algal eDNA database, such as... Figure 1As shown, it includes: S1. Automatically collect nucleic acid sequence data of cpc genes, which are specific markers related to cyanobacteria, from multiple public databases and / or authoritative literature, and construct a specific sequence dataset for cyanobacteria.

[0020] The data sources include: international public nucleic acid databases such as NCBI GenBank, ENA, and DDBJ; specialized databases such as SILVA and Greengenes; and algae-specific databases such as AlgaeBase. Authoritative literature can be experimentally validated sequence data extracted from published taxonomic and phylogenetic literature, especially those sequences that describe new species or revised taxonomic relationships.

[0021] For each data source, data collection was performed by writing or utilizing NCBI's E-utilities API and a customized web crawler. The target data collected included: nucleic acid sequences, corresponding metadata such as species classification information (kingdom, phylum, class, order, family, genus, species), collection location, environmental origin, literature source, and submitter information.

[0022] After data collection is complete, the data is preprocessed, including: The collected sequence data is formatted uniformly, the metadata of each sequence is extracted, and stored as a structured data table; Sequences containing N > 5%, those with obvious sequencing errors, or those whose length does not conform to the typical range of marker genes were screened using sequence quality assessment tools. Sequence clustering tools were used to cluster sequences of the same marker gene based on 99% sequence similarity. The longest sequence with the most complete metadata in each cluster was retained as the representative sequence to reduce the interference of redundant data on subsequent analysis.

[0023] The final selected specific sequences and their metadata are stored in a relational database or a dedicated sequence database, and a multi-dimensional index is created to facilitate rapid retrieval and comparison in the future.

[0024] S2. Construct a phylogenetic tree based on a specific sequence dataset.

[0025] Specifically, the maximum likelihood (ML) method is used as the main construction method, and IQ-TREE2 is used with the "model fusion" function enabled to automatically select the optimal model. Set up a fast bootstrap approximation to evaluate branch support, run 1000 repetitions, and combine it with the SH-like approximate likelihood ratio test to obtain multiple support metrics; For very large datasets (>1000 sequences), FastTree2 is used to construct an initial tree using fast approximate ML, and then IQ-TREE is imported for precise optimization, balancing computational efficiency and accuracy. A Bayesian inference (BI) tree was constructed simultaneously for validation. MrBayes v3.2.7 was used to run two independent MCMC chains, confirming convergence with a split frequency standard deviation of <0.01. Outgroup species were identified using NCBI Taxonomy and explicitly specified during tree construction to ensure correct tree orientation.

[0026] For ML trees, we integrate three metrics: UFBoot support (≥95% consider strong support, 80-94% consider moderate support, <80% consider weak support), aLRT support (≥80% consider reliable support), and standard nonparametric bootstrap (≥1000 repetitions). For a BI tree, record the posterior probability (PP), and a PP ≥ 0.95 is considered a strong support node; Identify branches with low support (UFBoot < 70% or PP < 0.9), mark them as "topological uncertainty regions," and reduce their weight in subsequent annotations; RogueNaRok analysis was performed to identify anomalous sequences that had the greatest impact on tree stability and to assess whether they needed to be removed from the final tree.

[0027] Use the iTOL (Interactive Tree Of Life) online tool or the ETE Toolkit Python package to annotate taxonomic information on tree branches using color coding or labels; Mark the support values ​​at key nodes and add dashed lines or transparency markers to branches with low support; Leaf nodes are marked with different shaped symbols according to the sequence origin to facilitate tracing the source; For the identified potential new taxonomic units, use an asterisk ( Highlight and associate the ID to be verified in step S3; Output vector graphics files (SVG / PDF) and standard tree files, with embedded metadata annotation blocks.

[0028] S3. Receive cyanobacterial eDNA query sequences, and perform similarity comparisons between the query sequences and the specific sequence dataset using the BLAST algorithm or a hidden Markov model constructed based on the conserved structural domain of the cyanobacterial-specific phycocyanin cpc gene. Screen candidate matching sequences, and combine them with the phylogenetic tree to construct a local phylogenetic tree. Perform joint annotation based on cluster positions and similarity comparison results. Construct a cyanobacterial-specific eDNA alignment and annotation database based on the query sequences, annotation results, specific sequence dataset, and phylogenetic tree. S4. Regularly perform redundancy removal and quality control on the cyanobacteria-specific eDNA alignment and annotation database, collect new data and import it into the database after processing according to S1-S3, and establish a data version management and traceability mechanism.

[0029] Furthermore, S3 specifically includes: Receive the cyanobacterial eDNA query sequence input by the user; The BLAST algorithm is used to compare the similarity with a specific sequence dataset, and reference sequences with a similarity greater than a first preset value are selected as candidate matching sequences. If no reference sequence with similarity greater than the first preset value is found, a hidden Markov model based on the conserved structural domain of the cyanobacterial phycocyanin cpc gene is used for comparison to screen out candidate matching sequences with a structural domain matching degree greater than the second preset value. By combining the phylogenetic tree, the query sequence and the candidate matching sequence are jointly reconstructed into a local phylogenetic tree. The final annotation result is determined based on the cluster position of the query sequence in the phylogenetic tree. By integrating query sequences, annotation results, specific sequence datasets, and phylogenetic trees, and establishing associations through unique sequence IDs, a cyanobacteria-specific eDNA alignment and annotation database is formed.

[0030] Furthermore, the BLAST algorithm is used to compare the similarity with a specific sequence dataset, and reference sequences with a similarity greater than a first preset value are selected as candidate matching sequences, including: Set the parameters for the BLAST algorithm; Perform a global alignment between the query sequence and the specific sequence dataset, and calculate global similarity; Based on global similarity ranking, reference sequences with a global similarity greater than a first preset value are selected as preliminary candidate matching sequences, with the first preset value being 85%-95%. Perform local alignment on the preliminary candidate matching sequences and calculate local similarity and similarity coverage; Select preliminary candidate matching sequences with local similarity greater than the first preset value and similarity coverage greater than the third preset value as the second candidate matching sequences; The second candidate matching sequence is verified by reverse BLAST. The second candidate matching sequence with a similarity difference of less than the fourth preset value and a coverage difference of less than the fifth preset value in the bidirectional alignment results is retained as the candidate matching sequence. The fourth preset value can be set to 5% and the fifth preset value can be 10%.

[0031] In another embodiment, a hidden Markov model constructed based on the conserved structural domain of the cyanobacterial-specific phycocyanin cpc gene is used for comparison to screen candidate matching sequences with a structural domain matching degree greater than a second preset value, including: The cyanobacterial specific phycocyanin cpc gene sequence was extracted from the specific sequence dataset. Sequence alignment was performed using a multiple sequence alignment tool to identify and extract the conserved structural domain regions of the cpc gene, which were then used as the model training dataset. The model was trained using a Hidden Markov Model (HMM) building tool based on the model training dataset to generate a HMM model for the conserved structural region of the cyanobacterial specific phycocyanin cpc gene. The HMM building tool can be the HMMER software package. During the model training process, the number of iterations should be set to no less than 100 to ensure the model's accuracy in capturing the features of the conserved structural domain. The query sequence is input into the trained Hidden Markov Model for structural domain matching analysis. The matching degree and matching coverage between the query sequence and the conservative structural region are calculated. The matching degree is evaluated by the built-in log-likelihood value of the model, and the matching coverage is the proportion of the length of the query sequence that matches the conservative structural region to the total length of the query sequence. Select reference sequences that simultaneously satisfy a matching degree greater than the second preset value and a matching coverage greater than the sixth preset value as candidate matching sequences; The selected candidate matching sequences are subjected to structural region boundary verification. By comparing the positional distribution of the matching region with the known conserved structural regions of the cpc gene, candidate matching sequences whose matching regions deviate from the conserved structural regions by more than a seventh preset value are eliminated, and the final candidate matching sequences are determined. The deviation can be set to 20%.

[0032] In another embodiment, a local phylogenetic tree is reconstructed by combining the query sequence and candidate matching sequences. The final annotation result is determined based on the cluster position of the query sequence within the phylogenetic tree, including: The query sequence is compared with the candidate matching sequence through multiple sequence alignment, and the local phylogenetic tree is reconstructed using the maximum likelihood method. Obtain the branch cluster where the query sequence belongs, and calculate the branch distance between the query sequence and each candidate matching sequence within the branch cluster; If the branch distance between the query sequence and a candidate matching sequence is less than the eighth preset value, and the self-expansion support rate of the query sequence and a candidate matching sequence in the branch cluster is greater than the ninth preset value, then the query sequence and the candidate matching sequence are determined to belong to the same classification unit. If the query sequence forms a branch on its own and the branch distance to multiple candidate matching sequences is greater than the eighth preset value, then based on the known classification information of the candidate matching sequences, the query sequence is determined to be a potential new taxonomic unit or a closely related unrecorded group, and marked as pending verification, prompting the user to further confirm it by combining it with other molecular markers.

[0033] In another embodiment, the method further includes: Establish a user feedback verification mechanism to receive user verification information on annotation results, including confirmation of correctness, classification correction, and suggestions for adding new classifications; Correction information that has been verified by at least three independent users or at least two different experiments is marked as verified and updated to the specific sequence dataset and phylogenetic tree, triggering database version iteration; For query sequences marked as pending verification or potential new taxonomic units, the latest public databases and literature data are automatically retrieved periodically, automated re-annotation is performed, and the updated annotation results are pushed to the original users.

[0034] On the other hand, the present invention provides a management system for an algal eDNA database, such as... Figure 2 As shown, it includes: The data acquisition unit is used to automatically collect nucleic acid sequence data of specific marker cpc genes related to cyanobacteria from multiple public databases and / or authoritative literature, and construct a specific sequence dataset for cyanobacteria. The phylogenetic tree building module is used to construct phylogenetic trees based on specific sequence datasets. The annotation module receives cyanobacterial eDNA query sequences, performs similarity comparisons with the specific sequence dataset using a BLAST algorithm or a hidden Markov model constructed based on the conserved structural domain of the cyanobacterial-specific phycocyanin cpc gene, filters candidate matching sequences, and constructs a local phylogenetic tree by combining the query sequence and candidate matching sequences with the phylogenetic tree. Joint annotation is performed based on cluster positions and similarity comparison results. A cyanobacterial-specific eDNA alignment and annotation database is constructed based on the query sequences, annotation results, specific sequence dataset, and phylogenetic tree. The quality control module is used to periodically remove redundancy and perform quality control on the cyanobacteria-specific eDNA alignment and annotation database, collect new data and import it into the database, and establish a data version management and traceability mechanism.

[0035] In another embodiment, the system further includes: The optimization module is used to establish a user feedback verification mechanism, receive user verification information on the annotation results, update the specific sequence dataset and phylogenetic tree, and trigger database version iteration.

[0036] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to the method section.

[0037] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for managing an algal eDNA database, characterized in that, include: S1. Automatically collect nucleic acid sequence data of cpc genes, which are specific markers related to cyanobacteria, from multiple public databases and / or authoritative literature, and construct a specific sequence dataset for cyanobacteria. S2. Construct a phylogenetic tree based on a specific sequence dataset; S3. Receive the cyanobacterial eDNA query sequence and use the BLAST algorithm or a cyanobacterial-specific phycocyanin-based algorithm. CPC The hidden Markov model constructed from the conserved gene domains is compared with the specific sequence dataset for similarity comparison, and candidate matching sequences are screened. Combined with the phylogenetic tree, a local phylogenetic tree is constructed by jointly constructing the query sequence and the candidate matching sequences. Joint annotation is performed based on the cluster position and similarity comparison results. A cyanobacteria-specific eDNA comparison and annotation database is constructed based on the query sequence, annotation results, specific sequence dataset, and phylogenetic tree. S4. Regularly perform redundancy removal and quality control on the cyanobacteria-specific eDNA alignment and annotation database, collect new data and import it into the database after processing according to S1-S3, and establish a data version management and traceability mechanism.

2. The method for managing an algal eDNA database according to claim 1, characterized in that, S3 includes: Receive the cyanobacterial eDNA query sequence input by the user; The BLAST algorithm is used to compare the similarity with a specific sequence dataset, and reference sequences with a similarity greater than a first preset value are selected as candidate matching sequences. If no reference sequence with a similarity greater than the first preset value exists, then a cyanobacteria-specific phycocyanin-based sequence will be used. CPC The hidden Markov model constructed from the conserved gene domains was compared to select candidate matching sequences with a domain matching degree greater than the second preset value. By combining the phylogenetic tree, the query sequence and the candidate matching sequence are jointly reconstructed into a local phylogenetic tree. The final annotation result is determined based on the cluster position of the query sequence in the phylogenetic tree. By integrating query sequences, annotation results, specific sequence datasets, and phylogenetic trees, and establishing associations through unique sequence IDs, a cyanobacteria-specific eDNA alignment and annotation database is formed.

3. The method for managing an algal eDNA database according to claim 2, characterized in that, The BLAST algorithm was used to compare the sequence with a specific sequence dataset, and reference sequences with a similarity greater than a first preset value were selected as candidate matching sequences, including: Set the parameters for the BLAST algorithm; Perform a global alignment between the query sequence and the specific sequence dataset, and calculate global similarity; Based on global similarity ranking, reference sequences with global similarity greater than a first preset value are selected as preliminary candidate matching sequences; Perform local alignment on the preliminary candidate matching sequences and calculate local similarity and similarity coverage; Select preliminary candidate matching sequences with local similarity greater than the first preset value and similarity coverage greater than the third preset value as the second candidate matching sequences; The second candidate matching sequence is verified by reverse BLAST. The second candidate matching sequence with a similarity difference of less than the fourth preset value and a coverage difference of less than the fifth preset value in the bidirectional alignment results is retained as the candidate matching sequence.

4. The method for managing an algal eDNA database according to claim 2, characterized in that, A hidden Markov model based on the conserved domain of the phycocyanin cpc gene, a specific gene of cyanobacteria, was used for comparison to screen candidate matching sequences with a domain matching degree greater than a second preset value, including: The cyanobacterial-specific phycocyanin cpc gene sequence was extracted from a specific sequence dataset, and sequence alignment was performed using a multiple sequence alignment tool to identify and extract the gene. CPC Conserved structural domains of genes are used as training datasets for the model. Hidden Markov Model (HMM) models were trained using a model training dataset using a hidden Markov model building tool to generate a hidden Markov model targeting the conserved structural region of the cyanobacterial-specific phycocyanin cpc gene. Input the query sequence into the trained Hidden Markov Model to perform structural domain matching analysis, and calculate the matching degree and matching coverage between the query sequence and the conservative structural region. Select reference sequences that simultaneously satisfy a matching degree greater than the second preset value and a matching coverage greater than the sixth preset value as candidate matching sequences; The selected candidate matching sequences are subjected to structural region boundary verification by comparing the matching regions with known... CPC The location and distribution of conserved structural regions of genes are analyzed, and candidate matching sequences that deviate from the conserved structural region by more than the seventh preset value are eliminated to finally determine the final candidate matching sequences.

5. The method for managing an algal eDNA database according to claim 1, characterized in that, By combining the phylogenetic tree, a local phylogenetic tree is reconstructed using the query sequence and candidate matching sequences. Based on the cluster position of the query sequence within the phylogenetic tree, the final annotation results are determined, including: The query sequence is compared with the candidate matching sequence through multiple sequence alignment, and the local phylogenetic tree is reconstructed using the maximum likelihood method. Obtain the branch cluster where the query sequence belongs, and calculate the branch distance between the query sequence and each candidate matching sequence within the branch cluster; If the branch distance between the query sequence and a candidate matching sequence is less than the eighth preset value, and the self-expansion support rate of the query sequence and a candidate matching sequence in the branch cluster is greater than the ninth preset value, then the query sequence and the candidate matching sequence are determined to belong to the same classification unit. If the query sequence forms a branch on its own and the branch distance to multiple candidate matching sequences is greater than the eighth preset value, then based on the known classification information of the candidate matching sequences, the query sequence is determined to be a potential new taxonomic unit or a closely related unrecorded group, and marked as pending verification, prompting the user to further confirm it by combining it with other molecular markers.

6. The method for managing an algal eDNA database according to claim 1, characterized in that, The method further includes: Establish a user feedback verification mechanism to receive user verification information on annotation results, including confirmation of correctness, classification correction, and suggestions for adding new classifications; Correction information that has been verified by at least three independent users or at least two different experiments is marked as verified and updated to the specific sequence dataset and phylogenetic tree, triggering database version iteration; For query sequences marked as pending verification or potential new taxonomic units, the latest public databases and literature data are automatically retrieved periodically, automated re-annotation is performed, and the updated annotation results are pushed to the original users.

7. A management system for an algal eDNA database, characterized in that, include: The data acquisition unit is used to automatically collect specific markers related to cyanobacteria from multiple public databases and / or authoritative literature. CPC Nucleic acid sequence data of genes were used to construct a specific sequence dataset for cyanobacteria. The phylogenetic tree building module is used to construct phylogenetic trees based on specific sequence datasets. The annotation module receives cyanobacterial eDNA query sequences, performs similarity comparisons with the specific sequence dataset using a BLAST algorithm or a hidden Markov model constructed based on the conserved structural domain of the cyanobacterial-specific phycocyanin cpc gene, filters candidate matching sequences, and constructs a local phylogenetic tree by combining the query sequence and candidate matching sequences with the phylogenetic tree. Joint annotation is performed based on cluster positions and similarity comparison results. A cyanobacterial-specific eDNA alignment and annotation database is constructed based on the query sequences, annotation results, specific sequence dataset, and phylogenetic tree. The quality control module is used to periodically remove redundancy and perform quality control on the cyanobacteria-specific eDNA alignment and annotation database, collect new data and import it into the database, and establish a data version management and traceability mechanism.

8. The management system for an algal eDNA database according to claim 7, characterized in that, The system also includes: The optimization module is used to establish a user feedback verification mechanism, receive user verification information on the annotation results, update the specific sequence dataset and phylogenetic tree, and trigger database version iteration.