Method for identifying species and functional information of marine eukaryotic microflora at high throughput
Through the collaborative use of multiple software and databases, the problem of accuracy and low efficiency of marine eukaryotic microbial annotation is solved, and high-throughput identification of marine eukaryotic microbial community species and functional information is achieved, providing comprehensive and reliable research data support.
Patent Information
- Application Number
- CN202510585878.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-12
Smart Images

Figure CN120472993A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of bioinformatics, and in particular to a method for high-throughput identification of species and functional information of marine eukaryotic microbial communities. Background Art
[0002] Eukaryotic microorganisms play a vital role in marine ecosystems. They participate in key ecological processes such as material circulation and energy flow, and are of great significance to maintaining the ecological balance of the ocean. However, due to the extreme complexity of the marine environment, species and genetic research on marine eukaryotic microorganisms faces many challenges. Traditional microbial research methods rely on the cultivation and isolation of microorganisms, which is not only time-consuming and labor-intensive, but also can only study a small number of culturable microorganisms, and cannot fully reveal the diversity and functions of marine eukaryotic microorganisms.
[0003] With the development of high-throughput sequencing technology, it has become possible to directly sequence microbial communities in environmental samples, providing a new approach for studying marine eukaryotic microorganisms. However, existing high-throughput sequencing-based methods suffer from low accuracy and efficiency when annotating marine eukaryotic microbial species and genes. For example, low reference genome coverage makes species annotation difficult, the complex genetic structure of eukaryotic microorganisms makes gene prediction and annotation accuracy difficult, and data processing is susceptible to false positives, seriously affecting the reliability and depth of research.
[0004] Existing methods for annotating microeukaryotic species and genes face the following technical challenges: 1) The high diversity of marine eukaryotic microorganisms and low reference genome coverage (<15%) lead to high miss rates with traditional alignment tools (e.g., BLAST); 2) The complex gene structure of eukaryotic microorganisms (e.g., introns, alternative splicing), making prokaryotic annotation pipelines (e.g., MetaPhlAn) poorly suited; 3) Existing tools (e.g., EukRep) rely on single-character classification, resulting in species annotation resolution limited to the phylum / class level; 4) The proportion of eukaryotic sequences in metagenomic data is typically <5%, resulting in a low signal-to-noise ratio; and contamination from cross-domain (bacteria-archaea-eukaryotic) sequences can lead to false positives. MetaEuk is a toolkit for mining eukaryotic protein-coding genes in metagenomic data. It can be affected by database and sequence similarity constraints, resulting in inaccurate annotations, inability to accurately recover splice sites, and the potential for false positives. These factors, to a certain extent, affect the accuracy and comprehensiveness of eukaryotic microbial annotation. In addition, EukRep is a software for splitting eukaryotic and prokaryotic sequences, but it has a false positive problem that interferes with the accurate annotation of eukaryotic microorganisms. When processing complex metagenomic samples, the reliability of the annotation results may be affected due to inaccurate sequence differentiation.
[0005] Therefore, it is of great significance to develop a new annotation method for marine eukaryotic microorganisms. Summary of the Invention
[0006] The technical problem to be solved by the present invention is the annotation method of microeukaryotes in existing high-throughput data, a method for high-throughput identification of species and functional information of marine eukaryotic microbial communities, which overcomes the shortcomings of the existing technology, improves the accuracy and efficiency of annotation, and provides strong technical support for in-depth research on the ecological functions and biodiversity of marine eukaryotic microorganisms.
[0007] The above-mentioned purpose of the present invention is achieved through the following technical solutions:
[0008] (1) Data collection: Obtaining raw metagenomic data from the Malaspina 2010 expedition cruise;
[0009] (2) Data quality control: Use BBDuk software to perform quality control trimming on the original metagenome reads. By setting appropriate parameters, low-quality bases, adapter sequences, and possible contaminating sequences are removed to obtain high-quality double-ended clean reads, providing reliable data for subsequent sequence analysis.
[0010] (3) Sequence assembly: Clean reads were assembled into contigs using the MEGAHIT metagenomic assembly software within MetaWRAP. MEGAHIT uses an optimized algorithm to efficiently assemble short sequences into long fragments, improving the accuracy and completeness of the assembly and laying the foundation for subsequent gene prediction and annotation.
[0011] (4) Eukaryotic sequence identification: Using the identification function of EukRep, contigs belonging to eukaryotes are identified from contigs. Based on a specific algorithm and database, EukRep can identify sequence features unique to eukaryotes, but this process may result in false positives, that is, the identified eukaryotic contigs may still contain contigs of bacteria and archaea.
[0012] (5) Taxonomic annotation: Kaiju software was used to perform microbial taxonomic annotation on the assembled contigs based on the nr-euk database. The nr-euk database contains rich eukaryotic sequence information. Kaiju software can determine the species classification status of each contig by comparing it with the database. However, due to the limitations of the nr-euk database and the characteristics of the algorithm, false positive results may be introduced, so further processing is required. Kaiju software can remove bacterial and archaeal contigs containing false positive contigs, ultimately leaving only false positive eukaryotic contigs, thereby improving the accuracy of species classification annotation.
[0013] (6) Gene prediction: MetaEuk is used to accurately predict protein gene sequences of eukaryotic microorganisms based on the uniref90 database and eukaryotic contig sequences. MetaEuk uses advanced algorithms and a rich reference database to accurately identify protein-coding genes in eukaryotic microorganisms, providing key information for subsequent functional studies.
[0014] (7) Gene function annotation: The KEGG database was used to perform functional annotation of protein gene sequences with the help of the DIAMOND tool. The DIAMOND tool has fast and efficient alignment capabilities, and combined with the rich functional information in the KEGG database, it can comprehensively and accurately annotate gene functions.
[0015] (8) Abundance calculation: CoverM was used to calculate the abundance of eukaryotic contigs, and Salmon was used to calculate the abundance of genes. By calculating the abundance, we can understand the relative content of different eukaryotic microbial species and genes in the sample, which is helpful for studying the structure of microbial communities.
[0016] (9) Providing quantitative data on structure and function. By matching species annotation information and abundance, we can obtain the abundance information of eukaryotic contigs species, thereby obtaining the composition of eukaryotic microorganisms in the microbial community; by matching gene annotation information and abundance, we can obtain gene abundance information, providing data support for in-depth research on gene function and microbial metabolic pathways.
[0017] Beneficial effects
[0018] (1) Improving annotation accuracy: Through the combined use of multiple software and databases, and the effective processing of false positive results using Kaiju software, the accuracy of marine eukaryotic microbial species and gene annotations has been significantly improved, erroneous annotations have been reduced, and a reliable data foundation has been provided for subsequent research.
[0019] (2) Efficiency: The software and tools used in this invention are highly efficient, such as the rapid assembly capability of MEGAHIT and the rapid comparison capability of DIAMOND, which can complete the processing and analysis of large amounts of data in a relatively short period of time, thereby improving research efficiency.
[0020] (3) Comprehensiveness: From data collection to species classification, gene prediction, functional annotation, and abundance calculation, the present invention provides a complete set of analysis processes, which comprehensively reveals the species and gene information of marine eukaryotic microorganisms, and helps to gain a deeper understanding of the ecological functions and diversity of marine eukaryotic microorganisms. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 Flowchart of the method for annotating species and genes of eukaryotic microorganisms.
[0022] Figure 2 Metagenomic marine eukaryotic microbial species annotation at the phylum level during the Malaspina 2010 expedition cruise.
[0023] Figure 3 Annotation of 18S rDNA marine eukaryotic microbial species at the phylum level during the Malaspina 2010 expedition cruise.
[0024] Figure 4 Table showing gene annotation of marine eukaryotic microbial metagenomics collected during the Malaspina 2010 expedition cruise. DETAILED DESCRIPTION
[0025] The following is a clear and complete description of the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.
[0026] Example 1
[0027] This implementation example 1 obtained the raw metagenomic data of samples collected at a depth of 3000m-4000m during the Malaspina 2010 expedition cruise, and included the following steps:
[0028] (1) Data collection, quality control, and assembly: The raw metagenomic data from the Malaspina 2010 expedition cruise were obtained. The raw metagenomic sequencing reads were quality-controlled and trimmed using BBDuk software (parameters: ktrim = rk = 28 mink = 12 hdist = 1 tbo -t tpe -t gtrim -rltrimg = 20 min length - 70) to remove low-quality bases, adapter sequences, and possible contaminating sequences, thereby obtaining high-quality double-end clean reads. The clean reads were then assembled and assembled to generate contigs using the MEGAHIT (v1.1.3; parameter: -m 500) metagenomic assembly software in MetaWRAP software. (2) Eukaryotic sequence separation and taxonomic annotation: The assembled contigs were used to identify eukaryotic sequences using the EukRep software. When the EukRep software was running, contigs that were likely to belong to eukaryotes (euk-contigs) were identified from the contigs based on its default parameter settings and built-in algorithms. Because EukRep identification results may contain false positives, microbial taxonomic annotation of the euk-contigs was performed using Kaiju (https: / / github.com / bioinformatics-centre / kaiju, code: kaiju-z 20-t / Kaiju / nodes.dmp\-f / kaiju_db_nr_euk.fmi\-i SRR3960572.fasta-o SRR3960572.kaiju.out; this code uses one sample as an example) software in conjunction with the nr-euk database. Kaiju software determines the taxonomic status of each contig by aligning sequences with those in the nr-euk database. To remove false positives, Kaiju's annotation function and a custom script were used to analyze and filter the annotation results, removing contigs belonging to bacteria and archaea, ultimately obtaining false-positive eukaryotic contigs. (3) Gene prediction: The false-positive eukaryotic contigs were used as input data for MetaEuk (https: / / github.com / soedingl ab / metaeuk) and combined with the uniref90 database to predict protein gene sequences of eukaryotic microorganisms. MetaEuk accurately predicted protein gene sequences of eukaryotic microorganisms through multi-step analysis and alignment of eukaryotic contigs, providing key gene information for subsequent functional annotation.
[0029] (4) Gene function annotation: With the help of the DIAMOND tool, the predicted protein gene sequence is compared with the Kyoto Encyclopedia of Genes and Genomes (KEGG) database to achieve gene function annotation. When running DIAMOND, set --evalue 1e-5 to control the significance of the comparison results; --max-target-seqs 1 to retain only the best comparison results. The DIAMOND tool quickly matches the protein gene sequence with the functional information in the KEGG database. According to the annotation system of the KEGG database, the functional annotation content of each gene is extracted, including information such as the metabolic pathways and biological processes involved in the gene, which provides an important basis for a deeper understanding of the functions of eukaryotic microorganisms.
[0030] (5) Abundance calculation: CoverM software was used to calculate the abundance of eukaryotic contigs. The clean reads data after quality control were compared with the assembled eukaryotic contigs, and CoverM software calculated the abundance of each eukaryotic contig based on the comparison results. During the calculation process, -min-identity90 was set to ensure that the consistency of the comparison was not less than 90%; --min-coverage80 was set to ensure that the coverage was not less than 80%, so as to obtain accurate eukaryotic contigs abundance data. Salmon software was used to calculate the abundance of genes. First, the corresponding environment was activated to establish an index (salmon index) for the Salmon software; then, based on the clean reads quantitative analysis (salmon quant), the standardized gene TPM value was obtained.
[0031] (6) The species annotation information was matched with the eukaryotic contigs abundance calculated by CoverM and organized into a detailed table to obtain the contigs species abundance information; the gene annotation information was matched with the gene abundance calculated by Salmon and organized into a gene abundance information table, which comprehensively displayed the species composition and gene function quantitative information of the marine eukaryotic microorganisms in this sea area.
[0032] Implementation Case 2
[0033] This case study selected 18S rDNA raw data from samples collected at depths of 3000-4000m during the Malaspina 2010 expedition cruise. The following steps were involved:
[0034] (1) Data collection and denoising: The raw metagenomic 18S rDNA data from the Malaspina 2010 expedition cruise were obtained. The obtained sequences were subjected to sequence quality control, denoising, linking, and chimera removal (UPARSE) using the QIIME2 (Quantitative Insights Into Microbial Ecology) DADA2 analysis pipeline. The sequences after DADA2 denoising are usually referred to as ASVs (i.e., amplicon sequence variants).
[0035] (2) Based on the classifier trained on the PR2 database (Protist ribosomal reference database), we can perform taxonomic annotation on the microbial sequences (ASVs) in the sample and determine the taxonomic status of each sequence, thereby understanding the species composition of the microorganisms in the sample. Run the code: time qiime feature-classifierclassify-sklearn\--iclassifier . / db / pr2_2024_classifier.qza\--i-reads rep-seqs.qza\--o-classification taxonomy.qza.
[0036] Based on the above examples, we can conclude that our proposed method for high-throughput identification of species and functional information of marine eukaryotic microbial communities has significant advantages.
[0037] (1) High species richness: The species annotation information of eukaryotic microorganisms was obtained by a high-throughput method for identifying species and functional information of marine eukaryotic microbial communities ( Figure 2 ) shows more species of eukaryotic microorganisms, which means it can present a more comprehensive picture of the species composition of the sample. Number of species categories displayed: Figure 2 It displays more than 20 eukaryotic microbial phyla, including Basidiomycota, Anisomycota, and Ascomycota, covering a wider range; Figure 3 Only three eukaryotic microbial phyla, namely Basidiomycota, Ascomycota, and Chlorophyta, as well as the "other" category, are displayed, and the species information presented is relatively limited.
[0038] (2) When studying eukaryotic microbial communities, rich species information helps to gain a deeper understanding of community structure and diversity, providing more sufficient data support for subsequent studies on the ecological relationships and functional roles of different species. For example, when analyzing the interactions of microorganisms in marine ecosystems, more species information can reveal more complex food chains and ecological networks.
[0039] (3) Figure 2The distribution of various marine eukaryotic microorganisms at the phylum level was shown, based on which gene function studies were conducted ( Figure 4 ), can encompass the gene functions of diverse marine eukaryotic microorganisms. From Basidiomycota to Anisomastigotes, diverse microorganisms play diverse roles in ecosystems, and their gene functions vary accordingly. This approach has the potential to uncover gene functions involved in diverse biological processes, such as cell motility, transport, metabolic regulation (such as lipid metabolism), and signal transduction, providing a rich source of information for a deeper understanding of the functions of marine eukaryotic microorganisms in ecosystems.
[0040] Overall, this innovative approach provides new perspectives and comprehensive insights, helping to deepen our understanding of eukaryotic microorganisms. Furthermore, these findings are valuable references for future studies of marine ecosystems. Overall, this patent holds enormous potential in the field of environmental ecology, providing strong support for research in ecology and environmental protection.
[0041] The above are only a few preferred embodiments of the present invention, and their description is relatively specific and detailed, but it should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art may make various modifications and improvements without departing from the scope of the present invention, and such modifications and improvements are within the scope of protection of the present invention.
Claims
1. A method for high-throughput identification of species and functional information of marine eukaryotic microbial communities, comprising the following steps: Step 1: Data collection: Obtain the raw metagenomic data from Malaspina's 2010 expedition cruise. Step 2: Data quality control: Use BBDuk software to perform quality control trimming on the original metagenome reads to obtain high-quality paired-end clean reads; Step 3. Sequence assembly: cleanreads were assembled into contigs using the metagenomic assembly software MEGAHIT included in the MetaWRAP software; Step 4. Eukaryotic sequence identification: Use the identification function of EukRep to identify contigs belonging to eukaryotes from contigs; Step 5: Taxonomic annotation: Use Kaiju software to perform microbial taxonomic annotation on the assembled contigs based on the nr-euk database; Step 6. Gene prediction: Use MetaEuk to predict the protein gene sequence of eukaryotic microorganisms; Step 7: Gene function annotation: Use the KEGG database to perform functional annotation on protein gene sequences; Step 8. Abundance calculation: Use CoverM and Salmon to calculate the abundance of contigs and genes respectively.
2. The method according to claim 1, characterized in that In the step three, BBDuk is used to perform quality filtering on the original reads obtained from the metagenomic sequencing to obtain high-quality double-end sequences; in the sequence assembly step, MEGAHIT is used to assemble the quality-filtered sequences into contigs.
3. The method according to claim 1, characterized in that In steps 4 and 5, eukaryotic contigs were identified using EukRep based on the nr-euk database, but the eukaryotic contigs identified by EukRep contained false positives, i.e., the eukaryotic contigs still contained bacterial and archaeal contigs.
4. The method according to claim 3, characterized in that Since the nr-euk database was used, the contigs obtained contained false positives, so the bacterial and archaeal contigs were removed using Kaiju software, leaving only the eukaryotic contigs without false positives.
5. The method according to step 6 of claim 1 and claim 4, characterized in that: Based on the eukaryotic contigs sequence, MetaEuk was used to accurately predict the gene sequence of eukaryotic microorganisms based on the uniref90 database.
6. The method according to claim 1, characterized in that In the step seven, the DIAMOND tool is used to perform functional annotation of genes based on the KEGG database.
7. The method according to claim 1, characterized in that In the step eight, CoverM is used to calculate the abundance of eukaryotic contigs; and Salmon is used to calculate the abundance of genes.
8. The method according to claim 3 and claim 6, characterized in that Using eukaryotic contig sequence information, the species annotation information and abundance are matched to obtain the abundance information of contig species, thereby obtaining the composition of eukaryotic microorganisms in the microbial community.
9. The method according to claim 3 and claim 7, characterized in that Using gene sequence information, gene annotation information and abundance are matched to obtain gene abundance information.
10. Use of the method according to any one of claims 1 to 9 in marine ecosystem research.