Method for detecting marine biological diversity based on eDNA metagenome and macro bar code

By combining metagenomics and metabarcoding technology, the problem of species composition deviation and low resolution of eDNA technology in marine biodiversity detection is solved, efficient and accurate monitoring of marine biodiversity and improving the ability of unknown biological detection.

CN120272579APending Publication Date: 2025-07-08江苏省工程咨询中心有限公司
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510361639.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

When detecting marine biodiversity, existing eDNA technologies have problems such as species composition deviation caused by PCR amplification process, weak detection capabilities for unknown biologicals and low species annotation resolution.

Method used

Combining metagenomic technology and metabarcoding technology, through metagenomic high-throughput sequencing, specific marker gene conserved region alignment and strict alignment, biodiversity detection methods for marine sediment samples are constructed, including pretreatment, eDNA extraction, metagenomic sequence quality control, assembly and specific marker gene database comparison, to achieve efficient and accurate species annotation.

Benefits of technology

It improves species annotation resolution, enhances detection capabilities for unknown organisms, reduces computing resource requirements, and reduces interference to ecosystems, providing more comprehensive monitoring of marine biodiversity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120272579A_ABST
    Figure CN120272579A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of bioinformatics, in particular to a method for detecting biodiversity based on eDNA metagenomes and macro bar codes. Comprising the following steps: collecting a sediment sample in a monitoring area and extracting eDNA; carrying out metagenome high-throughput sequencing on the extracted eDNA to obtain a metagenome original sequence; performing quality control on the metagenome original sequence to obtain a metagenome quality control sequence; assembling according to the repeated fragments to obtain a metagenome long sequence; and according to the specific marker gene, screening out a query sequence and carrying out annotation according to the query sequence. Compared with a traditional ecological system field investigation method, a molecular biology method is applied, investigation and monitoring of biological diversity are more efficient, more accurate and more comprehensive, and interference to the ecological system and organisms is small.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of bioinformatics, and particularly relates to a method for detecting biodiversity based on eDNA metagenome and metabarcoding. Background Art

[0002] The biological community structure of marine ecosystems is one of the key research focuses in the basic research of the ecological functions and biogeochemical processes of marine ecosystems. Marine ecosystems cover more than 70% of the Earth's surface area and provide humans with a variety of goods and services, such as providing substances, participating in biogeochemical cycles, providing habitats for a large number of marine organisms, purifying pollution, regulating climate change, and so on. In recent years, environmental changes and human activities such as seawater warming, ocean acidification, sea-level rise, habitat destruction, seawater pollution, and overexploitation are intensifying the activities of marine organisms such as inhabiting, migrating, and breeding, thereby affecting biodiversity and ecosystem functions.

[0003] eDNA (environmental DNA) is DNA directly extracted from environmental samples, which is a mixture of DNA from different species such as animals, plants, and microorganisms in the environment and records the species information of the ecosystem. In existing related technologies, eDNA technology is used to detect biodiversity. The patent application document with Chinese patent application number 202211567161.X and application date December 7, 2022 discloses a method for monitoring the biodiversity of reclaimed water-receiving rivers based on eDNA technology. This method performs PCR amplification on specific marker genes of animals and plants, conducts high-throughput sequencing, and performs bioinformatics analysis and species annotation, that is, the metabarcoding method, to obtain the composition of river biological species. However, the above method has certain deficiencies. Due to the inherent differences in the efficiency of PCR primers during the PCR amplification process, it is inevitable to cause deviations in community structure; at the same time, it relies on existing universal primers and has weak detection ability for unknown organisms; due to species annotation only based on specific marker gene fragments, the resolution of species annotation is low.

[0004] Metagenome is a technology for directly retrieving all genomic information from environmental samples, but its application in characterizing community organisms is still relatively few at present. Summary of the Invention

[0005] In order to improve the existing eDNA technology for monitoring biodiversity and construct a method for monitoring marine biodiversity using marine sediments, the present invention is hereby proposed.

[0006] A method for detecting marine biodiversity based on eDNA metagenome and metabarcoding, comprising the following steps:

[0007] Step 1: Set sampling points and collect sediment samples from the monitoring area;

[0008] Step 2: Pretreat the sediment samples and extract eDNA;

[0009] Step 3: Perform metagenomic high-throughput sequencing on the extracted eDNA to obtain the original metagenomic sequences;

[0010] Step 4: Perform quality control on the original metagenomic sequences to remove low-quality sequences and obtain the quality-controlled metagenomic sequences; Step 5: Assemble the quality-controlled metagenomic sequences based on repetitive fragments and screen out sequences with lengths greater than the length threshold to obtain long metagenomic sequences;

[0011] Step 6: Establish a database for aligning conserved regions of specific marker genes according to specific marker genes; the specific marker genes include COI gene, rbcL gene, and ITS2 gene;

[0012] Align and screen out the sequences in the long metagenomic sequences that contain the conserved regions of the specific marker genes of the target species to obtain query sequences;

[0013] Step 7: Establish a gene alignment database according to specific marker genes;

[0014] Align each query sequence with the gene alignment database; obtain the m database sequences with the highest consistency corresponding to each query sequence; where m is a preset parameter;

[0015] Step 8: For each query sequence, starting from the lowest taxonomic level, if the number of database sequences at the same taxonomic level is not less than n, annotate the query sequence to this taxonomic level, otherwise gradually increase the taxonomic level until the number of database sequences at the same taxonomic level is not less than n; where n is a preset parameter.

[0016] Preferably, the pretreatment is specifically to rinse with PBS solution.

[0017] Preferably, the specific method for performing quality control on the original metagenomic sequences to remove low-quality sequences is: shear the tail bases of the sequences with quality values lower than 20, and at the same time remove the sequences with more than 10% N bases.

[0018] Preferably, the length threshold is 1000bp.

[0019] Preferably, the method for establishing a database for aligning conserved regions of specific marker genes according to specific marker genes is: using a hidden Markov model.

[0020] Preferably, the alignment method for aligning each query sequence with the gene alignment database is the ublast algorithm.

[0021] The beneficial effects that can be achieved by this application specifically include:

[0022] (1) Compared with the existing environmental DNA technology, the present invention combines two methods of metagenomic technology and metabarcoding technology, incorporating the advantages of both technologies. The metagenomic technology has simpler processing of environmental DNA samples without the need for PCR amplification, avoiding species composition biases caused by PCR amplification. The assembled long sequence genes provide more complete and abundant species information, helping to improve species annotation resolution and detect unknown organisms. The metabarcoding technology has a rich database of specific marker genes, and the downstream data analysis method adopts a two-step approach of preliminary screening and strict alignment of specific marker gene fragments, further reducing the requirements for bioinformatics computing resources.

[0023] (2) Compared with traditional field investigation methods for ecosystems, the present invention applies molecular biology methods, which are more efficient, accurate, and comprehensive for the investigation and monitoring of biodiversity, and cause less interference to ecosystems and organisms. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flowchart of the invention;

[0025] Figure 2 is a graph of the percentage of species classification annotations of the query sequence;

[0026] Figure 3 is a graph of the species classification results. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0028] As Figure 1 shown, a method for detecting marine biodiversity based on eDNA metagenomics and metabarcoding includes the following steps:

[0029] Step 1: Set sampling points and collect sediment samples from the monitoring area;

[0030] In this embodiment, the sampling points are set as four stations in the coastal area of a certain place, including MS13 (Dapeng Bay), SS4 (Southern Waters), TS2 (Tolo Harbour), and VS3 (Victoria Harbour). 500 g of sediment samples are collected at each point. The sediment samples are stored at -20 °C before the next step of processing.

[0031] Step 2: Pretreat the sediment samples and extract eDNA; the specific pretreatment is to rinse with PBS solution;

[0032] In this example, within 12 hours after the sediment samples are collected, the sediment samples collected at each point are rinsed three times with PBS solution to reduce humic acid and improve the quality of eDNA. Use the DNeasy® PowerSoil® Kit to extract eDNA from the sediment samples, and use a Qubit fluorometer and a NanoDrop spectrophotometer to measure the concentration and quality (260 / 280, 260 / 230) of the extracted eDNA respectively. The extracted eDNA is stored at -80 °C.

[0033] Step 3: Perform metagenomic high-throughput sequencing on the extracted eDNA to obtain metagenomic raw sequences;

[0034] In this example, the extracted eDNA is sent to a commercial sequencing company for metagenomic high-throughput sequencing, and each sample generates 10 GB of metagenomic raw sequences.

[0035] Step 4: Perform quality control on the metagenomic raw sequences to remove low-quality sequences and obtain metagenomic quality-controlled sequences; among them, the specific method for performing quality control on the metagenomic raw sequences and removing low-quality sequences is: shear the tail bases of the sequences with a quality value lower than 20, and at the same time remove the sequences with more than 10% N bases;

[0036] In this example, use the fastp software to shear the tail bases of the sequences with a quality value (Q score) lower than 20 in the metagenomic raw sequences, and remove the sequences with more than 10% N bases in the sequencing data to obtain the quality-controlled metagenomic quality-controlled sequences.

[0037] Step 5: Assemble the metagenomic quality-controlled sequences according to the repetitive fragments, and screen out the sequences with a length greater than the length threshold to obtain metagenomic long sequences; the length threshold is 1000 bp;

[0038] In this example, use the megahit software to assemble the metagenomic quality-controlled sequences according to the repetitive fragments. Use seqkit to screen out the sequences with a length greater than 1000 bp after assembly to obtain metagenomic long sequences, and the metagenomic long sequences contain more comprehensive gene content.

[0039] Step 6: According to the specific marker genes, use the hidden Markov model to establish a database for aligning the conserved regions of the specific marker genes; the specific marker genes include COI gene, rbcL gene, ITS2 gene;

[0040] Align and screen the sequences in the metagenomic long sequences that contain the conserved regions of the specific marker genes of the target species to obtain query sequences;

[0041] In this embodiment, first, download the specific marker gene database of the target species. Download the COI and rbcL gene sequences from the NCBI website (https: / / www.ncbi.nlm.nih.gov / , the National Center for Biotechnology Information in the United States, which is an institution under the National Institutes of Health (NIH) in the United States and was established in 1988. The main task of NCBI is to develop and maintain databases and tools in the biomedical field and provide free bioinformatics resources and services for researchers around the world), download the COI sequence from the BOLD database (https: / / v4.boldsystems.org / , a DNA barcode database developed by the Canadian Centre for Biodiversity Genomics), and download the ITS gene sequence from the UNITE database (https: / / unite.ut.ee, a database for fungal DNA barcoding and taxonomy jointly developed by ten Nordic academic institutions). Among them, the COI gene is used for animal identification, the rbcL gene is used for plant identification, and the ITS2 gene is used for fungal identification.

[0042] Use the Hidden Markov Model (HMM) of the HMMER software to establish a database for aligning the conserved regions of specific marker genes. Hmmer is a sequence alignment tool based on a deep learning algorithm of the hidden Markov model, which is more sensitive to sequence fragments with higher similarity and can better capture information about unknown species.

[0043] Use the HMMER software to align and screen the sequences in the metagenomic long sequences that contain the conserved regions of the specific marker genes of the target species to obtain query sequences.

[0044] Step 7: Establish a gene alignment database according to the specific marker genes;

[0045] Use the ublast algorithm to align each query sequence with the gene alignment database; obtain the m database sequences with the highest consistency corresponding to each query sequence; where m is a preset parameter;

[0046] In this embodiment, the ublast algorithm in the USEARCH software is used to establish a comparison database for the specific marker genes of COI, rbcL, and ITS2. The ublast algorithm in the USEARCH software is used to compare the query sequences with the gene comparison database. The Ublast algorithm is a method for local sequence alignment search, which is used to compare and analyze the similarities and differences between sequences. In the comparison, the expected value (E-value, which refers to the number of local alignments with a given score that are expected to be found in a random sequence with the same length as the query sequence and the database) is set to 1e-5, and the 5 database sequences with the highest sequence identity to each query sequence are assigned (sequence identity refers to the percentage of the number of identical bases at corresponding positions in the aligned length where the query sequence and the database sequence are the same, and is used to measure the similarity between two sequences).

[0047] Step 8: For each query sequence, starting from the lowest taxonomic level, if the number of database sequences at the same taxonomic level is not less than n, then the query sequence is annotated to this taxonomic level; otherwise, the taxonomic level is gradually increased until the number of database sequences at the same taxonomic level is not less than n; where n is a preset parameter.

[0048] In this embodiment, starting from the lowest taxonomic level (species) of species classification, if among the 5 aligned database sequences, no less than 4 database sequences belong to the same species level A, then the species classification annotation of the query sequence is species A. If less than 4 database sequences belong to the same species level, then the species annotation is performed at the genus level of the previous taxonomic level. And so on, until no less than 4 database sequences belong to the same taxonomic level. In this way, all query sequences can obtain more accurate species annotation results.

[0049] A total of 2369 COI gene sequence variants, 733 rbcL gene sequence variants, and 1662 TIS2 gene sequence variants were detected in the four sediment samples. Gene sequence variants are high-resolution sequence variation units at the nucleotide level obtained through precise sequence alignment. Each gene sequence variant has its sequence uniqueness, retains the single nucleotide differences between sequences, can distinguish the subtle differences between species, and more accurately reflects microbial diversity. Each gene sequence variant represents a unique biological sequence, which may correspond to a specific microbial taxonomic unit (such as a species or a strain) and has clear biological significance.

[0050] The annotation percentages of the four sediment samples at different taxonomic levels are as Figure 2Among them, the proportions of query sequences that can be annotated to the levels of kingdom, phylum, class, order, family, genus, and species are 82%, 66%, 53%, 48%, 38%, 19%, and 5% respectively. The species classification results of eukaryotes (animals, plants, fungi) are as Figure 3 shown.

[0051] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for detecting marine biodiversity based on eDNA metagenomics and metabarcoding, characterized in that It includes the following steps: Step 1: Set sampling points and collect sediment samples in the monitoring area; Step 2: Pretreat the sediment samples and extract eDNA; Step 3: Perform metagenomic high-throughput sequencing on the extracted eDNA to obtain the original metagenomic sequences; Step 4: Perform quality control on the original metagenomic sequences to remove low-quality sequences and obtain the quality-controlled metagenomic sequences; Step 5: Assemble the quality-controlled metagenomic sequences based on repetitive fragments and screen out sequences with lengths greater than the length threshold to obtain long metagenomic sequences; Step 6: Establish a comparison database for the conserved regions of specific marker genes according to specific marker genes; the specific marker genes include COI gene, rbcL gene, and ITS2 gene; Compare and screen out the sequences containing the conserved regions of the specific marker genes of the target species in the long metagenomic sequences to obtain query sequences; Step 7: Establish a gene comparison database according to specific marker genes; Compare each query sequence with the gene comparison database; obtain the m database sequences with the highest consistency corresponding to each query sequence; where m is a preset parameter; Step 8: For each query sequence, starting from the lowest taxonomic level, if the number of database sequences at the same taxonomic level is not less than n, annotate the query sequence to this taxonomic level, otherwise gradually increase the taxonomic level until the number of database sequences at the same taxonomic level is not less than n; where n is a preset parameter.

2. The method for detecting marine biodiversity based on eDNA metagenome and metabarcoding according to claim 1, characterized in that The pretreatment is specifically to rinse with PBS solution.

3. The method for detecting marine biodiversity based on eDNA metagenome and metabarcoding according to claim 1, wherein The specific method for performing quality control on the original metagenomic sequences to remove low-quality sequences is: shear the tail bases of sequences with quality values lower than 20, and at the same time remove sequences with more than 10% N bases.

4. A method for detecting marine biodiversity based on eDNA metagenome and metabarcoding according to claim 1, characterized in that The length threshold is 1000bp.

5. A method for detecting marine biodiversity based on eDNA metagenome and metabarcoding according to claim 1, characterized in that, The method for establishing a comparison database for the conserved regions of specific marker genes according to specific marker genes is: using the hidden Markov model.

6. The method for detecting marine biodiversity based on eDNA metagenome and metabarcoding according to claim 1, characterized in that, The comparison method for comparing each query sequence with the gene comparison database is the ublast algorithm.

Citation Information

Patent Citations

  • Method for monitoring biodiversity of reclaimed water receiving river based on environmental DNA technology

    CN116103382A

  • B ephrin regulation of G-protein coupled chemoattraction, compositions, and methods of use

    US60280260P0

  • Method for evaluating diversity of marine zooplankton on basis of macro bar code technology

    CN111172258A

  • Environmental DNA macro bar code method for researching macrobenthic animal community structure

    CN112359119A

  • EDNA-based aquatic ecology analysis method and system

    CN112735533A