Method for analyzing diversity of biocenosis of ship sediment by utilizing eDNA (enhanced deoxyribonucleic acid) technology

Through the application of eDNA technology, the accuracy of detection of foreign species in ship ballast water is solved, efficient analysis of sediment biological communities and early warning of early invasion, reducing biological damage.

CN120290734APending Publication Date: 2025-07-11CHINA WATERBORNE TRANSPORT RES INST
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411931405.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to accurately detect foreign species in ship ballast water. Traditional morphological methods are difficult for small or low-density biological samples, and molecular biological methods have problems such as false positives, insufficient sensitivity and inability to distinguish living dead organisms.

Method used

EDNA technology was used to conduct biome diversity analysis of ship sediment, including sampling, design of primer sequences, PCR amplification, construction of Illumina PE250 library and sequencing, and data processing and species annotation were performed in combination with bioinformatics analysis.

Benefits of technology

Accurate analysis of ship sediment biomes is achieved, which can detect alien species invasions early, reduce biological damage, and provide efficient biological invasion control and early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120290734A_ABST
    Figure CN120290734A_ABST
Patent Text Reader

Abstract

The invention provides a method for analyzing the diversity of ship sediment biocenosis by using an eDNA technology, and belongs to the technical field of ship sediment biocenosis analysis. The method for analyzing the diversity of the marine sediment biocenosis by using the eDNA technology comprises the following steps: sampling a sample in a target port water area, designing and synthesizing a primer sequence, carrying out PCR amplification, product purification and PCR product quantification, constructing an Illumina PE250 library, carrying out Illumina PE250 sequencing, and analyzing a sequencing result. The primer sequences comprise at least 18 primer sequences in nucleotide sequences as shown in SEQ ID NO. 1 to SEQ ID NO. 24. By adopting the eDNA, the biocenosis diversity of the ship sediment can be accurately analyzed, and the purpose of early detection of foreign species invasion is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of ship sediment biocoenosis analysis, and particularly relates to a method for analyzing the biodiversity of ship sediment biocoenosis by using eDNA technology. Background Art

[0002] Ocean transportation plays a leading role in the global foreign trade cargo transportation, and more than 80% of the world's goods are transported by ships. To ensure normal navigation and safety, ships need to adjust the draft, transverse and longitudinal balance of the hull, and the safe metacentric height by ballasting during navigation and operation to ensure the navigation safety of the ship. During the navigation process, ships need to adjust the ballast water, and the adjustment method mainly relies on the discharge and loading of ballast water. When a ship loads ballast water at a certain port and discharges it at another port, the ballast water supply and discharge process of the ship will bring aquatic organisms in the sea area of one port to the sea area of another port. If the water quality conditions of the two sea areas are similar, these aquatic organisms will reproduce abnormally in the new sea area, causing great harm to the local marine ecosystem, fishery production, human health, etc. Therefore, the detection and research of alien invasive species in ballast water have attracted more and more attention from industry experts. Traditional species distribution analysis mainly uses on-site investigations and morphological methods for species identification, which has good results for species that are easy to sample and have relatively large individuals. However, in many environments, at certain specific life stages, and for groups with very low densities, on-site sampling is very difficult, and the morphological identification method has high requirements for the integrity of specimens and the professional qualities of classification and identification personnel, resulting in poor biological analysis results. Therefore, there is an urgent need to find an economical and efficient species identification and investigation method to provide technical support for the research on the types of alien species in ballast water.

[0003] In recent years, with the rapid development of molecular biology research, the species identification method has transitioned from traditional morphological classification to molecular biology, especially the development and application of DNA classification technology - Environmental DNA (eDNA) technology. Environmental DNA refers to free DNA molecules released from skin, mucus, saliva, secretions, eggs, feces, fruits, pollen, etc. Environmental DNA technology is a method of directly extracting DNA fragments from environmental samples (such as soil, sediment, and water body) and then using sequencing technology for qualitative or quantitative analysis. The environmental DNA analysis process is as Figure 1As shown. Currently, there are also some problems in the process of using environmental DNA for research, including how to accurately, fairly and detailedly record environmental DNA and how to best extract and analyze the genetic information in environmental DNA with modern technologies, because environmental DNA will degrade rapidly when exposed to oxygen, light, high temperature, DNA enzymes or water. For example, environmental DNA is detected but the species actually does not exist, that is, a false positive result appears. The appearance of such a result should be due to sample contamination; the PCR primers and probes do not have high enough specificity, resulting in the amplification of similar non-target species; the target species exists but the environmental DNA is not detected. This result may be due to insufficient sensitivity or the method not proceeding as expected; environmental DNA cannot distinguish dead and living organisms. Therefore, there is an urgent need in this field to provide a method for analyzing the microbial community diversity of ship sediments with high accuracy using eDNA technology. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide a method for analyzing the biodiversity of ship sediment communities using eDNA technology, which can accurately analyze the biodiversity of ship sediment communities.

[0005] To achieve the above invention purpose, the present invention provides the following technical solutions:

[0006] The present invention provides a method for analyzing the biodiversity of ship sediment communities using eDNA technology, including the following steps: taking samples in the target port waters, designing and synthesizing primer sequences for PCR amplification, product purification and quantification of PCR products, constructing an Illumina PE250 library, performing Illumina PE250 sequencing, and analyzing the sequencing results; the primer sequences include at least 18 of the nucleotide sequences shown in SEQ ID NO.1 to SEQ ID NO.24.

[0007] Preferably, the sample taken is the sediment at a distance of 0 - 10 cm from the bottom of the water.

[0008] Preferably, the reaction system of the PCR amplification is composed of, calculated as 20 μl: 5×FastPfu Buffer 4 μl, 2.5 mM dNTPs 2 μl, 5 μM forward primer 0.8 μl, 5 μM reverse primer 0.8 μl, FastPfu Polymerase 0.4 μl, template DNA 10 ng, and ddH2O, and the ddH2O is made up to 20 μl.

[0009] Preferably, the reaction program of the PCR amplification is: 95°C for 5 min; 95°C for 30 s, 58°C for 30 s, 72°C for 45 s for 40 cycles; 72°C for 10 min; and stored at 10°C.

[0010] Preferably, the construction of the IlluminaPE250 library includes the following steps: ligating a "Y"-shaped adaptor; using magnetic beads to screen and remove self-ligated adaptor fragments; enriching the library template by PCR amplification; denaturing with sodium hydroxide to generate single-stranded DNA fragments.

[0011] Preferably, the IlluminaPE250 sequencing includes the following steps: one end of the DNA fragment is complementary to the primer base and fixed on the chip; the other end randomly pairs with another primer and is also fixed to form a "bridge"; PCR amplification to generate DNA clusters; linearizing the DNA amplicons into single strands; adding a modified DNA polymerase and dNTPs with four fluorescent labels, and only one base is synthesized in each cycle; scanning the surface of the reaction plate with a laser to read the nucleotide types polymerized in the first-round reaction of each template sequence; chemically cleaving the "fluorescent group" and "termination group" to restore the 3'-end sticky ends and continue polymerizing the second nucleotide; counting the results of the fluorescent signals collected in each round to obtain the sequence of the template DNA fragment.

[0012] Preferably, the analysis includes clustering and species annotation, sample diversity analysis, species composition analysis, and / or inter-group diversity analysis.

[0013] Advantages of the present invention:

[0014] The present invention uses environmental DNA technology (eDNA) to accurately analyze the biodiversity of ship sediments, can be applied to the field of detecting alien species in ship ballast water, can detect certain specific life stages and groups with very low densities, achieve the purpose of early detection of alien species invasion, better control biological invasion, reduce biological damage, and play a warning effect on ships and shipping lanes with high-risk invasive species. Description of the Drawings

[0015] Figure 1 is the environmental DNA analysis process;

[0016] Figure 2 is the data analysis process;

[0017] Figure 3 is the number of sample sequences and average length obtained by primer 1 and primer 2, where the left figure is the result of primer 1 and the right figure is the result of primer 2;

[0018] Figure 4 is the sample dilution curve, the left figure is the result of primer 1 and the right figure is the result of primer 2;

[0019] Figure 5 is the sample Shannon-Wiener curve, the left figure is the result of primer 1 and the right figure is the result of primer 2;

[0020] Figure 6 is the sample Rank-Abundance curve. The left figure shows the result of primer 1, and the right figure shows the result of primer 2;

[0021] Figure 7 is the Specaccum species accumulation curve. The left figure shows the result of primer 1, and the right figure shows the result of primer 2;

[0022] Figure 8 is the species composition component map. The left figure shows the result at the phylum level, and the right figure shows the result at the genus level;

[0023] Figure 9 is the species clustering component map (phylum). The left figure shows the result of primer 1, and the right figure shows the result of primer 2;

[0024] Figure 10 is the community Heatmap. The left figure shows the result at the phylum level, and the right figure shows the result at the genus level;

[0025] Figure 11 is the dominant species community structure component map. The left figure shows the result of primer 1, and the right figure shows the result of primer 2;

[0026] Figure 12 is the distance heatmap between samples. The left figure shows the result of primer 1, and the right figure shows the result of primer 2;

[0027] Figure 13 is the principal component analysis map of samples. The left figure shows the result of primer 1, and the right figure shows the result of primer 2;

[0028] Figure 14 is the result of the similarity analysis of the sample community composition. The left figure shows the result of primer 1, and the right figure shows the result of primer 2;

[0029] Figure 15 is the analysis of the microbial evolutionary similarity of samples. The left figure shows the result of primer 1, and the right figure shows the result of primer 2;

[0030] Figure 16 is the similarity analysis of multiple samples. The left figure shows the result of primer 1, and the right figure shows the result of primer 2;

[0031] Figure 17 is the analysis of the evolutionary distance between samples. The left figure shows the result of primer 1, and the right figure shows the result of primer 2;

[0032] Figure 18 is the species composition component map. The left figure shows the result at the phylum level, and the right figure shows the result at the genus level;

[0033] Figure 19 is the community Heatmap. The left figure shows the result at the phylum level, and the right figure shows the result at the genus level. Detailed implementation manner

[0034] The present invention provides a method for analyzing the biodiversity of ship sediment biocommunities using eDNA technology, comprising the following steps: taking samples in the target port waters, designing and synthesizing primer sequences for PCR amplification, product purification, and quantification of the PCR products, constructing an Illumina PE250 library, performing Illumina PE250 sequencing, and analyzing the sequencing results; the primer sequences include at least 18 of the nucleotide sequences shown in SEQ ID NO.1 to SEQ ID NO.24.

[0035] In the present invention, the sample-taking is to take sediments at a distance of 0-10 cm from the bottom or the surface. After sampling, it is stored in a sterile sampling bag or an EP tube, transported on ice to the laboratory for nucleic acid extraction, or stored at -80 °C for later use. The storage of the samples taken is preferably to store the samples at room temperature under 95% alcohol concentration. The present invention has no special limitation on the specific method for extracting DNA from the samples, and conventional DNA extraction methods in the art can be used. In the present invention, the basic information of the primer sequences is shown in Table 1 below. In Table 1, R, W, Y, N, and D in the primer sequences all represent degenerate bases, where R represents A or G, W represents A or T, Y represents C or T, N represents A or T or C or G, and D represents G or A or T.

[0036] Table 1 Basic information of primer sequences

[0037]

[0038]

[0039] In the present invention, when performing PCR amplification using the primer sequences in Table 1, the primer sequences preferably need to include seawater fish, freshwater fish, eukaryotic invertebrates, benthic animals, diatoms, soil eukaryotes, fungi, and plant primer sequences at the same time. In the present invention, when performing PCR amplification, the reaction system for the PCR amplification is preferably composed of, based on 20 μl: 4 μl of 5×FastPfuBuffer, 2 μl of 2.5 mM dNTPs, 0.8 μl of 5 μM forward primer, 0.8 μl of 5 μM reverse primer, 0.4 μl of FastPfuPolymerase, 10 ng of template DNA, and ddH2O, and the ddH2O is made up to 20 μl. The present invention has no special limitation on the specific sources of the components in the reaction system. In the present invention, the reaction program for the PCR amplification is preferably: 95 °C for 5 min; 95 °C for 30 s, 58 °C for 30 s, 72 °C for 45 s for 40 cycles; 72 °C for 10 min; and stored at 10 °C.

[0040] In the present invention, the construction of an Illumina PE250 library preferably includes the following steps: ligating a "Y"-shaped adaptor; using magnetic beads to screen and remove self-ligated adaptor fragments; enriching the library template by PCR amplification; denaturing with sodium hydroxide to generate single-stranded DNA fragments. When performing Illumina PE250 sequencing in the present invention, it preferably includes the following steps: one end of the DNA fragment is base-complementary to a primer and fixed on the chip; the other end randomly complements another primer and is also fixed, forming a "bridge"; PCR amplification to generate DNA clusters; linearizing the DNA amplicons into single strands; adding a modified DNA polymerase and dNTPs with four fluorescent labels, and only one base is synthesized in each cycle; scanning the surface of the reaction plate with a laser to read the nucleotide types polymerized in the first-round reaction of each template sequence; chemically cleaving the "fluorescent group" and "termination group" to restore the 3'-end sticky ends and continue polymerizing the second nucleotide; counting the results of the fluorescent signals collected in each round to obtain the sequence of the template DNA fragment.

[0041] In the present invention, when analyzing the sequencing results, a bioinformatics analysis method is preferably adopted. The bioinformatics analysis process is preferably as follows: The PE reads obtained by Illumina PE250 sequencing are first assembled according to the overlap relationship, and at the same time, quality control and filtering of the sequences are performed. After differentiating the samples, OTU clustering analysis and species taxonomic analysis are carried out. Based on the OTU clustering analysis results, various diversity index analyses of the OTUs can be performed, as well as the detection of sequencing depth; based on the taxonomic information, statistical analysis of the community structure can be performed at each taxonomic level. On the basis of the above analysis, a series of in-depth statistical and visualization analyses such as community structure and phylogeny can be carried out. The above data analysis flow chart is as Figure 2 shown. In the present invention, the above analysis preferably includes clustering and species annotation, sample diversity analysis, species composition analysis, and / or inter-group diversity analysis.

[0042] The technical solutions provided by the present invention will be described in detail below in conjunction with the embodiments, but they should not be construed as limiting the protection scope of the present invention.

[0043] In the following embodiments, unless otherwise specified, they are all conventional methods.

[0044] In the following embodiments, the materials, reagents, etc. used, unless otherwise specified, can all be obtained from commercial sources.

[0045] Example 1

[0046] A method for analyzing the biodiversity of ship sediment biological communities using the eDNA technique, which consists of the following steps:

[0047] 1. Sampling of samples: Sampling is carried out in the waters of the target port. Specifically, 5 g of sediment at 5 cm and 9 cm from the bottom of the water are taken respectively. The sediment at 5 cm from the bottom is denoted as 2020-ND-1-1, and the sediment at 9 cm from the bottom is denoted as 2020-ND-9. They are respectively stored in sterile sampling bags and transported on ice to the laboratory for DNA extraction.

[0048] 2. PCR amplification of the extracted DNA using specific primers with barcodes. The nucleotide sequences of the primers are shown in SEQ ID NO.1~4 and SEQ ID NO.7~20, denoted as primer 1. PCR uses TransGen AP221-02: TransStart Fastpfu DNA Polymerase; PCR instrument: ABI Model 9700. The reaction system for PCR amplification is: 4 μl of 5×FastPfu Buffer, 2 μl of 2.5 mM dNTPs, 0.8 μl of 5 μM forward primer, 0.8 μl of 5 μM reverse primer, 0.4 μl of FastPfu Polymerase, 10 ng of template DNA, and ddH2O is added to make up to 20 μl. The reaction program for PCR amplification is: 95°C for 5 min; 95°C for 30 s, 58°C for 30 s, 72°C for 45 s for 40 cycles; 72°C for 10 min; stored at 10°C. Each sample has 3 replicates. The PCR products of the same sample are mixed and detected by 2% agarose gel electrophoresis. The PCR products are recovered by cutting the gel using the AxyPrep DNA Gel Extraction Kit (AXYGEN company) and eluted with Tris_HCl; detected by 2% agarose electrophoresis.

[0049] 3. Referring to the preliminary quantitative results of electrophoresis, the PCR products are detected and quantified using the QuantiFluor TM -ST blue fluorescence quantification system (Promega company).

[0050] 4. Construction of Illumina PE250 library: (1) Connect the "Y"-shaped adapter; (2) Use magnetic beads to screen and remove self-ligated adapter fragments; (3) Enrich the library template by PCR amplification; (4) Denature with sodium hydroxide to generate single-stranded DNA fragments.

[0051] 5. Illumina PE250 Sequencing: (1) One end of the DNA fragment is complementary to the primer base and is fixed on the chip; (2) The other end randomly hybridizes with another nearby primer and is also fixed, forming a "bridge"; (3) PCR amplification is performed to generate DNA clusters; (4) The DNA amplicons are linearized into single strands; (5) Modified DNA polymerase and dNTPs with four fluorescent labels are added, and only one base is synthesized in each cycle; (6) The surface of the reaction plate is scanned with a laser to read the types of nucleotides polymerized in the first-round reaction of each template sequence; (7) The "fluorescent group" and "terminating group" are chemically cleaved to restore the 3'-end sticky ends and continue to polymerize the second nucleotide; (8) The fluorescence signal results collected in each round are statistically analyzed to obtain the sequence of the template DNA fragment.

[0052] 6. Bioinformatics Analysis of Illumina PE250 Sequencing Results:

[0053] The PE reads obtained from Illumina PE250 sequencing are first assembled according to the overlap relationship. At the same time, quality control and filtering of the sequences are performed, and after differentiating the samples, OTU clustering analysis and species taxonomic analysis are carried out. Based on the results of OTU clustering analysis, various diversity index analyses of OTUs are performed, as well as the detection of sequencing depth; based on taxonomic information, statistical analysis of community structure is performed at each taxonomic level. The above analysis process is as Figure 2 shown.

[0054] 6.1 OTU Clustering and Species Annotation

[0055] 6.1.1 Data Optimization and Statistics

[0056] The Illumina PE250 sequencing sequences first need to obtain the valid sequences of all samples according to the barcode; then quality control filtering of the reads is performed; then, according to the overlap relationship between the PE reads, the paired reads are assembled (merged) into one sequence; finally, the high-quality sequences of each sample are obtained by splitting according to the barcode and primer sequences, and the sequence directions are corrected and chimeras are removed according to the forward and reverse barcodes and primer directions during the process.

[0057] Data Deburring Methods and Parameters:

[0058] Filter the bases with a quality value below 20 at the tail of the read. Set a 10-bp window. If the average quality value within the window is below 20, truncate the trailing bases from the start of the window, and filter the reads less than 50 bp after quality control.

[0059] According to the overlap relationship between PE reads, pair of reads are merged into one sequence, and the minimum overlap length is 10 bp;

[0060] The maximum mismatch ratio allowed in the overlap region of the merged sequence is 0.2, and sequences that do not meet the requirements are filtered out;

[0061] Samples are distinguished according to the barcodes and primers at both ends of the sequence, and the sequence direction is adjusted. The allowed number of mismatches for the barcode is 0, and the maximum number of primer mismatches is 2;

[0062] Use Usearch software and the gold database to remove chimeras by combining de novo and reference methods

[0063] Software used: Trimmomatic, FLASH, Usearch, qiime programs.

[0064] Optimize the data volume statistics and length distribution as shown in Table 2 and Figure 3 the left figure in:

[0065] Table 2 Number of sample sequences and average length

[0066] Primer Optimized data Number of samples Number of sequences Total number of bases in the sequences Average length of the sequences in the optimized data 1 jgHCO2198R 2 1368764 428522732 313.07 2 REV3R 2 1363520 525679915 385.53

[0067] 6.1.2 OTU Clustering

[0068] OTU (Operational Taxonomic Units) is a same flag artificially set for a certain taxonomic unit (strain, genus, species grouping, etc.) for the convenience of analysis in phylogenetic or population genetics research. To understand the number information of bacteria, genera, etc. in the sequencing results of a sample, it is necessary to classify (cluster) the sequences. Through the classification operation, the sequences are grouped into many small groups according to their similarities to each other, and one small group is an OTU. OTUs can be divided for all sequences according to different similarity levels. Usually, OTUs at a similarity level of 98.65% for the full-length sequences are used for bioinformatics statistical analysis.

[0069] The analysis steps are as follows:

[0070] Extract non-redundant sequences from the optimized sequences to reduce the redundant calculation amount in the intermediate process of analysis; remove single sequences without duplicates; perform OTU clustering on the non-redundant sequences (excluding single sequences) at a similarity of 98.65%, and remove chimeras during the clustering process to obtain the representative sequences of OTUs; map all optimized sequences to the OTU representative sequences, and select sequences with a similarity of more than 98.65% to the OTU representative sequences to generate an OTU table. The results are shown in Table 3.

[0071] Table 3 Sample OTU Clustering

[0072] OTUID 2020 - ND - 1 - 1 2020 - ND - 9 OTU1 112752 1 OTU2 104134 2 OTU3 2675 73029 OTU4 44022 0 OTU5 40566 0

[0073] 6.1.3 Taxonomic Analysis

[0074] To obtain the species taxonomic information corresponding to each OTU, the uclust algorithm was used to perform taxonomic analysis on the OTU representative sequences, and the community composition of each sample was statistically analyzed at each taxonomic level: domain, phylum, class, order, family, genus, and species.

[0075] 6.1.4 Taxonomic Statistics

[0076] The proportion of annotated species was statistically analyzed at each taxonomic level: domain, phylum, class, order, family, genus, and species.

[0077] The results of sample (2020-ND-1-1) are shown in Table 4 and Table 5.

[0078] Table 4 Total Species Information Statistics

[0079] Sample (2020 - ND - 1 - 1) Domain Phylum Class Order Family Genus Species OTU Primer 1 1 21 60 138 297 446 605 1184

[0080] Table 5 Taxonomic Information Statistics Below the Phylum Level

[0081] Phylum Class Order Family Genus Species OTU Annelida 2 3 3 3 3 9 Arthropoda 6 25 120 175 191 231 Bacillariophyta 4 11 13 16 23 31 Cercozoa 1 0 0 1 1 1 Chlorophyta 7 9 10 13 15 38

[0082] 6.2 Sample Diversity Analysis (α-Diversity)

[0083] 6.2.1 Diversity Index

[0084] In community ecology, the study of microbial diversity can reflect the abundance and diversity of microbial communities through single-sample diversity analysis (Alpha diversity), including a series of statistical analysis indices to estimate the species abundance and diversity of environmental communities.

[0085] Indices for calculating species abundance (Community richness) include: Chao; Richness: The species richness is the sum of the number of species with an abundance greater than 0 in the community, and the larger the value, the more species are present in the community.

[0086] Indices for calculating community diversity include: Shannon; Simpson; ACE index: an index for estimating species diversity using rare species, with a higher value indicating a more diverse community in terms of species. Evenness: also known as Pielou's evenness, which is the ratio of the actual Shannon index of a community to the maximum Shannon index that can be obtained in a community with the same species richness.

[0087] Sequencing depth indices include: Coverage.

[0088] The results are shown in Table 6.

[0089] Table 6 Analysis Results of Diversity Indices

[0090]

[0091]

[0092] 6.2.2 Rarefaction Curve

[0093] A rarefaction curve is constructed by randomly sampling a certain number of individuals from a sample, counting the number of species represented by these individuals, and using the number of individuals and the number of species. It can be used to compare the species richness in samples with different sequencing data volumes, and can also be used to illustrate whether the sequencing data volume of a sample is reasonable. By randomly sampling the sequences and constructing a rarefaction curve with the number of sequences drawn and the number of OTUs they can represent, when the curve flattens out, it indicates that the sequencing data volume is reasonable, and more data volume will only generate a small number of new OTUs. Conversely, it indicates that continued sequencing may still generate more new OTUs. Therefore, by constructing a rarefaction curve, the sequencing depth of the sample can be obtained. The results are as shown in the left figure of Figure 4 below.

[0094] 6.2.3 Shannon-Wiener Curve

[0095] The Shannon-Wiener index reflects the microbial diversity in a sample. A curve is constructed using the microbial diversity indices at different sequencing depths of each sample's sequencing volume to reflect the microbial diversity of each sample at different sequencing amounts. When the curve flattens out, it indicates that the sequencing data volume is large enough to reflect most of the microbial information in the sample. The results are as shown in the left figure of Figure 5 below.

[0096] 6.2.4 Rank-Abundance Curve

[0097] The Rank-abundance curve is a way to analyze diversity. The construction method is to count the number of sequences contained in each OTU in a single sample, rank the OTUs in descending order of abundance (the number of sequences contained), and then plot with the OTU rank as the abscissa and the number of sequences contained in each OTU (or the relative percentage content of the sequences in the OTU) as the ordinate. The Rank-abundance curve can be used to explain two aspects of diversity, namely species abundance and species evenness. Horizontally, the species abundance is reflected by the width of the curve. The higher the species abundance, the larger the range of the curve on the horizontal axis; the shape (smoothness) of the curve reflects the evenness of the species in the sample. The flatter the curve, the more evenly the species are distributed. The results are as Figure 6 shown in the left figure in Figure 6 In the left figure, the abscissa: OTU rank, "400" represents the 400th OTU ranked by abundance in the sample; the ordinate: the relative percentage content of the number of sequences in the OTU at this rank, that is, the number of sequences belonging to the OTU divided by the total number of sequences. The numbers on the vertical axis, for example, "100" represents a relative abundance of 100%, "10" represents a relative abundance of 10%, and so on.

[0098] 6.2.5 Specaccum species accumulation curve

[0099] Species accumulation curves are used to describe the situation of species increase as the sample size increases. They are effective tools for investigating the species composition of samples and predicting the species abundance in samples. In biodiversity and community surveys, they are widely used to judge whether the sample size is sufficient and to estimate species richness. Therefore, through species accumulation curves, not only can it be judged whether the sample size is sufficient, but on the premise of sufficient sample size, species accumulation curves can also be used to predict species richness. The results are as Figure 7 shown in the left figure in Figure 7 In the left figure in

[0100] 6.3 Species composition analysis

[0101] 6.3.1 Component diagram

[0102] According to the results of taxonomic analysis, the taxonomic comparison of one or more samples at each taxonomic level can be obtained. In the results, two pieces of information are included:

[0103] (1) What kinds of microorganisms are contained in the samples; (2) The sequence numbers of each microorganism in the samples, that is, the relative abundances of each microorganism.

[0104] Therefore, statistical analysis methods can be used to observe the community structure of samples at different taxonomic levels. When comparing the community structure analyses of multiple samples together, the changes can also be observed. Depending on whether the research object is a single or multiple samples, the results may be presented in different ways. Usually, intuitive forms such as pie charts or bar charts are used. The analysis of the community structure can be carried out at any taxonomic level.

[0105] Analysis levels: phylum, genus. The results are as Figure 8 shown.

[0106] 6.3.2 Combined Analysis of Sample Clustering Tree and Bar Chart

[0107] The hierarchical clustering analysis (Bray-Curtis algorithm) based on community composition among samples and the bar chart of the community structure of samples are shown in the left figure of Figure 9 . Analysis level: phylum.

[0108] 6.3.3 Community Heatmap

[0109] A Heatmap can reflect the data information in a two-dimensional matrix or table through color changes. It can intuitively represent the magnitude of data values with defined shades of color. According to the need, the data is clustered by species or abundance similarity among samples, and the clustered data is represented on the heatmap. High-abundance and low-abundance species can be grouped and aggregated, and the similarity and difference in community composition among multiple samples at each taxonomic level can be reflected through color gradients and similarity degrees.

[0110] Analysis levels: phylum, genus. The results are as Figure 10 shown.

[0111] 6.3.4 Component Diagram of Dominant Species Community Structure

[0112] Select some species with the highest abundances in a single or all samples to draw a bar chart and display the taxonomic information at the phylum level corresponding to the species. Species of the same color indicate that they come from the same phylum.

[0113] Analysis level: genus. The results are as Figure 11 shown in the left figure of

[0114] 6.4 β-Diversity Analysis between Groups (among Samples)

[0115] 6.4.1 Inter-group distance calculation

[0116] The difference degree of species or functional abundance distribution among samples can be quantitatively analyzed by the distance in statistics. The Bray-Curtis statistical algorithm is used to calculate the distance between pairwise samples to obtain a distance matrix, which can be used for subsequent beta diversity analysis and visualization statistical analysis. Representing the distance matrix using a heatmap can visually observe the high and low distribution of differences among samples.

[0117] Perform quartile calculation on the distances of multiple groups of samples in different classifications or environments, and compare the intra-group and inter-group distance distribution differences of different sample groups. At the same time, conduct multiple Student two-sample t-tests to judge the significant differences between sample groups.

[0118] Functions of the box plot: Identify data outliers; roughly estimate and judge data characteristics; compare the shapes of several batches of data. On the same number axis, the box plots of several batches of data are arranged in parallel, and the shape information such as the median, tail length, outliers, and distribution interval of several batches of data is clear at a glance.

[0119] The box plot (also known as the box-and-whisker plot) is a method of describing data using five statistics in the data: minimum value, first quartile, median, third quartile, and maximum value. It can also roughly see information such as whether the data is symmetric and the degree of dispersion of the distribution, especially for comparing several samples. A simple box plot consists of five parts, namely the minimum value, median, maximum value, and two quartiles.

[0120] The heatmap of the distance between samples is as shown in the left figure of Figure 12 the following.

[0121] 6.4.2 PCA analysis

[0122] PCA analysis (Principal Component Analysis), namely principal component analysis, is a technique for simplifying data analysis. This method can effectively identify the most "important" elements and structures in the data, remove noise and redundancy, reduce the dimensionality of the original complex data, and reveal the simple structure hidden behind the complex data. Its advantages are simplicity and no parameter restrictions. By analyzing the OTU (98.65% similarity) composition of different samples, the differences and distances between samples can be reflected. PCA uses variance decomposition to reflect the differences in multiple groups of data on a two-dimensional coordinate graph, and the coordinate axes take the two eigenvalues that can most reflect the variance value. For example, the more similar the sample composition, the closer the distance reflected in the PCA graph. Samples from different environments may show a scattered or aggregated distribution. The two or three components with the highest explanatory power for sample differences in the PCA results can be used to verify the hypothesized factors. The results are as Figure 13 shown in the left figure in. PC1 and PC2 respectively represent the suspected influencing factors for the deviation of the microbial composition of the two groups of samples. It is necessary to summarize and conclude in combination with the sample characteristic information. For example, if the samples of group A (red) and group B (blue) are separated in the direction of the pc1 axis, it can be analyzed that PC1 is the main factor causing the separation of group A and group B (which can be two locations or different acid-base conditions), and at the same time, it is verified that this factor has a high possibility of affecting the sample composition.

[0123] 6.4.3 PCoA analysis

[0124] PCoA analysis, namely principal co-ordinates analysis, is also a non-constrained data dimensionality reduction analysis method, which can be used to study the similarity or difference in the composition of sample communities, similar to PCA analysis; the main difference is that PCA is based on Euclidean distance, while PCoA is based on other distances except Euclidean distance, and potential principal components affecting the differences in the composition of sample communities are found through dimensionality reduction.

[0125] In PCoA analysis, a series of eigenvalues and eigenvectors are first sorted, and then the most important eigenvalues ranked among the top are selected and represented in the coordinate system. The result is equivalent to a rotation of the distance matrix, which does not change the relative position relationship between sample points, but only changes the coordinate system.

[0126] The results are as Figure 14As shown in the left figure in , PC1 and PC2 respectively represent the suspected influencing factors for the deviation of the microbial composition of two groups of samples, which need to be summarized in combination with the sample characteristic information. For example, if the samples of group A (red) and group B (blue) are separated in the direction of the PC1 axis, it can be analyzed that PC1 is the main factor causing the separation of group A and group B (which can be two locations or different acid-base levels), and at the same time, it is verified that this factor has a high probability of affecting the composition of the samples.

[0127] 6.4.4 PcoA Analysis Based on UniFrac

[0128] The distance matrix obtained from Unifrac analysis can be used for various analysis methods. Through the multivariate statistical method of PCoA analysis, it can visually display the similarities and differences in the evolution of microorganisms in different environmental samples.

[0129] PCoA (principal co-ordinates analysis) is a visualization method for studying data similarities or differences. After sorting through a series of eigenvalues and eigenvectors, the main eigenvalues ranked in the top few are selected. PCoA can find the most important coordinates in the distance matrix. The result is a rotation of the data matrix, which does not change the relative positions of the sample points, but only changes the coordinate system.

[0130] Unifrac PCoa is based on the evolutionary distance and explores the potential principal components that affect the differences in the community composition of samples at the evolutionary level.

[0131] The results are as Figure 15 As shown in the left figure in , PC1 and PC2 respectively represent the suspected influencing factors for the deviation of the microbial composition of two groups of samples, which need to be summarized in combination with the sample characteristic information. For example, if the samples of group A (red) and group B (blue) are separated in the direction of the PC1 axis, it can be analyzed that PC1 is the main factor causing the separation of group A and group B (which can be two locations or different acid-base levels), and at the same time, it is verified that this factor has a high probability of affecting the composition of the samples.

[0132] 6.4.5 Dendrogram of Multisample Similarity

[0133] Use the tree structure to describe and compare the similarity and difference relationships among multiple samples. First, use the algorithm that describes the community composition relationship and structure to calculate the distance between samples, that is, perform hierarchical clustering analysis based on the beta diversity distance matrix, and use the unweighted pair group method with arithmetic mean (UPGMA) algorithm to construct a tree structure, and obtain a tree-like relationship form for visual analysis.

[0134]

[0135] Among them:

[0136] SA,i represents the number of sequences contained in the i-th OTU in sample A;

[0137] SB,i represents the number of sequences contained in the i-th OTU in sample B.

[0138] The results are as Figure 16 shown in the left figure in. The branch length represents the distance between samples. Samples can be distinguished by different colors according to the pre-grouping, and the vertical lines at the ends indicate that the samples are clustered together with a relatively high similarity.

[0139] 6.4.6 Clustering number analysis based on UniFrac

[0140] The distance matrix obtained by Unifrac analysis can be used for various analysis methods. Through graphic visualization processing such as constructing phylogenetic trees by the unweighted pair group method with arithmetic mean (UPGMA) in hierarchical clustering, the similarities and differences in the evolution of microorganisms in different environmental samples can be intuitively displayed.

[0141] UPGMA (Unweighted pair group method with arithmetic mean) assumes that all nucleotides / amino acids have the same mutation rate during the evolution process, that is, there is a molecular clock. The evolutionary distance between samples can be observed through the distance of the branches and the proximity of the clustering.

[0142] The results are as Figure 17 shown in the left figure in.

[0143] Example 2

[0144] The difference from Example 1 is that primer 1 in step 2 is replaced by primer 2. The nucleotide sequences of the primer 2 are shown in SEQ ID NO.1-2, SEQ ID NO.5-16 and SEQ ID NO.21-24, and the rest are the same as in Example 1.

[0145] The corresponding results are shown in Table 2, Figure 3 the right figure in, Table 7, Table 8, Table 9, Figure 4 the right figure in, Figure 5 the right figure in, Figure 6 the right figure in, Figure 7 the right figure in, Figure 18 , Figure 9 the right figure in, Figure 19 , Figure 11 the right figure in, Figure 12 the right figure in,Figure 13 the right figure in Figure 14 the right figure in Figure 15 the right figure in Figure 16 the right figure in and Figure 17 shown by the right figure in

[0146] Table 7 Total Species Information Statistics

[0147] Sample (2020 - ND - 1 - 1) Domain Phylum Class Order Family Genus Species OTU Primer 2 1 23 67 151 226 395 570 4250

[0148] Table 8 Taxonomic Information Statistics below the Phylum Level

[0149] Phylum Class Order Family Genus Species OTU Annelida 2 3 3 3 3 3 Apicomplexa 2 4 8 9 12 18 Arthropoda 5 12 22 25 29 92 Bacillariophyta 4 15 20 33 53 183 Cercozoa 1 5 12 34 74 873

[0150] Table 9 Analysis Results of Diversity Index

[0151] Sample Reads Chao Richness Shannon Simpson ACE Evenness coverage 2020 - ND - 1 - 1 717989 4277.1615 4275 8.0495 0.0161 4287.6483 0.667366002958651 0.99993732494509 2020 - ND - 9 595668 1592.8139 1573 5.3706 0.0886 1601.5179 0.50574451027218 0.999879127299099

[0152] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A method for analyzing the biodiversity of ship sediment biocommunities using eDNA technology, characterized in that, It includes the following steps: taking samples in the target port waters, designing and synthesizing primer sequences for PCR amplification, product purification, and quantification of PCR products, constructing an Illumina PE250 library, performing Illumina PE250 sequencing, and analyzing the sequencing results; the primer sequences include at least 18 of the nucleotide sequences shown in SEQ ID NO.1 to SEQ ID NO.

24.

2. The method according to claim 1, characterized in that, The sample taking is to take sediments at a depth of 0 to 10 cm from the bottom.

3. The method according to claim 1, wherein The reaction system for the PCR amplification is composed of, based on 20 μl: 4 μl of 5×FastPfu Buffer, 2 μl of 2.5 mM dNTPs, 0.8 μl of 5 μM forward primer, 0.8 μl of 5 μM reverse primer, 0.4 μl of FastPfu Polymerase, 10 ng of template DNA, and ddH2O, and the ddH2O is supplemented to 20 μl.

4. The method according to claim 1, characterized in that, The reaction program for the PCR amplification is: 95°C for 5 min; 95°C for 30 s, 58°C for 30 s, 72°C for 45 s for 40 cycles; 72°C for 10 min; stored at 10°C.

5. The method according to claim 1, wherein The construction of the Illumina PE250 library includes the following steps: ligating a "Y"-shaped linker; using magnetic beads to screen and remove self-ligated linker fragments; enriching the library template by PCR amplification; denaturing with sodium hydroxide to generate single-stranded DNA fragments.

6. The method according to claim 1, wherein The Illumina PE250 sequencing includes the following steps: one end of the DNA fragment is base-complementary to a primer and fixed on the chip; the other end randomly complements another primer and is also fixed, forming a "bridge". PCR amplification to generate DNA clusters; linearizing the DNA amplicons into single strands; adding a modified DNA polymerase and dNTPs with 4 fluorescent labels, and only one base is synthesized in each cycle; scanning the surface of the reaction plate with a laser to read the nucleotide types polymerized in the first-round reaction of each template sequence; chemically cleaving the "fluorescent group" and "terminating group" to restore the 3'-end sticky ends and continue polymerizing the second nucleotide; counting the fluorescence signal results collected in each round to obtain the sequence of the template DNA fragment.

7. The method according to claim 1, wherein The analysis includes clustering and species annotation, sample diversity analysis, species composition analysis, and / or inter-group diversity analysis.

Citation Information

Patent Citations

  • High-throughput sequencing technology-based white fungus strain purity detection method and application thereof

    CN108486271A

  • Biological component analysis method of traditional Chinese medicine compound preparation

    CN110097976A

  • Operation method for influence of long-term coal pile on soil bacteria based on multi-dimensional indexes

    CN110714061A

  • Environmental DNA macro bar code method for researching macrobenthic animal community structure

    CN112359119A

  • Kit for detecting excrement DNA and bird feeding habit analysis method based on kit

    CN118932034A