A set of fish eDNA macro-barcoding PCGs primer group and its application
By designing COX2-eDNA primer sets and optimizing PCR reaction conditions, the problems of insufficient universality and species recognition rate of existing fish eDNA macrobarcode primer sets were solved, enabling more efficient fish diversity surveys and monitoring, and improving amplification capacity and species identification accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- YAZHOU BAY INNOVATION RESEARCH INSTITUTE HAINAN TROPICAL OCEAN UNIVERSITY
- Filing Date
- 2026-04-22
- Publication Date
- 2026-07-24
Smart Images

Figure CN122060877B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of molecular biology technology, and in particular to a set of fish eDNA macrobarcode PCG primers and their applications. Background Technology
[0002] Fish diversity is an important component of biodiversity. Traditional methods for surveying fish diversity mainly rely on netting, which has drawbacks such as low efficiency, time and labor consumption, significant damage to species and habitats, and difficulty in identifying juvenile individuals. In recent years, environmental DNA (eDNA) macrobarcoding technology has begun to be widely used in fish diversity surveys, showing promising application prospects in areas such as monitoring native fish species and detecting invasive fish species. eDNA macrobarcoding technology involves numerous steps, including baseline data setup, water environment sample collection, and sequencing result analysis. Significant progress has been made both domestically and internationally in these areas. For example, Chinese patent application number 202411845378.1, published on January 14, 2025, discloses a method for assessing aquatic ecological health based on large-scale experimental data of fish environmental DNA. This method includes the following steps: acquiring water samples from a target area; filtering the water samples to generate filtration data; centrifuging the water samples to generate centrifugation data; extracting fish environmental DNA from the water samples based on the filtration and centrifugation data to obtain regional fish environmental DNA data; identifying target DNA fragments from the regional fish environmental DNA data to obtain target DNA fragment information; and using data processing technology, high-throughput sequencing technology, and machine learning algorithms to perform flow parameter-water level correlation analysis on a specific water body area, and combining this with fish community structure to determine indicators, thereby improving the accuracy of aquatic ecological health assessment. For example, Chinese patent application No. 202211175107.0, published on May 23, 2023, discloses a method for quantifying fish based on environmental DNA macrobarcoding technology. This method, building upon existing technologies, adds *E. coli* transformed with a single-copy plasmid to the DNA extraction step. The extraction efficiency of the single-copy plasmid is used to correct for DNA extraction errors, resulting in a more accurate DNA copy number in the obtained sample. In the PCR amplification step, a gradient reference is added. After amplification, sequencing is performed to obtain the number of reads from the gradient reference. The linear slope between the DNA copy number and the number of reads is fitted, and combined with the number of reads from the sequenced sample DNA, the DNA copy number in the sample is calculated. This combined correction of DNA extraction, PCR amplification, and sequencing overcomes the shortcomings of existing methods that do not consider errors in the DNA extraction and PCR processes. This allows sequencing data to better reflect the absolute abundance of fish and has the potential for large-scale application.
[0003] Currently, fish eDNA macrobarcoding primer sets are mainly divided into two categories: ribosomal RNA gene (rRNA) primer sets and protein coding gene (PCG) primer sets. Among them, rRNA primer sets are the most commonly used for fish eDNA macrobarcoding. This is because rRNA primer sets have higher versatility and can amplify a wider range of fish species in practical applications. However, while rRNA primer sets have extremely high versatility, they suffer from low species recognition rates in amplified fragments. Fragments amplified by rRNA primer sets are more prone to sequence error merging during sequence alignment and species annotation, making them unsuitable for fish diversity surveys where high accuracy in species identification is required. In summary, there is currently no primer set in the field of fish eDNA macrobarcoding that simultaneously achieves high versatility and high species recognition rate. rRNA primer sets have extremely high versatility but low species recognition rates in amplified fragments; PCG primer sets have high species recognition rates in amplified fragments but poor versatility. The reason for this phenomenon is that the known fish eDNA metabarcode primer set is limited to four genes: mitochondrial 12S rRNAs, 16S rRNAs, COI, and Cytb. The advantages and disadvantages of other genes have not been explored, which limits the possibility of simultaneously solving the problems of universality and species identification. Therefore, it is urgent to develop a set of PCG primers with high universality and high species identification rate. Summary of the Invention
[0004] To address the aforementioned deficiencies in existing technologies, this invention proposes a set of fish eDNA macrobarcode PCG primers and their applications to solve the problems mentioned in the background.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A set of primers for fish eDNA macrobarcode PCGs, named COX2-eDNA, is provided. The forward primer COX2-eDNA-F is shown in SEQ ID NO.1, and the reverse primer COX2-eDNA-R is shown in SEQ ID NO.2.
[0006] The application method of fish eDNA meta-barcode PCGs primer sets includes the following steps: (1) Extract DNA from the sample to be tested; (2) Perform PCR amplification using the primer set described above; (3) Analyze the fragment size of the amplified product and perform fish identification analysis.
[0007] Preferably, the sample to be tested in step (1) is an environmental source sample and / or a biological source sample.
[0008] Preferably, the molar ratio of the forward primer to the reverse primer in the primer set in step (2) is 1:1.
[0009] Preferably, the annealing temperature in the PCR amplification reaction conditions in step (2) is 48~56℃.
[0010] Preferably, the annealing temperature in the PCR amplification reaction conditions in step (2) is 50°C.
[0011] Preferably, the fish identification analysis in step (3) includes high-throughput sequencing and bioinformatics analysis of the amplified products.
[0012] The present invention relates to the application of the fish eDNA macrobarcode PCG primer set in fish species identification or fish diversity analysis.
[0013] The present invention relates to the application of the fish eDNA macrobarcode PCG primer set in the evolutionary or phylogenetic analysis of fish.
[0014] The present invention relates to the application of the fish eDNA metabarcode PCG primer set in the monitoring of endangered fish species.
[0015] The present invention relates to the application of the fish eDNA metabarcode PCG primer set in the monitoring of invasive fish species.
[0016] Compared with the prior art, the beneficial effects of the present invention are: (1) This invention breaks through the limitations of traditional fish eDNA macrobarcoding technology, which relies on four classic gene regions: mitochondrial 12S rRNAs, 16S rRNAs, COI, and Cytb. It selects other protein-coding genes (PCGs) to design primer sets, and the provided PCG primer sets are different from any existing known eDNA macrobarcoding primer sets. It can better balance the needs of universality and species identification ability. Its universality is significantly better than the known existing commonly used PCG primer sets, and its species identification rate is significantly higher than the known existing commonly used rRNA primer sets. Moreover, the PCG primer sets have excellent amplification ability for both environmental and biological samples of fish. This lays a solid foundation for fish diversity surveys based on eDNA macrobarcoding technology and is of great significance for studying the species classification, phylogeny, phylogenetics, and conservation of endangered fish communities. (2) The present invention provides the optimal PCR reaction conditions for the primer set. The PCR reaction conditions are simple to operate and are very conducive to obtaining the best results. Attached Figure Description
[0017] Figure 1 Gel images showing PCR amplification of COX2-eDNA in 96 fish species; Figure 2 Gel images showing PCR amplification of MiFish-U / E on 96 fish species; Figure 3 The image shows a gel image of PS1 PCR amplification of 96 fish species. Figure 4 The amplification effect of three primer sets on aquatic environmental samples; Figure 5 The detection efficacy of three primer sets for fish species in aquatic environmental samples was studied. Detailed Implementation
[0018] To enable those skilled in the art to better understand the technical content of this invention, the technical solution of this invention will be further described in detail below with reference to specific embodiments. In the following embodiments, various processes and methods not described in detail are conventional methods known in the art.
[0019] Example 1 The universality comparison of PCG primer sets to fish tissue DNA included the following steps: 1. Design primers A novel set of PCG primers, named COX2-eDNA, was designed based on the complete mitochondrial genome sequences of 951 species of fish belonging to 2 classes, 56 orders, 260 families, and 656 genera. This primer set is located on the COX2 gene, and the amplified fragment length is 274 bp (without primers) [or 329 bp (with primers)]. The primer set was synthesized by Guangzhou Aiji Biotechnology Co., Ltd., and the forward and reverse primer sequences are shown below: COX2-eDNA-F: 5' ATTAAAGCCATAGGACACCAATGRTAYTG 3' (SEQ ID NO.1); COX2-eDNA-R: 5' AAGCTGTGGTTTGCTCCGCARATYTC 3' (SEQ ID NO.2); In the above forward and reverse primers, R and Y are degenerate nucleotide bases, where R represents A or G and Y represents C or T. The primers are prepared by mixing different base sequences corresponding to the degenerate sites in equal molar ratios.
[0020] 2. Selection of control primer set MiFish-U / E and PS1 were chosen as control primer sets because they are the most widely used rRNA and PCG primer sets in the field of fish eDNA macrobarcoding, respectively. MiFish-U / E is located on the 12S rRNA gene, with an amplified fragment length of approximately 168 bp (without primers) [or approximately 216 bp (with primers)]. PS1 is located on the COX1 gene, with an amplified fragment length of approximately 199 bp (without primers) [or 271 bp (with primers)]. These primer sets were also synthesized by Guangzhou Aiji Biotechnology Co., Ltd., and the specific forward and reverse primer sequences are shown below: MiFish-U / EF: 5' GTYGGTAAAWCTCGTGCCAGC 3' (SEQ ID NO.3); MiFish-U / ER: 5' CATAGTGGGGTATCTAATCCYAGTTTG 3' (SEQ ID NO.4); PS1-F: 5' ACCTGCCTGCCGTATTTGGYGCYTGRGCCGGRATAGT 3' (SEQ ID NO.5); PS1-R: 5' ACGCCACCGAGCCARAARCTYATRTTRTTYATTCG 3' (SEQ ID NO.6); In the above forward and reverse primers, Y, W, and R are degenerate nucleotide bases, where Y represents C or T, W represents A or T, and R represents A or G. The primers are prepared by mixing different base sequences corresponding to the degenerate sites in equal molar ratios.
[0021] 3. DNA extracted from 96 types of fish tissues Collecting sharp-nosed oblique-toothed sharks ( Scoliodon laticaudus ), eel ( Monopterus albus ), Yellowfin pufferfish ( Takifugu xanthopterus Ninety-six non-protected fish species were included, as shown in Table 1. One individual of each species was collected from various fish markets in Hainan Province. Genomic DNA was extracted from the genomic DNA of the 96 fish species using a Biosharp BL1043A instrument. Specific operating procedures were performed according to the instruction manual.
[0022] Table 1. 96 fish species used for primer set universality comparison.
[0023] 4. PCR amplification After extracting DNA from 96 fish tissues, PCR amplification was performed using the 2×Fast Taq plus Master Mix kit. The total volume of the PCR reaction system was 25 μL, consisting of 12.5 μL of 2×Fast Taq plus Master Mix, 9.5 μL of ultrapure water, 1 μL of forward primer, 1 μL of reverse primer, and 1 μL of fish tissue DNA template.
[0024] The optimal PCR reaction conditions for COX2-eDNA were: pre-denaturation at 95℃ for 5 min; denaturation at 95℃ for 30 s, annealing at 50℃ for 30 s, extension at 72℃ for 30 s, for a total of 35 cycles; and a final extension at 72℃ for 5 min. After the PCR products cooled to room temperature, they were subjected to gel electrophoresis with 1.0% agarose. The electrophoresis results were photographed and saved using a gel imaging system.
[0025] Referring to Miya et al. (2015) and Balasingham et al. (2018), the optimal annealing temperatures for MiFish-U / E and PS1 were set to 60℃ and 52℃, respectively, while the remaining reaction conditions and gel electrophoresis procedures were the same as those for COX2-eDNA.
[0026] Electrophoresis results are shown Figure 1 (COX2-eDNA) Figure 2 (MiFish-U / E) and Figure 3 (PS1). The well in the center of each image is a DNA marker. Figure 1 The PCR amplification product is approximately 329 bp in length, which translates to a molecular weight of approximately 217,140 g / mol. Figure 2 The PCR amplification product is approximately 216 bp in length, which translates to a molecular weight of approximately 142,560 g / mol. Figure 3 The PCR amplification product is approximately 271 bp in length, which translates to a molecular weight of approximately 178,860 g / mol.
[0027] Depend on Figures 1-3 It can be seen that the PCR amplification success rate of COX2-eDNA and MiFish-U / E for 96 fish species was 100%. COX2-eDNA showed highly consistent band position and brightness across all samples, stable amplification efficiency, and produced only a single band without any extraneous bands, tails, or diffusion, indicating precise primer binding. MiFish-U / E showed some samples with darker or missing bands, uneven amplification efficiency, multiple extraneous bands, diffuse bands, or tails, indicating non-specific amplification. PS1, however, failed to amplify in 6 fish species [specifically, the sharp-nosed scuttler shark (…]]. Scoliodon laticaudus ), striped bamboo shark ( Chiloscyllium plagiosum ), sharp-beaked ray ( Dasyatis zugei ), stingray ( Dasyatis akajei ), Tang's Fan Ray ( Platyrhina tangi ) and Shaolin ( Uranoscopus oligolepis It can be seen that the COX2-eDNA primer set has a good amplification ability for fish, and the universality of COX2-eDNA is higher than that of PS1. PS1 has a poor amplification effect on some fish (especially cartilaginous fish).
[0028] Example 2 The comparison of the amplification effects of primer sets on aquatic environmental samples includes the following steps: 1. Water environment sample collection Water samples were collected from 12 different aquariums (each with a volume of 25L) at the Atlantis Sanya Aquarium, including the "Lost Chambers," "Cartilaginous Fish Tank," and "Butterfly Fish Tank" aquariums. Each aquarium housed a variety of ornamental fish. Immediately after collection, the samples were returned to the laboratory for filtration at low temperatures (0–4℃). Three parallel samples were filtered from each plastic bucket (corresponding to three primer sets: COX2-eDNA, MiFish-U / F, and PS1), with a filtration volume of 5L per parallel sample. The filter membranes used were made of cellulose acetate with a pore size of 0.45μm and a diameter of 50mm.
[0029] 2. eDNA extraction Immediately after filtration, eDNA extraction was performed using the DNeasy Blood & Tissue Kit on all 36 filter membranes. To obtain the highest eDNA yield, the amount of reagents used during eDNA extraction was doubled, and the remaining steps were performed according to the instructions.
[0030] 3. PCR amplification After extracting the eDNA, PCR amplification and gel electrophoresis were performed according to the method in Example 1.
[0031] The gel electrophoresis results are shown in Figure 4 .Depend on Figure 4 It can be seen that the PCR success rate of all three primer sets for aquatic environmental samples was 100%, but the band brightness and width of PS1 were not as good as those of COX2-eDNA and MiFish-U / E. Therefore, the primer set provided by this invention has a higher amplification efficiency for aquatic environmental samples than PS1.
[0032] Example 3 The comparison of primer set detection efficacy for fish species in aquatic environmental samples included the following steps: 1. Water environment sample collection Refer to the method in Example 2, and use 4 plastic buckets to collect water environment samples from the drainage outlets of Yacheng Central Fishing Port Market, Lizhigou Seafood Market, Xinhonggang Seafood Market, and Sanya First Seafood Market respectively. Also refer to the method in Example 2 to filter the water environment samples.
[0033] 2. eDNA extraction Refer to the method in Example 2 for eDNA extraction.
[0034] 3. PCR amplification After extracting eDNA, perform PCR amplification and gel electrophoresis according to the method in Example 1.
[0035] 4. High-throughput sequencing and bioinformatics analysis Entrust Guangzhou Aiji Biotechnology Co., Ltd. to conduct quality inspection on the extracted eDNA samples. Then, use a commercial Illumina library construction kit to complete library construction for the amplification products of the three sets of primers. Subsequently, use an Illumina HiSeq 2000 sequencer for 2×250bp paired-end (PE250) on-machine sequencing. Use a high-fidelity enzyme (Q5® High-Fidelity Master Mix, New England Biolabs) for PCR amplification to ensure amplification efficiency and accuracy. Divide the original sequences that pass the initial quality screening into libraries, use the solexaQA software to perform quality control on the raw data, remove sequences that may be incorrect (<Q30), and at the same time remove low-quality sequences, adapter-contaminated sequences, sequences containing N, and low-complexity sequences to obtain Clean Data. Use the FLASH software to assemble and splice the paired-end sequencing data. Use the UCHIME software to detect chimeric sequences in the high-quality sequences, and finally remove the chimeric sequences to obtain valid sequences. Use the TagCleaner software to remove the primer sequences contained in the spliced sequences. Use the CROP software to perform set analysis on the same sequences in the spliced sequences to obtain initial operational taxonomic units (OTUs); use the BLAST program in the MitoFish database for species annotation, and the annotation conditions are set as follows: sequence similarity (Identity) is selected as ≥99%, and the expected value (E-value) is selected as ≤10 5 .
[0036] The results of bioinformatics analysis are shown in Figure 5Statistical results show that compared with the MiFish-U / E and PS1 primer sets, the COX2-eDNA primer set exhibits significant advantages in all detection indicators: COX2-eDNA yielded 458,302.75 fish sequences, 212.5 fish OTUs, and 95.50 fish species, respectively; MiFish-U / F yielded 506,491.5 sequences, 166.5 OTUs, and 95.25 species, respectively; and PS1 yielded 241,635.75 sequences, 151.75 sequences, and 70.75 species, respectively. It is evident that COX2-eDNA outperforms PS1 in all three detection indicators. Regarding the number of sequences, COX2-eDNA obtained 458,302.75 valid fish sequences, slightly fewer than MiFish-U / E's 506,491.5, but significantly more than PS1's 241,635.75, sufficient to meet the sequencing depth requirements for subsequent species annotation. Although the number of sequences in COX2-eDNA is slightly lower than that in MiFish-U / F, COX2-eDNA detected 212.5 fish OTUs, significantly more than the 166.5 in MiFish-U / E. Meanwhile, the number of fish species (95.5) is roughly the same as that in MiFish-U / E (95.25). This indicates that COX2-eDNA utilizes sequence data more efficiently, achieving more refined sequence variation analysis and comparable species detection capabilities at lower sequencing depths, resulting in superior overall performance.
[0037] Example 4 The species recognition capability of primer sets corresponding to amplified fragments includes the following steps: 1. Sequence Download The complete mitochondrial genome sequences of 935 species of fish belonging to 2 classes, 20 orders, 52 families, 106 genera were downloaded from the NCBI database. The specific species list, their order of origin, and corresponding sequence numbers are shown in Table 2.
[0038] 2. Sequence Processing Based on the forward and reverse primer sequences of three primer sets (COX2-eDNA, MiFish-U / E, and PS1), amplification regions were located in the whole mitochondrial genomes of 935 fish species, and the target amplification fragment sequences corresponding to each species were extracted. The SeqMan program in the DNAstar software package was used to truncate and align the sequences, with the default parameters.
[0039] 3. Phylogenetic tree construction Using MEGA 7.0 software, a phylogenetic tree was constructed based on the Neighbor-Joining method; evolutionary distance was calculated using the Kimura 2-parameter (K2P) model, and gaps and missing sites were determined using the pairwise deletion method; the reliability of all data was tested using the Bootstrap test, and the support rate of the branch tree nodes was obtained after 1000 repeated sampling tests.
[0040] 4. Comparison of species identification rates The species identification rate of 935 fish species was statistically analyzed. If a species formed an independent monophyletic branch on the phylogenetic tree and had no clustering overlap with other species (interspecific genetic distance > 0), it was considered correctly identified. If two or more different species clustered in the same branch (interspecific genetic distance = 0 or indistinguishable), the species identification was considered a failure. Finally, the species identification rates of the three sets of primers were statistically analyzed and compared.
[0041] Species identification results are shown in Table 2. The number of misidentified species for COX2-eDNA, MiFish-U / E, and PS1 were 168, 378, and 174, respectively, with corresponding correct identification rates of 82.03%, 59.57%, and 81.39%. The results indicate that COX2-eDNA had the highest species identification rate, followed by PS1, while MiFish-U / E had the lowest.
[0042] Table 2. Mitochondrial genome sequences and species identification of 935 fish species.
[0043] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. Applications of a set of fish eDNA metabarcode PCG primers in the following areas: (1) Fish species identification; (2) Fish diversity analysis; (3) Evolutionary or phylogenetic analysis of fish; (4) Monitoring of endangered fish species; (5) Monitoring of invasive fish species; The primer set is named COX2-eDNA, with the forward primer COX2-eDNA-F as shown in SEQ ID NO.1 and the reverse primer COX2-eDNA-R as shown in SEQ ID NO.2; In the forward and reverse primers, R and Y are degenerate nucleotide bases, where R represents A or G and Y represents C or T. Both the forward and reverse primers are prepared by mixing different base sequences corresponding to the degenerate sites in equal molar ratios.
2. The application according to claim 1, characterized in that, Includes the following steps: (1) Extract DNA from the sample to be tested; (2) Perform PCR amplification using the primer set described in claim 1; (3) Analyze the fragment size of the amplified product and perform fish identification analysis.
3. The application according to claim 2, characterized in that, The sample to be tested in step (1) is an environmental source sample and / or a biological source sample.
4. The application according to claim 2, characterized in that, In step (2), the molar ratio of the forward primer to the reverse primer in the primer set is 1:
1.
5. The application according to claim 2, characterized in that, In step (2), the annealing temperature for the PCR amplification reaction is 48~56℃.
6. The application according to claim 2, characterized in that, The fish identification analysis described in step (3) includes high-throughput sequencing and bioinformatics analysis of the amplified products.