A fish DNA barcode database and a method for constructing the same based on a complete mitochondrial genome sequence and application thereof

By constructing a fish DNA barcode database based on the mitochondrial whole genome, the problem of inaccurate species comparison in existing databases has been solved, achieving high-precision species identification and a simplified monitoring process, which is applicable to fish resource surveys in the field of ecological and environmental protection.

CN119132424BActive Publication Date: 2026-03-17FRESHWATER FISHERIES RES CENT OF CHINESE ACAD OF FISHERY SCI +2
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411156823.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2026-03-17
Estimated Expiration
2044-08-22

AI Technical Summary

Technical Problem

The existing public databases for fish are not complete enough, and the data sources are diverse and of varying quality, resulting in the loss of small population species comparisons, incorrect comparison results, and the inability to effectively exclude species from non-target areas, making it difficult to accurately identify closely related species.

Method used

We constructed a fish DNA barcode database based on the mitochondrial whole genome. By defining the fish species list, collecting and identifying samples, extracting DNA, sequencing and analyzing the mitochondrial whole genome, and constructing the DNA barcode library, we simplified the PCR amplification process and improved the accuracy of species identification.

Benefits of technology

It improved the accuracy of species comparison results, simplified the technical process, effectively eliminated false positive results, and enhanced the ability to identify closely related species.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119132424B_ABST
    Figure CN119132424B_ABST
Patent Text Reader

Abstract

This invention relates to the field of molecular biology and discloses a fish DNA barcode database and its construction method and application based on the mitochondrial genome whole sequence. The method includes: determining the fish list, fish collection and classification identification, DNA extraction, mitochondrial whole genome sequencing and analysis, DNA barcode library construction, and database application. Based on research foundations and historical data of fish in the target water area, this invention determines the fish list for the target water area and constructs a fish DNA barcode database using mitochondrial whole genome sequencing technology, providing precise database support for fish surveys and monitoring using environmental DNA technology. It solves problems such as insufficient completeness of existing public fish databases, diverse and inconsistent data sources and quality, insufficient regional specificity, potential loss of small population species in comparisons, erroneous comparison results, and the inability to effectively exclude false positive results such as species from non-target areas, as well as the difficulty in accurately identifying closely related species with high genetic similarity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of molecular biology, specifically to a fish DNA barcode database and its construction method and application based on the full mitochondrial genome sequence. Background Technology

[0002] Fish are an important component of aquatic ecosystems, and their species composition and genetic diversity are crucial indicators of the health of these ecosystems. Water pollution, habitat fragmentation, and invasive species lead to fish resource decline and community imbalance. Conducting fish resource surveys and monitoring can help us understand the species composition and diversity of fish in target waters, as well as the population status of rare and endangered fish species, thus supporting species conservation, resource protection, and habitat restoration.

[0003] Currently, fish resource surveys and monitoring primarily rely on net fishing for fish samples. This process requires significant manpower, causes some damage to fish resources, and disturbs fish habitats. Analysis of survey and monitoring results shows that net fishing is not ideal for catching fish with small populations, small individuals, and high concealment capabilities, and it also demands a high level of classification and identification skills from the survey personnel. Environmental DNA technology, with its advantages of non-destructive sampling, large-scale continuous sampling, high precision, and low cost, is gradually being widely applied in fish resource surveys and monitoring.

[0004] DNA barcoding is a core component of environmental DNA technology. A DNA barcode is a short, relatively conserved, and easily amplified DNA fragment that corresponds to a species, and can be used for species identification. The mitochondrial genome, characterized by rapid evolution, maternal inheritance, and lack of recombination, contains a high copy number and many conserved sequences that can distinguish species, making it excellent material for DNA barcoding. Commonly used mitochondrial genes include Cytb, COI, 12S, 16S, and D-loop, with fragment sizes ranging from 80-200 bp. Common primer designs for fish typically select mitochondrial gene fragments such as Cytb, 12S, and COI, with the 12S gene being the most widely used. When using primers for PCR amplification of gene sequences, errors can occur due to the formation of chimeras, heteroduplexes, and Taq DNA polymerase errors, so the amplification efficiency is largely affected by primer efficiency. Comparing the amplified sequence with the DNA barcode and determining the species' taxonomic position based on the sequence overlap rate can sometimes lead to difficulties in distinguishing closely related species with relatively short divergence periods. Mitochondrial whole-genome sequencing does not involve PCR, which can effectively avoid errors generated during amplification. At the same time, the sequences obtained by whole-genome sequencing are complete, and using them for sequence alignment can make up for the inability of local fragment alignment to distinguish closely related species.

[0005] The construction of barcode databases is fundamental to the implementation of environmental DNA technology. The National Center for Biotechnology (NCBI) in the United States, in conjunction with the China Bank for Biological Optics (CBOL), established the first database system, GenBank, in 1982. Later, in 2004, the Consortium for the Barcode of Life (CBOL) established the Barcode of Life Data Systems (BOLD), and in 2007, Canada established the DNA Barcode Identification Center. DNA barcoding technology compares target sequences used for species identification with sequences in a database to determine species information. It requires a high degree of database richness and accuracy. If a species' DNA sequence is not included in the database, the amount of data included is insufficient, or there is erroneous data, comparison failures and incorrect results will occur. Currently, the comprehensiveness of species and the accuracy of sequences in public databases cannot meet the needs of specific research. There are instances where gene sequences of species endemic to target waters are not included, and the accuracy of included sequences needs to be verified. Therefore, building accurate and traceable self-constructed databases based on research needs has become a trend.

[0006] The length of fish mitochondrial genomes is mostly around 15-20 kb, which is much longer than that of common DNA barcodes. Therefore, if the entire mitochondrial genome sequence of fish can be used as a fish DNA barcode, a wider range of sequence differences can be identified, greatly improving the accuracy of species identification. The method of constructing a fish DNA barcode database using the entire mitochondrial genome sequence can effectively improve the accuracy of environmental DNA species comparison results, simplify the technical steps, and is more conducive to the development of environmental DNA monitoring technology. Summary of the Invention

[0007] Technical problems addressed: Existing public fish databases suffer from insufficient completeness, diverse and inconsistent data sources, and inadequate regional targeting, potentially leading to the loss of small population species in comparisons, erroneous comparison results, and the inability to effectively exclude false positives such as species from non-target regions. Furthermore, they struggle to accurately identify closely related species with high genetic similarity. This invention provides a fish DNA barcode database and its construction method and application based on the mitochondrial genome whole sequence. The fish barcode database constructed using the mitochondrial genome is simple to operate and highly feasible.

[0008] Technical solution: A method for constructing a fish DNA barcode database based on the complete mitochondrial genome sequence, comprising the following steps:

[0009] (1) Determine the list of fish species;

[0010] (2) Collect and identify samples according to the fish list, and preserve the specimens;

[0011] (3) Extract DNA from the specimen;

[0012] (4) Perform mitochondrial whole genome sequencing and analysis on the extracted DNA samples;

[0013] (5) Construct a DNA barcode library.

[0014] Preferably, in step (1), the fish list is determined based on the survey data of the target waters and historical fish data.

[0015] As a preferred option, in step (2), samples are collected and identified according to the fish catalog and specimens are preserved, as follows: samples are collected according to the fish catalog, species are identified according to taxonomic criteria, morphologically intact individuals of each species are selected and photographed to capture their morphological characteristics, and muscle samples are collected and preserved as specimens.

[0016] Furthermore, priority is given to collecting individual fish from the target waters. If the number collected from the target waters is insufficient, individuals from neighboring waters are used to supplement the collection.

[0017] Preferably, in step (3), DNA is extracted from the specimen using SDS extraction combined with a purification column method.

[0018] As a preferred method, the extraction method is as follows: Dorsal muscle was taken from 3 individuals of each fish species, and total DNA was extracted by SDS (sodium dodecyl sulfate) extraction combined with purification column method. 500 mg of muscle was taken from each sample, and lysed with SDS lysis buffer. An appropriate amount of proteinase K and mercaptoethanol were added. The lysis process was gently inverted and mixed. After cooling to room temperature, the mixture was centrifuged. The supernatant was taken and extracted with chloroform / isoamyl alcohol (24:1) twice. DNA was precipitated with isopropanol. After gentle inversion and mixing, the mixture was centrifuged and the waste liquid was discarded. The precipitate was washed with 75% ethanol and repeated twice. After the DNA precipitate was dried, it was dissolved in 200 μL TE (Tris-EDTA), digested with RNase, and purified using a purification column (OMEGA). After purification, it was purified using the Ampure XPbeads kit. After purification, Nanodrop, Qbuit and electrophoresis quality checks were performed.

[0019] Preferably, in step (4), the extracted DNA sample undergoes mitochondrial whole-genome sequencing and analysis, specifically as follows: The sample genomic DNA is fragmented into 300-500 bp fragments using a Covaris ultrasonic disruptor; the DNA fragment ends are repaired, and an A base and sequencing adapter are added to the 3′ end; the ligation product is amplified by PCR, the product is recovered and purified using magnetic beads, and a library is constructed; after library construction, preliminary quantification is performed using Qubit 3.0, the library is diluted, and then the insert fragments in the library are detected using an Agilent 2100. Once the insert fragment size meets expectations, the effective concentration of the library is accurately quantified using Q-PCR; after the DNA library passes quality control, paired-end sequencing is performed using the MGISEQ-T7 high-throughput sequencing platform; the sequencing data is quality-controlled and filtered using UPARSE software (http: / / drive5.com / uparse / , version...). 7.1) The sequences were clustered into OTUs based on 97% similarity and chimeras were removed. After excluding erroneous sequences, the statistical yield, base content distribution, and data quality distribution were analyzed. The mitochondrial genome was then assembled and annotated.

[0020] Preferably, the DNA barcode entries in step (5) include at least one of the following information: Chinese name of the species, Latin name, taxonomic status, mitochondrial genome sequence, morphological image, specimen, muscle sample, DNA sample, latitude and longitude of the sampling point, and sampling time.

[0021] A fish DNA barcode database constructed using the method described above.

[0022] Application of the aforementioned fish DNA barcode database in determining the taxonomic rank and relative abundance of fish species in target waters.

[0023] As a preferred method, the steps are as follows: collect water samples from the target water area, enrich the DNA in the water samples using a microporous membrane and a vacuum filtration device to obtain environmental DNA samples, perform full-sequence short-read sequencing on the experimental samples, compare the sequencing results using BWA software based on the aforementioned DNA barcode database, directly count the reads of specific alignments, classify the reads of multiple alignments according to the taxonomic units of the aligned sequences, obtain species annotations through the corresponding taxonomic information in the database, and determine the taxonomic rank and relative abundance of fish species in the target water area.

[0024] Beneficial effects:

[0025] (1) The database provided by this invention contains a list of fish species based on research on the target waters and historical data collection, which effectively eliminates false positives in the comparison results and improves the reliability of the comparison results.

[0026] (2) This invention uses whole mitochondrial genome sequence alignment to make up for the inability of local fragment alignment to distinguish closely related species.

[0027] (3) This invention avoids the PCR amplification step required by conventional methods, shortens the technical process, simplifies the sequencing steps, and avoids the situation where species cannot be detected due to primer mismatch, etc. Attached Figure Description

[0028] Figure 1 This is a diagram showing the electrophoresis results of fish mitochondrial genomes in Example 1 of this patent.

[0029] Figure 2 This is a photograph of a fish identified as an anchovy in one of the morphological species of Embodiment 1 of the present invention;

[0030] Figure 3 This is the entry information for one type of anchovy in the database of Embodiment 1 of the present invention;

[0031] Figure 4 This is a diagram of the sequence alignment process in Embodiment 2 of the present invention. Detailed Implementation

[0032] This invention addresses the practical needs of fish resource survey and monitoring in the field of ecological environmental protection. It utilizes molecular biology methods to construct a fish DNA barcode database. The invention is further described below with reference to the accompanying drawings and specific embodiments. The embodiments described below are only a portion of the embodiments of this invention. Therefore, all other embodiments obtained by those skilled in the art without inventive effort are within the protection scope of this invention.

[0033] Unless otherwise specified, the raw materials and instruments used in the embodiments of this invention are all commercially available products.

[0034] Specific implementation examples are as follows:

[0035] Example 1

[0036] Establish a barcode database for fish species in the Jiangsu section of the Yangtze River.

[0037] (1) Determine the list of fish species: Collect historical data on fish species in the target waters (in this embodiment, refer to the "Fishes of the Yangtze River", "Fishes of Jiangsu", and the historical survey results of the lower reaches of the Yangtze River from 2017 to the sampling time) to determine the list of fish species in the lower reaches of the Yangtze River.

[0038] (2) Fish Collection and Classification: Samples were collected according to the fish catalog, prioritizing individuals from the Jiangsu section of the Yangtze River. If the number collected from the Jiangsu section was insufficient, individuals from nearby waters were used as supplements. Species were identified based on taxonomic criteria. For each species, three morphologically intact individuals were selected, photographed for their morphological characteristics, and preserved as specimens. This embodiment uses *Anchovy* as an example; the photographs of its morphological characteristics can be found in [reference needed]. Figure 2 .

[0039] (3) DNA Extraction: Dorsal muscle was collected from 3 individuals of each fish species. DNA was extracted using SDS (sodium dodecyl sulfate) combined with a purification column method. 500 mg of muscle was taken from each sample and lysed with 700 μL SDS lysis buffer. 10 μL proteinase K and 10 μL mercaptoethanol were added, and the mixture was gently inverted and mixed. After cooling to room temperature, the mixture was centrifuged at 12,000 rpm for 5 min. The supernatant was collected and extracted twice with chloroform / isoamyl alcohol (24:1). DNA was precipitated with isopropanol, gently inverted and mixed, centrifuged, and the waste liquid was discarded. The precipitate was washed with 800 μL 75% ethanol, and this process was repeated twice. After the DNA precipitate was dried, it was dissolved in 200 μL TE (Tris-EDTA), digested with 30 μL RNase, and purified using an OMEGA purification column. After purification, the DNA was purified using an Ampure XPbeads kit. Nanodrop, Qbuit, and electrophoresis were then performed for quality control (see [link to product details]). Figure 1 ).

[0040] (4) Mitochondrial whole-genome sequencing and analysis: Mitochondrial whole-genome sequencing was performed using extracted DNA samples. The genomic DNA was fragmented into 300-500 bp fragments using a Covaris ultrasonic disruptor. DNA fragment ends were repaired, and an A base and sequencing adapter were added to the 3′ end. The ligation products were amplified by PCR. The reaction mixture consisted of 20 μl product, 25 μl VAHTSHiFiAmplification Mix, and 5 μl PCR Primers. The reaction program was 95℃ for 3 min, 98℃ for 20 sec, 60℃ for 15 sec, and 72℃ for 30 sec, repeated 6 times, followed by 72℃ for 5 min, and storage at 4℃. The product was recovered and purified using magnetic beads, and a library was constructed. After library construction, preliminary quantification was performed using Qubit 3.0 to dilute the library. Subsequently, the insert fragments in the library were detected using an Agilent 2100. The effective concentration of the library was accurately quantified using Q-PCR for insert fragments of 400-600 bp. After the DNA library passed quality control, paired-end sequencing was performed using the MGISEQ-T7 high-throughput sequencing platform. The sequencing data underwent quality control filtering using UPARSE software (http: / / drive5.com / uparse / , version 7.1), with OTU clustering based on 97% similarity and chimeras removed. After excluding erroneous sequences, the statistical yield, base content distribution, and data quality distribution were analyzed. The mitochondrial genome was then assembled and annotated.

[0041] (5) DNA Barcode Database Construction: Information from three individuals collected for each species was compiled and summarized to establish a standardized species barcode database entry format. This format includes the species' Chinese name, Latin name, taxonomic classification, mitochondrial genome sequence, morphological images, specimen, muscle sample, DNA sample, sampling point latitude and longitude, and sampling time. A barcode database for fish in the Jiangsu section of the Yangtze River was then established. This embodiment uses *Anchovy* as an example; its entry information in the database can be found in [reference needed]. Figure 3 Due to space limitations, the DNA barcode in the image is specifically the sequence spliced ​​from SEQ ID NO: 1 and SEQ ID NO: 2. SEQ ID NO: 1 and SEQ ID NO: 2 are listed in the sequence listing.

[0042] Example 2

[0043] Species identification of environmental DNA samples using the fish barcode database of the Jiangsu section of the Yangtze River.

[0044] (1) Obtaining environmental DNA samples: Based on the principles of uniformity and representativeness, sampling sections were set up every 20 km along the Jiangsu section of the Yangtze River, for a total of 18 sampling sections. Surface water at a depth of approximately 0.5 m was collected using a sterilized water sampler and placed in 2L sampling bottles for storage in the dark and at low temperature. Six water samples were collected uniformly from the north bank to the south bank of each section. The water samples were filtered using a vacuum filtration pump and a 0.45 μm mixed cellulose microporous membrane to enrich the environmental DNA in the water samples. After filtration, the filter membrane was placed in a 5 mL sample tube containing anhydrous ethanol and stored at -20℃.

[0045] (2) DNA extraction: Extract DNA from the filter membrane. The extraction steps are the same as (3) in Example 1 (replace muscle with filter membrane DNA in this example).

[0046] (3) Mitochondrial whole genome sequencing and analysis: The extracted environmental DNA samples were subjected to mitochondrial whole sequence short read sequencing. The sequencing steps were the same as in (4) of Example 1 to obtain genomic DNA.

[0047] (4) Library Construction: Libraries were constructed using the MGIEAsy Universal DNA Library Preparation Kit V1.0 (CAT#1000005250, MGI). The procedure involved randomly fragmenting 1 μg of genomic DNA using Covaris, selecting DNA fragments with an average size of 200-400 bp. The selected fragments were then end-repaired, 3'-alkenylated, and ligated with adapters. DNA samples were amplified by PCR and further purified. Double-stranded PCR products were subjected to thermal denaturation and circularization. Single-stranded circular DNA (ssCir-DNA) was formatted into the final library and identified by QC. Qualified libraries were sequenced using PE150 sequencing on the DNBSEQ-T7RS platform.

[0048] (5) Species comparison: See Figure 4 For short reads obtained from sequencing, BWA was used to align them back to the previously constructed "Yangtze River Jiangsu Section Fish Barcode Database" to screen for high-quality alignment results. Reads that specifically aligned to unique sequences were assigned to the species corresponding to those sequences; for reads with multiple alignments, the species corresponding to the aligned sequences were classified at the genus / family / order level. Finally, the abundance of each species in each sample was determined. Through normalization correction of the multi-sample alignment results, the abundance of each species in each sample in the database was obtained. Based on the above data, the samples were further clustered to characterize the similarity relationships between samples regarding the included species; and the species were clustered to characterize the distribution similarity relationships of different species in the above samples.

[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for constructing a fish DNA barcode database based on the complete sequence of the mitochondrial genome, characterized in that, The steps are as follows: (1) determining the fish list: based on the investigation basis of the target water area and the historical data of fish, the fish list is determined; (2) sample collection and identification according to the fish list, and sample preservation, as follows: sample collection according to the fish list, identification of species according to taxonomic basis, selection of morphologically perfect individuals for each species to take morphological characteristic photos, collection of muscle samples, and preparation of specimens for preservation; (3) DNA extraction from the specimens: DNA extraction from the specimens, using SDS extraction combined with purification column method, the extraction method being as follows: taking the back muscle of 3 individuals of each fish species, lysing using SDS lysis solution, adding proteinase K and mercaptoethanol, gently inverting and mixing during the lysis process, centrifuging after cooling at room temperature, taking the supernatant, adding chloroform / isopentanol for extraction, extracting twice, precipitating DNA with isopropyl alcohol, centrifuging after gently inverting and mixing, discarding the waste liquid, washing the precipitate with 75% ethanol, repeating twice, dissolving the DNA precipitate with TE after air-drying, RNAse digestion, using a purification column for purification, purifying using an Ampure XP beads kit after purification, and performing Nanodrop and Qbuit and electrophoresis quality inspection, and being qualified; (4) mitochondrial whole genome sequencing and analysis of the extracted DNA samples: using a Covaris ultrasonic crusher to break the sample genomic DNA into 300-500 bp fragments, repairing the ends of the DNA fragments, adding A bases and sequencing adapters to the 3' ends; PCR amplification of the ligation products, recovery of the products using magnetic beads and purification, library construction; after library construction, using Qubit 3.0 for preliminary quantification, diluting the library, then using Agilent 2100 to detect the insert size of the library, and using Q-PCR method to accurately quantify the effective concentration of the library; after DNA library quality inspection, using the MGISEQ-T7 high-throughput sequencing platform for double-end sequencing; sequencing data quality control and filtering, using UPARSE software to cluster OTUs according to 97% similarity and remove chimeras, after excluding error sequences, counting data yield, base content distribution, data quality distribution, assembling mitochondrial genomes and annotating; (5) DNA barcode library construction, the items of the DNA barcode including at least one of the following information: Chinese name of species, Latin name, taxonomic status, mitochondrial genome sequence, morphological picture, specimen, muscle sample, DNA sample, sampling point longitude and latitude, and sampling time.

2. A fish DNA barcode database constructed by the method of claim 1.

3. Use of the fish DNA barcode database of claim 2 in determining the taxonomic rank and relative abundance of fish species in a target water area.

4. Use according to claim 3, characterized in that, The steps are as follows: collecting a water body sample in a target water area, enriching DNA in the water body sample by using a microporous filter membrane and a vacuum filtration device to obtain an environmental DNA sample, performing whole sequence short read sequencing on the experimental sample, performing alignment on the sequencing results by using BWA software based on the DNA barcode database of claim 2, directly counting the specific alignment reads, classifying the multiple alignment reads according to the classification units of the alignment sequences, obtaining species annotation through the overall taxonomic information corresponding to the database, and determining the classification rank of fish species in the target water area and the relative abundance.

Citation Information

Patent Citations

  • Method for amplifying mitochondrial gene total sequences of fish in Yangtze River based on degenerate primer combination and application thereof

    CN107557478A