Probe compositions, kits and applications for identifying or assisting in the identification of mammalian species

By designing probe compositions and second-generation sequencing technology for mammalian COI genes, the problems of cumbersome operations and unstable results in batch identification of mammalian species are solved, and rapid and accurate species identification is achieved.

CN115349020BActive Publication Date: 2025-08-01BEIJING ZOO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202180007269.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-28
Publication Date
2025-08-01
Estimated Expiration
2041-05-28

AI Technical Summary

Technical Problem

The prior art has cumbersome operations and unstable results when batch identification of mammalian species, making it difficult to achieve rapid and accurate COI gene sequence identification.

Method used

A probe composition for mammalian COI genes was designed. Through sequence clustering and second-generation sequencing technology, the Angiosperms353 method was used to set the genetic distance to 0.05, the coverage depth was 2X, and a single-stranded DNA probe with a GC content of >30% was designed to cover 2 probes per SNP site to form a COI gene capture kit to achieve batch identification.

Benefits of technology

It realizes rapid and accurate batch identification of mammalian species, improves the accuracy and efficiency of detection, and avoids the problems of cumbersome operations and unstable results in the first generation of sequencing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GDA0003875100370000021
    Figure GDA0003875100370000021
  • Figure GDA0003875100370000031
    Figure GDA0003875100370000031
  • Figure GDA0003875100370000041
    Figure GDA0003875100370000041
Patent Text Reader

Abstract

Probe compositions and kits for identifying or assisting in the identification of mammalian species are provided. The probe composition is a nucleotide probe shown in SEQ ID NO: 1-SEQ ID NO: 3590 in the sequence listing. The COI gene of the target species is captured by the probe, and the DNA sequence of the COI gene is obtained by using the method of high-throughput detection of next-generation sequencing for library construction. It can take into account the advantages of flexibility and convenience of first-generation sequencing and can obtain results quickly, while solving the problem of cumbersome operation of batch samples in first-generation sequencing, and avoiding the risk of inaccurate COI gene sequences caused by unstable partial results in first-generation sequencing. Developing the probe into a mammalian COI gene capture kit can achieve rapid and batch identification of mammalian species.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a probe composition, a kit thereof and an application for identifying or assisting in identifying mammalian species in the field of biotechnology. Background Art

[0002] Cytochrome C oxidase I gene (COI gene) is one of the three cytochrome oxidase subunits encoded by mitochondrial genes, and it is the gene with the largest molecular weight and the most conserved functional structure among them. The COI gene has characteristics such as multiple variations, being easily amplified by universal primers, and few insertions and deletions in its sequence itself. Therefore, the COI gene is selected as a marker gene for DNA classification (DNA barcode). The length of this coding gene segment is generally about 658bp. In addition to being used for DNA classification, it can also be used for the study of the phylogenetic relationship and molecular evolution of species.

[0003] In the past, to obtain the COI gene sequence, the first-generation sequencing method was generally used. This method has the advantages of being flexible, convenient, and capable of quickly obtaining results. However, at the same time, there are also risks such as cumbersome operation for batch samples and unstable results in part of the first-generation sequencing, resulting in the inability to obtain accurate COI gene sequences.

[0004] Publication of Invention

[0005] One technical problem to be solved by the present invention is how to batch identify or assist in identifying mammalian species.

[0006] To solve the above technical problems, the present invention provides a probe composition for identifying or assisting in identifying mammalian species.

[0007] The probe composition for identifying or assisting in identifying mammalian species provided by the present invention is a probe composition obtained by performing sequence clustering on the mammalian COI gene sequence to obtain a representative sequence, and designing probes to cover each SNP site for the representative sequence.

[0008] For the above probe composition, the sequence clustering is performed using the Angiosperms353 method, the genetic distance is set to 0.05 (i.e., the sequence similarity is 95%), and the coverage depth is 2X.

[0009] In the above probe composition, the probe design is carried out according to the standard that the GC content > 30%.

[0010] In the above probe composition, for each SNP site to be covered by probes, two probes are designed to cover each SNP site.

[0011] The probe composition for identifying or assisting in the identification of mammalian species provided by the present invention is specifically a combination consisting of 3,590 single-stranded DNAs shown in Sequence 1 - Sequence 3,590 of the sequence listing.

[0012] The present invention also provides a method for identifying or assisting in the identification of mammalian species using the above probe composition: including capturing the COI gene of the mammalian to be tested with the above probe composition, constructing a library, obtaining the DNA sequence of the COI gene through high-throughput detection by next-generation sequencing, and determining the mammalian species to be tested according to the obtained COI gene sequence.

[0013] In the above method, determining the mammalian species to be tested according to the obtained COI gene sequence is to compare it with the COI gene of known species, for example, comparing it with the COI gene in the mitochondrial whole gene data.

[0014] In the above method, mammalian species can be identified in batches.

[0015] The present invention also protects a reagent or kit for identifying or assisting in the identification of mammalian species, and the reagent or kit includes the above probe composition.

[0016] The application of the above probe composition, the above method and / or the above reagent or kit in identifying or assisting in the identification of mammalian species, or the application in preparing products for identifying or assisting in the identification of mammalian species belongs to the protection scope of the present invention.

[0017] The present invention applies the targeted sequencing genotyping technology, designs and synthesizes liquid-phase probes according to the selected sequences evaluated and analyzed, and tests the capture efficiency of the probes, and finally forms a COI gene capture kit to achieve the purpose of batch identification of mammals. Brief Description of the Drawings

[0018] Figure 1 It is a flow chart for the development of the probe combination in Example 1 of the present invention. Detailed Embodiments

[0019] The present invention will be further described in detail below in conjunction with the specific embodiments. The given embodiments are only for clarifying the present invention, rather than limiting the scope of the present invention. The following provided embodiments can be used as a guide for those of ordinary skill in the art to make further improvements, and do not constitute any limitation to the present invention in any way.

[0020] Unless otherwise specified, the experimental methods in the following embodiments are all conventional methods. Unless otherwise specified, the materials, reagents, etc. used in the following embodiments can all be obtained from commercial channels.

[0021] Example 1

[0022] Based on the polymorphism of the COI sequences among different species, the inventors designed capture probes to capture the COI genes of the target species, and then constructed a library to obtain the DNA sequences of the COI genes by means of high-throughput detection using next-generation sequencing. This solution can take into account the advantages of first-generation sequencing, such as flexibility, convenience, and the ability to obtain results quickly, while solving the problem of cumbersome operation of batch samples in first-generation sequencing, and avoiding the risk of inaccurate COI gene sequences caused by unstable partial results in first-generation sequencing. The capture probes were developed into a mammalian COI gene capture kit for DNA barcoding research on a large number of mammalian species, enabling the identification of mammalian species.

[0023] The specific development process is carried out according to Figure 1 the flowchart in

[0024] 1. Determination of technical route and solution

[0025] The comparison between two next-generation sequencing technologies, GenoBaits and GenoPlexs, is shown in Table 1:

[0026] Table 1 Comparison of GenoBaits and GenoPlexs technologies

[0027]

[0028]

[0029] The inventors applied the targeted sequencing genotyping (Genotyping By Target Sequencing, GBTS, also known as GenoBaits) technology, designed and synthesized liquid-phase probes based on the sequences selected through evaluation and analysis, and tested the capture efficiency of the probes. Finally, a COI gene capture kit was formed to achieve the purpose of batch identification of mammals.

[0030] Based on the COI gene sequences of more than 300 mammalian species and some COI genes extracted from mitochondrial whole-genome data, a total of 413 COI genes were subjected to sequence clustering. The purpose was to find the most representative COI sequence in each classification, and probe design was carried out based on the finally clustered sequences. Since different parameters have different effects during design, two sets of evaluation schemes are given as follows:

[0031] Scheme 1: Use the Angiosperms353 method for sequence clustering, set the genetic distance to 0.1 (i.e., the sequence similarity is 90%), and the coverage depth is 1X; perform clustering with these parameters. The final number of representative sequences obtained is 378, and the development and detection costs are relatively low. The risk is that when the target sequence mutates again (i.e., there are differences between the sequences captured during detection and the provided sequences), it may be impossible to capture the corresponding sequence, thus affecting the species identification result.

[0032] Scheme 2: Use the Angiosperms353 method for sequence clustering, set the genetic distance to 0.05 (i.e., the sequence similarity is 95%), and the coverage depth is 2X; perform clustering with these parameters. The final number of representative sequences is relatively large, which is 479. The development and detection costs are relatively high, but the probability of capturing the target sequence will increase significantly, and the accuracy of the final identification result is higher.

[0033] Finally, it was confirmed to develop the kit using Scheme 2 and perform probe design.

[0034] 2. Probe design and selection principles

[0035] 2.1 Probe design principle:

[0036] Use the GenoBaits Probe Designer software to design liquid-phase capture probes for the 479 representative sequences obtained in Scheme 2. The probe length is set to 110bp, the GC content > 30%, and each SNP site is covered by 2 capture probes.

[0037] This is the end of the translation. If you have any other questions, please feel free to let me know.

[0038] 1) Select probes with a content between 30% - 80%;

[0039] 2) Select those with the number of homologous regions < 10;

[0040] 3) Select probes whose regions do not contain SSR or N regions.

[0041] The probe design results are shown in Table 2:

[0042] Table 2 Probe design results

[0043]

[0044] Note: This evaluation result is the probe design result. It is not that as long as a probe is designed, it will definitely be able to capture the target site.

[0045] A total of 3,590 probes targeting mammals were selected, each with a length of 110 bp. The nucleotide probe sequences modified with biotin (B) (biotin is located at the 5' end of the probe) were synthesized using the in-situ chip synthesis technology.

[0046] 3. Primer synthesis and development testing

[0047] The 3,590 probes selected in step 2 of the synthesis were used to test a total of 10 samples from 7 test species. The 7 test species were yak (sample numbers qh421, qh565), argali (sample numbers M-15, M-11), Siberian ibex (sample numbers 661, 660), white-cheeked gibbon (sample number B12), Indian muntjac (sample number 62), Guizhou golden monkey (sample number 125), and black muntjac (sample number 188). The samples used were blood samples. All samples were obtained during their physical examinations, and the sample collection was reviewed and approved by the Academic Committee of Beijing Zoo.

[0048] 3.1 Sample DNA quality inspection

[0049] The DNA concentration of each test sample was measured using Qubit Fluorometric Quantitation (Thermo Fisher), and the integrity of the DNA was detected by 1% agarose gel electrophoresis. The qualified samples were placed in a 4°C refrigerator for storage and standby.

[0050] 3.2 Construction of sample DNA sequencing libraries

[0051] For each test sample, 12 μL of the DNA (200 ng) qualified in step 3.1 was placed in a 0.2 μL PCR tube. The tube was placed in an ultrasonic crusher to randomly physically fragment the DNA until the fragments were 200 - 400 bp. Then, 4 μL of GenoBaits EndRepairBuffer (Beijing BioDee Biotechnology Co., Ltd.) and 2.7 μL of GenoBaits End RepairEnzyme (Beijing BioDee Biotechnology Co., Ltd.) were added to the tube, and the volume was adjusted to 20 μL with water. The tube was placed in an ABI 9700 PCR instrument and incubated at 37°C for 20 minutes to complete the end repair and A-tailing of the fragmented DNA.

[0052] Take out the small tube from the PCR instrument, add 2 μL of GenoBaits Ultra DNA ligase (Breda Biotech Co., Ltd.), 8 μL of GenoBaits Ultra DNA Ligase Buffer (Breda Biotech Co., Ltd.), and 2 μL of GenoBaits Adapter (Breda Biotech Co., Ltd.), make up the volume to 40 μL with water, then place it on the ABI9700 PCR instrument and react at 22 °C for 30 minutes to complete the ligation of the sequencing adapter. Add 48 μL of Beackman AMPure XP Beads (Beackman Coulter) to the ligation product to purify the ligation product. After purification, perform fragment screening according to 0.65 + 0.2 times the magnetic beads, and retain the ligation product with an inserted fragment of 200 - 300 bp.

[0053] Add 5 μL of the sequencing adapter with a Barcode sequence (the used Barcode sequence is selected from the Barcode sequences in Sequence Listing Sequences 3591 - 3686, and different Barcodes are used to distinguish different samples), 1 μL of the P5 adapter, and 10 μL of GenoBaits PCR Master Mix (Breda Biotech Co., Ltd.) to the PCR tube from the previous step, and make up to 20 μL with pure water; perform amplification using the ABI9700 PCR instrument. The amplification program is: pre-denaturation at 95 °C for 5 min, denaturation at 95 °C for 30 s, annealing at 60 °C for 30 s, extension at 72 °C for 30 s; repeat steps 2 - 4 for a total of 8 cycles; extension at 72 °C for 5 min.

[0054] Add 24 μL of Beckmen AMPure XP Beads (Beckmen Coulter) to the second-round PCR product, pipette up and down evenly, place the 0.2 μL PCR tube on the magnetic stand until the solution becomes clear, discard the supernatant, wash the magnetic beads once with 75% ethanol, and elute the library DNA with Tris-HCl with a pH value of 8.0.

[0055] 3.3 Sample DNA Hybridization Capture

[0056] Take 500 ng of the completed sample DNA sequencing library, add 5 μL of GenoBaits Block I (Breda Biotech Co., Ltd.) and 2 μL of GenoBaits Block II (Breda Biotech Co., Ltd.), place it on an Eppendorf Concentratorplus (Eppendorf Co., Ltd.) vacuum concentrator, and evaporate it to dry powder at a temperature of ≤70 °C. Add 8.5 μL of GenoBaits 2x Hyb Buffer (Breda Biotech Co., Ltd.), 2.7 μL of GenoBaits Hyb Buffer Enhancer (Breda Biotech Co., Ltd.), and 2.8 μL of Nuclease-Free Water to the dry powder tube. After pipetting and mixing evenly, place it on an ABI 9700 PCR instrument and incubate at 95 °C for 10 minutes. Then take out the PCR tube, add 3 μL of the synthesized probe (60 ng / μL), vortex and mix evenly, and place it on an ABI 9700 PCR instrument and incubate at 65 °C for 2 hours to complete the probe hybridization reaction.

[0057] Add 100 μL of GenoBaits DNA Probe Beads (Breda Biotech Co., Ltd.) to the reaction system completed in the previous hybridization step, pipette up and down 10 times, and place it on an ABI 9700 PCR instrument and incubate at 65 °C for 45 minutes to allow the magnetic beads to bind to the probe. Wash the magnetic beads bound to the probe with 100 μL of GenoBaits Wash Buffer I (Breda Biotech Co., Ltd.) and 150 μL of GenoBaits Wash Buffer II (Breda Biotech Co., Ltd.) at 65 °C respectively, and then wash the magnetic beads with 100 μL of GenoBaits Wash Buffer I (Breda Biotech Co., Ltd.), 150 μL of GenoBaits Wash Buffer II (Breda Biotech Co., Ltd.), and 150 μL of GenoBaits Wash Buffer III (Breda Biotech Co., Ltd.) at room temperature respectively. Resuspend the washed magnetic beads with 20 μL of Nuclease-Free Water.

[0058] Take 13 μL of the resuspended DNA (with magnetic beads) and add it to a new 0.2 mL PCR tube. Then add 15 μL of GenoBaits PCR Master Mix (Breda Biotech Co., Ltd.) and 2 μL of GenoBaits Primer Mix (Breda Biotech Co., Ltd.) to configure the post-PCR system, and perform library amplification using an ABI 9700 PCR instrument. The amplification program is: pre-denaturation at 95 °C for 5 min, denaturation at 95 °C for 30 s, annealing at 60 °C for 30 s, extension at 72 °C for 30 s; repeat steps 2 - 4 for a total of 15 cycles; extension at 72 °C for 5 min.

[0059] Add 45 μL of Beckmen AMPure XP Beads (Beckmen) to the post-PCR product and pipette up and down to mix evenly. Then place the 0.2 mL PCR tube on a magnetic stand until the solution becomes clear. Discard the supernatant and wash the magnetic beads twice with 75% ethanol. Elute the library DNA with Tris-HCl at pH 8.0. Complete the hybridization capture of the test sample.

[0060] 3.4 Quality control and sequencing of the sample DNA hybridization capture library

[0061] Measure the DNA concentration of the library using Qubit Fluorometric Quantitation (Thermo Fisher), and then detect whether the fragment size of the library DNA is between 300 - 400 bp by agarose gel electrophoresis. Sequence the constructed DNA library using the Illumina Hiseq X ten sequencer.

[0062] 3.5 Test data analysis

[0063] The original sequenced reads obtained by sequencing, or raw reads, contain reads with adapters and low quality. For second-generation sequencing technology, the distribution of sequencing error rates has the following two main reasons: 1) Due to the consumption of chemical reagents during the sequencing process, the sequencing error rate will increase with the increase in the length of the sequenced reads. 2) Incomplete binding of random primers and DNA templates during the PCR process may lead to a relatively high sequencing error rate for the first few bases.

[0064] To ensure the quality of information analysis, it is necessary to filter the raw reads to obtain clean reads and use the clean reads for subsequent analysis. Use the software fastp (version 0.20.0, parameters: -n 10 -q 20 -u 40) to filter the raw reads. The steps of data processing are as follows:

[0065] 3.5.1 Remove adapter sequences;

[0066] 3.5.2 When the content of N in the sequencing read exceeds 10% of the length of this read, this pair of paired reads needs to be removed;

[0067] 3.5.3 When the number of low-quality (Q ≤ 20) bases in the sequencing read exceeds 40% of the length of this read, this pair of paired reads needs to be removed.

[0068] After obtaining the detection data, after performing full-length assembly on the sequencing results, analysis can be carried out again. It can be in the form of clustering or in the form of Blast. Ultimately, the sequence with the closest genetic relationship to the sequencing results is found to determine the species to which the target sample belongs.

[0069] A total of 10 samples of the above 7 test species were tested using a complete set of probe compositions for mammals (a total of 3,590 probes, each with a length of 110 bp. The specific probe sequences are shown in Sequence 1 - Sequence 3,590 of the sequence listing). In the test results, a coverage rate of 1 (i.e., 100%) indicates that the sample to be tested is the corresponding species, and a coverage rate not equal to 1 indicates that the sample to be tested is not the corresponding species. The test results are specifically shown in Table 1 below:

[0070] Table 1

[0071]

[0072]

[0073]

[0074] The results show that the species detected in the 10 samples are completely consistent with their actual species, and the accuracy rate is 100%. A COI gene capture kit produced with a probe composition composed of single-stranded DNA of Sequence 1 - Sequence 3,590 can be used for the detection of batch samples.

[0075] The above has detailed the present invention. For those skilled in the art, without departing from the purpose and scope of the present invention and without the need for unnecessary experiments, the present invention can be implemented within a relatively wide range under equivalent parameters, concentrations, and conditions. Although specific embodiments of the present invention are given, it should be understood that the present invention can be further improved. In short, according to the principle of the present invention, this application intends to include any changes, uses, or improvements to the present invention, including changes made using conventional techniques known in the art that are outside the scope disclosed in this application. Some basic features can be applied according to the scope of the following appended claims.

[0076] Industrial Application

[0077] The present invention discloses a probe composition, a kit and an application for identifying or assisting in identifying mammalian species. The probes are nucleotide probes shown in SEQ ID NO: 1-SEQ ID NO: 3590 in the sequence listing. By using the probes of the present invention to capture the COI gene of the target species and constructing a library to obtain the DNA sequence of the COI gene by means of high-throughput detection of next-generation sequencing, the advantages of flexibility, convenience and rapid result obtaining of first-generation sequencing can be taken into account at the same time, the problem of cumbersome operation of batch samples in first-generation sequencing can be solved, and the risk of inability to obtain accurate COI gene sequences caused by unstable partial results in first-generation sequencing can be avoided. Developing the probes into a mammalian COI gene capture kit for DNA barcoding research of a large number of mammalian species can achieve rapid and batch identification of unknown mammalian species.

Claims

1. A probe composition for identifying or assisting in the identification of mammalian species, characterized in that, The probe composition is a representative sequence obtained by sequence clustering of the mammalian COI gene sequence. For the representative sequence, probes are designed to cover each SNP locus, and the resulting probe composition; the probe composition is a combination consisting of 3,590 single-stranded DNAs shown in Sequence Listing Sequences 1-3,590.

2. The probe composition for identifying or assisting in identifying mammalian species according to claim 1, wherein: The sequence clustering is performed using the Angiosperms353 method, with a genetic distance set at 0.05 and a coverage depth of 2X.

3. The probe composition for identifying or assisting in the identification of mammalian species according to claim 2, characterized in that: The design of probes to cover each SNP locus means that 2 probes are designed to cover each SNP locus.

4. A method for identifying or assisting in the identification of a mammalian species using a probe composition, characterized in that: It includes capturing the COI gene of the mammalian to be tested with the probe composition according to any one of claims 1-3, constructing a library, obtaining the DNA sequence of the COI gene through high-throughput detection by next-generation sequencing, and determining the mammalian species to be tested based on the obtained COI gene sequence.

5. A reagent or kit for identifying or assisting in the identification of mammalian species, characterized in that: The reagent or kit includes the probe composition according to any one of claims 1-3.

6. Use of the probe composition according to any one of claims 1-3 in identifying or assisting in the identification of mammalian species.

7. Use of the probe composition according to any one of claims 1-3 in the preparation of a product for identifying or assisting in the identification of mammalian species.

8. Use of the method according to claim 4 in identifying or assisting in the identification of mammalian species.

9. Use of the reagent or kit according to claim 5 in identifying or assisting in the identification of mammalian species.

Citation Information

Patent Citations

  • Detection method and kit of mammalian and aves animal origin components

    CN107541566A

  • Method and means for identification of animal species

    US20170088903A1