SNP site combination for constructing molecular identity card of guo qin 2 and application thereof
By constructing a molecular identification code for Guoqin No. 2 and utilizing a combination of 12 highly consistent and stable SNP loci, the problem of accurate identification of new Scutellaria baicalensis varieties was solved, enabling rapid and accurate germplasm identification and market protection.
Patent Information
- Application Number
- CN202511669313.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-11-14
AI Technical Summary
Current technology lacks effective methods to accurately identify the new Scutellaria baicalensis variety Guoqin No. 2. Traditional traits and physiological and biochemical indicators are insufficient to distinguish germplasm resources of the same species from different sources, as well as similar varieties and adulterants.
A molecular identification system for the new Scutellaria baicalensis variety, Guoqin 2, was constructed by combining 12 highly homogeneous and stable SNP loci and using gene chips or reagent kits for identification. The genotype and coding results of the SNP loci were then used for accurate determination.
This has enabled rapid and accurate identification of Guoqin No. 2, ensuring the safety and reliability of the seed source, preventing counterfeit and substandard varieties from entering the market, and providing an important guarantee for the development of the Chinese medicinal materials industry.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of molecular biology technology, and specifically discloses a SNP site combination for constructing a molecular identity card of Guoqin No. 2 and application thereof. BACKGROUND
[0002] Scutellaria baicalensis Georgi is a common medicinal material, which has the effects of clearing heat and drying dampness, purging fire and detoxifying, and stopping bleeding. Guoqin No. 2 is a Scutellaria plant, which was awarded the plant new variety right by the State Forestry and Grassland Administration in April 2023, but there is still a lack of identification method for the germplasm. Although the traditional traits and physiological and biochemical indexes can identify the medicinal material germplasm resources to some extent, it is difficult for the traditional methods and physical and chemical methods to accurately identify the germplasm resources of the same species from different sources and the similar products and counterfeit products with close genetic relationship.
[0003] With the continuous deepening of the breeding of new varieties of medicinal materials, the accurate identification of new varieties of medicinal materials and mainstream family cultivation products and wild resources is more conducive to the healthy development of the medicinal material industry. The DNA molecular identity card has obvious advantages in identifying and evaluating germplasm resources due to its unique germplasm specific identification function. The effective protection and reasonable development of the new variety Guoqin No. 2 by using the molecular identity card technology provide an important basis for the preservation, research and development and utilization of Scutellaria baicalensis Georgi germplasm resources, and also ensure the safety and reliability of the introduction of Guoqin No. 2, which provides an important guarantee for the high-quality development of Scutellaria baicalensis Georgi industry. SUMMARY
[0004] In view of the lack of effective identification method for the new variety Guoqin No. 2 of Scutellaria baicalensis Georgi in the prior art, the application provides a SNP site combination for constructing a molecular identity card of Guoqin No. 2 and application thereof.
[0005] Firstly, this invention provides a SNP locus combination for constructing the molecular identity card of a new Scutellaria baicalensis variety, Guoqin 2. The SNP locus combination is located in the Scutellaria baicalensis reference genome, which is accessed by the National Center for Biotechnology Information (NCBI) under accession number GWHDEDD00000000. It is a set of SNP loci with high consistency, stability, and repetition among SNP varieties of Scutellaria baicalensis, consisting of the following 12 SNP loci: SNP1 is located on chromosome 1 at 57682107 bp, with a polymorphism of G / A; SNP2 is located on chromosome 2 at 21413410 bp, with a polymorphism of T / C; SNP3 is located on chromosome 2 at 22303165 bp, with a polymorphism of G / A; SNP4 is located on chromosome 2 at 32319969 bp, with a polymorphism of G / A; and SNP5 is located on chromosome 2 at 32319969 bp, with a polymorphism of G / A. The polymorphisms are G / T; SNP5 is located on chromosome 4 at 3623729 bp, with a polymorphism of T / G; SNP6 is located on chromosome 5 at 13129465 bp, with a polymorphism of C / T; SNP7 is located on chromosome 5 at 16767233 bp, with a polymorphism of A / C; SNP8 is located on chromosome 6 at 9220542 bp, with a polymorphism of C / T; SNP9 is located on chromosome 6 at 15964325 bp, with a polymorphism of C / G; SNP10 is located on chromosome 7 at 2489631 bp, with a polymorphism of G / A; SNP11 is located on chromosome 8 at 27198552 bp, with a polymorphism of C / T; SNP12 is located on chromosome 9 at 22225149 bp, with a polymorphism of C / T.
[0006] Secondly, this invention provides a product for detecting the SNP locus combination of a new Scutellaria baicalensis variety, Guoqin 2, as described above. In some embodiments, the product is a reagent kit; in other embodiments, the product is a gene chip.
[0007] Thirdly, the present invention provides the application of the aforementioned SNP site combinations or products in the identification of the new Scutellaria baicalensis variety Guoqin No. 2.
[0008] Fourthly, the present invention provides a method for identifying a new Scutellaria baicalensis variety, Guoqin 2, comprising: obtaining the genotypes of 12 SNP loci of the Scutellaria baicalensis to be tested, wherein the 12 SNP loci are SNP1 to SNP12 as described above; determining whether the Scutellaria baicalensis to be tested is the new Scutellaria baicalensis variety Guoqin 2 based on the genotype results of the 12 SNP loci of the Scutellaria baicalensis to be tested, wherein the determination method is as follows: if the genotypes of SNP1 to SNP12 are GG, TT, GG, GG, TT, CC, AA, CC, CG, GA, CT, CC respectively, then the Scutellaria baicalensis to be tested is the new Scutellaria baicalensis variety Guoqin 2, otherwise it is not the new Scutellaria baicalensis variety Guoqin 2.
[0009] Fifthly, the present invention provides a method for identifying a new Scutellaria baicalensis variety, Guoqin 2, comprising: obtaining the genotypes of 12 SNP loci of the Scutellaria baicalensis to be tested, wherein the 12 SNP loci are SNP1 to SNP12 as described above; arranging the genotypes of the 12 SNP loci of the Scutellaria baicalensis to be tested in the order of SNP1 to SNP12, and converting them into numerical codes according to the coding rule that bases N, A, G, C, and T correspond to numbers 0 to 4 respectively, to obtain the molecular identification code of the Scutellaria baicalensis to be tested; determining whether the Scutellaria baicalensis to be tested is the new Scutellaria baicalensis variety Guoqin 2 based on the molecular identification code result of the Scutellaria baicalensis to be tested, wherein the determination method is as follows: if the molecular identification code of the Scutellaria baicalensis to be tested is 224422224433113332213433, then the Scutellaria baicalensis to be tested is Guoqin 2, otherwise it is not Guoqin 2.
[0010] Sixthly, the present invention provides a method for constructing a molecular identity card for a new Scutellaria baicalensis variety, Guoqin 2, comprising: obtaining the genotypes of 12 SNP loci of Guoqin 2, wherein the 12 SNP loci are SNP1 to SNP12 as described above; arranging the genotypes of the 12 SNP loci in sequence, and converting them into numerical codes according to the coding rule that bases N, A, G, C, and T correspond to different numbers, to obtain the molecular identity card of Guoqin 2.
[0011] In some implementation cases, the aforementioned steps for obtaining the genotypes of the 12 SNP loci of the Scutellaria baicalensis to be tested include: extracting genomic DNA from the tissue sample of the Scutellaria baicalensis to be tested to obtain the DNA sample to be tested; performing whole-genome sequencing on the DNA sample to be tested, and performing quality control screening on the raw data obtained from the sequencing to obtain quality control qualified sequences; and comparing the quality control qualified sequences with the reference genome to obtain the genotypes of the 12 SNP loci of the Scutellaria baicalensis to be tested.
[0012] In some implementation cases, the aforementioned genomic DNA Nanodrop detection showed an OD260 / 280 ratio between 1.8 and 2.2, with no protein or visible contaminant contamination; the Qubit detection concentration was greater than 50 ng / μl.
[0013] In some implementation cases, the aforementioned quality control screening includes: filtering raw sequencing data to remove low-quality bases and adapter contamination, and screening data with a Q30 base ratio of ≥90%.
[0014] In some implementation cases, the aforementioned alignment includes: using BWA software to align the quality-controlled qualified sequences to the reference genome, calculating the genotype of the target site using samtools software based on the alignment results, and finally using Annovar software to annotate the mutation sites.
[0015] The beneficial effects of this invention are as follows:
[0016] The SNP locus combination provided by this invention can quickly and accurately identify the authenticity of the parent plants and their offspring seeds of the new Scutellaria baicalensis variety Guoqin 2. This is of great significance for protecting plant variety rights, preventing counterfeit and inferior varieties from entering the market, and ensuring the safety and reliability of superior seed sources.
[0017] Based on the provided SNP locus combination, this invention constructs a molecular identification card for the new Chinese medicinal herb Scutellaria baicalensis variety Guoqin No. 2, which broadens the application of molecular identification cards in the identification of Chinese medicinal herb varieties, lays the foundation for the establishment of molecular identification cards for the identification of other Chinese medicinal herb varieties and the establishment of a seed source traceability system, and provides theoretical support for the establishment of variety identification standards. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Example 1: Screening of SNP locus combinations for the new Scutellaria baicalensis variety, Guoqin 2
[0020] Step 1: Collect a total of 107 tissue samples, including those from non-Guoqin No. 2 Scutellaria baicalensis regions such as Hebei, Shaanxi, Inner Mongolia, and Liaoning, as well as tissue samples from Guoqin No. 2 Scutellaria baicalensis.
[0021] Step 2: Total DNA extraction, quality testing, and genotyping
[0022] 2.1 Total DNA was extracted from plant tissues using a universal genomic DNA extraction kit (magnetic bead method, China Migi-Yuhua kit);
[0023] 2.2 A DYY-6C electrophoresis apparatus was used, and 1% agarose gel electrophoresis was employed for detection.
[0024] 2.3 The concentration of DNA was detected using a Nanodrop UV spectrophotometer, with a retention absorbance ratio (OD260 / 280) between 1.8 and 2.2. The concentration was greater than 50 ng / μl and the total amount was greater than 2 μg using a Qubit 3.0 spectrophotometer.
[0025] 2.4 After the DNA samples passed the initial testing, they were randomly fragmented using a Covaris ultrasonic disruptor. The entire library preparation process then involved end repair, A-tailing, sequencing adapter addition, purification, and PCR amplification. After library construction, preliminary quantification was performed using Qubit 3.0 to dilute the library. Subsequently, the insert fragments were detected using an Agilent 2100. Once the insert fragment size met expectations, the effective concentration of the library was accurately quantified using Q-PCR to ensure library quality.
[0026] 2.5 After the libraries pass the detection, different libraries are pooled into flow cells according to the effective concentration and the required amount of data to be sequenced. After cBOT clustering, sequencing is performed using the Illumina NovaSeq high-throughput sequencing platform.
[0027] Step 3: Sequencing data quality control and sequence alignment
[0028] 3.1 Raw sequencing data (FASTQ format) was filtered using Trimmomatic, low-quality bases (HEADCROP:10) were removed, and adapter contamination was eliminated to generate high-quality Clean Reads. Data quality was assessed using FastQC, including base quality distribution, GC content, and sequence reproducibility, to ensure that the Q30 base ratio was ≥90% and there was no excessive shift or contamination.
[0029] 3.2 The quality-controlled sequences were aligned to the Scutellaria baicalensis reference genome with accession number GWHDEDD00000000, which is included in the National Center for Biotechnology Information, using BWA (v 0.7.13-r1126) software with the default parameters. Based on the alignment results, the genotype of the target site was calculated using samtools software (version 0.1.18). Finally, Annovar (2018-04-16) software was used to annotate the mutation sites.
[0030] Step 4: SNP site screening
[0031] 4.1 A total of 1,121,566 original SNP loci were obtained through raw data analysis. To ensure the accuracy and reliability of subsequent analysis results, all SNP loci underwent rigorous quality filtering, with deletion filtering <0.2 and minor allele frequency filtering >0.05. Dialleles and 4DTv loci were also screened. After multiple rounds of rigorous screening, 24,783 high-quality SNP loci remained.
[0032] 4.2 Candidate SNPs were filtered and screened from high-quality SNP loci using PLINK software. The specific criteria were as follows: a. biallelic loci; b. 4-fold degenerate SNPs; c. heterozygosity <0.2; d. deletion rate <0.2; e. minor allele frequency (MAF) >0.05; f. no InDel or SNPs in the 100bp flanking region.
[0033] 4.3 Screening results: A total of 12 SNP sites were finally obtained that can be used for the identification of Guoqin No. 2, as shown in Table 1.
[0034]
[0035] Note: Ref is the reference genome genotype; Alt is the validation result; PIC is the multi-site information content; MAF is the minor allele frequency; Pi is the nucleotide diversity.
[0036] The specific information of the nucleotide sequences of the SNP sites corresponding to SNP1-SNP12, SEQ ID NO.1~12, is as follows (underlined bases are SNP sites):
[0037] 1 Sbai1_57682107_57682207 (SEQ ID NO.1)
[0038] ACACAATTTCGTTCATGCCTATCATTGTTGTGTTCTGTCAAACCATGAATGTTTTAATGGAAGATTAGCAAATTTCATCATCCTTCAGGGAAGTGAAGAA [G|A] TTCTAAGAATGTGTGGTCTATTGTTCTTGATTATATGTCAAATTAAGTATGAGTAAGTAGATGAAGAAATCAATGTAACAGTTCTAGTCTTAATTAGGAA
[0039] 2 Sbai2_21413410_21413510 (SEQ ID NO.2)
[0040] TTTTTTACAATGGTCGGGATGAGTTACATTTTGATGTTAAGTAGGGAATAACCTGAAGCTCTGATGAAAGCTGGGGCTGTTTATAAATACGACTTGTCTCT [T|C]CCTCCAGAAAAGATGTATGATCTTGTTGAGGAAATGCGTGGACGCTTTGGTAATCGACTTAACATAATCTTTCAGAAAGATAAGACTTACTCAAATGTAT
[0041] 3 Sbai2_22303165_22303265(SEQ ID NO.3)
[0042] TGCATTCTGCTTTGGCGAGGTCTCTTCCAGTAGCTTCCTGTAGTCATCAGTTTACCCGATTCTTCTTCACCATTCTCTTCAGTATCCTTCCTCTGTTCTT [G|A] TAATCTCTTGCAGTAACTCTGTTGCAGTAACTTATCCTAAGATTGGTTATGCTTGTTAAAGTTACTTAAAGTTTTGTCTCATCAAAACTTTCTTCAAATG
[0043] 4 Sbai2_32319969_32320069(SEQ ID NO.4)
[0044] CCGTTGTCTGGTAATTAGCTTCTTCAGCATGCATATCATTATACGTCTTCCTGCGTTTATTGTGCTTATTTTCTGTTTTCTGGTGCTAGGGATACAGGAA [G|T] GCCTCAGCTTCTGAAACGTAAGAAAATTTTCTGGTGATGCTTTACTTTGTGTTTCCACCATCCTTGCATACTTATGCTGGGTTGTACAAGTTTGCTAGTT
[0045] 5 Sbai4_3623729_3623829(SEQ ID NO.5)
[0046] CTGTTTGTTTGAAATACTCATTCTCTGTGGTTGAAAATAGACTAGATTTTATTGCCATGAGACCTGCTTCATGTTAGATGCAGTAGAAGGAAGTTACTGG [T|G]ATGGAAAGTTTCAGATTTTCTAGGCTTTGCTGTTATCGGCAATATAATAGTCATAGGCTTGATGTGACCTGCTTAATCATGACTTTGATCATG
[0047] 6 Sbai4_13129465_13129565 (SEQ ID NO.6)
[0048] AAGCCAACCTTCACTACAAAAGCACAAATGTAATTAATTGCCCTCACCTCAAGATTTGATTCATCAAGTGCATCCTTGACATCTGTACTGTACTT [C|T] CACCTTATCTTTGTGCATCTTTGATAATCAAAGCCCTGAACCAGCAACCCCTGATGCCACTATCTTTGAGAGAACCTCAACTTGGGTGCCAATGCCCAAA
[0049] 7 Sbai5_16767233_16767333(SEQ ID NO.7)
[0050] TTTTCTCTTTCTAATATTGTGGCTTGGATAAATAATATTTCGTTTCAAAATTGTTGTTTGTATAAAATTTGTTCAATCATTTGTTGTTAT [A|C] CATCHGTTATTTCTTTATGAATTTGTTATACTCTCGTGTACTGTGATTTTATGAAAATCTCTGATTGAAGTTTTTTTAAAATGAGA
[0051] 8 Sbai6_9220542_9220642(SEQ ID NO.8)
[0052] TTATTAAATTTGGGAACAAACCACATACGAATTACTCGAACCCTAATCTTAAAAATCTCCTTCGATCAATCTCTGCGAAAAAAGATATTCTTGTT [C|T]CATATTCCTACACACTATTCAAATGAGAATGAGAGTGATTTGATTTTGTTGTGTCATATTTTCAATGGTGAATTTCTTTAAATAAGACCATCAATTGAAA
[0053] 9 Sbai6_15964325_15964425(SEQ ID NO.9)
[0054] TATTGAACGAGTCACTATTAAAAAAATTAACCGAGAGACTCGTTGTGAGATTATACAAGTAACCAAGTGACTCGTATAACCTTAAACGTGTCACTCGTTCA [C|G] CTCAACGAGTGACTCGTTTAAGGCTCAAAGAGTCACTCGACTAATTGCATAGAATTACTCGGTCACCATCTCATAGAGTCTCTCGGCTAATAGTGGCTCG
[0055] 10 Sbai7_2489631_2489731(SEQ ID NO.10)
[0056] AAACCAGCATCAGGAGGAAGCAAATCCAGGCTGGCAAAGTATAGAAACAAAACTAGTTTGAAGAAGGCGAGAACAAATTATACAACAGGAAGGATACGTG [G|A] AACTTCCCTGCGAGTGCAATTTGAAGGCAAGCCCTCTACAAACAATGTGCTGCTAGCATCAGGAGGAAGCGGAATTTCAGGCCGCAACCCAAGTGCAACA
[0057] 11 Sbai4_27198552_27198652(SEQ ID NO.11)
[0058] TCTTGTCTTCATCTCTTTTCATTGTTATTATATATATTGTTTTGTTATATTGATCCATAGAAAGTGGTGCACACGATATCATATTTTACAACACATTATC [C|T]GCACGAGTCTCTACCTGATCTGAGTGTTTGCCCTATGAAAGAATCGTGTCTTTCCTGCAAGGTATTTTTCTTTGATAAAATTCAAATATCCAATTTATTA
[0059] 12 Sbai9_22225149_22225249 (SEQ ID NO.12)
[0060] AAATTACTTTTAATTCCTTAATTAGTTTATTAATTTTTTAATTTGTTAATTAGTTAACTTTAATTCGTTAGATTCCTAGTTAATTTTTAATTAGCTTAGT [C|T] AATTTATAATTTGTTAGTTTATTATTTTTTTAATTAGTTAGTTAATTTTTTTAATTCCTAGTTAATTTAGTTAATTTTTAATTCGTTAGTTTTTCGTTAAT
[0061] Specifically, the genotypes of the 12 SNP loci of the Guoqin No. 2 variety are as follows: GG, TT, GG, GG, TT, CC, AA, CC, CG, GA, CT, CC.
[0062] The genotypes of the 12 SNP loci were converted into binary codes (using numbers 0–4 to encode the bases N, A, G, C, and T), resulting in the molecular identification code: 224422224433113332213433.
[0063] Example 2 uses the SNP locus combination obtained from Example 1 to identify the new Scutellaria baicalensis variety Guoqin No. 2.
[0064] Step 1: Collect Scutellaria baicalensis seeds and seedlings of the SX1, SX2, HB and GS series (not the germplasm of Guoqin No. 2) sold on the market, and add the selected Guoqin No. 2 sample GQ2 series. See Table 2 for sample information.
[0065]
[0066]
[0067] Step 2: Total DNA was extracted using a plant tissue genomic DNA extraction kit. After extraction, the DNA was detected by 1% agarose gel electrophoresis using a DYY-6C electrophoresis apparatus, and the DNA concentration was detected by a Nanodrop UV spectrophotometer. The retention absorbance ratio (OD260 / 280) was between 1.8 and 2.2, and the concentration was detected by Qubit 3.0 with a concentration greater than 50 ng / μl and a total amount greater than 2 μg.
[0068] Step 3: After the DNA sample passes the initial testing, it is randomly fragmented using a Covaris ultrasonic disruptor. The entire library preparation process then involves end repair, A-tailing, sequencing adapter addition, purification, and PCR amplification. After library construction, preliminary quantification is performed using Qubit 3.0 to dilute the library. Subsequently, the insert fragments are detected using an Agilent 2100. Once the insert fragment size meets expectations, the effective concentration of the library is accurately quantified using Q-PCR to ensure library quality.
[0069] Step 4: After the libraries pass the test, different libraries are pooled into flow cells according to the effective concentration and the required amount of data to be sequenced. After cBOT clustering, sequencing is performed using the Illumina NovaSeq high-throughput sequencing platform.
[0070] Step 5: The raw sequencing data (FASTQ format) is filtered using Trimmomatic to remove low-quality bases (HEADCROP:10) and adapter contamination, generating high-quality Clean Reads. Data quality is assessed using FastQC, including base quality distribution, GC content, sequence reproducibility, etc., ensuring that the Q30 base percentage is ≥90% and there is no excessive shift or contamination.
[0071] Step 6: Use BWA (v 0.7.13-r1126) software to align the quality-controlled sequences to the reference genome with the default parameters. Based on the alignment results, use samtools software (version 0.1.18) to calculate the genotype of the target site. Finally, use Annovar (2018-04-16) software to annotate the mutation sites.
[0072] Step 7: Determine the genotype of the sample to be tested at the same SNP locus location based on the relevant location information of S. 2 in Table 1.
[0073] Step 8: Encode the obtained sites according to the numbers 0–4 for the bases N, A, G, C, and T to obtain the molecular identification code. See Table 3 for specific sample SNP site genotype information and codes.
[0074]
[0075]
[0076] Step 9: Compare the molecular identification code of the obtained test sample with the molecular identification code of Guoqin No. 2 SNP. The results show that the Scutellaria baicalensis seeds and seedlings collected in the market, SX1, SX2, HB and GS series, do not match the molecular identification code of Guoqin No. 2 and do not belong to Guoqin No. 2 sample. The selected sample GQ2 series matches the molecular identification code of Guoqin No. 2 and belongs to Guoqin No. 2 sample. This result is consistent with the actual situation, which shows that the accuracy of the SNP site combination identification result of the present invention is 100%.
[0077] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.
Claims
1. Application of SNP locus combinations in the identification of the Scutellaria baicalensis variety Guoqin No. 2; The SNP locus combination is located in the Scutellaria baicalensis reference genome, which is accessed by the National Center for Biotechnology Information (GWHDEDD00000000) and consists of the following 12 SNP loci: SNP1 is located at 57682107 bp on chromosome 1, and its polymorphism is G / A; SNP2 is located on chromosome 2 at 21413410 bp, and its polymorphism is T / C. SNP3 is located on chromosome 2 at 22303165 bp, and its polymorphism is G / A. SNP4 is located at 32319969 bp on chromosome 2, and its polymorphism is G / T. SNP5 is located at 3623729 bp on chromosome 4, and its polymorphism is T / G; SNP6 is located on chromosome 5 at 13129465 bp, and its polymorphism is C / T. SNP7 is located on chromosome 5 at 16767233 bp, and its polymorphism is A / C. SNP8 is located on chromosome 6 at 9220542 bp, and its polymorphism is C / T. SNP9 is located on chromosome 6 at 15964325 bp, and its polymorphism is C / G; SNP10 is located at 2489631 bp on chromosome 7, and its polymorphism is G / A. SNP11 is located on chromosome 8 at 27198552 bp, and its polymorphism is C / T. SNP12 is located on chromosome 9 at 22225149 bp, and its polymorphism is C / T. The application includes: obtaining the genotypes of the 12 SNP loci of the Scutellaria baicalensis to be tested; and determining whether the Scutellaria baicalensis to be tested is the Scutellaria baicalensis variety Guoqin 2 based on the genotypes of the 12 SNP loci of the Scutellaria baicalensis to be tested. The determination method is as follows: if the genotypes of SNP1 to SNP12 are GG, TT, GG, GG, TT, CC, AA, CC, CG, GA, CT, and CC, respectively, then the Scutellaria baicalensis to be tested is the Scutellaria baicalensis variety Guoqin 2; otherwise, it is not the Scutellaria baicalensis variety Guoqin 2.
2. A method for identifying the Scutellaria baicalensis variety Guoqin No. 2, characterized in that, The method includes: Genotypes of 12 SNP loci in *Scutellaria baicalensis* were obtained. The reference genome of *Scutellaria baicalensis* containing these 12 SNP loci was included in the National Center for Biotechnology Information (accession number GWHDEDD00000000). These SNPs consist of SNP1 located at 57682107 bp on chromosome 1, SNP2 located at 21413410 bp on chromosome 2, SNP3 located at 22303165 bp on chromosome 2, SNP4 located at 32319969 bp on chromosome 2, and SNP5 located at 36 bp on chromosome 4. It consists of SNP5 (23729 bp), SNP6 (13129465 bp on chromosome 5), SNP7 (16767233 bp on chromosome 5), SNP8 (9220542 bp on chromosome 6), SNP9 (15964325 bp on chromosome 6), SNP10 (2489631 bp on chromosome 7), SNP11 (27198552 bp on chromosome 8), and SNP12 (22225149 bp on chromosome 9); Based on the genotype results of the 12 SNP loci of the Scutellaria baicalensis to be tested, it is determined whether the Scutellaria baicalensis to be tested is the Scutellaria baicalensis variety Guoqin 2. The determination method is as follows: if the genotypes of SNP1 to SNP12 are GG, TT, GG, GG, TT, CC, AA, CC, CG, GA, CT, CC respectively, then the Scutellaria baicalensis to be tested is the Scutellaria baicalensis variety Guoqin 2; otherwise, it is not the Scutellaria baicalensis variety Guoqin 2.
3. A method for identifying the Scutellaria baicalensis variety Guoqin No. 2, characterized in that, The method includes: Genotypes of 12 SNP loci in *Scutellaria baicalensis* were obtained. The reference genome of *Scutellaria baicalensis* containing these 12 SNP loci was included in the National Center for Biotechnology Information (accession number GWHDEDD00000000). These SNPs consist of SNP1 located at 57682107 bp on chromosome 1, SNP2 located at 21413410 bp on chromosome 2, SNP3 located at 22303165 bp on chromosome 2, SNP4 located at 32319969 bp on chromosome 2, and SNP5 located at 36 bp on chromosome 4. It consists of SNP5 (23729 bp), SNP6 (13129465 bp on chromosome 5), SNP7 (16767233 bp on chromosome 5), SNP8 (9220542 bp on chromosome 6), SNP9 (15964325 bp on chromosome 6), SNP10 (2489631 bp on chromosome 7), SNP11 (27198552 bp on chromosome 8), and SNP12 (22225149 bp on chromosome 9); The genotypes of the 12 SNP sites of the tested Scutellaria baicalensis were arranged in the order of SNP1 to SNP12, and converted into numerical codes according to the coding rule that the bases N, A, G, C, and T correspond to the numbers 0 to 4, respectively, to obtain the molecular identification code of the tested Scutellaria baicalensis. Based on the molecular identification code of the Scutellaria baicalensis to be tested, it is determined whether the Scutellaria baicalensis to be tested is the Scutellaria baicalensis variety Guoqin No.
2. The determination method is as follows: if the molecular identification code of the Scutellaria baicalensis to be tested is 224422224433113332213433, then the Scutellaria baicalensis to be tested is Guoqin No. 2, otherwise it is not Guoqin No.
2.
4. A method for constructing the molecular identity card of the Scutellaria baicalensis variety Guoqin 2, characterized in that, The method includes: Genotypes of 12 SNP loci in *Scutellaria baicalensis* var. *Guoqin* No. 2 were obtained. The reference genomes of these 12 SNP loci in *Scutellaria baicalensis* were accessed by the National Center for Biotechnology Information (accession number GWHDEDD00000000), and consist of SNP1 located at 57682107 bp on chromosome 1, SNP2 located at 21413410 bp on chromosome 2, SNP3 located at 22303165 bp on chromosome 2, SNP4 located at 32319969 bp on chromosome 2, and SNP5 located at 36 bp on chromosome 4. It consists of SNP5 (23729 bp), SNP6 (13129465 bp on chromosome 5), SNP7 (16767233 bp on chromosome 5), SNP8 (9220542 bp on chromosome 6), SNP9 (15964325 bp on chromosome 6), SNP10 (2489631 bp on chromosome 7), SNP11 (27198552 bp on chromosome 8), and SNP12 (22225149 bp on chromosome 9); The genotypes of the 12 SNP sites were arranged in sequence and converted into numerical codes according to the coding rule that N, A, G, C, and T correspond to different numbers, thus obtaining the molecular identity card of Guoqin No.
2.
5. The method according to any one of claims 2 to 3, characterized in that, The steps for obtaining the genotypes of the 12 SNP loci of the Scutellaria baicalensis to be tested include: Genomic DNA was extracted from the tissue sample of Scutellaria baicalensis to obtain the DNA of the sample to be tested; Whole-genome sequencing was performed on the DNA of the sample to be tested, and the raw data obtained from the sequencing were screened for quality control to obtain sequences that passed quality control. The quality-controlled sequences were compared with the reference genome to obtain the genotypes of the 12 SNP sites of the Scutellaria baicalensis to be tested.
6. The method according to claim 5, characterized in that, The genomic DNA was detected by Nanodrop with an OD260 / 280 ratio between 1.8 and 2.2, and was free of protein and visible contaminants; the Qubit detection concentration was greater than 50 ng / μl.
7. The method according to claim 5, characterized in that, The quality control screening includes: filtering raw sequencing data to remove low-quality bases and adapter contamination, and screening data with a Q30 base ratio ≥90%; the alignment includes: using BWA software to align the quality-controlled qualified sequences to the reference genome, calculating the genotype of the target site based on the alignment results using samtools software, and finally using Annovar software to annotate the mutation sites.
Citation Information
Patent Citations
SNP molecular marker for identifying production place of radix scutellariae and method and application thereof
CN112342313A