Identification method and application of traditional Chinese medicine based on Shizhen method

Through the Shizhen method combining nuclear, chloroplast and mitochondrial genome analysis and target sequence detection methods, the problem of molecular identification of traditional Chinese medicine in the prior art was solved, and the accuracy and efficiency of traditional Chinese medicine identification were improved.

CN118345184BActive Publication Date: 2025-05-20INST OF MEDICINAL PLANT DEV CHINESE ACADEMY OF MEDICAL SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410528276.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-03
Publication Date
2025-05-20
Estimated Expiration
2043-03-03

AI Technical Summary

Technical Problem

The prior art is difficult to identify traditional Chinese medicine molecules through nuclear, chloroplast and mitochondrial genomes, and the specificity of species-specific target sequences is low, resulting in identification difficulties.

Method used

The Shizhen method is used to combine nuclear, chloroplast and mitochondrial genome analysis with different target sequence detection methods to screen candidate target sequences, and select detection methods based on application scenarios to achieve the accuracy of Chinese medicine identification.

Benefits of technology

Through the Shizhen method, the identity of the Chinese medicine sample to be tested and the designated Chinese medicine base species can be accurately determined, and the risk of off-target errors can be eliminated, which improves the accuracy and efficiency of Chinese medicine identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118345184B_ABST
    Figure CN118345184B_ABST
Patent Text Reader

Abstract

The present application discloses a rhubarb identification method based on the Shizhen method and its application, wherein the standard specific target sequence is obtained by using the nuclear, chloroplast and mitochondrial genomes of the Chinese medicine base species, so that the species identification of the sample to be detected can be realized based on the standard specific target sequence. The Shizhen method can screen and obtain the standard specific target sequence for determining the identity of any Chinese medicine sample to be detected with the specified Chinese medicine base species, eliminating the risk of errors such as off-target, so that the Shizhen method can accurately determine the identity of any Chinese medicine sample to be detected with the specified Chinese medicine base species.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of Chinese medicine identification, and in particular to a method for identifying Chinese medicine based on the Shizhen method (Analysis of whole-GEnome, AGE, also known as whole genome analysis or A-Ge method) and its application. Background Art

[0002] Traditional Chinese medicine is a great treasure trove that has been passed down for thousands of years. It contains the precious experience and wisdom accumulated by countless predecessors in the fight against diseases. For thousands of years, traditional Chinese medicine has safeguarded the health and safety of the people. From the "Yellow Emperor's Internal Classic" to the "New Coronavirus Pneumonia Diagnosis and Treatment Plan", traditional Chinese medicine has played an important role in it, and the identification of traditional Chinese medicine is the key to ensuring the safety and effectiveness of traditional Chinese medicine clinical use. Traditional identification methods of traditional Chinese medicine (such as visual inspection, hand touch, and oral taste) can no longer meet the needs of traditional Chinese medicine identification with high timeliness, and physical and chemical identification methods based on the detection of index components often cannot solve identification problems such as artificial adulteration. DNA molecular identification methods identify species-specific DNA sequences for identification. They have the advantages of objectivity, sensitivity, strong specificity, simplicity, and ease of standardization. They are playing an increasingly important role in the field of traditional Chinese medicine identification.

[0003] The nuclear, chloroplast and mitochondrial genomes, as carriers of all the genetic information of a species, can provide a large number of species-specific DNA sequences and are ideal databases for the identification of traditional Chinese medicine. However, due to technical limitations, there is currently no molecular identification method for traditional Chinese medicine based on the nuclear, chloroplast and mitochondrial genomes.

[0004] A variety of methods have been developed for identifying species-specific DNA sequences, including the CRISPR / Cas system, TaqMan probe-based real-time PCR system, PCR system, and Sanger sequencing system. The CRISPR / Cas system relies on the trans-cleavage activity of Cas proteins (such as Cas12a and Cas13a): crRNA specifically recognizes the target sequence, and the trans-cleavage activity of the Cas protein is activated, non-specifically cleaving single-stranded DNA fluorescent signal molecules, generating detectable fluorescence. This reaction is performed at 37°C and is simple to operate, requiring only a constant temperature and fluorescence detection instrument. The TaqMan probe-based real-time PCR system adds a fluorescent probe that specifically recognizes the target sequence during the PCR reaction, allowing species identification by detecting the fluorescent signal. This method requires a real-time fluorescence quantitative PCR instrument and provides relatively accurate results. It is also a laboratory test for COVID-19. The PCR system uses primers designed to specifically recognize the target sequence, followed by PCR and species identification by electrophoresis. It is the simplest, but less specific. The Sanger sequencing system offers the most accurate results, but requires a sequencing instrument, making the test more cumbersome and time-consuming. In addition to the aforementioned specific DNA sequence detection methods, other sequence detection methods are also available and suitable for different application scenarios. However, the species-specific target sequences that have been screened often have disadvantages such as low specificity and high off-target risk. The construction of a library of suitable species-specific target sequences has become a key factor restricting the widespread application of existing DNA sequence recognition methods in the molecular identification of traditional Chinese medicines. Summary of the Invention

[0005] The species identification method of the present invention (hereinafter referred to as the Shizhen method) combines nuclear, chloroplast and mitochondrial genome analysis with different target sequence detection methods, thereby realizing the identification of traditional Chinese medicine from the nuclear, chloroplast and mitochondrial genome levels. Compared with the prior art, the Shizhen method screens candidate target sequences for identification from the nuclear, chloroplast and mitochondrial genomes, selects different detection methods according to the application scenario, and screens the corresponding species standard specific target sequences, fully tapping the potential of the nuclear, chloroplast and mitochondrial genomes for species identification. Taking into account the huge information contained in the nuclear, chloroplast and mitochondrial genomes and the diversity of target sequence selection provided by different target sequence detection methods, in theory, the Shizhen method can determine whether the genome of the sample to be tested has a specific target sequence through a variety of relatively simple and lower-cost methods. It does not need to sequence the whole genome of the sample to determine the identity of any traditional Chinese medicine sample to be tested with the specified traditional Chinese medicine base species, eliminating the risk of errors such as off-target. Therefore, the Shizhen method can accurately determine the identity of any traditional Chinese medicine sample to be tested with the specified traditional Chinese medicine base species.

[0006] In one aspect, the present invention provides a method for identifying traditional Chinese medicine based on Shizhen method, comprising:

[0007] (1) Obtain the nuclear, chloroplast, and mitochondrial genome sequences of the original species of traditional Chinese medicine and construct a small fragment genome library;

[0008] (2) extracting candidate target sequences from the small fragment genomic library and analyzing them, and screening candidate sequences that meet at least one of the screening conditions (a) to (d) from the candidate target sequences as standard specific target sequences, wherein the screening conditions include:

[0009] (a) The GC content of the candidate sequence is 40%-60%;

[0010] The candidate sequence cannot contain four or more nucleotide repeats, consecutive trinucleotide repeats, or discontinuous three or more trinucleotide repeats;

[0011] The candidate sequence cannot be complementary to the crRNA repeat sequence;

[0012] The G+C content of the 6 nucleotides at the 5' end of the candidate sequence is 30%-80%;

[0013] The candidate sequence contains at least three different nucleotides when compared with the nuclear, chloroplast, and mitochondrial genome sequences of the admixture and closely related species; and / or

[0014] There are no four consecutive sequences within the region from -50bp to +300bp or from -300bp to +50bp where the candidate sequence is located.

[0015] Repeat on A or GT;

[0016] (b) the G+C content of the candidate sequence and the primers that amplify the region where the candidate sequence is located is 30%-80%;

[0017] The annealing temperatures of the fluorescent probes designed based on the candidate sequences and the primers for the upstream and downstream amplification regions of the candidate sequences are 55°C-60°C and 68°C-70°C, respectively, and the annealing temperature difference between the primer pairs does not exceed 2°C;

[0018] The length of the candidate sequence and the primers that amplify the region where the candidate sequence is located is 15-30 bp;

[0019] The primers for the upstream and downstream amplification regions where the candidate sequence is located are spaced 50-150 bp apart and the forward primer is close to the fluorescent probe designed based on the candidate sequence;

[0020] The fluorescent probe designed based on the candidate sequence and the primers for the upstream and downstream amplification regions where the candidate sequence is located do not contain a hairpin structure and the primer pair cannot form dimers;

[0021] There should be no more than two G / C bases within 5 bp of the 3' end of the primers in the region where the upstream and downstream candidate sequences are amplified;

[0022] The fluorescent probe designed based on the candidate sequence contains more C bases than G bases; and / or

[0023] Candidate sequences were compared with nuclear, chloroplast and mitochondrial genomes of at least

[0024] Contains 2 differential nucleotides;

[0025] (c) The GC content of the amplification primers designed based on the candidate sequence is 40%-60%, and the four nucleotides are evenly distributed, without polypurines and polypyrimidines, and without GC-rich regions;

[0026] The length of the amplification primers designed based on the candidate sequence is 18-30 bp, and the difference between primer pairs does not exceed 3 bases;

[0027] The amplification primers designed based on the candidate sequence do not contain inverted repeat sequences and self-complementary sequences greater than 3 bp, and the primers cannot form dimers;

[0028] The annealing temperature of the amplification primers designed based on the candidate sequence is 55-60°C, and the annealing temperature difference between primer pairs does not exceed 2°C;

[0029] The amplification primer designed based on the candidate sequence contains more than one but no more than three G / C bases in the 5 bases at the 3' end; and / or

[0030] The candidate sequence contains at least two different nucleotides when compared with the nuclear, chloroplast, and mitochondrial genomes of the admixture and closely related species;

[0031] or, (d) the region where the candidate sequence is located is homozygous;

[0032] The GC content of the region where the candidate sequence is located is 30%-80%;

[0033] The region where the candidate sequence is located cannot contain more than four consecutive repeated nucleotides;

[0034] The region where the candidate sequence is located cannot contain methylated sequences;

[0035] There should be no hairpin structure in the region where the candidate sequence is located; and / or

[0036] The candidate sequence contains at least two different nucleotides when compared with the nuclear, chloroplast, and mitochondrial genomes of the admixture and closely related species;

[0037] (3) extracting genomic DNA from the Chinese medicine sample to be tested, optionally amplifying it, using the genomic DNA or its amplified product as a DNA substrate, and using a target sequence detection system to detect whether the standard specific target sequence exists in the DNA substrate, wherein the target sequence detection system includes a CRISPR / Cas12a system, a TaqMan probe-based real-time PCR system, a PCR system or a Sanger sequencing system, wherein, for the CRISPR / Cas12a system, if an obvious fluorescent signal is generated by the detection, the Chinese medicine sample to be tested is identical to the specified Chinese medicine origin species, otherwise it is not; for the TaqMan probe-based real-time For the PCR system, if the CT value is greater than 37 through the detection, the Chinese medicine sample to be detected is identical to the specified Chinese medicine origin species, otherwise it is not; for the PCR system, if the electrophoresis result produces obvious bands through the detection, the Chinese medicine sample to be detected is identical to the specified Chinese medicine origin species, otherwise it is not; for the Sanger sequencing system, if the sequencing result is the same as the standard specific target sequence, the Chinese medicine sample to be detected is identical to the specified Chinese medicine origin species, otherwise it is not.

[0038] In some embodiments, the nuclear, chloroplast and mitochondrial genome sequences of the traditional Chinese medicine-derived species are obtained by constructing a genome map or shallow sequencing.

[0039] In some embodiments, the nuclear, chloroplast, and mitochondrial genome sequences of the TCM-derived species are divided into L-K+1 fragments of length K to form the small-fragment genomic library. The copy number of each fragment is calculated, and the genomic position of each fragment is determined by alignment with the genome, where L represents the genomic sequence length and K represents the library fragment length. In some embodiments, K is 15-750 bp, e.g., 18-30 bp, 20-750 bp, or 25-28 bp.

[0040] In some embodiments, the Chinese herbal medicine origin species designated for the above method is Rheum palmatum, Rheum tanguticum, or Rheum officinale, but is not limited thereto. Herein, the Chinese herbal medicine origin species refers to the species to which the sample to be tested may belong. For example, for a sample to be tested suspected of being derived from rhubarb, Rheum palmatum, Rheum tanguticum, or Rheum officinale can be selected as the origin species. Furthermore, if one wishes to determine whether the sample to be tested is identical to Rheum officinale, then Rheum officinale is designated as the Chinese herbal medicine origin species.

[0041] In some embodiments, in step (2), the candidate target sequence is extracted according to the detection system for detecting the standard specific target sequence used in step (4). Those skilled in the art are familiar with the requirements of various sequence detection systems for the sequence to be tested. For example, for the CRISPR / Cas12a system, a motif with TTTV at the 5' end or VAAA (PAM) at the 3' segment can be selected (preferably, the PAM motif is detected for each fragment in the small fragment genomic library, and the candidate target sequence with PAM is extracted to construct a candidate target sequence library); for the TaqMan probe-based real-time PCR system, a motif that does not contain more than 3 consecutive repeated bases and 3 consecutive G bases can be selected; for the PCR system, a motif that does not contain GC-rich regions can be selected; or, for example, for the Sanger sequencing system, a motif suitable for sequencing can be selected, but the scope of protection of the present invention is not limited thereto.

[0042] In some embodiments, the specificity of the standard-specific target sequence can be further improved by increasing the number of nucleotide differences between the candidate sequence and the nuclear, chloroplast, and mitochondrial genome sequences of the fakes and closely related species; or the standard-specific target sequence can be screened within a predetermined number range by adjusting the number of nucleotide differences between the candidate sequence and the nuclear, chloroplast, and mitochondrial genome sequences of the fakes and closely related species.

[0043] In a preferred embodiment, the designated Chinese medicine-based species is Rheum officinale, the target sequence detection system is a CRISPR / Cas12a system, and the standard-specific target sequence includes the target nucleotide sequences shown in SEQ ID NOs: 1 and 4. In a preferred embodiment, the designated Chinese medicine-based species is Rheum officinale, the target sequence detection system is a CRISPR / Cas12a system, and the standard-specific target sequence includes the target nucleotide sequence shown in SEQ ID NO: 2. In a preferred embodiment, the designated Chinese medicine-based species is Rheum officinale, the target sequence detection system is a CRISPR / Cas12a system, and the standard-specific target sequence includes the target nucleotide sequence shown in SEQ ID NO: 3.

[0044] In a preferred embodiment, the designated Chinese medicinal herb species is Rheum officinale, the target sequence detection system is a TaqMan probe-based real-time PCR system, and the standard specific target sequence includes the target nucleotide sequence shown in SEQ ID NO: 17.

[0045] In a preferred embodiment, the designated Chinese medicinal origin species is Rheum officinale, the target sequence detection system is a PCR system, and the standard specific target sequence includes the target nucleotide sequence shown in SEQ ID NO: 19.

[0046] In a preferred embodiment, the designated Chinese medicinal origin species is Rheum officinale, the target sequence detection system is a Sanger sequencing system, and the standard specific target sequence includes the target nucleotide sequence shown in SEQ ID NO: 20.

[0047] In step (3), those skilled in the art are familiar with the requirements of various sequence detection systems for detection sequences, and based on the standard specific target sequence obtained by the present invention, they can use known molecular biology websites or tools to conventionally design and synthesize corresponding detection sequences. In some embodiments, for example, for the CRISPR / Cas12a system, CRISPR RNA (crRNA) can be designed and synthesized based on the standard specific target sequence; preferably, a crRNA sequence library matching the standard specific target sequence library can be constructed by constructing a standard specific target sequence library of a specified traditional Chinese medicine-based species relative to its mixed products and closely related species. For example, for the TaqMan probe-based real-time PCR system, primers and fluorescent probes can be designed and synthesized based on the standard specific target sequence and the region where the standard specific target sequence is located; for example, for the PCR system, amplification primers can be designed and synthesized based on the standard specific target sequence; for example, for the Sanger sequencing system, sequencing primers can be synthesized based on the selected standard specific target sequence. It will be understood by those skilled in the art that the target sequence detection system is not limited to those exemplified above, but can be based on any detection system known in the art, and the corresponding detection sequence can be designed thereby.

[0048] In a preferred embodiment, the CRISPR / Cas12a system is used to detect whether the standard specific target sequence is present in the DNA substrate using crRNA shown in SEQ ID NOs: 5, 6, 7 or 8.

[0049] In a preferred embodiment, Taqman probe-based real-time PCR is used with a primer pair represented by SEQ ID NOs: 18 and 16 and a Taqman probe represented by FAM-GCTTGAATGAAAGTCAGGCA CTCCGCCA-BHQ to detect whether the standard specific target sequence is present in the DNA substrate.

[0050] In a preferred embodiment, a PCR system is used with the primer pair shown as SEQ ID NOs: 19 and 16 to detect the presence of the standard specific target sequence in the DNA substrate.

[0051] In a preferred embodiment, the sequencing primer pair represented by SEQ ID NOs: 19 and 21 is used in a Sanger sequencing system to detect whether the standard specific target sequence is present in the DNA substrate.

[0052] In some embodiments, in step (4), for example, a primer pair that specifically amplifies a standard-specific target sequence can be used to amplify the genomic DNA of the traditional Chinese medicine to be tested and the amplified standard-specific target sequence can be recovered as a DNA substrate; or a primer pair that specifically amplifies a DNA sequence containing a standard-specific target sequence can be used to amplify the genomic DNA to be tested and the amplified DNA sequence containing the standard-specific target sequence can be recovered as a DNA substrate.

[0053] For example, when the CRISPR / Cas12a system is selected as the target sequence detection system, the primer pairs shown in SEQ ID NOs: 9-10, SEQ ID NOs: 11-12, SEQ ID NOs: 13-14 or SEQ ID NOs: 15-16 can be used to amplify the genomic DNA of the traditional Chinese medicine to be detected and recover the amplified standard specific target sequence as a DNA substrate.

[0054] In some preferred embodiments, the CRISPR / Cas12a system includes: gene editing buffer, Cas12a, crRNA, nuclease-free water, DNA substrate and fluorescent signal molecule (such as ssDNA reporter gene).

[0055] As an example, for the CRISPR / Cas12a system, NEBuffer 2.1 and Lba Cas12a (Cpf1) can be selected, and the fluorescent signal molecule can be selected as Poly_C_FQ (5'-FAM-CCCCCCCCCC-BHQ-3'), and the reaction is carried out as follows:

[0056] (1) Prepare the following reaction system (final concentrations are in brackets):

[0057]

[0058] (2) Incubate at room temperature for 10 minutes.

[0059] (3) The standard specific target sequence recovered after amplification was used as the DNA substrate. 10 μL of the standard specific target sequence recovered after amplification (1 ng / μL) and 4 μL of Poly_C_FQ (400 nM) were added and incubated at 37°C. The samples were marked with a microplate reader at λ at 0, 3, 6, 9, 12, 15, 25, 35, and 45 minutes. ex 483nm / λ em The fluorescence value is detected at 535 nm (determined by the selected fluorescent signal molecule). If the test result is significantly different from the blank control (P < 0.01), it can be determined that the sample to be tested is identical to the specified species, otherwise it is not identical.

[0060] In some preferred embodiments, the Taqman probe-based real-time PCR system comprises: a buffer, primers, nuclease-free water, a DNA substrate, and a Taqman probe.

[0061] As an example, for a Taqman probe-based real-time PCR system, select Probe qPCRMix, select FAM as the Taqman probe fluorophore, and select BHQ as the quencher. Perform the reaction as follows:

[0062] (1) Prepare the following reaction system (final concentrations are in brackets):

[0063]

[0064] (2) The reaction conditions were 95°C for 30 seconds; 40 cycles of: 95°C for 5 seconds; 60°C for 30 seconds.

[0065] (3) If the CT value of the test result is greater than 37, it can be determined that the sample to be tested is identical to the specified species; otherwise, it is not identical.

[0066] In some preferred embodiments, the PCR system comprises: a buffer, forward / reverse primers, nuclease-free water, and a DNA substrate.

[0067] As an example, for the PCR system, Taq PCR MasterMix can be selected and the reaction can be performed as follows:

[0068] (1) Prepare the following reaction system (final concentrations are in brackets):

[0069]

[0070] (2) The reaction conditions were 95°C for 30 seconds; 30 cycles of 95°C for 5 seconds; T m30sec; 72℃ 30sec; 72℃ 10min.

[0071] (3) The PCR products were electrophoresed on a 2% agarose gel at 120 V for 30 min. If a specific band was found in the electrophoresis results, it was determined that the species being tested was identical to the designated species; otherwise, they were not identical.

[0072] In some preferred embodiments, the Sanger sequencing system comprises: a buffer, sequencing primers, nuclease-free water and a DNA substrate.

[0073] As an example, in the case of the Sanger sequencing system, Taq PCR MasterMix can be selected and the reaction can be performed as follows:

[0074] (1) Prepare the following reaction system (final concentrations are in brackets):

[0075]

[0076] (2) The reaction conditions were 95°C for 30 seconds; 30 cycles of 95°C for 5 seconds; T m 30sec; 72℃ 30sec; 72℃ 10min.

[0077] (3) The PCR products were subjected to Sanger sequencing using sequencing primers.

[0078] (4) Analyze the sequencing results. If the sequencing results contain the standard specific target sequence, it can be determined that the species to be detected is identical to the specified species. Otherwise, it is not identical.

[0079] In one aspect, the present invention provides a standard specific target nucleotide for rhubarb, wherein the target nucleotide is at least one selected from the following: (1) a target nucleotide (AATATGGTTATGTTATATTAATAAA) shown in SEQ ID NO: 1; (2) a target nucleotide (TTTATATTGATTGTTTTATATTGAT) shown in SEQ ID NO: 2; (3) a target nucleotide (TTTCGCCAGTATCATTATTATTTAATTT); (4) a target nucleotide (TTTCTTGTGGCGGAGTGCCTGACTT) shown in SEQ ID NO: 4; (5) a target nucleotide (GCTTGAATGAAAGTCAGGCACTCCGCCA) shown in SEQ ID NO: 17; (6) a target nucleotide (AAGCTGGCTGTCATTCAGCT) shown in SEQ ID NO: 19; or (7) a target nucleotide (ATAACCTGCATTCTATGGTTTGGTT) shown in SEQ ID NO: 20.

[0080] In another aspect, the present invention provides the use of the above-mentioned standard specific target nucleotides in species identification of rhubarb or materials derived from rhubarb, or in distinguishing rhubarb from its closely related species.

[0081] The present invention first discovered the presence of the species-specific standard target nucleotide sequences listed above in rhubarb. The target nucleotide sequences shown in SEQ ID NOs: 1, 4, 17, 19, and 20 are found in Rheum officinale, the target nucleotide sequence shown in SEQ ID NO: 2 is found in Rheum palmatum, and the target nucleotide sequence shown in SEQ ID NO: 3 is found in Rheum tanguticum. In some embodiments, the rhubarb can be selected from Rheum palmatum, Rheum tanguticum, or Rheum officinale.

[0082] In one aspect, the present invention provides a primer pair for species identification of rhubarb or material derived from rhubarb, comprising:

[0083] (1) Ro_cp_1F: CCAAATTGCCCGAAGCCTATG (SEQ ID NO: 9), and Ro_cp_1R: ATCGCTTTCCGACCCACAAT (SEQ ID NO: 10);

[0084] (2) Rp_cp_1F: GTTTAGGCGGTACGTACATAGA (SEQ ID NO: 11), and Rp_cp_1R: GATCTCAGTAAGAAGGGTTTACGA (SEQ ID NO: 12);

[0085] (3) Rt_cp_1F: CGCTTTCGCCAGTATCATTAT (SEQ ID NO: 13); and Rt_cp_1R: CCATTCCACAAAGGGATCC (SEQ ID NO: 14);

[0086] (4) Ro_wg_1F: ATGGCGAGAGAGGTGTTCCTAAA (SEQ ID NO: 15); and Ro_wg_1R: GTTGTGAATCCGACACGACCAATAT (SEQ ID NO: 16);

[0087] (5) Ro_wg_2F: GCTGTCATTCAGCTGTTCTCTGT (SEQ ID NO: 18); and Ro_wg_2R: GTTGTGAATCCGACACGACCAATAT (SEQ ID NO: 16);

[0088] (6) Ro_wg_3F: AAGCTGGCTGTCATTCAGCT (SEQ ID NO: 19); and Ro_wg_3R: GTTGTGAATCCGACACGACCAATAT (SEQ ID NO: 16); or

[0089] (7) Ro_wg_4F: AAGCTGGCTGTCATTCAGCT (SEQ ID NO: 19); and Ro_wg_4R: ATATTGGTCGTGTCGGATTCACAAC (SEQ ID NO: 21).

[0090] The primer pairs of the present invention can be used to amplify species-specific target sequences of rhubarb, thereby enabling species identification of rhubarb. For example, primer pairs represented by SEQ ID NOs: 9-10, 15-16, 18 and 16, 19 and 16, and 19 and 21 can be preferably used for species identification of materials derived from medicinal rhubarb; primer pairs represented by SEQ ID NOs: 11-12 can be preferably used for species identification of materials derived from Rheum palmatum; and primer pairs represented by SEQ ID NOs: 13-14 can be preferably used for species identification of materials derived from Rheum tanguticum. In some embodiments, the rhubarb can be selected from Rheum palmatum, Rheum tanguticum, or Rheum gracile.

[0091] In another aspect, the present invention relates to a kit for species identification of rhubarb or material derived from rhubarb, wherein the kit comprises the above-mentioned primer pair.

[0092] In some embodiments, the kit further comprises PCR reaction reagents. In some preferred embodiments, the PCR reaction reagents comprise: PCR amplification buffer, dNTPs, Taq DNA polymerase, MgCl2 and sterile ultrapure water.

[0093] In some embodiments, the kit further comprises at least one selected from the following: a reaction reagent of a TaqMan probe system, a reaction reagent of a CRISPR / Cas12a system.

[0094] In some preferred embodiments, the reaction reagents of the Taqman probe system include: DNA polymerase, dNTPs, buffer, MgCl2, nuclease-free water, and a Taqman probe. In a specific embodiment, the fluorescent group of the Taqman probe includes FAM, and the quencher group includes BHQ. In a specific embodiment, the Taqman probe is represented by FAM-GCTTGAATGAAAGTCAGGCACTCCGCCA-BHQ.

[0095] In some preferred embodiments, the CRISPR / Cas12a system reaction reagents include: gene editing buffer, Cas protein, crRNA, nuclease-free water and fluorescent signal molecules (such as ssDNA fluorescent reporter gene). In a specific embodiment, the Cas12a is Lba Cas12a (Cpf1); the fluorescent signal molecule includes Poly_C_FQ (5'-FAM-CCCCCCCCCC-BHQ-3'). In a specific embodiment, the crRNA can be any one of the crRNAs shown in SEQ ID NOs: 5, 6, 7 or 8.

[0096] On the other hand, the present invention provides the use of the above-mentioned primer pair or kit in the following aspects: identifying the rhubarb component in a sample to be tested, performing species identification on rhubarb or materials derived from rhubarb, distinguishing Chinese medicinal materials derived from rhubarb from their mixed products, or for safety testing of food, medicine or health products.

[0097] In some embodiments, the sample to be tested can be rhubarb, rhubarb tissue or organ (such as rhubarb root or stem), a Chinese medicinal material containing rhubarb, or rhubarb counterfeit.

[0098] In the present invention, the genomic DNA of the sample to be tested (for example, rhubarb, rhubarb tissue or organ, adulterated products, Chinese medicinal materials, decoction pieces, Chinese patent medicines, foods, and health products) is obtained, and whether the genomic DNA contains the standard-specific target sequence of a specified species (for example, medicinal rhubarb) is detected, thereby determining the identity of the sample to be tested with the specified species.

[0099] Herein, the sample to be tested may be a sample from rhubarb. In some exemplary embodiments, the rhubarb sample to be tested may be a sample from Rheum officinale, Rheum palmatum, or Rheum tanguticum. In some embodiments, the rhubarb sample to be tested is a sample from a single species or a mixture of samples from multiple species.

[0100] Herein, exemplary embodiments of the present invention are described in the following numbered paragraphs:

[0101] 1. A method for identifying traditional Chinese medicine based on Shizhen method, comprising:

[0102] (1) Obtain the nuclear, chloroplast, and mitochondrial genome sequences of the original species of traditional Chinese medicine and construct a small fragment genome library;

[0103] (2) extracting candidate target sequences from the small fragment genomic library and analyzing them, and screening candidate sequences that meet at least one of the screening conditions (a) to (d) from the candidate target sequences as standard specific target sequences, wherein the screening conditions include:

[0104] (a) The GC content of the candidate sequence is 40%-60%;

[0105] The candidate sequence cannot contain four or more nucleotide repeats, consecutive trinucleotide repeats, or discontinuous three or more trinucleotide repeats;

[0106] The candidate sequence cannot be complementary to the crRNA repeat sequence;

[0107] The G+C content of the 6 nucleotides at the 5' end of the candidate sequence is 30%-80%;

[0108] The candidate sequence contains at least three different nucleotides when compared with the nuclear, chloroplast, and mitochondrial genome sequences of the admixture and closely related species; and / or

[0109] There are no four consecutive sequences within the region from -50bp to +300bp or from -300bp to +50bp where the candidate sequence is located.

[0110] Repeat on A or GT;

[0111] (b) the G+C content of the candidate sequence and the primers that amplify the region where the candidate sequence is located is 30%-80%;

[0112] The annealing temperatures of the fluorescent probes designed based on the candidate sequences and the primers for the upstream and downstream amplification regions of the candidate sequences are 55°C-60°C and 68°C-70°C, respectively, and the annealing temperature difference between the primer pairs does not exceed 2°C;

[0113] The length of the candidate sequence and the primers that amplify the region where the candidate sequence is located is 15-30 bp;

[0114] The primers for the upstream and downstream amplification regions where the candidate sequence is located are spaced 50-150 bp apart and the forward primer is close to the fluorescent probe designed based on the candidate sequence;

[0115] The fluorescent probe designed based on the candidate sequence and the primers for the upstream and downstream amplification regions where the candidate sequence is located do not contain a hairpin structure and the primer pair cannot form dimers;

[0116] There should be no more than two G / C bases within 5 bp of the 3' end of the primers in the region where the upstream and downstream candidate sequences are amplified;

[0117] The fluorescent probe designed based on the candidate sequence contains more C bases than G bases; and / or

[0118] Candidate sequences were compared with nuclear, chloroplast and mitochondrial genomes of at least

[0119] Contains 2 differential nucleotides;

[0120] (c) The GC content of the amplification primers designed based on the candidate sequence is 40%-60%, and the four nucleotides are evenly distributed, without polypurines and polypyrimidines, and without GC-rich regions;

[0121] The length of the amplification primers designed based on the candidate sequence is 18-30 bp, and the difference between primer pairs does not exceed 3 bases;

[0122] The amplification primers designed based on the candidate sequence do not contain inverted repeat sequences and self-complementary sequences greater than 3 bp, and the primers cannot form dimers;

[0123] The annealing temperature of the amplification primers designed based on the candidate sequence is 55-60°C, and the annealing temperature difference between primer pairs does not exceed 2°C;

[0124] The amplification primer designed based on the candidate sequence contains more than one but no more than three G / C bases in the 5 bases at the 3' end; and / or

[0125] The candidate sequence contains at least two different nucleotides when compared with the nuclear, chloroplast, and mitochondrial genomes of the admixture and closely related species;

[0126] or, (d) the region where the candidate sequence is located is homozygous;

[0127] The GC content of the region where the candidate sequence is located is 30%-80%;

[0128] The region where the candidate sequence is located cannot contain more than four consecutive repeated nucleotides;

[0129] The region where the candidate sequence is located cannot contain methylated sequences;

[0130] There should be no hairpin structure in the region where the candidate sequence is located; and / or

[0131] The candidate sequence contains at least two different nucleotides when compared with the nuclear, chloroplast, and mitochondrial genomes of the admixture and closely related species;

[0132] (3) extracting genomic DNA from the Chinese medicine sample to be tested, optionally amplifying it, using the genomic DNA or its amplified product as a DNA substrate, and using a target sequence detection system to detect whether the standard specific target sequence exists in the DNA substrate, wherein the target sequence detection system includes a CRISPR / Cas12a system, a TaqMan probe-based real-time PCR system, a PCR system or a Sanger sequencing system, wherein, for the CRISPR / Cas12a system, if an obvious fluorescent signal is generated by the detection, the Chinese medicine sample to be tested is identical to the specified Chinese medicine origin species, otherwise it is not; for the TaqMan probe-based real-time For the PCR system, if the CT value is greater than 37 through the detection, the Chinese medicine sample to be detected is identical to the specified Chinese medicine origin species, otherwise it is not; for the PCR system, if the electrophoresis result produces obvious bands through the detection, the Chinese medicine sample to be detected is identical to the specified Chinese medicine origin species, otherwise it is not; for the Sanger sequencing system, if the sequencing result is the same as the standard specific target sequence, the Chinese medicine sample to be detected is identical to the specified Chinese medicine origin species, otherwise it is not.

[0133] 2. The method of paragraph 1, wherein the nuclear, chloroplast and mitochondrial genome sequences of the original species of traditional Chinese medicine are obtained by constructing a genome map or shallow sequencing.

[0134] 3. The method of paragraph 1 or 2, wherein the nuclear, chloroplast, and mitochondrial genome sequences of the Chinese medicine-based species are divided into L-K+1 fragments of length K to form the small fragment genomic library, and the copy number of each fragment is calculated, and the genomic position of each fragment is determined by comparison with the genome, wherein L represents the length of the genome sequence and K represents the length of the library fragment, and K is 15-750 bp.

[0135] 4. The method as described in any one of paragraphs 1 to 3, wherein the Chinese medicinal base species is Rheum palmatum, Rheum tanguticum or Rheum officinale.

[0136] 5. The method of any of paragraphs 1 to 4, wherein the Chinese medicinal herb is Rheum officinale, the target sequence detection system is a CRISPR / Cas12a system, and the standard specific target sequence comprises the target nucleotide sequences shown in SEQ ID NOs: 1 and 4;

[0137] The Chinese medicinal material is rhubarb, the target sequence detection system is a TaqMan probe-based real-time PCR system, and the standard specific target sequence includes the target nucleotide sequence shown in SEQ ID NO: 17;

[0138] The Chinese medicinal material is Rheum officinale, the target sequence detection system is a PCR system, and the standard specific target sequence includes the target nucleotide sequence shown in SEQ ID NO: 19; or

[0139] The traditional Chinese medicine base species is medicinal rhubarb, the target sequence detection system is a Sanger sequencing system, and the standard specific target sequence includes a target nucleotide sequence shown in SEQ ID NO: 20.

[0140] 6. The method of any of paragraphs 1-4, wherein the Chinese medicinal base species is Rheum palmatum, the target sequence detection system is a CRISPR / Cas12a system, and the standard specific target sequence includes the target nucleotide sequence shown in SEQ ID NO: 2.

[0141] 7. The method of any of paragraphs 1-4, wherein the Chinese medicinal origin species is Rheum tanguticum, the target sequence detection system is a CRISPR / Cas12a system, and the standard specific target sequence includes a target nucleotide sequence shown in SEQ ID NO: 3.

[0142] 8. The method of any one of paragraphs 1-7, wherein the CRISPR / Cas12a system is used to detect the presence of the standard-specific target sequence in the DNA substrate using crRNA represented by SEQ ID NOs: 5, 6, 7 or 8.

[0143] 9. The method of any of paragraphs 1 to 5, wherein the presence of the standard-specific target sequence in the DNA substrate is detected using a TaqMan probe-based real-time PCR system with a primer pair represented by SEQ ID NOs: 18 and 16 and a TaqMan probe represented by FAM-GCTTGAATGAAAGTCAGGCACTCCGCCA-BHQ.

[0144] 10. The method of any of paragraphs 1-5, wherein a PCR system is used to detect the presence of the standard-specific target sequence in the DNA substrate using the primer pair shown as SEQ ID NOs: 19 and 16.

[0145] 11. The method of any one of paragraphs 1 to 5, wherein the Sanger sequencing system is used to detect whether the standard-specific target sequence is present in the DNA substrate using the sequencing primer pair shown as SEQ ID NOs: 19 and 21.

[0146] 12. The method of any one of paragraphs 1 to 11, wherein, in step (4), a primer pair that specifically amplifies a standard-specific target sequence is used to amplify the genomic DNA of the Chinese medicine to be tested, and the amplified standard-specific target sequence is recovered as a DNA substrate; or a primer pair that specifically amplifies a DNA sequence containing the standard-specific target sequence is used to amplify the genomic DNA to be tested, and the amplified DNA sequence containing the standard-specific target sequence is recovered as a DNA substrate.

[0147] 13. The method of any one of paragraphs 1-8, wherein the CRISPR / Cas12a system comprises: a gene editing buffer, Cas12a, crRNA, nuclease-free water, a DNA substrate, and a fluorescent signal molecule.

[0148] 14. The method of any one of paragraphs 1-5 and 9, wherein the Taqman probe-based real-time PCR system comprises a buffer, primers, nuclease-free water, a DNA substrate, and a Taqman probe.

[0149] 15. The method of any of paragraphs 1-5 and 10, wherein the PCR system comprises a buffer, a forward primer, a reverse primer, nuclease-free water, and a DNA substrate.

[0150] 16. The method of any one of paragraphs 1-5 and 11, wherein the Sanger sequencing system comprises a buffer, a sequencing primer, nuclease-free water, and a DNA substrate.

[0151] 17. A standard specific target nucleotide for rhubarb, wherein the target nucleotide is at least one selected from the following: (1) a target nucleotide shown in SEQ ID NO: 1; (2) a target nucleotide shown in SEQ ID NO: 2; (3) a target nucleotide shown in SEQ ID NO: 3; (4) a target nucleotide shown in SEQ ID NO: 4; (5) a target nucleotide shown in SEQ ID NO: 17; (6) a target nucleotide shown in SEQ ID NO: 19; or (7) a target nucleotide shown in SEQ ID NO: 20.

[0152] 18. Use of the standard specific target nucleotide described in paragraph 17 for species identification of rhubarb or materials derived from rhubarb, or for distinguishing rhubarb from its closely related species.

[0153] 19. The use according to paragraph 18, wherein the rhubarb is selected from Rheum palmatum, Rheum tanguticum or Rheum officinale.

[0154] 20. A primer pair for species identification of rhubarb or material derived from rhubarb, comprising:

[0155] (1) SEQ ID NO: 9 and SEQ ID NO: 10;

[0156] (2) SEQ ID NO: 11 and SEQ ID NO: 12;

[0157] (3) SEQ ID NO: 13 and SEQ ID NO: 14;

[0158] (4) SEQ ID NO: 15 and SEQ ID NO: 16;

[0159] (5) SEQ ID NO: 18 and SEQ ID NO: 16;

[0160] (6) SEQ ID NO: 19 and SEQ ID NO: 16; or

[0161] (7) SEQ ID NO: 19 and SEQ ID NO: 21.

[0162] 21. The primer pair of paragraph 20, wherein the rhubarb is selected from Rheum palmatum, Rheum tanguticum, or Rheum officinale.

[0163] 22. A kit for species identification of rhubarb or material derived from rhubarb, wherein the kit comprises the primer pair of paragraph 20 or 21.

[0164] 23. The kit according to paragraph 22, further comprising a PCR reaction reagent.

[0165] 24. The kit according to paragraph 23, wherein the PCR reaction reagents comprise: PCR amplification buffer, dNTPs, Taq DNA polymerase, MgCl2 and sterile ultrapure water.

[0166] 25. The kit of any one of paragraphs 22-24, further comprising at least one selected from the group consisting of a TaqMan probe system reaction reagent and a CRISPR / Cas12a system reaction reagent.

[0167] 26. The kit according to paragraph 25, wherein the reaction reagents of the Taqman probe system comprise: DNA polymerase, dNTPs, buffer, MgCl2, nuclease-free water and Taqman probe.

[0168] 27. The kit of paragraph 26, wherein the Taqman probe is represented by FAM-GCTTGAATGAAAGTCAGGCACTCCGCCA-BHQ.

[0169] 28. The kit of paragraph 25, wherein the CRISPR / Cas12a system reaction reagents comprise: gene editing buffer, Cas protein, crRNA, nuclease-free water, and fluorescent signal molecules.

[0170] 29. The kit of paragraph 28, wherein the crRNA is any one of the crRNAs shown in SEQ ID NOs: 5, 6, 7, or 8.

[0171] 30. Use of the primer pair described in paragraph 20 or 21 or the kit described in any one of paragraphs 22-29 for identifying rhubarb components in a sample to be tested, for species identification of rhubarb or materials derived from rhubarb, for distinguishing Chinese medicinal materials derived from rhubarb from their adulterants, or for safety testing of foods, drugs, or health products.

[0172] 31. The use according to paragraph 30, wherein the sample to be tested is rhubarb, rhubarb tissue or organ, a Chinese medicinal material containing rhubarb, or a rhubarb counterfeit.

[0173] 32. The use according to paragraph 30 or 31, wherein the rhubarb sample to be tested is a sample from Rheum officinale, Rheum palmatum, or Rheum tanguticum.

[0174] 33. The use according to any one of paragraphs 30 to 32, wherein the rhubarb sample to be tested is a sample originating from a single species or a mixture of samples originating from multiple species.

[0175] The following will further illustrate the exemplary embodiments of the present invention in conjunction with the accompanying drawings to fully illustrate the purpose, technical features and technical effects of the present invention. The following drawings are only examples of the present disclosure and the scope of protection of the present invention is not limited thereto. BRIEF DESCRIPTION OF THE DRAWINGS

[0176] Figure 1 is a schematic flow chart of the Shizhen method disclosed herein;

[0177] Figure 2 The results of the exemplary CRISPR / Cas12a system-based Shizhen method disclosed herein applied to the identification of medicinal rhubarb (using Ro_cp_target1 as the standard specific target sequence);

[0178] Figure 3 The results of applying the exemplary CRISPR / Cas12a system-based Shizhen method disclosed herein to the identification of Rheum palmatum are shown in FIG.

[0179] Figure 4 The results of applying the exemplary CRISPR / Cas12a system-based Shizhen method disclosed herein to the identification of Rheum tanguticum L.

[0180] Figure 5 The results of the exemplary CRISPR / Cas12a system-based Shizhen method disclosed herein applied to the identification of medicinal rhubarb (using Row_wg_target1 as the standard specific target sequence);

[0181] Figure 6 The results of applying the exemplary TaqMan probe-based real-time PCR system of the present disclosure to the identification of medicinal rhubarb are shown in FIG.

[0182] Figure 7 The results of applying the exemplary PCR-based Shizhen method disclosed herein to identify medicinal rhubarb are shown in FIG.

[0183] Figure 8 The results of applying the exemplary Shizhen method disclosed herein to the identification of the rhubarb medicinal material to be tested;

[0184] Figure 9 The result of applying the exemplary Shizhen method disclosed herein to the identification of the rhubarb slices to be tested;

[0185] Figure 10 The following is the result of applying the exemplary Shizhen method disclosed in the present invention to the identification of Chinese patent medicine containing rhubarb.

[0186] Figure 11The results of the Ro_cp_medicine, Rp_cp_medicine, and Rt_cp_medicine groups are shown when the Shizhen method is applied to the detection of Chinese patent medicines containing rhubarb. DETAILED DESCRIPTION

[0187] Figure 1 The schematic flow chart of the Shizhen method of the present application is shown. The identification process of rhubarb Chinese medicine is used as a specific implementation example to further illustrate the identification method of Chinese medicine based on the Shizhen method of the present disclosure. The experimental methods without specific conditions in the following examples are all implemented according to conventional conditions.

[0188] Example 1: Identification of Rhubarb Plants Using the CRISPR / Cas12a System

[0189] The 2020 edition of the Chinese Pharmacopoeia defines rhubarb as the dried roots and rhizomes of Rheum palmatum, Rheum tanguticum, or Rheum officinale, all belonging to the Polygonaceae family. Rhubarb is a traditional and commonly used Chinese medicinal herb with the properties of purging, clearing away heat and purging fire, cooling the blood and detoxifying, removing blood stasis and promoting menstruation, and promoting the removal of dampness and jaundice. Chemical analysis has shown differences in the active ingredient content of the three rhubarb varieties. Distinguishing these three varieties is essential for ensuring safe and effective clinical use. However, commonly used molecular identification methods, such as DNA barcoding, cannot distinguish between the three parent species of rhubarb.

[0190] In this example, the CRISPR / Cas12a system-based Shizhen method was used to identify rhubarb plants according to the following procedures:

[0191] (1) The nuclear, chloroplast, and mitochondrial genome sequences of the three rhubarb species were obtained by shallow sequencing.

[0192] (2) The nuclear, chloroplast, and mitochondrial genome sequences (L = genome sequence length) of three rhubarb species were selected and divided into (L-25+1) sequences of 25 bp in length or (L-28+1) sequences of 28 bp in length using Jellyfish (v1.1.12) to construct small fragment genomic libraries.

[0193] (3) Sequences with PAMs were extracted from a small fragment genomic library of rhubarb (in this example, the CRISPR / Cas12a system was used as the target sequence detection system, and sequences with TTTV at the 5' end or VAAA (PAM) at the 3' end were selected); and standard specific target sequences for identifying rhubarb plants were selected using Bowtie (v1.1.0) according to the following screening principles:

[0194] The GC content of the candidate sequence is 40%-60%; the candidate sequence cannot contain tetranucleotide or more repeat sequences, continuous trinucleotide repeat sequences, or discontinuous 3 or more trinucleotide repeat sequences; the candidate sequence cannot be complementary to the crRNA repeat sequence (5'-UAAUUUCUACUAAGUGUAGAU-3'; SEQ ID NO: 22); the G+C content of the 6 nucleotides at the 5' end of the candidate sequence is 30%-80%; the candidate sequence contains at least 3 different nucleotides when compared with the nuclear, chloroplast and mitochondrial genome sequences of fake products and closely related species; there are no more than 4 consecutive A or GT repeats in the region from -50bp to +300bp or from -300bp to +50bp where the candidate sequence is located.

[0195] To demonstrate the versatility of the Shizhen method, this example screened a standard specific target sequence from the chloroplast genomes of three rhubarb species, named Ro_cp_target1 (rhea), Rp_cp_target1 (rhea palmata), and Rt_cp_target1 (rhea tangutica). A standard specific target sequence was screened from the nuclear genome of rhubarb, named Ro_wg_target1. The sequence is as follows (5'→3'): Ro_cp_target1: AATATGGTTATGTTATATTAATAAA (SEQ ID NO: 1)

[0196] Rp_cp_target1: TTTATATTGATTGTTTTATATTGAT (SEQ ID NO: 2)

[0197] Rt_cp_target1:TTTCCGCCAGTATCATTATTATTTAATTT(SEQ ID NO:3) Ro_wg_target1:TTTCTTGTGGCGGAGTGCCTGACTT(SEQ ID NO:4)

[0198] Based on the selected CRISPR / Cas system and crRNA design principles, crRNAs matching the standard specific target sequences were designed. The crRNAs corresponding to the standard specific target sequences selected from the chloroplast genomes of three rhubarb species were named Ro_cp_crRNA1, Rp_cp_crRNA1, and Rt_cp_crRNA1, respectively. The crRNA corresponding to the standard specific target sequence selected from the nuclear genome of medicinal rhubarb was named Ro_wg_crRNA1, and the sequence is as follows (5'→3'):

[0199] Ro_cp_crRNA1: UAAUUUCUACUAAGUGUAGAUUUAAUAUAACAUAACCAUAUU (SEQ ID NO: 5)

[0200] Rp_cp_crRNA1: UAAUUUCUACUAAGUGUAGAUUAUUGAUUGUUUUAUAUUGAU (SEQ ID NO: 6)

[0201] Rt_cp_crRNA1: UAAUUUCUACUAAGUGUAGAUGCCAGUAUCAUUAUUAUUUAAUUU (SEQ ID NO: 7)

[0202] Ro_wg_crRNA1: UAAUUUCUACUAAGUGUAGAUUUGUGGCCGGAGUGCCUGACUU (SEQ ID NO: 8)

[0203] (4) Amplify and purify the standard specific target sequence. Rhubarb was collected from Mianning, Sichuan and numbered Y6. Rhubarb palmatum was collected from Dangchang, Gansu and numbered Z5. Rhubarb tanguticus was collected from Hongya, Sichuan and numbered T1. The collected plant samples were dried and crushed with a ball mill, and then genomic DNA was extracted according to the instructions of the Plant Genomic DNA Kit of TIANGEN. The integrity of the genomic DNA was detected by 0.8% agarose gel electrophoresis, and then its purity and concentration were detected by Nanodrop2000C spectrophotometer. Primers were designed according to the region where the standard specific target sequence was located to amplify the standard specific target sequence. The primers were named Ro_cp_1F, Ro_cp_1R; Rp_cp_1F, Rp_cp_1R; Rt_cp_1F, Rt_cp_1R; Ro_wg_1F, Ro_wg_1R, and their sequences are as follows (5'→3'):

[0204] Ro_cp_1F: CCAAATTGCCCGAAGCCTATG (SEQ ID NO: 9)

[0205] Ro_cp_1R: ATCGCTTTCCGACCCACAAT (SEQ ID NO: 10)

[0206] Rp_cp_1F:GTTTAGGCGGTACGTACATAGA (SEQ ID NO: 11)

[0207] Rp_cp_1R:GATCTCAGTAAGAAGGGTTTACGA (SEQ ID NO: 12)

[0208] Rt_cp_1F: CGCTTTCGCCAGTATCATTAT (SEQ ID NO: 13)

[0209] Rt_cp_1R: CCATTCCACAAAGGGATCC (SEQ ID NO: 14)

[0210] Ro_wg_1F:ATGGCGAGAGAGGTTCCTAAA (SEQ ID NO: 15)

[0211] Ro_wg_1R:GTTGTGAATCCGACACGACCAATAT (SEQ ID NO: 16)

[0212] The total volume of the PCR reaction was 50 μL: 25 μL 2× Taq MasterMix, 2 μL primers (F / R) (400 nM), 2 μL total DNA sample, and nuclease-free water was added to 50 μL. The PCR reaction conditions were: 95°C for 30 seconds; 35 cycles of 95°C for 5 seconds; T m 30 seconds; 72°C for 2 minutes; 72°C for 10 minutes; and stored at 10°C. PCR products were recovered and purified using the TIANGEN Universal DNA Purification Kit according to the manufacturer's instructions. The integrity of the standard-specific target sequence was verified by 2% agarose gel electrophoresis, and the purity and concentration were determined using a Nanodrop 2000C spectrophotometer. The recovered fragments containing the standard-specific target sequence were used as DNA substrates in subsequent experiments.

[0213] (5) Identification of rhubarb plant samples by Shizhen method. Ro_cp_crRNA1, Rp_cp_crRNA1, Rt_cp_crRNA1 and Ro_wg_crRNA1 were used as crRNA, and the region fragments of Ro_cp_target1, Rp_cp_target1, Rt_cp_target1 and Ro_wg_target1 amplified above were used as DNA substrates to set up Ro_cp_plant, Rp_cp_plant, Rt_cp_plant and Ro_wg_plant groups respectively. EnGen Lba Cas12a (Cpf1) from NEB was used for the experiment. The total reaction volume was 100 μL: 10 μL 10×NEBuffer 2.1, 2 μL Lba Cas12a (20 nM), 3 μL crRNA (300 nM), 10 μL DNA substrate (1 ng / μL), 4 μL Poly_C_FQ (400 nM) and 71 μL nuclease-free water. NEBuffer 2.1, Lba Cas12a, crRNA and nuclease-free water were first added to the reaction system and incubated at room temperature for 30 minutes. Then, DNA substrate and Poly_C_FQ were added and incubated at 37°C. The reaction was detected by microplate reader at λ at 0, 3, 6, 9, 12, 15, 25, 35 and 45 minutes. ex 483nm / λ em Fluorescence was detected at 535 nm.

[0214] The results of the Ro_cp_plant, Rp_cp_plant, Rt_cp_plant, and Ro_wg_plant groups are shown in Figure 2 、 Figure 3 、 Figure 4 and Figure 5 . In the Ro_cp_plant group, using Ro_cp_crRNA1 as crRNA, only medicinal rhubarb produced a fluorescent signal, while Rheum palmatum and Rheum tanguticum produced no fluorescent signal. In the other three groups, only the combination of crRNA and the corresponding sample produced a fluorescent signal, which was significantly different from the CK group (blank control group, no DNA substrate was added, and the rest were the same as the above groups) (P>0.01). This result shows that the Shizhen method based on the CRISPR / Cas12a system can accurately and quickly identify the three rhubarb plants.

[0215] Example 2: Identification of Rhubarb Plants Using a TaqMan Probe-based Real-time PCR System

[0216] (1) The nuclear, chloroplast, and mitochondrial genome sequences of three rhubarb species were obtained through shallow sequencing.

[0217] (2) The nuclear, chloroplast, and mitochondrial genome sequences (L = genome sequence length) of three rhubarb strains were selected and divided into (L-28+1) sequences of 28 bp in length (the length is optional, 18-30 bp is recommended) using Jellyfish (v1.1.12) to construct small fragment genomic libraries.

[0218] (3) Sequences without more than three consecutive repeated bases and three consecutive G bases were extracted from the small fragment genomic library of Rheum officinale; and standard specific target sequences for identifying Rheum officinale plants were selected using Bowtie (v1.1.0) software according to the following screening principles:

[0219] The G+C content of the candidate sequence and the primers for amplifying the region where the candidate sequence is located is 30%-80%; the annealing temperatures of the fluorescent probe designed according to the candidate sequence and the primers for amplifying the regions where the upstream and downstream candidate sequences are located are 55°C-60°C and 68°C-70°C, respectively, and the annealing temperature difference between the primer pairs does not exceed 2°C; the length of the candidate sequence and the primers for amplifying the region where the candidate sequence is located is 15-30bp; the interval between the primers for amplifying the regions where the upstream and downstream candidate sequences are located is 50-150bp, and the forward primer is close to the fluorescent probe designed according to the candidate sequence; the fluorescent probe designed according to the candidate sequence and the primers for amplifying the regions where the upstream and downstream candidate sequences are located do not contain a hairpin structure and the primer pairs cannot form dimers; there are no more than two G / C bases within 5bp of the 3' end of the primers for amplifying the regions where the upstream and downstream candidate sequences are located; the C bases in the fluorescent probe designed according to the candidate sequence are more than the G bases; and the candidate sequence contains at least two different nucleotides when compared with the nuclear, chloroplast and mitochondrial genomes of the adulterated products and closely related species.

[0220] In this example, a standard specific target sequence was screened from the nuclear genome of Rhubarb (Euphrasia officinalis), named Ro_wg_target2. The sequence is as follows (5'→3'):

[0221] Ro_wg_target2:GCTTGAATGAAAGTCAGGCACTCCGCCA (SEQ ID NO: 17)

[0222] Based on the selected TaqMan probe-based real-time PCR system and fluorescent probe design principles, a fluorescent probe matching the standard specific target sequence was designed and named Ro_wg_probe2. The sequence is as follows (5'→3'):

[0223] Ro_wg_probe2:FAM-GCTTGAATGAAAGTCAGGCACTCCGCCA-BHQ

[0224] Amplification primers were designed based on the upstream and downstream fragments of the standard specific target sequence and named Ro_wg_2F and Ro_wg_2R, respectively. The sequences are as follows (5'→3'):

[0225] Ro_wg_2F:GCTGTCATTCAGCTGTTCTCTGT (SEQ ID NO: 18)

[0226] Ro_wg_2R:GTTGTGAATCCGACACGACCAATAT (SEQ ID NO: 16)

[0227] (4) The extraction and detection of rhubarb plant samples and plant DNA were the same as in Example 1.

[0228] (5) The Shizhen method was used to identify rhubarb plant samples. Ro_wg_probe2 was used as a fluorescent probe, Ro_wg_2F and Ro_wg_2R were used as primers, and three types of rhubarb genomic DNA were used as DNA substrates (i.e., total DNA samples). The Ro_wg_probe group was set. The experiment was performed using TaKaRa's Probe qPCR Mix. The total qPCR reaction volume was 20 μL: 10 μL 2× TaqMasterMix, 0.4 μL primers (F / R) (200 nM), 2 μL total DNA sample, 0.8 μL fluorescent probe, and nuclease-free water was used to make up to 20 μL. The PCR reaction conditions were: 95°C for 30 seconds; 35 cycles of: 95°C for 5 seconds; 60°C for 30 seconds; 72°C for 2 minutes; 72°C for 10 minutes; and storage at 10°C.

[0229] See the results Figure 6 Among the Rowg_probe group, only Rhubarb medica had a Ct value less than 37, showing a significant difference from the CK (blank control) (P>0.01). Rhubarb palmatum and Rhubarb tanguticus had Ct values ​​identical to the CK, both exceeding 37. This result demonstrates that the Shizhen method based on TaqMan probe-based real-time PCR can accurately and rapidly identify the three rhubarb species.

[0230] Example 3: Identification of Rhubarb Plants Using a PCR-Based Method

[0231] (1) The nuclear, chloroplast, and mitochondrial genome sequences of three rhubarb species were obtained through shallow sequencing.

[0232] (2) The nuclear, chloroplast, and mitochondrial genome sequences (L = genome sequence length) of three rhubarb strains were selected and divided into (L-20+1) sequences of 20 bp in length (the length is optional, 18-30 bp is recommended) using Jellyfish (v1.1.12) to construct small fragment genomic libraries.

[0233] (3) A candidate target sequence library for rhubarb was constructed, and sequences not rich in GC regions were extracted from a small fragment genomic library of rhubarb. Standard specific target sequences for identifying rhubarb plants were selected using Bowtie (v1.1.0) software according to the following screening principles:

[0234] The GC content of the amplification primers designed based on the candidate sequence is 40-60%, and the four nucleotides are evenly distributed (no polypurines and polypyrimidines, no GC-rich regions); the length of the amplification primers designed based on the candidate sequence is 18-30 bp, and the difference between primer pairs does not exceed 3 bases; the amplification primers designed based on the candidate sequence do not contain inverted repeat sequences and self-complementary sequences greater than 3 bp, and primers cannot form dimers; the annealing temperature of the amplification primers designed based on the candidate sequence is 55-60°C, and the difference in annealing temperature between primer pairs does not exceed 2°C; the G / C bases in the 5 bases at the 3' end of the amplification primers designed based on the candidate sequence are more than one but no more than three; the candidate sequence contains at least two different nucleotides when compared with the nuclear, chloroplast and mitochondrial genomes of the fake and closely related species.

[0235] In this example, a standard specific target sequence was screened from the nuclear genome of Rhubarb (Radix Rhei), named Ro_wg_target3. The sequence is as follows (5'→3'):

[0236] Ro_wg_target3: AAGCTGGCTGTCATTCAGCT (SEQ ID NO: 19)

[0237] Based on the selected PCR system and primer design principles, detection primers matching the standard specific target sequence were designed and named Ro_wg_3F and Ro_wg_3R, respectively. The sequences are as follows (5'→3'):

[0238] Ro_wg_3F: AAGCTGGCTGTCATTCAGCT (SEQ ID NO: 19)

[0239] Ro_wg_3R:GTTGTGAATCCGACACGACCAATAT (SEQ ID NO: 16)

[0240] (4) The extraction and detection of rhubarb plant samples and plant DNA were the same as in Example 1

[0241] (5) The Shizhen method was used to identify rhubarb plant samples. Using Ro_wg_3F and Ro_wg_3R as primers, and three rhubarb genomic DNAs as DNA substrates (i.e., total DNA samples), the Ro_wg_PCR group was set up. Taq MasterMix from Adelaide was used for the experiment. The total PCR reaction volume was 25 μL: 12.5 μL 2× Taq MasterMix, 0.5 μL primers (F / R) (200 nM), 2 μL total DNA sample, and nuclease-free water was used to make up to 25 μL. PCR reaction conditions were: 95°C for 30 seconds; 40 cycles of: 95°C for 5 seconds; 60°C for 30 seconds. PCR products were electrophoresed on a 2% agarose gel at 120 V for 30 minutes.

[0242] See the results Figure 7 In the Rowg_PCR group, only Rheum officinale produced a bright single band, while Rheum palmatum and Rheum tanguticum produced no bright single bands, similar to the CK. This result indicates that the Shizhen method based on the PCR system can accurately and quickly identify the three rhubarb plants.

[0243] Example 4: Identification of Rhubarb Plants Using the Sanger Sequencing System

[0244] (1) The nuclear, chloroplast, and mitochondrial genome sequences of three rhubarb species were obtained through shallow sequencing.

[0245] (2) The nuclear, chloroplast, and mitochondrial genome sequences (L = genome sequence length) of three rhubarb strains were selected and divided into (L-25+1) sequences of 25 bp in length (the length is optional, recommended to be 20-750 bp) using Jellyfish (v1.1.12) to construct small fragment genomic libraries.

[0246] (3) Since the Sanger sequencing system has low sequence requirements, standard specific target sequences for identifying rhubarb plants were selected directly based on the small fragment genomic library of rhubarb according to the following screening principles using Bowtie (v1.1.0) software:

[0247] The candidate sequence region must be homozygous; the GC content of the candidate sequence region must be around 50%, for example, 30%-80%, but not too high; the candidate sequence region must not contain more than four consecutive repeated nucleotides; the candidate sequence region must not contain methylated sequences; the candidate sequence region must not contain hairpin structures; and the candidate sequence must contain at least two different nucleotides when compared with the nuclear, chloroplast, and mitochondrial genomes of adulterated samples and closely related species.

[0248] In this example, a standard specific target sequence was screened from the nuclear genome of Rhubarb (Radix Rhei), named Ro_wg_target4. The sequence is as follows (5'→3'):

[0249] Ro_wg_target4: ATAACCTGCATTCTATGGTTTGGTT (SEQ ID NO: 20)

[0250] Based on the selected PCR system and primer design principles, sequencing primers matching the standard specific target sequence were designed and named Ro_wg_4F and Ro_wg_4R, respectively. The sequences are as follows (5'→3'):

[0251] Ro_wg_4F: AAGCTGGCTGTCATTCAGCT (SEQ ID NO: 19)

[0252] Ro_wg_4R: ATATTGGTCGTGTCGGATTCACAAC (SEQ ID NO: 21)

[0253] (4) The extraction and detection of rhubarb plant samples and plant DNA were the same as in Example 1.

[0254] (5) The Shizhen method was used to identify rhubarb plant samples. Using Ro_wg_4F and Ro_wg_4R as primers, three rhubarb genomic DNAs were used as DNA substrates (i.e., total DNA samples), and the Ro_wg_sequencing group was set up. TaqMasterMix from Adelaide was used for the experiment. The total PCR reaction volume was 25 μL: 12.5 μL 2× Taq MasterMix, 0.5 μL primer (F / R) (200 nM), 2 μL total DNA sample, and 25 μL of nuclease-free water. The PCR reaction conditions were: 95°C for 30 seconds; 40 cycles of 95°C for 5 seconds; 60°C for 30 seconds. The PCR products were spliced ​​and sequenced using the Sanger sequencing system.

[0255] See the results Figure 8 In the Rowg_sequencing group, only Rheum officinale contained the standard specific target sequence, while Rheum palmatum and Rheum tanguticum did not have sequences completely identical to the standard specific target sequence. This result shows that the Shizhen method based on the Sanger sequencing system can accurately and quickly identify the three rhubarb plants.

[0256] Example 5: Results of applying the Shizhen method to the identification of rhubarb medicinal materials

[0257] This example uses the CRISPR / Cas system as a detection method. The three rhubarb standard specific target sequences have been obtained in Example 1. This example directly uses Ro_cp_crRNA1, Rp_cp_crRNA1 and Rt_cp_crRNA1 as crRNA, and uses the following primer pairs as primers for the corresponding amplification standard specific target sequences: Ro_cp_1F, Ro_cp_1R; Rp_cp_1F, Rp_cp_1R; Rt_cp_1F and Rt_cp_1R.

[0258] The medicinal rhubarb to be tested was collected from Mianning, Sichuan, and numbered Y10, and the Tangut rhubarb was collected from Ruoergai, Sichuan, and numbered T12. The medicinal material samples to be tested were pulverized with a ball mill, and then the genomic DNA was extracted according to the instructions of the Plant Genomic DNA Kit provided by TIANGEN. The integrity of the genomic DNA was detected by 0.8% agarose gel electrophoresis, and then its purity and concentration were detected by Nanodrop 2000C spectrophotometer. The genomic DNA of the two obtained rhubarb medicinal materials was used as a substrate, and the above three pairs of primers were used to amplify the standard specific target sequence. The total volume of the PCR reaction was 50μL: 25μL 2×Taq MasterMix, 2μL primers (F / R) (200nM), 2μL total DNA sample, and nuclease-free water was added to make up to 50μL. The PCR reaction conditions were: 95℃30sec; 35 cycles: 95℃5sec; T m 30 sec; 72℃ 2 min; 72℃ 10 min; store at 10℃.

[0259] The PCR product was recovered and purified using the Universal DNA Purification Kit provided by TIANGEN according to the instructions. The integrity of the standard-specific target sequence was detected by 2% agarose gel electrophoresis, and then its purity and concentration were detected by Nanodrop2000C spectrophotometer. The recovered fragments of the region where the standard-specific target sequence was located were used as DNA substrates in subsequent experiments.

[0260] Ro_cp_crRNA1, Rp_cp_crRNA1, and Rt_cp_crRNA1 were used as crRNAs, and the region fragments of the standard specific target sequences recovered above were used as DNA substrates to set up Ro_cp_herb, Rp_cp_herb, and Rt_cp_herb groups, respectively. EnGen Lba Cas12a (Cpf1) from NEB was used for the experiment, and the total reaction volume was 100 μL: 10 μL 10×NEBuffer 2.1, 2 μL Lba Cas12a (20 nM), 3 μL crRNA (300 nM), 10 μL DNA substrate (1 ng / μL), 4 μL Poly_C_FQ (400 nM), and 71 μL nuclease-free water. NEBuffer 2.1, LbaCas12a, crRNA and nuclease-free water were first added to the reaction system and incubated at room temperature for 30 minutes. Then, DNA substrate and Poly_C_FQ were added, and the reaction was incubated at 37°C. The fluorescence was detected at λex 483nm / λem535nm using a microplate reader at 0, 3, 6, 9, 12, 15, 25, 35, and 45 minutes.

[0261] The results of Ro_cp_herb, Rp_cp_herb and Rt_cp_herb groups are shown in Figure 9 In the Ro_cp group, using Ro_cp_crRNA1 as crRNA, only medicinal rhubarb produced a fluorescent signal, while Tangut rhubarb did not. In the other two groups, only the combination of crRNA and the corresponding sample produced a fluorescent signal, which was significantly different from the CK group (P>0.01). This result shows that the Shizhen method based on the CRISPR / Cas12a system can accurately and quickly identify the three rhubarb medicinal materials.

[0262] Example 6: Results of applying Shizhen method to the identification of rhubarb slices

[0263] This example uses the CRISPR / Cas system as a detection method. The three rhubarb standard specific target sequences have been obtained in Example 1. This example directly uses Ro_cp_crRNA1, Rp_cp_crRNA1 and Rt_cp_crRNA1 as crRNA, and uses the following primer pairs as primers for the corresponding amplification standard specific target sequences: Ro_cp_1F, Ro_cp_1R; Rp_cp_1F, Rp_cp_1R; Rt_cp_1F and Rt_cp_1R.

[0264] Three packages of slices were purchased from Beijing, Anguo, Hebei, and Bozhou, Anhui, respectively. One slice was randomly selected from each of them as a sample, numbered YP1, YP2, and YP3, and the slice samples were pulverized with a ball mill, and then genomic DNA was extracted according to the instructions for use of the Plant Genomic DNA Kit provided by TIANGEN. The integrity of the genomic DNA was detected by 0.8% agarose gel electrophoresis, and then its purity and concentration were detected by Nanodrop 2000C spectrophotometer. The genomic DNA of the three rhubarb slices obtained was used as a substrate, and the above three pairs of primers were used to amplify the standard specific target sequence. The total volume of the PCR reaction was 50 μL: 25 μL 2×Taq MasterMix, 2 μL primer (F / R) (400nM), 2 μL total DNA sample, and nuclease-free water was added to make up to 50 μL. The PCR reaction conditions were: 95°C 30 sec; 35 cycles: 95°C 5 sec; T m 30 sec; 72℃ 2 min; 72℃ 10 min; store at 10℃.

[0265] The PCR product was recovered and purified using the Universal DNA Purification Kit provided by TIANGEN according to the instructions. The integrity of the target sequence was detected by 2% agarose gel electrophoresis, and then its purity and concentration were detected by Nanodrop2000C spectrophotometer. The recovered fragments in the region where the standard specific target sequence was located were used as DNA substrates in subsequent experiments.

[0266] Ro_cp_crRNA1, Rp_cp_crRNA1, and Rt_cp_crRNA1 were used as crRNAs, and the region fragments of the standard specific target sequences recovered above were used as DNA substrates to set up Ro_cp_decoction, Rp_cp_decoction, and Rt_cp_decoction groups, respectively. EnGen Lba Cas12a (Cpf1) from NEB was used for the experiment, and the total reaction volume was 100 μL: 10 μL 10×NEBuffer 2.1, 2 μL Lba Cas12a (20 nM), 3 μL crRNA (300 nM), 10 μL DNA substrate (1 ng / μL), 4 μL Poly_C_FQ (400 nM), and 71 μL nuclease-free water. NEBuffer 2.1, Lba Cas12a, crRNA and nuclease-free water were first added to the reaction system and incubated at room temperature for 30 minutes. Then, DNA substrate and Poly_C_FQ were added and incubated at 37°C. The reaction was detected by microplate reader at λ at 0, 3, 6, 9, 12, 15, 25, 35 and 45 minutes. ex 483nm / λ emFluorescence was detected at 535 nm.

[0267] The results of Ro_cp_decoction, Rp_cp_decoction and Rt_cp_decoction groups are shown in Figure 10 In the Ro_cp_decoction group, no obvious fluorescence signal was generated, indicating that none of the three pieces of rhubarb were medicinal rhubarb. In the Rp_cp_decoction group, pieces 1 and 2 (YP1 and YP2) produced fluorescence signals, indicating that pieces 1 and 2 were Rheum palmatum, and piece 3 (YP3) was not Rheum palmatum. In the Rt_cp_decoction group, only piece 3 produced a fluorescence signal, indicating that piece 3 was Rheum tanguticum, which was consistent with the results of the other two groups of experiments. This result shows that the Shizhen method based on the CRISPR / Cas12a system can accurately and quickly identify the three rhubarb pieces.

[0268] Example 7: Results of applying the Shizhen method to Chinese patent medicines containing rhubarb

[0269] This example uses the CRISPR / Cas system as a detection method. The three rhubarb standard specific target sequences have been obtained in Example 1. This example directly uses Ro_cp_crRNA1, Rp_cp_crRNA1 and Rt_cp_crRNA1 as crRNA, and uses the following primer pairs as primers for the corresponding amplification standard specific target sequences: Ro_cp_1F, Ro_cp_1R; Rp_cp_1F, Rp_cp_1R; Rt_cp_1F and Rt_cp_1R.

[0270] Three rhubarb-containing traditional Chinese medicines (TCMs) were selected: Daochi Pills, Maren Runchang Tablets, and Ruyi Jinhuang Powder. Daochi Pills were purchased from Beijing, while Maren Runchang Tablets and Ruyi Jinhuang Powder were purchased from Shijiazhuang, Hebei Province. 3-5g of each of Daochi Pills, Maren Runchang Tablets, and Ruyi Jinhuang Powder were pulverized in a ball mill and washed three times with nuclear isolation buffer. Genomic DNA was then extracted according to the instructions for the Plant Genomic DNA Kit provided by TIANGEN. Genomic DNA integrity was verified by 0.8% agarose gel electrophoresis, and purity and concentration were determined using a Nanodrop 2000C spectrophotometer. Genomic DNA from each of the three TCMs was used as substrate to amplify the standard specific target sequences using the three primer pairs described above. The total PCR reaction volume was 50μL: 25μL 2× TaqMasterMix, 2μL primers (F / R) (400nM), 2μL total DNA sample, and the total volume was made up to 50μL with nuclease-free water. PCR reaction conditions were: 95°C for 30 seconds; 35 cycles of 95°C for 5 seconds; T m 30 sec; 72℃ 2 min; 72℃ 10 min; store at 10℃.

[0271] The PCR product was recovered and purified using the Universal DNA Purification Kit provided by TIANGEN according to the instructions. The integrity of the standard-specific target sequence was detected by 2% agarose gel electrophoresis, and then its purity and concentration were detected by Nanodrop2000C spectrophotometer. The recovered fragments of the region where the standard-specific target sequence was located were used as DNA substrates in subsequent experiments.

[0272] Ro_cp_crRNA1, Rp_cp_crRNA1, and Rt_cp_crRNA were used as crRNAs, and the region fragments of the standard specific target sequences recovered above were used as DNA substrates to set up Ro_cp_medicine, Rp_cp_medicine, and Rt_cp_medicine groups, respectively. EnGen Lba Cas12a (Cpf1) from NEB was used for the experiment, and the total reaction volume was 100 μL: 10 μL 10×NEBuffer 2.1, 2 μL Lba Cas12a (20 nM), 3 μL crRNA (300 nM), 10 μL DNA substrate (1 ng / μL), 4 μL Poly_C_FQ (400 nM), and 71 μL nuclease-free water. NEBuffer 2.1, Lba Cas12a, crRNA and nuclease-free water were first added to the reaction system and incubated at room temperature for 30 minutes. Then, DNA substrate and Poly_C_FQ were added and incubated at 37°C. The reaction was detected by microplate reader at λ at 0, 3, 6, 9, 12, 15, 25, 5 and 45 minutes. ex 483nm / λ em Fluorescence was detected at 535 nm.

[0273] The results of the Ro_cp_medicine, Rp_cp_medicine, and Rt_cp_medicine groups are shown in Figure 11 In the Ro_cp_medicine group, Daochi Pills and Ruyi Jinhuang Powder produced fluorescence signals, indicating that Daochi Pills and Ruyi Jinhuang Powder contain medicinal rhubarb, while Maren Runchang Pian does not. In the Rp_cp_medicine group, Maren Runchang Pian produced fluorescence signals, indicating that it contains Rheum palmatum, while the other two medicinal pieces do not. In the Rt_cp_medicine group, there was no significant fluorescence signal, indicating that none of the three Chinese patent medicines contained Rheum tangutum, consistent with the results of the other two experimental groups. This result demonstrates that the Shizhen method based on the CRISPR / Cas12a system can accurately and rapidly identify Chinese patent medicines containing rhubarb.

Claims

1. A method for identifying rhubarb based on Shizhen method, comprising: (1) obtaining the nuclear, chloroplast and mitochondrial genome sequences of the original species of traditional Chinese medicine by constructing a genome map or shallow sequencing, dividing the nuclear, chloroplast and mitochondrial genome sequences of the original species of traditional Chinese medicine into L-K+1 fragments of length K to construct a small fragment genome library, and calculating the copy number of each fragment, and then determining the genomic position of each fragment by comparing with the genome, wherein L represents the genome sequence length, K represents the library fragment length, wherein K is 18-30 bp, and the original species of traditional Chinese medicine is medicinal rhubarb; (2) extracting candidate target sequences from the small fragment genomic library and analyzing them according to the Taqman probe-based real-time PCR system, and screening candidate sequences that meet the following screening conditions from the candidate target sequences as standard specific target sequences, wherein the screening conditions include: The G+C content of the candidate sequence and the primers amplifying the region where the candidate sequence is located is 30%-80%; The annealing temperatures of the fluorescent probes designed according to the candidate sequences and the primers for the upstream and downstream amplification candidate sequences are 55°C-60°C and 68°C-70°C, respectively, and the annealing temperature difference between the primer pairs does not exceed 2°C; The length of the candidate sequence and the primers that amplify the region where the candidate sequence is located is 15-30 bp; The primers for the upstream and downstream amplification regions where the candidate sequence is located are spaced 50-150 bp apart and the forward primer is close to the fluorescent probe designed based on the candidate sequence; The fluorescent probe designed according to the candidate sequence and the primers for upstream and downstream amplification of the candidate sequence region do not contain a hairpin structure and the primer pair cannot form a dimer; The primers in the region where the upstream and downstream amplification candidate sequences are located have no more than two G / C bases within 5 bp of the 3' end; The fluorescent probe designed based on the candidate sequence contains more C bases than G bases; and / or the candidate sequence contains at least 2 different nucleotides when compared with the nuclear, chloroplast and mitochondrial genomes of the admixture and closely related species; The standard specific target sequence consists of the target nucleotide sequence shown in SEQ ID NO: 17; (3) extracting genomic DNA from the Chinese medicine sample to be tested, optionally amplifying it, using the genomic DNA or its amplified product as a DNA substrate, and using a target sequence detection system to detect whether the standard specific target sequence exists in the DNA substrate, wherein the target sequence detection system Taqman probe-based real-time PCR system passes the detection. If the CT value is greater than 37, the Chinese medicine sample to be tested is identical to the specified Chinese medicine origin species, otherwise it is not.

2. The method of claim 1, wherein: The presence of the standard specific target sequence in the DNA substrate was detected using a TaqMan probe-based real-time PCR system with a primer pair represented by SEQ ID NOs: 18 and 16 and a TaqMan probe represented by FAM-GCTTGAATGAAAGTCAGGCACTCCGCCA-BHQ.

3. The method of claim 1, wherein: In step (3), a primer pair for specifically amplifying a standard specific target sequence is used to amplify the genomic DNA of the Chinese medicine to be tested and the amplified standard specific target sequence is recovered as a DNA substrate; Alternatively, the genomic DNA to be detected is amplified using primers that specifically amplify a DNA sequence containing a standard specific target sequence, and the amplified DNA sequence containing a standard specific target sequence is recovered as a DNA substrate.

4. The method of claim 1, wherein: The Taqman probe-based real-time PCR system includes a buffer, primers, nuclease-free water, a DNA substrate and a Taqman probe.

5. The method according to any one of claims 1 to 4, wherein the standard specific target nucleotide is used to identify the species of rhubarb or materials derived from rhubarb, or to distinguish rhubarb from its closely related species.

6. The method according to claim 3, wherein the primer pair comprises: SEQ ID NO:18 and SEQ ID NO:

16.

7. The method of claim 3, wherein: The primer pair is used for the following aspects: identifying rhubarb components in a sample to be tested, performing species identification on rhubarb or materials derived from rhubarb, distinguishing Chinese medicinal materials derived from rhubarb from their mixed products, or for safety testing of food, medicine or health products.

8. The method of claim 3, wherein: The sample to be tested is rhubarb, rhubarb tissue or organ, Chinese medicinal materials containing rhubarb, or rhubarb mixed with counterfeit products.

9. The method of claim 3, wherein: The sample to be detected is a sample from medicinal rhubarb, palmate rhubarb or tangutic rhubarb.

10. The method of claim 3, wherein: The sample to be detected is a sample originating from a single species or a mixture of samples originating from multiple species.

Citation Information

Patent Citations

  • Method for detecting pseudo-ginseng in Chinese patent medicine and application

    CN114182002A

  • Eukaryote species identification method based on whole genome analysis and application

    CN115087750A