DNA bar code system and method for identifying glycyrrhiza medicinal plant and hybrid complex thereof

By constructing a DNA barcoding system based on nuclear genes ITS, trnH-psbA, and ndhA, the problems of insufficient resolution and difficulty in tracing the paternal parent in the identification of medicinal plants of the genus Glycyrrhiza were solved, realizing efficient and accurate identification of medicinal plants of the genus Glycyrrhiza and their hybrid complexes, simplifying the operation process and reducing costs.

CN121975969APending Publication Date: 2026-05-05SHIHEZI UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-05
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing DNA barcoding technology has insufficient resolution in identifying medicinal plants of the genus Glycyrrhiza and their hybrid complexes. It is difficult to distinguish three closely related medicinal licorice plants, especially hybrid offspring and complex introgression types. Moreover, the operation process is cumbersome, costly, and difficult to trace the paternal lineage.

Method used

A high-resolution DNA barcoding system was constructed using a combination of the nuclear gene ITS, the chloroplast gene intergenic region trnH-psbA, and the newly screened chloroplast gene ndhA. The haplotype information of the three gene fragments was integrated to identify the chloroplast genome by utilizing the paternal inheritance characteristics of the chloroplast genome.

Benefits of technology

It has achieved efficient and accurate identification of three medicinal licorices and their hybrids and hybrid infiltration complexes, with a success rate of up to 98.08%, which simplifies the operation process and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121975969A_ABST
    Figure CN121975969A_ABST
Patent Text Reader

Abstract

The invention discloses a DNA bar code system and method for identifying glycyrrhiza medicinal plants and hybrid complexes thereof, and belongs to the technical field of molecular biology and traditional Chinese medicine identification. The system is formed by combining three bar codes of a nuclear gene segment ITS, a chloroplast gene segment ndhA and a chloroplast gene spacer trnH-psbA. Wherein ndhA is a core bar code which is screened and verified for the first time, and the specific variation site of the ndhA can be used for accurately distinguishing three kinds of medicinal liquorice, namely, Glycyrrhiza uralensis Fischhex DC., Glycyrrhiza glabra. And Glycyrrhiza inflata Batalin. The system fully utilizes the characteristics of chloroplast paternal heredity of glycyrrhiza, not only can identify homozygous species, but also can efficiently analyze complex hybridization and introgression types, and traces the source of a male parent. Compared with a traditional multi-bar code combination, the method has the advantages of being high in identification success rate (up to 98.08%), easy and convenient to operate and low in cost, and is suitable for liquorice germplasm resource identification, medicinal material market rapid screening and improved variety breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of molecular biology and traditional Chinese medicine identification technology, specifically relating to a DNA barcoding system and method for identifying medicinal plants of the genus Glycyrrhiza and their hybrid complexes, particularly relating to a method for species identification, hybrid origin analysis and paternal tracing using a combination of nuclear gene ITS, chloroplast gene intercalation region trnH-psbA and newly screened chloroplast gene ndhA. Background Technology

[0002] Over the past two decades, wild licorice resources have declined sharply, prompting a gradual shift in the main source of medicinal licorice products towards cultivated varieties. However, the quality of cultivated licorice remains unstable, reflecting serious challenges in terms of germplasm consistency and genetic stability. Extensive hybridization and gene introgression exist among the three main medicinal licorice varieties, leading to significant genetic differentiation between their seeds and seedlings. This hybridization not only increases the difficulty of standardized cultivation but also causes seed source mixing problems during the harvesting stage.

[0003] Therefore, establishing reliable morphological identification traits applicable to these hybrid complexes is of significant practical importance. However, due to the high complexity of morphological differentiation and genetic structure of licorice from different sources and origins, "recessive species" with similar phenotypes but vastly different genetic backgrounds sometimes appear. This further increases the difficulty of germplasm resource classification and identification, and also restricts the efficient development and utilization of wild licorice resources. Against this backdrop, systematically elucidating the origin, phylogenetic relationships, and accurate identification methods of medicinal licorice hybrid complexes is particularly urgent.

[0004] Molecular identification techniques, especially DNA barcoding, provide a powerful supplement to species identification. Current technologies commonly use barcodes such as ITS, trnH-psbA, rbcL, matK, and trnV-ndhC, or combinations thereof, to identify licorice. However, these barcodes have the following limitations: 1. Insufficient resolution: Existing barcodes have limited variation sites, making it difficult to effectively distinguish three closely related medicinal licorice species, especially their hybrid offspring and complex introgression types. 2. System complexity: To improve the identification rate, 4-5 barcodes are usually required, leading to cumbersome procedures, high costs, and low efficiency, making it unsuitable for large-scale germplasm resource screening and rapid identification in the medicinal materials market. 3. Difficulty in tracing paternal lineage: Traditionally, chloroplasts are considered maternally inherited, but in the genus *Glycyrrhiza*, they have been proven to be paternally inherited. Current technologies fail to fully utilize this key genetic characteristic to trace the paternal origin of hybrid offspring.

[0005] Therefore, there is an urgent need in this field to develop a novel DNA barcoding system that combines high resolution, high efficiency, and low cost to solve the problem of accurate identification of medicinal plants in the genus *Glycyrrhiza* and their hybrid complexes. Summary of the Invention

[0006] The purpose of this invention is to overcome the shortcomings of insufficient DNA barcode resolution in existing licorice identification techniques, and to provide a method for screening high-resolution, specific DNA barcodes from the chloroplast genome of the *Glycyrrhiza* genus, and to establish an efficient, accurate, and convenient molecular identification system to solve the classification and identification problems of three medicinal licorice species and their hybrids and hybrid introgression complexes. This invention provides the following technical solution: By sequencing, comparing, and analyzing multiple *Glycyrrhiza* chloroplast genomes, including those with clearly defined artificial hybrid offspring, highly variable regions are systematically screened, and specific DNA barcode fragments capable of stably distinguishing the three medicinal licorice species are identified. Based on this core concept, this invention provides the following specific solutions: To achieve the above objectives, the present invention provides the following technical solution: Firstly, this invention provides a method for screening high-resolution specific DNA barcodes. This method involves sequencing and comparing multiple chloroplast genomes of the *Glycyrrhiza* genus, including those of known artificially hybridized offspring, to systematically screen for hypervariable regions and identify specific DNA barcode fragments capable of stably distinguishing three medicinal licorice species. This method provides an effective approach for rapidly and accurately locating key identifying markers from massive amounts of genomic data.

[0007] Secondly, this invention provides a DNA barcoding system for identifying hybrids of medicinal licorice plants. This system is based on a combination of three gene fragments: the internal transcribed spacer region (ITS) of the nuclear gene, the chloroplast gene spacer region (trnH-psbA), and a newly screened chloroplast gene fragment (ndhA) from this invention. This combination fully utilizes the different genetic characteristics of the nuclear genome and the chloroplast genome. By integrating and analyzing the haplotype information of the three fragments, it enables comprehensive identification of three homozygous medicinal licorice species, their hybrid offspring, and more complex genetic introgression types.

[0008] Thirdly, the present invention provides the application of the DNA barcoding system in the identification of species of Glycyrrhiza, analysis of hybridization origins, and identification of the source of Chinese medicinal materials.

[0009] (III) Beneficial Effects Compared with the prior art, the present invention has the following significant advantages: Original innovation: For the first time, a complete method for screening high-resolution DNA barcodes from chloroplast genomes was proposed and implemented, and the key marker ndhA was successfully discovered.

[0010] Strong identification capability: The constructed identification system can effectively identify hybrid and introgression complexes that are difficult to distinguish using traditional barcodes, with a high success rate.

[0011] Strong systematicity: Through the combination of multiple genes and multiple inheritance patterns, a complementary identification network is formed, which significantly improves the reliability and depth of the identification results. Attached Figure Description

[0012] Figure 1 Analysis of hypervariable regions in the chloroplast genome of the genus Glycyrrhiza and screening of ndhA barcodes.

[0013] Figure 2 The results of species identification of the genus Glycyrrhiza based on the new molecular identification system.

[0014] Figure 3 DNA barcodes used in 2016, 2020, and 2023. Detailed Implementation

[0015] To enable those skilled in the art to better understand and implement the technical solutions of the present invention, the key technical means, operating procedures, and technical effects involved in the present invention will be described in detail below in a non-limiting manner with reference to the accompanying drawings, tables, and specific embodiments. Experimental methods in the following embodiments that do not specify specific conditions are generally performed under conventional conditions or according to the conditions recommended by the reagent company.

[0016] Example 1: Preliminary morphological observation, chloroplast genome analysis, and screening of the core barcode ndhA in a Licorice hybrid complex. 1. Experimental Materials and Morphological Basis The research of this invention began with the observation of complex phenotypes in the natural hybridization regions of the genus Glycyrrhiza. We collected a total of 345 samples from nine natural hybridization regions in Xinjiang and Gansu, including three basic medicinal licorice species (G. uralensis, G. inflata, G. glabra) and their suspected hybrid offspring.

[0017] By systematically measuring and comparing key traits of pods and compound leaves in these samples (such as pod shape, curvature, and whether they are swollen, as well as the number, shape, and leaf margin characteristics of leaflets), we have preliminarily identified five typical hybrid types with intermediate morphological or characteristic combinations (see...). Figure 1 ).like Figure 1 As shown, the pod morphology of the three homozygous licorice species differed significantly: *G. uralensis* had long, curved, ring-shaped or sickle-shaped pods; *G. inflata* had short, noticeably inflated pods; and *G. glabra* had straight or slightly curved pods that were not inflated. The pods and leaves of the five hybrid types inherited characteristics of their respective parents to varying degrees. This work provides a morphological reference and a preliminary classification basis for subsequent molecular identification.

[0018] 2. Chloroplast genome sequencing and hypervariable region analysis To develop high-resolution molecular identification markers, we sequenced and aligned the genomes of 12 Glycyrrhiza chloroplasts, including artificially hybridized parents and offspring. Specific methods included DNA extraction, second / third-generation sequencing, genome assembly and annotation (see previous draft for details), ultimately yielding high-quality genome sequences.

[0019] To locate the most variable regions across the entire genome, we performed a sliding window analysis using DnaSP 6 software (window length 600 bp, step size 200 bp). Figure 1 As shown in Part A, this analysis clearly demonstrates the distribution of nucleotide diversity (Pi values) across the entire chloroplast genome. The X-axis represents the center of the window, and the Y-axis represents the Pi values. We set Pi > 0.004 as the threshold (shown by the dashed line in the figure) and selected seven prominent hypervariable peaks, providing direction for subsequent screening of specific barcodes.

[0020] 3. Screening and Verification of Core Barcode ndhA We conducted a detailed sequence alignment analysis on the seven initially selected hypervariable regions. For example... Figure 1 As shown in Part B, through parallel comparison of the features of multiple hypervariable regions, we found that although most regions are polymorphic, it is difficult to find a simple combination of sites that can stably and clearly distinguish the three types of licorice.

[0021] Ultimately, our focus was on the ndhA gene region. For example... Figure 1 As shown in section C, the alignment diagram of this sequence visually reveals its strong potential as a core barcode. The arrows in the diagram clearly indicate the specific variant sites located at positions 122-131 (a 10 bp fragment) and position 168 (a single base site). These site combinations exhibit distinct differences in three medicinal licorice varieties, forming three very different haplotypes: Haplotype NA-1: corresponds to Glycyrrhiza uralensis. Sequence characteristics: positions 122-131 are AGCTATCTCC, and position 168 is G.

[0022] Haplotype NA-2: corresponds to Glycyrrhiza inflata. Sequence characteristics: positions 122-131 are AGCTATCTCC, and position 168 is A.

[0023] Haplotype NA-3: corresponds to Glycyrrhiza glabra. Sequence characteristics: positions 122-131 are GGAGATAGCT, and position 168 is A.

[0024] Conclusion: This embodiment utilizes whole-genome sliding window analysis ( Figure 1Part A) systematically screens high-variability regions and compares sequence features ( Figure 1 Part B) and fine sequence alignment ( Figure 1 (Part C of the DNA sequence), and finally successfully identified and verified ndhA as the core DNA barcode capable of distinguishing 100% of the three medicinal licorices.

[0025] Example 2: Construction, application, and result visualization of a new molecular identification system 1. Construction of molecular identification systems and haplotype networks Based on the findings of Example 1, we constructed a novel molecular identification system consisting of the nuclear gene ITS, the chloroplast interstitial region trnH-psbA, and the core barcode ndhA. This system fully utilizes the different genetic characteristics of the nuclear genome and the chloroplast genome (which has been confirmed to be paternally inherited).

[0026] like Figure 2 As shown in the figure, this atlas illustrates the final identification results of the molecular identification system of this invention. It visualizes how complex sample sets can be clearly classified by integrating the haplotype information of three barcodes. The figure clearly shows: The three basic species branches are defined based on the ndhA haplotypes (NA-1, NA-2, NA-3).

[0027] Different hybridization types (such as I×U, U×I, etc.) and introgression types (such as G(+U), etc.) were further analyzed using ITS and trnH-psbA haplotypes.

[0028] Figure 2 This demonstrates intuitively that the system can not only identify homozygous species, but also efficiently analyze the genetic composition of hybrid and introgression complexes.

[0029] 2. Phylogenetic cluster analysis To verify the reliability of the identification results from the perspective of system evolution, we constructed an adjacency tree based on the tandem sequences of ITS, trnH-psbA, and ndhA for cluster analysis.

[0030] like Figure 1 As shown, the neighbor-joining tree clearly displays three main branches, corresponding to taxa centered on *G. uralensis*, *G. inflata*, and *G. glabra*, respectively. More importantly, all hybrids and introgressive individuals cluster strictly according to their paternal origin (identified by chloroplast barcodes) under branches centered on the paternal species. This result strongly confirms the accuracy of the identification system of this invention.

[0031] The key role of paternal chloroplast inheritance in Glycyrrhiza hybridization events is that the chloroplast haplotypes of the hybrid offspring inherit the characteristics of their fathers.

[0032] The system's powerful classification capabilities can clearly reveal the evolutionary relationships between samples.

[0033] 3. Actual identification process and results In practice, after DNA extraction, PCR amplification (using the ndhA primers verified in Example 1 and other barcode primers) and sequencing of the sample to be tested, the obtained sequence is compared with the haplotype database established in this invention (the core of which can be found in...). Figure 2 The following can be compared with the legend. By analyzing the haplotype combinations of the three, the following can be determined step by step: the basal species (determined by ndhA); whether hybridization exists and which species is involved (indicated by ITS); and the paternal origin and whether introgression exists (confirmed jointly by trnH-psbA and ndhA).

[0034] The system was applied to identify 52 difficult-to-identify licorice samples, achieving a success rate of 98.08%. The results have been approved. Figure 2 and Figure 1 It provided a clear and powerful demonstration.

[0035] In summary, the present invention achieves this through ( Figure 1 Morphological observation laid the foundation for real-world problems, through ( Figure 1 A systematic analysis of the data revealed the core barcode ndhA, which was ultimately discovered through (…). Figure 2 )and( Figure 1 The identification results and cluster analysis fully validated the high resolution, high accuracy, and powerful practicality of the constructed molecular identification system. The attached figures provide irrefutable visual evidence for each part of the specification.

[0036] It should be noted that, Figure 1 (A) A sliding window analysis diagram based on the chloroplast genomes of 12 Glycyrrhiza genus (window length: 600 bp; step size: 200 bp), showing nucleotide diversity (Pi value) across the genome. The X-axis represents the center position of the window, the Y-axis represents the nucleotide diversity value, and the peak region represents the selected hypervariable region. (B) A comparison diagram of hypervariable region characteristics used to distinguish three medicinal licorices. (C) Sequence alignment diagram of the core DNA barcode ndhA in the three medicinal licorices, with arrows clearly indicating the specific variant sites located at positions 122-131 and 168.

[0037] The table compares the success rates of species identification with the old and new molecular identification systems. Species I, U, and G correspond to *G. inflata*, *G. uralensis*, and *G. glabra*, respectively. H represents a hybrid complex, and the parental origin of the hybrid type is indicated in parentheses. In summary, through Examples 1 and 2, this invention fully demonstrates the entire process from core barcode discovery and genetic principle verification to practical identification applications. The provided technical solution is clear and specific, and experiments have proven that it has significant advantages such as high resolution, ease of operation, and low cost. It can effectively solve the problem of accurate identification of medicinal plants of the genus Glycyrrhiza and their hybrid complexes, and has good prospects for widespread application.

Claims

1. A method for identifying species, hybrids, or introgression types of Glycyrrhiza genus plants, characterized in that, Includes the following steps: a) Obtain DNA from the licorice sample to be tested; b) Determine the sequences of the nuclear gene internal transcription spacer (ITS), chloroplast gene spacer (trnH-psbA), and chloroplast gene ndhA in the DNA; c) The measured ITS, trnH-psbA, and ndhA sequences were compared with the haplotype database; d) Based on the haplotype combination of the three, determine the species identity of the sample, or identify its hybridization type and introgression origin.

2. The method according to claim 1, characterized in that, The criteria for determining haplotype combinations mentioned in step d) include: The ndhA sequence is used to distinguish the basal species, wherein: Haplotype NA-1 corresponds to Glycyrrhiza uralensis; Haplotype NA-2 corresponds to Glycyrrhiza inflata; The haplotype NA-3 corresponds to Glycyrrhiza glabra.

3. The ITS sequence is used to indicate hybridization events, wherein haplotype I-3 indicates hybridization of Glycyrrhiza uralensis with other species.

4. The trnH-psbA sequence is used to assist in species identification and identification of specific introgression types.

5. The method according to claim 2, characterized in that, The haplotypes NA-1, NA-2, and NA-3 are defined by the nucleotide sequences at positions 122–131 and 168 of the ndhA gene, wherein: The sequence characteristics of NA-1 are that positions 122-131 are AGCTATCTCC, and position 168 is G; The sequence characteristics of NA-2 are that positions 122-131 are AGCTATCTCC, and position 168 is A; The sequence characteristics of NA-3 are that positions 122-131 are GGAGATAGCT and position 168 is A.

6. A combination of DNA molecular markers for identifying plants of the genus Glycyrrhiza, characterized in that, The combination consists of the internal transcriptional spacer (ITS) of the nuclear gene, the chloroplast gene spacer (trnH-psbA), and the chloroplast gene ndhA.

7. A primer pair for amplifying the ndhA fragment in the DNA molecular marker combination of claim 4, characterized in that, Its nucleotide sequence is as follows: Upstream primer ndhA-F: 5'-AATCAAGCAATACTCCCC-3' (SEQ ID NO: 1), Downstream primer ndhA-R: 5'-GAATTATCAATAACCCCAT-3' (SEQ ID NO: 2).

8. The use of the DNA molecular marker combination of claim 4 in the preparation of identification reagents for identifying species, hybrids or introgression types of Glycyrrhiza genus plants.