Markers for identifying Yinghong No. 9 tea trees and their applications

By using next-generation sequencing and specific DNA barcoding technology, the standardization and accuracy issues of Yinghong No. 9 tea tree identification have been solved, enabling seedling identification, fresh leaf detection, and traceability of processed varieties, providing a highly specific and accurate identification method.

CN121137255BActive Publication Date: 2026-01-30TEA RES INST GUANGDONG ACAD OF AGRI SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511709260.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-01-30
Estimated Expiration
2045-11-20

AI Technical Summary

Technical Problem

Existing technologies cannot achieve standardized and precise varietal identification of Yinghong No. 9 tea trees. Traditional morphological and chemical marker methods are greatly affected by the environment, PCR-based molecular markers are prone to non-specific amplification and have poor reproducibility, and DNA barcoding technology has low consistency among tea tree varieties.

Method used

Using a second-generation sequencing-based approach, specific nucleotide fragment DNA barcodes (SEQ ID NO.1 and/or SEQ ID NO.2) were designed. Through high-throughput sequencing and alignment technologies, accurate identification of Yinghong No. 9 tea trees was achieved, including seedling identification, fresh leaf detection, traceability of processed varieties, and survey of tea garden distribution.

Benefits of technology

It achieves high specificity, wide applicability and high accuracy in the identification of Yinghong No. 9 tea trees, reduces the probability of false positives, is suitable for mixed sample detection, reduces sequencing costs, and improves the standardization and repeatability of results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121137255B_ABST
    Figure CN121137255B_ABST
Patent Text Reader

Abstract

This invention provides a biomarker for identifying the Yinghong No. 9 tea plant, which consists of two specific DNA sequences. This invention also provides a method for identifying the Yinghong No. 9 tea plant using this biomarker. After extracting total DNA from the sample, high-throughput sequencing is performed. The sequencing reads are then aligned to the biomarker. If the alignment completely covers the biomarker and generates a sequence that is identical to the biomarker sequence, the species of the sample is determined to be the Yinghong No. 9 tea plant.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of biotechnology, specifically relating to a marker for the Yinghong No. 9 tea tree and its application. Background Technology

[0002] Yinghong 9 is a superior large-leaf tea variety bred by the Tea Research Institute of Guangdong Academy of Agricultural Sciences. It is the core variety of Yingde black tea geographical indication product, with specific adaptability, unique quality, and extremely high economic value.

[0003] Tea variety identification has long faced multiple challenges. As a cross-pollinated plant, the tea genome exhibits polyploid characteristics with a high proportion of repetitive sequences. Traditional morphological identification relies on indicators such as leaf morphology and bud characteristics, which not only requires waiting for the tea tree to grow to maturity (3-5 years) but is also easily affected by environmental factors such as light and soil, resulting in a high misidentification rate for closely related varieties. Chemical markers (such as the content of tea polyphenols and theaflavins) are significantly affected by processing techniques and harvesting season, with characteristic component differences of up to 15%-20% between different batches of the same variety, making them unreliable as a stable basis for identification.

[0004] Existing identification techniques all have significant limitations. While PCR-based molecular markers (such as SSR and AFLP) are more accurate than traditional methods, primer design is extremely difficult—the tea genome contains many repetitive sequences, and closely related varieties show little difference in conserved regions, making non-specific amplification easy, with a heterogeneous band ratio of 20%-30%. Furthermore, primers with high specificity are easily affected by DNA template degradation and polyphenol residues. Simultaneously, the PCR reaction is highly sensitive to annealing temperature and Mg²⁺. + Highly sensitive to conditions such as concentration, the results show poor repeatability among different laboratories and operators, and the inter-laboratory validation pass rate is insufficient, making standardized application difficult. Traditional DNA barcoding technology relies on universal fragments such as rbcL and matK, which have over 99% sequence identity among tea varieties. However, the nuclear gene ITS exhibits multi-copy heterogeneity, making it impossible to generate a uniform and consistent sequence, thus rendering variety-level identification completely ineffective.

[0005] In summary, the complex genetic background of tea trees and the inherent limitations of existing technologies have resulted in a lack of standardized and precise methods for identifying the Yinghong No. 9 variety. Developing an identification method based on the complete comparison of second-generation sequencing reads and specific DNA barcodes, achieving precise variety determination through consistent sequence matching, can not only address the pain points of existing technologies but also meet the industry's urgent need for varietal authenticity verification, thus possessing significant practical value and technological innovation significance. Summary of the Invention

[0006] To overcome the aforementioned difficulties, we used next-generation sequencing to obtain two specific nucleotide fragments from the Yinghong 9 tea plant (Camellia sinensis 'Yinghong 9'). Analysis showed that these fragments did not exhibit high homology with the genomes of any existing tea plants, making them suitable as identification reference sequences (DNA barcodes) for the Yinghong 9 tea plant. Furthermore, we designed a reliable method for identifying the Yinghong 9 tea plant using the sequencing data. The technical solution adopted in this invention is as follows:

[0007] On the one hand, the present invention provides a marker for identifying or assisting in the identification of Yinghong No. 9 tea trees, the marker being a DNA barcode, the nucleic acid sequence of which is shown in SEQ ID NO. 1 and / or SEQ ID NO. 2.

[0008] Preferably, the DNA barcode can be used alone or in combination with SEQ ID NO.1 or SEQ ID NO.2; when used in combination, the alignment conditions of the two sequences must be met simultaneously, which can significantly reduce the probability of false positive identification.

[0009] On the other hand, another object of the present invention is to provide the application of the above-mentioned DNA barcode in the identification or auxiliary identification of Yinghong No. 9 tea tree.

[0010] This invention provides the application of the above-mentioned DNA barcode in the identification or auxiliary identification of Yinghong No. 9 tea trees, specifically including but not limited to:

[0011] Authenticity identification of tea tree seedling varieties: used for early identification of Yinghong No. 9 tea trees in the seedling stage, solving the defect of traditional morphological identification that requires waiting for maturity;

[0012] Fresh leaf raw material purity testing: used to screen for variety mixing during the harvesting of Yinghong No. 9 fresh leaves, ensuring the variety uniformity of the processed raw materials;

[0013] Processed product traceability and identification: It can trace the origin of processed products such as Yinghong No. 9 black tea and green tea (even if the DNA is partially degraded) to exclude the possibility of non-Yinghong No. 9 raw materials being used as substitutes;

[0014] Tea Variety Distribution Survey: Rapidly detect the mixing ratio of Yinghong No. 9 in large tea gardens to provide data support for variety purification.

[0015] On the other hand, the present invention provides a method for identifying or assisting in the identification of Yinghong No. 9 tea trees, comprising the following steps:

[0016] 1) Collect tea tree tissues to be tested as samples;

[0017] 2) Extract total DNA from the sample to be tested;

[0018] 3) Perform high-throughput sequencing on the total DNA of the above samples to obtain sequencing reads;

[0019] 4) Assemble the sequencing data and compare the assembled contigs with the aforementioned DNA barcodes;

[0020] 5) The contigs that completely cover the DNA barcode and have the highest consistency are sequence-aligned with the DNA barcode. The results are used to determine whether the sample is from Yinghong No. 9 tea tree. If the sequence consistency between the contigs and the target DNA barcode is 100% (no base mismatches, insertions or deletions), the sample is identified as Yinghong No. 9 tea tree. If there is only partial coverage or the sequence consistency is <100%, it is identified as not Yinghong No. 9 tea tree.

[0021] Preferably, contigs can be obtained by assembling data using software such as SPAdes, Velvet, or SOAPdenovo2.

[0022] On the other hand, the present invention provides a method for identifying or assisting in the identification of Yinghong No. 9 tea trees, comprising the following steps:

[0023] 1) Collect tea tree tissue as the sample to be tested;

[0024] 2) Extract total DNA from the sample to be tested;

[0025] 3) Perform high-throughput sequencing on the total DNA of the above samples to obtain sequencing reads;

[0026] 4) Use software to align the above sequencing reads to the DNA barcode described in this invention;

[0027] 5) Determine whether the sample contains Yinghong No. 9 tea tree based on the coverage of sequencing reads on the DNA barcode; count the coverage of sequencing reads on the target DNA barcode and generate a consistency sequence using software; if the coverage is 100% and the consistency sequence is completely identical to SEQ ID NO. 1 and / or SEQ ID NO. 2 (no base difference), then the sample is determined to contain Yinghong No. 9 tea tree; if the coverage is <100% or the consistency sequence has base differences from SEQ ID NO. 1 and / or SEQ ID NO. 2, then the sample is determined not to contain Yinghong No. 9 tea tree.

[0028] Furthermore, for a large number of samples, a group screening strategy can be adopted:

[0029] a) Divide the n samples to be tested into 2 groups, and mix the sequencing reads of each sample in each group;

[0030] b) Test each group for the presence of Yinghong No. 9 tea trees according to steps 4)-5) above, and locate the positive group (the group containing Yinghong No. 9);

[0031] c) For positive samples, a binary grouping method is used to repeatedly group and mix sequencing reads and detect until a single sample containing Yinghong No. 9 tea tree is identified; this strategy is suitable for large-scale sample screening (such as batch testing in nurseries) and can reduce sequencing costs by more than 50%.

[0032] In one embodiment, the high-throughput sequencing described in step 3) is second-generation sequencing or third-generation sequencing.

[0033] In one embodiment, step 4) uses any one of the following software: Minimap2, Geneious, Bowtie, Tophat, or HISAT to perform read comparison.

[0034] In one embodiment, SEQ ID NO.1 or SEQ ID NO.2 is used alone.

[0035] In one embodiment, SEQ ID NO.1 and SEQ ID NO.2 are used together.

[0036] Preferably, the aforementioned high-throughput sequencing data reaches 10~15×.

[0037] The beneficial effects of this invention are:

[0038] High specificity: The newly discovered SEQ ID NO.1 and SEQ ID NO.2 are specific fragments exclusive to Yinghong No. 9. After comparison and verification with the genome database, they show no high homology with other tea varieties, which solves the problem that traditional markers cannot distinguish closely related varieties.

[0039] Wide applicability: It can directly detect mixed samples (such as mixed fresh leaves and processed broken tea) without separating individual plants, making it especially suitable for "batch rapid screening" scenarios in industry.

[0040] Simple to operate: Based on high-throughput sequencing and direct alignment, there is no need to design specific primers, which avoids the defects of PCR technology such as difficult primer design, condition sensitivity (such as amplification failure caused by annealing temperature fluctuations), and poor reproducibility, and has a high degree of standardization.

[0041] High accuracy: Combined dual-marker verification further reduces errors and meets the stringent requirements for variety identification. Attached Figure Description

[0042] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0043] Figure 1 This is the comparison result of SEQ ID NO.1, the marker of Yinghong No. 9 tea tree, in the NCBI nr / nt database.

[0044] Figure 2 It is the sequence with the highest alignment consistency between SEQ ID NO.1, the marker of Yinghong No. 9 tea tree, and the sequence in the NCBI nr / nt database.

[0045] Figure 3 This is the comparison result of SEQ ID NO.2, the marker of Yinghong No. 9 tea tree, in the NCBI nr / nt database.

[0046] Figure 4 It is the sequence with the highest alignment consistency between SEQ ID NO.2, the marker of Yinghong No. 9 tea tree, and the sequence in the NCBI nr / nt database.

[0047] Figure 5 This is the coverage result of the first group of tea tree samples containing Yinghong No. 9 on the SEQ ID NO.1 sequence. The figure shows the result after the line break, with several bases overlapping at the beginning and end of the two lines.

[0048] Figure 6 This is the coverage result of the first group of tea tree samples containing Yinghong No. 9 on the SEQ ID NO. 2 sequence. The figure shows the result after the line break, with several bases overlapping at the beginning and end of the two lines.

[0049] Figure 7 This is the coverage result of the second group of tea tree samples containing Yinghong No. 9 on the SEQ ID NO.1 sequence. The figure shows the result after the line break, with several bases overlapping at the beginning and end of the two lines.

[0050] Figure 8 This is the coverage result of the second group of tea tree samples containing Yinghong No. 9 on the SEQ ID NO. 2 sequence. The figure shows the result after the line break, with several bases overlapping at the beginning and end of the two lines.

[0051] Figure 9 This is the coverage result of the first group of tea tree samples that do not contain Yinghong No. 9 on the SEQ ID NO.1 sequence. The figure shows the result after the line break, with several bases overlapping at the beginning and end of the two lines.

[0052] Figure 10 This is the coverage result of the first group of tea tree samples that do not contain Yinghong No. 9 on the SEQ ID NO. 2 sequence. The figure shows the result after the line break, with several bases overlapping at the beginning and end of the two lines.

[0053] Figure 11This is the coverage result of the second group of tea tree samples that do not contain Yinghong No. 9 on the SEQ ID NO.1 sequence. The figure shows the result after the line break, with several bases overlapping at the beginning and end of the two lines.

[0054] Figure 12 This is the coverage result of the second group of tea tree samples that do not contain Yinghong No. 9 on the SEQ ID NO. 2 sequence. The figure shows the result after the line break, with several bases overlapping at the beginning and end of the two lines.

[0055] Figure 13 This is the comparison result of the contig assembled from the first group of tea tree samples containing Yinghong No. 9 with SEQ ID NO. 1.

[0056] Figure 14 This is the comparison result of the contig assembled from the first group of tea tree samples containing Yinghong No. 9 with SEQ ID NO. 2.

[0057] Figure 15 This is the comparison result of the contig assembled from the second group of tea tree samples containing Yinghong No. 9 with SEQ ID NO. 1.

[0058] Figure 16 This is the comparison result of the contig assembled from the second group of tea tree samples containing Yinghong No. 9 with SEQ ID NO. 2. Detailed Implementation

[0059] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these examples are for illustrative purposes only and are not intended to limit the scope of the invention. Unless otherwise specified, specific experimental methods can be used in the following examples.

[0060] Unless otherwise defined, all technical and scientific terms used in this invention have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0061] Example 1: Markers for Yinghong No. 9 Tea Trees

[0062] This invention, through high-throughput sequencing assembly and comparative analysis, determined the standard detection sequence for biomarkers in the Yinghong No. 9 tea plant. The specific sequence is as follows:

[0063] SEQ ID NO.1:

[0064] CGTGCCCTACGACGAGTAGGCCACTCACTCCCTTTAGAGTTGGTTTAACACTAATAAGGGCGCTGCGAGTGCTATCGATAAAGGCTGCACATGAAGCTGTGGGACTTGAGCTAGTGGCATGAATCGTCCGCTTCTTCCTAGCTGCGTCAGTCCCGCTTTCAGCGTAGAATCGGA ACGGAAGGCCTAGTATTAGCGCTAGCAGACTACCGGTTTTGATCTACTTACTTTTGCCCCCTATCTACTTTTGAACAGAGGCTTCGTAGCCCATTTTCCTTCCCTAAACAAAGGTCTTAAGAATCTCATTCAACCAGGTAACTAAATTATGCCACTTTTTAAGCGCCATAAGCA

[0065] SEQ ID NO.2:

[0066] ATTAATAAGTGAGTTATAAAACCCCATTCCTTTTCTTATACTATGTATATGTTTATATTCTAAAATAAATTATATTCTAAAGAGTACTCAAAGGGTATAGGTACTTTGTGTTTATATTTCCTTTTCTTATACTATGTATATGTTTTATATTCTAAAATAAATTATATTCTAAAGAGTACTCAAAGGGTATATGTACTTTTTTTTCTTATTCTAAAGAGTATATGTCCTTTGTGTTTATATTCTAAAGAGTACTCAAAGGG

[0067] Example 2: Sequence alignment of biomarkers of Yinghong No. 9 tea tree with closely related species

[0068] To determine the species specificity of the biomarkers described in this invention, the biomarkers were submitted to NCBI for online BLAST comparison. The nr / nt library was selected, species were not limited, and the "More dissimilar sequences" parameter was selected. Figures 1-4 As shown, the marker described in this invention was found to be closely related to Camellia species in the NCBI nr / nt database. Figure 1 The highest identity for SEQ ID NO.1 is found in *Camellia nitidissima*, although the identity is 100%. However, *Camellia nitidissima* is distributed across three different locations in the genome. Figure 2 The same applies to other closely related species. Figure 3 Only six closely related species showed homology with SEQ ID NO.2, with the highest similarity being Camellia tachangensis, which, although showing 99.07% similarity, is distributed across two different locations in the Camellia tachangensis genome. Figure 4 This is also true for other closely related species. This indicates that the markers described in this invention are highly specific.

[0069] Example 3: Further Validation of Biomarker Specificity

[0070] To further verify the specificity of the biomarker of the present invention, 93 tea tree samples were selected as the samples to be tested. The second-generation sequencing reads (sequencing depth 10×~15×) were mapped to SEQ ID NO.1 and SEQ ID NO.2 (the two sequences were mixed as a reference sequence) using minimap2, and the total number of aligned reads was counted.

[0071] The results are shown in Table 1 below:

[0072] Table 1. Total number of reads compared between different tea varieties

[0073]

[0074] It is evident that the total number of sequencing reads of other varieties on SEQ ID NO.1 and SEQ ID NO.2 is much lower than that of Yinghong No.9 (by 2 to 3 orders of magnitude; the table shows the combined statistical results after mapping the two sequences), indicating that SEQ ID NO.1 and SEQ ID NO.2 have extremely high specificity within the genome of Yinghong No.9.

[0075] Example 4: Identification Method of Yinghong No. 9 Tea Tree

[0076] Based on the markers described in the above embodiments, this embodiment provides a specific method for identifying Yinghong No. 9 tea trees, which specifically includes the following steps:

[0077] 1) Collect tea tree tissue as the sample to be tested;

[0078] 2) Extract total DNA from the sample to be tested;

[0079] 3) Perform next-generation high-throughput sequencing on the total DNA of the above samples to obtain sequencing reads;

[0080] 4) Using the default parameters of Geneious software, align the above sequencing reads to the biomarkers described in this invention;

[0081] 5) Determine whether the sample contains Yinghong No. 9 tea tree based on the coverage of sequencing reads on the marker; if the marker is completely covered and a consistent sequence that is completely identical to the marker can be generated, it is determined that the species of the sample to be tested contains Yinghong No. 9 tea tree.

[0082] 6) If multiple sample sequencing reads are grouped and mixed in step 4), then step 6) is further included to further divide the samples containing Yinghong No. 9 tea trees into two groups, mix the sequencing reads separately, and repeat step 5) until the Yinghong No. 9 tea tree in a single sample is identified.

[0083] Test results as follows Figures 5-12 As shown, to ensure clarity, the image is displayed in a sequence-wrapped format, with several bases overlapping at the beginning and end of the two lines of reference sequences (marker sequences). Figures 5-8 The sample contains Yinghong No. 9 tea trees. The figure shows that the sequencing reads completely cover the reference sequence (i.e., the markers SEQ ID NO. 1~2 described in this invention) and generate a consistent sequence that is exactly the same as the marker sequence. Figures 9-12 For samples that do not contain Yinghong No. 9 tea trees (data from Example 3 and a random mix of all second-generation sequencing reads of non-Yinghong No. 9 tea trees published on NCBI as of November 11, 2025), the read coverage on the reference sequence is incomplete, and a consensus sequence that is completely identical to the marker sequence cannot be generated (the ? in the consensus sequence in the figure indicates that the site is a gap), indicating that the method of the present invention can accurately identify Yinghong No. 9.

[0084] Example 5: Identification Method Two for Yinghong No. 9 Tea Tree

[0085] Based on the markers described in the above embodiments, this embodiment provides another specific method for identifying Yinghong No. 9 tea trees, which specifically includes the following steps:

[0086] 1) Collect tea tree tissues to be tested as samples;

[0087] 2) Extract total DNA from the sample to be tested;

[0088] 3) Perform high-throughput sequencing on the total DNA to obtain sequencing reads;

[0089] 4) Assemble the sequencing data and align the assembled contigs with the biomarker BLAST;

[0090] 5) Align the contigs that can completely cover the marker and have the highest consistency with the marker multiple sequence (using software such as MAFFT, MUSCLE, MEGA, DNAMAN, etc.), and determine whether the sample contains Yinghong No. 9 tea tree based on the alignment results; if they are consistent with the marker sequence, it is determined that the sample to be tested contains Yinghong No. 9 tea tree sample.

[0091] Test results as follows Figures 13-16 As shown, both groups of samples containing Yinghong 9 can generate contigs that are completely consistent with the marker (DNA barcode) sequence. Other samples cannot assemble a consistent sequence (or are too short and dispersed across different contigs, making it impossible to define a homologous sequence), so they are not shown.

[0092] Based on the disclosure and teachings of the foregoing specification, those skilled in the art can make appropriate changes and modifications to the above embodiments. Therefore, the present invention is not limited to the specific embodiments disclosed and described above, and some modifications and changes to the present invention should also fall within the protection scope of the claims of the present invention. Furthermore, although some specific terms are used in this specification, these terms are only for convenience of explanation and do not constitute any limitation on the present invention.

Claims

1. A marker for identifying or assisting in the identification of Camellia sinensis cv. 'Keemun 9', characterised in that, The marker is a DNA barcode, and the nucleotide sequence of the DNA barcode is shown in SEQ ID NO. 1 and / or SEQ ID NO.

2.

2. A method of identifying or assisting in the identification of Camellia sinensis cv. 'English Breakfast No. 9' characterised by, The method comprises the following steps: 1) Collecting tea tree tissues to be detected as a sample to be detected; 2) Extracting total DNA of the sample to be detected; 3) High-throughput sequencing of the total DNA to obtain sequencing reads; 4) Assembling the sequencing data, and performing blast comparison of the assembled contigs with the marker of claim 1; 5) Performing multiple sequence alignment of the contigs that can completely cover the marker of claim 1 and have the highest consistency with the marker of claim 1, and judging whether the sample contains the Camellia sinensis cv. Yinghong No. 9 according to the comparison results; if the sequence is completely consistent with the sequence of the marker of claim 1, it is judged that the sample to be detected contains the Camellia sinensis cv. Yinghong No.

9.

3. A method of identifying or assisting in the identification of Camellia sinensis cv. 'English Breakfast No. 9' characterised by, The method comprises the following steps: 1) Collecting tea tree tissues to be detected as a sample to be detected; 2) Extracting total DNA of the sample to be detected; 3) High-throughput sequencing of the total DNA to obtain sequencing reads; 4) Comparing the sequencing reads to the marker of claim 1; 5) Judging whether the sample contains the Camellia sinensis cv. Yinghong No. 9 according to the coverage of the sequencing reads on the marker; if the sequencing reads can completely cover the marker and generate a consistent sequence identical to the marker, it is judged that the sample to be detected contains the Camellia sinensis cv. Yinghong No.

9.

4. The method of identifying or assisting in the identification of Camellia sinensis var. assamica cv. 'Kejriwal Red' as claimed in claim 3, wherein, If the sequencing reads are mixed after grouping a plurality of samples in step 4), further comprising step 6) grouping the samples containing the Camellia sinensis cv. Yinghong No. 9 into two groups, mixing the sequencing reads respectively, repeating step 5), and identifying a single sample containing the Camellia sinensis cv. Yinghong No.

9.

5. A method of identifying or assisting in the identification of Camellia sinensis var. assamica cv. 'Kejriwal Red' as claimed in claim 3 or claim 4, wherein, The high-throughput sequencing in step 3) is second-generation sequencing or third-generation sequencing.

6. A method of identifying or assisting in the identification of Camellia sinensis var. assamica cv. 'Kejriwal Red' as claimed in claim 3 or claim 4, wherein, In step 4), any one of Geneious, Minimap2, Bowtie, Tophat or HISAT is used for reads comparison.

Citation Information

Patent Citations

  • SSR (Simple Sequence Repeat) core primer combination and kit for identifying tea tree variety 'Yinghong No.9' and application of SSR core primer combination and kit

    CN118186143A

  • Primers for identifying of green tea species and specific-identification methods of green tea species using the primers

    KR1020090130606A