Non-targeted identification and analysis method for plant source components in honey based on DNA bar code technology
By using DNA barcoding technology and PCR amplification with rbcL primers, combined with Qiime analysis, the problem of rapid and accurate identification of plant-derived components in honey has been solved, achieving high-throughput screening and supporting the development of the honey industry and regulatory enforcement.
Patent Information
- Application Number
- CN202510849589.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-11-07
AI Technical Summary
Existing technologies are insufficient for quickly and accurately identifying plant-derived components in honey, and suffer from problems such as time consumption, expensive equipment, or low sensitivity.
Using DNA barcoding technology, PCR amplification was performed with rbcL primers, combined with Qiime analysis, to establish a high-throughput non-targeted screening method for honey samples, screen out the optimal DNA barcode, and identify plant-derived components in honey.
It enables rapid and accurate identification of plant-derived components in honey, protects consumer rights, promotes the development of the honey industry, and provides technical support for government regulation.
Smart Images

Figure CN120905370A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of detection, and particularly relates to a method for non-targeted identification of plant-derived components in honey based on DNA barcoding technology. BACKGROUND
[0002] Honey is a natural sweet substance with medicinal and edible properties, rich in various amino acids, vitamins, minerals and other nutrients, and has good effects on improving intestinal function and regulating the stomach and spleen, and has been favored by consumers. However, the phenomenon of substandard honey and fake honey occurs frequently. The Codex Alimentarius Commission (CAC) of the United States states that honey is one of the most adulterated foods, and the European Union also points out that honey is a high-risk adulterated food. Honey adulteration usually involves adding low-price honey to high-price honey for sale, such as adding low-price rape honey to high-price jujube flower honey and linden honey; or adding sugar syrup, such as rice syrup, corn syrup, cassava syrup, peanut syrup, sugarcane syrup, and mixed syrup, to honey or directly using it as honey; or feeding bees with sugar syrup or sucrose during the honey collection period. Therefore, it is crucial to establish an effective honey identification method.
[0003] In recent years, various methods have been established to evaluate the quality of different honey products, mainly including chromatography, mass spectrometry detection technology, spectroscopy detection technology, protein-based detection technology, and nucleic acid-based detection technology. Chromatography and mass spectrometry technology can efficiently identify plant-derived components in honey, but are time-consuming and expensive. Spectroscopy technology requires little or no sample preparation, is environmentally friendly, easy to operate, and can be used for online detection or online process control, but has low sensitivity, high signal-to-noise ratio, and difficulty in accurately determining the authenticity of plant components. Protein-based detection methods, including high-performance liquid chromatography, electrophoresis technology, and enzyme-linked immunosorbent assay, have shown satisfactory results in species identification, but this method depends on the detection of polypeptide targets, which is always tissue-dependent, and the sensitivity significantly decreases when subjected to heat treatment. In addition, pollen analysis can be used to identify the flower source of honey, but this method is time-consuming, requires professional knowledge, and involves a tedious counting process, making the identification of plant sources challenging. Due to the high stability of DNA and the high sensitivity and specificity of PCR amplification, DNA-based detection technology is considered a more reliable and stable method. DNA barcoding technology is a rapid, accurate, and automated identification technology based on DNA that has developed rapidly in recent years. It uses a standard DNA sequence as a marker to identify different species based on the chromosome set. This method has many advantages, such as reliability, speed, and non-targeting, and has been widely applied in many scientific fields. Therefore, there is an urgent need for a method that can identify plant-derived components in honey using DNA barcoding technology. SUMMARY
[0004] Therefore, the present application aims to provide a DNA barcode technology-based non-targeted identification analysis method for plant-derived components in honey, to identify honey samples through DNA barcode technology, to establish a high-quality DNA extraction and evaluation method for honey samples, to screen the optimal DNA barcode, and to establish a high-throughput non-targeted screening method for plant-derived components in honey samples.
[0005] To achieve the above-mentioned purposes, the technical scheme of the present application is as follows:
[0006] A DNA barcode technology-based non-targeted identification analysis method for plant-derived components in honey, comprising the following steps:
[0007] 1) Extracting DNA from the honey sample, performing PCR amplification using rbcL primers, the rbcL primer sequences being shown in SEQ ID No. 5 and SEQ ID No. 6, and obtaining a single band for subsequent sequencing analysis;
[0008] 2) Measuring the sequencing quality by the base quality value Q phred , wherein Q phred is the error probability of the recognized base in the base recognition process, and the greater Q phred is, the lower the error probability is: Q phred =-10log 10 (e); wherein e represents the sequencing error rate;
[0009] 3) Using Qiime to analyze the species classification information corresponding to the OTU, and analyzing the plant-derived components contained in the honey.
[0010] Further, the DNA concentration extracted by the kit used for extracting the DNA of the honey sample is greater than 80 mg / L, and the OD 260 / OD 280 value is between 1.8 and 2.0.
[0011] Further, the kit used for extracting the DNA of the honey sample is any one of DNLT0-DNeasy LT, Plant Genomic DNA Extraction Kit-DiO, Plant Genomic DNA Extraction Kit-Solebao, and High-efficiency Plant Genomic DNA Extraction Kit.
[0012] Furthermore, to obtain higher quality and more accurate bioinformatics analysis results, the raw sequencing data were optimized. First, the two sequences were aligned, and then spliced according to the overlapping regions at the ends of the alignment, removing sequences containing N from the spliced results. Next, adapter sequences were removed, as were bases with a quality value lower than 20 at both ends and sequences shorter than 200 bp. Finally, the spliced and filtered sequences were compared with the database to remove chimeric sequences, yielding the final valid data.
[0013] Furthermore, the reaction parameters for PCR amplification were as follows: the total reaction volume was 25 μL, containing: 12.5 μL of Premix Ex Taq, 0.5 μL each of forward and reverse primers, with a concentration of 10 μmol / L for both forward and reverse primers, and 2 μL of honey sample DNA with a concentration of 10 μg / mL-100 μg / mL. ddH2O was added to bring the total volume to 25 μL.
[0014] Furthermore, the PCR amplification conditions were as follows: 95℃ pre-denaturation for 30s; 95℃ denaturation for 30s, 50℃ annealing for 30s, 72℃ extension for 30s, 35 cycles; 72℃ extension for 10min; storage at 4℃.
[0015] Furthermore, sequencing quality analysis parameters include Q 20 Q 30 Q 20 The probability of a base error is 1%, Q 30 The probability of a base error is 0.1%, Q 20 Reaching over 90% or Q 30 A success rate of 85% or higher indicates good sequencing results and high reliability for subsequent analyses.
[0016] Compared with existing technologies, the non-targeted identification and analysis method for plant-derived components in honey based on DNA barcoding technology described in this invention has the following advantages:
[0017] The present invention describes a non-targeted identification and analysis method for plant-derived components in honey based on DNA barcoding technology. This method utilizes specifically selected DNA barcodes to establish a high-throughput non-targeted screening method for plant-derived components in honey samples. This research is of great significance for protecting consumer rights and promoting the development of the honey industry. Simultaneously, it provides strong technical support for government departments to regulate and enforce regulations in the honey industry and combat illegal and irregular activities such as adulteration. Attached Figure Description
[0018] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an undue limitation of the invention. In the drawings:
[0019] Figure 1PCR electrophoresis map of 4 gene sequence fragments of positive sample 4; M is Marker, from top to bottom, 2000, 1000, 750, 500, 250, 100 bp; A is matK gene, B is ITS2 gene, C is trnH-psbA gene, D is rbcL gene; numbers 1-10 represent respectively: wolfberry, black locust, jujube, Chinese milkvetch, Vitex, linden, rape, loquat, milkvetch, Codonopsis;
[0020] Figure 2 PCR electrophoresis map of rbcL gene fragments of different samples; M is Marker, from top to bottom, 2000, 1000, 750, 500, 250, 100 bp; numbers 1-10 represent respectively: apple, grape, kiwi, apricot, strawberry, pear, peach, hawthorn, mango, orange;
[0021] Figure 3 PCR electrophoresis map of rbcL gene fragments of honey samples; M is Marker, from top to bottom, 2000, 1000, 750, 500, 250, 100 bp; numbers 1-12 represent respectively: 12 honey samples (the same as Table 4);
[0022] Figure 4 Statistical diagram of effective sequence length distribution;
[0023] Figure 5 OTU abundance clustering heat map;
[0024] Figure 6 Species distribution heat map. DETAILED DESCRIPTION
[0025] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0026] The technical solutions in the embodiments of the present application will be described clearly and completely below with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0027] Example 1 DNA extraction method screening
[0028] Plant Genome DNA Extraction Kit (Solabio), Fast Plant Genome DNA Extraction System (DP321), DNLT0-DNeasy LT, CTAB Method Plant Genome DNA Extraction Kit, High-efficiency Plant Genome DNA Extraction Kit, Plant Genome DNA Extraction Kit (CTAB Crude Extraction Method), PureLink TM Genomic DNA Mini Kit, Plant Genome DNA Extraction Kit (Dio), Universal DNA Extraction Kit, Rapid Plant Genome DNA Extraction Kit, were used to extract DNA from honey samples. The DNA extraction steps were performed according to the instructions, and each sample was repeated 3 times. The extracted DNA was measured for concentration using a nucleic acid protein quantifier and stored at -20°C for later use.
[0029] The quality of the results of DNA barcode detection technology depends on the extraction of high-quality DNA from food samples. Honey has a high sugar content, so it is crucial to extract high-quality honey DNA.
[0030] In this example, honey samples were first treated, and DNA was extracted from the samples using different kits, and the extraction effects were compared. In terms of operation steps and time, the CTAB method kit took the longest time, and the centrifugal column type kit had a similar operation process. In terms of DNA extraction concentration and purity (Table 1), for honey samples, the DNA extracted by the centrifugal column type plant genome DNA extraction kit had a higher concentration, and the OD 260 / OD 280 was between 1.8-1.9, indicating high DNA purity, while the DNA extracted by the CTAB method kit had a lower concentration, and the OD 260 / OD 280 was lower than 1.80, indicating possible contamination, and the DNA extracted by the universal DNA extraction kit had a higher purity but a lower concentration. Therefore, DNLT0-DNeasy LT, Plant Genome DNA Extraction Kit (Dio), Plant Genome DNA Extraction Kit (Solabio), and High-efficiency Plant Genome DNA Extraction Kit can be used for subsequent experiments.
[0031] Table 1 Quality and concentration of genomic DNA extracted from honey samples using different methods
[0032]
[0033]
[0034] Example 2 Screening and determination of DNA barcode
[0035] 1. DNA extraction from other samples
[0036] Other plant samples except honey were taken 100 mg to 200 mg, and DNA was extracted by plant genomic DNA extraction kit. The concentration was measured by nucleic acid protein quantifier, and stored at-20℃ for later use.
[0037] In this example, DNA of 10 kinds of honey plant, such as wolfberry, black locust, jujube, Chinese milk vetch, Vitex, linden, rape, loquat, milk vetch and codonopsis, was extracted.
[0038] 2. Primer design
[0039] The ITS2, MatK, rbcL and trnH-psbA gene sequences of common honey plant species were retrieved from GenBank database, and multiple sequence alignment was performed by MegAlign software. Then, specific primers were designed by Premier 5.0 software. The primer sequence information is shown in Table 2, and the primers were synthesized by Suzhou Jinyuzhi Biological Technology Co., Ltd.
[0040] Table 2 Gene primer sequence
[0041]
[0042]
[0043] PCR amplification was performed by matK, ITS2, rbcL and trnH-psbA primers, respectively, and sterile ddH2O was used as negative control. The amplification products were detected by 2% agarose gel electrophoresis.
[0044] PCR amplification: the total volume of reaction system was 25 μL, which contained: Premix Ex Taq 12.5 μL, upstream and downstream primers (10 μmol / L) 0.5 μL each, sample DNA (10 μg / mL-100 μg / mL) 2 μL, and ddH2O to make up the total volume of 25 μL.
[0045] The matK primer amplification conditions were: 95℃ pre-denaturation for 30 s; 95℃ denaturation for 30 s, 56℃ annealing for 30 s, 72℃ extension for 30 s, 35 cycles; 72℃ extension for 10 min; 4℃ storage.
[0046] The ITS2 primer amplification conditions were: 95℃ pre-denaturation for 30 s; 95℃ denaturation for 30 s, 54℃ annealing for 30 s, 72℃ extension for 30 s, 35 cycles; 72℃ extension for 10 min; 4℃ storage.
[0047] The rbcL primer amplification conditions were: 95℃ pre-denaturation for 30 s; 95℃ denaturation for 30 s, 50℃ annealing for 30 s, 72℃ extension for 30 s, 35 cycles; 72℃ extension for 10 min; 4℃ storage.
[0048] trnH-psbA primer amplification conditions: 95℃ pre-denaturation 30s; 95℃ denaturation 30s, 64℃ annealing 30s, 72℃ extension 30s, 35 cycles; 72℃ extension 10min; 4℃ preservation.
[0049] Electrophoresis: prepare 2% agarose gel, take 2g agarose, heat in 100mL electrophoresis buffer, add 10μL 4SGelred nucleic acid dye (10000x), prepare gel. Add electrophoresis buffer in the electrophoresis tank, so that the liquid surface just covers the gel. Mix 5μL PCR amplification product with appropriate loading buffer, spot sample, and load 5μL DNA Marker at the same time, 9V / cm constant voltage, electrophoresis until the bromophenol blue indicator migrates to 2 / 3 of the gel. Observe the electrophoresis result by gel imaging analysis system and record.
[0050] The detection results are shown in Table 2. Figure 1 As shown in the figure, the matK sequence primer amplification efficiency is poor, and the wolfberry, black locust, rape, loquat are not amplified with obvious target bands, and the rest of the samples such as jujube appear multiple non-specific bands. The ITS2 sequence amplification band size is about 500bp, among which about 250bp non-specific bands are also amplified in wolfberry and black locust. The trnH-psbA amplifies about 600bp bands in wolfberry, black locust, jujube and Chinese milkvetch, while about 300bp bands are obtained in wattle and basswood, about 500bp bands are obtained in rape, loquat, milkvetch and codonopsis, the amplification band size is different, and non-specific amplification bands appear. The rbcL sequence universal primer amplification efficiency is good, and clear single 300bp bands are amplified in 10 positive samples, and no non-specific amplification appears.
[0051] According to the PCR results of matK, ITS2, rbcL and trnH-psbA 4 pairs of primers, the rbcL primer amplification product is sent to the sequencing company for sequencing, and the amplification sequencing result is submitted to GenBank for BLAST comparison analysis, and the comparison result is shown in Table 3. The 10 positive samples are sequenced successfully, and the comparison result in GenBank is consistent with the known species, and the similarity rate is more than 99%. Based on the above analysis, rbcL can be used as a DNA barcode sequence fragment for subsequent experiments.
[0052] Table 3 Different species DNA sequence GenBank database comparison identification results
[0053]
[0054]
[0055] Example 3 primer universality detection
[0056] In order to further verify the universality of the rbcL primer, the DNA of 10 kinds of angiosperms such as apple, grape, kiwi, apricot, strawberry, pear, peach, hawthorn, mango, orange, etc. was subjected to PCR amplification, and sterile water was used as a blank control. The electrophoresis detection results verified the universality of the primer.
[0057] The results, as shown in Table 1, showed that a single band of about 300 bp was amplified from the 10 plants, indicating that the rbcL primer had good universality. Figure 2
[0058] Example 4 Analysis of non-targeted screening results of commercially available honey
[0059] 1. DNA extraction and amplification of honey samples
[0060] Twelve commercially available honey samples (Table 4) were treated, and honey DNA was extracted using the DNLT0-DNeasy LT kit. After detecting the genomic concentration and purity, PCR amplification was performed using the rbcL primer. As shown in Table 5, a single band of about 300 bp was obtained from the 12 honey samples, and subsequent sequencing analysis could be performed. Figure 3
[0061] Table 4 Honey sample information table
[0062]
[0063]
[0064] 2. Analysis of sequencing data quality control results
[0065] Sequencing quality is measured by base quality value Q. Q refers to the error probability of the recognized base in the base recognition process. The larger the Q, the lower the error probability. If the sequencing error rate is represented by e, and the base quality value is represented by Q, then the following relationship exists: phred
[0066] Q phred = -10log 10 (e);
[0067] Q 20 indicates that the probability of base error is 1%, and Q 30 indicates that the probability of base error is 0.1%. Generally, Q 20 reaches more than 90% or Q 30 reaches more than 85% indicates that the sequencing result is very good.
[0068] Table 5 Quality statistics of raw data
[0069]
[0070]
[0071] GC content is a very important indicator in sequencing technology, which represents the bias in the sequencing process and is used to determine whether the sequencing process is random enough. It is generally believed that when the GC content is less than 35% or more than 65%, the sequencing difficulty will increase and the subsequent result analysis will be affected, which needs to be solved by constructing different length sequencing library and increasing sequencing amount. As can be seen from Table 5, the GC content in the sequence is about 50%. 20 The minimum proportion is 93.83%, and the average is 96.12%. 30 The minimum proportion is 77.45%, which appears in the sequence of sample 8, and the overall average is 85.80%. This indicates that the sequencing quality is good and the reliability of subsequent analysis is high.
[0072] In order to obtain higher quality and more accurate biological information analysis results, the raw sequencing data is optimized. First, the two sequences are aligned, and the sequences containing N in the alignment end overlap region are removed. Next, the adapter sequence is removed, the bases with quality value less than 20 at both ends are removed, and the sequences with length less than 200 bp are removed. Finally, the filtered sequences are aligned with the database, and the chimeric sequences are removed to obtain the final effective data. As shown in Table 6, most of the sequence lengths are concentrated around 280 bp. Figure 4
[0073] 3. Sequencing data result analysis
[0074] (1) OTU analysis
[0075] OTU (Operational Taxonomic Units) is a unified symbol set by a person for a certain classification unit in systematics or population genetics research for easy analysis. The clustering analysis of 12 samples (FM1-FM12) produced 357 OTUs, Figure 5 The 30 OTUs with the highest abundance are shown. It can be seen that samples 1-5, 7, 9, 10 of honey mainly enrich in OTU1 and OTU2; sample 6 mainly enriches in OTU5 and OTU6, followed by OTU1 and OTU2; sample 8 mainly enriches in OTU9 and OTU11; samples 11 and 12 mainly enrich in OTU3 and OTU4.
[0076] (2) Species annotation result analysis
[0077] Using Qiime to analyze the species classification information corresponding to OTUs, the number of species in the 12 honey samples at the seven levels of Kingdom, Phylum, Class, Order, Family, Genus, and Species is shown in Table 6. It was found that all 12 honey samples contained multiple species. At the species level, honey sample 1 had 21 species, honey samples 2 and 3 had 29 species, honey sample 4 had 28 species, honey sample 5 had 25 species, honey sample 6 had 39 species, honey sample 7 had 32 species, honey sample 8 had 30 species, honey sample 9 had 43 species, honey sample 10 had 37 species, honey sample 11 had 68 species, and honey sample 12 had as many as 121 species.
[0078] Table 6. Statistical table of species number under different classification levels for each sample.
[0079]
[0080] Clustering is performed at each level to obtain a clustering heatmap. Figure 6 As shown in the figure, the 12 honey samples contain rich species information and multiple species were identified. Among them, the honey samples with the highest abundance were peanut (1-5, 7, 9, and 10), honey sample with the highest abundance was honeysuckle (6), honey sample with the highest abundance was soybean (8), and honey samples with the highest abundance were rapeseed (11 and 12). In addition, honey sample No. 1 also contained astragalus, sappanwood, soybean, and wolfberry, but not loquat; sample No. 2 also contained locust, soybean, cherry, codonopsis, astragalus, morning glory, and wolfberry, but not vitex or linden; sample No. 3 also contained honeysuckle, astragalus, soybean, morning glory, codonopsis, and sappanwood, but not jujube flower; sample No. 4 also contained honeysuckle, goldenrain tree, cassava, wolfberry, morning glory, soybean, and rapeseed, but not locust; sample No. 5 contained locust, and also contained astragalus, poplar, soybean, and morning glory; sample No. 6 also contained peanut, sappanwood, astragalus, soybean, and... Honey sample 7 contained melon and other components, but no Vitex negundo was identified; sample 8 contained wolfberry, as well as Astragalus membranaceus and melon; sample 9 contained soybean, Astragalus membranaceus, cassava, honeysuckle, morning glory, and peanut; sample 10 contained Astragalus membranaceus, soybean, and honeysuckle; sample 11 contained Codonopsis pilosula; and samples 12 contained Astragalus membranaceus, as well as Locust tree, broad bean, Koelreuteria paniculata, Photinia serratifolia, cherry, pea, morning glory, and shrub bean.
[0081] The above merely provides the preferred embodiment of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A method for non-targeted identification of plant-derived components in honey based on DNA barcoding technology, characterized in that: The method comprises the following steps: 1) Extracting DNA of the honey sample, performing PCR amplification by using rbcL primers, and performing subsequent sequencing analysis after obtaining a single band; 2) sequencing quality is measured by base quality values Q phred Q phred is the probability of error given for a base identified during base calling, Q phred the greater the probability of error: Q phred = -10 log 10 (e); where e represents the sequencing error rate; 3) Using Qiime to analyze the species classification information corresponding to the OTU to analyze the plant-derived components contained in the honey.
2. The method for non-targeted identification of plant-origin components in honey based on DNA barcoding technology according to claim 1, characterized in that: The DNA concentration extracted by the kit used to extract the honey sample DNA was > 80 mg / L, OD 260 / OD 280 values were between 1.8 and 2.
0.
3. The method for non-targeted identification of plant-origin components in honey based on DNA barcoding technology according to claim 2, characterized in that: The kit for extracting DNA of the honey sample is any one of DNLT0-DNeasy LT, plant genome DNA extraction kit-Diogen, plant genome DNA extraction kit-Solepia, and high-efficiency plant genome DNA extraction kit.
4. The method for non-targeted identification of plant-origin components in honey based on DNA barcoding technology according to claim 1, characterized in that: In order to obtain higher quality and more accurate bioinformation analysis results, the original sequencing data can also be optimized. First, the two sequences are aligned, and the sequences containing N in the splicing results are removed according to the terminal overlapping region of the alignment. Next, the adapter sequence is removed, the bases with a quality value lower than 20 at both ends are removed, and the sequences with a length less than 200 bp are removed. Finally, the spliced and filtered sequences are aligned with the database, and the chimeric sequences are removed to obtain the final effective data.
5. The method for non-targeted identification of plant-origin components in honey based on DNA barcoding technology according to claim 1, characterized in that: The reaction parameters of PCR amplification are as follows: the total volume of the reaction system is 25 μL, which contains 12.5 μL of Premix Ex Taq, 0.5 μL of each of the upstream and downstream primers, the concentration of the upstream and downstream primers is 10 μmol / L, the concentration of the honey sample DNA is 10 μg / mL-100 μg / mL, the volume is 2 μL, and ddH2O is added to make up the total volume of 25 μL.
6. The method for non-targeted identification of plant-origin components in honey based on DNA barcoding technology according to claim 1, characterized in that: The PCR amplification conditions are as follows: 95℃ pre-denaturation for 30 s; 95℃ denaturation for 30 s, 50℃ annealing for 30 s, 72℃ extension for 30 s, 35 cycles; 72℃ extension for 10 min; and 4℃ preservation.
7. The method for non-targeted identification of plant-origin components in honey based on DNA barcoding technology according to claim 1, characterized in that: The sequencing quality analysis parameters include Q 20 , Q 30 , Q 20 , Q 30 , Q 20 , and Q 30 . Q 20 indicates that the probability of base error is 1%, Q 30 indicates that the probability of base error is 0.1%, Q 20 indicates that the probability of base error is 0.01%, Q 30 indicates that the probability of base error is 0.001%, Q 20 indicates that the probability of base error is 0.0001%, and Q 30 indicates that the probability of base error is 0.00001%. When Q 30 reaches more than 90% or Q 30 reaches more