A method for identifying native tea plant varieties using characteristic metabolites
By constructing a multivariate statistical clustering model based on characteristic metabolite vectors and using UPLC-QTOF/MS to detect fresh tea leaf samples, the problem of morphological methods being unable to distinguish the Fuliang Zhuye group from other tea varieties was solved. This enabled reliable traceability and variety identification of Fuliang tea raw materials, and improved the quality stability of geographical indication products.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-03-31
AI Technical Summary
Existing morphological methods are insufficient to effectively distinguish the Fuliang Zhuye group from other tea varieties, affecting the standardization of Fuliang tea raw material sources and the quality stability of geographical indication products.
By constructing a multivariate statistical clustering model based on sixteen characteristic metabolites, and using UPLC-QTOF/MS to detect the metabolite peak signals of fresh tea leaf samples, characteristic metabolite vectors are constructed and clustering is performed to achieve automatic identification of variety identity.
This enables quantifiable and repeatable determination of the authenticity of the source of Fuliang tea raw materials, improves the reliability of raw material control for geographical indication products and the objectivity of variety identification, avoids interference from external factors, and ensures consistency in identification across batches and seasons.
Smart Images

Figure CN121385159B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of identifying tea tree varieties. More specifically, the present invention relates to a method for identifying native tea tree varieties using characteristic metabolites. Background Art
[0002] The Fuliang Zhuye population variety is the main raw material variety for producing the geographical indication product "Fuliang Tea". The Fuliang Tea made from its fresh leaves has quality characteristics consistent with traditional processes.
[0003] However, in the actual fresh leaf acquisition and processing and circulation links, fresh leaves of other tea tree varieties are often mixed in or冒充 as the Fuliang Zhuye population variety.
[0004] Since the differences in external morphology between different tea tree varieties are not sufficient to support reliable determination, the existing morphological identification methods are difficult to meet the needs of variety authenticity discrimination.
[0005] This problem directly affects the standardization of the raw material source of Fuliang Tea and the quality stability of geographical indication products.
[0006] Therefore, it is necessary to establish an identification method based on the internal chemical composition characteristics of fresh leaves of the Fuliang Zhuye population variety to achieve the traceability and determination of the raw material source and provide technical support for the geographical indication protection system of Fuliang Tea. Summary of the Invention
[0007] To overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a method for identifying native tea tree varieties using characteristic metabolites. Based on the stable numerical distribution pattern of sixteen characteristic metabolites in fresh leaves of the Fuliang Zhuye population variety, a characteristic metabolite vector is constructed and the variety identity determination is completed through multivariate statistical clustering to solve the problems raised in the above background art.
[0008] To achieve the above object, the present invention provides the following technical solution: A method for identifying native tea tree varieties using characteristic metabolites, comprising:
[0009] S1. Sample preparation: Obtain fresh leaf samples of tea trees to be identified, standard samples of the Fuliang Zhuye population variety, and control samples of non-Fuliang Zhuye population varieties. After microwave drying and pulverizing each sample, weigh 50 mg of the powder, add 1200 μL of a 70% methanol extraction solution pre-cooled to -20 °C for metabolite extraction, centrifuge and filter to obtain the待测 extraction solution.
[0010] S2. Metabolite detection: Inject the待测 extraction solution obtained in S1 into an Agilent 1290 ultra-high performance liquid chromatography-6545 QTOF / MS mass spectrometry system (UPLC-QTOF / MS), and detect the sample under preset chromatographic conditions and mass spectrometry conditions to obtain the mass spectrometry peak signals of the metabolites in the sample.
[0011] S3. Data processing: The raw mass spectrometry data was converted into mzXML format using ProteoWizard. Peak extraction, peak alignment and retention time correction were performed using the XCMS program. The peak area was corrected using the SVR method, and metabolites with a missing rate greater than 50% were filtered out.
[0012] S4. Construction of characteristic metabolites: In the calibration peak set of S3, 16 characteristic metabolites were obtained by searching public databases and annotating them, and the peak intensity of each characteristic metabolite was used as the relative content to construct a metabolite feature vector.
[0013] S5. Clustering discrimination: Input the 16-dimensional characteristic metabolite vectors of the sample to be identified, the standard sample, and the control sample into the multivariate statistical clustering model. If the sample to be identified and the standard sample are clustered in the same class, the sample is determined to be a species of Castanopsis fargesii; if the sample is clustered in the same class as the control sample, the sample is determined not to be a species of Castanopsis fargesii.
[0014] In a preferred embodiment, metabolite extraction in S1 includes: vortexing once every 30 min for 30 s each time, for a total of 6 vortexes; centrifugation conditions of 13400×g for 3 min; and filtration using a microporous membrane with a pore size of 0.22 μm.
[0015] In a preferred embodiment, the chromatographic and mass spectrometric conditions used in S2 include:
[0016] (1) The chromatographic column includes ACQUITY-UPLC-HSS-T3-C18 type chromatographic column, with a packing particle size of 1.8 micrometers, an inner diameter of 2.1 mm, and a column length of 100 mm;
[0017] (2) Mobile phase A includes a 0.1% formic acid aqueous solution, and mobile phase B is a 0.1% formic acid-acetonitrile solution;
[0018] (3) The gradient elution program includes: 0-11 min, phase B 5%-90%; 11-12 min, phase B 90%; 12-12.1 min, phase B 90%-5%; 12.1-14 min, phase B 5%;
[0019] (4) Column temperature 40℃, flow rate 0.4mL / min, injection volume 2μL;
[0020] (5) The mass spectrometer uses ESI positive and negative ion mode, with ion source voltage of 2500 / 1500V, auxiliary gas flow rate of 8L / min, ion source temperature of 325℃, and rupture voltage of 135V.
[0021] In a preferred embodiment, the 16 characteristic metabolites in S4 include:
[0022] Cellotetrasaccharide, theophylline, proanthocyanidin C1, vitexin 2′′-O-β-L-rhamnoside, delphinidin-3-O-morula disaccharide, epicatechin, epigallocatechin, naringenin chalcone, gallatechin, 2-methylaminoadenosine, L-alanyl-adenosine, cytosine, guanine, cytidine-2′,3′-monophosphate, 8-hydroxy-2′-deoxyguanosine, and valine-aspartic acid.
[0023] In a preferred embodiment, the feature vector composed of the 16 characteristic metabolites maintains a stable numerical distribution pattern in samples collected in spring, summer and autumn; the discrimination results obtained based on the feature vector are consistent between samples collected in Fuliang County and samples collected in Nanchang County.
[0024] In a preferred embodiment, the cluster analysis in S5 is based on a hierarchical clustering method using Euclidean distance, and the branch affiliation of a sample in the cluster tree is used as the identification criterion, wherein:
[0025] (1) When the branching of the sample is consistent with that of the standard sample of the *Castanopsis fargesii* population, it is judged to be the *Castanopsis fargesii* population.
[0026] (2) When the branches of the sample are consistent with those of the control sample of the non-floating sago palm species, it is judged as a non-population species.
[0027] In a preferred embodiment, a method for identifying native tea plant varieties using characteristic metabolites further includes an application:
[0028] The method for identifying native tea tree varieties using characteristic metabolites is used for tracing the source of raw materials for Fuliang tea geographical indication products, judging the acceptance of fresh leaves, obtaining evidence for market supervision, or identifying the authenticity of varieties.
[0029] The technical effects and advantages of this invention are as follows:
[0030] This invention uses a metabolite feature vector composed of sixteen characteristic metabolites to replace external morphology as the discrimination criterion, which directly solves the problem that existing morphological methods cannot distinguish between the Fuliang Zhuye group species and other tea tree varieties, and realizes a quantifiable and repeatable determination of the authenticity of the variety of Fuliang tea raw materials, thereby improving the reliability of the control of raw materials for geographical indication products.
[0031] This invention uses high-dimensional metabolite peak intensity information obtained by LC-MS to reflect the intrinsic chemical composition of tea varieties, avoiding interference from external factors such as light, leaf age, or harvesting location, so that the discrimination criteria are established at a stable molecular level, and improving the consistency of variety identification under different environments and production conditions.
[0032] This invention employs data processing procedures such as peak extraction, retention time correction, and SVR correction, which can effectively correct systematic biases and instrument drift in mass spectrometry data, making the obtained metabolite intensities comparable across batches and seasons, thereby ensuring the robustness of the final cluster analysis results.
[0033] This invention uses a hierarchical clustering model based on Euclidean distance to perform unsupervised clustering of samples. It can automatically identify differences in metabolite patterns between samples without relying on subjective experience, enabling natural separation of population species and non-population species of Castanopsis fargesii in terms of branching structure, thereby improving the objectivity and interpretability of variety determination.
[0034] The sixteen-dimensional characteristic metabolites of this invention exhibit a stable distribution pattern in samples from different production areas (Fuliang County and Nanchang County) and different seasons, verifying the cross-regional and cross-seasonal stability of the characteristic vectors, making them applicable to various application scenarios such as raw material acquisition, production area supervision, and geographical indication protection. Attached Figure Description
[0035] Figure 1 This invention presents a multivariate statistical hierarchical clustering dendrogram constructed based on samples collected in Fuliang County.
[0036] Figure 2 This invention presents a multivariate statistical hierarchical clustering dendrogram constructed based on samples collected in Nanchang County.
[0037] Figure 3 This is a comparative graph showing the content of characteristic metabolites of different tea varieties in Jinggongqiao Town, Fuliang County, under spring sampling conditions, as presented in this invention.
[0038] Figure 4 This is a comparative graph showing the content of characteristic metabolites of different tea varieties in Xihu Township, Fuliang County, under summer sampling conditions, as presented in this invention.
[0039] Figure 5 This is a comparative graph showing the content of characteristic metabolites of different tea varieties in Ehu Town, Fuliang County, under autumn sampling conditions, as presented in this invention.
[0040] Figure 6 This is a comparative graph showing the content of characteristic metabolites of different tea varieties in Huangma Township, Nanchang County, under spring sampling conditions, which is the subject of this invention.
[0041] Figure 7 This diagram illustrates the clustering method based on principal component analysis proposed in this invention.
[0042] Figure 8 This diagram illustrates the hierarchical clustering method based on Pearson correlations in this invention.
[0043] Figure 9 This diagram illustrates the cosine-based hierarchical clustering method of this invention. Detailed Implementation
[0044] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0045] Reference Figures 1-9 An embodiment of the present invention provides a method for identifying native tea tree varieties using characteristic metabolites, comprising:
[0046] S1. Sample preparation: Obtain fresh tea leaf samples to be identified, standard samples of the Fuliang Zhuye population, and control samples of non-Fuliang Zhuye population. After microwave drying, each sample is pulverized. 50 mg of powder is weighed and added to 1200 μL of 70% methanol extraction solution pre-cooled to -20℃ for metabolite extraction. The extract is centrifuged and filtered to obtain the test extract.
[0047] S2, Metabolite Detection: The extract obtained in S1 was injected into a 1290 ultra-high performance liquid chromatography-6545 QTOF / MS mass spectrometry system (UPLC-QTOF / MS). The sample was detected under preset chromatographic and mass spectrometric conditions to obtain the mass spectrometric peak signals of metabolites in the sample.
[0048] S3. Data processing: The raw mass spectrometry data was converted into mzXML format using ProteoWizard. Peak extraction, peak alignment and retention time correction were performed using the XCMS program. The peak area was corrected using the SVR method, and metabolites with a missing rate greater than 50% were filtered out.
[0049] S4. Construction of characteristic metabolites: In the calibration peak set of S3, 16 characteristic metabolites were obtained by searching public databases and annotating them, and the peak intensity of each characteristic metabolite was used as the relative content to construct a metabolite feature vector.
[0050] S5. Clustering discrimination: Input the 16-dimensional characteristic metabolite vectors of the sample to be identified, the standard sample, and the control sample into the multivariate statistical clustering model. If the sample to be identified and the standard sample are clustered in the same class, the sample is determined to be a species of Castanopsis fargesii; if the sample is clustered in the same class as the control sample, the sample is determined not to be a species of Castanopsis fargesii.
[0051] Metabolite extraction in S1 included: vortexing once every 30 min for 30 s each time, for a total of 6 vortexes; centrifugation conditions were 13400×g for 3 min; and filtration was performed using a microporous membrane with a pore size of 0.22 μm.
[0052] The chromatographic and mass spectrometric conditions used in S2 include:
[0053] (1) The chromatographic column includes ACQUITY-UPLC-HSS-T3-C18 type chromatographic column, with a packing particle size of 1.8 micrometers, an inner diameter of 2.1 mm, and a column length of 100 mm;
[0054] (2) Mobile phase A includes a 0.1% formic acid aqueous solution, and mobile phase B is a 0.1% formic acid-acetonitrile solution;
[0055] (3) The gradient elution program includes: 0-11 min, phase B 5%-90%; 11-12 min, phase B 90%; 12-12.1 min, phase B 90%-5%; 12.1-14 min, phase B 5%;
[0056] (4) Column temperature 40℃, flow rate 0.4mL / min, injection volume 2μL;
[0057] (5) The mass spectrometer uses ESI positive and negative ion mode, with ion source voltage of 2500 / 1500V, auxiliary gas flow rate of 8L / min, ion source temperature of 325℃, and rupture voltage of 135V.
[0058] The 16 characteristic metabolites in S4 include:
[0059] Cellotetrasaccharide, theophylline, proanthocyanidin C1, vitexin 2′′-O-β-L-rhamnoside, delphinidin-3-O-morula disaccharide, epicatechin, epigallocatechin, naringenin chalcone, gallatechin, 2-methylaminoadenosine, L-alanyl-adenosine, cytosine, guanine, cytidine-2′,3′-monophosphate, 8-hydroxy-2′-deoxyguanosine, and valine-aspartic acid.
[0060] The feature vector composed of the 16 characteristic metabolites maintained a stable numerical distribution pattern in samples collected in spring, summer and autumn; the discrimination results obtained based on the feature vector were consistent in samples collected in Fuliang County and samples collected in Nanchang County.
[0061] In the cluster analysis of S5, a hierarchical clustering method based on Euclidean distance is used, and the branch affiliation of a sample in the cluster tree is used as the identification criterion, wherein:
[0062] (1) When the branching of the sample is consistent with that of the standard sample of the *Castanopsis fargesii* population, it is judged to be the *Castanopsis fargesii* population.
[0063] (2) When the branches of the sample are consistent with those of the control sample of the non-floating sago palm species, it is judged as a non-population species.
[0064] An application of a method for identifying native tea tree varieties using characteristic metabolites includes: tracing the source of raw materials for Fuliang tea geographical indication products, judging the acceptance of fresh leaves, obtaining evidence for market supervision, or identifying the authenticity of varieties;
[0065] This scheme can be used for tracing the source of raw materials for Fuliang tea, inspecting fresh leaves, supervising the market, and identifying the authenticity of the variety because the feature vector composed of sixteen characteristic metabolites can stably characterize the intrinsic chemical composition of the Fuliang oak leaf population. This metabolic characteristic remains consistent under different seasons and production conditions, giving it variety specificity and environmental robustness. The identification at the metabolite level does not rely on appearance, origin markings, or subjective judgment, but is based on quantifiable, repeatable, and comparable chemical characteristics. Therefore, it can verify whether the fresh leaves are of the designated variety at the acquisition stage, conduct authenticity checks on raw materials before processing, and provide traceable and verifiable objective evidence in regulatory or law enforcement scenarios, thereby achieving the protection of raw materials and identification of the variety of Fuliang tea geographical indication products.
[0066] It should be noted that:
[0067] In the metabolite extraction process of step S1, the fresh leaf samples are first microwave-dried to quickly terminate the activity of endogenous enzymes in the fresh leaf tissue, preventing the degradation, oxidation, or redistribution of metabolites such as catechins, flavonoids, and nucleosides during post-harvest processing, thus maintaining the original metabolic state of the sample at the time of harvest. The dried sample is then pulverized to homogenize the particles, improving the diffusion efficiency of metabolites in the solvent and enhancing the reproducibility of the extraction. A fixed sample volume of 50 mg of powder is used to maintain a consistent solid-liquid ratio across different samples, avoiding extraction deviations due to differences in sample volume. Pre-cooled 70% methanol to -20°C is used as the extraction solution because 70% methanol can dissolve both polar and moderately polar metabolites, covering substances such as polyphenols, nucleosides, amino acids, and dipeptides. The low-temperature environment inhibits the degradation of heat-sensitive metabolites, ensuring the stability of the extracted metabolite components.
[0068] In the specific extraction operation, the setting of vortexing once every 30 minutes, lasting 30 seconds each time, for a total of 6 vortexes, aims to enhance the degree of cell disruption through intermittent mechanical disturbance, so that metabolites are fully released into the extract, while avoiding prolonged continuous oscillation that could cause the solution temperature to rise and damage sensitive metabolites. The centrifugation conditions are set at 13400×g for 3 minutes to ensure that fine particles and cell debris are completely settled in a short time, so that the supernatant reaches a cleanliness suitable for direct injection. A microporous membrane with a pore size of 0.22μm is used for final filtration, which can effectively remove residual particles and prevent them from entering the chromatographic column and mass spectrometry ion source, avoiding system contamination or retention time drift, thereby ensuring the stability of metabolite peak signals and the repeatability of the entire detection process.
[0069] In section S2, it's important to note that metabolite detection relies on the combined action of liquid chromatography and mass spectrometry. Therefore, each parameter in both chromatographic and mass spectrometric conditions corresponds to a specific separation and detection mechanism, and their meaning and application must be clearly understood. The 1290 ultra-high performance liquid chromatography-6545 QTOF / MS mass spectrometry system consists of two parts: the 1290 liquid chromatography system at the front end separates the mixed metabolites sequentially according to their hydrophobicity, polarity, and structural characteristics, allowing different metabolites to elute from the column at different times; the 6545 QTOF / MS mass spectrometry system at the back end ionizes each metabolite molecule eluting from the column into charged ions and detects them based on their mass-to-charge ratio. Therefore, chromatography handles "time-based separation," while mass spectrometry handles "molecular detection," and their combined action generates the corresponding mass spectrometric peak signals for each metabolite.
[0070] The ACQUITY-UPLC-HSS-T3-C18 column used in this system is a reversed-phase column. C18 represents the octadecyl bonded phase, which is suitable for separating common polyphenols, flavonoids, nucleosides, and various weakly or moderately polar small molecule compounds in tea. T3 indicates that it has enhanced retention capacity for polar compounds, allowing for distinguishable retention behavior of both polar and moderately polar substances. The combination of a column particle size of 1.8 μm, an inner diameter of 2.1 mm, and a length of 100 mm gives the system high column efficiency and rapid separation capability. The purpose is to ensure that each metabolite forms an independent elution peak on the time axis, avoiding peak overlap and misinterpretation.
[0071] Mobile phase A and mobile phase B refer to the two solvents used in the chromatographic system. Mobile phase A is a 0.1% formic acid aqueous solution, used to provide a polar environment to ensure sufficient retention of highly polar metabolites in the initial stage. The addition of formic acid helps increase ionization efficiency in subsequent mass spectrometry electrospray ionization. Mobile phase B is a 0.1% formic acid-acetonitrile solution, where acetonitrile is an organic solvent used to elute moderately polar and hydrophobic metabolites. Formic acid is also used to improve ionization stability. The so-called "gradient elution program" refers to adjusting the proportion of mobile phase B (organic phase) over time so that compounds with different degrees of polarity and hydrophobicity are eluted sequentially within a set time range. "Phase B 5%-90%" in the program means gradually increasing from the initial 5% acetonitrile to 90%, with the aim of eluting polar substances first and hydrophobic substances later. The 90% phase B holding stage is used to thoroughly elute the most difficult-to-elute hydrophobic metabolites. Then, the concentration is quickly restored to 5% phase B and held for a period of time to bring the chromatographic column back to its initial state to ensure consistent injection conditions for the next time.
[0072] The column temperature was set to 40 degrees Celsius to reduce the viscosity of the mobile phase and stabilize the chromatographic retention behavior; the flow rate was 0.4 mL / min to achieve a balance between separation speed, separation efficiency and system pressure; the injection volume was set to 2 μL to avoid peak distortion caused by a large amount of sample, while ensuring the sensitivity of mass spectrometry detection.
[0073] The mass spectrometry section employs electrospray ionization (ESI) in both positive and negative modes, meaning the same sample can be detected in both positive and negative ion modes to cover metabolite categories that may respond with different charges. An ion source voltage of 2500 / 1500 volts drives the droplets to become charged and form an electrospray jet, atomizing the sample into charged microdroplets in the ion source region. An auxiliary gas flow rate of 8 liters per minute accelerates solvent evaporation, causing the droplets to gradually shrink and eventually release individual charged molecules. An ion source temperature of 325 degrees Celsius accelerates solvent evaporation without damaging thermistor molecules, improving ionization efficiency. A fragmentation voltage of 135 volts breaks down some parent ions into characteristic fragment ions in the collision chamber, enhancing the ability to determine metabolite structures.
[0074] Under these combined conditions, the sample is eluted sequentially from the chromatographic column along with the mobile phase and enters the electrospray ionization source. After being ionized, it enters the QTOF analyzer one by one, and the corresponding ion signals are obtained according to the mass-to-charge ratio. These signals form independent mass spectrometry peaks in the time and intensity dimensions, and each peak corresponds to the presence and relative content of a specific metabolite, thereby obtaining the mass spectrometry peak signals of all characteristic metabolites in the sample.
[0075] In section S3, it's important to note that the raw mass spectrometry data obtained from metabolite detection is typically in a proprietary file format provided by the instrument manufacturer. This format cannot be directly parsed by subsequent data processing software. Therefore, ProteoWizard is used to convert the data first. ProteoWizard is a mass spectrometry data conversion tool that transforms the raw files generated by the chromatography-mass spectrometer into the publicly available standard mzXML format. mzXML is a structured data format that includes the mass spectrometry scan sequence, the mass-to-charge ratio distribution (m / z value) for each scan, the corresponding ion signal intensity, and retention time information. Through this conversion, the raw data is standardized into a format that can be read by various open-source metabolite analysis software, making subsequent processing steps more universal and reproducible.
[0076] After format conversion, the XCMS program was used to perform peak extraction, peak alignment, and retention time correction on the mzXML data. Peak extraction refers to XCMS scanning the ion signals in the chromatographic-mass spectrometry data, identifying signal abrupt change points in the time and m / z dimensions, defining signal regions with continuity and morphological characteristics as metabolite peaks, and measuring their peak area and peak height. The purpose of peak extraction is to transform the continuous raw mass spectrometry scan into discrete metabolite variables, thereby forming a feature table that can be used for statistical analysis.
[0077] Peak alignment is used to address the issue of slight shifts in the retention times of metabolites between different samples. Since retention time drift inevitably occurs during long-term operation of liquid chromatography systems, the elution times of the same metabolite will vary slightly in different samples. The role of peak alignment is to register these shifts using algorithms, so that the corresponding metabolite peaks in all samples are identified as the same variable. The peak alignment algorithm of XCMS is usually based on the similarity of characteristic peaks between samples to ensure the consistency of variables in subsequent statistical analysis.
[0078] Retention time correction is a modeling adjustment of the global retention time after peak alignment, so that the retention time distribution of the same substance in different samples converges to a unified reference curve. This process eliminates systematic deviations caused by changes in temperature, flow rate or column efficiency in the chromatographic system, and keeps the time position of metabolite peaks consistent in all samples. After this step, the metabolite peaks of all samples have a unified time coordinate system, which facilitates subsequent peak matching and quantitative analysis.
[0079] After peak extraction and retention time correction, the peak area was corrected using the SVR method. SVR stands for Support Vector Regression, which aims to correct for non-biological differences such as instrument drift and signal strength fluctuations caused by different batches or different running times. In mass spectrometry detection, phenomena such as signal attenuation and changes in ionization efficiency may occur over time. SVR fits these systematic changes by constructing a regression model and then performs inverse correction on the peak area, so that the corrected peak area can better represent the true metabolite content and reduce the impact of batch effects on the results.
[0080] In the corrected metabolite peak matrix, variables with a missing rate greater than 50% are filtered out. The missing rate refers to the proportion of a certain metabolite that is missing in all samples. Peaks with a high missing rate are often unstable in origin, incomplete in shape, or have too low a signal, and cannot be used as reliable analytical variables. By removing metabolites with a missing rate of more than 50%, it can be ensured that the variables finally used to construct the feature vector have good detection consistency and biological significance.
[0081] After the above data processing steps, the initial raw mass spectrometry scan data is converted into a structured, aligned and corrected metabolite peak intensity matrix, providing a reliable data foundation for the subsequent identification of 16 characteristic metabolites and the construction of feature vectors.
[0082] In S4, it should be noted that the calibration peak set obtained in S3 already contains the metabolite peak signals of all samples at a uniform retention time coordinate. However, these peaks are still just a "signal set" and have not been identified as specific compounds. Therefore, these peaks need to be annotated first. This involves using public metabolite databases to compare the precise mass-to-charge ratio, isotope distribution, fragment ion mode, and chromatographic retention time of each peak to determine the identity of the compound corresponding to the peak. Commonly used databases include HMDB, KEGG, and MassBank. These databases provide a large amount of characteristic ion information of known compounds. By matching the m / z values of the calibration peaks with the theoretical m / z values in the database, and further improving the accuracy of annotation through fragment feature comparison, the identifiable metabolite species can finally be screened from all peaks.
[0083] After database matching was completed, sixteen metabolites that could characterize the metabolic features of the *Zhugellium fuliangense* population were selected from all successfully annotated metabolites. These sixteen metabolites are cellotetrasaccharide, theophylline, proanthocyanidin C1, vitexin 2′′-O-β-L-rhamnoside, delphinidin-3-O-morula disaccharide, epicatechin, epigallocatechin, naringenin chalcone, gallatechin, 2-methylaminoadenosine, L-alanyl-adenosine, cytosine, guanine, cytidine-2′,3′-monophosphate, 8-hydroxy-2′-deoxyguanosine, and valine-aspartic acid. Each metabolite was supported by its specific m / z value, fragmentation pattern, and chromatographic retention behavior. These metabolites were selected into the feature set because they showed consistent accumulation patterns in multiple batches, locations, and seasons of sampling and could stably distinguish the *Zhugellium fuliangense* population from other tea varieties.
[0084] After identifying the metabolites, the peak intensity of each metabolite is used as the relative abundance of that metabolite. The peak intensity is derived from the ion signal area or peak height recorded after XCMS peak extraction, which can reflect the relative abundance of the metabolite in the sample under the same detection conditions. By arranging the peak intensities of the sixteen metabolites in a fixed order, a sixteen-dimensional metabolite feature vector can be constructed. This vector uniformly describes the metabolic characteristics of the sample and serves as the input variable for the subsequent clustering discrimination model.
[0085] It is important to emphasize that this feature vector exhibits a stable numerical distribution pattern in samples collected in different seasons. For example, in samples collected in spring, summer, and autumn, these sixteen metabolites consistently show a relatively consistent strength relationship and their overall distribution pattern is not disrupted by changes in environmental temperature, rainfall, or light. Furthermore, in samples collected from two different geographical areas, Fuliang County and Nanchang County, the discrimination results based on the same feature vector remain consistent. That is, samples belonging to the Fuliang Castanopsis fuliangense species group always show similar numerical configurations in the feature vector space, while varieties not belonging to the Fuliang Castanopsis fuliangense species group always show a different distribution.
[0086] It is precisely because of the stability of the aforementioned feature vectors across seasons and regions that they possess the reliability to serve as a basis for identifying tea tree varieties, thereby ensuring that subsequent cluster analysis can maintain consistent discrimination results under different conditions.
[0087] In S5, it should be noted that the sixteen-dimensional metabolite vector obtained in the preceding steps has converted the metabolite information of each sample into structured numerical variables. However, these variables themselves cannot directly indicate the species classification. Therefore, a multivariate statistical clustering model is needed to analyze the numerical relationships between samples. The role of cluster analysis is to automatically classify samples with similar metabolite distribution patterns into the same category based on the similarity of each sample in the sixteen-dimensional vector space, without the need for preset labels. The core steps of implementing cluster analysis include constructing a distance matrix, applying a clustering algorithm, forming a cluster tree, and determining the category of the sample based on the branch classification.
[0088] When constructing the distance matrix, Euclidean distance is used to measure the difference in metabolite feature vectors between samples. Euclidean distance represents the geometric distance between two samples in sixteen-dimensional space. It is obtained by calculating the square root of the sum of the squares of the differences in the intensity of corresponding metabolite peaks, which can reflect the degree of difference in the overall metabolite content distribution between samples. The reason for choosing Euclidean distance is that it can directly compare numerical features in multi-dimensional space without the need for additional encoding of variables, and can better maintain the contribution of the intensity difference of each metabolite to the overall distance.
[0089] After obtaining the distance matrix, the samples are clustered using hierarchical clustering. Hierarchical clustering is a stepwise sample merging technique. Its basic process is to first treat each sample as an independent category, and then merge the two most similar samples into a new category based on the minimum distance in the distance matrix. Subsequently, the distance between the merged category and other categories is updated, and the nearest category pair is found and merged again. This process is repeated until all samples gradually converge into a whole. This recursive merging process records the order and distance of each merging, thus forming a clustering tree.
[0090] Each branch of the clustering tree represents a set of samples whose metabolite distribution patterns are sufficiently similar to be classified into the same category by the algorithm. The branching structure of the clustering tree provides a visualization method that can intuitively reflect the differences in metabolite characteristics among samples. For example, if a sample and a standard sample from the *Castanopsis fargesii* population appear on the same branch in the clustering tree, it means that its 16-dimensional feature vector is highly similar to that of the standard sample. If it is always on the same branch as the control sample from the non-*Castanopsis fargesii* population, it means that its metabolite distribution pattern is similar to that of the control sample.
[0091] In application, the sample to be identified, the standard sample, and the control sample are simultaneously input into the hierarchical clustering model. The species identity is determined based on the specific branch position of the sample in the clustering tree. When the branch of the sample to be identified is consistent with that of the standard sample of the *Castanopsis fargesii* population, it means that its overall metabolite distribution pattern is highly consistent with the standard sample, and it can be identified as a *Castanopsis fargesii* population species. When the branch of the sample to be identified is consistent with that of the control sample of a non-*Castanopsis fargesii* population, it can be identified as a non-population species. This method achieves classification without human intervention through branch assignment driven by numerical similarity, which can avoid identification errors caused by appearance, subjective judgment, or environmental factors, and ensure that consistent discrimination results can be obtained for different batches of samples at different detection times and different sampling locations.
[0092] Furthermore, the identification method of this scheme is explained as follows:
[0093] Sixteen characteristic metabolites are specifically accumulated in the fresh leaves of the *Castanopsis fargesii* population: cellotetrasaccharide, theophylline, proanthocyanidin C1, vitexin 2′′-O-β-L-rhamnoside, delphinidin-3-O-morula disaccharide, epicatechin, epigallocatechin, naringenin chalcone, gallatechin, 2-methylaminoadenosine, L-alanyl-adenosine, cytosine, guanine, cytidine-2′,3′-monophosphate, 8-hydroxy-2′-deoxyguanosine, and valine-aspartic acid.
[0094] One standard sample (from the *Zanthoxylum bungeanum* cultivar), one non-standard sample (any variety other than the *Zanthoxylum bungeanum* cultivar), and the variety to be identified were collected. The peak intensities of 16 characteristic metabolites in all samples were detected by liquid chromatography-mass spectrometry (LC-MS). Multivariate statistical cluster analysis was performed. Those clustered with the standard sample belonged to the *Zanthoxylum bungeanum* cultivar, while those clustered with the non-standard sample did not belong to the *Zanthoxylum bungeanum* cultivar.
[0095] like Figure 1As shown in Example 1 of this scheme: Fresh leaves of 7 tea varieties were picked in Fuliang County for identification. Fuliang Zhuye Group Species - Xihu and Fuliang Zhuye Group Species - Jinggong were grouped with Fuliang Zhuye Group Species, indicating that these two are Fuliang Zhuye Group Species. Zhenong 117, Zhongcha 108, Wuniuzao, Gancha No. 2 and Zhenong 1 were grouped with non-Fuliang Zhuye Group Species, indicating that they are not Fuliang Zhuye Group Species.
[0096] like Figure 2 As shown in Example 2 of this scheme: Fresh leaves of two tea tree varieties were picked in Nanchang County for identification. The Fuliang Zhuye Group Species-Jinggong and the Fuliang Zhuye Group Species clustered together, indicating that the Fuliang Zhuye Group Species-Jinggong is a Fuliang Zhuye Group Species. The Zhenong 117 and the non-Fuliang Zhuye Group Species clustered together, indicating that the Zhenong 117 is not a Fuliang Zhuye Group Species.
[0097] exist Figure 3 It should be noted that the metabolite content of all samples was obtained by liquid chromatography-mass spectrometry (UPLC-QTOF / MS). The values in the table represent the mass spectrometry peak intensities of 16 characteristic metabolites, reflecting the relative abundance of each metabolite in the fresh leaves of different varieties. By comparing the metabolite peak intensities of *Zhuyeqi*, *Fuding Dabaicha*, *Wuniuzao*, and *Zhuye* cultivars from Fuliang, it can be seen that the *Zhuye* cultivars from Fuliang exhibit significantly higher accumulation levels of multiple metabolites, including cellotetrasaccharide, theophylline, proanthocyanidin C1, delphinidin-3-O-morum disaccharide glycoside, epicatechin, epigallocatechin, gallocatechin, cytidine-2′,3′-monophosphate, and 8-hydroxy-2′-deoxyguanosine, forming a metabolite distribution pattern that is significantly different from other varieties. This characterization indicates that the *Zhuye* cultivars from Fuliang have stable differences in the characteristic vectors constructed from the 16 metabolites, providing a data foundation for subsequent variety identification using metabolite characteristic vectors.
[0098] exist Figure 4It should be noted that all values in the table are peak intensities of metabolites obtained by liquid chromatography-mass spectrometry (UPLC-QTOF / MS), reflecting the relative abundance levels of sixteen characteristic metabolites in the fresh leaves of each variety. A comparison of Zhenong 117, Baihaozao, Yingshuang, Zhongcha 108, and the Fuliang Zhuye population shows that the Fuliang Zhuye population exhibits significantly higher peak intensities for multiple metabolites, including cellotetrasaccharide, theophylline, proanthocyanidin C1, delphinidin-3-O-morula disaccharide, epicatechin, epigallocatechin, gallocatechin, cytidine-2′,3′-monophosphate, and 8-hydroxy-2′-deoxyguanosine, forming a metabolic characteristic distribution distinct from other tea varieties. This figure illustrates that under summer environmental conditions, the Fuliang Zhuye population maintains a stable metabolite accumulation pattern, and the differences between it and other varieties remain significant, providing cross-seasonal data support for variety identification using characteristic metabolite vectors.
[0099] exist Figure 5 It should be noted that all values in the table are derived from metabolite peak intensities detected by UPLC-QTOF / MS, and peak intensities are used to represent the relative abundance of each metabolite in fresh leaves of different varieties. A comparison of Fuding Da Bai Cha, Zhongcha 108, Hongqi No. 1, Zhenong 113, and the Fuliang Zhuye population shows that the Fuliang Zhuye population exhibits higher levels of cellotetrasaccharides, theophylline, proanthocyanidins C1, vitexin 2′′-O-β-L-rhamnoside, delphinidin-3-O-morula disaccharide, and [unclear text - possibly a specific type of metabolite]. The peak intensities of multiple metabolites, including catechin, epigallocatechin, gallocatechin, cytidine-2′,3′-monophosphate, and 8-hydroxy-2′-deoxyguanosine, were significantly higher than those of other varieties, exhibiting a stable and distinct metabolite distribution pattern. This figure illustrates that under autumn growing conditions, the *Castanopsis fargesii* population still maintains its unique metabolite accumulation characteristics, forming a distinguishable metabolite difference from other tea varieties. This provides effective cross-seasonal data support for variety identification using metabolite feature vectors.
[0100] exist Figure 6It should be noted that all values in the table are the mass spectrometric peak intensities of the sixteen characteristic metabolites detected by liquid chromatography-mass spectrometry (UPLC-QTOF / MS), reflecting the relative content distribution of each metabolite in the fresh leaves of different tea varieties. By comparing the Gancha No. 2, Gancha No. 3, Ningzhou No. 2, Meizhan, and Fuliang Zhuye populations, it can be observed that the Fuliang Zhuye population has higher levels of cellotetrasaccharides, theophylline, proanthocyanidins C1, vitexin 2′′-O-β-L-rhamnoside, delphinidin-3-O-morula disaccharide, epicatechin, and epigallocatechin. Several characteristic metabolites, including guarcatechin, gallocatechin, cytidine-2′,3′-monophosphate, and 8-hydroxy-2′-deoxyguanosine, showed significantly higher peak intensities than other varieties. The distribution pattern of these metabolites was consistent with the results measured at other locations and in different seasons, indicating that the accumulation levels of these metabolites in the *Castanopsis fargesii* population are stable and variety-specific. The data shown in the figure further validates that the sixteen-dimensional metabolite feature vector can still maintain significant differentiation under cross-regional sampling conditions, providing reliable data support for subsequent variety identification using this feature vector.
[0101] Furthermore, refer to Figure 1 , Figure 7 , Figure 8 and Figure 9 Comparative analysis of species discrimination in *Castanopsis fargesii* populations based on different clustering analysis methods:
[0102] Experimental samples: fresh leaves of nine varieties, including: Fuliang Castanopsis spp. - standard sample, Fuliang Castanopsis spp. - Xihu Township, Fuliang Castanopsis spp. - Jinggongqiao Town, non-Fuliang Castanopsis spp. - standard sample (Castanopsis spp.), Zhenong 117, Zhongcha 108, Wuniuzao, Gancha No. 2, and Zhenong-1.
[0103] Detection data: Peak intensities of 16 characteristic metabolites in the fresh leaves of the above 9 varieties were detected;
[0104] Experimental Objective: To cluster the *Castanopsis fuliangensis* population species - standard sample, *Castanopsis fuliangensis* population species - Xihu Township, and *Castanopsis fuliangensis* population species - Jinggongqiao Town into one group in the cluster analysis diagram; and to cluster the non-*Castanopsis fuliangensis* population species - standard sample (Castanopsis fuliangensis), Zhenong 117, Zhongcha 108, Wuniuzao, Gancha 2, and Zhenong-1 into another group; effectively distinguishing the *Castanopsis fuliangensis* population species from other varieties.
[0105] Clustering analysis methods: Based on four clustering analysis methods: principal component analysis, Euclidean distance analysis, Pearson correlation analysis, and cosine correlation analysis.
[0106] Experimental results:
[0107] Figure 7While the clustering method based on principal component analysis can effectively distinguish the *Castanopsis fuliangensis* population (*Castanopsis fuliangensis* population - standard sample, *Castanopsis fuliangensis* population - Xihu Township, *Castanopsis fuliangensis* population - Jinggongqiao Town) from other varieties (non-*Castanopsis fuliangensis* population - standard sample (Castanopsis fuliangensis), Zhenong 117, Zhongcha 108, Wuniuzao, Gancha 2, Zhenong-1), the results seen on the graph are not intuitive enough and may lead to errors in actual use.
[0108] Figure 1 The hierarchical clustering method based on Euclidean distance effectively clustered the *Castanopsis fuliangensis* species-standard sample, *Castanopsis fuliangensis* species-Xihu Township, and *Castanopsis fuliangensis* species-Jinggongqiao Town into one class, and clustered the non-*Castanopsis fuliangensis* species-standard sample (*Castanopsis fuliangensis*), *Zhenong 117*, *Zhongcha 108*, *Wuniuzao*, *Gancha 2*, and *Zhenong-1* into another class; effectively identifying the *Castanopsis fuliangensis* species; this method is effective and intuitive;
[0109] Figure 8 The hierarchical clustering method based on Pearson correlations contained an error, clustering the *Castanopsis fuliangensis* population species-standard sample, the *Castanopsis fuliangensis* population species-Xihu Township, and Gancha No. 2 into one class, and clustering non-*Castanopsis fuliangensis* population species-standard sample (Castanopsis fuliangensis), Zhenong 117, Zhongcha 108, Wuniuzao, the *Castanopsis fuliangensis* population species-Jinggongqiao Town, and Zhenong-1 into another class; this method cannot effectively distinguish the *Castanopsis fuliangensis* population species.
[0110] Figure 9 The cosine-based hierarchical clustering method in this study contained an error, clustering non-Fuliang Castanopsis spp. - standard sample (Castanopsis spp. Qi), Fuliang Castanopsis spp. - Xihu Township and Gancha No. 2 into one class, and Fuliang Castanopsis spp. - standard sample, Zhenong 117, Zhongcha 108, Wuniuzao, Fuliang Castanopsis spp. - Jinggongqiao Town and Zhenong-1 into another class; this method cannot effectively distinguish Fuliang Castanopsis spp. populations.
[0111] Therefore, through comparison of the four methods, the hierarchical clustering method based on Euclidean distance is the most accurate and intuitive method.
[0112] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for discriminating between native tea plant varieties using characteristic metabolites, characterized in that, Comprise: S1, sample preparation: obtain Camellia sinensis fresh leaf sample to be identified, Castanopsis fargesii leaf population standard sample and non-Castanopsis fargesii leaf population control sample, crush the samples after microwave drying, take 50 mg of powder, add 1200 μL of 70% methanol extract solution pre-cooled to-20℃ for metabolite extraction, centrifuge and filter to obtain the test extract; S2, metabolite detection: the test extract obtained in S1 is injected into a liquid chromatography-QTOF / MS mass spectrometry system, and the sample is detected under preset chromatographic conditions and mass spectrometry conditions to obtain the mass spectrometry peak signal of the metabolites in the sample; The chromatographic conditions and mass spectrometry conditions used in S2 include: The chromatographic column is an ACQUITY-UPLC-HSS-T3-C18 type chromatographic column, the packing particle size is 1.8 microns, the inner diameter is 2.1 millimeters, and the column length is 100 millimeters; The mobile phase A is 0.1% formic acid aqueous solution, and the mobile phase B is 0.1% formic acid-acetonitrile solution; The gradient elution program is: 0-11 min, B phase 5%-90%; 11-12 min, B phase 90%; 12-12.1 min, B phase 90%-5%; 12.1-14 min, B phase 5%; The column temperature is 40℃, the flow rate is 0.4 mL / min, and the injection amount is 2 μL; The mass spectrometry adopts ESI positive and negative ion modes, the ion source voltage is 2500 / 1500V, the auxiliary gas flow is 8L / min, the ion source temperature is 325℃, and the fragmentation voltage is 135V; S3, data processing: convert the original mass spectrometry data into mzXML format by ProteoWizard, perform peak extraction, peak alignment and retention time correction by XCMS program, correct the peak area by SVR method, and filter out metabolites with a missing rate greater than 50%; S4, characteristic metabolite construction: in the corrected peak set of S3, search the public database and annotate to obtain 16 characteristic metabolites, and construct a metabolite feature vector with the peak intensity of each characteristic metabolite as the relative content; The 16 characteristic metabolites in S4 include: Cellotetraose, theabromine, procyanidin C1, vitexin 2''-O-beta-L-rhamnoside, delphinidin-3-O-myrtiloside, epicatechin, epigallocatechin, naringenin chalcone, gallocatechin, 2-methylaminoadenosine, L-alanyl-adenylic acid, cytosine, guanine, cytidine-2', 3'-monophosphate, 8-hydroxy-2'-deoxyguanosine, and valine-aspartate; S5, clustering discrimination: input the 16-dimensional characteristic metabolite vectors of the sample to be identified, the standard sample and the control sample into a multivariate statistical clustering model, if the sample to be identified and the standard sample are clustered into the same class, it is determined that the sample is Castanopsis fargesii leaf population; if it is clustered into the same class as the control sample, it is determined that the sample is not Castanopsis fargesii leaf population; In the clustering analysis of S5, the hierarchical clustering method based on Euclidean distance is used, and the branch attribution of the sample in the clustering tree is used as the identification basis, wherein: The sample is determined to be Castanopsis fargesii leaf population when it is consistent with the Castanopsis fargesii leaf population standard sample branch; The sample is determined as non-Castanopsis fargesii leaf population species when the sample and the control sample are consistent in branch. 2.The method according to claim 1, characterized in that: The metabolite extraction in S1 includes: vortexing for 30 s each time for 6 times, 30 min each time; the centrifugal condition is 13400×g, and the centrifugal time is 3 min; and the filtration uses a microporous filter film with a pore size of 0.22 μm.
3. A method of identifying indigenous tea plant varieties using characteristic metabolites as claimed in claim 1, wherein: The characteristic vector composed of the 16 characteristic metabolites maintains a stable numerical distribution mode in the samples collected in spring, summer and autumn; the discrimination results obtained based on the characteristic vector are consistent in the samples collected in Fuliang County and the samples collected in Nanchang County.
Citation Information
Patent Citations
Method for identifying tea species and determining contents of 21 characteristic components
CN104914190A
Method for identifying wild Jianghua bitter tea based on wide targeting metabonomics technology
CN115308318A