Method for determining the origin of tea based on the stable isotope ratio of EGCG monomer
Patent Information
- Application Number
- CN202610721529.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-25
- Publication Date
- 2026-08-18
- Estimated Expiration
- 2046-05-25
AI Technical Summary
[0005]本发明的一个目的是提供一种基于EGCG单体稳定同位素比值的茶叶产地判别方法,针对现有茶叶产地鉴别方法仅依靠常规指标或茶叶整体稳定同位素比值,存在针对性不强、准确性不足的问题,解决的是如何提供一种全新的鉴别思路,通过结合茶叶整体与EGCG单体稳定同位素比值,搭建基础鉴别框架,为后续精准鉴别提供技术基础的问题
本发明打破了现有茶叶产地鉴别仅依赖整体指标的局限,将茶叶整体稳定同位素比值与EGCG单体稳定同位素比值相结合,构建了全新的鉴别基础框架。EGCG作为茶叶中的特征性活性成分,其同位素比值受产地环境影响更为显著,结合两者数据能显著提升鉴别针对性。同时,明确了EGCG单体的制备核心步骤,为后续同位素比值测定提供了可靠保障,为整个鉴别方法的实施奠定了坚实基础,可有效弥补现有方法针对性和准确性不足的缺陷。
Smart Images

Figure CN122262927B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of food traceability and quality identification technology. More specifically, this invention relates to a method for determining the origin of tea based on the stable isotope ratio of EGCG monomers. Background Technology
[0002] Green tea is a traditional Chinese tea, ranking first in both production and sales among the six major tea categories, and is widely consumed for its flavor and nutritional value. Tea trees are perennial economic crops, and the quality and efficacy of tea vary depending on factors such as the growing environment (climate, soil conditions, topography, etc.), variety, and processing methods. The market frequently sees teas from other regions being passed off as geographical indication products. Therefore, establishing effective methods for tea origin certification is of practical significance for protecting regional brands and ensuring tea quality and safety. Existing tea origin identification techniques can be mainly divided into three categories: sensory evaluation, chemical composition analysis, and stable isotope ratio analysis.
[0003] Stable isotope ratio analysis is a source tracing technique developed in recent years. Most stable isotopes originate naturally and are minimally affected by human activity; the analytical basis is the isotope fractionation effect. Stable isotopes in tea (such as δ¹⁸O₂)... 13 C、δ 15 N, δ 2 H, δ 18 The composition of isotopes (O) is related to its growth environment. Current techniques typically measure the stable isotope ratios of the entire tea tissue (e.g., dried tea powder) and combine this with statistical models such as linear discriminant analysis to determine the origin. Practice shows that this method is effective in identifying origins with significant geographical and climatic differences. However, when the latitude, altitude, and rainfall of the tea-producing areas are similar, the ranges of the overall stable isotope ratios overlap, significantly reducing the accuracy of identification and making it difficult to meet traceability requirements. The reason for this is that the overall isotope value of tea is a mixed average of the isotopic signals of all components in the tea (including cellulose, proteins, and small molecule metabolites). Different compounds exhibit significant differences in their synthetic pathways, physiological turnover rates, and sensitivity to changes in the external environment. Overall averaging may dilute or mask the specific isotopic information of certain characteristic components that are more sensitive to the origin environment. Therefore, relying solely on the overall stable isotope ratios of tea has limitations in distinguishing origins from geographically similar areas.
[0004] Epigallocatechin gallate (EGCG) is a unique and abundant catechin active ingredient in tea. However, there is currently no technical solution that combines the stable isotope ratio of EGCG monomers with the overall isotope ratio of tea leaves for accurate identification of tea origin. Existing methods rely solely on overall indicators or conventional chemical components, resulting in insufficient accuracy when dealing with tea-producing regions with similar environmental backgrounds. Therefore, there is an urgent need for a tea origin traceability method that can integrate characteristic monomer isotope information to improve the accuracy and reliability of identification. EGCG is composed of elements such as carbon, hydrogen, and oxygen, and its biosynthetic pathway is significantly regulated by environmental factors (such as light, moisture, and temperature). Summary of the Invention
[0005] One objective of this invention is to provide a method for identifying the origin of tea based on the stable isotope ratio of EGCG monomers. This addresses the problem that existing methods for identifying the origin of tea rely solely on conventional indicators or the overall stable isotope ratio of tea leaves, which suffer from insufficient specificity and accuracy. The invention aims to provide a novel identification approach by combining the overall tea leaf with the stable isotope ratio of EGCG monomers to establish a basic identification framework, thus providing a technical foundation for subsequent accurate identification.
[0006] To address the problem that relying on stable isotope ratios cannot achieve accurate origin identification, this invention analyzes the detected isotope ratios using a linear discriminant model, transforming abstract detection data into directly identifiable origin results, thus improving the operability of the identification process.
[0007] To address the issues of misjudging origin by linear discriminant function values, which are prone to misjudgment due to similar values and lack a discrimination threshold and secondary verification mechanism, this invention uses difference comparison to screen out suspected samples, avoiding the limitations of a single discrimination method, reducing the probability of misjudgment, and improving the accuracy of identification.
[0008] To address the problem that suspected samples cannot be accurately identified when the difference value does not reach the threshold due to the lack of effective secondary discrimination methods, this invention further filters suspected samples by calculating difference features and introducing a random forest classifier. At the same time, it clarifies the subsequent processing method when the votes are consistent, thus improving the identification process.
[0009] Regarding the issue of identical voting ratios and Y in the secondary discrimination, i In extreme cases where values are the same, the lack of a final judgment basis leads to the inability to identify or misjudgment of some samples. This invention addresses the problem of how to introduce new judgment parameters, clarify the judgment logic and anomaly handling methods in extreme cases, fill judgment loopholes, and ensure that all samples can obtain reasonable judgment results or clear labeling.
[0010] To address the problem that the lack of a scientific and unified method for determining the discrimination threshold T leads to differences in threshold settings by different users, affecting the consistency and accuracy of identification results, this invention uses leave-one-out cross-validation, combined with the difference data of correctly identified samples, to determine a reasonable threshold and ensure the uniformity of identification standards.
[0011] To achieve these objectives and other advantages according to the present invention, a method for determining the origin of tea based on the stable isotope ratio of EGCG monomers is provided, comprising the following steps: Includes the following steps: Determination of the overall stable isotope ratio δ in tea samples 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 and the δ¹² value of stable isotope ratios of EGCG monomers 13 C EGCG、 δ 2 H EGCG and δ 18 O EGCG The preparation method of EGCG monomer is as follows: tea sample is made into tea powder, epigallocatechin gallate ester EGCG monomer is extracted and purified from tea powder, and then the stable isotope ratio of EGCG monomer is determined. Based on the above stable isotope ratios, the origin of tea samples can be identified. The method for identifying the origin of tea is as follows: The overall stable isotope ratio δ0.05 is calculated. 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 and the δ¹² value of stable isotope ratios of EGCG monomers 13 C EGCG δ 2 H EGCG and δ 18 O EGCG The input is a pre-built linear discriminant model for different production areas. The pre-built linear discriminant model for different production areas is as follows: Y i =a i δ 13 C 茶叶 +b i δ 15 N 茶叶 +c i δ 2 H 茶叶 +di δ 18 O 茶叶 +e i δ¹³C EGCG +f i δ 2 H EGCG +g i δ 18 O EGCG +h i ; Among them, a i b i c i d i e i f i g i h i Let be a constant obtained by training from tea samples from known origins using Fisher linear discriminant analysis, where i represents the origin number; During discrimination, the seven isotope ratios of the sample to be tested are substituted into the discrimination model for each origin to obtain Y. i Value, based on Y i The value is used to determine the place of origin.
[0012] Preferably, the δ of the tea sample to be tested is... 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 δ¹³C EGCG δ 2 H EGCG δ 18 O EGCG Substituting the seven stable isotope ratios into the linear discriminant function for each origin yields the discriminant function value Y for each origin. i ; Calculate the difference ΔY between the largest and second largest discriminant function values; compare the difference ΔY with the discrimination threshold T. If ΔY is greater than T, then the origin corresponding to the largest discriminant function value is determined to be the origin of the tea sample to be tested; if ΔY is less than or equal to T, then proceed to the second-level discrimination step; wherein, the discrimination threshold T is obtained by leave-one-out cross-validation on tea samples from known origins, and T is taken as the 10th percentile of the ΔY value of all tea samples from known origins when correctly discriminated.
[0013] Preferably, the second-level discrimination step includes: S1. Calculate the three difference characteristics: △ 13 C、△ 2 H and △ 18 O, where: △13 C=δ 13 C EGCG -δ 13 C 茶叶 ; △ 2 H=δ 2 H EGCG -δ 2 H 茶叶 ; △ 18 O=δ 18 O EGCG -δ 18 O 茶叶 ; S2. Calculate the three difference characteristics and the overall stable isotope ratio δ. 15 N tea leaves have four features. The input is a pre-built second-level random forest classifier, consisting of 200 decision trees. The output is the voting ratio for each production area. If two or more production areas have the same voting ratio and all have the highest voting ratio, then the linear discriminant function value Y calculated for these production areas is obtained. i Choose Y i The region of origin with the highest value is used as the final judgment result; The second-level random forest classifier is trained using only tea samples from known origins that satisfy the difference ΔY ≤ T.
[0014] Preferably, in step S2, if Y from different origins i If the values are the same, perform the following steps: A1. Calculate the difference in deuterium excess parameter Δd of the tea sample to be tested. Δd is calculated according to the following formula: △d=(δ 2 H EGCG -k2δ 18 O EGCG )-(δ 2 H 茶叶 -k1δ 18 O 茶叶 ); Wherein, coefficient k1 is the δ value obtained from tea samples from known origins through linear regression analysis. 18 O 茶叶 and δ 2 H 茶叶 The slope obtained from the fitted data, and the coefficient k2, are the δ values obtained from all tea samples of known origins through linear regression analysis. 2 H EGCG and δ 18 O EGCG The slope obtained by fitting the data; A2. For each candidate origin, a subset of boundary samples that satisfy the discriminant function value ΔY≤T is pre-selected from the training samples of that candidate origin. The median M of the Δd values of all samples in this subset is then calculated. j If the number of samples in the boundary sample subset of a candidate origin is less than 3, then the candidate origin is removed from the candidate list. A3. Compare the Δd of the tea sample to be tested with the median M of each candidate origin. j The absolute difference between them |△d-M j | The candidate origin with the smallest absolute difference is selected as the final discrimination result; If two or more candidate production areas have the same absolute difference and both are the minimum, or if all candidate production areas are removed due to insufficient sample size, the tea sample to be tested will be marked as pending verification.
[0015] Preferably, the discrimination threshold T is determined through the following steps: B1. Perform leave-one-out cross-validation on all tea samples from known origins: Each time, remove one sample as a validation sample, use the remaining samples to train the linear discriminant function, and calculate the δ of the validation sample. 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 δ 13 C EGCG δ 2 H EGCG δ 18 O EGCG The maximum discriminant function value Y obtained by substituting the seven stable isotope ratios into the discriminant function for each origin is... max And the second largest discriminant function value Y second The difference between them is △Y'=Y max -Y second ; B2. Record the true origin of the verification sample and determine Y. max Does the corresponding place of origin match the actual place of origin? If they match, mark the verification sample as a single-level correctly judged sample and record the ΔY' value of the sample. If they do not match, mark the verification sample as a single-level incorrectly judged sample and do not record its ΔY' value. B3. Iterate through all tea samples from known origins, collect the ΔY' values of all correctly judged samples at the single level, calculate the 10th percentile of these ΔY' values, and use this percentile as the discrimination threshold T; Among them, single-level correct discrimination means that the origin corresponding to the maximum discriminant function value in leave-one-out cross-validation is consistent with the true origin.
[0016] Preferably, in step one, the method for extracting EGCG monomers is as follows: The crude extract was obtained by heating and ultrasonic-assisted extraction of tea powder using an ethanol-water solution. The crude extract was extracted sequentially with dichloromethane and ethyl acetate, and the supernatant was collected from each extraction to obtain the extract. After purification, the EGCG monomer was obtained.
[0017] The present invention has at least the following beneficial effects: This invention breaks through the limitations of existing methods that rely solely on overall indicators for tea origin identification. It combines the overall stable isotope ratios of tea with the stable isotope ratios of EGCG monomers, constructing a novel identification framework. As a characteristic active ingredient in tea, the isotope ratio of EGCG is significantly affected by the origin environment; combining data from both significantly improves the specificity of the identification. Furthermore, the core steps for EGCG monomer preparation are clearly defined, providing a reliable guarantee for subsequent isotope ratio determination and laying a solid foundation for the implementation of the entire identification method. This effectively overcomes the shortcomings of existing methods in terms of specificity and accuracy.
[0018] This invention provides a specific and operable method for determining the origin of a site. It transforms abstract stable isotope ratio data into directly identifiable Y origin values, determining the origin by comparing the maximum values. The logic is clear, the operation is simple, and the difficulty of identification is reduced. This method fully utilizes the detected ratio data of the seven isotopes, avoiding the limitations of a single indicator, improving the intuitiveness and operability of origin identification, ensuring that different users can perform identification according to a unified logic, guaranteeing the consistency of identification results, and further improving identification efficiency.
[0019] This invention clarifies the construction criteria and discriminant function form of the linear discriminant model. Specific constants are obtained through Fisher linear discriminant analysis training, providing a unified and standardized calculation basis for the discrimination of different production areas, avoiding subjectivity and arbitrariness in model construction. Seven isotope ratios are used as discriminant features, comprehensively reflecting the regional correlation of tea and EGCG monomers, ensuring the scientific validity and accuracy of the discriminant function. Simultaneously, the calculation process during discrimination is clarified, making the entire identification process more standardized, reducing human error, and providing standardized technical support for subsequent accurate discrimination.
[0020] This invention effectively avoids the problem caused by Y by setting a discrimination threshold T and a two-stage discrimination step. iThe problem of misjudgment caused by close differences in values has been significantly improved in terms of identification accuracy. The discrimination threshold T is obtained through leave-one-out cross-validation, combined with actual discrimination data of samples from known origins. The value is scientifically reasonable and can accurately screen out suspected samples. For samples whose differences do not reach the threshold, a two-stage discrimination step is introduced, realizing stratified identification of preliminary discrimination, suspected screening, and secondary verification. This further reduces the probability of misjudgment, makes the identification results more reliable, and meets the needs of accurate traceability in practical applications.
[0021] This invention improves the two-level discrimination system by calculating the difference characteristics, further exploring the correlation differences between EGCG monomers and the overall isotope ratios of tea leaves, combined with δ¹⁸O₂. 15 The N-tea index enhances the targeting and accuracy of the secondary discrimination. The random forest classifier, composed of 200 decision trees, boasts high classification accuracy and strong anti-interference capabilities. Furthermore, it is trained using only boundary samples, precisely adapting to the discrimination needs of suspected samples. Simultaneously, the handling method for cases with identical voting ratios is clarified, filling gaps in the secondary discrimination process and ensuring accurate identification of suspected samples, further improving the integrity and reliability of the entire identification process.
[0022] This invention addresses the extreme cases in the secondary discrimination by introducing the deuterium excess parameter difference Δd as a supplementary discrimination index, further improving discrimination accuracy and resolving the issues related to voting ratio and Y. i The challenge lies in identifying samples with identical values. This is addressed by filtering a boundary subset of samples and calculating the median M. j This ensures the scientific rigor and relevance of the judgment criteria, while setting sample quantity screening conditions to avoid judgment bias caused by insufficient sample size. The handling methods for abnormal situations are clearly defined, and situations requiring further verification are explicitly marked, ensuring the rigor of the identification results, avoiding misjudgments caused by forced judgment, and further improving the practicality and reliability of the entire identification method.
[0023] This invention provides a standardized and scientific process for determining the discrimination threshold T. Through leave-one-out cross-validation, it fully utilizes actual data from samples of known origins to ensure that the threshold value closely matches actual identification needs, avoiding deviations caused by manually setting the threshold. Calculating the 10th percentile using only the ΔY' value of correctly identified samples at a single level effectively eliminates interference from erroneous data, making the threshold more reasonable and reliable. A unified threshold determination standard ensures consistency in identification standards across different users and batches of samples, improving the comparability and authority of identification results and providing a guarantee for the standardized implementation of the entire identification method.
[0024] Other advantages, objectives and features of the present invention will become apparent in part from the following description, and in part from those skilled in the art through study and practice of the invention. Attached Figure Description
[0025] Figure 1 The tea leaves in Embodiment 1 of the present invention 2 H tea oxygen isotope δ 18 O tea The linear fitting plot, where, Figure 1 δ 2 H tea For δ 2 H 茶叶 δ 18 O tea For δ 18 O 茶叶 ; Figure 2 The tea leaves in Embodiment 1 of the present invention 2 H tea With δ in EGCG monomer 2 H EGCG The linear fitting plot, where, Figure 2 δ 2 H tea For δ 2 H 茶叶 . Detailed Implementation
[0026] The present invention will be further described in detail below with reference to embodiments, so that those skilled in the art can implement it based on the description.
[0027] This invention provides a method for determining the origin of tea based on the stable isotope ratio of EGCG monomers, comprising the following steps: Determination of the overall stable isotope ratio δ in tea samples 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 and the δ¹² value of stable isotope ratios of EGCG monomers 13 C EGCG、 δ 2 H EGCG and δ 18 O EGCG The preparation method of EGCG monomer is as follows: tea sample is made into tea powder, epigallocatechin gallate ester EGCG monomer is extracted and purified from tea powder, and then the stable isotope ratio of EGCG monomer is determined. Based on the above stable isotope ratios, the origin of tea samples can be identified.
[0028] In the above technical solution, during the pretreatment of tea samples, fresh or dried tea leaves can be selected as the experimental subject. After removing impurities, the tea sample is pulverized in a pulverizing device. The pulverized tea leaves can be passed through a 40-80 mesh sieve to obtain uniform tea powder. The powder can be stored in a desiccator for later use to avoid moisture affecting subsequent detection. During the preparation of EGCG monomer, the amount of tea powder used can be 5-20g, and an ethanol aqueous solution with a volume fraction of 50%-80% can be used for extraction.
[0029] Stable isotope ratios can be determined using a stable isotope ratio mass spectrometer. This equipment can be placed in a temperature- and humidity-controlled laboratory, with the temperature maintained at 20-25℃ and the humidity at 40%-60%. The equipment needs to be preheated for 30-60 minutes to ensure detection accuracy. (The last sentence appears to be incomplete and refers to the determination of the overall stable isotope ratio δ in tea leaves.) 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 Tea samples can be freeze-dried at temperatures ranging from -40°C to -20°C for 12-24 hours. After processing, 1-5 mg of sample is placed in a sample vial for testing. The stable isotope ratio δ of EGCG monomers is then determined. 13 C EGCG δ 2 H EGCG and δ 18 O EGCG The extracted and purified EGCG monomer was freeze-dried to obtain a solid powder, which was then detected using the same method as the overall stable isotope determination method for tea leaves.
[0030] When identifying the place of origin, first organize the data of the seven stable isotope ratios measured above, remove outliers, and then organize and archive them.
[0031] By adopting this technical solution, the present invention can achieve preliminary identification of tea origin, providing a foundation for subsequent accurate identification, ensuring that the identification process has clear steps to support it, and obtaining accurate and reliable stable isotope ratio data that can reflect the origin-related characteristics of tea and EGCG monomers. It solves the problem of insufficient targeting of existing identification methods. At the same time, the EGCG monomer preparation method is simple and feasible, and the equipment and materials used are all existing commercially available products, which facilitates practical promotion and application.
[0032] In another technical solution, the method for identifying the origin of tea is: using the overall stable isotope ratio δ 13 C 茶叶 δ 15 N 茶叶 δ2 H 茶叶 δ 18 O 茶叶 and the δ¹² value of stable isotope ratios of EGCG monomers 13 C EGCG δ 2 H EGCG and δ 18 O EGCG The input is fed into a pre-built linear discriminant model for different production areas to obtain the Y values for different production areas. 产地 Value, with the largest Y 产地 The place of origin corresponding to the value is used as the discrimination result.
[0033] In the above technical solution, when constructing the linear discriminant model, computer programming software can be used to construct the model. The training samples required for model construction can be selected from tea samples from different production areas, with 30-60 samples from each production area to ensure the reliability of model training. During the model construction process, the isotope ratio data of the training samples can be standardized using the Z-score standardization method.
[0034] During origin determination, the processed isotope ratio data of the samples to be tested are input into a pre-constructed linear discrimination model. The model can be installed in a regular computer, which can be placed on a laboratory workbench. It can be connected to a stable isotope ratio mass spectrometer via a data cable to achieve direct data transmission. When the model is running, the computer processor should be an Intel Core i5 or higher, and the memory should be 8GB or higher to ensure smooth model operation. After inputting the data, the model can calculate the Y origin values for different origins and output all Y origin values and corresponding origin information after the calculation is complete.
[0035] The process is as follows: First, the data of the seven stable isotope ratios of the tea sample to be tested are organized. After ensuring that the data is accurate, the data is input into a linear discriminant model that has been built in advance on the computer. The model processes and calculates the input data and outputs the Y origin value corresponding to each candidate origin. The operator compares all the Y origin values to determine the origin of the tea sample to be tested.
[0036] By adopting this technical solution, the present invention can provide a specific and operable method for determining the origin of the product, transforming abstract isotope data into directly identifiable results, reducing the difficulty of identification, ensuring that different operators can perform identification according to a unified process, improving identification efficiency and consistency, and at the same time, the equipment and software used for model construction are all existing products, requiring no additional research and development, reducing implementation costs, and effectively solving the problem of poor operability of existing identification methods.
[0037] In another technical solution, the linear discriminant model for different origins uses δ 13 C茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 δ¹³C EGCG δ 2 H EGCG δ 18 O EGCG As a discriminant feature, a linear discriminant model of the following form is pre-established for each candidate origin: Y i =a i δ 13 C 茶叶 +b i δ 15 N 茶叶 +c i δ 2 H 茶叶 +d i δ 18 O 茶叶 +e i δ¹³C EGCG +f i δ 2 H EGCG +g i δ 18 O EGCG +h i ; Among them, a i b i c i d i e i f i g i h i Y is a constant obtained by training tea samples from known origins using Fisher linear discriminant analysis, where i represents the origin number. During discrimination, the seven isotope ratios of the sample are substituted into the discriminant function for each origin to obtain Y. i Value, based on Y i The value is used to determine the place of origin.
[0038] In the above technical solution, when constructing the linear discriminant model, for each candidate origin, tea samples from that origin are pre-selected as training samples. The training samples can undergo three parallel tests, and the average of the test data is taken as the training data. The training process can use Fisher's linear discriminant analysis method, implemented through computer software such as SPSS or MATLAB. The constant 'a' of the discriminant function for each origin is obtained through this method. i b i c i di e i f i g i h i These constants can range from -10 to 10. After training, the discriminant functions for each origin are stored in the linear discriminant model for easy subsequent use. The discriminant model takes the following form: Y i =a i δ 13 C 茶叶 +b i δ 15 N 茶叶 +c i δ 2 H 茶叶 +d i δ 18 O 茶叶 +e i δ¹³C EGCG +f i δ 2 H EGCG +g i δ 18 O EGCG +h i , where i represents the place of origin number, which can be set sequentially as 1, 2, 3...
[0039] The general process is as follows: First, seven stable isotope ratios are determined as discriminant features. Tea samples from various producing areas are selected as training samples. The isotope ratio data of the training samples are parallelly tested and processed. The linear discriminant model for each producing area is trained using Fisher's linear discriminant analysis method. The constants in the model are determined. During discrimination, the seven isotope ratios of the sample to be tested are substituted into the discriminant model for each producing area to calculate the Yg for each area. i Value, and then based on Y i The value is used to determine the place of origin.
[0040] By adopting this technical solution, the present invention can clarify the construction criteria and discrimination logic of the linear discriminant model, so that the discrimination of different production areas has a unified calculation basis, avoids the subjectivity of model construction, and improves the accuracy and standardization of the identification results. At the same time, the training method is mature and the software used is readily available, which can reduce the difficulty of model construction and provide reliable technical support for subsequent accurate discrimination.
[0041] In another technical solution, the δ of the tea sample to be tested is... 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ18 O 茶叶 δ 13 C EGCG δ 2 H EGCG δ 18 O EGCG Substituting the seven stable isotope ratios into the linear discriminant function for each origin yields the discriminant function value Y for each origin. i ; Calculate the difference ΔY between the largest and second largest discriminant function values; compare the difference ΔY with the discrimination threshold T. If ΔY is greater than T, then the origin corresponding to the largest discriminant function value is determined to be the origin of the tea sample to be tested; if ΔY is less than or equal to T, then proceed to the second-level discrimination step; wherein, the discrimination threshold T is obtained by leave-one-out cross-validation on tea samples from known origins, and T is taken as the 10th percentile of the ΔY value of all tea samples from known origins when correctly discriminated.
[0042] In the above technical solution, the δ of the tea sample to be tested is... 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 δ 13 C EGCG δ 2 H EGCG δ 18 O EGCG The seven stable isotope ratios are substituted into the linear discriminant function for each origin. The accuracy of the data must be verified before each ratio is substituted to avoid input errors. The calculation process can be automated using computer software. Each Y... i The value can be calculated in three parallel steps, and the average value is taken as the final Y. i value.
[0043] When calculating △Y, first start from all Y... i Filter the values to find the largest discriminant function value Y. max and the second largest discriminant function value Y second , △Y=Y max -Y second If there are multiple Y i If the values are the same and the maximum value is reached, the data is rechecked and recalculated. The discrimination threshold T is obtained through leave-one-out cross-validation. The value of T is the 10th percentile of the ΔY value of all tea samples from known origins when correctly discriminated. Through actual verification, the specific value of T can be 0.8-1.2.
[0044] By adopting this technical solution, the present invention can effectively avoid the problem caused by Y. iTo prevent misjudgments caused by similar values, a scientific and reasonable discrimination threshold is set to screen out suspected samples for secondary verification, thereby improving the accuracy of identification. The method for determining the threshold is standardized to ensure that the identification standards are consistent across different batches and users. At the same time, the calculation process is simple and can be completed automatically by software, reducing human error and further improving the identification process.
[0045] In another technical solution, the second-level discrimination step includes: S1. Calculate the three difference characteristics: △ 13 C、△ 2 H and △ 18 O, where: △ 13 C=δ 13 C EGCG -δ 13 C 茶叶 ; △ 2 H=δ 2 H EGCG -δ 2 H 茶叶 ; △ 18 O=δ 18 O EGCG -δ 18 O 茶叶 ; S2. Calculate the three difference characteristics and the overall stable isotope ratio δ. 15 N 茶叶 There are four features. The input is a pre-built second-level random forest classifier, which consists of 200 decision trees. The output is the voting ratio of each production area. If two or more production areas have the same voting ratio and all of them have the highest voting ratio, then the linear discriminant function value Y calculated for these production areas is obtained. i Choose Y i The region of origin with the highest value is used as the final judgment result; The second-level random forest classifier is trained using only tea samples from known origins that satisfy the difference ΔY ≤ T.
[0046] In the above technical solution, the δ of the tea sample to be tested is... 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 δ 13 C EGCG δ 2 H EGCG δ18 O EGCG Data, calculate △ 13 C、△ 2 H and △ 18 O, the calculation process can be completed automatically using Excel software, and then compared with δ. 15 N 茶叶 These four features serve as the four discriminative features. In practical applications, a sufficient number of samples from known production areas (e.g., more than 100 samples per production area) can be collected to ensure the quantity of the boundary sample subset, thereby training a stable random forest classifier. Those skilled in the art can also adjust the parameters of the random forest (e.g., reduce the number of decision trees, limit the tree depth) according to the actual number of boundary samples to avoid overfitting. When constructing the random forest classifier, Python programming software can be used. The classifier consists of 200 decision trees, with a depth of 5-10 layers, a minimum number of sample splits of 2, and a minimum number of leaf nodes of 1. The training samples of the classifier are selected only from tea samples from known production areas that satisfy the difference ΔY≤T. The number of training samples for each production area can be 10-20 (e.g., with 60 total samples per production area, the number of boundary samples (ΔY≤T) selected is 8-15, which meets the training requirement of 10-20 samples). Five-fold cross-validation can be set during the training process. The classifier can be installed on a regular computer, on the same device as the linear discriminant model, which facilitates data transfer and retrieval. When applying it, four discriminant features are input into the classifier, and the classifier outputs the voting ratio of each production area after running for 1-3 minutes.
[0047] After entering the second level of discrimination, the three difference features Δ are calculated first. 13 C、△ 2 H and △ 18 O, the calculation process can be completed automatically using Excel software, and then compared with δ. 15 N 茶叶 The data is input into a pre-built random forest classifier. The classifier analyzes the data using 200 decision trees and outputs the voting percentage for each region. If two or more regions have the same number of votes and all receive the highest number of votes, then the Y-values of these regions are retrieved. i Value, select Y i The value with the largest value is used as the final judgment result.
[0048] By adopting this technical solution, the present invention can improve the two-level discrimination system, further refine the screening of suspected samples, and the random forest classifier has strong anti-interference ability and high classification accuracy. Training with only boundary samples can improve the specificity. At the same time, it clarifies the handling method when the votes are consistent, fills the discrimination loopholes, ensures that suspected samples can be accurately identified, and improves the reliability of the entire identification method.
[0049] In another technical solution, in step S2, if Y from different places of origin i If the values are the same, perform the following steps: A1. Calculate the difference in deuterium excess parameter Δd of the tea sample to be tested. Δd is calculated according to the following formula: △d=(δ 2 H EGCG -k2δ 18 O EGCG )-(δ 2 H 茶叶 -k1δ 18 O 茶叶 ); Wherein, coefficient k1 is the δ value obtained from tea samples from known origins through linear regression analysis. 18 O 茶叶 and δ 2 H 茶叶 The slope obtained from the fitted data, and the coefficient k2, are the δ values obtained from tea samples from known origins through linear regression analysis. 2 H EGCG and δ 18 O EGCG The slope obtained by fitting the data; A2. For each candidate origin, pre-select a subset of boundary samples from the training samples of that candidate origin that satisfy the discriminant function value ΔY≤T (here, for the training samples, ΔY refers to the difference between the largest and second largest discriminant function values obtained after substituting the sample into the discriminant functions of each origin). Calculate the median M of the Δd values of all samples in this subset of boundary samples. j If the number of samples in the boundary sample subset of a candidate origin is less than 3, then the candidate origin is removed from the candidate list. A3. Compare the Δd of the tea sample to be tested with the median M of each candidate origin. j The absolute difference between them |△d-M j | The candidate origin with the smallest absolute difference is selected as the final discrimination result; If two or more candidate production areas have the same absolute difference and both are the minimum, or if all candidate production areas are removed due to insufficient sample size, the tea sample to be tested will be marked as pending verification.
[0050] In the above technical solution, Δd is calculated based on the measured δ. 2 H EGCG δ 18 O EGCG δ 2 H 茶叶 δ 18 O 茶叶 Data, according to the formula △d=(δ 2 H EGCG -k2δ18 O EGCG )-(δ 2 H 茶叶 -k1δ 18 O 茶叶 ); Calculations are performed. Methods for obtaining k1 and k2: Through the δ... 2 H 茶叶 With δ 18 O 茶叶 δ 2 H EGCG With δ 18 O EGCG The data were obtained through linear regression analysis. As an example of this invention, multiple samples covering major tea-producing areas in Sichuan (Ya'an, Yibin, Dujiangyan, etc.), Zhejiang, and Fujian, across different seasons, were used to obtain constants k1=5.37 and k2=6.01. Verification showed that within the covered production areas, the δ0.01 values in tea samples from different production areas and seasons were consistent. 2 H and δ 18 The linear relationship of O is highly consistent (slope standard deviation < 0.15), and the discrimination results are not sensitive to fluctuations in the slope within a reasonable range. Therefore, using this slope as a fixed constant simplifies the model, avoids repeated fitting calculations, and ensures the accuracy and reproducibility of origin discrimination, making it suitable for standardized testing and widespread application.
[0051] When selecting the boundary sample subset for candidate origins, for each candidate origin, samples satisfying the discriminant function value ΔY≤T are selected from its training samples to form a boundary sample subset. If the number of samples in this subset is less than 3, the candidate origin is removed from the candidate list to avoid discrimination bias caused by insufficient sample size. Median M j During calculation, the Δd values of each valid boundary sample subset are sorted, and the median value after sorting is taken as M. j If the number of samples in the subset is even, then the average of the two middle values is taken as M. j M j Keep 4 decimal places. The calculation process can be completed using SPSS software.
[0052] By adopting this technical solution, the present invention can solve the identification problem under extreme conditions, introduce △d as a supplementary indicator to improve the discrimination accuracy, set sample size screening conditions to ensure the reliability of the discrimination basis, clarify the handling method for abnormal situations, ensure the rigor of the identification results, avoid misjudgment caused by forced discrimination, and further improve the practicality of the identification method.
[0053] In another technical solution, the discrimination threshold T is determined through the following steps: B1. Perform leave-one-out cross-validation on all tea samples from known origins: Each time, remove one sample as a validation sample, use the remaining samples to train the linear discriminant function, and calculate the δ of the validation sample. 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 δ 13 C EGCG δ 2 H EGCG δ 18 O EGCG The maximum discriminant function value Y obtained by substituting the seven stable isotope ratios into the discriminant function for each origin is... max And the second largest discriminant function value Y second The difference between them is △Y'=Y max -Y second ; B2. Record the true origin of the verification sample and determine Y. max Does the corresponding place of origin match the actual place of origin? If they match, mark the verification sample as a single-level correctly judged sample and record the ΔY' value of the sample. If they do not match, mark the verification sample as a single-level incorrectly judged sample and do not record its ΔY' value. B3. Iterate through all tea samples from known origins, collect the ΔY' values of all correctly judged samples at the single level, calculate the 10th percentile of these ΔY' values, and use this percentile as the discrimination threshold T; Among them, single-level correct discrimination means that the origin corresponding to the maximum discriminant function value in leave-one-out cross-validation is consistent with the true origin.
[0054] In the above technical solution, when implementing the leave-one-out cross-validation method, all tea samples from known production areas are selected as validation objects. The number of samples from each production area can be 30-60. One sample is removed each time as the validation sample, and the linear discriminant function is retrained using the remaining samples. Then, the seven stable isotope ratios of the validation samples are substituted into the discriminant function of each production area to calculate Y. max With Y second The difference △Y' = Y max - Y second Each verification sample was verified three times, and the average value was taken as the final ΔY' value.
[0055] When identifying sample markings, record and verify the true origin of the sample, and compare it with Y. maxIf the corresponding origin matches the actual origin, it is marked as a single-level correctly identified sample, and its ΔY' value is recorded. If they do not match, it is marked as a single-level incorrectly identified sample, and its ΔY' value is not recorded. The marking process can be recorded using Excel software for easy data processing later. When collecting ΔY' values, all tea samples from known origins are traversed, and the ΔY' values of all single-level correctly identified samples are summarized. After removing outliers, the 10th percentile of these ΔY' values is calculated using SPSS software. This percentile is the discrimination threshold T, and the specific value of T can be between 0.8 and 1.2.
[0056] By adopting this technical solution, the present invention can provide a standardized threshold determination process, ensuring that the value of T is scientific and reasonable, meets actual identification needs, eliminates interference from erroneous identification data, makes the identification standards of different users and different batches of samples consistent, improves the comparability and authority of identification results, and provides a guarantee for the standardized implementation of the entire identification method.
[0057] In another technical solution, the extraction method of EGCG monomers in step one is as follows: The crude extract was obtained by heating and ultrasonic-assisted extraction of tea powder using an ethanol-water solution. The crude extract was extracted sequentially with dichloromethane and ethyl acetate, and the supernatant was collected from each extraction to obtain the extract. After purification, the EGCG monomer was obtained.
[0058] In the above technical solution, when extracting crude EGCG extract, take 5-20g of tea powder and add 50-200mL of ethanol aqueous solution with a volume fraction of 50%-80%. The ethanol aqueous solution can be prepared by mixing food-grade ethanol and deionized water. During extraction, a constant temperature water bath can be used for heating, and the heating temperature can be 40-60℃. At the same time, ultrasonic equipment is used for assisted extraction, and the ultrasonic power can be 100-300W, and the ultrasonic time can be 30-60min.
[0059] By adopting this technical solution, the present invention can provide a simple and feasible method for the extraction and purification of EGCG monomers. The equipment and materials used are all existing commercially available products. The operation is simple and the cost is low. It can obtain high-purity EGCG monomers, ensuring the accuracy of subsequent isotope ratio determination and providing reliable sample support for the implementation of the entire identification method.
[0060] <Example 1> 1. Sample collection and related information Green tea buds and leaves (one bud and two leaves) were collected from Yibin, Ya'an, and Dujiangyan cities in Sichuan Province in spring (March), summer (June), and autumn (October) of 2024. Twenty tea samples were collected from each production area per season, for a total of 180 samples. Longitude, latitude, altitude, average annual temperature, and average annual precipitation data were recorded and listed in Table 1. The tea samples were rinsed with deionized water and then dried to constant weight in an oven at 60℃. The dried samples were then ground and passed through a 100-mesh nylon sieve to obtain a uniform powder.
[0061] surface Geographical, climatic, and soil information of different tea-producing regions Note: Annual average temperature and annual rainfall data are from the local meteorological bureau; soil type data are from the China Soil Database. 2. Preparation and purity verification of catechin monomers 2.1 Preparation of EGCG monomers Weigh 5 g of green tea powder, add 50 ml of 80% ethanol, extract in a 60℃ constant temperature water bath for 30 min, sonicate in an ultrasonic instrument for 30 min, cool, and filter to separate the residue to obtain crude extract. Recover the ethanol using a rotary evaporator at 55℃ and concentrate to near dryness. Dissolve in 10 ml of water, extract three times with dichloromethane to remove caffeine, collect the upper layer (aqueous layer), and extract three more times with ethyl acetate. Collect the upper ethyl acetate layer, combine the extracts, concentrate by rotary evaporation, and dissolve in the lower phase (20-30 mL) of the solvent system. Inject the extract into a high-speed countercurrent chromatograph. The mobile phase solvent system is ethyl acetate: anhydrous ethanol: water = 10:1:10, and the flow rate is 10 mL / min.
[0062] 2.2 EGCG monomer purity verification Catechins in tea leaves were determined using a high-performance liquid chromatograph (HPLC) equipped with an Agilent C18 column (250 mm × 4.6 mm, 5.0 μm). 0.2 g of EGCG monomer was weighed into a 15 mL centrifuge tube, 10 mL of 99.9% methanol was added, and the tube was sonicated at 55 °C for 30 min, followed by centrifugation at 8000 r / min for 10 min. The supernatant was then filtered through a 0.45 μm organic membrane. The detection parameters were as follows: Mobile phase: A: 0.1% formic acid aqueous solution; B: acetonitrile solution; Column temperature: 20 °C; Flow rate: 1 mL / min; Detection wavelength: 278 nm; Injection volume: 10 μL; Gradient program: 0 min (92% A + 8% B), 5 min (92% A + 8% B), 25 min (89% A + 11% B). Twenty-seven tea samples were randomly selected (three from each production area and season) for EGCG purity verification. Table 2 shows the purity of EGCG extracted from the samples. By comparing the peak times of the extracted EGCG with those of the standard and analyzing the purity of EGCG in the samples, the results showed that the purity of the extracted EGCG was higher than 95%, reaching a maximum of 98.2%, laying the foundation for monomeric stable isotope analysis.
[0063] Table 2 Purity of EGCG extracted from tea leaves 3. Differences in stable isotopes in green tea from different producing areas The average value and standard deviation of stable isotopes in tea from different origins are shown in Table 3. Table 3. Stable isotope ratios and organic component content of green tea from different origins 4. Correlation analysis Pearson correlation analysis was used to study the correlation between tea leaves and stable isotopes of EGCG. The results are shown in Table 4. 茶叶 With δ 13 C EGCG It shows a highly significant positive correlation ( P <0.01), and δ 18 O EGCG Significant positive correlation ( P <0.05). Furthermore, δ 2 H 茶叶 δ 18 O 茶叶 δ 2 H EGCG With δ 18 O EGCG Several isotopes show highly significant correlations with each other ( P<0.01 indicates that the variation trends of hydrogen and oxygen isotopes are highly consistent across different production sites. Figure 1 Furthermore, the sources of hydrogen and oxygen in EGCG are highly consistent with those in tea. Figure 1 The water mainly comes from local water sources.
[0064] Table 4. Correlation analysis of stable isotopes of tea leaves and EGCG 5. Discriminant Analysis To verify the discriminative contribution of EGCG stable isotopes, discriminant analysis was conducted on tea samples from three major tea-producing areas: Ya'an, Yibin, and Dujiangyan. Discriminant models were established using both overall stable isotopes and EGCG monomeric stable isotopes. Discriminant Model 1 was established using overall stable isotopes. To assess the robustness of the model, cross-validation was used to classify and predict the tea samples from different producing areas. The results are shown in Table 5. The initial discrimination and cross-validation discriminant rates for the three producing areas were 89.5% and 83.9%, respectively. Simultaneously, Discriminant Model 2 was established using EGCG monomeric stable isotopes. Cross-validation was used to classify and predict the tea samples from different producing areas. The results are shown in Table 6. The initial discrimination and cross-validation discriminant rates for the three producing areas were 99.4% and 98.9%, respectively. In summary, combining EGCG monomeric stable isotopes significantly increases the discrimination accuracy of green tea samples from different producing areas.
[0065] Table 5. Results of linear discriminant analysis of green tea from different origins (overall stable isotopes) Based on the results of linear discriminant analysis, we obtained the first discriminant model for different production areas, and the results are as follows: Y 雅安 = -106.787 δ 13 C 茶叶 + 33.096 δ 15 N 茶叶 - 10.312 δ 2 H 茶叶 + 40.731 δ 18 O 茶叶 -2103.824 Y 宜宾 = -105.333 δ 13 C 茶叶 + 33.248 δ 15 N 茶叶 - 10.300 δ 2 H 茶叶 + 40.616 δ 18 O 茶叶-2062.343 Y 都江堰 = -105.333 δ 13 C 茶叶 + 30.513 δ 15 N 茶叶 - 9.967 δ 2 H 茶叶 + 39.336 δ 18 O 茶叶 -1976.795; Table 6. Results of linear discriminant analysis of green tea from different origins (overall + EGCG stable isotopes) Based on the results of linear discriminant analysis, a second discriminant model for different production areas was obtained, and the results are as follows: Y 雅安 = -198.618 δ 13 C 茶叶 - 57.787 δ 15 N 茶叶 - 44.365 δ 2 H 茶叶 + 205.054 δ 18 O 茶叶 -650.957 δ 13 C EGCG - 16.113 δ 2 H EGCG + 76.695δ 18 O EGCG -16557.977 Y 宜宾 = -192.666 δ 13 C 茶叶 - 53.730 δ 15 N 茶叶 - 42.863 δ 2 H 茶叶 + 197.751 δ 18 O 茶叶 -622.965 δ 13 C EGCG - 15.304 δ 2 H EGCG + 72.879δ 18 O EGCG -15266.761 Y 都江堰 = -193.627 δ 13 C 茶叶 - 59.346 δ 15 N茶叶 - 43.554 δ 2 H 茶叶 + 201.530 δ 18 O 茶叶 -643.104 δ 13 C EGCG - 15.809 δ 2 H EGCG + 75.065δ 18 O EGCG -16040.476; 6. Verification of constants k1 and k2 To verify the applicability of the universal constants k1 = 5.37 and k2 = 6.01, the δ values of 180 samples (from three production areas and three seasons) collected in this embodiment were used. 2 H 茶叶 With δ 18 O 茶叶 δ 2 H EGCG With δ 18 O EGCG Substituting the data into the aforementioned constants for linear fitting, the R-squared value of the resulting regression equation is... 2 The values were 0.62 and 0.86 respectively (P<0.01), indicating that the sample data of this embodiment are in high agreement with the predetermined constant, verifying the applicability of the constant in the production areas and seasons covered by this application.
[0066] 7. Application of the discriminant model In the spring of the following year, nine fresh tea samples were collected from three production areas. The overall stable isotope and EGCG monomeric stable isotope composition of the samples were analyzed. Substituted into the discriminant model 2 established above, Y was calculated. 雅安 Y 宜宾 and Y 都江堰 The magnitude of the value is shown in Table 7. Table 7 shows Y i Value Result Based on the calculation results, the Y values of the three Ya'an samples... 雅安 The values are all greater than Y 宜宾 and Y 都江堰 Values of Y from 3 Yibin samples 宜宾 The values are all greater than Y 雅安 and Y 都江堰 Values of Y from three samples in Dujiangyan 都江堰 The values are all greater than Y 雅安 and Y 宜宾 The value indicates that all samples can be accurately identified.
[0067] Although embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the specification and embodiments. They can be applied to various fields suitable for the present invention. For those skilled in the art, other modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and embodiments shown and described herein.
Claims
1. A method for determining the origin of tea based on the stable isotope ratio of EGCG monomers, characterized in that, Includes the following steps: Determination of overall stable isotope ratio δ 13 C 茶叶 , δ 15 N 茶叶 , δ 2 H 茶叶 , δ 18 O 茶叶 and δ of EGCG monomer stable isotope ratio in tea samples 13 C EGCG 、 δ 2 H EGCG and δ 18 O EGCG ; wherein the preparation method of the EGCG monomer is: preparing tea samples into tea powder, extracting and purifying to obtain epigallocatechin gallate (EGCG) monomer from the tea powder, and then determining the stable isotope ratio of the EGCG monomer. Based on the above stable isotope ratios, the origin of tea samples can be identified. The method for identifying the origin of tea is as follows: The overall stable isotope ratio δ0.05 is calculated. 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 and the δ¹² value of stable isotope ratios of EGCG monomers 13 C EGCG δ 2 H EGCG and δ 18 O EGCG The input is a pre-built linear discriminant model for different production areas. The pre-built linear discriminant model for different production areas is as follows: Y i =a i d 13 C 茶叶 +b i d 15 N 茶叶 +c i d 2 H 茶叶 +d i d 18 The 茶叶 +e i δ¹³C EGCG +f i d 2 H EGCG +g i d 18 The EGCG +h i ; Among them, a i b i c i d i e i f i g i h i Let be a constant obtained by training from tea samples from known origins using Fisher linear discriminant analysis, where i represents the origin number; During discrimination, the seven isotope ratios of the sample to be tested are substituted into the discrimination model for each origin to obtain Y. i Value, based on Y i The value is used to determine the place of origin; The δ of the tea sample to be tested 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 δ¹³C EGCG δ 2 H EGCG δ 18 O EGCG Substituting the seven stable isotope ratios into the linear discriminant function for each origin yields the discriminant function value Y for each origin. i Calculate the difference ΔY between the largest and second-largest discriminant function values; compare the difference ΔY with the discrimination threshold T. If ΔY is greater than T, the origin corresponding to the largest discriminant function value is determined to be the origin of the tea sample to be tested; if ΔY is less than or equal to T, proceed to the second-level discrimination step; wherein, the discrimination threshold T is obtained by leave-one-out cross-validation on tea samples from known origins, and T is taken as the 10th percentile of the ΔY value of all tea samples from known origins when correctly discriminated; The second-level discrimination steps include: S1. Calculate the three difference characteristics: △ 13 C、△ 2 H and △ 18 O, where: △ 13 C=δ 13 C EGCG -d 13 C 茶叶 ; △ 2 H=d 2 H EGCG -d 2 H 茶叶 ; △ 18 O=d 18 The EGCG -d 18 The 茶叶 ; S2. Calculate the three difference characteristics and the overall stable isotope ratio δ. 15 N tea leaves have four features. The input is a pre-built second-level random forest classifier, consisting of 200 decision trees. The output is the voting ratio for each production area. If two or more production areas have the same voting ratio and all have the highest voting ratio, then the linear discriminant function value Y calculated for these production areas is obtained. i Choose Y i The region of origin with the highest value is used as the final judgment result; The second-level random forest classifier is trained using only tea samples from known origins that satisfy the difference ΔY≤T.
2. The method for determining the origin of tea based on the stable isotope ratio of EGCG monomers as described in claim 1, characterized in that, In step S2, if Y from different places of origin i If the values are the same, perform the following steps: A1. Calculate the difference in deuterium excess parameter Δd of the tea sample to be tested. Δd is calculated according to the following formula: △d=(δ 2 H EGCG -k2δ 18 The EGCG )-(d 2 H 茶叶 -k1δ 18 The 茶叶 ); Wherein, coefficient k1 is the δ value obtained from tea samples from known origins through linear regression analysis. 18 O 茶叶 and δ 2 H 茶叶 The slope obtained from the fitted data, and the coefficient k2, are the δ values obtained from all tea samples of known origins through linear regression analysis. 2 H EGCG and δ 18 O EGCG The slope obtained by fitting the data; A2. For each candidate origin, a subset of boundary samples that satisfy the discriminant function value ΔY≤T is pre-selected from the training samples of that candidate origin. The median M of the Δd values of all samples in this subset is then calculated. j If the number of samples in the boundary sample subset of a candidate origin is less than 3, then the candidate origin is removed from the candidate list. A3. Compare the Δd of the tea sample to be tested with the median M of each candidate origin. j The absolute difference between them |△d-M j | The candidate origin with the smallest absolute difference is selected as the final discrimination result; If two or more candidate production areas have the same absolute difference and both are the minimum, or if all candidate production areas are removed due to insufficient sample size, the tea sample to be tested will be marked as pending verification.
3. The method for determining the origin of tea based on the stable isotope ratio of EGCG monomers as described in claim 2, characterized in that, The discrimination threshold T is determined through the following steps: B1. Perform leave-one-out cross-validation on all tea samples from known origins: Each time, remove one sample as a validation sample, use the remaining samples to train the linear discriminant function, and calculate the δ of the validation sample. 13 C 茶叶 δ 15 N 茶叶 δ 2 H 茶叶 δ 18 O 茶叶 δ 13 C EGCG δ 2 H EGCG δ 18 O EGCG The maximum discriminant function value Y obtained by substituting the seven stable isotope ratios into the discriminant function for each origin is... max And the second largest discriminant function value Y second The difference between them is △Y'=Y max -Y second ; B2. Record the true origin of the verification sample and determine Y. max Does the corresponding place of origin match the actual place of origin? If they match, mark the verification sample as a single-level correctly judged sample and record the ΔY' value of the sample. If they do not match, mark the verification sample as a single-level incorrectly judged sample and do not record its ΔY' value. B3. Iterate through all tea samples from known origins, collect the ΔY' values of all correctly judged samples at the single level, calculate the 10th percentile of these ΔY' values, and use this percentile as the discrimination threshold T; Among them, single-level correct discrimination means that the origin corresponding to the maximum discriminant function value in leave-one-out cross-validation is consistent with the true origin.
4. The method for determining the origin of tea based on the stable isotope ratio of EGCG monomers as described in claim 1, characterized in that, In step one, the method for extracting EGCG monomers is as follows: The crude extract was obtained by heating and ultrasonic-assisted extraction of tea powder using an ethanol-water solution. The crude extract was extracted sequentially with dichloromethane and ethyl acetate, and the supernatant was collected from each extraction to obtain the extract. After purification, the EGCG monomer was obtained.
Citation Information
Patent Citations
Method for estimating production location
CN111868507A
Method for analyzing oxygen stable isotope ratio of water in highly alcoholic beverage
CN121830875A