White tea quality identification method
By combining Raman spectroscopy with machine learning models, and utilizing AgNCs SERS substrates and GBR models, the time-consuming and expensive equipment problems of white tea quality identification have been solved. This has enabled accurate detection and quality grading of EGCG content in white tea, providing an efficient and scientific quality identification method.
Patent Information
- Application Number
- CN202511481183.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2026-02-03
AI Technical Summary
Existing methods for identifying the quality of white tea are time-consuming, require expensive equipment, and rely on experience, lacking efficient and sensitive rapid testing methods.
By combining Raman spectroscopy with machine learning models, the quality of white tea can be identified through EGCG content detection. Feature extraction and regression quantitative analysis are performed using Raman substrate enhancer AgNCs SERS substrate and machine learning models such as gradient boosting regression (GBR).
This method enables precise detection and quality grading of EGCG content in white tea, providing an objective and scientific quality assessment method that improves the speed and accuracy of detection.
Smart Images

Figure CN121453740A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of tea identification, in particular to a white tea quality identification method. BACKGROUND
[0002] Fuding white tea is a kind of white tea product made of fresh leaves of local characteristic tea tree varieties such as Fuding Dahuo tea and Fuding Dabaicha, etc. by unique process of wilting and drying without frying and rubbing. According to the differences in raw material tenderness and process, Fuding white tea is mainly divided into categories such as baihao yinzhen, baimudan, gongmei and shoumei, and its quality identification becomes the core link to guarantee the authenticity and grade of the product.
[0003] At present, the mainstream identification methods can be divided into sensory evaluation and physicochemical index detection; among them, sensory evaluation relies on experienced tea tasters, which has great limitations; physicochemical index detection methods include high performance liquid chromatography (HPLC), gas chromatography-mass spectrometry (GC-MS), inductively coupled plasma mass spectrometry (ICP-MS), colorimetry and fluorescence sensor, Fourier transform infrared reflection (FTIR) spectrum and nuclear magnetic resonance (NMR) spectrum, which are reliable, but have the problems of time-consuming, long analysis time and complex sample pretreatment, and the machines required are expensive, large and heavy. Therefore, it is particularly important to develop an efficient and sensitive rapid detection method. SUMMARY
[0004] In view of the problems of the prior art, the present application provides a white tea quality identification method, which can accurately detect the EGCG content in white tea and realize accurate grading of white tea quality based on the EGCG content.
[0005] A white tea quality identification method, comprising the following steps: (1) providing EGCG standard solution of different concentrations, respectively dropping on the Raman substrate, then performing laser irradiation and collecting Raman spectrum data, introducing the Raman spectrum data into different machine learning models for feature extraction and regression quantitative analysis, and screening to obtain the best analysis model of EGCG; (2) providing white tea samples of different quality grades, preparing tea water samples from each white tea sample, and respectively dropping on the Raman substrate, then performing laser irradiation and collecting Raman spectrum data, introducing the Raman spectrum data into the best analysis model, detecting the EGCG content in each white tea sample, and constructing a retrieval database of different quality grades and EGCG content; (3) preparing a white tea tea water sample to be tested, dropping the white tea tea water sample to be tested on the Raman substrate, then performing laser irradiation and collecting Raman spectrum data, introducing the Raman spectrum data into the best analysis model for prediction, and matching in the retrieval database, and outputting the EGCG content in the white tea to be tested and / or the quality grade of the white tea to be tested.
[0006] The following also provides several optional modes, but not as an additional limitation of the above general scheme, just a further supplement or preferred, without technical or logical contradiction, each optional mode can be combined alone for the above general scheme, but also can be combined between multiple optional modes.
[0007] Optionally, the solvent of the EGCG standard solution is water, and the concentration range is 1x10 1 ~1x10 4 μM.
[0008] Optionally, the optimal analysis model is a gradient boosting regression model.
[0009] Optionally, the parameters of the gradient boosting regression model are: learning rate: 0.1; number of weak learners: 100; sample ratio per tree: 1; maximum depth per tree: 3; minimum number of samples required for splitting internal nodes: 2; minimum number of samples required for leaf nodes: 1.
[0010] Optionally, the white tea quality identification method comprises: providing white tea samples of different years, preparing tea water samples from each white tea sample, and adding each tea water sample dropwise on a Raman substrate, then performing laser irradiation and collecting Raman spectrum data, importing the Raman spectrum data into different machine learning models, performing feature extraction and classification training, and screening to obtain an optimal classification model; importing the Raman spectrum data of the white tea tea water sample to be tested into the optimal classification model, and outputting the year identification result.
[0011] Optionally, the optimal classification model is a random forest model.
[0012] Optionally, the parameters of the random forest model are: number of trees: 200; minimum number of samples required for controlling internal node splitting: 5; minimum number of samples required for controlling leaf nodes: 1; random seed: 42; maximum number of features considered during partitioning: square root of the total number of features.
[0013] Optionally, the preparation method of the tea water sample comprises: mixing white tea sample powder and water in a mass ratio of 0.05~0.15:5, and incubating at a temperature of 75~85℃ for 5~15min; After incubation, centrifugation, filtration, and taking the supernatant are performed to obtain the tea water sample.
[0014] Preferably, the mass ratio of the white tea sample powder to water is 0.1:5.
[0015] Preferably, the incubation temperature is 80℃ and the time is 10 min.
[0016] Optionally, the Raman substrate is an AgNCs SERS enhancement substrate.
[0017] Optionally, the preparation method of the AgNCs SERS enhancement substrate comprises: The silver nitrate (AgNO3) solution and the fish sperm DNA (FSDNA) solution are mixed in a certain proportion, and then the mixed solution is dropped on a copper sheet and left to stand for 3-5 min; after cleaning and drying, the AgNCs SERS enhancement substrate is prepared.
[0018] Optionally, the concentration of the silver nitrate solution is 0.5-1.5 μg / μL, preferably 1 μg / μL.
[0019] Optionally, the concentration of the fish sperm DNA solution is 50-100 mM, preferably 80 mM.
[0020] Optionally, the volume ratio of the silver nitrate solution to the fish sperm DNA solution is 1:1.
[0021] Preferably, the mixed solution is dropped on the copper sheet and left to stand for 5 min.
[0022] Compared with the prior art, the present application combines Raman spectroscopy technology and a machine learning model to achieve simple and rapid precise detection of the key component EGCG in white tea, and based on the significant positive correlation between the EGCG content and the tea quality, the present application takes EGCG as an objective index to establish an objective and scientific white tea quality identification method. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 is a field emission scanning electron microscope (SEM) image of the AgNCs SERS substrate at 5 min; Figure 2A is an energy dispersive X-ray spectroscopy (EDX) element analysis point scanning image of the AgNCs SERS substrate; Figure 2B is an energy dispersive X-ray spectroscopy (EDX) element analysis area scanning image of the AgNCs SERS substrate; Figure 3 is a Raman spectrum test image of the AgNCs SERS substrate at different oxidation times (0-2 h); Figure 4 is a machine learning flowchart; Figure 5 is a flowchart of the identification of white tea quality by combining SERS technology and machine learning; Figure 6Radar chart of evaluation indexes of each classification model of machine learning, including cross-validation accuracy (Cv-score), test set accuracy (Accuracy), precision (Precision), recall (Recall), and F1 score (F1-score); Figure 7A Confusion matrix chart of machine learning RF model in the implementation process of the white tea year identification method of Example 2; Figure 7B ROC curve chart and AUC value of machine learning RF model in the implementation process of the white tea year identification method of Example 2; Figure 8A Line chart of performance indicators of each regression model of machine learning, including mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE); Figure 8B Column chart of performance indicators of each regression model of machine learning, including coefficient of determination (R 2 ); Figure 9 Chart of top 12 proportions of feature weights of the optimal GBR model in the machine learning regression model; Figure 10A Column chart of EGCG content of each subtype of white tea in 2020 detected by GBR model; Figure 10B Column chart of EGCG content of each subtype of white tea in 2021 detected by GBR model; Figure 10C Column chart of EGCG content of each subtype of white tea in 2022 detected by GBR model; Figure 10D Column chart of EGCG content of each subtype of white tea in 2023 detected by GBR model; Figure 10E Column chart of EGCG content of each subtype of white tea in 2024 detected by GBR model; Figure 11A Bland-Altman chart of HPLC and machine learning (ML) methods; Figure 11B Comparison chart of actual value of HPLC and predicted value of ML; Figure 11C Error column chart of actual value of HPLC and predicted value of ML. DETAILED DESCRIPTION
[0024] The technical solutions described in the present application will be further described below in conjunction with specific embodiments, but the present application is not limited thereto.
[0025] Preparation of tea water samples White tea from five years (2020-2024) was selected, with each year including four different grades (Silver Needle, Peony, Gongmei and Shoumei). Tea samples from three brands (Dingye, Pinpinxiang and Lanxichun) were selected for each grade, for a total of 60 tea samples.
[0026] A small, high-efficiency, multi-functional pulverizer with a power of 1300W, a speed of 25000r / min, and a fineness of 30~300 mesh was selected to pulverize each tea sample into powder. The pulverizing conditions were: 30 seconds each time, 10 seconds interval, repeated 5 times. Each tea powder was then placed in an aluminum foil sealed bag and stored in a 4℃ refrigerator for later use.
[0027] Take 0.1 g of tea powder into a 10 mL glass bottle, add 5 mL of ultrapure water, heat in an 80℃ water bath for 10 min, then centrifuge at 4000 rpm for 10 min in a high-speed refrigerated centrifuge, collect the supernatant, and obtain 60 tea water samples.
[0028] Subsequently, white tea samples from different brands of the same year and grade were mixed in a 1:1 ratio of tea to water volume to create a standardized tea sample library: a total of 5 years × 4 grades × 6 samples (as shown in Table 1, including 3 single-brand samples and 3 mixed-brand samples), totaling 120 tea samples.
[0029] Table 1. Tea samples from different brands
[0030] Preparation Example 2: Preparation of EGCG Standard Solution Epigallocatechin gallate (EGCG) standard was dissolved in ultrapure water to prepare EGCG solutions of different concentrations: 1×10 1 μM, 5×10 1 μM, 1×10 2 μM, 5×10 2 μM, 1×10 3 μM, 5×10 3 μM, 1×10 4 μM. After preparation, the solution was sealed and stored in a refrigerator at 4°C, protected from light, for later use.
[0031] Preparation Example 3 3.1 Preparation of AgNCs SERS-enhanced substrate The copper sheet was pre-cleaned with ethanol and dried before use. A 1 μg / μL silver nitrate (AgNO3) solution and an 80 mM fish sperm DNA (FSDNA) solution were prepared using 25 mM Tris-HCl buffer at pH 8.
[0032] Take 5 μL of AgNO3 solution and mix with 5 μL of 80 mM protamine DNA solution, immediately add the mixture evenly to the surface of the copper sheet, and let it react for 5 minutes. Then, rinse the substrate with ultrapure water and dry it with nitrogen or air to remove unbound reactants and impurities, thus preparing the AgNCs SERS enhancement substrate.
[0033] 3.2 Characterization Detailed observation was carried out using field emission scanning electron microscopy (SEM). The SEM images ( Figure 1 ) show that the surface of the prepared substrate is very uniform, indicating that the distribution of nanostructures is consistent and good.
[0034] In addition, point scanning ( Figure 2A ), area scanning ( Figure 2B ) analysis by energy dispersive X-ray spectroscopy (EDX) revealed the presence of elements such as carbon (C), nitrogen (N), oxygen (O), copper (Cu), and silver (Ag) on the AgNCs SERS substrate. These results indicate that FSDNA was indeed involved in the growth of silver nanocoral, further verifying the reliability of the growth mechanism.
[0035] After adding a silver needle tea sample to the AgNCs SERS substrate, the oxidation time was systematically tested ( Figure 3 ). The results show that within 1 hour of reaction time, the characteristic peak intensity of the sample remains almost consistent with the initial peak, indicating that the oxidation process is relatively stable within this time period. However, after 2 hours of reaction, a weakening of the peak intensity was observed.
[0036] Therefore, in order to ensure the accuracy and stability of the experimental results, it was decided to limit the reaction time of subsequent experiments to within 1 hour. This choice not only helps to maintain the activity of the sample, but also provides a more reliable basis for subsequent data analysis.
[0037] Example 1 1.1 Data acquisition 120 tea samples were each added to the AgNCs SERS enhancement substrate, with each tea sample being added in an amount of 10 μL, and then left to stand for 10 minutes to ensure that the tea sample and the AgNCs SERS substrate interacted fully; after completion, they were rinsed with ultrapure water and dried. A portable Raman spectrometer was used for measurement, with an excitation wavelength of 785 nm, an excitation power of 100 mW, an integration time of 3 seconds, and a spectral measurement range covering 82 cm -1 to 3200 cm -1, to obtain high quality SERS spectral data. 10 random detections were performed in each category of samples, and a total of 1200 data points were obtained, i.e. dataset 1. All data were processed by baseline correction and labeled.
[0038] 10 μL EGCG standard solution at 7 different concentrations were respectively dropped on the AgNCs SERS substrate, and after standing for 10 minutes, they were rinsed with ultrapure water and blown dry. The portable Raman spectrometer was also used for data acquisition, and the measurement conditions were consistent with before. At each concentration, 60 random samples were collected, a total of 420 data points were obtained, i.e. dataset 2, and these data were processed by baseline correction and labeled.
[0039] 1.2 Machine learning process Figure 4 The various stages of the machine learning project are clearly shown. First, the full spectral signal intensity of the Raman data after baseline correction is extracted as a feature value, and the dataset is input into the pre-written program for preprocessing, i.e. transposing and saving the data.
[0040] Next, the preprocessed data samples are randomly divided into 80% training data set and 20% test data set, and multiple models are used for feature extraction. According to the test results of the model, the optimal model is selected, or the optimal model is further optimized, and the program automatically establishes a database from the data set. Finally, the optimal model is used to retrieve the database to detect tea samples, thereby classifying their years or dividing their grades.
[0041] Example 2 Machine learning assisted SERS platform for white tea year identification The preprocessed dataset 1 was trained for classification model, and 7 different classification algorithms were used, including logistic regression (LR), support vector machine (SVM), decision tree (DT), random forest (RF), K-nearest neighbors (KNN), Gaussian naive Bayes (NB), and multilayer perceptron classifier (MLPC). In the feature extraction and recognition analysis process, a variety of evaluation indicators were used, including cross-validation accuracy (Cv-score), test set accuracy (Accuracy), precision (Precision), recall (Recall), and F1 score (F1-score). The results showed that the performance of the RF model was the best ( Figure 6 ), with an accuracy of up to 94% (Table 2).
[0042] Table 2 Qualitative results of each classification model
[0043] As previously mentioned, Dataset 1 was randomly divided into 80% training set and 20% test set, and the confusion matrix and ROC plot of the RF model on the training set are shown in FIGS. 1 and 2, respectively. Figure 7B The results show that the classification task of Fuding White Tea achieved 100% recognition rate during 2020-2024, fully demonstrating the efficiency and accuracy of the RF model in this field, see FIGS. 3 and 4. Figure 5
[0044] Although the confusion matrix and ROC curve results of the test set show that the RF model has the highest accuracy among the seven classification models, it is still desirable to further improve the prediction accuracy of the RF model. To this end, the hyperparameters of the RF model were systematically optimized. Through the method of grid search, the number of model trees, the maximum depth, the minimum number of samples required for internal node redivision, the minimum number of samples in leaf nodes, the maximum number of features for decision tree splitting, and other parameter settings were adjusted, and the optimal parameter combination was selected through cross-validation. The optimal parameters are: number of trees: 200; tree depth: no limit; minimum number of samples required for internal node splitting: 5; minimum number of samples required in leaf nodes: 1; random seed: 42; maximum number of features considered when splitting: square root of the total number of features; whether to use bootstrap method to create a subset of the dataset: do not use.
[0045] After determining the optimal parameters, Dataset 1 was reimported into the RF model with optimized parameter settings, and the performance of the model was then evaluated. Through the results of various evaluation indicators and the confusion matrix of the test dataset ( Figure 7A ), it was observed that the accuracy of the RF model in the classification and identification of Fuding White Tea from 2020 to 2024 was significantly improved, reaching 96%. In addition, the analysis of the ROC curve ( Figure 7B ) showed that the area under the curve (AUC) value of the model was infinitely close to 1, which further proved the high recognition performance of the RF model in the classification task.
[0046] This high recognition rate not only reflects the effectiveness of the model on the dataset, but also indicates that the RF model optimized by parameters has excellent generalization ability when dealing with complex data. In this way, the overall performance of the model is successfully improved, providing more reliable prediction results for subsequent practical applications.
[0047] Example 3 Machine learning assisted SERS platform for quantitative detection of EGCG in white tea The pretreated dataset 2 was imported into the constructed 7 regression models, namely, logistic regression (LR), gradient boosting regression (GBR), decision tree regression (DTR), random forest regression (RFR), elastic net regression (EN), extreme gradient boosting regression (XGBR), and light gradient boosting regression (LGBM), to perform feature extraction and regression quantitative analysis on the EGCG samples with different concentrations in dataset 2. First, these regression models were used to predict the test set, and the relevant evaluation indicators were output. Generally speaking, the closer the values of mean squared error (MSE), root mean squared error (RMSE), and mean absolute error (MAE) to 0, the smaller the prediction error of the model, and the closer the coefficient of determination (R 2 ) to 1, the better the fitting degree of the model to the data. From Figure 8A -B, it can be seen that the values of MSE, RMSE, and MAE of the GBR model are more superior than those of other models, tending to 0, and its R 2 value is also close to 1, reaching 98% (Table 3). This indicates that the GBR model has extremely high accuracy and reliability in predicting the concentration of EGCG.
[0048] In addition, the present application also establishes an EGCG concentration database, providing important data support and reference for subsequent analysis and research. Through this series of work, not only the performance differences of different regression models in quantitative analysis are demonstrated, but also a foundation is laid for the accurate determination of EGCG concentration in white tea.
[0049] Table 3 Quantitative results of different regression models
[0050] According to the above results, the GBR model with the best performance was selected for subsequent analysis. Next, the performance of the GBR model in feature peak extraction was verified, and dataset 2 was imported into the feature weight prediction program of the constructed GBR model. The output data showed that the weights of feature peaks 1236-1238 Figure 9 accounted for the majority in the prediction process of the GBR model, which is consistent with the characteristics of the characteristic peaks of EGCG, further verifying the effectiveness and reliability of the model.
[0051] To further improve the prediction accuracy of the GBR model, its parameters were optimized, and the optimal parameters were as follows: learning rate: 0.1; number of weak learners: 100; sample ratio per tree: 1; maximum depth of each tree: 3; minimum number of samples required for splitting internal nodes: 2; minimum number of samples required for leaf nodes: 1.
[0052] Referring to Figure 5The samples from each year in Dataset 1 were imported into the GBR model for prediction, and by retrieving the established EGCG concentration database, it was found that the EGCG content in white tea varied from different years, and the EGCG content in tea of the same year was significantly positively correlated with its quality, see Figure 10A -E. Specifically, the EGCG content of Yinjian was higher than that of Mudan, Mudan was higher than Gongmei, and Gongmei was higher than Shoumei, i.e., Yinjian > Mudan > Gongmei > Shoumei. This quantitative identification method based on EGCG content is highly consistent with the traditional experience classification (tea quality: Yinjian > Mudan > Gongmei > Shoumei), which not only verifies the reliability of the method, but also provides an objective and quantifiable scientific basis for traditional experience classification from the chemical substance level.
[0053] To further verify the accuracy of the experimental results, the EGCG content of the 6 tea samples was experimentally verified using the national standard method - high performance liquid chromatography (HPLC). Through statistical analysis, the results of HPLC and machine learning (ML) were compared using Bland-Altman plots ( Figure 11A ), and it was observed that all data points were located between the two green dashed lines. This indicates that the measurement difference between the two methods is within an acceptable range, further confirming that their measurement results have good correlation.
[0054] In addition, by comparing the actual measurement values of HPLC with the predicted values of the ML model ( Figure 11B ), combined with the analysis of the error bar chart ( Figure 11C ), it was found that the error between the detection results of the two methods was small, further proving the consistency and reliability of the prediction results. These results not only enhance the confidence in the credibility of the GBR model in EGCG concentration prediction, but also provide a solid data foundation and experimental support for future related research.
[0055] The above-described embodiments only express several embodiments of the present application, and the description is more specific and detailed, but it should not be understood as limiting the scope of the patent of the present application. It should be noted that for ordinary skilled persons in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are within the scope of protection of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A method of identifying white tea quality, characterized by, The method comprises the following steps: (1) providing EGCG standard solution with different concentrations, respectively dropping on the Raman substrate, then performing laser irradiation and collecting Raman spectrum data, introducing the Raman spectrum data into different machine learning models, performing feature extraction and regression quantitative analysis, and screening the best analysis model of EGCG; (2) providing white tea samples with different quality grades, preparing tea water samples from each white tea sample, and respectively dropping on the Raman substrate, then performing laser irradiation and collecting Raman spectrum data, introducing the Raman spectrum data into the best analysis model, detecting the EGCG content in each white tea sample, and constructing a retrieval database of different quality grades and EGCG content; (3) preparing a to-be-tested white tea tea water sample, dropping the to-be-tested white tea tea water sample on the Raman substrate, then performing laser irradiation and collecting Raman spectrum data, introducing the Raman spectrum data into the best analysis model for prediction, and matching in the retrieval database, outputting the EGCG content in the to-be-tested white tea and / or the quality grade of the to-be-tested white tea.
2. The white tea quality identification method according to claim 1, characterized in that, The solvent of the EGCG standard solution is water, and the concentration range is 1 x 10 1 ~1 x 10 4 μM.
3. The white tea quality identification method according to claim 1, characterized in that, The best analysis model is a gradient boosting regression model.
4. The method of claim 3, wherein the white tea is selected from the group consisting of Silver Needle, White Peony, and White Darjeeling. The parameters of the gradient boosting regression model are as follows: learning rate: 0.1; number of weak learners: 100; sample ratio per tree: 1; maximum depth of each tree: 3; minimum number of samples required for splitting internal nodes: 2; minimum number of samples required for leaf nodes:
1.
5. The white tea quality identification method according to claim 1, characterized in that, It comprises: providing white tea samples of different years, preparing tea water samples from each white tea sample, and respectively dropping on the Raman substrate, then performing laser irradiation and collecting Raman spectrum data, introducing the Raman spectrum data into different machine learning models, performing feature extraction and classification training, and screening the best classification model; introducing the Raman spectrum data of the to-be-tested white tea tea water sample into the best classification model, and outputting the year identification result.
6. The white tea quality identification method according to claim 5, characterized in that, The best classification model is a random forest model.
7. The white tea quality identification method according to claim 6, characterized in that, The parameters of the random forest model are as follows: number of trees: 200; minimum number of samples required for controlling internal node splitting: 5; minimum number of samples required for controlling leaf nodes: 1; random seed: 42; maximum number of features considered during partitioning: square root of the total number of features.
8. The method of claim 1 or 5, wherein the white tea is a tea produced in the Wuyi Mountains of China. The preparation method of the tea water sample comprises: mixing white tea sample powder and water in a mass ratio of 0.05-0.15:5, and incubating at a temperature of 75-85℃ for 5-15 min; after incubation, centrifugation, filtration, and taking the supernatant, the tea water sample is obtained.
9. The method of claim 1 or 5, wherein the white tea is a tea leaf product. The Raman substrate is an AgNCs SERS enhancement substrate.