Method for determining storage year of Pu'er tea based on packaging cotton paper near infrared spectrum

By collecting near-infrared spectra of Pu'er tea packaging paper in different regions and constructing an EPO-corrected PLS-DA model, the accuracy and stability issues of Pu'er tea storage year identification were solved, achieving rapid and accurate year identification.

CN121027031APending Publication Date: 2025-11-28YUNNAN AGRI VOCATIONAL & TECH COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511078249.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing technologies for identifying the storage age of Pu-erh tea are affected by the complexity of raw material varieties, the diversity of origins, the differences in processing techniques, and the changes in storage conditions, resulting in insufficient accuracy and stability of near-infrared spectroscopy analysis.

Method used

By dividing the cotton paper used to package Pu'er tea into a full contact area and a partial contact area, near-infrared spectra were collected and a variance spectral matrix was constructed. The EPO correction method was used to eliminate background information differences, and a PLS-DA model was established to extract information related to the storage year.

Benefits of technology

The accuracy and stability of identifying the storage age of Pu'er tea have been improved. The accuracy and sensitivity of the model in interactive validation and external testing have been significantly improved, while the error rate has been reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121027031A_ABST
    Figure CN121027031A_ABST
Patent Text Reader

Abstract

The invention discloses a method for determining the storage year of Pu'er tea based on packaging cotton paper near infrared spectroscopy, which comprises the following steps of: dividing Pu'er tea packaging cotton paper into a Pu'er tea full contact area and a Pu'er tea partial contact area according to the contact area of the cotton paper and Pu'er tea in the Pu'er tea storage process; collecting near infrared spectrums of the two regions and calculating a variance spectrum of each sample, constructing a variance spectrum matrix of cotton paper at different storage times as a disturbance matrix of EPO correction, eliminating the difference of background information of Pu'er tea packaging cotton paper from different sources, extracting information related to the storage year of Pu'er tea in near infrared spectrum information of the packaging cotton paper, and calculating the storage year of the Pu'er tea. Establishing PLS-DA models of the Pu'er tea cotton paper in different storage years; the model is used for predicting the storage year of a to-be-tested sample; the method has the advantages of high measurement speed, high efficiency, excellent accuracy and stability, and good popularization and application values.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a method for determining the storage years of Pu'er tea based on near-infrared spectrum of packaging cotton paper, and belongs to the technical field of food detection. BACKGROUND

[0002] Pu'er tea is a kind of tea formed by using Yunnan large-leaf species sun-cured tea as raw material and planting and processing within the geographical indication protection range in China. As a kind of flavor tea, Pu'er tea has the health care functions of regulating human intestinal flora and reducing blood lipids, and is favored by consumers, forming a large-scale consumer market. According to the processing technology and quality characteristics, Pu'er tea can be divided into raw tea and ripe tea. As a kind of post-fermented tea, the types and contents of chemical components (such as tea polyphenols, theaflavins and free amino acids) of Pu'er tea change with the influence of microbial fermentation and time during the storage process, thereby affecting the health care functions and flavor characteristics. The storage years are also one of the important factors for determining the price of Pu'er tea in the market. Therefore, it is necessary to identify the storage years of Pu'er tea. The identification of the storage years of Pu'er tea is carried out by using chromatography-mass spectrometry analysis, spectral analysis (terahertz spectrum, Raman spectrum, infrared spectrum and near-infrared) and taste information analysis (electronic tongue). The above analysis technologies are mainly realized by detecting the differences in the types and contents of chemical components of Pu'er tea with different storage years. However, the raw material variety of Pu'er tea is complex, the sources are various, the planting methods are different, the processing technology parameters are various, and the altitudes, temperatures and humidities and other storage conditions in different storage areas are quite different, and many other “background” factors will affect the accuracy and stability of the above analysis methods. After being pressed into cakes, the cotton paper packaging storage is the most common storage method for compacted Pu'er tea. With the increase of storage years, the color of the packaging cotton paper of Pu'er tea changes. The color change of the paper is caused by the change of chemical components.

[0003] As an accurate, rapid, green and non-destructive product quality qualitative and quantitative analysis method, the near-infrared spectrum analysis technology is used for the detection of physical performance indexes (strength, whiteness and quantity) of paper, chemical components (moisture, pH, potassium permanganate value and alkali reserve) and comprehensive performance indexes (fiber polymerization degree and thermal aging degree). The near-infrared spectrum analysis technology is also applied to the identification of different types and processing technology paper. The accuracy and stability of the near-infrared spectrum analysis technology are determined by the performance of the established qualitative and quantitative models. The accuracy, stability and adaptability of the model are the keys to the popularization and application of the near-infrared spectrum analysis technology. Representative sample selection, sample pretreatment, spectrum pretreatment, characteristic wavelength screening and internal cross-validation are common methods for improving the accuracy, stability and adaptability of the model.

[0004] When a method for identifying the storage years of Pu'er tea is established by using the near-infrared spectrum information of the wrapping cotton paper, the selection of the representative area of the cotton paper during the storage process of the Pu'er tea is the key to the establishment of the analysis technology, and the influence of the background near-infrared spectrum information of the cotton paper from different sources on the accuracy, stability and adaptability of the near-infrared spectrum model is eliminated. SUMMARY

[0005] The application provides a method for determining the storage years of Pu'er tea based on the near-infrared spectrum of wrapping cotton paper, which has good accuracy, stability and adaptability. According to the wrapping characteristics of the Pu'er tea cake, the method divides the wrapping cotton paper of the Pu'er tea into a full contact area of the Pu'er tea and a partial contact area of the Pu'er tea by the area of the cotton paper in contact with the Pu'er tea during the storage process of the Pu'er tea, collects the near-infrared spectrum of the two areas and calculates the variance spectrum of each sample, constructs the variance spectrum matrix of the cotton paper with different storage times as a disturbance matrix of EPO (External Parameter Orthogonalization) correction, eliminates the difference in the background information of the wrapping cotton paper of the Pu'er tea from different sources, extracts the information related to the storage years of the Pu'er tea from the near-infrared spectrum information of the wrapping cotton paper, establishes a PLS-DA (partial least-squares discrimination analysis) model of the cotton paper of the Pu'er tea with different storage years, and provides an effective analysis method for the accurate and rapid identification of the storage years of the Pu'er tea.

[0006] The method for determining the storage years of the Pu'er tea based on the EPO correction of the variance spectrum matrix of the cotton paper in different areas according to the application has the following steps:

[0007] 1. Collect the Pu'er tea samples (raw tea and ripe tea) wrapped by cotton paper with different storage years, at least one sample per month per year, unpack the wrapping cotton paper of the Pu'er tea as samples, and record the storage year information of the samples;

[0008] 2. Divide each wrapping cotton paper sample in step 1 into two areas, namely a full contact area of the Pu'er tea and a partial contact area of the Pu'er tea, and cut n pieces of cotton paper from each area, wherein n is greater than 5;

[0009] The cut cotton paper pieces are circular pieces with a radius of 25 mm-30 mm;

[0010] The full contact area of the Pu'er tea and the partial contact area of the Pu'er tea are as follows: the full contact area of the Pu'er tea is a circular area in the wrapping cotton paper of the cake-shaped Pu'er tea which is not folded and is in full contact with the Pu'er tea, and the partial contact area of the Pu'er tea is a folded paper area;

[0011] 3. Collecting the diffuse reflectance spectral data of the pieces of tissue paper obtained from each of the packaged tissue paper samples in step 2, the n pieces of tissue paper from each region of each of the packaged tissue paper samples being longitudinally overlapped and stacked together, and then being sequentially contacted with the spectral collection probe of the near-infrared spectrometer to perform near-infrared spectral detection, n x 2 pieces of spectral data being obtained for each of the packaged tissue paper samples;

[0012] 4. Preprocessing the spectral data of the packaged tissue paper samples collected in step 3 by using second-order derivation and Norris smoothing, and selecting the spectral data in the wavelength range of 8500 cm -1 ~ 4000 cm -1 for each of the packaged tissue paper samples, n x 2 pieces of spectral data in the wavelength range of 8500 cm -1 ~ 4000 cm -1 being obtained for each of the packaged tissue paper samples;

[0013] Each piece of spectral data contains the absorbance at m wavelengths, and n x 2 pieces of spectral data form an n x 2 row and m column spectral matrix;

[0014] 5. Based on the data obtained in step 4, calculating the standard deviation of the absorbance at each wavelength of the n x 2 pieces of spectral data in the wavelength range of 8500 cm -1 ~ 4000 cm -1 for each of the packaged tissue paper samples, and obtaining the standard deviation spectrum for each of the packaged tissue paper samples, wherein when there are multiple samples in a month, the average of the standard deviation spectra of all the samples is calculated and taken as the standard deviation spectrum of the samples in the month;

[0015] The standard deviation of the absorbance wherein x i is the absorbance of the i-th sample at the m-th wavelength, and μ is the average of the absorbance of the n x 2 pieces of spectral data of the i-th sample at the m-th wavelength; and a 1 row and m column standard deviation spectrum is obtained for each sample;

[0016] According to the length of the storage time of each sample, a standard deviation spectrum matrix is constructed in the order of shorter storage time first and longer storage time last;

[0017] 6. Taking the standard deviation spectrum matrix in step 5 as a disturbance spectrum matrix, and correcting the spectral data of the packaged tissue paper samples obtained in step 4 by using the EPO algorithm to obtain the corrected spectral data and the projection matrix of the disturbance spectrum matrix;

[0018] The steps of correcting the spectral data of the packaged tissue paper samples by using the EPO algorithm are as follows:

[0019] (1) Calculating the difference spectrum matrix D difThe values of each column in the second row to the last row in the perturbation spectrum matrix are subtracted from the values of the corresponding column in the first row respectively, and then the values are placed in the original positions, and the data in the first row remains unchanged, to obtain a difference spectrum matrix D dif

[0020] (2) Calculate the covariance matrix of the difference spectrum matrix D dif Wherein represents the transpose of D dif

[0021] (3) Singular value decomposition is performed on the covariance matrix D dif-cov using the following formula:

[0022] [U, S, V] = svd(D dif-cov ), wherein svd represents singular value decomposition of the matrix, U represents the left singular vector obtained by decomposition, S is the singular value, and V represents the right singular vector;

[0023] (4) When the extraction rate of the singular value decomposition on the original matrix information is greater than 99%, the corresponding principal factor is k, and the submatrix V k corresponding to the first k factors of the right singular matrix is taken;

[0024] (5) The projection matrix Q of the perturbation spectrum matrix is calculated, Wherein represents the transpose of V k

[0025] (6) The spectrum data x of the packaged cotton paper sample obtained in step 4 is corrected using x EPO = xQ to obtain corrected spectrum data x EPO .

[0026] 7. The storage years of the sample in step 1 are taken as the classification label, and the corrected spectrum data of the corresponding sample in step 6 are combined to establish a classification and identification model of the packaged cotton paper of Pu'er tea with different storage years;

[0027] 8. The sample to be tested is operated according to steps 1-4 to obtain n×2 spectrum data in the wavelength range of 8500cm -1 ~4000cm -1 , the average value of the absorbance value of each wavelength of the n×2 spectrum data is calculated to obtain the average near-infrared spectrum data of the sample to be tested, the data is multiplied by the projection matrix of the perturbation spectrum matrix in step 6 for EPO correction, and the corrected result is input into the classification and identification model in step 7 to obtain the storage years of the sample to be tested.

[0028] The EPO correction of the spectrum data x new of the sample to be tested is performed using x new-EPO = x​​​​new Q.

[0029] Advantages and technical effects of the present application:

[0030] The present application is based on the EPO-PLS-DA model for Pu'er tea year identification using the variance near-infrared spectrum of two different regions of the packaging cotton paper as the disturbance spectrum matrix, which has the advantages of high accuracy and stability. In the Pu'er tea (raw tea) model, the correct rate NER, error rate ER, sensitivity TPR, missed diagnosis rate FNR, specificity TNR, false positive rate FPR, precision Pr, classification efficiency EFF and F1 score of the training set samples in the leave-one-out cross-validation are 0.955, 0.045, 0.964, 0.036, 0.945, 0.055, 0.946, 0.955 and 0.959, respectively. The correct rate NER, error rate ER, sensitivity TPR, missed diagnosis rate FNR, specificity TNR, false positive rate FPR, precision Pr, classification efficiency EFF and F1 score of the external test set samples are 0.900, 0.100, 0.925, 0.075, 0.875, 0.125, 0.881, 0.900 and 0.912, respectively. In the EPO-PLS-DA model for Pu'er tea (cooked tea) year identification, the correct rate NER, error rate ER, sensitivity TPR, missed diagnosis rate FNR, specificity TNR, false positive rate FPR, precision Pr, classification efficiency EFF and F1 score of the training set samples in the leave-one-out cross-validation are 0.957, 0.043, 0.962, 0.038, 0.952, 0.048, 0.953, 0.957 and 0.960, respectively. The correct rate NER, error rate ER, sensitivity TPR, missed diagnosis rate FNR, specificity TNR, false positive rate FPR, precision Pr, classification efficiency EFF and F1 score of the external test set samples are 0.900, 0.100, 0.920, 0.080, 0.880, 0.120, 0.885, 0.900 and 0.910, respectively.

[0031] The method of the present application provides a new method for determining the storage year of Pu'er tea, which has the advantages of fast determination speed, high efficiency, high accuracy and stability, and good application value. BRIEF DESCRIPTION OF DRAWINGS

[0032] Figure 1 The sample number of the packaging cotton paper of Pu'er tea (raw tea and cooked tea) with different storage times;

[0033] Figure 2 Schematic diagram for dividing the complete contact area and the partial contact area of the packaging cotton paper of Pu'er tea;

[0034] Figure 3 The original near-infrared spectrum of the packaging cotton paper of Pu'er tea (raw tea);

[0035] Figure 4 Original near infrared spectrum of the cotton paper for Pu'er tea (ripe tea) packaging;

[0036] Figure 5 Pretreated near infrared spectrum of the cotton paper for Pu'er tea (raw tea) packaging;

[0037] Figure 6 Pretreated near infrared spectrum of the cotton paper for Pu'er tea (ripe tea) packaging;

[0038] Figure 7 Standard deviation spectrum data obtained after calculation of the pretreated spectrum data of the cotton paper for Pu'er tea (raw tea) packaging;

[0039] Figure 8 Standard deviation spectrum data obtained after calculation of the pretreated spectrum data of the cotton paper for Pu'er tea (ripe tea) packaging;

[0040] Figure 9 Classification effect of PLS-DA model of the near infrared spectrum of the cotton paper for Pu'er tea (raw tea) and Pu'er tea (ripe tea);

[0041] Figure 10 PLS-DA model of Pu'er tea (raw tea) year before EPO correction;

[0042] Figure 11 PLS-DA model of Pu'er tea (raw tea) year after EPO correction;

[0043] Figure 12 PLS-DA model of Pu'er tea (ripe tea) year before EPO correction;

[0044] Figure 13 PLS-DA model of Pu'er tea (ripe tea) year after EPO correction;

[0045] Figure 14 Result of 200 cross-validations of the PLS-DA model of Pu'er tea (raw tea) year after EPO correction;

[0046] Figure 15 Result of 200 cross-validations of the PLS-DA model of Pu'er tea (ripe tea) year after EPO correction. DETAILED DESCRIPTION

[0047] The application will be described in further detail below with reference to specific embodiments, but the scope of the application is not limited to the described content. The methods in the examples are all conventional methods unless otherwise specified, and the reagents used are all conventional commercially available reagents or reagents prepared according to conventional methods unless otherwise specified;

[0048] Instruments used in the examples: Near infrared spectrometer: Nicolet Antaris II type including integral sphere diffuse reflection analysis module with InGaAs detector, rotating sampling device, 5 cm quartz cup and other sampling accessories and RESULT TM Spectrum collection software (Thermo, USA), near infrared spectrum collection range: 10000cm -1 ~4000cm -1 , spectral resolution: 8cm -1 , the number of spectral scans is 64 times;

[0049] Example 1: Collecting Pu'er tea samples packaged by Yunnan Dayi Tea Group Co., Ltd., Yunnan Zhongcha Tea Co., Ltd., Yunnan Haiguan Tea Co., Ltd. and Yunnan State Farms Group Menghai Bajiaoting Tea Co., Ltd. in 2018 and 2019, wherein the raw tea samples are 250 / year x 2 years = 500, the ripe tea samples are 130 / year x 2 years = 260, and the sample quantity of each month is shown in Figure 1 .

[0050] 1. Pretreatment of Pu'er tea packaging cotton paper

[0051] Collect the above packaged Pu'er tea samples, unpack the packaging cotton paper, lay flat, record the packaging storage year information of the sample, and press the cotton paper flat with a heavy object of the same area for 48 hours; use a steel round mold with a sharp edge (diameter 47mm), and cut the cotton paper under the action of external force to divide the cotton paper into a complete contact area of Pu'er tea and a partial contact area of Pu'er tea (see Figure 2 ), 6 cotton paper pieces are cut from each area of each cotton paper, and 12 cotton paper pieces are obtained from one packaging cotton paper sample;

[0052] 2. Collection of near infrared spectrum of sample

[0053] Collect the diffuse reflectance spectrum data of the cotton paper pieces in step 1, and stack the 6 samples in each area vertically together when collecting the spectrum data, and contact the near infrared spectrometer spectrum collection probe in turn, 2 regions for each sample, a total of 12 spectrum data are collected, and the results are shown in Figure 3 、 Figure 4 From Figure 3 and Figure 4 , it can be seen that the near infrared spectra of different areas of the cotton paper (complete contact area and partial contact area) have certain differences, indicating the necessity of classifying the cotton paper into two areas; the PLS-DA classification results of the near infrared spectra of Pu'er tea (raw tea) and Pu'er tea (ripe tea) cotton paper are shown in Figure 9 , and Figure 9It can be seen that the near-infrared spectra of Pu'er tea (raw tea) and Pu'er tea (cooked tea) cotton paper are obviously classified in the PLS-DA first and second principal component score space, indicating that it is very necessary to distinguish the year identification model of Pu'er tea into raw tea and cooked tea;

[0054] 3. Spectral pretreatment

[0055] The near-infrared spectra of the samples were pretreated by second derivative and Norris smoothing using TQ Analyst software, and the spectral data in the wavelength range of 8500 cm -1 ~ 4000 cm -1 in the pretreated data were selected for subsequent calculation. Each packaging cotton paper sample obtained 12 spectral data in the wavelength range of 8500 cm -1 ~ 4000 cm -1 , and the results are shown in Figure 5 、 6 ;

[0056] 4. Construction of standard deviation spectral matrix

[0057] (1) Each spectral data of each packaging cotton paper sample contains absorbance at m wavelengths, and 12 spectral data form a spectral matrix of 12 rows and m columns;

[0058] The standard deviation of each column data in the above 12-row m-column spectral matrix was calculated, and a standard deviation spectrum of 1 row and m columns was obtained for each sample;

[0059] The standard deviation of the absorbance at the mth wavelength point where x i is the absorbance of the ith sample at the mth wavelength point, μ is the average of the absorbance of the 12 samples at the mth wavelength point, and i = 1, 2, 3, …;

[0060] (2) The average value of the standard deviation spectrum of the samples in each month of 2018 and 2019 was calculated and taken as the standard deviation spectrum of the samples in that month, and the results are shown in Figure 7 、 8 ;

[0061] (3) The standard deviation spectrum of December 2019 was taken as the first row, followed by November-December 2019, December 2018-January 2018, and sequentially placed in the 2nd to 24th rows of the matrix to form a 24-row m-column standard deviation spectral matrix, which was taken as the perturbation spectral matrix;

[0062] 5. Calibration of near-infrared spectra of Pu'er tea packaging cotton paper in 2018 and 2019 based on the perturbation spectral matrix

[0063] (1) The difference spectrum matrix D difThe values of each column in the second row to the last row in the disturbance spectrum matrix D are subtracted by the values of the corresponding column in the first row respectively, and then put into the original position, and the data in the first row remains unchanged, to form a difference spectrum matrix D with 24 rows and m columns dif ;

[0064] (2) Calculate the covariance matrix of the difference spectrum matrix D dif where denotes the transpose of D dif ;

[0065] (3) Singular value decomposition is performed on the covariance matrix D dif-cov using the following formula:

[0066] [U, S, V] = svd(D dif-cov ), where svd denotes singular value decomposition of the matrix, U denotes the left singular vector obtained by decomposition, S is the singular value, and V denotes the right singular vector;

[0067] (4) When the extraction rate of the singular value decomposition on the original matrix information is greater than 99%, the corresponding principal factor is K, and the submatrix V k of the right singular matrix corresponding to the first K factors is taken;

[0068] (5) Calculate the projection matrix Q of the disturbance spectrum matrix, where denotes the transpose of V k ;

[0069] (6) The spectrum data x of the packaged cotton paper sample obtained in step 3 is corrected using x EPO = xQ to obtain corrected spectrum data x EPO ;

[0070] 6. Establishment and evaluation of PLS-DA model for determining the storage years of Pu'er tea based on near-infrared spectrum of packaged cotton paper

[0071] Among the near-infrared spectra of 500 Pu'er tea (raw tea) and 260 Pu'er tea (cooked tea) samples, 80 Pu'er tea (raw tea) samples (40 samples from 2018 and 40 samples from 2019) and 50 Pu'er tea (cooked tea) samples (25 samples from 2018 and 25 samples from 2019) were randomly selected as the model test set and were not involved in modeling, and the remaining 420 Pu'er tea (raw tea) samples and 210 Pu'er tea (cooked tea) samples were used as the model training set.

[0072] ​Based on the near-infrared spectra (independent variables X) of Pu'er tea (raw tea) and Pu'er tea (cooked tea) packaging cotton paper samples in 2018 and 2019 after EPO correction in step 5 and the corresponding packaging storage years (dependent variables Y), the partial least squares-discriminant analysis (PLS-DA) was used to establish the year identification model of Pu'er tea (raw tea) and Pu'er tea (cooked tea), respectively, as shown in Figure 11 、 Figure 13 , Figure 11 and Figure 13 It can be seen that the samples of Pu'er tea (raw tea) and Pu'er tea (cooked tea) stored in 2018 and 2019 are classified into two categories in the model space, indicating that the model has good accuracy;

[0073] In order to compare the advantages of using the standard deviation spectrum matrix of the two regions of the cotton paper as the EPO perturbation matrix to correct the original sample spectrum matrix, the original near-infrared spectra of Pu'er tea (raw tea) and Pu'er tea (cooked tea) packaging cotton paper in the range of 8500 cm -1 ~ 4000 cm -1 were pretreated by the same second derivative and Norris smoothing, combined with the year label, and the PLS-DA year identification models of Pu'er tea (raw tea) and Pu'er tea (cooked tea) were established, as shown in Figure 10 、 12 The stability and accuracy of the training set and test set of the model were evaluated; and Figure 10 and Figure 11 、 Figure 12 and Figure 13 It can be seen that EPO correction improves the accuracy of model classification.

[0074] The "leave-one-out" cross-validation method was used to evaluate whether the model was overfitting. The accuracy, error rate, sensitivity, misdiagnosis rate, specificity, precision, classification efficiency and F1 score calculated by the confusion matrix of the test set of the model were used to evaluate the model;

[0075] The confusion matrix and model parameters of the training set and test set of the model before and after EPO correction are shown in Tables 1 to 8. As shown in the results of Table 8, EPO correction using the variance spectrum as the perturbation matrix can improve the accuracy and stability of the model. The results of 200 times of interactive validation (as shown in Figure 14 and Figure 15 ) show that the EPO correction model does not appear overfitting.

[0076] The model output result (prediction confusion matrix) of the Puer tea (raw tea) training set "leave-one-out" cross-validation is shown in Table 1, and the model accuracy and stability evaluation parameters calculated from the confusion matrix result are shown in Table 2. The model output result (prediction confusion matrix) of the 80 external prediction set prediction is shown in Table 3, and the model accuracy and stability evaluation parameters calculated from Table 3 are shown in Table 4. As can be seen from Tables 2 and 4, the EPO correction significantly improves the prediction accuracy and stability of the Puer tea (raw tea) identification model for cross-validation and prediction set samples;

[0077] Table 1 Confusion matrix of Puer tea (raw tea) training set cross-validation PLS-DA prediction result

[0078]

[0079] Table 2 Puer tea (raw tea) training set PLS-DA model parameters

[0080]

[0081] Table 3 Confusion matrix of Puer tea (raw tea) prediction set PLS-DA prediction result

[0082]

[0083] Table 4 Puer tea (raw tea) prediction set PLS-DA model parameters

[0084]

[0085]

[0086] The model output result (prediction confusion matrix) of the Puer tea (raw tea) training set "leave-one-out" cross-validation is shown in Table 1, and the model accuracy and stability evaluation parameters calculated from the confusion matrix result are shown in Table 2. The model output result (prediction confusion matrix) of the 80 external prediction set prediction is shown in Table 3, and the model accuracy and stability evaluation parameters calculated from Table 3 are shown in Table 4. As can be seen from Tables 2 and 4, the EPO correction significantly improves the prediction accuracy and stability of the Puer tea (raw tea) identification model for cross-validation and prediction set samples;

[0087] Table 5 Confusion matrix of Puer tea (raw tea) training set cross-validation PLS-DA prediction result

[0088]

[0089] Table 6 Puer tea (raw tea) training set cross-validation PLS-DA model parameters

[0090]

[0091] Table 7 Confusion matrix of PLS-DA prediction results for Pu-erh tea (ripened tea) prediction set

[0092]

[0093] Table 8 PLS-DA model parameters for Pu-erh tea (ripened tea) prediction set

[0094]

Claims

1. A method for determining the storage age of Pu-erh tea based on near-infrared spectroscopy of packaging cotton paper, characterized in that, The steps are as follows: (1) Collect Pu'er tea samples with intact cotton paper packaging from different storage years. There should be at least one sample for each month of each year. Open the cotton paper packaging of the Pu'er tea as a sample and record the packaging and storage year information of the sample. (2) Divide each packaging cotton paper sample in step (1) into two areas: one is the area where the Pu'er tea is in complete contact, and the other is the area where the Pu'er tea is in partial contact. Cut n pieces of cotton paper from each area, where n is greater than 5. (3) Acquisition step (2) Diffuse reflectance spectral data of cotton paper pieces obtained from each packaging cotton paper sample. When acquiring spectral data, n cotton paper pieces of each region of each packaging cotton paper sample are stacked together longitudinally and then contacted with the spectral acquisition probe of the near-infrared spectrometer in turn for near-infrared spectral detection. Each packaging cotton paper sample obtains n×2 spectral data. (4) The spectral data of the packaging cotton paper samples collected in step (3) were preprocessed using second-order differentiation and Norris smoothing. The wavelength range of each preprocessed data was selected to be within 8500 cm⁻¹. -1 ~4000cm -1 Spectral data within the range were obtained for each packaging cotton paper sample, with n×2 wavelengths in the range of 8500 cm⁻¹. -1 ~4000cm -1 Spectral data within the range; (5) Based on the data obtained in step (4), calculate the wavelength range of each packaging paper sample within 8500 cm. -1 ~4000cm -1 The standard deviation of absorbance at each wavelength of n×2 spectral data within the range is used to obtain the standard deviation spectrum of each packaging cotton paper sample. When there are multiple samples in a month, the average of the standard deviation spectra of all samples is calculated and used as the standard deviation spectrum of the samples in that month. Based on the length of storage time of each sample, a standard deviation spectral matrix is ​​constructed with the shorter storage time listed above and the longer storage time listed below. (6) Using the standard deviation spectral matrix of step (5) as the perturbation spectral matrix, the EPO algorithm is used to correct the spectral data of the packaging cotton paper sample obtained in step (4) to obtain the corrected spectral data and the projection matrix of the perturbation spectral matrix. (7) Using the storage year of the sample in step (1) as the classification label, and combining the corrected spectral data of the corresponding sample in step (6), a classification and identification model for Pu'er tea packaging cotton paper is established. (8) The sample to be tested for storage years was obtained by following steps (1) to (4) to obtain a wavelength range of 8500 cm⁻¹. -1 ~4000cm -1 The average near-infrared spectral data of the sample to be tested is obtained by calculating the average absorbance value of each wavelength of the n×2 spectral data within the range. This data is multiplied by the projection matrix of the perturbation spectral matrix in step (6) for EPO correction, and the corrected result is input into the classification and identification model in step (7) to obtain the storage year of the sample to be tested.

2. The method for determining the storage age of Pu-erh tea based on near-infrared spectroscopy of packaging paper according to claim 1, characterized in that: Pu-erh tea includes ripe Pu-erh tea and raw Pu-erh tea.

3. The method for determining the storage age of Pu-erh tea based on near-infrared spectroscopy of packaging paper according to claim 1, characterized in that: The cut cotton paper pieces are circular pieces with a radius of 25mm-30mm.

4. The method for determining the storage age of Pu-erh tea based on near-infrared spectroscopy of packaging paper according to claim 1, characterized in that: Absorbance standard deviation Where x i Let μ be the absorbance of the i-th sample at the m-th wavelength, and μ be the average absorbance of the n×2 spectral data of the i-th sample at the m-th wavelength.

5. The method for determining the storage age of Pu-erh tea based on near-infrared spectroscopy of packaging paper according to claim 1, characterized in that: The steps for correcting the spectral data of packaging cotton paper samples using the EPO algorithm are as follows: A. Calculate the difference spectral matrix D of the perturbation spectral matrix. dif The difference spectral matrix D is obtained by subtracting the corresponding column number from the value in the first row from the value in each column of the second to the last row of the perturbation spectral matrix, and then placing the result back in its original position. The data in the first row remains unchanged. dif ;; B. Calculate the difference spectral matrix D dif covariance matrix in D represents dif Transpose of; C. Use the following formula to evaluate the covariance matrix D. dif-cov Perform singular value decomposition; [U, S, V] = svd(D dif-cov In the formula, svd represents the singular value decomposition of the matrix, U represents the left singular vector obtained by the decomposition, S is the singular value, and V represents the right singular vector. D. When the singular value decomposition extracts more than 99% of the information from the original matrix, the corresponding principal factor is k. Take the submatrix V of the right singular matrix corresponding to the first k factors. k ; E. Calculate the projection matrix Q of the perturbation spectral matrix. in V represents k Transpose of; F. Using x EPO =xQ Corrects the spectral data x of the packaging cotton paper sample obtained in step (4) to obtain the corrected spectral data x. EPO .

6. The method for determining the storage age of Pu-erh tea based on near-infrared spectroscopy of packaging paper according to claim 5, characterized in that: In step (8), x is used new-EPO =x new Q is the spectral data of the sample to be tested. new Perform EPO calibration.