Method for discriminating quality of white tea and related equipment
By using ultraviolet-visible diffuse reflectance spectroscopy and chemometrics, a set of characteristic wavelengths was selected to identify the origin and grade of white tea. This solved the problems of high cost and low accuracy in white tea quality identification, and achieved efficient and accurate identification of the authenticity of the origin and grade.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI ACAD OF AGRI SCI
- Filing Date
- 2026-04-15
- Publication Date
- 2026-05-12
AI Technical Summary
The quality assessment of white tea suffers from high costs and low accuracy, especially in terms of balancing the authenticity of the place of origin and the grade assessment. Furthermore, the accuracy in identifying low-proportion adulteration and adjacent grades is not high.
The spectral data of white tea samples were obtained and preprocessed using ultraviolet-visible diffuse reflectance spectroscopy combined with chemometrics. The characteristic wavelength set was screened using variable projection importance analysis, and the authenticity of the place of origin and grade were determined. The determination process was optimized by logical constraints and risk labeling.
It reduces the cost of judging the quality of white tea, improves the accuracy of judgment, and achieves high-precision identification of the authenticity of the place of origin and grade, especially the identification of adulteration and subtle differences in grade.
Smart Images

Figure CN122016688A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tea quality assessment technology, and in particular to a method and related equipment for assessing the quality of white tea. Background Technology
[0002] White tea is a lightly fermented tea, rich in active enzymes and polyphenols, making it highly popular among consumers. Among them, geographical indication products such as Fuding white tea have higher market value due to their specific production environment and traditional processing methods. The quality characteristics of white tea are mainly reflected in two aspects: first, the origin attribute—white tea from different production areas exhibits significant differences in intrinsic quality and market price due to variations in climate, soil, and other natural conditions; second, the grading attribute—based on harvesting standards and the tenderness of the raw materials, white tea can be divided into several grades, such as Silver Needle, White Peony, and Shoumei, with substantial price differences between different grades.
[0003] However, driven by high profits, the white tea market is currently facing a severe "trust crisis," with problems such as counterfeiting of origin, difficulty in tracing origins across production areas, and blurred sensory characteristics of adjacent grades leading to misjudgments of white tea quality.
[0004] There is currently no effective solution to the aforementioned problems in the relevant technologies. Summary of the Invention
[0005] The present invention provides a method and related equipment for judging the quality of white tea, which at least partially solves the problems of high cost and low accuracy in judging the quality of white tea in related technologies.
[0006] To address the aforementioned problems, one aspect of this invention provides a method for determining the quality of white tea, comprising: Acquire the ultraviolet-visible diffuse reflectance spectral data of the white tea sample to be tested; wherein, the wavelength range of the spectral data covers the ultraviolet region to the near-infrared short-wave region; The spectral data is subjected to a first preprocessing. Using chemometric methods, based on the spectral data after the first preprocessing and a first characteristic wavelength set, the authenticity of the origin of the white tea sample to be tested is determined, and the authenticity of the origin is obtained. When the authenticity of the origin is determined to be pure material from the target production area, the spectral data undergoes a second preprocessing. Using chemometric methods, based on the spectral data after the second preprocessing and the second characteristic wavelength set, the grade of the white tea sample to be tested is determined, and the grade determination result is obtained.
[0007] In some embodiments, when the origin authenticity determination result is a pure or blended sample from a non-target origin area, the method further includes: Termination of execution level determination; or, Perform a risk assessment and attach a risk label to the assessment results.
[0008] In some embodiments, the method further includes: The ultraviolet-visible diffuse reflectance spectral data of white tea samples with known origin and grade were obtained as the training dataset. Based on variable projection importance analysis, the first contribution value of each wavelength variable in the full-band spectrum of the training dataset to the determination of the authenticity of the place of origin is calculated, and wavelength variables with a first contribution value higher than a first preset threshold are selected as the first feature wavelength set; the first feature wavelength set includes at least wavelength variables located in the ultraviolet region and the visible light region. Based on variable projection importance analysis, the second contribution value of each wavelength variable in the full-band spectrum of the training dataset to the grade discrimination is calculated, and wavelength variables with a second contribution value higher than a second preset threshold are selected as the second feature wavelength set; the second feature wavelength set includes at least wavelength variables located in the near-infrared shortwave region; Wherein, both the first contribution value and the second contribution value are variable projection importance values.
[0009] In some embodiments, the training dataset also includes white tea samples with different blending ratios, which are obtained by mixing white tea samples from at least two known origins in a preset ratio. The authenticity determination of origin is configured to identify pure tea from the target production area, pure tea from non-target production areas, and blended samples of different proportions among the white tea samples to be tested.
[0010] In some of these embodiments, the first preprocessing includes sequentially performing scattering correction and derivative processing; The second preprocessing includes smoothing, scattering correction, and derivative processing in sequence.
[0011] In some of these embodiments, the chemometric method includes at least one of partial least squares discriminant analysis, principal component analysis, and linear discriminant analysis.
[0012] In some embodiments, the wavelength range of the ultraviolet-visible diffuse reflectance spectral data is 190-1100 nm; the white tea sample to be tested is dry tea powder that has been pulverized and sieved; wherein the sieve mesh number corresponding to the sieve treatment is 60-100 mesh.
[0013] To address the aforementioned problems, one aspect of this invention provides a device for determining the quality of white tea, comprising: The spectral data acquisition module is used to acquire the ultraviolet-visible diffuse reflectance spectral data of the white tea sample to be tested; wherein, the wavelength range of the spectral data covers the ultraviolet region to the near-infrared short-wave region; The origin identification module is used to perform a first preprocessing on the spectral data, and using chemometric methods, based on the spectral data after the first preprocessing and a first characteristic wavelength set, to identify the authenticity of the origin of the white tea sample to be tested, and obtain the origin authenticity identification result. The grading module is used to perform a second preprocessing on the spectral data when the authenticity of the origin is determined to be pure material from the target production area. Using chemometric methods, based on the spectral data after the second preprocessing and the second characteristic wavelength set, the module performs grading on the white tea sample to be tested and obtains the grading result.
[0014] In some embodiments, the spectral data acquisition module includes a full-band spectrophotometer, or a combination of a discrete narrowband light source and a photoelectric sensor; wherein the discrete narrowband light source includes at least an ultraviolet light source emitting wavelengths in the ultraviolet region, a visible light source emitting wavelengths in the visible light region, and a near-infrared light source emitting wavelengths in the near-infrared short-wave region.
[0015] To address the aforementioned problems, one aspect of this invention provides a non-transitory machine-readable medium storing computer instructions for causing a computer to execute any of the aforementioned methods for determining the quality of white tea.
[0016] The beneficial effects of this invention are as follows: By acquiring the ultraviolet-visible diffuse reflectance spectral data of the white tea sample to be tested; wherein the wavelength range of the spectral data covers the ultraviolet region to the near-infrared short-wave region; the spectral data undergoes a first preprocessing, and a chemometric method is used to determine the authenticity of the origin of the white tea sample based on the first preprocessed spectral data and a first characteristic wavelength set, obtaining the authenticity determination result; when the authenticity determination result indicates that the tea is pure material from the target production area, the spectral data undergoes a second preprocessing, and a chemometric method is used to determine the authenticity of the origin of the white tea sample based on the second .... The second characteristic wavelength set is a technical means to classify the grade of the white tea sample under test and obtain the grade classification result. It overcomes the problems of high cost and low accuracy of white tea quality classification in related technologies. By collecting wide-band spectral data, it provides a raw data foundation covering polyphenols, pigments and leaf tissue structure for subsequent analysis. Then, the first preprocessing and the first characteristic wavelength set are used to classify the authenticity of the place of origin, and the second preprocessing and the second characteristic wavelength set are used to classify the grade. This achieves the technical effect of reducing the cost of white tea quality classification and improving the accuracy of classification.
[0017] Details of one or more embodiments of the present invention are set forth in the following drawings and description, so that other features, objects and advantages of the invention will be more readily understood. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a schematic diagram of the main process of a method for judging the quality of white tea according to one embodiment of the present invention; Figure 2 This is a schematic diagram showing the distribution of characteristic wavelengths corresponding to the spectral data of white tea samples from different production areas after preprocessing, which is one embodiment of the present invention. Figure 3 This is a schematic diagram of the main modules of a white tea quality discrimination device according to one embodiment of the present invention; Figure 4 This is a schematic diagram of the electronic device of the present invention. Detailed Implementation
[0020] Embodiments of the present invention will now be described in more detail with reference to the accompanying drawings. While some embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the invention. It should be understood that the accompanying drawings and embodiments are for illustrative purposes only and are not intended to limit the scope of protection of the invention.
[0021] The quality judgment of white tea mainly relies on the following technical means: (1) Sensory evaluation method: It relies on the subjective experience of the review experts to comprehensively evaluate the appearance, aroma, liquor color, taste and other aspects of white tea. This method has obvious limitations: on the one hand, it is highly subjective and has poor reproducibility, and the evaluation results of different review experts often differ; on the other hand, for low-proportion adulteration (such as 10%-20% blending) or subtle differences between adjacent grades (such as between White Peony and Shoumei), sensory evaluation is difficult to accurately judge. (2) Physicochemical analysis method: such as high performance liquid chromatography (HPLC) and gas chromatography-mass spectrometry (GC-MS), which can accurately determine the chemical components in white tea. However, these methods require complex sample pretreatment, long detection time, high cost, and are destructive, making it difficult to meet the market's demand for rapid and non-destructive screening of large batches of samples. (3) Spectroscopic analysis techniques: Near-infrared spectroscopy (NIR) has been attempted for tea quality detection. However, NIR spectroscopy mainly reflects the overtone and combination frequency absorption of hydrogen-containing groups (CH, OH, NH), and is not sensitive enough to the direct response of characteristic components such as polyphenols and pigments in tea. At the same time, existing spectroscopic analysis methods mostly focus on single-dimensional qualitative discrimination, such as independent studies of origin traceability or grade classification, and lack the ability to collaboratively discriminate the authenticity of origin (including blending status) and grade attributes.
[0022] In summary, the relevant technologies have the following technical problems in judging the quality of white tea: First, it is difficult to take into account the multi-dimensional judgment requirements of authenticity of origin and grade attributes; second, the ability to identify blending of origin (especially low proportion blending) is insufficient; and third, there is a lack of refined data processing strategies for the differences in characteristics of different quality dimensions, resulting in low judgment accuracy of adjacent grade samples.
[0023] To address the aforementioned problems, embodiments of the present invention provide a method for judging the quality of white tea, such as... Figure 1 As shown, the methods for judging the quality of this white tea mainly include: Step S101: Obtain the ultraviolet-visible diffuse reflectance spectral data of the white tea sample to be tested; wherein, the wavelength range of the spectral data covers the ultraviolet region to the near-infrared short-wave region; Step S102: Perform a first preprocessing on the spectral data, and use chemometric methods to determine the authenticity of the origin of the white tea sample to be tested based on the spectral data after the first preprocessing and the first characteristic wavelength set, and obtain the authenticity determination result of the origin. Step S103: When the authenticity of the origin is determined to be pure material from the target production area, the spectral data is preprocessed a second time. Using chemometric methods, based on the spectral data after the second preprocessing and the second characteristic wavelength set, the grade of the white tea sample to be tested is determined, and the grade determination result is obtained.
[0024] Based on the above setup, the acquisition of broadband spectral data provides a raw data foundation covering polyphenols, pigments, and leaf tissue structure for subsequent analysis. Then, the authenticity of the place of origin is determined by the first preprocessing and the first characteristic wavelength set matched with it, and the grade is determined by the second preprocessing and the second characteristic wavelength set. This achieves the technical effect of reducing the cost of white tea quality determination and improving the accuracy of determination.
[0025] Specifically, the ultraviolet (UV) region exhibits characteristic absorption of polyphenols in white tea, the visible light region is sensitive to pigments, and the near-infrared short-wave region reflects leaf tissue structure, moisture state, and vibrational characteristics of organic functional groups. The combination of these three allows the spectral data to comprehensively cover the chemical and physical information related to origin and grade. Based on step S101, broad-band spectral information covering the UV, visible, and near-infrared short-wave regions is obtained, providing rich material basis data for subsequent identification. Specifically, the wavelength range corresponding to the UV region is 200-400 nm, the visible light region is 400-780 nm, and the near-infrared region is 780-1100 nm.
[0026] In some embodiments, the first preprocessing (such as scattering correction and derivative processing) can eliminate the influence of physical factors such as sample particle size and packing state on the spectrum, highlighting the spectral changes caused by differences in the composition of substances such as polyphenols and pigments. The first characteristic wavelength set is screened from the entire spectrum based on variable projection importance analysis, mainly concentrated in the ultraviolet and visible light regions. These wavelengths correspond precisely to the key chemical information required for origin identification. Chemometric methods (such as partial least squares discriminant analysis) use these features to establish a robust classification model, thereby achieving the identification of the authenticity of the origin and the blending ratio. Based on the above step S102, through preprocessing and characteristic wavelength screening for origin identification, physical interference (such as particle scattering and baseline drift) is effectively suppressed, and the spectral differences related to the origin are enhanced, achieving accurate identification of whether the sample is pure material from the target production area and whether there is blending.
[0027] According to embodiments of the present invention, grade differences (such as tenderness and bud-to-leaf ratio) are typically weak features with small amplitudes in the spectrum, easily masked by noise. The second preprocessing (such as smoothing, scattering correction, and derivative processing sequentially) can reduce random noise, eliminate scattering effects, and amplify subtle absorption changes, effectively extracting grade-related information. The second feature wavelength set is mainly distributed in the near-infrared shortwave region and other bands related to leaf tissue structure and moisture status, focusing on key spectral regions for grade discrimination. Limiting grade discrimination to pure samples from the target production area not only conforms to practical application scenarios (such as geographical indication product protection) but also avoids spectral feature confusion caused by blended samples, thereby improving the accuracy and reliability of grade discrimination. Based on the above step S103, a logical constraint of first determining the production area and then determining the grade is established. Grade discrimination is only performed when the sample is confirmed to be pure material from the target production area (without blending risk), avoiding interference from non-target production areas or blended samples. Simultaneously, the second preprocessing and second feature wavelength set matched to the grade discrimination are used to enhance the weak spectral features related to the grade, achieving high-precision discrimination of processing grades.
[0028] In some embodiments, when the origin authenticity determination result is a pure or blended sample from a non-target production area, the method further includes: terminating the grade determination; or, performing the grade determination and attaching a risk label to the obtained grade determination result.
[0029] Based on the above settings, differentiated processing of samples from non-target production areas or blended samples is achieved, which avoids invalid grade discrimination operations, retains the possibility of obtaining grade information when needed, and warns users of potential doubts about the reliability of the grade results through risk marking.
[0030] According to embodiments of the present invention, the above-mentioned technical features provide two optional processing paths for situations where the origin determination result does not meet the "target origin pure material" condition. The first path, "terminating the grade determination," directly stops the subsequent process, avoiding the computational resources and time costs consumed by continuing to perform grade determination on samples from non-target origins or blended samples. The second path, "performing grade determination and attaching a risk label to the obtained grade determination result," retains the complete determination process, but uses a risk label to indicate to the user that the origin background of the sample is abnormal, therefore the reference value of its grade determination result should be carefully evaluated.
[0031] Specifically, in the discrimination process, the origin authenticity discrimination result has already revealed the sample's source attribute—pure raw material from a non-target production area means that the sample originates entirely from a non-target production area, while blended samples mean that they are partly from the target production area and partly from a non-target production area. In both cases, the spectral characteristics of the sample differ from those of the pure raw material sample from the target production area used to train the grading discrimination model. If grading is performed directly without any processing, the grading results may be biased or misjudged due to differences in origin background. By "terminating the execution of grading discrimination," the process is directly cut off through logical judgment, fundamentally eliminating the possibility of grading due to inconsistent origin background, ensuring that only samples that meet the conditions enter the grading discrimination stage. On the other hand, the method of "performing grading discrimination and attaching risk labels to the results" allows the process to continue, but the information of origin anomalies is presented in association with the grading results through risk labels, allowing users to know the background information when using the grading results, thereby making more accurate judgments. The coexistence of the two methods provides options for different application scenarios: in scenarios where efficiency is the priority, termination can be chosen, while in scenarios where comprehensive information is required (even with risks), execution and marking can be chosen.
[0032] In some embodiments, the method further includes: acquiring ultraviolet-visible diffuse reflectance spectral data of white tea samples with known origins and grades as a training dataset; calculating the first contribution value of each wavelength variable in the full-band spectrum of the training dataset to the determination of origin authenticity based on variable projection importance analysis, and selecting wavelength variables with the first contribution value higher than a first preset threshold as a first feature wavelength set; the first feature wavelength set includes at least wavelength variables located in the ultraviolet and visible light regions; calculating the second contribution value of each wavelength variable in the full-band spectrum of the training dataset to the determination of grade based on variable projection importance analysis, and selecting wavelength variables with the second contribution value higher than a second preset threshold as a second feature wavelength set; the second feature wavelength set includes at least wavelength variables located in the near-infrared shortwave region; wherein, the first contribution value and the second contribution value are both variable projection importance values.
[0033] Based on the above settings, objective and quantitative screening of the characteristic wavelength set for origin identification and grade identification is achieved, ensuring that the extracted characteristic wavelengths are intrinsically related to the corresponding identification tasks, and that the screening results have clear physical meaning and interpretability.
[0034] Specifically, this technique first constructs a training dataset containing known origins and grades, providing a sample foundation for supervised learning in subsequent analyses. Based on this, variable projection importance analysis is used to calculate the contribution of each wavelength variable to origin and grade determination, and wavelength variables with contributions exceeding a preset threshold are selected as the feature wavelength set. This process ensures that the first feature wavelength set relied upon for origin determination is primarily concentrated in the ultraviolet and visible light regions, while the second feature wavelength set relied upon for grade determination includes at least wavelength variables in the near-infrared shortwave region.
[0035] Figure 2 This diagram illustrates the distribution of characteristic wavelengths in the spectral data of white tea samples from different production areas, according to one embodiment of the present invention; for example... Figure 2 As shown, the characteristic wavelength distribution diagrams of the spectral data of white tea samples from six major white tea producing areas—Fujian (FJ), Yunnan (YN), Guizhou (GZ), Guangxi (GX), and Sichuan (SC)—are illustrated. This shows that there are differences in the characteristic wavelengths of different producing areas, providing a basis for determining the place of origin.
[0036] Specifically, the aforementioned technical effects stem from the combined effect of the inherent characteristics of the Variable Importance in the Projection (VIP) method and the threshold screening mechanism. VIP is an indicator used in partial least squares discriminant analysis to measure the explanatory power of each independent variable (wavelength variable) on the dependent variable (origin category or grade category). A higher VIP value indicates a more significant contribution of the wavelength variable to the discrimination. By performing VIP analysis on the spectral data of samples from known origins in the training dataset, the correlation strength between each wavelength and the origin attribute can be quantified; similarly, performing VIP analysis on the spectral data of samples from known grades can quantify the correlation strength between each wavelength and the grade attribute. This quantification process ensures that the selection of characteristic wavelengths no longer relies on subjective experience or trial and error, but is based on a statistically significant contribution assessment.
[0037] According to a specific embodiment of the present invention, the method for judging the quality of white tea provided by the present invention further includes: (1) Feature wavelength screening and mechanism: Based on chemometric methods, feature wavelengths related to the judgment task are extracted. These feature wavelengths include, but are not limited to: the polyphenol and pigment composition difference bands in the ultraviolet and visible light regions (i.e., the first feature wavelength set), used for judging the origin and blending; and the bands in the visible light region and near-infrared region that reflect the differences in bud and leaf structure, tenderness and leaf tissue (i.e., the second feature wavelength set), used for judging the grade. (2) Construction and integration of multidimensional discrimination models: Independent origin authenticity discrimination models (used to identify the origin and blending ratio) and grade classification models (used to identify the grade of white tea) are constructed respectively. In actual discrimination, the spectral data of the white tea sample to be tested (after corresponding preprocessing) are input into the above independent models respectively, and the origin, blending risk and grade results of the sample are output at the same time; or the logical constraint mode is adopted, the origin authenticity discrimination is performed first, and when the result is the target production area and there is no blending risk, the grade classification model is triggered; when there is a blending risk, the grade discrimination result is marked or the output is restricted. (3) Model construction and verification: The collected raw spectral data are randomly divided into training set and test set according to a preset ratio to ensure the uniformity of sample distribution. Based on the training set data, a supervised learning classification model (such as partial least squares discriminant analysis) is constructed, and the generalization ability of the model is externally verified using the test set to ensure the robustness of the discrimination results.
[0038] In some embodiments, the introduction of a first preset threshold and a second preset threshold further enhances the objectivity of the screening. By setting thresholds, only wavelength variables with VIP values higher than the threshold are retained, while wavelengths that contribute little to the discrimination or may introduce noise are eliminated, thereby ensuring that the selected feature wavelength set has high discriminative value and stability.
[0039] In some examples, the first characteristic wavelength set includes wavelength variables located in the ultraviolet and visible light regions. This distribution characteristic is consistent with the inherent logic of VIP analysis results—the absorption of polyphenols in the ultraviolet region and the absorption of pigments in the visible light region are the main material basis for the differences in white tea from different origins. Therefore, VIP analysis naturally points the wavelengths with higher contribution to these regions. The second characteristic wavelength set includes wavelength variables located in the near-infrared shortwave region, which also stems from the VIP analysis's identification of grade-related characteristics—the near-infrared shortwave region is sensitive to physical characteristics related to tenderness, such as leaf tissue structure and moisture status. Therefore, wavelengths with higher contribution naturally concentrate in this region. This limitation of the distribution area is not artificially predetermined but rather an objective reflection of the VIP analysis results, and it also provides material-level interpretability for the screening results.
[0040] In some embodiments, the training set also includes white tea samples with different blending ratios, which are obtained by mixing white tea samples from at least two known origins in a preset ratio; wherein, the origin authenticity determination is configured to identify pure white tea from the target origin, pure white tea from a non-target origin, and blended samples with different ratios in the white tea samples to be tested.
[0041] Based on the above settings, the ability to quantitatively identify the blending status of white tea samples was realized, expanding the determination of the authenticity of the place of origin from a simple binary "yes / no" classification to a refined classification that can distinguish between pure materials and different blending ratios, providing a more accurate screening basis for the selective execution of subsequent grade determination.
[0042] According to embodiments of the present invention, the aforementioned technical features first introduce blended samples obtained by mixing at least two samples from known origins in a preset ratio during the training dataset construction stage. This ensures that the training dataset covers a complete range of sample types, from pure raw materials from the target origin to different blending ratios and even pure raw materials from non-target origins. Based on this, the origin authenticity determination is configured to output three discrimination results: pure raw materials from the target origin, pure raw materials from non-target origins, and blended samples with different ratios, achieving a refined distinction of the sample's source status. Furthermore, the aforementioned technical effect stems from the matching between the completeness of the training dataset and the output capability of the discrimination model. In conventional origin determination methods, the training dataset typically only contains pure raw material samples from each origin. The model can only learn the spectral characteristics of pure raw material samples, thus its output capability is limited to determining which origin a sample belongs to, and cannot identify the blending status. When encountering blended samples, such models often classify them as belonging to a category similar to the origin with a higher proportion of the blending components, but cannot inform the user whether the sample is pure or blended, let alone indicate the existence of the blending ratio.
[0043] In some embodiments, gradient blended samples are introduced into the training dataset, allowing the model to be exposed to continuous spectral changes from pure raw materials to different blending ratios and then to pure raw materials from another origin during the learning process. The spectral characteristics of these blended samples are not a simple linear superposition of the pure raw material spectra, but rather exhibit a continuous spectral pattern that varies with the blending ratio. In variable projection importance analysis, when calculating the contribution of wavelength variables to origin discrimination, because the training dataset includes blended samples, the first set of feature wavelengths selected not only distinguishes the origin of pure raw materials but also has the ability to distinguish different blending ratios.
[0044] In some examples, origin authenticity determination is configured to identify three scenarios, essentially expanding the output space of the discrimination model from binary (yes / no) to multi-valued classification or continuous value prediction. This configuration allows the model to utilize spectral feature variations learned from training on blended samples to perform more refined pattern matching on the spectrum of the sample under test, thereby outputting its specific origin status—whether it is pure material from a single origin, entirely from another origin, or a blend in some proportion somewhere in between. This capability provides crucial information for subsequent processing: only samples identified as pure material from the target origin are considered valuable for grading; samples from non-target origin pure materials or blended samples, due to their impure origin background, either have their grading terminated or require additional risk labeling.
[0045] In some embodiments, the first preprocessing includes sequentially performing scattering correction and derivative processing; the second preprocessing includes sequentially performing smoothing, scattering correction, and derivative processing.
[0046] Based on the above settings, precise adaptation processing is achieved to address the differences in spectral characteristics between origin determination and grade determination. This ensures that the preprocessing method matches the inherent requirements of the determination task, thereby maximizing the extraction of effective information relevant to each task while suppressing irrelevant interference. Specifically, a preprocessing flow of "sequential scattering correction and derivative processing" is set up for origin determination, while a preprocessing flow of "sequential smoothing, scattering correction, and derivative processing" is set up for grade determination. The two preprocessing methods differ in their step composition and execution order. This differentiated design ensures that the spectral data is optimized for the characteristics of different tasks before entering the determination model.
[0047] In some embodiments, for origin identification, the spectral differences of white tea from different origins mainly stem from differences in chemical composition such as polyphenols and pigments. These differences are typically manifested in the original spectrum as noticeable changes in the intensity or position of absorption peaks. Scatter correction can eliminate physical interference, making chemical differences more prominent; derivative processing further amplifies the spectral changes caused by these chemical differences. Therefore, the sequential combination of scatter correction followed by derivative processing can effectively remove physical noise while enhancing origin-related chemical characteristics, without the need for an intermediate smoothing step—because the origin-related spectral signal itself is strong, the noise impact is relatively limited, and smoothing is not necessary.
[0048] In some examples, for grading, the spectral differences between different grades of white tea mainly stem from variations in tissue structure caused by tenderness, bud-leaf ratio, and moisture binding state. These differences typically manifest as subtle features with small amplitudes in the spectrum, easily masked by noise. Smoothing, as the first step, effectively reduces random noise, providing a cleaner signal foundation for subsequent processing. Then, scattering correction eliminates interference from physical factors such as particle size, highlighting information related to tissue structure. Finally, derivative processing further amplifies the subtle absorption changes related to tenderness. The underlying logic of this specific order is that the grade characteristic signal is weak, necessitating smoothing to remove noise that might obscure the signal, followed by scattering correction to eliminate physical interference, and finally, derivative processing to amplify the subtle differences in chemical and physical characteristics. If the order is reversed—for example, performing derivative processing before smoothing—derivative processing will simultaneously amplify both the signal and noise, making subsequent smoothing ineffective in removing the amplified noise, leading to a deterioration in the signal-to-noise ratio.
[0049] In some of these embodiments, the chemometric methods include at least one of partial least squares discriminant analysis, principal component analysis, and linear discriminant analysis.
[0050] Based on the above settings, a variety of chemometric methods suitable for classification modeling of high-dimensional spectral data are provided, enabling those skilled in the art to select the most suitable analytical tool according to data characteristics and application requirements, thereby ensuring the robustness and accuracy of the discrimination model. The inherent characteristics and adaptability to spectral data of the three listed chemometric methods are as follows: Partial Least Squares Discriminant Analysis (PLS-DA) is one of the most commonly used classification methods in the field of spectral analysis. It can perform dimensionality reduction and classification modeling while processing high-dimensional spectral data, making it particularly suitable for situations where the number of variables far exceeds the number of samples. PLS-DA effectively utilizes classification information across the entire spectral band by extracting latent variables that are most correlated with the class labels, and can also output variable projection importance values, providing a basis for feature wavelength selection.
[0051] Principal Component Analysis (PCA) is an unsupervised dimensionality reduction method that transforms the original spectral variables into a few principal components through orthogonal transformation. These principal components retain the variance information of the original data to the greatest extent possible. Before discriminant analysis, PCA can be used for data visualization, outlier detection, and dimensionality reduction preprocessing, providing simplified data input for subsequent classification modeling. Its advantage lies in its independence from class labels, objectively revealing the inherent structure of the data.
[0052] Linear Discriminant Analysis (LDA) is a supervised dimensionality reduction classification method that aims to find the projection direction that maximizes between-class scatter and minimizes within-class scatter. LDA is effective at extracting discriminative information between classes when processing spectral data with clearly defined class labels, and its model interpretability is relatively strong. When spectral data meets certain statistical assumptions, LDA often achieves good classification results.
[0053] Although the three methods described above differ in their algorithmic principles, they can all process preprocessed spectral data and be used in conjunction with a selected set of characteristic wavelengths to construct classification models for origin or grade identification. Based on this setup, it not only covers the PLS-DA method used in the embodiments of this application but also reserves space for other potentially applicable methods. This allows those skilled in the art to flexibly select the most suitable chemometric tools when faced with different data characteristics or application requirements, thereby ensuring the adaptability and reliability of the discrimination model.
[0054] In some of these embodiments, the wavelength range of the ultraviolet-visible diffuse reflectance spectral data is 190-1100 nm; the white tea sample to be tested is dry tea powder that has been pulverized and sieved; wherein the sieve mesh corresponding to the sieve treatment is 60-100 mesh.
[0055] Based on the above settings, the integrity, reproducibility, and stability of the spectral data were ensured, providing a high-quality and highly comparable data foundation for subsequent differential preprocessing and chemometric analysis, while eliminating interference from non-chemical factors introduced by differences in sample physical states. Specifically, the wavelength range was first limited to 190-1100 nm, which fully covers the ultraviolet, visible, and near-infrared short-wave regions, ensuring that spectral information related to white tea quality, such as polyphenols (UV absorption), pigments (visible absorption), and leaf tissue structure and moisture state (near-infrared short-wave response), could be comprehensively captured. Secondly, the samples to be tested were limited to pulverized and sieved dry tea powder, and the powder particle size was controlled by a sieve mesh size of 60-100 mesh, ensuring that the samples were physically consistent.
[0056] According to embodiments of the present invention, the selection of the wavelength range of 190-1100 nm is based on the spectral response characteristics of key quality components in white tea: the ultraviolet region (especially 200-400 nm) is the characteristic absorption region of polyphenols such as catechins, which are closely related to the production environment and processing technology of white tea; the visible light region (400-780 nm) is the absorption region of pigments such as chlorophyll and carotenoids, and white teas of different origins and grades show significant differences in pigment composition and content due to differences in raw material tenderness and processing technology; the near-infrared short-wave region (780-1100 nm) can respond to the overtone absorption of hydrogen-containing groups such as CH and OH, and this information is related to leaf tissue structure, water binding state, and vibrational characteristics of organic functional groups, providing a physical structural basis for grade determination. Continuous coverage of 190-1100 nm ensures that all the above-mentioned information related to origin and grade can be completely acquired in a single acquisition, avoiding the omission of key information due to missing wavelengths.
[0057] In some embodiments, the sample to be tested is prepared as dry tea powder and then sieved. This process aims to eliminate spectral variations caused by differences in the physical morphology of the sample. During spectral acquisition, factors such as surface morphology, bulk density, and orientation distribution of whole leaves or coarse-grained samples can lead to uncertainties in light scattering paths and intensities. These physical factors often mask or interfere with spectral features related to chemical composition. Through pulverization, the sample is transformed into fine, uniform particles; sieving further controls the particle size within the 60-100 mesh range. A 60-mesh (approximately 250 micrometers) sieve ensures that sufficiently large particles are retained, avoiding uneven light scattering due to excessively large particles; a 100-mesh (approximately 150 micrometers) sieve avoids potential changes in sample composition or electrostatic adsorption caused by over-pulverization. This particle size range ensures sample uniformity while maintaining the integrity of the original tea's chemical characteristics.
[0058] Understandably, the sample powders after the above standardization process exhibit a concentrated particle size distribution and stable packing density, resulting in good consistency in light scattering effects during spectral acquisition. This ensures that the measured spectral data primarily reflects differences in the chemical composition of the samples rather than differences in their physical morphology. This high-quality, highly reproducible data provides a reliable foundation for subsequent preprocessing, characteristic wavelength screening, and chemometric modeling, ensuring the stability and generalization ability of the discrimination model.
[0059] The method for judging the quality of white tea provided in this embodiment of the invention utilizes the acquisition of ultraviolet-visible diffuse reflectance spectral data of the white tea sample to be tested. The wavelength range of the spectral data covers the ultraviolet region to the near-infrared short-wave region. The spectral data undergoes a first preprocessing step, and a chemometric method is used to determine the authenticity of the origin of the white tea sample based on the first preprocessed spectral data and a first characteristic wavelength set, yielding an authenticity determination result. When the authenticity determination result indicates pure material from the target production area, a second preprocessing step is performed on the spectral data. A chemometric method is then used to determine the grade of the white tea sample based on the second preprocessed spectral data and a second characteristic wavelength set, yielding a grade determination result. By acquiring broad-band spectral data, a raw data foundation covering polyphenols, pigments, and leaf tissue structure is provided for subsequent analysis. Then, matching first preprocessing and first characteristic wavelength set are used for origin authenticity determination, and second preprocessing and second characteristic wavelength set are used for grade determination. This achieves the technical effect of reducing the cost of judging white tea quality and improving the accuracy of judgment.
[0060] According to an embodiment of the present invention, a method for judging the quality of white tea based on ultraviolet-visible diffuse reflectance spectroscopy (which can be used to determine the place of origin and blending ratio) is provided, including: (1) Sample preparation First, in order to construct a widely representative adulteration detection model, this embodiment collected Fuding white tea (Baihao Yinzhen) from 2022, 2024, and 2025 as base samples. At the same time, Baihao Yinzhen samples from other producing areas during the same period were collected as adulteration raw materials, covering Guizhou (such as Liupanshui, Weng'an, and Fanjingshan in Guizhou), Sichuan (such as Meishan and Ya'an), and Guangxi (such as Sanjiang).
[0061] Using Fuding Baihao Yinzhen as the base, white tea from non-Fuding producing areas was used as an admixture to construct a gradient model with different blending ratios. The specific design is as follows: 2022 group: Guizhou, Meishan and Ya'an white tea from Sichuan were blended into Fuding white tea at ratios of 0%, 10%, 20%, ..., 100%, resulting in 31 blended samples. 2024 group: Guangxi, Guizhou and Sichuan were blended into Fuding white tea at the same ratio, resulting in 41 blended samples. (3) 2025 group: Guizhou Fanjingshan white tea was blended into Fuding white tea at the same ratio, resulting in 11 samples. In addition, pure Fuding Baihao Yinzhen samples from 2022, 2024 and 2025 were collected, totaling 9 samples.
[0062] A total of 92 samples were collected. All samples were ground and passed through an 80-mesh sieve to obtain uniform powder, which was used for subsequent spectroscopic determination and origin identification experiments.
[0063] (2) Spectroscopic determination The spectral acquisition conditions are as follows: incident angle: 8°, scanning wavelength: 190-1100 nm, spectral bandwidth: 1 nm, scanning speed: 800 nm / min, each sample is scanned three times, and the average value is taken as the sample spectral data.
[0064] (3) Spectral data processing and place of origin identification The raw spectral data were imported into SIMCA software, and a partial least squares discriminant analysis (PLS-DA) model was established for all samples to preliminarily determine their origin. Cross-validation results showed a prediction accuracy of 86.96%, indicating that the raw spectra can be used to distinguish between Fuding white tea and non-Fuding white tea, but there are still some misclassifications.
[0065] The original spectral data and corresponding labels were imported into the Python environment; a PLS-DA model was built for the entire sample, with a prediction accuracy of 91.3%; the samples were randomly divided into a training set (75%) and a test set (25%) to ensure uniform sample distribution; the training set was used for spectral preprocessing, PLS-DA modeling and latent variable (LV) optimization, and the test set was used for external validation of the model's discriminative performance.
[0066] The training set sample spectra underwent various preprocessing methods: multivariate scattering correction (MSC), standard normal transformation (SNV), Savitzky-Golay smoothing (SG), first derivative (1st), second derivative (2nd), and combinations such as MSC+1st, SG+1st, MSC+2nd, SG+2nd, MSC+SG+1st, and MSC+SG+2nd. Preprocessing parameters, such as the MSC baseline spectrum and scattering correction coefficients, and the mean and standard deviation of SNV, were fitted using the training set as a baseline.
[0067] The preprocessed parameters obtained from fitting the training set are applied to the spectrum of the test set to ensure that the training set and the test set are processed in a consistent manner and to prevent data leakage.
[0068] A PLS-DA model was constructed using the training set after preprocessing the spectra and corresponding labels. The number of LVs was determined through cross-validation. The trained model was then used to predict on the test set, and the prediction results were binarized to distinguish between Fuding white tea and non-Fuding white tea.
[0069] Under the original spectral conditions (without preprocessing), the accuracy on the training set was 97.1%, and the accuracy on the test set was 91.3%. The results show that the preprocessed spectral data significantly improves the discrimination performance, with the MSC+1st, SNV, and 1st methods all achieving 100% accuracy on the test set.
[0070] Furthermore, in this specific implementation, the step of calculating the VIP value includes: calculating the variable importance in projection (VIP) value of each wavelength in the model of the PLS-DA model established by the spectral data of the training set after MSC+1st preprocessing, which is used to measure the contribution of each wavelength to the classification.
[0071] The steps for key wavelength screening (i.e., obtaining the first set of characteristic wavelengths) include: dividing wavelengths into different levels based on their VIP values: wavelengths with a VIP value > 1 are important characteristic wavelengths of the model and can be used to determine the origin of samples; wavelengths with a VIP value between 1 and 2 are moderately important wavelengths and contribute to the determination; wavelengths with a VIP value between 2 and 3 are highly important wavelengths and contribute significantly to the determination, totaling 41 wavelengths; wavelengths with a VIP value > 3 are extremely important wavelengths and have the strongest effect on the determination (not observed in this experiment). It should be noted that in practical applications, the specific value of the first preset threshold corresponding to the first set of characteristic wavelengths can be adjusted according to the required accuracy.
[0072] The discriminant contribution analysis shows that wavelengths with high VIP values are mainly concentrated in the ultraviolet region (approximately 200-400 nm) and the visible light region (approximately 400-780 nm), corresponding to the absorption characteristics of polyphenols and chlorophyll in white tea, respectively. The selected key wavelengths maintain high discriminant accuracy on the external test set, demonstrating that this method can effectively identify the characteristic spectral regions of Fuding white tea and non-Fuding white tea.
[0073] In another specific embodiment, the VIP value for each wavelength is calculated, and the characteristic wavelength regions that significantly contribute to the discrimination are selected as follows: 200-330 nm, 360-390 nm, 420-435 nm, 460-475 nm, 510-555 nm, 580-608 nm, 650-680 nm, 700-728 nm, and 820-1087 nm. These wavelengths mainly correspond to the optical responses related to polyphenols, chlorophyll, and leaf structure in white tea, providing key basis for grade discrimination. It is understood that key wavelength screening can reduce spectral dimensions, improve model stability, and maintain high accuracy. The method in this specific embodiment of the invention can quickly and non-destructively achieve accurate classification of different grades of white tea (Silver Needle, White Peony, Spring Shoumei, Autumn Shoumei, and Gongmei) in the Fuding region.
[0074] It should be noted that the specific values involved in the above-mentioned specific timing methods are merely examples and are not intended to limit the embodiments of the present invention.
[0075] Compared with related technologies, the embodiments of the present invention have at least the following advantages: (1) They enable rapid, non-destructive, and independent identification of the multidimensional attributes of white tea; (2) They significantly improve the accuracy of identification of origin, blending, and grade through task-driven differentiated preprocessing and characteristic wavelength screening; (3) They do not rely on complex sample processing or organic solvents, have a fast detection speed, and are suitable for large-scale and portable detection applications; (4) They provide an scalable intelligent detection system platform, which is convenient for integration and practical application.
[0076] According to another specific embodiment of the present invention, a method for judging the quality of white tea based on ultraviolet-visible diffuse reflectance spectroscopy (which can be used for grade determination) includes: (1) Select Fuding white tea harvested from 2013 to 2025, and classify it into five categories according to grade: Baihao Yinzhen, Bai Mudan, Chun Shoumei, Qiu Shoumei, and Gongmei.
[0077] (2) The spectral acquisition conditions are as follows: incident angle: 8°, scanning wavelength 190-1100 nm, spectral bandwidth 1 nm, scanning speed 800 nm / min, each sample is scanned three times, and the average value is taken as the sample spectral data.
[0078] (3) Spectral data preprocessing: Various preprocessing methods were compared. The results showed that the combined preprocessing method of Savitzky-Golay smoothing (SG), multivariate scattering correction (MSC) and first derivative (1st) can achieve a PLS-DA discrimination rate of 100%.
[0079] (4) Feature wavelength screening (i.e., screening to obtain the second feature wavelength set): Under the SG+MSC+1st preprocessing conditions, the VIP value of each wavelength is calculated, and the feature wavelength regions that contribute significantly to the discrimination are screened as follows: 200-330 nm, 360-390 nm, 420-435 nm, 460-475 nm, 510-555 nm, 580-608 nm, 650-680 nm, 700-728 nm, 820-1087 nm. These wavelengths mainly correspond to the optical responses related to polyphenols, chlorophyll and leaf structure in white tea, providing key basis for grade discrimination.
[0080] (5) Judgment Results and Implementation Effects. Under the SG+MSC+1st conditions, the PLS-DA model can achieve 100% correct discrimination for the five types of white tea on both the training and test sets. Key wavelength screening can reduce spectral dimensions, improve model stability, and maintain high accuracy. The specific implementation method provided in this embodiment of the invention can quickly and non-destructively achieve accurate classification of different grades of white tea (Silver Needle, White Peony, Spring Shoumei, Autumn Shoumei, and Gongmei) in the Fuding area.
[0081] Based on the white tea quality discrimination method provided in the embodiments of the present invention, the embodiments of the present invention also provide a white tea quality discrimination device; Figure 3 As shown, the white tea quality judging device 300 includes: The spectral data acquisition module 301 is used to acquire the ultraviolet-visible diffuse reflectance spectral data of the white tea sample to be tested; wherein, the wavelength range of the spectral data covers the ultraviolet region to the near-infrared short-wave region; The origin identification module 302 is used to perform a first preprocessing on the spectral data and, using chemometric methods, based on the spectral data after the first preprocessing and the first characteristic wavelength set, to identify the authenticity of the origin of the white tea sample to be tested and obtain the origin authenticity identification result. The grading module 303 is used to perform a second preprocessing on the spectral data when the authenticity of the origin is determined to be pure material from the target production area. Using chemometric methods, based on the spectral data after the second preprocessing and the second characteristic wavelength set, the grading of the white tea sample to be tested is determined, and the grading result is obtained.
[0082] Based on the above setup, the acquisition of broadband spectral data provides a raw data foundation covering polyphenols, pigments, and leaf tissue structure for subsequent analysis. Then, the authenticity of the place of origin is determined by the first preprocessing and the first characteristic wavelength set matched with it, and the grade is determined by the second preprocessing and the second characteristic wavelength set. This achieves the technical effect of reducing the cost of white tea quality determination and improving the accuracy of determination.
[0083] In some embodiments, the grade discrimination module 303 is further configured to: when the origin authenticity discrimination result is a pure material or blended sample from a non-target production area, the method further includes: terminating the grade discrimination; or, performing the grade discrimination and attaching a risk label to the obtained grade discrimination result.
[0084] Based on the above settings, differentiated processing of samples from non-target production areas or blended samples is achieved, which avoids invalid grade discrimination operations, retains the possibility of obtaining grade information when needed, and warns users of potential doubts about the reliability of the grade results through risk marking.
[0085] In some embodiments, the white tea quality discrimination device 300 further includes a feature wavelength set construction module, used for: acquiring ultraviolet-visible diffuse reflectance spectral data of white tea samples with known origin and grade as a training set; calculating the first contribution value of each wavelength variable in the full-band spectrum of the training dataset to the authenticity of the origin based on variable projection importance analysis, and selecting wavelength variables with the first contribution value higher than a first preset threshold as a first feature wavelength set; the first feature wavelength set includes at least wavelength variables located in the ultraviolet and visible light regions; calculating the second contribution value of each wavelength variable in the full-band spectrum of the training dataset to the grade discrimination based on variable projection importance analysis, and selecting wavelength variables with the second contribution value higher than a second preset threshold as a second feature wavelength set; the second feature wavelength set includes at least wavelength variables located in the near-infrared shortwave region; wherein, the first contribution value and the second contribution value are both variable projection importance values.
[0086] Based on the above settings, objective and quantitative screening of the characteristic wavelength set for origin identification and grade identification is achieved, ensuring that the extracted characteristic wavelengths are intrinsically related to the corresponding identification tasks, and that the screening results have clear physical meaning and interpretability.
[0087] In some of these embodiments, the training dataset also includes white tea samples with different blending ratios, which are obtained by mixing white tea samples from at least two known origins in a preset ratio; wherein, the origin authenticity determination is configured to identify pure white tea from the target origin, pure white tea from a non-target origin, and blended samples with different ratios in the white tea samples to be tested.
[0088] Based on the above settings, the ability to quantitatively identify the blending status of white tea samples was realized, expanding the determination of the authenticity of the place of origin from a simple binary "yes / no" classification to a refined classification that can distinguish between pure materials and different blending ratios, providing a more accurate screening basis for the selective execution of subsequent grade determination.
[0089] In some embodiments, the first preprocessing includes sequentially performing scattering correction and derivative processing; the second preprocessing includes sequentially performing smoothing, scattering correction, and derivative processing.
[0090] Based on the above settings, precise adaptation processing is achieved to address the differences in spectral characteristics between origin determination and grade determination. This ensures that the preprocessing method matches the inherent requirements of the determination task, thereby maximizing the extraction of effective information relevant to each task while suppressing irrelevant interference. Specifically, a preprocessing flow of "sequential scattering correction and derivative processing" is set up for origin determination, while a preprocessing flow of "sequential smoothing, scattering correction, and derivative processing" is set up for grade determination. The two preprocessing methods differ in their step composition and execution order. This differentiated design ensures that the spectral data is optimized for the characteristics of different tasks before entering the determination model.
[0091] In some of these embodiments, the chemometric methods include at least one of partial least squares discriminant analysis, principal component analysis, and linear discriminant analysis.
[0092] Based on the above settings, a variety of chemometric methods suitable for classification modeling of high-dimensional spectral data are provided, enabling those skilled in the art to select the most suitable analytical tool according to data characteristics and application requirements, thereby ensuring the robustness and accuracy of the discrimination model.
[0093] In some of these embodiments, the wavelength range of the ultraviolet-visible diffuse reflectance spectral data is 190-1100 nm; the white tea sample to be tested is dry tea powder that has been pulverized and sieved; wherein the sieve mesh corresponding to the sieve treatment is 60-100 mesh.
[0094] Based on the above settings, the integrity, reproducibility, and stability of the spectral data were ensured, providing a high-quality and highly comparable data foundation for subsequent differential preprocessing and chemometric analysis, while eliminating interference from non-chemical factors introduced by differences in sample physical states. Specifically, the wavelength range was first limited to 190-1100 nm, which fully covers the ultraviolet, visible, and near-infrared short-wave regions, ensuring that spectral information related to white tea quality, such as polyphenols (UV absorption), pigments (visible absorption), and leaf tissue structure and moisture state (near-infrared short-wave response), could be comprehensively captured. Secondly, the samples to be tested were limited to pulverized and sieved dry tea powder, and the powder particle size was controlled by a sieve mesh size of 60-100 mesh, ensuring that the samples were physically consistent.
[0095] In some embodiments, the spectral data acquisition module includes a full-band spectrophotometer, or a combination of a discrete narrowband light source and a photoelectric sensor; wherein the discrete narrowband light source includes at least an ultraviolet light source emitting wavelengths in the ultraviolet region, a visible light source emitting wavelengths in the visible light region, and a near-infrared light source emitting wavelengths in the near-infrared short-wave region.
[0096] Based on the above settings, a spectral acquisition hardware system that balances versatility and specificity is constructed by providing two optional hardware implementation methods for the spectral data acquisition module: a full-band spectrophotometer or a combination of discrete narrowband light sources and photoelectric sensors, and by specifically defining the wavelength range of the discrete narrowband light sources. This enables flexible configuration of the spectral acquisition hardware, allowing for the acquisition of continuous broad spectra using a full-band spectrophotometer to meet research or diverse application needs, and also enabling miniaturized, low-cost, and high-speed dedicated detection using a combination of discrete narrowband light sources customized for characteristic wavelengths. This provides adaptable hardware solutions for different application scenarios.
[0097] It should be noted that the specific modules in the aforementioned white tea quality assessment device are defined primarily based on the corresponding operations performed, and are not intended to limit the specific modules.
[0098] This invention also provides a non-transitory machine-readable medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of this invention.
[0099] This invention also provides a computer program product, including a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform the methods of embodiments of this invention. The computer program product should be understood as a software product that primarily implements the methods of this invention through a computer program.
[0100] This invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method of this invention.
[0101] refer to Figure 4 The present invention will now be described in the form of a structural block diagram of an electronic device that can serve as an embodiment of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0102] like Figure 4As shown, the electronic device includes a computing unit 401, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. The RAM 403 may also store various programs and data required for the operation of the electronic device. The computing unit 401, ROM 402, and RAM 403 are interconnected via a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.
[0103] Multiple components in the electronic device are connected to I / O interface 405, including: input unit 406, output unit 407, storage unit 408, and communication unit 409. Input unit 406 can be any type of device capable of inputting information into the electronic device. Input unit 406 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of the electronic device. Output unit 407 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 408 may include, but is not limited to, disks and optical discs. Communication unit 409 allows the electronic device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, and / or wireless communication transceivers, such as Bluetooth devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0104] The computing unit 401 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, CPUs, graphics processing units (GPUs), various special-purpose artificial intelligence (AI) computing units, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. The computing unit 401 performs the various methods and processes described above. For example, in some embodiments, the method embodiments of the present invention can be implemented as a computer program tangibly contained in a machine-readable medium, such as storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on an electronic device via ROM 402 and / or communication unit 409. In some embodiments, the computing unit 401 can be configured to perform the methods described above by any other suitable means (e.g., by means of firmware).
[0105] Computer programs for implementing the methods of embodiments of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0106] In the context of embodiments of the present invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable signal medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, or infrared systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0107] It should be noted that the term "comprising" and its variations used in the embodiments of the present invention are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The modifications of "one" and "multiple" mentioned in the embodiments of the present invention are illustrative and not restrictive. Those skilled in the art should understand that, unless explicitly indicated otherwise in the context, they should be understood as "one or more".
[0108] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in the embodiments of the present invention are all information and data that have been permitted by the user or have been fully agreed upon by all parties. Furthermore, the collection, use and processing of the relevant data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to agree or refuse.
[0109] The steps described in the method embodiments provided by this invention can be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of protection of this invention is not limited in this respect.
[0110] The term "embodiment" in this specification refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of the invention. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply independence or alternativeity from other embodiments. The various embodiments in this specification are described in a related manner, with reference to each other for similar or identical parts. In particular, for apparatus, device, and system embodiments, since they are substantially similar to method embodiments, the description is relatively simple, and relevant details are referred to in the description of the method embodiments.
[0111] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.
Claims
1. A method for judging the quality of white tea, characterized in that, include: Acquire the ultraviolet-visible diffuse reflectance spectral data of the white tea sample to be tested; wherein, the wavelength range of the spectral data covers the ultraviolet region to the near-infrared short-wave region; The spectral data is subjected to a first preprocessing. Using chemometric methods, based on the spectral data after the first preprocessing and a first characteristic wavelength set, the authenticity of the origin of the white tea sample to be tested is determined, and the authenticity of the origin is obtained. When the authenticity of the origin is determined to be pure material from the target production area, the spectral data undergoes a second preprocessing. Using chemometric methods, based on the spectral data after the second preprocessing and the second characteristic wavelength set, the grade of the white tea sample to be tested is determined, and the grade determination result is obtained.
2. The method according to claim 1, characterized in that, When the result of the origin authenticity determination is that the sample is a pure material or blended sample from a non-target production area, the method further includes: Termination of execution level determination; or, Perform a risk assessment and attach a risk label to the assessment results.
3. The method according to claim 1, characterized in that, The method further includes: The ultraviolet-visible diffuse reflectance spectral data of white tea samples with known origin and grade were obtained as the training dataset. Based on variable projection importance analysis, the first contribution value of each wavelength variable in the full-band spectrum of the training dataset to the determination of the authenticity of the place of origin is calculated, and wavelength variables with a first contribution value higher than a first preset threshold are selected as the first feature wavelength set; the first feature wavelength set includes at least wavelength variables located in the ultraviolet region and the visible light region. Based on variable projection importance analysis, the second contribution value of each wavelength variable in the full-band spectrum of the training dataset to the grade discrimination is calculated, and wavelength variables with a second contribution value higher than a second preset threshold are selected as the second feature wavelength set; the second feature wavelength set includes at least wavelength variables located in the near-infrared shortwave region; Wherein, both the first contribution value and the second contribution value are variable projection importance values.
4. The method according to claim 3, characterized in that, The training dataset also includes white tea samples with different blending ratios, which are obtained by mixing white tea samples from at least two known origins in a preset ratio. The authenticity determination of origin is configured to identify pure tea from the target production area, pure tea from non-target production areas, and blended samples of different proportions among the white tea samples to be tested.
5. The method according to claim 1, characterized in that, The first preprocessing includes sequentially performing scattering correction and derivative processing; The second preprocessing includes smoothing, scattering correction, and derivative processing in sequence.
6. The method according to claim 1, characterized in that, The chemometric methods include at least one of partial least squares discriminant analysis, principal component analysis, and linear discriminant analysis.
7. The method according to claim 1, characterized in that, The wavelength range of the ultraviolet-visible diffuse reflectance spectral data is 190-1100 nm; the white tea sample to be tested is dry tea powder that has been pulverized and sieved; wherein, the sieve mesh corresponding to the sieve treatment is 60-100 mesh.
8. A device for judging the quality of white tea, characterized in that, include: The spectral data acquisition module is used to acquire the ultraviolet-visible diffuse reflectance spectral data of the white tea sample to be tested; wherein, the wavelength range of the spectral data covers the ultraviolet region to the near-infrared short-wave region; The origin identification module is used to perform a first preprocessing on the spectral data, and using chemometric methods, based on the spectral data after the first preprocessing and a first characteristic wavelength set, to identify the authenticity of the origin of the white tea sample to be tested, and obtain the origin authenticity identification result. The grading module is used to perform a second preprocessing on the spectral data when the authenticity of the origin is determined to be pure material from the target production area. Using chemometric methods, based on the spectral data after the second preprocessing and the second characteristic wavelength set, the module performs grading on the white tea sample to be tested and obtains the grading result.
9. The apparatus according to claim 8, characterized in that, The spectral data acquisition module includes a full-band spectrophotometer, or a combination of a discrete narrowband light source and a photoelectric sensor; wherein the discrete narrowband light source includes at least an ultraviolet light source emitting wavelengths in the ultraviolet region, a visible light source emitting wavelengths in the visible light region, and a near-infrared light source emitting wavelengths in the near-infrared short-wave region.
10. A non-transitory machine-readable medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 7.