A taiping monkey king tea near infrared spectrum FS-SC semi-supervised feature selection method
Through the FS-SC semi-supervised feature selection method, the Fisher Score and Silhouette Coefficient scores were used to calculate and select feature variables, and a partial least squares prediction and discrimination model was established. This solved the problem of insufficient feature mining of near-infrared spectral data and achieved efficient distinction and accurate identification of the Taiping Houkui tea producing area.
Patent Information
- Application Number
- CN202310921911.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-24
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2043-07-24
AI Technical Summary
Existing technologies rely on a small number of labeled near-infrared spectral samples and are unable to fully explore key characteristic variables, resulting in poor predictive performance of the correction model and an inability to effectively distinguish between the core and non-core production areas of Taiping Houkui tea.
The FS-SC semi-supervised feature selection method of Taiping Houkui tea near-infrared spectroscopy was adopted. The spectral data were preprocessed by standard normal transformation. The supervised Fisher Score and unsupervised Silhouette Coefficient score calculations were combined, and weight factors were added to select the optimal feature variables. A partial least squares prediction and discrimination model was established.
The generalization ability and accuracy of the model are improved, the complexity of the model is simplified, and it can distinguish the core and non-core production areas of Taiping Houkui tea with high precision, reducing the influence of spectral noise and baseline drift.
Smart Images

Figure CN117115460B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of near-infrared spectroscopy and relates to a FS-SC semi-supervised feature selection method for Taiping Houkui tea. BACKGROUND
[0002] Taiping Houkui tea belongs to green tea and is mainly produced in the Fenghuang, Shitong and Jigongjianshan areas of the Sanhecun Houkeng, Hougang and Yanjia natural villages in Huangshan District, Huangshan City, Anhui Province. Taiping Houkui tea is favored by consumers because of its unique processing technology and use of special persimmon tea as the main tea variety. Taiping Houkui tea is listed as one of the top ten famous teas in China. However, the quality of Taiping Houkui tea is closely related to the soil environment, variety and production process of tea trees. The differences in the characteristic components of tea leaves result in a large difference in market prices of Taiping Houkui tea from different producing areas. The difference in market prices leads to phenomena such as "substituting inferior for good" and "adulteration", which seriously damages the healthy and stable development of the Taiping Houkui tea market.
[0003] Traditional identification methods mainly realize tea origin and grade identification analysis through artificial experience. However, the geographical positions of the main producing areas of Taiping Houkui tea are similar, and the types of ingredients contained in the tea products from adjacent producing areas are basically the same. In addition, the appearance of tea products prepared by traditional manual methods has small differences, which makes it impossible to realize rapid and high-precision origin identification analysis based on sensory evaluation methods relying on artificial experience. At present, the origin identification analysis is mainly realized through chemical analysis of mineral element content. In the prior art, the application public number CN106560697A of the Chinese invention patent application "Joint near-infrared spectroscopy and trace element Wuyi rock tea origin identification method" proposed by Ye Zhihong realizes the identification of Wuyi rock tea origin by combining near-infrared spectroscopy and mineral element content. Although the identification method combining mineral element content analysis can distinguish tea from certain areas, the mineral element analysis belongs to chemical analysis, and the detection process is costly, time-consuming, complex and environmentally destructive. In addition, there is no unified tea mineral element content detection standard in China. For specific tea producing areas such as Taiping Houkui, the distribution of mineral elements between samples from different producing areas in similar production areas is basically the same, and it is impossible to effectively identify the origin by analyzing the mineral element content.
[0004] Near infrared spectrum is electromagnetic radiation wave between visible light and mid-infrared, and its wavelength range is 780nm to 2526nm. Near infrared spectrum has strong penetration ability, and the absorption area of near infrared spectrum is consistent with the vibration of hydrogen-containing groups in organic molecules, which can carry the characteristic information of hydrogen-containing groups in the sample, and is more suitable for the analysis of organic matter composition. In recent years, near infrared spectrum technology has been favored by various fields because of its relative accuracy, rapidity, simplicity and other advantages, and has become a new analysis and research method. One of the most common qualitative and quantitative analysis model establishment methods in near infrared spectrum analysis is partial least squares method, which belongs to factor analysis method like principal component regression. In the modeling process, the spectrum matrix needs to be decomposed, a few latent variables are extracted to represent most of the information of the original spectrum, and these variables are called principal components in partial least squares regression. However, in the principal component extraction process of partial least squares regression, not only the detection target vector is considered, but also the covariance between the extracted principal components and the detection target vector is maximized, which ensures that the principal components have the maximum correlation with the detection target vector.
[0005] Before establishing the correction model, the spectrum pretreatment method is used to correct the original near infrared spectrum data, and to weaken or even eliminate the phenomena such as baseline drift, spectral peak overlap and noise. The currently widely used near infrared spectrum processing methods mainly include standard normal transformation, multivariate scatter correction, smoothing filter, baseline correction and the like. The existing pretreatment can eliminate the noise information contained in the original spectrum data, highlight the differences between the spectrum signals of different samples, simplify the complexity of the subsequently established model, and improve the model performance and generalization ability.
[0006] Due to the high characteristic dimension of near infrared spectrum data, in the actual detection process, the labeled data is rare and the labeling processing is expensive, it is more difficult to obtain complete labeled data set, and in practice, it is often necessary to process the data set with partial label information missing. The supervised feature selection method cannot utilize all samples, so it is difficult to fully mine the characteristics of near infrared spectrum data categories only by using a small amount of labeled data. SUMMARY
[0007] The technical scheme of the present application is used to solve the problem that the existing technology cannot fully mine the key variables of characteristics due to the small amount of labeled near infrared spectrum samples, and the poor prediction performance of the correction model.
[0008] The present application solves the above technical problems by the following technical scheme:
[0009] A near infrared spectrum FS-SC semi-supervised feature selection method of Taiping Houkui tea, comprising the following steps:
[0010] S1, collect the near-infrared spectrum intensity data of Taiping Houkui tea to form an original near-infrared spectrum data set of Taiping Houkui tea containing N near-infrared spectrum data samples, each sample containing M characteristic variables;
[0011] S2, performing preprocessing on the original near-infrared spectrum data set by standard normal transformation ;
[0012] S3, randomly dividing the preprocessed near-infrared spectrum data set into labeled samples and unlabeled samples, the label being: core producing area and non-core producing area; further randomly dividing the labeled samples into a calibration set and a test set according to a proportion;
[0013] S4, performing supervised Fisher Score score calculation on the labeled samples and unsupervised Silhouette Coefficient score calculation on the unlabeled samples, adding a weight factor between the supervised Fisher Score score and the unsupervised Silhouette Coefficient score to obtain a final semi-supervised score, and sorting the semi-supervised score, the highest score being the characteristic variable to be selected;
[0014] S5, establishing a partial least squares prediction discriminant model, training the partial least squares prediction discriminant model with the characteristic variable obtained in step S4, and correcting and testing the model with the calibration set and the test set, and finally predicting the unlabeled samples.
[0015] Further, the formula of the standard normal transformation in step S2 is as follows:
[0016]
[0017] wherein, is the average spectrum, X j is the jth near-infrared spectrum data sample in the original near-infrared spectrum data set.
[0018] Further, the formula of the supervised Fisher Score score calculation on the labeled samples in step S4 is as follows:
[0019]
[0020] wherein, n + is the number of samples of the core producing area in the labeled samples, n - is the number of samples of the non-core producing area in the labeled samples; represents the ith characteristic value of the sample of the core producing area, represents the ith characteristic value of the sample of the non-core producing area, represents the average value of the i-th eigenvalue; represents the ith eigenvalue of the kth sample in the core production area sample, represents the i-th eigenvalue of the k-th sample in the non-core production area samples;
[0021] Furthermore, the formula for calculating the unsupervised Silhouette Coefficient score for the unlabeled sample described in step S4 is as follows:
[0022]
[0023] Among them, N u is the number of unlabeled samples, b j (i) represents the average shortest distance between the i-th feature of the j-th sample and the features of other samples in other clusters, a j (i) represents the average shortest distance between the i-th feature of the j-th sample and the features of other samples in the same cluster, max{a j (i),b j (i)} Select the largest value; finally, add up the values of the i-th feature of all samples and take the average as the score value of the feature.
[0024] Furthermore, a weight factor is added between the supervised Fisher Score and the unsupervised Silhouette Coefficient score in step S4 to obtain the final semi-supervised score using the following calculation formula:
[0025]
[0026] Among them, final_score(i) is the semi-supervised score.
[0027] Furthermore, the formula of the partial least squares prediction discriminant model described in step S5 is as follows:
[0028] W=TP+E (5)
[0029] V=UQ+F (6)
[0030] Where W is the independent variable matrix and V is the dependent variable matrix. Where T=[t1,t2,…,t L ] is the potential score vector, P = [p1,p2,…,p L ] and Q=[q1,q2,…,q L ] are the loading matrices of W and V respectively, E and F are the PLS residuals corresponding to W and V, and the number of L is usually determined by cross-validation, which gives the maximum prediction ability based on the data of the prediction set and ensures that there is no underfitting or overfitting.
[0031] The present application has the advantages of:
[0032] (1) The technical scheme of the present application performs supervised Fisher Score score calculation on labeled samples, unsupervised Silhouette Coefficient score calculation on unlabeled samples, adds a weight factor between the supervised Fisher Score score and the unsupervised Silhouette Coefficient score, obtains the final semi-supervised score, sorts the semi-supervised score, and the highest score is the feature variable to be selected. The feature variable is used to train a partial least squares prediction discriminant model, which simplifies the model complexity while improving the generalization ability, can improve the model precision, and improve the robustness of the model.
[0033] (2) By pre-processing the original near-infrared spectrum data set, irrelevant and redundant features are removed, the influence of near-infrared spectrum noise and baseline drift is weakened, and the influence of spectrum disturbance on model establishment is avoided. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 Flow chart of FS-SC semi-supervised feature selection method for Taiping Houkui tea near-infrared spectrum;
[0035] Figure 2 Raw spectrum data for Taiping Houkui tea near-infrared spectrum analysis;
[0036] Figure 3 Spectrum data after SNV preprocessing;
[0037] Figure 4 Results of model prediction based on FS-SC and PLS;
[0038] Figure 5 Results of model prediction based on FS and PLS. DETAILED DESCRIPTION
[0039] To make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0040] The technical solutions of the present application will be further described below in conjunction with the drawings in the specification and specific embodiments:
[0041] Embodiment one
[0042] As shown in Figure 1 .
[0043] To achieve the above purpose, the present application is realized by the following technical solutions:
[0044] A near-infrared spectrum FS-SC semi-supervised feature selection method of Taiping Houkui tea, comprising the following steps:
[0045] 1. Collecting near-infrared spectrum absorbance data of Taiping Houkui tea to form an original near-infrared spectrum data set of Taiping Houkui tea Containing N near-infrared spectrum data samples, wherein the number of labeled samples is N l , the number of unlabeled samples is N u , and each sample contains M column feature variables.
[0046] 2. Preprocessing the original near-infrared spectrum data set by using standard normal transformation (Standard Normal Variate Transform, SNV); the formula of the standard normal transformation is as follows:
[0047]
[0048] Wherein, is the average spectrum, X j is the jth near-infrared spectrum data sample in the original near-infrared spectrum data set.
[0049] 3. Randomly dividing the preprocessed near-infrared spectrum data set into 20% labeled samples and 80% unlabeled samples, and the label is: core production area and non-core production area; The labeled samples are randomly divided into a calibration set and a test set according to a 7:3 ratio.
[0050] 4. Supervised Fisher Score calculation is performed on the labeled samples, and unsupervised Silhouette Coefficient score calculation is performed on the unlabeled samples, a weight factor is added between the supervised Fisher Score and the unsupervised Silhouette Coefficient score, and the final semi-supervised score is obtained. The highest score of the semi-supervised score is the selected feature variable.
[0051] (1) The supervised Fisher Score calculation of the labeled samples is as follows:
[0052]
[0053] Wherein, n + n is the number of samples in the core region in the labeled samples - n is the number of samples in the non-core region in the labeled samples is the i-th feature value representing the samples in the core region, is the i-th feature value representing the samples in the non-core region, is the average value of the i-th feature value; is the i-th feature value representing the k-th sample in the core region, is the i-th feature value representing the k-th sample in the non-core region,
[0054] (2) The unsupervised Silhouette Coefficient score calculation for the unlabeled samples is as follows:
[0055]
[0056] where b j (i) represents the average shortest distance between the i-th feature of the j-th sample and the features of other samples in other clusters, a j (i) represents the average shortest distance between the i-th feature of the j-th sample and the features of other samples in the same cluster, max{a j (i), b j (i)} is selected as the maximum value; finally, the values of the i-th feature of all samples are accumulated and averaged to obtain the score value of the feature.
[0057] (3) A weight factor is added between the supervised Fisher Score and the unsupervised Silhouette Coefficient score to obtain the final semi-supervised score calculation formula as follows:
[0058]
[0059] where final_score is the semi-supervised score.
[0060] 5. Establish a partial least squares discriminant analysis (PLS-DA) model, use the feature variables obtained in step 4 to train the partial least squares discriminant analysis model, then use the calibration set and the test set to calibrate and test the model, and finally use the model to predict the unlabeled samples.
[0061] The formula of the partial least squares discriminant analysis model is as follows:
[0062] W = TP + E (5)
[0063] V = UQ + F (6)
[0064] where W is the independent variable matrix, and V is the dependent variable matrix. Where T = [t1, t2, …, tn] is the score matrix of the calibration set, P is the loading matrix, E is the residual matrix, U = [u1, u2, …, un] is the score matrix of the test set, Q is the loading matrix, and F is the residual matrix.L ] is the potential score vector, P = [p1,p2,…,p L ] and Q=[q1,q2,…,q L ] are the loading matrices of W and V respectively, E and F are the PLS residuals corresponding to W and V. Under the condition of ensuring no underfitting or overfitting, the number of L is usually determined by cross-validation, which gives the maximum prediction ability based on the data of the prediction set.
[0065] 6. Test verification
[0066] The spectrum data of Taiping Houkui was collected as follows: like Figure 2 As shown in the figure, the Taiping Houkui production area is divided into core production area and non-core production area, 20% of the samples are used as labeled data sets, and the remaining 80% are used as unlabeled data sets. Then the labeled data sets are randomly divided into calibration set and prediction set according to the ratio of 7:3; the original near-infrared spectral data are preprocessed using standard normal transformation, and the near-infrared spectrum after preprocessing is shown in the figure below. Figure 3 As shown in the figure, the FS-SC semi-supervised feature selection method is used to calculate the scores of the preprocessed spectral data, and the calculated scores are arranged in descending order to finally obtain the optimal number of feature variables; based on the finally selected spectral feature variables, a discriminant analysis model is established using PLS to analyze the model performance. Under the condition that the optimal PLS principal component is 5, the results of 100 Monte Carlo simulation experiments are shown in the figure. Figure 4 As shown, the coefficient of determination R of the prediction set 2 The median is 0.8333, and the coefficient of determination R 2 The median is 0.8438; the PLS model under the FisherScore (FS) supervised feature selection method is used as a comparison, and the results of 100 Monte Carlo simulation experiments are as follows Figure 5 As shown, under the condition that the optimal PLS principal component is 5, the determination coefficient R of the prediction set 2 The median is 0.75; by comparison, it can be seen that the method of the present invention can achieve high-precision classification and identification of Taiping Houkui core production areas and non-core production areas through a small amount of labeled near-infrared spectral data and a large amount of unlabeled near-infrared spectral data.
[0067] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A semi-supervised feature selection method for near-infrared spectroscopy of Taiping Houkui tea, characterized in that: The steps include: S1. Collect the near-infrared spectral intensity data of Taiping Houkui tea to form the original near-infrared spectral data set of Taiping Houkui tea , which contains NIR spectral data samples, each sample contains Column feature variables; S2. Standard normal transformation of the original near-infrared spectral dataset Perform pretreatment; S3. Randomly divide the preprocessed near-infrared spectral dataset into labeled samples and unlabeled samples, with the labels being: core production area and non-core production area; the labeled samples are further randomly divided into a calibration set and a test set according to the proportion; S4. Perform supervised Fisher Score calculation on labeled samples and unsupervised Silhouette Coefficient calculation on unlabeled samples. Add a weight factor between the supervised Fisher Score and the unsupervised Silhouette Coefficient to obtain the final semi-supervised score. Sort the semi-supervised scores, and the one with the highest score is the feature variable to be selected. The calculation formula for adding a weight factor between the supervised Fisher Score and the unsupervised Silhouette Coefficient to obtain the final semi-supervised score is as follows: (4) in, final_score ( i ) is the semi-supervised score, is the number of unlabeled samples, is the number of samples from the core production areas among the labeled samples, is the number of samples from non-core production areas among the labeled samples; represents the ith eigenvalue of the core production area sample, The first sample represents non-core production areas. i eigenvalues, Representative i The average value of the eigenvalues; Representative core production area samples k The first sample i eigenvalues, Representative non-core production area samples k The first sample i eigenvalues; Representative The first sample The average shortest distance between a feature and other samples in other clusters, Representative The first sample The average shortest distance between a feature and the features of other samples in the same cluster, Indicates selection and The largest value among S5. Establish a partial least squares prediction and discrimination model, use the feature variables obtained in step S4 to train the partial least squares prediction and discrimination model, then use the calibration set and test set to calibrate and test the model, and finally use it to predict unlabeled samples.
2. The semi-supervised feature selection method for near-infrared spectroscopy FS-SC of Taiping Houkui tea according to claim 1, characterized in that, The formula for the standard normal transformation described in step S2 is as follows: (1) in, is the average spectrum, , The first Near infrared spectral data samples.
3. The semi-supervised feature selection method for near-infrared spectrum of Taiping Houkui tea FS-SC according to claim 2 is characterized in that, The formula for calculating the supervised Fisher Score for labeled samples described in step S4 is as follows: (2) in, is the number of samples from the core production areas among the labeled samples, is the number of samples from non-core production areas among the labeled samples; represents the ith eigenvalue of the core production area sample, represents the i-th eigenvalue of the non-core production area sample, represents the average value of the i-th eigenvalue; represents the ith eigenvalue of the kth sample in the core production area sample, Represents the i-th eigenvalue of the k-th sample in the non-core production area samples.
4. The semi-supervised feature selection method for near-infrared spectrum of Taiping Houkui tea FS-SC according to claim 3 is characterized in that, The formula for calculating the unsupervised Silhouette Coefficient score for the unlabeled sample described in step S4 is as follows: (3) in, Representative The first sample The average shortest distance between a feature and other samples in other clusters, Representative The first sample The average shortest distance between a feature and the features of other samples in the same cluster, Indicates selection and The largest value among all the samples; finally, The values of the features are accumulated and the average is taken as the score of the feature.
5. The semi-supervised feature selection method for near-infrared spectrum of Taiping Houkui tea according to claim 4, characterized in that, The formula of the partial least squares prediction discriminant model described in step S5 is as follows: (5) (6) Among them, W is the independent variable matrix, V is the dependent variable matrix; among them, is the potential score vector, and are the loading matrices of W and V respectively, E and F are the PLS residuals corresponding to W and V, and the number of L is usually determined by cross-validation, which gives the maximum predictive ability based on the data of the prediction set and ensures that there is no underfitting or overfitting.
Citation Information
Patent Citations
Method for identifying producing area of Wuyi rock tea through combination of near infrared spectroscopy and trace element detection
CN106560697A
Hyperspectral image classification and wave band selection method based on multi-target immune cloning
CN103914705A
Tea classification method based on near infrared spectrum feature selection and parameter optimization
CN114720419A