A spectral sample screening method based on a multi-sensor portable spectrometer
By combining spectral data and multidimensional models to screen samples using a two-level screening method, the problem of excessive elimination of marginal samples in multi-sensor portable spectrometers was solved, and the stability and predictive ability of the model were improved.
Patent Information
- Application Number
- CN202311231124.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-22
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2043-09-22
AI Technical Summary
In multi-sensor portable spectrometers, existing technologies fail to effectively consider the positions and combination information of different sensors, resulting in excessive elimination of edge samples, affecting the stability and generalization ability of the spectral model.
A two-stage screening method is adopted. The initial screening screens samples by combining spectral data and sample calibration value information. The final screening uses a multi-dimensional model to screen out abnormal points, and uses spectral characteristics and model predictions to determine whether the sample is an abnormal point.
The stability and generalization ability of the spectral model are improved, marginal samples are retained, and the predictive ability and reliability of the model are enhanced.
Smart Images

Figure CN117292767B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of spectral screening, and more particularly to a spectral sample screening method based on a multi-sensor portable spectrometer. Background Art
[0002] In recent years, near-infrared spectroscopy has developed rapidly, finding applications in a wide range of fields, including chemical engineering, pharmaceuticals, military, and food. Near-infrared spectroscopy, a molecular spectroscopy technique, can reveal the composition and properties of substances at the molecular level. It has achieved significant economic and social benefits and holds great potential for development.
[0003] In near-infrared spectroscopy analysis technology, the multivariate corrected regression modeling method has been applied in various quantitative analysis fields. In the process of multivariate corrected regression modeling, the performance of the constructed model depends to a large extent on the training set used. Therefore, it is extremely important to select representative samples from a large number of samples to construct a high-quality training set that is conducive to improving model performance.
[0004] The most commonly used calibration set sample screening method at present is the deviation value screening method. The internal cross-prediction of the sample set is performed by leaving one out method, and samples with large deviations between the predicted values and the actual physical and chemical values are eliminated, thereby reducing the root mean square error of the spectral model. The root mean square error (RMSECV) is the most important measurement indicator of the performance of the partial least squares model. The smaller its value, the better the performance of the spectral model.
[0005] However, when applied to the modeling of multi-sensor portable spectrometers, the position and combination information of different sensors are not taken into account. If the deviation value screening method is completely used to select the correction set samples, the edge samples are likely to be eliminated. Since the edge sample points have a greater impact on the generalization ability of the spectral model, eliminating too many edge sample points will seriously affect the stability and generalization ability of the spectral model, thereby reducing the predictive ability of the spectral model. Summary of the Invention
[0006] The purpose of the present invention is to overcome the shortcomings of traditional spectral sample screening methods when applied to multi-sensor portable spectrometers, and to provide a spectral sample screening method based on a multi-sensor portable spectrometer. The method is a two-stage screening method. The initial screening preliminarily selects the samples to be screened out by combining spectral data and sample calibration value information.
[0007] In order to achieve the above object, the present invention adopts the following technical solutions:
[0008] A spectral sample screening method based on a multi-sensor portable spectrometer, comprising:
[0009] Import M samples as the original sample set, set the initial screening ratio, calculate the number of samples for initial screening, and then set the number of initial screening samples C based on this number;
[0010] Calculate the average spectral value and average physicochemical value of the original sample set respectively, and calculate the SPXY value of each sample in the training set for the average spectral value and average physicochemical value according to the SPXY formula;
[0011] Sort the original sample set according to the SPXY value, and select C samples with high SPXY values to form the initial screening sample set;
[0012] Set the model building scheme, which includes different sensor combinations, different preprocessing methods and different modeling algorithms, and set the total number of model building schemes;
[0013] The model is established using the set modeling scheme and cross-validated on the samples in the primary screening sample set to obtain a set of cross-validated prediction values for each sample in the primary screening sample set under different model establishment schemes;
[0014] Perform a t-test on the cross-validation prediction value set of each sample and the corresponding calibration value, and calculate the t-test statistic;
[0015] Set the significance level factor and obtain the threshold of the t-test based on the number of samples in the initial screening set;
[0016] The t-test statistic and threshold of each sample are compared in turn to determine the degree of similarity between the cross-validation predicted value and the calibration value of the sample, and outliers are screened out to obtain the final training set for modeling.
[0017] In some embodiments, the number of samples included in the original sample set is at least greater than 220. Different retention ratios are set for different numbers of original sample sets. If the number of samples to be screened out is N, then in different cases, the number of samples to be screened out is:
[0018]
[0019] Among them, floor() is the symbol for rounding down, and ceiling() is the symbol for rounding up.
[0020] In some embodiments, when calculating the SPXY value of a single sample in the original sample set, the average spectral value and the average physicochemical value of the original sample set are used as reference values.
[0021] In some embodiments, the number of spectral wavelength points of each spectrum of the M samples included therein is set to N, and the single spectral data are x i =(x i1 , x i2 ,……x iN), where i∈[1, M], and their physical and chemical values are Y=(y1, y2, ... y M ), calculate the average spectrum X of the original sample set mean and average physicochemical value y mean They are:
[0022]
[0023] The SPXY value of the average spectral value and the average physicochemical value of a single sample in the original sample set is calculated according to the SPXY formula, where the SPXY formula is:
[0024]
[0025]
[0026]
[0027] Among them, X i Represents the spectral value of a single sample, Y i Indicates the calibration value of a single sample, SPXY i Represents the SPXY value of a single sample, M represents the number of samples in the original sample set, N represents the number of spectral wavelength points of the sample, and max is the maximum value symbol; the SPXY value of the average spectral value and average physicochemical value of each sample in the original sample set is calculated.
[0028] In some embodiments, performing a t-test on the cross-validation prediction value set of each sample and the corresponding calibration value to calculate the t-test statistic includes:
[0029] The null hypothesis of the t-test is that the average of the multi-model cross-validation prediction values of a sample is equal to the calibration value of the sample itself. The degree of similarity between the cross-validation prediction value and the calibration value is determined by whether the null hypothesis is rejected, and the t-test statistic is used to characterize the above degree of similarity.
[0030] In some embodiments, the single-sample t-test statistic is represented by the following formula:
[0031]
[0032]
[0033]
[0034] in is the average value of the single sample cross-validation prediction value set, s a is the variance of the cross-sample prediction value set of a single sample, T aC is the number of samples in the initial screening sample set, and k is the number of cross-validation prediction values for a single sample.
[0035] In some embodiments, the outlier screening includes: when the t-test statistic is greater than a threshold, it is considered that the cross-validation prediction value of the sample is far from the calibration value, and the sample is selected into the final screening sample set and screened out as an outlier.
[0036] Compared with the prior art, the present invention has the following beneficial effects:
[0037] A spectral sample screening method based on a multi-sensor portable spectrometer is provided. The method includes two levels of screening. The initial screening preliminarily selects samples to be screened out by combining spectral data and sample calibration value information. The final screening identifies samples that are far from the average spectral calibration value. The final screening removes abnormal samples by using a multidimensional model to predict the initial screening sample set. This method uses the sample's own spectral characteristics and the multidimensional model's prediction of the sample to determine whether the sample should be screened out as an outlier. This not only ensures a higher reliability of the training set after outliers are screened out, but also maximizes the retention of marginal samples, effectively improving the generalization ability and stability of the spectral model. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 It is a schematic diagram of a spectral sample screening method based on a multi-sensor portable spectrometer of the present invention;
[0039] Figure 2 This is a schematic diagram of the sensor combination of the multi-sensor portable spectrometer described in the present invention;
[0040] Figure 3 It is a schematic diagram of a characteristic setting model establishment scheme of the multi-sensor portable spectrometer in the present invention. DETAILED DESCRIPTION
[0041] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the embodiments of the present application will be described in more detail below in conjunction with the drawings in the preferred embodiments of the present application. In the drawings, the same or similar reference numerals throughout represent the same or similar parts or parts with the same or similar functions. The described embodiments are part of the embodiments of the present application, not all of the embodiments. The embodiments described below with reference to the drawings are exemplary and are intended to be used to explain the present application, and should not be understood as limitations on the present application. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0042] The embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0043] The following will be combined Figure 1-3 , a spectral sample screening method based on a multi-sensor portable spectrometer involved in an embodiment of the present application is described in detail. The process is to first set a preliminary screening ratio based on the number of samples in the original sample set, and then set the preliminary screening number based on this ratio. The SPXY value of each sample in the original sample set relative to the average spectral value and the average calibration value is calculated according to the spectral-calibration value co-occurrence distance (SPXY) formula. The original sample set is then screened based on the SPXY value to select samples that meet the preliminary screening number, forming a preliminary screening sample set for final screening.
[0044] The final screening uses a multi-dimensional model to predict the initial screening sample set to find abnormal samples. The process is to first set the model building scheme according to the characteristics of the multi-sensor portable spectrometer, and use the model building schemes with different sensor combinations, different preprocessing methods and different modeling algorithms to cross-validate the initial screening sample set to obtain the cross-validation prediction value of different models for each sample in the initial screening sample set. Then, a t-test is performed on the cross-validation prediction value and calibration value of each sample, and the t-test statistic is calculated. By setting the significance level factor, the degree of similarity between the cross-validation prediction value and the calibration value is judged. The samples that are closer to each other constitute the final screening sample set for outlier screening, and finally constitute the training set for modeling.
[0045] This method combines the training set sample's own information with the model prediction information, ensuring the reliability of the samples selected for final modeling, while retaining marginal samples to the greatest extent, effectively improving the generalization ability and stability of the spectral model.
[0046] Embodiment 1:
[0047] See also Figure 1 As shown, a spectral sample screening method based on a multi-sensor portable spectrometer includes the following steps:
[0048] Step 101: Import M samples as the original sample set, set the initial screening ratio, calculate the number of samples for initial screening, and then set the number of samples for initial screening C based on the number;
[0049] Step 102: Calculate the average spectral value and average physicochemical value of the training set samples respectively, and calculate the SPXY value of each sample in the training set for the average spectral value and average physicochemical value according to the SPXY formula;
[0050] Step 103: Sort the training sample set according to the SPXY value, and select C samples with high SPXY values to form a preliminary screening sample set;
[0051] Step 104: setting a model building scheme based on the characteristics of the multi-sensor portable spectrometer, wherein the scheme includes different sensor combinations, different preprocessing methods, and different modeling algorithms, and setting the total number of model building schemes;
[0052] Step 105: Establish a model using the set modeling scheme and perform cross-validation on samples in the initial screening sample set to obtain a set of cross-validated prediction values for a single sample under the multidimensional model;
[0053] Step 106: Perform a t-test on the cross-validation prediction value set of a single sample and the corresponding calibration value, and calculate the t-test statistic;
[0054] Step 107: Set the significance level factor and obtain the threshold value of the t-test based on the number of samples in the initial screening set by looking up the table;
[0055] Step 108: Compare the t-test statistic and the threshold of each sample in turn to determine the degree of similarity between the cross-validation predicted value and the calibration value of the sample. When the t-test statistic is greater than the threshold, it is considered that the cross-validation predicted value of the sample is far from the calibration value, and it is selected into the final screening sample set and screened out as an outlier to obtain the final training set for modeling.
[0056] As a further optimization, in step a, M samples are imported as the original spectral sample set, a retention screening ratio is set, and based on the retention screening ratio, the number of samples to be screened out is calculated, and then the number of primary screening samples is set based on this number. In order to ensure the performance of the spectral model, different primary screening ratios are set for different original sample sets. When the original sample set contains a small number of samples, a higher primary screening ratio is set to ensure that the calibration set samples have as large a coverage range as possible, thereby ensuring the stability and universality of the spectral model; when the original sample set contains a large number of samples, a lower primary screening ratio is set to ensure that as many abnormal samples as possible are eliminated, thereby improving the predictive ability of the spectral model.
[0057] Figure 1 In step 101, M samples are imported as the original sample set, the initial screening ratio is set, the number of samples for the initial screening is calculated, and then the number of initial screening samples C is set based on this number. To ensure the performance of the spectral model, the initial screening ratio is different for different numbers of original sample sets. When the original sample set contains a small number of samples, a lower initial screening ratio is set to ensure that the remaining training set samples have the largest possible coverage after screening, thereby ensuring the stability and universality of the spectral model; when the original sample set contains a large number of samples, a higher initial screening ratio is set to ensure that as many abnormal samples as possible are eliminated, thereby improving the predictive ability of the spectral model.
[0058] In the present embodiment, in the actual application of near-infrared spectral analysis, too few samples will make the spectral model highly random and unstable, and it will be very easy to overfit, which will seriously affect the predictive ability of the spectral model; too many samples will have a weaker ability to improve the effect of the spectral model, and collecting too many samples will cause a lot of waste of manpower and material resources. When the number of training set samples reaches about 200, the spectral model established by the training set has preliminary stability and predictive ability; when the number of training set samples reaches about 400, the spectral model established by the training set has good stability and predictive ability; when the number of training set samples reaches 800, continuing to increase the number of samples will have a weaker ability to improve the spectral model. If the initial screening ratio is set to 10% when the number of samples is small; when the number of samples is large, the initial screening ratio is 20%. It can be seen that the number of samples contained in the training set sample set is at least greater than 220. Different retention ratios are set for different numbers of original sample sets. If the number of samples to be screened is C, then in different cases, the number of samples to be screened is:
[0059]
[0060] Among them, floor() is the symbol for rounding down, and ceiling() is the symbol for rounding up.
[0061] In this embodiment, the currently common sample selection principles include those based on spectral distance and those based on spectral-physical and chemical value co-occurrence distance. The goal of the selection principle based on spectral distance is to calculate the Euclidean distance of the sample data from the uniform spectrum, and according to the set screening ratio, the samples with larger Euclidean distance values are screened out as abnormal points in proportion. The defect is that only the characteristics of the spectrum are considered, and the influence of the component physicochemical value Y of the sample is not considered. Therefore, there are limitations and irrationalities, and it is difficult to obtain a model with stable performance and strong applicability. The biggest difference between the SPXY division method based on the spectral-physical and chemical value co-occurrence distance selection principle and the sample screening method based on the spectral distance selection principle is that it fully considers both the spectral information and the influence of the component physicochemical values. This type of method improves the distance selection criterion for data set division, takes into account the factors of component physicochemical values, and can effectively improve the stability and applicability of the spectral model.
[0062] The SPXY formula is used to calculate the SPXY value of each sample in the original sample set for the average spectral value and average physicochemical value. When calculating the SPXY value of a single sample in the original sample set, the average spectral value and average physicochemical value of the original sample set are used as the benchmark value. This can not only best reflect the distribution range of the spectral data in the original sample set, but also minimize the error caused by randomly selecting a single sample data as the benchmark value, effectively improving the accuracy of the SPXY value of each sample in the original sample set.
[0063] For the original sample set of this patent, it is assumed that there are M samples contained therein, the number of spectral wavelength points of each spectrum is N, and the single spectral data are X i =(x i1 , x i2 ,……x iN ), where i∈[1, M], and their physical and chemical values are Y=(y1, y2, ... y M ), calculate the average spectrum X of the original sample set mean and average physicochemical value y mean They are:
[0064]
[0065] The SPXY value of the average spectral value and average physicochemical value of a single sample in the original sample set is further calculated according to the SPXY formula, where the SPXY formula is:
[0066]
[0067]
[0068]
[0069] Among them, X i Represents the spectral value of a single sample, Y i Indicates the calibration value of a single sample, SPXY i represents the SPXY value of a single sample, M represents the number of samples in the original sample set, N represents the number of spectral wavelength points of the sample, and max is the symbol for taking the maximum value. The above formula can be used to calculate the SPXY value of each sample in the original sample set for the average spectral value and average physicochemical value.
[0070] As a further optimization, C samples with high SPXY values are selected in step 103 to form the initial screening sample set. The SPXY values fully consider both spectral information and the influence of component physicochemical values. These two variables are used to simultaneously calculate inter-sample distances to ensure maximum characterization of sample distribution, effectively covering the multidimensional vector space, and retaining marginal samples to the greatest extent possible, thereby increasing sample diversity and representativeness.
[0071] Step 104 describes a modeling scheme based on the characteristics of the multi-sensor portable spectrometer. The scheme includes different sensor spectral band combinations, different preprocessing methods, and different modeling algorithms. The different sensor combinations refer to the different spectral bands and sensor spectral combinations at different optical path design positions included in the multi-sensor portable spectrometer. These combinations can be configured based on the actual application scenario.
[0072] Figure 1104 is a model establishment scheme set according to the characteristics of the multi-sensor portable spectrometer, which includes different sensor combinations, different preprocessing methods and different modeling algorithms. The total number of model establishment schemes set is k. The specific model establishment scheme is set according to the characteristics of the multi-sensor portable spectrometer, which includes different sensor spectral band combinations, different preprocessing methods and different modeling algorithms. The different sensor combinations refer to the different spectral bands contained in the multi-sensor portable spectrometer and the sensor spectral combinations at different optical path design positions. The combination can be set according to the actual application scenario needs. Figure 2 Design schematic diagram for different band sensor positions of multi-sensor portable spectrometer, according to Figure 2 The number of each sensor in the figure sets the spectral band combination of different sensors, where 201 is the arrangement position of the four sensors. Figure 3 This is a schematic diagram of the model building scheme, which lists the sub-option combinations of different model building schemes.
[0073] As a further optimization, the model described in step 105 is established using the set modeling scheme and cross-validated on the samples in the primary screening sample set. The specific method is the leave-one-out method, that is, each sample in the primary screening sample set is extracted from the original sample set in turn, and a model is constructed for other samples in the original sample set using the set modeling scheme. Then, the constructed model is used to predict the currently extracted sample to obtain a set of predicted values of multiple models for the sample. The above method is repeated until a set of predicted values for all samples in the primary screening sample set is obtained.
[0074] Figure 1 In step 105, a model is established using a set modeling scheme and cross-validated on samples in the primary screening sample set to obtain a set of cross-validated prediction values for a single sample in the primary screening sample set under different model establishment schemes, wherein the cross-validation value set of a single sample contains k data, and the specific method used for each prediction value is the leave-one-out method, selecting all samples except the current sample as training sets, and constructing a model using the set modeling scheme. Then, the constructed model is used to predict the current sample to obtain a set of prediction values for the k models of the sample, and the above method is used until a set of prediction values Q for all samples in the primary screening sample set is obtained.
[0075] As a further optimization, the t-test described in step f has a null hypothesis that the average of the multi-model cross-validation prediction values of a sample is equal to the calibration value of the sample itself. The degree of similarity between the cross-validation prediction value and the calibration value is determined by whether the null hypothesis is rejected.
[0076] Figure 1In step 106, a t-test is performed on the cross-validation prediction value set of each sample and the corresponding calibration value, and the t-test statistic is calculated. The null hypothesis of the t-test is that the average of the multi-model cross-validation prediction values of a sample is equal to the calibration value of the sample itself. The degree of similarity between the cross-validation prediction value and the calibration value is determined by whether the null hypothesis is rejected. The t-test statistic is used to represent the above-mentioned similarity. In this example, the t-test statistic for a single sample is represented by the following formula:
[0077]
[0078]
[0079]
[0080] in is the average value of the single sample cross-validation prediction value set, s a is the variance of the cross-sample prediction value set of a single sample, T a C is the number of samples in the initial screening sample set, and k is the number of cross-validation prediction values for a single sample.
[0081] Figure 1 In the example, 107 is the setting of the significance level α, and according to the number of samples in the initial screening set, the threshold T of the t-test is obtained by looking up the table. The significance level α is generally set to 0.05 or 0.01. In this example, it is set to 0.05. Since the t-test statistic obeys the t-distribution, and the number of single sample cross-validation prediction values is known to be k, the threshold R is obtained by looking up the table to test the above null hypothesis.
[0082] Figure 1 In step 108, the t-test statistic and the threshold of each sample are compared in turn to determine the degree of similarity between the cross-validation predicted value and the calibration value of the sample. When the t-test statistic is greater than the threshold T, it is considered that the cross-validation predicted value of the sample is far from the calibration value, and it is selected into the final screening sample set and screened out as an outlier to obtain the final training set for modeling.
[0083] The above description is only a preferred embodiment of the present invention and is used to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A spectral sample screening method based on a multi-sensor portable spectrometer, characterized in that: include: Import M samples as the original sample set, set the initial screening ratio, calculate the number of samples for initial screening, and then set the number of initial screening samples C based on this number; Calculate the average spectral value and average physicochemical value of the original sample set respectively, and calculate the SPXY value of each sample in the training set for the average spectral value and average physicochemical value according to the SPXY formula; The SPXY formula is: ; ; ; in, represents the spectral value of a single sample, represents the calibration value of a single sample, Represents the SPXY value of a single sample, M represents the number of samples in the original sample set, Indicates the number of spectral wavelength points of the sample, To take the maximum value symbol; calculate the SPXY value of the average spectral value and average physicochemical value of each sample in the original sample set; Average spectrum; Average physicochemical values; Sort the original sample set according to the SPXY value, and select C samples with high SPXY values to form the initial screening sample set; Set the model building scheme, which includes different sensor combinations, different preprocessing methods and different modeling algorithms, and set the total number of model building schemes; The model is established using the set modeling scheme and cross-validated on the samples in the primary screening sample set to obtain a set of cross-validated prediction values for each sample in the primary screening sample set under different model establishment schemes; Perform a t-test on the cross-validation prediction value set of each sample and the corresponding calibration value, and calculate the t-test statistic; Set the significance level factor and obtain the threshold of the t-test based on the number of samples in the initial screening set; The t-test statistic and threshold of each sample are compared in turn to determine the degree of similarity between the cross-validation predicted value and the calibration value of the sample, and outliers are screened out to obtain the final training set for modeling.
2. The spectral sample screening method based on a multi-sensor portable spectrometer according to claim 1, characterized in that: The number of samples contained in the original sample set is at least greater than 220. Different retention ratios are set for different numbers of original sample sets. If the number of samples to be screened out is N, then in different cases, the number of samples to be screened out is: ; Among them, floor() is the symbol for rounding down, and ceiling() is the symbol for rounding up.
3. The spectral sample screening method based on a multi-sensor portable spectrometer according to claim 1, characterized in that: When calculating the SPXY value of a single sample in the original sample set, the average spectral value and average physicochemical value of the original sample set are used as the benchmark value.
4. The spectral sample screening method based on a multi-sensor portable spectrometer according to claim 3, characterized in that: Assume that the number of spectral wavelength points of each spectrum of the M samples contained therein is N, and the single spectrum data are ,in , and their physical and chemical values are , calculate the average spectrum of the original sample set and average physicochemical values They are: 。 5. The spectral sample screening method based on a multi-sensor portable spectrometer according to claim 1, characterized in that: The method of performing a t-test on the cross-validation prediction value set of each sample and the corresponding calibration value and calculating the t-test statistic comprises: The null hypothesis of the t-test is that the average of the multi-model cross-validation prediction values of a sample is equal to the calibration value of the sample itself. The degree of similarity between the cross-validation prediction value and the calibration value is determined by whether the null hypothesis is rejected, and the t-test statistic is used to characterize the above degree of similarity.
6. The spectral sample screening method based on a multi-sensor portable spectrometer according to claim 5, characterized in that: The t-test statistic for a single sample is expressed as follows: ; ; ; in is the average of the single sample cross-validation prediction value set, is the variance of the set of cross-sample predictions for a single sample, C is the number of samples in the initial screening sample set, and k is the number of cross-validation prediction values for a single sample.
7. The spectral sample screening method based on a multi-sensor portable spectrometer according to claim 1, characterized in that: The outlier screening includes: when the t-test statistic is greater than a threshold, it is considered that the cross-validation prediction value of the sample is far from the calibration value, and the sample is selected into the final screening sample set and screened out as an outlier.
Citation Information
Patent Citations
Multi-sensor spectral data processing method based on multi-component sample
CN114609081A
Traditional Chinese medicine component analysis method and system based on spectral clustering
CN114783539A