Portable near infrared quantitative model and unknown sample matching degree discrimination method

By constructing a PLS quantitative model for a portable near-infrared spectrometer and using a multivariate t-test method, the matching degree of unknown samples was determined, solving the problems of stability and accuracy of spectrometer spectral data, and achieving high-accuracy prediction of unknown samples and adaptive updating of the model.

CN116166973BActive Publication Date: 2026-03-31四川启睿克科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-01
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Portable near-infrared spectrometers are susceptible to the influence of light sources, detectors, detection methods, and environmental conditions, resulting in poor stability and low accuracy of spectral data, which affects the prediction accuracy of unknown samples, and the predicted values ​​are abnormal when the spectrum is abnormal.

Method used

By constructing a partial least squares (PLS) quantitative model, the score thresholds of principal component I and principal component II are calculated using the multivariate t-test method. An elliptic polar coordinate expression is constructed to determine the matching degree between unknown samples and the quantitative model, identify abnormal samples and prompt for re-acquisition of spectra, and update the model when necessary.

Benefits of technology

It improves the prediction accuracy of unknown samples, ensures the stability and accuracy of the spectrometer, and reduces the impact of human operation and environmental changes on prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116166973B_ABST
    Figure CN116166973B_ABST
Patent Text Reader

Abstract

The application discloses a portable near-infrared spectrum quantitative model and unknown sample matching degree discrimination method, constructs a quantitative model by using a partial least square method, obtains a load matrix and a score matrix of PLS principal components, respectively calculates score thresholds of the first principal component and the second principal component, constructs an elliptical polar coordinate expression by using the two thresholds, maps collected unknown sample spectrum to the principal components to obtain score values, substitutes the score value of the first principal component into the elliptical polar coordinate expression to obtain two predicted score values of the second principal component to form a judgment interval, detects whether the score value of the second principal component of the unknown sample is in the judgment interval, and thus the matching degree of the unknown sample and the quantitative model is judged. Through the application, the matching degree of the unknown sample in a prediction process and the quantitative model is judged, abnormal samples are discriminated, the model can be updated in time, and the stability and accuracy of the spectrometer during long-time use are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of near-infrared anomalous spectral sample discrimination technology, and in particular to a method for judging the matching degree between a portable near-infrared spectral quantitative model and unknown samples. Background Technology

[0002] Near-infrared spectroscopy is widely used for the detection of material components, characterized by its speed, non-destructive nature, high efficiency, and low cost. Portable near-infrared spectrometers refer to near-infrared analytical instruments that can be handheld or carried. These instruments are lightweight, portable, low-power, fast-detecting, and inexpensive, and are widely used in agriculture, medicine, geology, and environmental fields.

[0003] However, portable near-infrared spectrometers are susceptible to the influence of light sources, detectors, detection methods, and environmental conditions, resulting in poor stability and low accuracy of the acquired spectral data, which in turn affects their spectral prediction and analysis capabilities. In practical applications, due to the influence of spectrometer inherent factors and sample variations, the spectral data acquired by portable near-infrared spectrometers is prone to poor matching with the model, thus affecting the spectrometer's prediction accuracy for unknown samples. Furthermore, when using portable spectrometers to acquire spectral samples, spectral anomalies may occur during the acquisition of unknown samples due to factors such as manual operation, sample condition, and instrument condition. Predicting unknown samples with spectral anomalies will lead to abnormal predicted values. Summary of the Invention

[0004] The purpose of this invention is to provide a portable near-infrared spectroscopy quantitative model matching degree discrimination method for judging the matching degree between the quantitative model and the unknown sample, thereby improving the accuracy of predicting unknown samples using the quantitative model.

[0005] The present invention solves the above problems through the following technical solution:

[0006] A method for determining the matching degree between a portable near-infrared spectroscopy quantitative model and an unknown sample includes the following steps:

[0007] Step a. Based on the spectral data of the training set, construct a quantitative model using partial least squares (PLS) to obtain the loading matrix X_loadings of each variable in principal component one and principal component two of PLS ​​and the score matrix Xt_scores of each training set sample;

[0008] Step b. Calculate the score thresholds for principal component one and principal component two using the multivariate t-test method, and substitute the two thresholds as the major and minor radii into the elliptical polar coordinate formula to construct the elliptical polar coordinate expression;

[0009] Step c. Using the loading matrix X_loadings obtained in step a, map the collected unknown sample spectrum to principal component one and principal component two to obtain the score values ​​of principal component one and principal component two corresponding to the unknown sample.

[0010] Step d. Substitute the score of principal component one into the elliptical polar coordinate expression constructed in step b to obtain the predicted scores of principal component two with the same value but opposite signs. The two predicted scores of principal component two are used to form the judgment interval.

[0011] Step e. Check whether the score of principal component 2 of the unknown sample is within the judgment interval, so as to determine the matching degree between the unknown sample and the quantitative model.

[0012] As a further improvement of the present invention, step e, determining the matching degree between the unknown sample and the quantitative model includes:

[0013] (1) When the principal component score of an unknown sample is within the judgment interval, it is judged that the unknown sample has a high degree of matching with the quantitative model, and the unknown sample is predicted normally.

[0014] (2) When the principal component score of an unknown sample is not within the judgment interval, it is determined that the unknown sample has a low matching degree with the quantitative model, and the spectral acquisition of the unknown sample is repeated.

[0015] As a further improvement of the present invention, the matching degree discrimination method further includes:

[0016] Step f. After secondary spectral acquisition, if the unknown sample is also determined to have a low matching degree with the quantitative model, then the unknown sample is regarded as an edge sample and normal prediction is performed.

[0017] Step g. After predicting unknown samples for a certain period of time, detect the matching degree between the quantitative model and the unknown samples in the unknown sample set of the collection period, and set a threshold to judge the overall matching degree.

[0018] As a further improvement of the present invention, in step f, the edge sample is a sample whose own calibration value exceeds the range of variation of the calibration value of the training set samples.

[0019] As a further improvement of the present invention, step g includes:

[0020] 1) When the overall matching degree of the unknown sample set in a given time period is greater than the threshold, the original quantitative model will be used to identify and predict the unknown samples normally in the next time period.

[0021] 2) When the overall matching degree of the unknown sample set in this time period is less than the threshold, the quantitative model is updated.

[0022] As a further improvement of the present invention, step b includes:

[0023] b1) According to the multivariate t-test method, the scores of each sample point for principal component one and principal component two after the changes follow an F distribution;

[0024] b2) Set the significance level to obtain the score thresholds for principal component one and principal component two. The formula for calculating the score threshold is:

[0025] threshold=(s h ·F 2,n-2,α ·2(n 2 -1) / (n(n-2))) 0.5 ;

[0026] Where threshold is the score threshold; s h Let F be the variance of the h-th principal component, where h is 1 or 2; 2,n-2,α The rejection region boundary of the F-distribution can be found by looking up a table; α is the significance level; n is the number of samples.

[0027] 7. The portable near-infrared spectroscopy quantitative model and unknown sample matching method according to claim 6, characterized in that, in step b, the scores of each sample point to principal component one and principal component two, calculated using the multivariate t-test method, follow an F-distribution after transformation, as shown in the formula:

[0028]

[0029] Among them, t hi The score of the i-th sample in the training set to the h-th principal component, where h is 1 or 2, s h Let be the variance of the h-th principal component, n be the number of samples, and m be the number of variables.

[0030] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0031] This invention provides a portable near-infrared spectroscopy quantitative model matching degree discrimination method for unknown samples. This method can identify abnormal samples by judging the matching degree between the unknown sample prediction process and the quantitative model, prompting the re-acquisition of spectra, avoiding human acquisition factors. When the sample state, internal and external environment, and spectrometer state change, the proportion of samples with low matching degree in the unknown sample set is large, which can prompt timely model updates and ensure the stability and accuracy of the spectrometer for long-term use. Attached Figure Description

[0032] Figure 1 This is a schematic diagram of a portable near-infrared spectroscopy quantitative model and a method for determining the matching degree of unknown samples according to the present invention.

[0033] Figure 2 This is a schematic diagram illustrating the matching degree between unknown samples and quantitative models in this invention. Detailed Implementation

[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0035] Example

[0036] Figure 1 A portable near-infrared spectroscopy quantitative model and a method for determining the matching degree of unknown samples are proposed, comprising the following steps:

[0037] S101. Based on the spectral data of the training set, a quantitative model is constructed using partial least squares (PLS) to obtain the loading matrix X_loadings of each variable of principal component one and principal component two of PLS ​​and the score matrix Xt_scores of each training set sample.

[0038] Specifically, Partial Least Squares (PLS) is used to transform potentially correlated population variables into a set of linearly uncorrelated variables, which become the principal components. The training set contains the spectrum and calibration values ​​(true values) of each sample, while unknown samples only contain the spectrum. By performing eigenvalue decomposition on the correlation coefficient matrix between the spectrum and calibration values ​​in the training set, the loading matrix and score matrix of the principal components are obtained. Lower-order principal components are retained while higher-order principal components are ignored, thus preserving highly correlated features while reducing dimensionality. In this embodiment, only two principal components, Principal Component 1 and Principal Component 2, are retained.

[0039] In constructing a quantitative model, multiple samples need to be processed simultaneously. This processing is represented by a matrix. The spectral set of multiple samples in the training set is used as the spectral matrix. After PLS processing, the spectral matrix can be expressed by the following formula:

[0040]

[0041] Where X is the n×m training set spectral matrix, n is the number of samples, m is the number of variables, and Xt is the training set spectral matrix. scores Let X be an n×p score matrix, where p is the principal component number. Since only principal component 1 and principal component 2 are retained here, p is 2. loadings It is an m×p load matrix, E is the residual matrix, and T is the transpose.

[0042] S102. Calculate the score thresholds of principal component one and principal component two using the multivariate t-test method, and substitute the two thresholds as the major and minor radii into the elliptical polar coordinate formula to construct the elliptical polar coordinate expression.

[0043] The specific steps include:

[0044] b1) According to the multivariate t-test method, the scores of each sample point on principal component one and principal component two, after transformation, follow an F-distribution, that is:

[0045]

[0046] Among them, t hi The score of the i-th sample in the training set to the h-th principal component, where h is 1 or 2, s h Let be the variance of the h-th principal component.

[0047] b2) Setting the significance level to 0.05 yields the threshold scores for Principal Component 1 and Principal Component 2, i.e.:

[0048] threshold=(s h ·F 2,n-2,0.05 ·2(n 2 -1) / (n(n-2))) 0.5 ;

[0049] Where threshold is the score threshold; s h Let F be the variance of the h-th principal component, where h is 1 or 2; 2,n-2,α The rejection region boundary of the F-distribution can be found by looking up a table; α is the significance level; n is the number of samples.

[0050] b3) The obtained score threshold matrix contains the score thresholds of principal component one and principal component two. Substitute this matrix into the elliptic polar coordinate matrix to construct the elliptic polar coordinate expression.

[0051] Substitute it into the elliptic polar coordinate matrix:

[0052] x = threshold(1) × cos(θ);

[0053] y = threshold(2) × sin(θ);

[0054] Where θ is the angle, and its value range is [0°, 360°]; threshold(1) is the score threshold of principal component one; threshold(2) is the score threshold of principal component two; x and y represent the independent and dependent variables of the formula, specifically the score of principal component one and the score of principal component two.

[0055] Substituting the principal component 1 of the unknown sample into the formula as x, we get two y values ​​of the same size but opposite signs, which can be used as a range threshold. When the score of the true principal component 2 of the unknown sample exceeds this range threshold, it is judged as an anomaly.

[0056] S103. Using the loading matrix X_loadings of the principal components obtained earlier, the collected unknown sample spectrum Xp is mapped to principal component one and principal component two to obtain the principal component one score Xp_scores1 and principal component two score Xp_scores2 corresponding to the sample.

[0057] The mapping method is as follows:

[0058] XP scores =Xp×X loadings ;

[0059] S104. Substitute the score value Xp_scores1 of principal component one into the elliptic polar coordinate expression to obtain the predicted score values ​​of principal component two, which have the same value but opposite signs. The judgment interval formed by these two values ​​can be represented as: [-Xk_scores,Xk_scores];

[0060] S105. Detect whether the score of the true principal component II of the unknown sample is within the judgment interval, thereby determining the matching degree between the unknown sample and the quantitative model.

[0061] Specifically:

[0062] (1) When Xp_scores2∈[-Xk_scores,Xk_scores] for a single unknown sample, it indicates that the sample has a high degree of matching with the quantitative model and can be directly predicted.

[0063] (2) When a single unknown sample If the sample does not match the quantitative model well, the spectrum should be re-acquired to eliminate the interference of human factors.

[0064] S106. After secondary spectral acquisition, if the sample is also judged to have a low degree of matching with the quantitative model, it means that the influence of external factors such as manual operation and experimental environment can be ruled out. Then, the unknown sample is treated as a marginal sample and normal prediction operation is performed.

[0065] In this embodiment, edge samples are those whose self-calibration values ​​exceed the range of variation of the calibration values ​​of the training set samples.

[0066] S107. After predicting unknown samples for a period of time, detect the matching degree between the quantitative model and the samples in the unknown sample set for that period of time, and set a threshold to judge the overall matching degree. The period of time can be set to one month according to the frequency of the predicted samples, and the overall matching degree threshold is set to 80%.

[0067] S108. Use a threshold to determine the overall matching degree in the unknown sample set. When the overall matching degree is greater than the threshold, use the original model to make judgments and predictions in the next time period. When the overall matching degree is less than the threshold, update the model to enhance the model's adaptability to the unknown samples.

[0068] Figure 2 This diagram illustrates the matching degree between unknown samples and the quantitative model. The horizontal and vertical coordinates of each dot represent the principal component one and principal component two scores for each unknown sample. The ellipse is drawn using polar coordinates based on the score thresholds. Dots inside the ellipse indicate a higher matching degree between the corresponding unknown sample and the model, while dots outside the ellipse indicate a lower matching degree.

[0069] The method of this invention can identify abnormal samples by judging the matching degree between the unknown sample prediction process and the quantitative model, prompting the re-acquisition of spectra, avoiding human factors in the acquisition process. When the sample state, internal and external environment, and spectrometer state change, the proportion of samples with low matching degree in the unknown sample set is relatively large, which can prompt timely model updates and ensure the stability and accuracy of the spectrometer during long-term use.

[0070] In addition, the present invention also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described portable near-infrared spectroscopy quantitative model and unknown sample matching degree discrimination method.

[0071] Although the present invention has been described herein with reference to illustrative embodiments, the above embodiments are merely preferred embodiments of the present invention, and the implementation of the present invention is not limited to the above embodiments. It should be understood that those skilled in the art can devise many other modifications and implementations, which will fall within the scope and spirit of the principles disclosed in this application.

Claims

1. A portable near-infrared spectroscopy quantitative model and unknown sample matching degree discrimination method, characterized in that, The method comprises the following steps: Step a. Constructing a quantitative model by using partial least squares (PLS) based on the spectral data of the training set to obtain the loading matrix X_loadings of each variable of the first principal component and the second principal component and the score matrix Xt_scores of each training set sample; Step b. Calculating the score threshold of the first principal component and the second principal component by using a multivariate t-test method, and bringing the two thresholds into an elliptical polar coordinate formula as a long radius and a short radius to construct an elliptical polar coordinate expression; In step b, the score threshold of the first principal component and the second principal component is calculated by using a multivariate t-test method, and the specific method comprises: b1) According to the multivariate t-test method, the score values of each sample point on the first principal component and the second principal component after being changed are subjected to F distribution; b2) Setting a significant level to obtain the score threshold on the first principal component and the second principal component, and the calculation formula of the score threshold is: ; wherein, is a score threshold value; is a first variance of the principal component, is 1 or 2; is distribution rejection region boundary, which can be obtained by looking up a table; α is a significance level; n is a sample size. Step c. Mapping the collected unknown sample spectrum into the first principal component and the second principal component by using the loading matrix X_loadings obtained in step a to obtain the score values of the unknown sample on the first principal component and the second principal component; Step d. Substituting the score value of the first principal component into the elliptical polar coordinate expression constructed in step b to obtain two predicted score values of the second principal component which are the same in value but opposite in sign, and the two predicted score values of the second principal component form a judgment interval; Step e. Detecting whether the score value of the unknown sample on the second principal component is in the judgment interval one by one to judge the matching degree of the unknown sample with the quantitative model.

2. The method according to claim 1, wherein, In step e, the matching degree of the unknown sample with the quantitative model comprises: (1) When the score value of the unknown sample on the second principal component is in the judgment interval, it is judged that the matching degree of the unknown sample with the quantitative model is high, and the unknown sample is normally predicted; (2) When the score value of the unknown sample on the second principal component is not in the judgment interval, it is judged that the matching degree of the unknown sample with the quantitative model is low, and the spectrum collection of the unknown sample is re-performed.

3. The method according to claim 2, wherein the method is characterized by, The matching degree discrimination method further comprises: Step f. After the second spectrum collection, if it is also judged that the matching degree of the unknown sample with the quantitative model is low, the unknown sample is regarded as an edge sample and is normally predicted; Step g. After predicting the unknown samples in a time period, detecting the matching degree of the quantitative model and the unknown samples in the unknown sample set in the time period, and setting a threshold to judge the overall matching degree.

4. The method according to claim 3, wherein the method is characterized by, In step f, the edge sample is a sample whose own calibration value exceeds the change range of the calibration values of the training set samples.

5. The method according to claim 3, wherein the method is characterized by, In step g, it comprises: 1) When the overall matching degree of the unknown sample set in the time period is greater than the threshold, the original quantitative model is used to discriminate and normally predict the unknown samples in the next time period; 2) When the overall matching degree of the unknown sample set in the time period is less than the threshold, the quantitative model is updated.

6. The method according to any one of claims 1-5, wherein, In step b, the score values of each sample point on the first principal component and the second principal component after being changed are subjected to F distribution according to the multivariate t-test method, and the formula is: ; in, For the training set The nth sample pair The scores of each principal component, It is 1 or 2. For the first The variance of the principal components, where n is the number of samples and m is the number of variables.

Citation Information

Patent Citations

  • Gasoline property evaluation method and device thereof

    CN110987866A

  • Methods for analysis of spectral data and their applications osteoporosis

    US20050037515A1