A sample identification method based on near-infrared spectroscopy
Patent Information
- Application Number
- CN202311524753.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2026-09-08
- Estimated Expiration
- 2043-11-14
AI Technical Summary
[0004]本发明的主要目的在于解决现有技术中依靠外观评价和传统化学分析方法对烟叶产地进行鉴别所存在的主观性强、准确性低、费时费力的问题
[0040] The sample identification method based on near-infrared spectroscopy provided by this invention establishes a general model of all samples in the modeling sample set and a classification model of each category of samples in the modeling sample set. Then, the near-infrared spectrum of the sample to be tested is projected onto the general model and each classification model respectively to obtain the length of the residual vector of the sample to be tested orthogonal to the model. After sorting by the size of the residual vector, the sample to be tested is classified, which can reduce time and cost and improve detection efficiency.
Smart Images

Figure CN117554326B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of near-infrared spectroscopy analysis technology, and in particular to a sample identification method based on near-infrared spectroscopy. Background Technology
[0002] Tobacco leaves are the foundation of the cigarette industry and have always been a focus of industry attention and research, especially as the basis for cigarette formulation design. Due to the influence of factors such as internal genetic composition, external environmental conditions, and cultivation practices during the growth process of tobacco leaves, the quality of tobacco leaves from different planting regions, different parts, and different grades varies greatly, resulting in distinct style characteristics. Therefore, different types of tobacco leaves can be blended during the cigarette formulation design process to meet the different positioning requirements of the products.
[0003] In the tobacco industry, the planting area is a crucial basis for the classification and management of tobacco leaf quality. Currently, most companies rely on appearance evaluation and traditional chemical analysis methods to identify the origin of tobacco leaves, which suffers from drawbacks such as strong subjectivity, low accuracy, and being time-consuming and labor-intensive. Therefore, establishing a rapid and effective non-artificial method for identifying the origin of tobacco leaves is of great significance. Summary of the Invention
[0004] The main objective of this invention is to address the problems of high subjectivity, low accuracy, and time-consuming and labor-intensive methods in existing technologies that rely on appearance evaluation and traditional chemical analysis to identify the origin of tobacco leaves. To achieve this objective, this invention provides a sample identification method based on near-infrared spectroscopy. This method can rapidly identify the origin of tobacco leaves using near-infrared spectroscopy technology. Furthermore, this invention employs a two-step identification process to determine the category of the sample, significantly improving the accuracy of sample identification.
[0005] One embodiment of the present invention provides a sample identification method based on near-infrared spectroscopy, comprising:
[0006] The system obtains a total recognition model and multiple classification models. The total recognition model is obtained based on the near-infrared spectra of modeling samples of multiple different categories. The multiple classification models correspond to different categories, and each classification model is obtained based on the near-infrared spectra of modeling samples of the same category.
[0007] Determine the first threshold;
[0008] The near-infrared spectrum of the sample to be tested is obtained as the spectrum to be tested, and the length of the first residual vector orthogonal to the total recognition model space is obtained.
[0009] Based on the fact that the length of the first residual vector is less than the first threshold, for each classification model, the length of the second residual vector that is orthogonal to the space of the spectrum to be tested and the classification model is obtained respectively.
[0010] The category of the sample to be tested is determined based on the length of each second residual vector.
[0011] As a specific embodiment of the present invention, determining the category of the sample to be tested based on the length of each second residual vector includes:
[0012] The length of each second residual vector is standardized to obtain the standard length of each second residual vector;
[0013] The category corresponding to the second residual vector with the smallest standard length is determined as the category of the sample to be tested.
[0014] As a specific embodiment of the present invention, the lengths of each second residual vector are standardized to obtain the standard lengths of each second residual vector, including:
[0015] For each category, the upper and lower limits of the standard are determined based on the classification model corresponding to that category and the near-infrared spectra of each modeled sample in that category.
[0016] The standard length of each second residual vector is obtained by standardizing each second residual vector based on the following formula.
[0017]
[0018] Among them, C p r represents the standard length of the second residual vector corresponding to the p-th category. dp The length of the second residual vector corresponding to the p-th category is max(r). p,q ) and min(r p,q ) represent the upper and lower limits of the standard corresponding to the p-th category, respectively.
[0019] As a specific embodiment of the present invention, based on the classification model corresponding to the category and the near-infrared spectra of each modeled sample of the category, the upper and lower limits of the standard corresponding to the category are determined, including:
[0020] Project the near-infrared spectra of all modeling samples belonging to the category onto the classification model corresponding to the category to obtain the length of the third residual vector orthogonal to the classification model space of the near-infrared spectra of each modeling sample of the category. Then, obtain the maximum and minimum values of the lengths of each third residual vector as the standard upper limit and standard lower limit corresponding to the category.
[0021] As a specific embodiment of the present invention, determining the first threshold includes:
[0022] Project the near-infrared spectra of all modeling samples onto the overall recognition model to obtain the length of the fourth residual vector that is orthogonal to the space of the near-infrared spectra of each modeling sample and the overall recognition model.
[0023] Obtain the length of each fourth residual vector;
[0024] Determine the average length and standard deviation of each fourth residual vector;
[0025] The first threshold is determined based on the mean and standard deviation.
[0026] As a specific embodiment of the present invention, the first threshold is equal to the sum of the average value and 3 to 5 times the standard deviation.
[0027] As a specific embodiment of the present invention, the steps for obtaining the overall recognition model include:
[0028] Determine the modeling sample set and obtain the near-infrared spectra of all modeling samples in the modeling sample set; wherein, the modeling sample set includes modeling samples of multiple different categories;
[0029] A total spectral matrix is constructed based on the near-infrared spectra of all modeling samples. Singular value decomposition is performed on the total spectral matrix, and the total spectral matrix is reconstructed using the first set number of principal component factors obtained from the decomposition to obtain the total reconstructed spectral matrix.
[0030] A total recognition model is constructed based on the total reconstructed spectral matrix.
[0031] As a specific embodiment of the present invention, the overall recognition model is as follows:
[0032]
[0033] Among them, H t Let I represent the overall recognition model, and let X represent the identity matrix. tnew Represents the total reconstructed spectral matrix. For X tnew The generalized inverse matrix.
[0034] As a specific embodiment of the present invention, the first set quantity is obtained based on interactive verification.
[0035] As a specific embodiment of the present invention, obtaining multiple classification models includes performing the following steps on the modeling samples of each category:
[0036] For all modeled samples of the same category, obtain their near-infrared spectra;
[0037] A classification spectral matrix is constructed based on the near-infrared spectra of all modeled samples of each category. Singular value decomposition is performed on the classification spectral matrix, and the spectral reconstruction is performed on the classification spectral matrix using the second set number of principal component factors obtained from the decomposition, thus obtaining the classification reconstructed spectral matrix.
[0038] A classification model corresponding to each category is constructed based on the reconstructed spectral matrix.
[0039] Compared with the prior art, the present invention has at least the following technical effects:
[0040] The sample identification method based on near-infrared spectroscopy provided by this invention establishes a general model of all samples in the modeling sample set and a classification model of each category of samples in the modeling sample set. Then, the near-infrared spectrum of the sample to be tested is projected onto the general model and each classification model respectively to obtain the length of the residual vector of the sample to be tested orthogonal to the model. After sorting by the size of the residual vector, the sample to be tested is classified, which can reduce time and cost and improve detection efficiency. Attached Figure Description
[0041] Figure 1 A flowchart illustrating the sample identification method based on near-infrared spectroscopy provided in an embodiment of the present invention is shown;
[0042] Figure 2 The near-infrared spectra of five samples from different origins after preprocessing, as provided in an embodiment of the present invention, are shown.
[0043] Figure 3 This diagram shows the distribution of residual vector lengths for the overall model built from all modeling samples provided in this embodiment of the invention.
[0044] Figure 4 This diagram illustrates the distribution of the residual vector lengths projected from the test sample onto the overall model, as provided in an embodiment of the present invention.
[0045] Figure 5 This diagram illustrates the distribution of the residual vector lengths of the test sample projected onto the first-class classification model, as provided in an embodiment of the present invention.
[0046] Figure 6 This diagram illustrates the distribution of residual vector lengths of a test sample obtained in different classification models according to an embodiment of the present invention. Detailed Implementation
[0047] The following specific embodiments illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention will be presented in conjunction with preferred embodiments, this does not mean that the features of the invention are limited to this embodiment. On the contrary, the purpose of describing the invention in conjunction with embodiments is to cover other options or modifications that may be derived based on the claims of the present invention. To provide a deep understanding of the invention, many specific details will be included in the following description. The invention may also be implemented without using these details. Furthermore, to avoid confusion or obscuring the focus of the invention, some specific details will be omitted in the description. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0048] It should be noted that in this specification, similar reference numerals and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0049] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0050] Near-infrared spectroscopy (NIRS) technology, with its advantages of speed, efficiency, and non-destructive operation, has been widely used in the tobacco industry. However, due to the severe peak overlap in NIRS data, it is often necessary to combine it with multiple chemometric methods to extract key spectral information, establish a link between NIRS and tobacco-producing regions, and thus identify tobacco-growing areas.
[0051] A specific embodiment of the present invention provides a sample identification method based on near-infrared spectroscopy, specifically, as follows: Figure 1 As shown, the method includes the following steps:
[0052] Step S101: Obtain the overall recognition model and multiple classification models.
[0053] Specifically, the overall identification model is obtained based on the near-infrared spectra of modeling samples of multiple different categories, and the multiple classification models correspond to different categories, with each classification model obtained based on the near-infrared spectra of modeling samples of the same category.
[0054] For example, methods for constructing the overall recognition model include:
[0055] Step S1011: Determine the modeling sample set and obtain the near-infrared spectra of all modeling samples in the modeling sample set; wherein, the modeling sample set includes modeling samples of multiple different categories.
[0056] Step S1012: Construct the total spectral matrix X based on the near-infrared spectra of all modeled samples. S For the total spectral matrix X S Singular value decomposition is performed, and the top A principal component factors obtained from the decomposition are used to reconstruct the total spectral matrix, resulting in the total reconstructed spectral matrix X. tnew .
[0057] Taking a total of m samples for modeling and n as an example, the total spectral matrix X of the modeling sample set is... S It can be represented as:
[0058]
[0059] Wherein, the data element x located in the i-th row and j-th column of the matrix ij This represents the spectral data of the i-th modeling sample in the modeling sample set at the j-th wave point.
[0060] For example, in the total spectral matrix X S Before performing singular value decomposition, the total spectral matrix X can be analyzed first. S Preprocessing is performed to obtain the preprocessed total spectral matrix X. t Preprocessing includes Standard Normal Variation (SNV) or Multiplicative Scatter Correction (MSC), which are used to eliminate the influence of solid particle size, surface scattering and optical path variation on the spectrum of the modeled sample, thereby improving the accuracy of subsequent identification.
[0061] Then, the preprocessed total spectral matrix X t Singular value decomposition is performed, and the number of principal component factors used to construct the spectrum is obtained using cross-validation, i.e., a first predetermined number A. Then, the total reconstructed spectral matrix of all samples is reconstructed using A principal component factors. Where U and V are orthogonal matrices, and S is a diagonal matrix. The matrix X... tnew It covers the effective information of the spectra of all samples in the modeling sample set, reducing the interference of other information (such as noise, background, etc.).
[0062] Step S1013: Based on the total reconstructed spectral matrix X tnew Construct the overall recognition model. For example, the expression for the overall recognition model is:
[0063]
[0064] Among them, H t Let I represent the overall recognition model, and let X represent the identity matrix. tnew Represents the total reconstructed spectral matrix. For X tnew The generalized inverse matrix.
[0065] For example, the construction methods for each classification model include the following steps:
[0066] For each category, a classification spectral matrix is constructed based on the near-infrared spectra of all modeling samples belonging to that category. This classification spectral matrix is then preprocessed using SNV or MSC to obtain a preprocessed classification spectral matrix. Singular value decomposition is performed on the preprocessed classification spectral matrix, and the number of principal component factors used to construct the spectrum (i.e., the second predetermined number B) is obtained using cross-validation. Then, the classification reconstruction spectral matrix for that category is reconstructed using B principal component factors. Finally, the classification spectral matrix is reconstructed using the second predetermined number of principal component factors obtained from the decomposition to obtain the classification reconstruction spectral matrix. Based on the classification reconstruction spectral matrix, a classification model corresponding to the category is constructed.
[0067] Furthermore, in order to improve the near-infrared spectral differences between categories, classification vectors for each category can be constructed for the modeling samples of each category. For example, the classification vectors can be constructed using trigonometric functions, such as using a sine function for the first category and a cosine function for the second category. That is, the classification vectors between each category can be set to have large differences, which will help to further increase the distinguishability between each category.
[0068] Then, based on the constructed classification vectors and their classification spectral matrices for each category, a new matrix for each category is obtained. For example, using h... p and X p Let denot n be the classification vector and classification spectrum matrix of the p-th category, respectively. Then, the new matrix NX for the p-th category... p It can be represented as NX p =[X p m p *h p ], where m p Let be a column vector, which can be represented as (1,1,...,1). T In the form of m p The number of elements and the classification spectral matrix X p The number of rows is equal; h p It is a row vector.
[0069] After obtaining the new matrix, singular value decomposition is performed on the constructed new matrix. The number of principal component factors used to construct the spectrum is obtained by cross-validation, i.e., the second set number B. Then, the classification reconstruction spectrum matrix of this category is reconstructed using B principal component factors. The new matrix is then spectrally reconstructed using the second set number of principal component factors obtained from the decomposition to obtain the classification reconstruction spectrum matrix.
[0070] Optionally, the classification model for each category can proceed sequentially from the first category to the last category, as follows:
[0071] (a) The pre-processed classification spectral matrix of each modeling sample in the first category is denoted as X1 (with a size of m1 samples and n variables). To improve the near-infrared spectral differences between categories, a classification vector h1 is constructed for this category, resulting in a new matrix for the first category samples:
[0072]
[0073] The number of principal components B is obtained using cross-validation, and then the classification and reconstruction spectral matrix of the first category is reconstructed using B principal component factors. (U and V are orthogonal matrices, and S is a diagonal matrix); Construct the classification model H1 for the first category: Where I represents the identity matrix, For matrix NX 1new The generalized inverse matrix.
[0074] (b) Repeat the above steps in a loop until a classification model for all categories is established, denoted as H1, H2, H3, etc.
[0075] Step S102: Determine the first threshold.
[0076] For example, the near-infrared spectra of each modeling sample in the modeling sample set are sequentially fed into the overall recognition model H. t Projection is performed to obtain the fourth residual vector, which is orthogonal to the near-infrared spectrum of each modeling sample and the overall recognition model space. in, Here, we distinguish the preprocessed near-infrared spectrum of the i-th modeling sample, which can be represented in vector form.
[0077] Calculate the length of the fourth residual vector for each modeled sample. i = 1, 2, 3, ... m. The average and standard deviation of the length of each fourth residual vector are calculated, and the sum of the average and 3-5 times the standard deviation is used as the first threshold.
[0078] Step S103: Obtain the near-infrared spectrum of the sample to be tested as the spectrum to be tested, and obtain the length of the first residual vector that is orthogonal to the total recognition model space.
[0079] Specifically, the formula for calculating the first residual vector is as follows: in, The spectrum to be measured, having undergone the same preprocessing as the modeling sample, can be represented as a vector. The first residual vector... The length is the L2 norm of the first residual vector.
[0080] Step S104: Compare the length of the first residual vector obtained in step S103 with the first threshold. If the length of the first residual vector is greater than or equal to the first threshold, it means that the sample to be tested does not belong to any of the categories in the above modeling sample set. If the length of the first residual vector is less than the first threshold, continue to execute steps S105 and S106 to further identify the specific category of the sample to be tested.
[0081] Step S105: Obtain the length of the second residual vector that is orthogonal to the classification model space of each category of the spectrum to be measured. in, H represents the second residual vector, which is orthogonal to the classification model space of the p-th category and is the spectrum to be measured. p Let represent the classification model for the p-th category.
[0082] Furthermore, when each classification model is constructed based on a new matrix for each category, it is also necessary to generate a classification vector h for the test sample. p The construction of a new spectra of the sample to be tested is used to obtain the new spectra of the sample to be tested. The formula for calculating the length of the second residual vector, which is orthogonal to the classification model space of each category, is as follows: Where, r dp It represents the length of the second residual vector that is orthogonal to the classification model space of the p-th category of the new spectrum to be measured, that is, the length of the second residual vector corresponding to the p-th category.
[0083] Step S106: Determine the category of the sample to be tested based on the length of each second residual vector.
[0084] For example, step S106 may include the following steps:
[0085] Step S1061: Standardize the length of each second residual vector to obtain the standard length of each second residual vector.
[0086] Specifically, after obtaining the classification models for each category, the near-infrared spectra of each modeling sample in that category are projected onto the classification model to obtain the length of the third residual vector orthogonal to the classification model space for the near-infrared spectra of each modeling sample in that category.
[0087]
[0088] Where q = 1, 2, 3, ..., m p ; Let r represent the third residual vector, which is orthogonal to the classification model space of the q-th modeling sample of the p-th class. p,q Represents the third residual vector The length of Hp This represents the classification model for the p-th category. This represents the near-infrared spectrum of the q-th modeled sample in the p-th category.
[0089] Among the lengths of the third residual vectors obtained above, the maximum and minimum values are taken as the standard upper limit and standard lower limit for that category, respectively.
[0090] Similarly, when each classification model is constructed based on a new matrix for each category, the near-infrared spectra of each modeled sample for each category are... Replace with The length of the third residual vector at this point is
[0091] Then, based on the following formula (1), the length r of each second residual vector is... dp Perform standardization to obtain the standard length C of each second residual vector. p .
[0092]
[0093] Among them, C p Let r represent the standard length of the second residual vector corresponding to the p-th category (i.e., the second residual vector orthogonal to the classification model space of the p-th category from the near-infrared spectrum of the sample to be tested). dp The length of the second residual vector corresponding to the p-th category is max(r). p,q ) and min(r p,q ) represent the upper and lower limits of the standard corresponding to the p-th category, respectively.
[0094] Those skilled in the art should understand that, in addition to the max-min standardization method described above, other standardization methods, such as mean standardization and Z-score standardization, can be used to adjust the length r of each second residual vector. dp This application does not impose any restrictions on the standardization process.
[0095] Step S1062: Determine the category corresponding to the second residual vector with the smallest standard length, which is the category of the sample to be tested.
[0096] The sample identification method based on near-infrared spectroscopy of the present invention establishes a general model of all samples in the modeling sample set and a classification model of each category of samples in the modeling sample set. Then, the near-infrared spectrum of the sample to be tested is projected onto the general model and each classification model respectively to obtain the length of the spectral residual vector of the sample to be tested orthogonal to the model. After sorting by the size of the residual vector, the sample to be tested is classified, which can reduce time and cost and improve detection efficiency.
[0097]
Example
[0098] Taking the classification of tobacco leaf samples as an example, the specific process of the method in this application will be explained.
[0099] Step 1): Tobacco Leaf Sample Processing
[0100] 459 tobacco leaf samples were provided by Guizhou Tobacco Industry Co., Ltd., originating from Guangdong, Henan, Hunan, Sichuan, and Yunnan provinces. Before collecting the sample spectra, the tobacco leaf samples were placed in a 40℃ oven for two hours according to "YCT 31-1996 Preparation of Tobacco and Tobacco Product Samples and Determination of Moisture Content by Oven Method". The samples were then removed and cooled to room temperature. After pulverizing the samples in a plant pulverizer, they were passed through a 40-mesh sieve to separate samples with a particle size less than 40 mesh (≤0.45mm). Once the samples had cooled to room temperature, they were placed in disposable sealed bags and stored at a low temperature and protected from light.
[0101] Step 2): Experimental instruments and spectral acquisition
[0102] The laboratory temperature was controlled between 22±2℃ and the relative humidity between 40%±1%. The near-infrared instrument was preheated for at least 1 hour and then calibrated using ValPro before use. An appropriate amount of prepared tobacco powder was placed in a sample cup for scanning. The near-infrared spectral wavenumber scanning range was 4000-10000 cm⁻¹. -1 The resolution is 8cm. -1 ; 64 scans.
[0103] Step 3): Establishment and application of the classification model
[0104] Near-infrared spectra of tobacco leaf powder samples, such as Figure 2 As shown, from Figure 2 It can be seen that the tobacco leaf samples from the five production areas show certain differences, with the main spectral differences occurring between 4200 and 7000 cm⁻¹. -1 The wavenumber range. Following the steps above, a total model of 299 modeling samples was established. The near-infrared spectra of the total model were reconstructed using 10 principal components, resulting in the total model matrix H. t And calculate the length of the residual vector for all samples. Figure 3 It can be seen that the length distribution of its vectors lies in the interval between 0.02 and 0.08. The thresholds for all modeled samples under this overall model are calculated. If calculated with a standard deviation of 3, threshold 1 is 0.0787; if calculated with a standard deviation of 5, threshold 2 is 0.1004. Following the above method, a classification model with 5 categories is established, and the matrices of the reconstructed classification models using 8 principal components are recorded as H1, H2, H3, H4, and H5, respectively. The length of the residual vector of the modeled samples in each class is calculated, and its maximum and minimum values are determined.
[0105] First, the near-infrared spectra of 160 unsampled samples are projected onto the overall model to obtain the residual optical vector spectrum of each sample to be tested. The length of this vector spectrum is then calculated. (See...) Figure 4 As shown in the figure, the lengths of the spectral residual spectral vectors of samples 69#, 70#, 71#, 72#, 73#, and 131# to 160# far exceed thresholds 1 and 2. Therefore, the spectral space of these samples does not belong to the overall model space, and thus the model cannot be used to further determine the class of these samples. The remaining samples with spectral residual vector lengths below the thresholds are projected into the classification model, as shown below. Figure 5 As shown, the remaining 125 samples were projected onto the Class 1 classification model to obtain the lengths of their spectral residual vectors. It was found that samples 1# to 21# had relatively low lengths. Similarly, projecting them onto other classification models yielded... Figure 6 The results show that samples 1#-21# belong to category 1; samples 22#-32# belong to category 2; samples 33#-68# and 74#-85# belong to category 3; samples 86#-121# belong to category 4; and samples 122#-130# belong to category 5.
[0106] This invention uses a two-step identification method to determine the category of the sample to be tested, which greatly improves the accuracy of sample classification.
[0107] Although the present invention has been illustrated and described with reference to embodiments thereof, those skilled in the art should understand that the above description is a further detailed explanation of the invention in conjunction with specific embodiments, and should not be construed as limiting the specific implementation of the invention to these descriptions. Those skilled in the art can make various changes in form and detail, including several simple deductions or substitutions without departing from the spirit and scope of the invention.
Claims
1. A sample identification method based on near-infrared spectroscopy, characterized in that, include: A total identification model and multiple classification models are obtained, wherein the total identification model is obtained based on the near-infrared spectra of modeling samples of multiple different categories, and the multiple classification models correspond to different categories, and each classification model is obtained based on the near-infrared spectra of modeling samples of the same category; Determine the first threshold; The near-infrared spectrum of the sample to be tested is obtained as the spectrum to be tested, and the length of the first residual vector orthogonal to the spectrum to be tested and the total recognition model space is obtained. Based on the fact that the length of the first residual vector is less than the first threshold, for each classification model, the length of the second residual vector that is orthogonal to the space of the spectrum to be measured and the classification model is obtained respectively; The lengths of each of the second residual vectors are standardized to obtain the standard lengths of each of the second residual vectors. The category corresponding to the second residual vector with the smallest standard length is determined as the category of the sample to be tested.
2. The sample identification method based on near-infrared spectroscopy as described in claim 1, characterized in that, The lengths of each of the second residual vectors are standardized to obtain the standard lengths of each of the second residual vectors, including: For each category, the upper and lower standard limits corresponding to the category are determined based on the classification model corresponding to the category and the near-infrared spectra of each modeled sample of the category. The standard length of each second residual vector is obtained by standardizing each second residual vector based on the following formula; Among them, C p This represents the standard length of the second residual vector corresponding to the p-th category. This represents the length of the second residual vector corresponding to the p-th category. and Let represent the upper and lower limits of the standard corresponding to the p-th category, respectively.
3. The sample identification method based on near-infrared spectroscopy as described in claim 2, characterized in that, Based on the classification model corresponding to the category and the near-infrared spectra of each modeled sample of that category, determine the upper and lower standard limits corresponding to the category, including: The near-infrared spectra of all modeling samples belonging to the category are projected onto the classification model corresponding to the category to obtain the length of the third residual vector orthogonal to the space of the classification model for the near-infrared spectra of each modeling sample in the category. The maximum and minimum values of the lengths of each third residual vector are then obtained as the upper and lower limits of the standard corresponding to the category.
4. The sample identification method based on near-infrared spectroscopy as described in claim 1, characterized in that, Determining the first threshold includes: Project the near-infrared spectra of all the modeling samples onto the overall recognition model to obtain the length of the fourth residual vector that is orthogonal to the space of the near-infrared spectra of each modeling sample and the overall recognition model. Obtain the length of each of the fourth residual vectors; Determine the average length and standard deviation of each of the fourth residual vectors; The first threshold is determined based on the average value and the standard deviation.
5. The sample identification method based on near-infrared spectroscopy as described in claim 4, characterized in that, The first threshold is equal to the sum of the average value and 3 to 5 times the standard deviation.
6. The sample identification method based on near-infrared spectroscopy as described in claim 1, characterized in that, The steps for obtaining the overall recognition model include: A modeling sample set is determined, and the near-infrared spectra of all modeling samples in the modeling sample set are obtained; wherein, the modeling sample set includes modeling samples of multiple different categories; A total spectral matrix is constructed based on the near-infrared spectra of all modeling samples. Singular value decomposition is performed on the total spectral matrix, and the total spectral matrix is spectrally reconstructed using a first set number of principal component factors obtained from the decomposition to obtain the total reconstructed spectral matrix. The overall recognition model is constructed based on the overall reconstructed spectral matrix.
7. The sample identification method based on near-infrared spectroscopy as described in claim 6, characterized in that, The overall identification model is as follows: in, Let I represent the overall recognition model, and let I represent the identity matrix. Represents the total reconstructed spectral matrix. The generalized inverse matrix.
8. The sample identification method based on near-infrared spectroscopy as described in claim 6, characterized in that, The first set quantity is obtained based on interactive verification.
9. The sample identification method based on near-infrared spectroscopy as described in claim 1, characterized in that, Obtaining the multiple classification models involves performing the following steps on the modeling samples for each category: For all the modeled samples of the same category, obtain their near-infrared spectra; A classification spectral matrix is constructed based on the near-infrared spectra of all modeled samples of the category. Singular value decomposition is performed on the classification spectral matrix, and the spectral reconstruction is performed on the classification spectral matrix using a second set number of principal component factors obtained from the decomposition, to obtain the classification reconstructed spectral matrix. The classification model corresponding to the category is constructed based on the reconstructed spectral matrix.
Citation Information
Patent Citations
Near-infrared model maintenance method
CN116049617A
Method for stabilizing near-infrared models and determining their applicability
US5668374A