A method for identifying abnormal samples in a sample set

By constructing the projection matrix and calculating the 2-norm identification of anomaly samples of pure spectral signals, the problem of not being able to effectively identify the abnormal samples in the sample set in the prior art is solved, and the accuracy and robustness of the model are improved.

CN116046717BActive Publication Date: 2025-07-25CHINA TOBACCO GUIZHOU IND
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310139085.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-20
Publication Date
2025-07-25
Estimated Expiration
2043-02-20

AI Technical Summary

Technical Problem

In the prior art, abnormal samples in the sample set cannot be effectively identified, resulting in the impact of the quality and accuracy of the correction model.

Method used

By constructing the projection matrix corresponding to the sample, identify the abnormal samples using the 2 norms and abnormal eigenvalues of the pure spectral signal. The specific steps include obtaining the spectral data and concentration values of the sample, constructing the projection matrix, calculating the 2 norms of the pure spectral signal, and setting a threshold based on the mean and standard deviation to judge the abnormal eigenvalue of the sample.

Benefits of technology

Effectively identifying and eliminating abnormal samples improves the accuracy and prediction ability of the correction model and enhances the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116046717B_ABST
    Figure CN116046717B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for identifying abnormal samples in a sample set, including: obtaining spectral data and concentration values of the components to be measured of each sample in the sample set; for each sample in the sample set, constructing a projection matrix corresponding to the sample based on the spectral data and concentration values of the other samples in the sample set; for each sample in the sample set, respectively obtaining each target value of the spectral data of the sample under the projection matrices corresponding to the other samples in the sample set; and determining the abnormal samples in the sample set based on the target values of each sample in the sample set. This method has a simple process and can effectively identify abnormal samples.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of near-infrared spectroscopy analysis, and particularly relates to a method for identifying abnormal samples in a sample set. Background Art

[0002] As a green analysis technology, near-infrared spectroscopy analysis technology has the advantages of fast analysis speed, simple operation, and can realize in-situ, non-destructive, and on-line qualitative and quantitative analysis, etc., and has gradually become the preferred method for quality control and evaluation of tobacco and tobacco products. By applying near-infrared technology, the contents of chemical components (such as total alkaloids, total sugars, etc.) in tobacco and tobacco products can be obtained quickly, enabling the evaluation of cigarette quality to shift from sensory evaluation to the combination with internal quality, so as to achieve the mutual unity of appearance quality and internal quality.

[0003] In order to accurately predict the content of chemical components in tobacco based on near-infrared spectroscopy analysis technology, a calibration model needs to be constructed. The establishment of the calibration model is often based on a large number of samples, and the quality of the established calibration model depends to a large extent on the accuracy of the sample data participating in the modeling. If there are abnormal samples in the model, it will directly destroy the similarity between samples, thereby affecting the quality and accuracy of the model. Therefore, the identification of abnormal samples is the premise and foundation for establishing a reliable and accurate calibration model, which plays a very important role in improving the prediction ability of the model. However, there is currently no method that can effectively identify abnormal samples in a sample set. Summary of the Invention

[0004] The main purpose of the present invention is to solve the problem that abnormal samples cannot be effectively identified in the prior art.

[0005] To achieve the above object, an embodiment of the present invention provides a method for identifying abnormal samples in a sample set, which can accurately identify abnormal samples in the sample set based on the spectral data of each sample in the sample set and the concentration values of the components to be measured, thereby facilitating the establishment of a reliable and accurate calibration model. Specifically, the method includes:

[0006] Obtain the spectral data and concentration values of the components to be measured of each sample in the sample set;

[0007] For each sample in the sample set, based on the spectral data and concentration values of other samples in the sample set, construct a projection matrix corresponding to the sample;

[0008] For each sample in the sample set, respectively obtain each target value of the spectral data of the sample under the projection matrices corresponding to other samples in the sample set;

[0009] Based on the respective target values of each sample in the sample set, determine the abnormal samples in the sample set.

[0010] Specifically, the component to be measured in this method can be the chemical components of tobacco or tobacco products, such as total alkaloids, total sugars, total nitrogen, etc.

[0011] Among them, the projection matrix corresponding to each sample is constructed based on the other remaining samples in the sample set after removing this sample. That is, for sample i, its corresponding projection matrix is constructed based on all other samples in the sample set except sample i.

[0012] This solution first analyzes the spectral data of each sample in the sample set and the concentration of the component to be measured in each sample, so as to obtain the target value of each sample (where the target value is related to the pure spectral signal), and then identifies the abnormal samples in the sample set by comparing the target values of each sample. Specifically, this solution uses the method of excluding one by one to calculate the target value of each sample. That is, when calculating the target value of a certain sample, the data of this sample needs to be excluded, and the data of other samples in the sample set except this sample are used to calculate this target value. If there are abnormal samples among the other samples in the sample set at this time, then this target value will be different from the target value obtained when there are no abnormal samples. Therefore, useful information can be obtained from the distribution according to the target value of each sample, and then the abnormal samples in the sample set can be identified from it. This method is simple in process and can effectively identify the abnormal samples among them.

[0013] As a specific embodiment of the present invention, based on the spectral data and concentration values of other samples in the sample set, a projection matrix corresponding to the sample is constructed, including:

[0014] Reconstruct the first spectral matrix according to the spectral data of other samples in the sample set. In the first spectral matrix, the elements in the same row represent the spectral data of the same sample;

[0015] Determine the first concentration vector according to the first spectral matrix and the concentration values of other samples in the sample set;

[0016] Determine the second spectral matrix according to the first spectral matrix and the first concentration vector, where the second spectral matrix is used to represent the space composed of other information orthogonal to the subspace of the component to be measured;

[0017] Construct the projection matrix corresponding to the sample based on the second spectral matrix and the generalized inverse matrix of the second spectral matrix.

[0018] As a specific embodiment of the present invention, reconstructing the first spectral matrix according to the spectral data of other samples in the sample set includes:

[0019] Perform singular value decomposition on the residual matrix composed of the spectral data of other samples in the sample set, and use the first p principal components obtained by the decomposition to perform spectral reconstruction on the residual matrix to obtain the first spectral matrix.

[0020] As a specific embodiment of the present invention, the expression of the first concentration vector is:

[0021]

[0022] Wherein, represents the first concentration vector, and X SR represents the first spectral matrix, represents the generalized inverse matrix of the first spectral matrix, represents the concentration vector composed of the concentration values of other samples in the sample set.

[0023] As a specific embodiment of the present invention, the expression of the projection matrix is:

[0024] H = I - X SR,-k (X SR,-k ) +

[0025] Wherein, H represents the projection matrix, I represents the identity matrix, and X SR,-k represents the second spectral matrix, and (X SR,-k ) + represents the generalized inverse matrix of the second spectral matrix.

[0026] As a specific embodiment of the present invention, for each sample in the sample set, the respective target values of the spectral data of the sample under the projection matrices corresponding to the other samples in the sample set are respectively obtained, including:

[0027] For each sample in the sample set, the spectral data of the other samples in the sample set are respectively projected based on the projection matrix corresponding to the sample, so as to obtain the respective pure spectral signals of the other samples in the sample set under the projection matrix corresponding to the sample;

[0028] Determine the 2-norm of the respective pure spectral signals of the samples in the sample set as their target values.

[0029] As a specific embodiment of the present invention, for each sample in the sample set, the respective target values of the spectral data of the sample under the projection matrices corresponding to the other samples in the sample set are respectively obtained, including;

[0030] For each sample in the sample set, the elements of each row in the residual matrix corresponding to the sample are respectively projected by using the projection matrix corresponding to the sample, so as to obtain the respective pure spectral signals of the elements of each row in the residual matrix corresponding to the sample under the projection matrix corresponding to the sample, and determine the 2-norm of the respective pure spectral signals as the target values of the other samples under the projection matrix corresponding to the sample; wherein, the residual matrix is composed of the spectral data of the other samples in the sample set.

[0031] As a specific embodiment of the present invention, determining the abnormal samples in the sample set based on the target values of each sample in the sample set includes:

[0032] For each sample in the sample set, determining the mean and standard deviation of the target values of the sample, and determining the threshold of the sample according to the mean and standard deviation of the sample;

[0033] For each sample in the sample set, determining the abnormal feature value of the sample according to the number of times the target values of the sample exceed the threshold of the sample;

[0034] Determining that the sample is an abnormal sample according to the abnormal feature value of the sample exceeding the set value.

[0035] As a specific embodiment of the present invention, determining the abnormal feature value of the sample according to the number of times the target values of the sample exceed the threshold of the sample includes:

[0036] Determining the number of times the target values of the sample exceed the threshold of the sample;

[0037] Based on the ratio of the number determined in the previous step to the number of samples in the sample set, determining the abnormal feature value.

[0038] As a specific embodiment of the present invention, obtaining the spectral data of the components to be measured of each sample in the sample set includes:

[0039] Respectively obtaining the near-infrared spectral data of the components to be measured of each sample in the sample set, and constructing a spectral matrix of the sample set according to the obtained near-infrared spectral data; in this spectral matrix, the elements in the same row represent the spectral data of the same sample at different wavenumber points, and the elements in the same column represent the spectral data of each sample at the same wavenumber point.

[0040] As a specific embodiment of the present invention, it further includes: preprocessing the spectral matrix, and the preprocessing includes standard normal variate transformation or multiplicative scatter correction. Description of the Drawings

[0041] Figure 1 Showing the flow of the method for identifying abnormal samples in the sample set provided by the embodiment of the present invention Figure 1 ;

[0042] Figure 2 Showing the flow of the method for identifying abnormal samples in the sample set provided by the embodiment of the present invention Figure 2 ;

[0043] Figure 3 Showing the schematic diagram of the near-infrared original spectra of each tobacco leaf sample in the tobacco leaf sample set provided by the embodiment of the present invention;

[0044] Figure 4Show the near-infrared original spectra of each tobacco leaf sample in the preprocessed tobacco leaf sample set provided by the embodiments of the present invention;

[0045] Figure 5 Show the pure spectral signals of each tobacco leaf sample in the tobacco leaf sample set based on total alkaloids provided by the embodiments of the present invention;

[0046] Figure 6 Show the normal distribution fitting histogram of the 2-norm of the pure spectral signal corresponding to one tobacco leaf sample in the tobacco leaf sample set provided by the embodiments of the present invention;

[0047] Figure 7 Show the 2-norm and threshold corresponding to one tobacco leaf sample in the projection matrices corresponding to other samples in the tobacco leaf sample set provided by the embodiments of the present invention;

[0048] Figure 8 Show the occurrence probability distribution diagram of the abnormal eigenvalues of each tobacco leaf sample in the tobacco leaf sample set provided by the embodiments of the present invention. Detailed implementation manners

[0049] The following specific embodiments illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. Although the description of the present invention will be introduced in conjunction with the preferred embodiments, this does not mean that the features of this invention are limited to this implementation manner. On the contrary, the purpose of introducing the invention in conjunction with the implementation manner is to cover other alternatives or modifications that may be extended based on the claims of the present invention. In order to provide a deep understanding of the present invention, many specific details will be included in the following description. The present invention can also be implemented without these details. In addition, in order to avoid confusing or obscuring the key points of the present invention, some specific details will be omitted in the description. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0050] It should be noted that in this specification, similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0051] To make the purpose, technical solutions and advantages of the present invention clearer, the implementation manners of the present invention will be further described in detail below with reference to the drawings.

[0052] A specific implementation manner of the present invention provides a method for identifying abnormal samples in a sample set. An abnormal sample refers to a sample in which the spectrum and content of the component to be measured fall outside the overall. The reasons for abnormal samples may be changes in sample properties, errors generated during data measurement or recording, or the existence of samples with properties completely different from the overall. Specifically, such asFigure 1 As shown in the figure, the method includes the following steps:

[0053] S101: Obtain the spectral data and concentration values of the components to be measured for each sample in the sample set.

[0054] Specifically, the spectral data can be near-infrared spectral data, and the components to be measured can be chemical components of tobacco or tobacco products, such as total alkaloids, total sugars, total nitrogen, etc.

[0055] In this step, first, obtain the near-infrared spectral data of each sample in the sample set respectively, and then construct the spectral matrix X of the sample set according to the obtained near-infrared spectral data; in this spectral matrix X, the elements in the same row represent the spectral data of the same sample at different wavenumber points, and the elements in the same column represent the spectral data of each sample at the same wavenumber point.

[0056] Suppose the number of samples in the sample set is m, and the spectrum of each sample consists of n variables, then the spectral matrix corresponding to this sample set is an m×n matrix, that is, this matrix has m rows and n columns, and n corresponds to the wavenumber point, that is, the matrix elements in the same row correspond to the spectral information of n wavenumber points of the same sample.

[0057] Figure 3 Shows the near-infrared spectra of each sample in a possible implementation manner of the present application, as Figure 3 shown. Each curve in the figure represents the spectrum of a sample, and its wavenumber is 7500 cm -1 -4000 cm -1 . Figure 3 The size m of the spectral matrix X composed of the near-infrared spectra shown is 209 samples, and n is 1557 variables. For example, the first point corresponds to the wavenumber 3999.6 cm -1 , the second point corresponds to the wavenumber 4003.5 cm -1 , and the nth point corresponds to 10001 cm -1 .

[0058] S102: Preprocess the spectral matrix X. The preprocessing includes standard normal variate transformation (SNV) or multiplicative scatter correction (MSC). Its function is to eliminate the influence of the size of sample solid particles, surface scattering, and optical path changes on the spectrum, and improve the accuracy of subsequent identification.

[0059] S103: Based on the spectral data and concentration values of the components to be measured for each sample, solve the pure spectral signal of each sample, and determine each target value of each sample according to the pure spectral signal.

[0060] Taking the example that the sample set has m samples, during the interactive detection process, the remaining m - 1 samples in the sample set are often used to build a model to predict the remaining i-th sample. Therefore, in a possible implementation manner of the present application, the pure spectral signals of the samples can be obtained by constructing pure spectral signals under different sets through the elimination method. Assuming that the elimination method is executed multiple times, each sample will obtain a pure spectral signal. If there are abnormal samples, the pure spectral signal will be different from the pure spectral signal generated when there are no abnormal samples. Through calculation, the pure spectral signal of each sample can be obtained, and then useful information can be obtained from the distribution of the pure spectral signals of each sample.

[0061] Specifically, in step S103, first, for each sample in the sample set, a projection matrix H corresponding to the sample is constructed based on the spectral data and concentration values of the other samples in the sample set; then, according to the projection matrix H of the sample, the spectral data of the other samples in the sample set are processed respectively to obtain the pure spectral signals of the other samples in the sample set under the projection matrix H.

[0062] That is, the projection matrix corresponding to each sample in the sample set is constructed based on the other remaining samples after removing this sample from the sample set. That is, for sample i, its corresponding projection matrix is constructed based on all other samples except sample i in the sample set, and sample i does not participate in the construction process of its corresponding projection matrix.

[0063] Further, as Figure 2 shown, step S103 may include:

[0064] S1031: Reconstruct the first spectral matrix X according to the spectral data of the other samples in the sample set SR , in the first spectral matrix X SR , the elements in the same row represent the spectral data of the same sample.

[0065] Specifically, according to the residual matrix X S composed of the spectral data of the other samples in the sample set, perform singular value decomposition [u, s, v] = svd(X S ), and use the first p principal components obtained from the decomposition to perform spectral reconstruction on the residual matrix to obtain the first spectral matrix X SR = u p * s p * v p '. Using the principal component spectrum for reconstruction can reduce the influence of errors such as the background and noise of the measured spectrum. It is a relatively mature technology and will not be elaborated here.

[0066] It should be noted that the residual matrix X SIt refers to the matrix obtained by excluding the target samples from the spectral matrix X of the original sample set. Taking m samples and n variables as an example, the representation form of the spectral matrix X of this sample set is as follows:

[0067]

[0068] Among them, the data element x at the i-th row and j-th column in this matrix ij represents the spectral data of the i-th sample in the j-th wave point of this sample set. When calculating the projection matrix corresponding to the a-th sample, it is necessary to first remove each element in the a-th row from the spectral matrix X to obtain the residual matrix X S corresponding to the a-th sample. For example, when a = 1, that is, for the first sample, the corresponding residual matrix X S is:

[0069]

[0070] Similarly, for the second sample, the corresponding residual matrix X S is:

[0071]

[0072] S1032: Determine the first concentration vector SR according to the first spectral matrix X

[0073] and the concentration values of other samples in the sample set Specifically, the expression of the first concentration vector

[0074]

[0075] is: represents the generalized inverse matrix of the first spectral matrix, represents the concentration vector composed of the concentration values of the components to be measured of other samples in the sample set.

[0076] S1033: Determine the second spectral matrix X SR according to the first spectral matrix X and the first concentration vector SR,-k where the second spectral matrix X SR,-k can be used to characterize the space composed of other information orthogonal to the subspace of the component to be measured.

[0077] Specifically, the expression of the second spectral matrix X SR,-k is:

[0078]

[0079] Among them, XSR is the first spectral matrix corresponding to the sample, which represents the spectral information of the component to be measured and can be replaced by the average spectral table of the first spectral matrix X SR . Further, the average spectrum of the first spectral matrix X SR is obtained by adding up all the column elements of the matrix X SR and then dividing by the number of samples. α is a scalar with a value of

[0080] S1034: Construct the projection matrix H corresponding to the sample based on the second spectral matrix and the generalized inverse matrix of the second spectral matrix. Specifically, the expression of the projection matrix H is:

[0081] H = I - X SR,-k (X SR,-k ) +

[0082] where I represents the identity matrix, and (X SR,-k ) + represents the generalized inverse matrix of the second spectral matrix.

[0083] S1035: Project each row element in the residual matrix X S corresponding to the sample according to the projection matrix H corresponding to the sample, so as to obtain the pure spectral signals S of each row element in the residual matrix X

[0084] corresponding to the sample under the projection matrix H corresponding to the sample. Specifically, the pure spectral signal of the sample can be expressed in vector form, and its expression is: where represents the vector composed of each element in the i-th row of the residual matrix X S .

[0085] S1036: Calculate the 2-norm r of each of the above-obtained pure spectral signals i of the samples. Specifically, Furthermore, the full-spectrum vector spectral information of other samples in the sample set is simplified to a determined scalar information r i . Repeat the above steps S1031 to S1036, that is, exclude the first sample, the second sample, until the last sample in turn, calculate the 2-norm of the pure spectral signal of each sample, and record it as matrix R. Among them, the first row of matrix R represents the 2-norm of the pure spectral signal of the first sample in multiple calculations, the second row represents the 2-norm of the pure spectral signal of the second sample in multiple calculations, and so on.

[0086] S104: Determine the abnormal samples in the sample set based on the target values of each sample in the sample set.

[0087] It should be noted that each target value of each sample is the 2-norm r of the pure spectral signal of the sample. of the pure spectral signal i . Specifically, step S104 includes: for each sample in the sample set, determine the mean and standard deviation of the target values of the sample, and determine the threshold of the sample according to the mean and standard deviation of the sample; for each sample in the sample set, determine the abnormal characteristic value of the sample according to the number of times that the target values of the sample exceed the threshold of the sample; determine that the sample is an abnormal sample according to the abnormal characteristic value of the sample exceeding the set value.

[0088] Specifically, the abnormal characteristic value of each sample is equal to the ratio of the number of times that the target values of the sample exceed the threshold of the sample to the number of samples in the sample set. For example, if the total number of samples in the sample set is m, and for the i-th sample, the number of times that its target values exceed the threshold of the i-th sample is a, then the abnormal characteristic value of the i-th sample is a / m.

[0089] Furthermore, the determination process of abnormal samples can also be based on matrix data. As mentioned above, the 2-norms of the pure spectral signals of each sample can be constructed into a matrix R. Among them, each element in the i-th row of the matrix R represents the 2-norms of the pure spectral signals of the i-th sample in the sample set in multiple calculations (it should be noted that taking the b-th sample as an example, since for this sample, only the pure spectral signals of the sample under the projection matrices corresponding to other samples are calculated, and the pure spectral signal of the sample under its corresponding projection matrix is not calculated, that is, for the b-th sample, the corresponding element r bb in the matrix R is not calculated. For the convenience of calculation, this element is defined as 0, that is, in the matrix R, r ii =0, where i = 1, 2, 3...). Taking the matrix R represented as follows as an example:

[0090]

[0091] When determining abnormal samples based on the above matrix R, first calculate the mean and standard deviation of the elements in each row of the matrix as the mean and standard deviation of the corresponding sample in that row. For example, the first row in the matrix corresponds to the first sample in the sample set. Therefore, calculate the mean and standard deviation of the elements in the first row of the matrix as the mean and standard deviation of the first sample in the sample set, and take "mean ± 3 * standard deviation" as the threshold for the first sample. Similarly, calculate the mean and standard deviation of the elements in the second row of the matrix as the mean and standard deviation of the second sample in the sample set, and then take "the mean of the second sample ± 3 * the standard deviation of the second sample" as the threshold for the second sample, and so on until the m-th sample.

[0092] Then, compare each element in the first row with the threshold of the first row respectively to determine the number of times a1 that the element value is greater than the threshold, and take the ratio of this number a1 to the total number of samples m, i.e., a1 / m, as the abnormal feature value of the first sample. Determine whether the abnormal feature value of the first sample is greater than the set value. If it is greater than the set value, then consider this sample as an abnormal sample; otherwise, this sample is not an abnormal sample. Similarly, compare each element in the second row with the threshold of the second row to determine the number of times a2 that the element value is greater than the threshold of the second row, and take the ratio of this number a2 to the total number of samples m, i.e., a2 / m, as the abnormal feature value of the first sample. Determine whether the abnormal feature value of the second sample is greater than the set value. If it is greater than the set value, then consider this sample as an abnormal sample; otherwise, this sample is not an abnormal sample. The same method can be used to determine whether the remaining samples are abnormal samples, which will not be elaborated here one by one.

[0093] Specifically, the set value can be set according to the actual situation. For example, it can be set to 10%, that is, if the abnormal feature value of a certain sample exceeds 10%, then it can be defined as an abnormal sample.

[0094]

Embodiment

[0095] Taking the identification of abnormal samples from tobacco leaf samples and then establishing a calibration model as an example, the specific process of the identification method of the present application will be described below. It should be noted that in this embodiment, the components to be measured in the cigarette cut tobacco samples are the contents of total plant alkaloids and total sugars.

[0096] Step 1): Treatment of tobacco leaf samples

[0097] 209 tobacco leaf samples were provided by China Tobacco Guizhou Industrial Co., Ltd. and originated from production areas in Guangdong, Henan, Heilongjiang, Hunan, Liaoning, Shaanxi, Sichuan, and Yunnan provinces. Before collecting the spectra of the samples, the tobacco leaf samples were placed in an oven at 40 °C for two hours according to "YCT 31-1996 Tobacco and Tobacco Products - Preparation of Test Samples and Determination of Moisture - Oven Method"; then the samples were taken out and cooled to room temperature. The samples were poured into a plant grinder for pulverization and then passed through a 40-mesh sieve to separate the samples with a particle size less than 40 mesh (≤0.45 mm). After the samples were cooled to room temperature, they were put into disposable sealed bags and stored at low temperature and in the dark.

[0098] Step 2): Experimental Instruments and Spectrum Collection

[0099] The laboratory temperature was controlled between 22 ± 2 °C, and the relative humidity was controlled between 40% ± 10%. The near-infrared instrument was turned on and preheated for at least 1 hour, and then used after being calibrated and qualified by the ValPro program. An appropriate amount of the prepared tobacco leaf powder was put into a sample cup for scanning, with a scanning range of 4000 - 10000 cm -1 , and the resolution was 8 cm -1 ; the number of scans was 64 times.

[0100] Step 3): Determination of Chemical Values

[0101] According to the tobacco industry standard, the content of the total alkaloid chemical component in the tobacco leaf samples was determined. "YC / T 468-2013 Tobacco and Tobacco Products - Determination of Total Alkaloids - Continuous Flow (Potassium Thiocyanate) Method".

[0102] Step 4): Identification of Abnormal Samples

[0103] The near-infrared spectra of the tobacco leaf powder samples were as Figure 3 shown, and the drift range of the sample spectra was relatively large. The standard normal variate transformation method was selected to preprocess the spectra, and the Figure 4 shown spectrogram was obtained. According to the above calculation process, the pure spectral signals of the samples (as Figure 5 shown) were obtained, and their norms and thresholds were calculated. Among them, Figure 6 shows the normal distribution fitting histogram of the 2-norm of the pure spectral signals corresponding to one tobacco leaf sample in the projection matrices corresponding to other samples, Figure 7 shows the 2-norm and threshold corresponding to one tobacco leaf sample.

[0104] Figure 8The distribution diagram of the abnormal characteristic values of each sample in the tobacco leaf sample set is shown. Specifically, when there are no abnormal samples, the 2-norms of the spectral pure signals of each sample in the calibration set are close. Only when abnormal samples are included, the 2-norms of their spectral pure signals will be quite different from the values obtained without including abnormal samples. It is precisely through this numerical difference that abnormal samples are identified. For example Figure 6 In Figure 6 , the mean of the pure spectral signals of the samples is 0.4285, and the standard deviation is 0.0019. Therefore, the calculated range of the norm values of the spectral pure signals is between 0.4229 and 0.4341. Among them, the norms of samples No. 5, 9, 45, 48, 49, 51, and 95 are all lower than 0.4229, indicating that the appearance of these abnormal samples causes a large change in the norm values.

[0105] Such as Figure 8 As shown in Figure 8 , for the total alkaloid, the target values of 16 samples exceed the range. Through calculation, the abnormal characteristic values of a total of 8 samples, namely sample numbers 5, 9, 45, 48, 49, 51, 95, and 140, exceed the range (the set value of the range is 10%), so they are identified as abnormal samples.

[0106] Before removing the abnormal samples, the determination coefficient of the near-infrared quantitative model constructed using the above samples is 0.964. The root mean square error of prediction (RMSEC) of the calibration set, the root mean square error of prediction (RMSECV) of the cross-validation set, and the root mean square error of prediction (RMSEP) of the prediction set are 0.0815, 0.0907, and 0.0998 respectively. After removing the abnormal samples, the determination coefficient of the constructed model is increased to 0.966, and the corresponding RMSEC, RMSECV, and RMSEP values are reduced to 0.0788, 0.0847, and 0.0858. The value of the model robustness evaluation parameter SEP / SEC generally needs to be less than 1.2 to indicate that the model has good robustness for the samples to be measured. The calculated SEP / SEC values before and after removing the abnormal samples are 1.22 and 1.09 respectively, indicating that the near-infrared calibration model established by removing the abnormal samples through this method is more robust and accurate.

[0107] The identification method of the embodiment of the present application can obtain the pure spectral signal of the sample based on the content of the component to be measured, calculate the abnormal characteristic value of each sample according to the 2-norm of the pure spectral signal, and then judge the abnormal sample, so as to ensure the accuracy of the established calibration model and improve its prediction ability.

[0108] Although the present invention has been illustrated and described by reference to the embodiments thereof, those of ordinary skill in the art should understand that the above content is a further detailed description of the present invention in conjunction with specific embodiments, and it cannot be determined that the specific implementation of the present invention is limited only to these descriptions. Those skilled in the art can make various changes in form and detail, including making several simple deductions or substitutions without departing from the spirit and scope of the present invention.

Claims

1. A method for identifying abnormal samples in a sample set, characterized in that, Including: Obtaining spectral data and concentration values of the components to be measured for each sample in the sample set; For each sample in the sample set, constructing a projection matrix corresponding to the sample based on the spectral data and concentration values of other samples in the sample set; For each sample in the sample set, respectively obtaining each target value of the spectral data of the sample under the projection matrices corresponding to other samples in the sample set, including: for each sample in the sample set, using the projection matrix corresponding to the sample to project each row element in the residual matrix corresponding to the sample, so as to obtain each pure spectral signal of each row element in the residual matrix corresponding to the sample under the projection matrix corresponding to the sample, and determining the 2-norm of each pure spectral signal as the target value of other samples under the projection matrix corresponding to the sample; wherein, the residual matrix is composed of the spectral data of other samples in the sample set; Based on the target values of each sample in the sample set, determining the abnormal samples in the sample set.

2. The method according to claim 1, wherein Constructing a projection matrix corresponding to the sample based on the spectral data and concentration values of other samples in the sample set, including: Reconstructing a first spectral matrix according to the spectral data of other samples in the sample set, in the first spectral matrix, the elements in the same row represent the spectral data of the same sample; Determining a first concentration vector according to the first spectral matrix and the concentration values of other samples in the sample set; Determining a second spectral matrix according to the first spectral matrix and the first concentration vector, wherein the second spectral matrix is used to represent the space composed of other information orthogonal to the subspace of the component to be measured; Constructing the projection matrix corresponding to the sample based on the second spectral matrix and the generalized inverse matrix of the second spectral matrix.

3. The method according to claim 2, wherein Reconstructing a first spectral matrix according to the spectral data of other samples in the sample set, including: Performing singular value decomposition on the residual matrix composed of the spectral data of other samples in the sample set, and using the first p principal components obtained by the decomposition to perform spectral reconstruction on the residual matrix to obtain the first spectral matrix.

4. The method according to claim 2, wherein The expression of the first concentration vector is: Among them, represents the first concentration vector, X SR represents the first spectral matrix, represents the generalized inverse matrix of the first spectral matrix, represents the concentration vector composed of the concentration values of other samples in the sample set.

5. The method according to claim 2, characterized in that, The expression of the projection matrix is: H = I - X SR,-k (X SR,-k ) + Among them, H represents the projection matrix, I represents the identity matrix, and X SR,-k represents the second spectral matrix, and (X SR,-k ) + represents the generalized inverse matrix of the second spectral matrix.

6. The method according to claim 1, wherein Based on the target values of each sample in the sample set, determining the abnormal samples in the sample set, including: For each sample in the sample set, determining the mean value and standard deviation of each target value of the sample, and determining the threshold value of the sample according to the mean value and the standard deviation of the sample; For each sample in the sample set, determining the abnormal characteristic value of the sample according to the number of times that each target value of the sample exceeds the threshold value of the sample; Determining that the sample is an abnormal sample according to that the abnormal characteristic value of the sample exceeds the set value.

7. The method according to claim 6, wherein Determining the abnormal characteristic value of the sample according to the number of times that each target value of the sample exceeds the threshold value of the sample, including: Determining the number of times that each target value of the sample exceeds the threshold value of the sample; Determining the abnormal characteristic value based on the ratio of the number determined in the previous step to the number of samples in the sample set.

8. The method according to claim 1, wherein Obtain the spectral data of the components to be measured for each sample in the sample set, including: Respectively obtain the near-infrared spectral data of the components to be measured for each sample in the sample set, and construct the spectral matrix of the sample set according to the obtained near-infrared spectral data of each sample; in the spectral matrix, the elements in the same row represent the spectral data of the same sample at different wavenumber points, and the elements in the same column represent the spectral data of each sample at the same wavenumber point.

9. The method according to claim 8, wherein It also includes: Perform preprocessing on the spectral matrix, and the preprocessing includes standard normal variate transformation or multiplicative scatter correction.

Citation Information

Patent Citations

  • Unknown pollutant early-warning method based on ultraviolet-visible spectrum

    CN103776789A

  • Qualitative analysis method for improving identification result on basis of near-infrared mode

    CN104374738A