A method for evaluating weak difference signals of near infrared spectroscopy data

By preprocessing near-infrared spectral data using standard normal transformation, first-order differentiation, and the maximum-minimum rule, and combining multiple similarity calculation methods, the problem of difficulty in identifying subtle spectral differences in existing technologies has been solved, enabling accurate classification of tobacco leaf samples.

CN116432051BActive Publication Date: 2025-12-19CHINA TOBACCO YUNNAN IND
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310560109.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-18
Publication Date
2025-12-19
Estimated Expiration
2043-05-18

AI Technical Summary

Technical Problem

Existing near-infrared spectroscopy technology, when subjected to improper preprocessing or unsuitable algorithms, cannot effectively identify and reflect subtle differences in the spectra of tobacco leaves, resulting in an inability to accurately distinguish between different categories of test samples.

Method used

Near-infrared spectral data are preprocessed using standard normal transformation, first-order differentiation, and the maximum-minimum rule. Spectral similarity is calculated by combining Euclidean distance, correlation coefficient, and information divergence to eliminate scattering effects, noise, and dimensional differences, thereby enhancing spectral comparability.

Benefits of technology

It effectively identifies subtle signal differences between near-infrared spectra, enabling accurate differentiation of different types of samples and improving the comparability and difference mining capabilities of spectral data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116432051B_ABST
    Figure CN116432051B_ABST
Patent Text Reader

Abstract

The application discloses a kind of near infrared spectrum data weak difference signal evaluation method, namely SSMS (Standard normal variate transform+Savitzky golay+Min max+Spectral similarity) method.The method is scattered correction to near infrared spectrum data using standard normal variate transform, eliminate the scattering influence caused by uneven distribution of sample;Using first derivative to remove noise in spectrum, improve the signal-to-noise ratio of spectrum and enhance the distinguish degree of overlapping peak;Using maximum minimum rule method, eliminate the dimension of spectrum and enhance data comparability;Finally, the similarity of near infrared spectrum data is evaluated by combining euclidean distance, correlation coefficient, divergence and other evaluation information.The application can effectively identify the weak signal difference of near infrared spectrum, and then accurately distinguish different categories of test samples, and can be used as an effective tool for accurately identifying and distinguishing differences between test samples using near infrared technology.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the field of near-infrared spectrum qualitative analysis, and particularly relates to a method for evaluating weak difference signals of near-infrared spectrum data. BACKGROUND

[0002] The near-infrared technology has been widely applied due to its advantages of rapidness, low cost, high precision and the like. However, the near-infrared spectrum needs to be properly pretreated so as to effectively identify the overall information of various chemical components of tobacco leaves and interpret the overall difference and similarity of the tobacco leaves, which is influenced by spectral peak overlap, background noise, baseline drift and the like.

[0003] If the selected pretreatment mode is improper, the subtle difference between the near-infrared spectrums cannot be mined, and if the algorithm for calculating the similarity value between the near-infrared spectrums is improper, the final near-infrared similarity value cannot reflect the subtle difference between the near-infrared spectrums. SUMMARY

[0004] The application develops a method for evaluating weak difference signals of near-infrared spectrum data, namely, the SSMS method. The method adopts the standard normal variable transformation to perform scattering correction on the near-infrared spectrum data, eliminates the scattering influence caused by uneven distribution of samples, adopts the first-order derivation to remove the noise in the spectrum, improves the signal-to-noise ratio of the spectrum and enhances the distinguishability of the overlapping peaks, adopts the maximum-minimum rule method to eliminate the dimension of the spectrum and enhance the comparability of the data, and finally combines the Euclidean distance, the correlation coefficient, the divergence and the like to evaluate the similarity of the near-infrared spectrum data. The application can effectively identify the weak signal difference between the near-infrared spectrums, and further accurately distinguish different categories of detection samples, and can be used as an effective tool for accurately identifying and distinguishing the difference between the detection samples by using the near-infrared technology.

[0005] To achieve the above object, the application adopts the following technical scheme:

[0006] A method for evaluating weak difference signals of near-infrared spectrum data, comprising the following steps:

[0007] Step 1: respectively performing infrared spectrum determination on sample A and sample B to obtain two pieces of near-infrared spectrum data;

[0008] Step 2: respectively performing scattering correction on the two pieces of near-infrared spectrum data by using the standard normal variable transformation to eliminate the scattering influence caused by uneven distribution of samples;

[0009] Step 3: respectively performing noise processing on the two pieces of near-infrared spectrum data which have completed the scattering correction by using the first-order derivation method to remove the noise in the spectrum, improve the signal-to-noise ratio of the spectrum and enhance the distinguishability of the overlapping peaks;

[0010] Step 4: The maximum-minimum rule method is used to normalize the two near-infrared spectral data after removing noise in the spectrum, to enhance the comparability of the data.

[0011] Step 5: The similarity of the two near-infrared spectral data is calculated by combining the Euclidean distance, correlation coefficient, and information divergence.

[0012] Further, the specific method of step 1 is as follows:

[0013] The near-infrared spectral data of sample A and sample B are collected and denoted as SpecA and SpecB, respectively, as shown in equations (1) and (2):

[0014]

[0015]

[0016] where m is the number of wavelength points, represents the absorbance of the i-th wavelength point of the near-infrared spectral data SpecA of sample A, represents the absorbance of the i-th wavelength point of the near-infrared spectral data SpecB of sample B.

[0017] Further, the specific method of step 2 is as follows:

[0018] The standard normal variable transformation method is used to eliminate the influence of sample particle surface scattering and optical path changes on near-infrared diffuse reflectance spectra during near-infrared spectral collection. Unlike the standardization algorithm, the standard normal variable transformation method can process a single spectrum.

[0019] The standard normal variable transformation processing method for the near-infrared spectral data SpecA of sample A is as follows:

[0020]

[0021] where, represents the absorbance of the i-th wavelength point of the near-infrared spectral data SpecA of sample A, represents the value of the near-infrared spectral data SpecA after standard normal variable transformation processing, is the average value of all wavelength points of the near-infrared spectral data SpecA, and m is the number of wavelength points.

[0022] The spectral data of the near-infrared spectral data SpecA of sample A after standard normal variable transformation processing is represented as follows:

[0023]

[0024] According to the same procedure, the spectrum data of sample B after the standard normal variable transformation is shown as follows:

[0025]

[0026] Further, the specific method of step 3 is as follows:

[0027] For the near-infrared data using the standard normal variable transformation to eliminate the effect of near-infrared diffuse reflection, the first derivative method is used to smooth filter the near-infrared spectrum data, reduce the interference of noise data, and the first derivative method is based on the improvement of moving smoothing algorithm;

[0028] wherein the data SpecA of sample A after the standard normal variable transformation processing method is SpecA 1 The specific process of denoising is as follows:

[0029] The filter window length 2k+1 (k is a constant, generally the wavelength point number m in the spectrum is ≤2000, and the value of k=5; when the wavelength point number m in the spectrum is >2000, the value of k=8) is set, and for the wavelength point absorbance 1 in the near-infrared spectrum data SpecA The filter window is expressed as

[0030]

[0031] Wherein a=min(i-k,0), b=min(i+k,m), and l=b-a represents the number of spectral measurement points in the filter window ;

[0032] The k-1 order polynomial shown in formula (7) is used to fit the l data points:

[0033]

[0034] Wherein j=(a,a+1,…,b);

[0035] For each spectral measurement point in the filter window , the equation is constructed based on formula (7), and finally a k-linear equation group composed of l equations is constructed, and after fitting the k-linear equation group by the least square method, the parameters A={a'0,a'1,…,a' k-1} of the polynomial are determined, and the absorbance of the wavelength point in SpecA 1 is filtered by the following formula:

[0036]

[0037] Absorbance of wavelength point in near infrared spectrum data SpecA 1 Absorbance of wavelength point in near infrared spectrum data SpecA After the above processing, the near infrared spectrum data SpecA 1 The spectrum data after smoothing filtering is shown in the following formula.

[0038]

[0039] At this point, SpecA 2 is the near infrared spectrum data SpecA of sample A after the standard normal variable transformation processing method and noise processing;

[0040] Through the above same process, SpecB 1 is denoised to obtain the near infrared spectrum data SpecB 2 .

[0041] Further, the specific method of step 4 is as follows:

[0042] For the near infrared spectrum data after the near infrared diffuse reflection effect is eliminated by using the standard normal variable transformation and the noise in the spectrum is removed by using the first derivative, the maximum and minimum rule is used to eliminate the dimension of the spectrum to enhance the comparability between the spectra;

[0043] The near infrared spectrum data SpecA 2 is an example, and the specific processing process of the maximum and minimum rule for eliminating the dimension of the spectrum is shown in the following formula:

[0044]

[0045] Among them, is the absorbance of the i-th (i=1, 2, …, m) wavelength point in SpecA 2 which has eliminated the dimension of the spectrum by the maximum and minimum rule, is the absorbance of the i-th (i=1, 2, …, m) wavelength point in SpecA 2 , and m is the number of wavelength points of SpecA 2 . 2

[0046] SpecA 3 is the near infrared spectrum data SpecA of sample A after the standard normal variable transformation processing method, the first derivative denoising operation and the maximum and minimum rule for eliminating the dimension;

[0047] Through the above same process, SpecB 2 ​After the de-dimension processing, the near-infrared spectrum data SpecB is obtained 3 ;

[0048] Further, the step 5 is specifically as follows:

[0049] The specific method of calculating the similarity of the near-infrared spectrum data in combination with the Euclidean distance, the correlation coefficient and the information divergence is as follows:

[0050] wherein, SpecA 3 is the near-infrared spectrum data SpecA of the sample A after the standard normal variable transformation processing method, the first-order derivative denoising operation and the maximum-minimum rule dimension elimination, SpecB 3 is the near-infrared spectrum data SpecB of the sample B after the standard normal variable transformation processing method, the first-order derivative denoising operation and the maximum-minimum rule dimension elimination;

[0051] Firstly, in the calculation of the Euclidean distance of the two near-infrared spectrums SpecA 3 and SpecB 3 , the distance size of the two near-infrared spectrum vectors is calculated in the Euclidean space according to the following formula:

[0052]

[0053] wherein, EDM(SpecA 3 , SpecB 3 ) represents the Euclidean distance value of the near-infrared spectrum data SpecA 3 and SpecB 3 , represents the absorbance of the i-th (i=1, 2, …, m) wavelength point of the near-infrared spectrum data SpecA 3 , represents the absorbance of the i-th (i=1, 2, …, m) wavelength point of the near-infrared spectrum data SpecB 3 , and m is the wavelength point number of the near-infrared spectrum data SpecA 3 and SpecB 3 ;

[0054] Secondly, in the calculation of the correlation coefficient of the two near-infrared spectrums, the correlation of the two near-infrared spectrum vectors is calculated through the following formula:

[0055]

[0056] wherein, SCM(SpecA 3 , SpecB 3 ) is the correlation coefficient of the near-infrared spectrum data SpecA 3 and SpecB3 The correlation coefficient, SpecA represents near-infrared spectral data 3 The absorbance at the i-th wavelength (i = 1, 2, ..., m) SpecB represents near-infrared spectral data 3 The absorbance at the i-th (i = 1, 2, ..., m) wavelength point, where m is the near-infrared spectral data SpecA. 3 and SpecB 3 The number of wavelength points, Near-infrared spectral data SpecA 3 and SpecB 3 Average absorbance;

[0057] Then, in calculating SpecA 3 and SpecB 3 When considering the scattering information of two near-infrared spectra, based on information measurement theory, SpecA... 3 and SpecB 3 The two near-infrared spectra are each considered as information elements with probabilistic statistical characteristics, and the absorbance probability of each wavenumber in the two spectra is described by the following formula:

[0058]

[0059]

[0060] in, For SpecA 3 The probability value of absorbance at the i-th wavelength point (i = 1, 2, ..., m) in the spectrum. For SpecB 3 The probability value of absorbance at the i-th wavelength point (i = 1, 2, ..., m) in the spectrum.

[0061] Accordingly, SpecA 3 and SpecB 3 The formula for calculating the relative entropy of the two near-infrared spectra is as follows:

[0062]

[0063]

[0064] Among them, D(SpecA) 3 ||SpecB 3 ) is SpecA 3 Compared to SpecB 3 The relative entropy, D(SpecB) 3 ||SpecA 3SpecB 3 Relative Entropy with respect to SpecA 3

[0065] According to the relative entropy of SpecA 3 and SpecB 3 , the information divergence of the two spectra is calculated according to the following formula:

[0066] SID(SpecA 3 ,SpecB 3 ) = D(SpecA 3 ||SpecB 3 ) + D(SpecB 3 ||SpecA 3 ) (17);

[0067] Wherein, SID(SpecA 3 ,SpecB 3 ) represents the information divergence of SpecA 3 and SpecB 3 two near infrared spectra;

[0068] As above, EDM(SpecA 3 ,SpecB 3 ) represents the Euclidean distance of near infrared spectrum data SpecA 3 and SpecB 3 , SCM(SpecA 3 ,SpecB 3 ) represents the correlation coefficient of near infrared spectrum data SpecA 3 and SpecB 3 , SID(SpecA 3 ,SpecB 3 ) represents the information divergence of near infrared spectrum data SpecA 3 and SpecB 3 , according to the following formula, to finally represent the similarity of near infrared spectrum data SpecA 3 and SpecB 3 :

[0069]

[0070] Wherein, SS(SpecA 3 ,SpecB 3 ) is the similarity of the two near infrared spectrum data SpecA 3 and SpecB 3 described in the present application, to represent the similarity of sample A and sample B. ​

[0071] Further, the smaller the value of EDM (SpecA 3 , SpecB 3 ) and the larger the value of SCM (SpecA 3 , SpecB 3 ) and the smaller the value of SID (SpecA 3 , SpecB 3 ), the higher the similarity of the samples characterized by the near infrared spectrum data SpecA 3 and SpecB 3 .

[0072] Advantages of the present application:

[0073] The standard normal variable transformation can eliminate the scattering effect caused by uneven distribution of samples in the experiment process, the first-order derivation can effectively remove the high-frequency noise in the spectrum data, and the maximum-minimum rule can eliminate the dimension of the spectrum and enhance the comparability of the data. Through the processing of the standard normal variable transformation, the first-order derivation and the maximum-minimum rule, the noise in the spectrum can be eliminated, the baseline and scattering interference factors in the near infrared spectrum can be excluded, the comparability of the near infrared spectrum can be enhanced, and the subtle differences between the near infrared spectra can be mined. At the same time, the similarity of the near infrared spectrum in the spectrum amplitude, the spectrum form, the spectrum information divergence and other aspects is comprehensively considered. Compared with the similarity calculation method of only single difference information (for example, the Pearson similarity only considers the difference in the spectrum form), the evaluation method of the weak difference signal of the near infrared spectrum data provided by the present application can better reflect the subtle differences between the near infrared spectra. BRIEF DESCRIPTION OF DRAWINGS

[0074] Figure 1 The steps of the evaluation method of the near infrared spectrum data similarity of the present application;

[0075] Figure 2 The near infrared spectrum similarity evaluation method research scheme of Example 1;

[0076] Figure 3 The near infrared spectrum data collected in the experiment in Example 1;

[0077] Figure 4 The near infrared spectrum data after pretreatment by the standard normal variable transformation, the first-order derivation and the maximum-minimum rule method in Example 1;

[0078] Figure 5 The experimental results of different spectrum similarity evaluation methods in Example 1, only the experimental results of 352 in the figure are shown, the average similarity of the samples of the same category is greater than 0.9, and the average similarity of the samples between different categories is less than 0.7;

[0079] Figure 6 The two near-infrared spectrum data examples before pretreatment are shown in the following table:

[0080] Figure 7 The result spectrum data after the spectrum is processed by the standard normal variable transformation, first-order derivation and maximum-minimum rule method in sequence is shown in the following table, which is used for comparison with the near-infrared spectrum without processing, and the mining effect of the pretreatment method on the spectrum subtle difference is shown graphically. DETAILED DESCRIPTION

[0081] The application will be further described in detail below in combination with the drawings and specific embodiments:

[0082] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs.

[0083] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, the singular form is intended to include the plural form unless the context clearly indicates otherwise, and it should also be understood that when the terms "comprise" and / or "include" are used in the specification, there is a feature, step, operation, device, component and / or combination thereof.

[0084] Example 1

[0085] Fifty-one reprocessed tobacco samples of different varieties and parts from different production areas in Yunnan Province were collected, and the experimental samples were divided according to production area, variety and part as shown in the following table.

[0086] Table 1 Classification of experimental samples

[0087] Category Origin, Variety, Part Sample Number Category 1 Baoshan, K326, Middle 4 Category 2 Baoshan, Hongda, Upper 3 Category 3 Honghe, K326, Middle 6 Category 4 Honghe, Yun 87, Upper 5 Category 5 Honghe, Yun 87, Middle 10 Category 6 Kunming, Hongda, Upper 3 Category 7 Kunming, Hongda, Middle 8 Category 8 Qujing, K326, Middle 4 Category 9 Qujing, Yun Series, Middle 4 Category 10 Qujing, Yun Series, Lower 4

[0088] The near-infrared spectrum data of the 51 experimental samples were collected by a thermoelectric antaris II near-infrared spectrometer under the same experimental environment, and the specific sample preparation method and experimental environment are shown in Table 2.

[0089] Table 2 Near-infrared spectrum collection sample preparation specifications and experimental conditions

[0090]

[0091] The experimental scheme is shown in Figure 2 The collected data is shown in Figure 3The collected data is preprocessed by combining the currently commonly used scatter correction, denoising, data enhancement near-infrared spectral pretreatment methods shown in Table 3, and the total near-infrared spectral data pretreatment scheme is 2x4x4=32.

[0092] Table 3 Commonly used near-infrared spectral pretreatment methods

[0093]

[0094] The near-infrared spectral data after pretreatment by the near-infrared data pretreatment scheme (SNV+SG1D+MM) combining standard normal variable transformation, first-order derivation, and maximum-minimum rule is shown in Table 2. Figure 4

[0095] On the basis of spectral data pretreatment, different spectral similarity calculation methods shown in Table 4 are used, and a total of 32x11=352 near-infrared spectral similarity evaluation methods are used for comparative analysis to verify the advancement of the method.

[0096] Table 4 Near-infrared spectral similarity calculation methods

[0097]

[0098]

[0099] The spectral similarity evaluation method Sim i (i=1,2,...,352) is shown as follows:

[0100]

[0101]

[0102] Among them, Within_Category i is the spectral similarity evaluation method Sim i The mean value of the similarity calculation value of the spectrum of all intra-class samples is used to evaluate Sim i The evaluation result of similar tobacco sample spectrum; Between_Categories i is the spectral similarity evaluation method Sim i The mean value of the similarity calculation of all inter-class sample spectra is used to evaluate Sim i The evaluation result of dissimilar tobacco sample spectrum; CLASS represents all 10 sample categories, P and Q are the spectral data corresponding to samples p and samples q, and n and m represent the number of intra-class and inter-class similarity calculation numbers, respectively. ​

[0103] The above-mentioned methods are evaluated, and a similarity calculation scheme capable of better distinguishing similar samples in the same category (Within_Category>0.9) and non-similar samples between different categories (Between_Categories<0.7) is selected, as shown in Table 5 and Figure 5 .

[0104] Table 5 Experimental results of different spectral similarity evaluation methods

[0105] Analysis Scheme Between_Categories Within_Category Sim_Estimate SNV+SG1D+MM+SS 0.6073 0.9051 0.6489 SG1D+MM+SS 0.6188 0.9027 0.6420 SG+MM+ED / COD 0.6744 0.9051 0.6154 MM+ED / COD 0.6744 0.9051 0.6154 MC+ED / COD 0.6926 0.9125 0.6100 SG+MC+ED / COD 0.6926 0.9125 0.6100

[0106] The calculation formula of Sim_Estimate is as follows, which represents the final evaluation result of the spectral similarity evaluation method.

[0107] Sim_Estimate=(Within_Category+(1-Between_Categories)) / 2 (3)

[0108] As shown in Table 5 and Figure 5 , in the similarity analysis scheme of Within_Category=>0.9 and Between_Categories<0.7, which can better distinguish the similarity of tobacco samples in the same type, the Sim_Estimate value of the method studied in the application (i.e., the "SNV+SG1D+MM+SS" scheme of using standard normal variable transformation, first-order derivation, maximum-minimum rule for spectral pretreatment, and combining Euclidean distance, correlation coefficient, and divergence for similarity evaluation) is the largest (0.6489), indicating that the comprehensive performance of the method scheme studied in the application is better, the similarity of samples in the same category is higher, and the similarity of samples between different categories is smaller, which can effectively evaluate the similarity of samples in different categories.

[0109] Example 2

[0110] Example data: actual cigarette formula samples are recorded as A, and sample A contains 17 tobacco raw materials; tobacco raw material samples B1, B2, B3, C4, C5, and C6 are not contained in A, and B1 and C4 are tobacco raw materials of the same origin, the same category, and the same part, B2 and C5 are tobacco raw materials of the same origin, the same category, and the same part, and B3 and C6 are tobacco raw materials of different origins, different categories, and different parts. After mixing the cigarette samples A, B1, B2, B3, C1, C2, and C3 in different proportions, mixed samples A1, A2, A3, A4, A5, and A6 are obtained, and the specific mixing relationship is shown in Table 6.

[0111] Table 6 Tobacco sample mixing scheme

[0112] Blended Sample Number Blended Sample Composition A1 Cigarette sample A and tobacco raw material sample B1 were blended at a ratio of 95:5. A2 Cigarette sample A and tobacco raw material sample C4 were blended at a ratio of 95:5. A3 Cigarette sample A and tobacco raw material sample B2 were blended at a ratio of 75:25. A4 Cigarette sample A and tobacco raw material sample C5 were blended at a ratio of 75:25. A5 Cigarette sample A and tobacco raw material sample B3 were blended at a ratio of 95:5. A6 Cigarette sample A and tobacco raw material sample C6 were blended at a ratio of 95:5.

[0113] Take 100 grams of tobacco raw material samples B1, B2, B3, C4, C5, C6, and mixed samples A1, A2, A3, A4, A5, A6 respectively, and collect spectral data of each sample using a thermoelectric antaris II near-infrared spectrometer according to the experimental conditions shown in Table 2, and record them as Spec_B1, Spec_B2, Spec_B3, Spec_C4, Spec_C5, Spec_C6, and Spec_A1, Spec_A2, Spec_A3, Spec_A4, Spec_A5, Spec_A6 respectively.

[0114] The collected Spec_B1_Before, Spec_B2_Before, Spec_B3_Before, Spec_C4_Before, Spec_C5_Before, Spec_C6_Before, and Spec_A1_Before, Spec_A2_Before, Spec_A3_Before, Spec_A4_Before, Spec_A5_Before, Spec_A6_Before spectral data are preprocessed using a near-infrared data preprocessing scheme composed of standard normal variable transformation method, first-order derivative method, and maximum minimum rule method, to obtain corresponding preprocessed spectral data, which are recorded as Spec_B1_Behind, Spec_B2_Behind, Spec_B3_Behind, Spec_C4_Behind, Spec_C5_Behind, Spec_C6_Behind, and Spec_A1_Behind, Spec_A2_Behind, Spec_A3_Behind, Spec_A4_Behind, Spec_A5_Behind, Spec_A6_Behind respectively.

[0115] Among them, Figure 6 is a graphical plot of the Spec_B1_Before spectral data and the Spec_C4_Before spectral data obtained by the experiment, Figure 7 is a graphical plot of the Spec_B1_Behind spectral data and the Spec_C4_Behind spectral data after being processed by the "SNV+SG1D+MM" preprocessing scheme, and in Figure 6The two spectral data with close shape and distance are preprocessed by the standard normal variable transformation method, the first derivative method and the maximum and minimum rule, and great differences are generated in the shape and distance, indicating that the spectral preprocessing method of the standard normal variable transformation method, the first derivative method and the maximum and minimum rule in the application can well mine the subtle differences between spectral data.

[0116] For the 12 processed near-infrared spectra, the similarity calculation method of near-infrared spectral data combining the Euclidean distance, correlation coefficient and divergence information is used to calculate the similarity between the spectra, and the results are shown in Table 7.

[0117] Table 7: Experimental results of near-infrared spectral similarity calculation

[0118]

[0119] In the examples, when similar tobacco leaves (similarity: 0.9606) are used for small proportion (5%) formula tobacco leaf replacement, the similarity of the replaced formula is still very similar (similarity: 0.9975); when similar tobacco leaves (similarity: 0.9790) are used for large proportion (25%) formula tobacco leaf replacement, the similarity of the replaced formula is still high, but compared with the small ratio high similarity tobacco leaf replacement, there is a more obvious decrease (similarity: 0.9846); when dissimilar tobacco leaves (similarity: 0.1358) are used for small proportion (5%) formula tobacco leaf replacement, the similarity of the replaced formula has a more obvious decrease (similarity: 0.9202).

[0120] The above examples not only introduce the specific application process of the application, but also verify the evaluation method of the weak difference signal of near-infrared spectral data provided by the application, which can identify the weak signal difference of near-infrared spectral data and accurately distinguish different categories of detection samples.

Claims

1. A method for evaluating a weak difference signal of near infrared spectroscopy data, characterized by, The method comprises the following steps: Step 1: Preparation of samples and samples respectively, two near infrared spectrum data were obtained. Step 2: scatter correction is performed on the two pieces of near-infrared spectrum data respectively by using a standard normal variable transformation, so as to eliminate the scattering effect caused by uneven distribution of samples; Step 3: noise processing is performed on the two pieces of near-infrared spectrum data which have completed the scatter correction respectively by using a first-order derivation method, so as to remove noise in the spectrum, improve the signal-to-noise ratio of the spectrum and enhance the distinguishability of overlapping peaks; Step 4: normalization processing is performed on the two pieces of near-infrared spectrum data which have removed noise in the spectrum respectively by using a maximum-minimum rule, so as to enhance data comparability; Step 5: similarity of the two pieces of near-infrared spectrum data is calculated by combining the Euclidean distance, the correlation coefficient and the information divergence, and the specific method is as follows: wherein, is the sample near infrared spectrum data the near infrared spectrum data after standard normal variate transformation processing method, first order derivative denoising operation and maximum minimum rule dimension elimination, is the sample near infrared spectrum data the near infrared spectrum data after standard normal variate transformation processing method, first order derivative denoising operation and maximum minimum rule dimension elimination; First, in calculating the Euclidean distance of two near infrared spectra and In the Euclidean space, the distance between two near infrared spectrum vectors is calculated according to the following formula: (11) in: Representing near-infrared spectral data and The Euclidean distance value, Representing near-infrared spectral data The Absorbance at each wavelength point Representing near-infrared spectral data The Absorbance at each wavelength point Near-infrared spectral data and The number of wavelength points; Secondly, when calculating the correlation coefficient of the two pieces of near-infrared spectrum data, the correlation of the two pieces of near-infrared spectrum vectors is calculated by the following formula: (12) in: Near-infrared spectral data and The correlation coefficient, Representing near-infrared spectral data The Absorbance at each wavelength point Representing near-infrared spectral data The Absorbance at each wavelength point Near-infrared spectral data and The number of wavelength points, , Near-infrared spectral data and Average absorbance; Then, in calculating the scattering degree information of the two near-infrared spectra, based on the information measure theory, the two near-infrared spectra are respectively regarded as information elements with probability statistical characteristics, and the absorbance probability of each wavelength point in the two spectra is described according to the following formula: and and the two near-infrared spectra are respectively regarded as information elements with probability statistical characteristics, and the absorbance probability of each wavelength point in the two spectra is described according to the following formula:​ (13) (14) wherein is the absorbance probability value of the th wavelength point in the is the absorbance probability value of the th wavelength point in the Accordingly, and The relative entropy calculation formula of two near-infrared spectra is expressed as follows: (15) (16) wherein is relative to the relative entropy, is relative to the relative entropy, According to and The relative entropy of two near infrared spectra is calculated as the information divergence of the two spectra according to the following formula: (17); wherein represents and the information divergence of two near-infrared spectra; As above, representing near infrared spectroscopy data and Euclidean distance, representing near infrared spectroscopy data and correlation coefficient, representing near infrared spectroscopy data and information divergence, to finally characterize the similarity of near infrared spectroscopy data and according to the following formula: (10) wherein, the similarity of two near infrared spectral data described in the present invention and to characterize the sample the similarity of the sample to the sample.

2. The method of evaluating a weak difference signal of near infrared spectroscopic data according to claim 1, characterized by, Step 1 is specifically as follows: The samples were collected respectively and the near infrared spectrum data of the samples were recorded as and , respectively, which were represented as formula (1) and formula (2) respectively: (1) (2) wherein, is the number of wavelength points, represents the near infrared spectrum data of a sample represents the absorbance of the nth wavelength point of the near infrared spectrum data of a sample represents the near infrared spectrum data of a sample represents the absorbance of the nth wavelength point of the near infrared spectrum data of a sample represents the near infrared spectrum data of a sample represents the absorbance of the nth wavelength point of the near infrared spectrum data of a sample represents the near infrared spectrum data of a sample represents the absorbance of the nth wavelength point of the near infrared spectrum data of a sample 3. The method of evaluating a weak difference signal of near infrared spectroscopic data according to claim 1, characterized by, Step 2 is specifically as follows: The standard normal variable transformation method is used to eliminate the influence of sample particle surface scattering and light path change on the near-infrared diffuse reflection spectrum during the collection of the near-infrared spectrum, and the standard normal variable transformation method is different from the standardization algorithm in that the standard normal variable transformation method can process a spectrum individually; The near-infrared spectrum data of the sample The standard normal variable transformation processing method of the near-infrared spectrum data is as follows: The standard normal variable transformation processing method of the near-infrared spectrum data is as follows: (3) in, Indicates sample Near-infrared spectral data The Absorbance at each wavelength point Representing near-infrared spectral data The value after standard normal variable transformation. Near-infrared spectral data The average absorbance at all wavelengths. Number of wavelength points; Samples Near infrared spectroscopy data The spectrum data processed by standard normal variate transformation is expressed as follows: (4) Following the same procedure, the near infrared spectral data of the samples are represented as follows: The spectral data after standard normal variate transformation are represented as follows: (5)。 4. The method of evaluating a weak difference signal of near infrared spectroscopic data according to claim 3, characterized by, Step 3 is specifically as follows: For the near-infrared data in which the influence of near-infrared diffuse reflection is eliminated by using the standard normal variable transformation, a first-order derivation method is used to perform smoothing filtering on the near-infrared spectrum data, so as to reduce the interference of noise data, and the first-order derivation method used is an improvement based on a moving smoothing algorithm; The near-infrared spectrum data of the sample The near-infrared spectrum data of the sample The data after the standard normal variable transformation processing method The denoising specific process is as follows: Set the filter window length ( The number of wavelength points in a typical spectrum is a constant. Values When the number of wavelength points in the spectrum Values For near-infrared spectral data absorbance at wavelength Its filtering window is represented as (6) wherein , , denotes the number of spectral measurement points of the filter window ; The polynomial is fitted to the data points using a least squares method as shown in equation (7) The polynomial is fitted to the data points using a least squares method as shown in equation (7)​ (7) wherein ; For each spectral measurement point within the filter window an equation is constructed based on equation (7), ultimately forming a system of linear equations, which is fitted by the least squares method. After fitting the system of linear equations, the parameters of the polynomial are determined and the absorbance of the wavelength points in is filtered by the following equation:​ (8) The absorbance of the wavelength point in the near infrared spectrum data The absorbance of the wavelength point in the near infrared spectrum data After the above processing, the near infrared spectrum data The spectrum data after smoothing filtering is shown in the following formula: (9) To date, For samples Near infrared spectroscopy data Near infrared spectroscopy data after noise processing after standard normal variable transformation processing method By the same process as above, the near-infrared spectrum data of the sample is obtained after denoising .

5. The method of evaluating a weak difference signal of near infrared spectroscopic data according to claim 4, characterized by, Step 4 is specifically as follows: For the near-infrared spectrum data in which the influence of near-infrared diffuse reflection is eliminated by using the standard normal variable transformation and noise in the spectrum is removed by using the first-order derivation, a maximum-minimum rule is further used to eliminate the dimension of the spectrum, so as to enhance the comparability between the spectra; The near infrared diffuse reflectance effect was eliminated using the standard normal variate transformation and the first derivative was used to remove noise from the near infrared spectral data For example, the specific process of the maximum-minimum rule to eliminate the spectral dimension is shown in the following formula: (10) wherein, is the absorbance of the wavelength point in the spectrum, wherein, is the absorbance of the wavelength point in the spectrum, the absorbance of the wavelength point in the spectrum, wherein, is the number of wavelength points; for the sample near infrared spectroscopy data near infrared spectroscopy data after standard normal variate transformation method, first derivative denoising operation and maximum minimum rule dimensionless elimination By the same procedure as described above, the following compounds were prepared. After de-dimensioning, near-infrared spectral data was obtained.

6. The method of evaluating a weak difference signal of near infrared spectroscopic data according to claim 1, characterized by, The smaller the value, The larger the value, The smaller the value, indicates that the near infrared spectral data and characterized sample similarity will be higher.

Citation Information

Patent Citations

  • Evaluation method and system

    CN106248621A

  • A near infrared spectrum signal processing method and system based on wavelet packet analysis

    CN109558843A