An adaptive near infrared spectrum transformation method based on correlation and gaussian curve fitting
By employing an adaptive near-infrared spectral transformation method, combined with correlation analysis and Gaussian curve fitting, the problem of quantitative analysis error caused by spectral peak overlap was solved, achieving efficient information decomposition and recombination of near-infrared spectra and reducing prediction errors.
Patent Information
- Application Number
- CN202310236906.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-13
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2043-03-13
AI Technical Summary
The overlapping of spectral peaks in near-infrared spectra leads to significant errors in the quantitative analysis of analytes. Existing methods struggle to effectively decompose and reconstruct overlapping information, and are also susceptible to noise and non-unique parameters.
An adaptive near-infrared spectral transformation method based on correlation and Gaussian curve fitting is adopted. By determining the number of Gaussian functions, the center wavelength and the optimal bandwidth, and solving the Gaussian function height by combining a system of linear equations, a full-rank fitting integral transformation is performed to achieve effective decomposition and recombination of overlapping information.
Without introducing fitting error, the quantitative analysis error was reduced, the utilization rate of spectral information was improved, the influence of noise was reduced, and the effective decomposition and recombination of overlapping information was achieved, resulting in a prediction error reduction of at least 25%.
Smart Images

Figure CN116559110B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an adaptive near-infrared spectral transformation method based on correlation and Gaussian curve fitting, and relates to the field of near-infrared spectral transformation technology. Background Technology
[0002] Near-infrared spectroscopy is widely used for quantitative analysis of analytes in food, chemical, and environmental fields due to its advantages such as no secondary pollution and fast detection speed. [1]-[3] Near-infrared spectroscopy is formed by the absorption of photon energy in the wavelength range of 780–2526 nm by polar chemical bonds in molecules. Because each chemical bond has multiple overtone peaks, its spectral band is usually broad, and the peaks of different chemical bonds overlap significantly, which can lead to errors in quantitative analysis. [4] To reduce the impact of overlapping peaks and improve the utilization of spectral information, appropriate transformation or preprocessing of near-infrared spectra is crucial.
[0003] Near-infrared spectroscopy usually cannot be directly quantitatively analyzed based on the peak values. It is necessary to establish a prediction model for the content of the analyte. In this process, different preprocessing methods are used to improve the accuracy of the prediction model. Reference [5] proposes a generalized multiplicative scattering correction method to better eliminate the influence of the baseline on the near-infrared spectrum and improve the prediction accuracy of the oil content in oil palm fruit. Reference [6] combines multiple preprocessing methods and proposes a selective ensemble preprocessing strategy, which achieves better preprocessing effects on the near-infrared spectra of corn, blood, and edible oil. In addition, some band selection methods have emerged to screen effective information in the original spectral data. Reference [7] proposed an iterative window reduction self-service soft shrinking algorithm, which accurately selects near-infrared spectral bands by continuously reducing the window, thereby improving the accuracy of corn protein content prediction. Reference [8] proposed a dual-competition adaptive reweighted sampling algorithm, which effectively reduces the band selection differences between near-infrared spectra collected by multiple instruments and batches, and verified it on corn, pharmaceutical and wheat spectral datasets. Reference [9] proposed a three-step hybrid strategy, which combines the advantages of coarse selection, fine selection and optimal selection, and achieves effective selection of near-infrared spectral bands for tobacco and beer. Such methods can improve the accuracy of prediction models, but they rarely pay attention to the influence of spectral overlapping peaks on quantitative analysis, and only select information that is more conducive to modeling on the original spectral data. They cannot fundamentally change the spectral peak overlap of near-infrared spectra, and the utilization rate of information is low.
[0004] Current methods to reduce the influence of overlapping peaks on the spectrum include improving the instrument's own detection resolution and separating overlapping peaks using mathematical methods. However, instrument improvements are often limited by funding and research and development time, so more researchers choose to use mathematical methods.
[10] Reference
[11] proposes a near-infrared spectral feature extraction method based on Gaussian curve fitting, which decomposes the original spectrum into three Gaussian peaks and establishes a more accurate prediction model for maize chlorophyll content, verifying the effectiveness of the curve fitting method in resolving near-infrared spectral overlapping peaks and improving model performance. Curve fitting of various function forms achieves peak resolution by determining the peak shape parameters. Before this, it is usually necessary to determine the peak position. Reference
[12] proposes an overlapping peak separation method combining second derivative of spectral lines and Voigt function curve fitting. First, the peak position is determined by finding the second derivative, and then other peak parameters are determined by Voigt function curve fitting. This method is applied to the infrared and Raman spectra of four globulins, achieving a more accurate estimation of globulin structure. Reference
[13] proposes a method combining second derivative and Lorentz force function curve fitting to resolve the infrared spectra of four different asphalts, achieving peak resolution of different chemical bonds. The above methods determine the spectral position by improving spectral resolution, but ignore the possible noise influence. To reduce the influence of noise during differentiation, reference
[14] proposed a method for determining the spectral peak position by combining Fourier deconvolution and second derivative. The noise was controlled by the SG convolution smoothing method, thus realizing accurate qualitative and quantitative analysis of protein secondary structure using infrared spectroscopy. In addition, applying wavelet transform to the detection of overlapping peak positions can also reduce the influence of noise.
[15] Reference
[16] proposes a method combining image segmentation and wavelet transform to better eliminate the influence of noise on peak position determination, and realizes accurate detection of peak positions in simulated spectra and actual matrix-assisted laser desorption / ionization time-of-flight mass spectra. With the development of neural networks, deep learning has also been applied to overlapping peak analysis. Reference
[17] proposes a method to determine spectral peak parameters using deep neural networks. The deep model is trained using a synthesized nuclear magnetic resonance dataset and verified on the nuclear magnetic resonance spectra of complex protein and metabolomics mixtures, achieving good results.
[0005] The methods described above achieve overlapping peak separation by determining actual spectral peak parameters, exhibiting strong interpretability. However, they are rarely applied to quantitative analysis in near-infrared spectroscopy and have certain limitations. On one hand, to reduce the impact of noise on peak position determination, data denoising is necessary, which may result in a loss of usable information.
[18] On the other hand, the values of spectral peak parameters are usually related to the initial values, so the analysis results may not be unique.
[0006] To avoid the problems encountered in finding actual spectral peak parameters, some scholars have proposed a method based on the overall near-infrared spectrum. Reference
[19] applied the self-modeling mixture analysis method to the single-substance analysis of the mixture, and used matrix decomposition to separate the surface-enhanced Raman scattering spectra of different substances. At the same time, it performed qualitative and quantitative analysis on each component in the mixed pesticide. Reference
[20] made the infrared spectra of pure substances and mixtures into a dataset, and applied the neural network to the quantitative analysis of functional groups of mixed hydrocarbon fuels, which helps to accurately predict various properties of the mixture. The above methods can reduce the influence of overlapping spectral peaks on quantitative analysis without determining the actual spectral peaks. However, the method based on the overall near-infrared spectrum has high requirements for spectral quality. It is usually based on the superposition theory, so it is not suitable for near-infrared spectra with small signal energy, sensitive to external interference, and complex composition.
[0007] Based on the problems of current methods, this paper proposes a Gaussian curve fitting method that does not require determining the actual spectral peaks. By combining correlation analysis and solving the system of equations, the parameters of the Gaussian function are determined, so that the parameters have a unique solution. It is imperative to achieve effective decomposition and recombination of overlapping information in the original near-infrared spectrum without introducing fitting errors. Summary of the Invention
[0008] The technical problem to be solved by this invention is that the overlapping of spectral peaks in near-infrared spectra causes large errors in the quantitative analysis of the analyte. To address this problem, an adaptive near-infrared spectral transformation method based on correlation and Gaussian curve fitting is proposed.
[0009] The technical solution adopted by the present invention to solve the above-mentioned technical problems is as follows:
[0010] An adaptive near-infrared spectral transformation method based on correlation and Gaussian curve fitting is described below:
[0011] (1) Determine the number of Gaussian functions based on the number of discrete points n of the original spectrum, so that each discrete wavelength corresponds to one Gaussian function and each Gaussian function corresponds to one center wavelength;
[0012] (2) The original near-infrared spectrum is combined with formula (9) to adaptively determine the optimal bandwidth of each Gaussian function to ensure that the subsequent transformation does not amplify noise and effectively displays the information of overlapping peaks;
[0013]
[0014] In the formula, P(δ, λ) represents the correlation between the local original spectral data centered at wavelength λ and a Gaussian function with bandwidth δ. The greater the correlation, the more Gaussian function components with bandwidth δ are present at that wavelength. S(x) represents the original spectral vector corresponding to the original spectral wavelength x.
[0015] The correlation of each bandwidth of each Gaussian function is calculated by equation (9), and the bandwidth δ corresponding to the maximum correlation value is selected as the optimal bandwidth of the Gaussian function; the number of optimal bandwidths is consistent with the number of Gaussian functions.
[0016] (3) Substitute the center wavelength from step (1) and the optimal bandwidth from step (2) into the linear equation system shown in equation (5) to obtain the height A of each Gaussian function.
[0017]
[0018] In the formula, A is the height vector of the Gaussian function. To fit the spectral vector, S is the original spectral vector; n is the number of discrete spectral points; λ(1) represents the wavelength corresponding to 1 discrete spectral point, and λ(n) represents the wavelength corresponding to n discrete spectral points, and so on. s(1) represents the fitted value corresponding to the first discrete point of the spectrum; a(1) represents the original value of the first discrete point of the spectrum; the others are similar; the matrix on the left side of formula (5) is a full-rank matrix;
[0019] The near-infrared spectrum transformation is completed by integrating the area of each Gaussian function based on the bandwidth δ and the height A (column vector) of the Gaussian function.
[0020] An application of near-infrared spectroscopy, wherein the near-infrared spectrum is a transformed near-infrared spectrum obtained by the transformation method described in claim 1, and the transformed near-infrared spectrum is used for quantitative analysis.
[0021] A quantitative analysis method based on near-infrared spectroscopy, the implementation process of which is as follows:
[0022] Quantitative analysis consists of two parts: modeling and detection.
[0023] During the modeling phase, the spectral data in the training set need to be compared with the true values of the corresponding analyte content to establish a quantitative analysis model, and partial least squares method is used for modeling.
[0024] During the detection phase, the test set or the actual collected spectral data is input into the quantitative analysis model to obtain the predicted content of the analyte.
[0025] During the modeling phase:
[0026] (1) Input the original near-infrared spectral matrix as the training set; the number of rows in the matrix is the number of training set samples, and the number of columns is the number of spectral discrete points;
[0027] (2) Calculate the average spectrum. Since the near-infrared spectral shapes of the same substance are not significantly different, the average spectrum of the original near-infrared spectral matrix can be used to determine the bandwidth of the Gaussian function, which can reduce the processing time.
[0028] (3) The bandwidth of the Gaussian function is adaptively determined according to equation (9) to obtain the optimal bandwidth;
[0029] (4) Perform an adaptive full-rank fitting integral transformation on the original near-infrared spectral matrix (training set) according to equations (3) and (9) to obtain the transformed data matrix (training set); perform the adaptive near-infrared spectral transformation on the original near-infrared spectral matrix as described in claim 1.
[0030] (5) Construct a content prediction model based on equation (1), and establish the content prediction model by combining the transformed data matrix with partial least squares method;
[0031] During the testing phase:
[0032] (1) Input the raw near-infrared spectrum as the test set, or the actual collected data as the test set;
[0033] (2) Perform an adaptive full-rank fitting integral transform on the original spectrum (test set or actual collected data) to obtain the transformed data;
[0034] (3) Content prediction: The transformed data is input into the regression model for quantitative analysis, and the content prediction results are output to achieve quantitative analysis.
[0035] The quantitative analysis method based on near-infrared spectroscopy is used to detect the moisture content of grains.
[0036] The quantitative analysis method based on near-infrared spectroscopy is used to detect the COD content in wastewater.
[0037] The present invention has the following beneficial technical effects:
[0038] This invention addresses the significant errors in the quantitative analysis of analytes caused by overlapping peaks in near-infrared spectra. The proposed method eliminates the need to determine the actual peaks during Gaussian curve fitting. Instead, it uses correlation analysis and equation solving to determine the parameters of the Gaussian function, ensuring a unique solution. The area integral of the Gaussian function is then used to replace the original spectrum, achieving effective decomposition and reconstruction of overlapping information in the original near-infrared spectrum without introducing fitting errors.
[0039] This invention determines the number of Gaussian functions involved in Gaussian curve fitting by using the number of discrete points in the near-infrared spectrum of the analyte, and determines the center position of the Gaussian function by using the wavelength position of the discrete points. Correlation analysis is used to determine the optimal bandwidth of the Gaussian function that is conducive to extracting overlapping information. Based on this, a set of equations for curve fitting is constructed. The height of the Gaussian function is determined by solving the equations, and the area integral of the Gaussian function is performed to obtain the transformed result of the analyte spectrum, thereby constructing a content prediction model for the analyte. The proposed method has been applied to the prediction of COD content in wastewater and moisture content in corn, and the prediction mean square error has been reduced by at least 25% compared to before the transformation. This indicates that the Gaussian function involved in curve fitting does not need to correspond to the true spectral peak information to achieve effective decomposition and recombination of overlapping information in the original spectrum, thereby reducing quantitative analysis errors. Attached Figure Description
[0040] Figure 1 The process of Gaussian curve fitting and integration in near-infrared spectroscopy is shown: (a) original spectrum, (b) Gaussian curve fitting, (c) area integral of Gaussian function; Figure 2 The original spectrum and the Gaussian functions corresponding to the wavelengths λ(i) and λ(j) are given. Figure 3 A schematic diagram illustrating the process of adaptively determining the bandwidth of a Gaussian function; Figure 4 This is a schematic diagram of adaptive full-rank fitting integral transform; Figure 5 A flowchart for applying adaptive full-rank fitting integral transform to quantitative analysis; Figure 6 Figures show the near-infrared spectroscopy module and acquisition device: (a) near-infrared spectroscopy module, (b) near-infrared spectroscopy acquisition device for water samples; Figure 7 The plot of the Gaussian function fitted to full rank; Figure 8 Fitted spectrum and original spectrum; Figure 9 Comparison of near-infrared spectra and bandwidth of the optimal Gaussian function ( Figure 9 (This shows a comparison of the adaptively determined optimal Gaussian function bandwidth and near-infrared spectrum). Figure 10 The image shows the near-infrared spectrum of the original wastewater and the transformation results. Figure 11 Pearson correlation coefficients at different wavelengths before and after near-infrared spectral matrix transformation of wastewater; Figure 12 The image shows the near-infrared spectrum and transformation results of the original moisture content of the grain (corn); Figure 13 Pearson correlation plots of different wavelengths before and after near-infrared spectral matrix transformation of moisture-content grains (corn). Detailed Implementation
[0041] The following is in conjunction with the appendix Figures 1 to 13This paper describes the implementation of a novel near-infrared spectral transformation method based on correlation and Gaussian curve fitting. Part 1 presents the implementation process of an adaptive near-infrared spectral transformation method based on correlation and Gaussian curve fitting. Part 2 introduces the quantitative analysis of near-infrared spectroscopy and the relevant theory of Gaussian curve fitting. Part 3 explains the principle of the proposed adaptive full-rank fitting integral transformation method. Part 4 describes the process of applying the proposed transformation method to quantitative near-infrared spectroscopy analysis. Part 5 experimentally verifies the proposed method using two near-infrared spectra.
[0042] 1. The implementation process of the adaptive near-infrared spectral transformation method based on correlation and Gaussian curve fitting described in this invention is as follows:
[0043] (1) Determine the number of Gaussian functions based on the number of discrete points n of the original spectrum, so that each discrete wavelength corresponds to one Gaussian function and each Gaussian function corresponds to one center wavelength;
[0044] (2) The original near-infrared spectrum is combined with formula (9) to adaptively determine the optimal bandwidth of each Gaussian function to ensure that the subsequent transformation does not amplify noise and effectively displays the information of overlapping peaks;
[0045]
[0046] In the formula, P(δ, λ) represents the correlation between the local original spectral data centered at wavelength λ and a Gaussian function with bandwidth δ. The greater the correlation, the more Gaussian function components with bandwidth δ are present at that wavelength. S(x) represents the original spectral vector corresponding to the original spectral wavelength x.
[0047] The correlation of each bandwidth of each Gaussian function is calculated by equation (9), and the bandwidth δ corresponding to the maximum correlation value is selected as the optimal bandwidth of the Gaussian function; the number of optimal bandwidths is consistent with the number of Gaussian functions.
[0048] (3) Substitute the center wavelength from step (1) and the optimal bandwidth from step (2) into the linear equation system shown in equation (5) to obtain the height A of each Gaussian function.
[0049]
[0050] In the formula, A is the height vector of the Gaussian function. To fit the spectral vector, S is the original spectral vector; n is the number of discrete spectral points; λ(1) represents the wavelength corresponding to 1 discrete spectral point, and λ(n) represents the wavelength corresponding to n discrete spectral points, and so on. s(1) represents the fitted value corresponding to the first discrete point of the spectrum; a(1) represents the original value of the first discrete point of the spectrum; the others are similar; the matrix on the left side of formula (5) is a full-rank matrix;
[0051] The near-infrared spectrum transformation is completed by integrating the area of each Gaussian function based on the bandwidth δ and the height A (column vector) of the Gaussian function.
[0052] 2. Relevant Theories
[0053] 2.1 Quantitative Analysis of Near-Infrared Spectroscopy
[0054] Due to the unique properties of near-infrared spectroscopy, such as its broad bands and numerous overtone peaks, the content information of the analyte cannot be directly obtained from the spectrum. Therefore, a content prediction model for the analyte needs to be established based on the spectral matrix and the corresponding true content values.
[21] Partial least squares (PLS) is a commonly used near-infrared spectroscopy modeling method. This method determines the regression coefficient between the independent variable (near-infrared spectral matrix) and the dependent variable (analyte content) through multiple iterations. Its expression is:
[0055] U=Tβ (1)
[0056] In the formula, U is the score matrix of the dependent variable, T is the score matrix of the independent variable, and β is the regression coefficient matrix. U and T are determined according to equation (2):
[0057]
[0058] In the formula, X and T are the spectral matrix and the analyte content matrix, respectively; P and E are the weight matrix and residual matrix of the independent variable, respectively; Q and F are the weight matrix and residual matrix of the dependent variable, respectively; and a is the number of principal components determined by iteration. The iteration stops when increasing the number of principal components fails to improve the model performance.
[0059] 2.2 Gaussian Curve Fitting and Integration
[0060] Peak overlap analysis is a method to reduce the influence of overlapping spectral peaks. This method separates peaks that were originally overlapping, effectively improving spectral resolution. Gaussian curve fitting is widely used in peak overlap analysis; combined with integral transform, it can further improve the resolution of the original spectrum. Assuming the original spectrum is formed by the superposition of several independent peaks, each represented by a Gaussian function, the process of superimposing these Gaussian functions and approximating the original spectrum is called Gaussian curve fitting.
[22] Its expression is
[0061]
[0062] In the formula, The fitted spectral data is a set of discrete values, where x is the wavelength and A is the wavelength.i δ i , λ i Let be the height, bandwidth, and center position of the i-th Gaussian function, respectively, representing the peak height, peak width, and peak position of the actual spectral peak, and m be the number of Gaussian functions involved in the fitting. If the mean squared error is used to measure the fitting error, then the expression for the fitting error is:
[0063]
[0064] In the formula, S represents the original spectrum, and n is the number of points in the discrete spectral data. Gaussian curve fitting typically uses a least squares algorithm to iteratively adjust the Gaussian function parameters, narrowing down the fitted spectral data. The difference between the original spectral data S and the original spectral data S, MSE, is used to determine the parameter adjustment. The parameter adjustment ends when the MSE is less than the set value.
[0065] In near-infrared spectroscopy, the spectrum usually refers to the absorbance spectrum. According to Beer's Law, absorbance is directly proportional to the number of molecules of the analyte. Since absorbance can be expressed as peak area, the amount or concentration of the analyte can be analyzed using the area of the fitted Gaussian function.
[15] It is worth noting that the area calculation is an integration process. Integration transforms data that originally had bandwidth information into data without bandwidth information. Replacing the original spectral data with the integrated data can improve the resolution of the original data. Figure 1 The process of fitting and integrating the Gaussian curve of the near-infrared spectrum is demonstrated.
[0066] 3 Adaptive Full-Rank Fitting Integral Transform
[0067] 3.1 Full-rank Gaussian curve fitting and integration
[0068] As described in Chapter 2, in traditional Gaussian curve fitting and integration, the Gaussian function is used to represent the true spectral peak, so the determination of the peak position is crucial. However, on the one hand, the determination of the peak position is affected by factors such as noise, and on the other hand, the actual spectral peak shape usually does not completely conform to the standard Gaussian function. Fitting the original spectrum with a small number of Gaussian functions will inevitably lead to fitting errors. In order to address the problems of traditional methods, this paper proposes a transformation method that combines full-rank Gaussian curve fitting (full-rank fitting) with integral operations.
[0069] If the number of discrete points in the original spectrum is n, then this method requires n Gaussian functions to fit the original spectrum, that is, each wavelength corresponds to a Gaussian function. If the bandwidth δ of the Gaussian function is fixed, when performing full-rank fitting, only the height parameter of the Gaussian function needs to be adjusted. Each fitted discrete point corresponds to a linear equation as shown in equation (3), then all spectral data correspond to a system of equations. The system of equations can be written in matrix form as shown in equation (5):
[0070]
[0071] In the formula, A is the height vector of the Gaussian function. To fit the spectral vector, S is the original spectral vector.
[0072] As can be seen from equation (5), unlike traditional Gaussian curve fitting, the result of full-rank fitting can be equal to the original spectrum, which is determined by the characteristics of this linear equation system. In the linear equation system shown in equation (5), the different column vectors of the coefficient matrix (matrix size: n×n) are linearly independent, and the rank of the augmented matrix formed by it and the spectral vector S is n, which is equal to the dimension of the unknown A. Therefore, the height vector A of the Gaussian function has only one solution, so that the Gaussian function completely fits the original spectrum and there is no fitting error. After the height A is determined, the area of each Gaussian function is integrated, and the integrated data is used to replace the original spectrum. This transformation method is called full-rank fitting integral transformation. Taking any two adjacent points in the original spectral data (excluding the starting point and the ending point) as the research object, their corresponding wavelengths (horizontal coordinates) are λ(i) and λ(j), and their energies (vertical coordinates) are S(i) and S(j), respectively. The Gaussian function corresponding to full-rank fitting is as follows. Figure 2 As shown:
[0073] Assumption Figure 2 Since the fitting parameters of the two Gaussian functions are only affected by the adjacent Gaussian functions on the left and right, after full-rank fitting, S(i) and S(j) can be expressed as the sum of the three Gaussian functions, as shown in equation (6):
[0074]
[0075] Based on this, by integrating the areas of the two Gaussian functions, we obtain the energies of these two discrete points after the full-rank fitted integral transform:
[0076]
[0077] Comparing equations (6) and (7), it can be seen that this transformation method treats each discrete data point as an independent transformation object. In the fitting stage, a relatively dense Gaussian function with a certain bandwidth decomposes the data information at different locations; in the integration stage, the information within the bandwidth of the Gaussian function is gathered at the center of the Gaussian function, realizing the decomposition and recombination of the original information, thereby revealing the information covered by the overlapping spectral peaks. As can be seen from equation (7), the transformation result is directly related to the height and bandwidth, and the height is obtained by solving the system of equations after the bandwidth is determined. Therefore, the key to achieving a good transformation effect for the full-rank fitting integral is to determine the optimal bandwidth of each Gaussian function before the transformation, so that the transformation can better eliminate the influence of overlapping peaks.
[0078] 3.2 Adaptive Determination of Gaussian Function Bandwidth
[0079] Near-infrared spectral waveforms are highly variable and exhibit complex overlap patterns. Different Gaussian function bandwidths are suitable for different overlap scenarios: a larger bandwidth allows for full-rank fitting integral transform to decompose and reassemble the data over a wider range, which is useful when spectral peak overlap is significant; conversely, a smaller bandwidth has the opposite effect. Manually analyzing and determining the bandwidth of each Gaussian function for different scenarios is clearly impractical. Therefore, this paper proposes an adaptive method for determining the bandwidth of Gaussian functions.
[0080] Since full-rank fitting can perfectly fit the original spectrum to the curve after superposition of Gaussian functions, this paper uses the bandwidth of the superposition components contained in the original spectrum as the bandwidth of the Gaussian function during full-rank fitting. The analysis of superposition components can be achieved through correlation calculation. Specifically, the correlation is calculated by performing inner product operations between Gaussian function components with different bandwidths and the original spectrum. The process of determining the bandwidth using correlation is called adaptive determination of Gaussian function bandwidth, as shown in equation (8):
[0081]
[0082] In the formula, P(δ, λ) represents the correlation between the local original spectral data centered at λ and the Gaussian function with bandwidth δ. The greater the correlation, the more Gaussian function components with bandwidth δ are present at that wavelength.
[0083] In the process of calculating correlation, it is necessary to ensure that the calculation result is only related to the bandwidth. As can be seen from equation (8), the bandwidth δ and height A of the Gaussian function will affect the inner product result simultaneously. There are two solutions to this problem: 1. Fix the height of the Gaussian function; 2. Represent the height parameter with bandwidth. In solution 1, since the integral of the Gaussian function is always positive, if the function height is fixed, as δ gradually increases, the number of data points affecting the inner product result gradually increases, resulting in the inner product value of the spectrum and the Gaussian function gradually increasing, and the process of finding the optimal δ does not converge and cannot be solved. Therefore, this paper adopts solution 2, which fixes the area of the Gaussian function so that the height of the Gaussian function is represented by the bandwidth. At this time, the inner product result is only related to the bandwidth of the Gaussian function. The expression for adaptively determining the bandwidth of the Gaussian function is updated as shown in equation (9):
[0084]
[0085] After improvement, the bandwidth of the Gaussian function with the strongest correlation at each wavelength position is adaptively determined by equation (9). The specific process is as follows: Figure 3 As shown in the figure, the dark boxes represent the maximum values of the correlation values in each column, and the corresponding Gaussian function bandwidth is the optimal bandwidth at that wavelength position.
[0086] Since actual near-infrared spectra exhibit diverse waveforms, the smoothed portion of the spectrum will have a stronger correlation with Gaussian functions possessing a larger δ value. This larger δ value allows the subsequent full-rank fitting integral transform to decompose and reconstruct the original data over a wider range, maximizing the extraction of information obscured by overlapping peaks in the smoothed portion. Conversely, the drastically changing portion of the spectrum has a higher data resolution and a stronger correlation with Gaussian functions possessing a smaller δ value. Therefore, decomposition and reconstruction are performed only within a small range, preserving the original information and avoiding noise amplification. In summary, this method utilizes the correlation with local spectral waveforms to determine the bandwidth of the Gaussian function, giving it significant advantages under various near-infrared spectral waveform conditions.
[0087] 3.3 Adaptive Full-Rank Fitting Integral Transform
[0088] Combining the full-rank fitting integral transform from 3.1 with the adaptive determination of the Gaussian function bandwidth from 3.2, a novel near-infrared spectral transformation method for quantitative analysis is proposed—the adaptive full-rank fitting integral transform. The transformation steps are as follows:
[0089] (1) Determine the number of Gaussian functions based on the number of discrete points in the original spectrum, so that each discrete wavelength corresponds to one Gaussian function;
[0090] (2) The original near-infrared spectrum is combined with formula (9) to adaptively determine the optimal bandwidth of each Gaussian function to ensure that the subsequent transformation does not amplify noise and effectively displays the information of overlapping peaks;
[0091] (3) Based on the center wavelength of step (1) and the optimal bandwidth of step (2), construct a linear system of equations as shown in equation (5) and perform a full-rank fitting integral transformation on the original spectrum.
[0092] The transformation process of the proposed method is as follows: Figure 4 As shown: 4. Quantitative Analysis Based on Adaptive Full-Rank Fitting Integral Transform
[0093] Quantitative analysis consists of two parts: modeling and detection. In the modeling stage, the spectral data in the training set needs to be compared with the true values of the corresponding analyte content to establish a quantitative analysis model. This paper uses partial least squares method for modeling. In the detection stage, the test set or the actual collected spectral data is input into the quantitative analysis model to obtain the predicted results of the analyte content.
[0094] The flowchart for applying the proposed transformation method to quantitative analysis is as follows: Figure 5 As shown, the specific steps are as follows:
[0095] Modeling phase:
[0096] (1) Input the original near-infrared spectral matrix (training set). The number of rows in the matrix is the number of training set samples, and the number of columns is the number of discrete spectral points;
[0097] (2) Calculate the average spectrum. Since the near-infrared spectral shapes of the same substance do not differ much, the average spectrum of the original near-infrared spectral matrix can be used to determine the Gaussian function bandwidth, which can reduce processing time;
[0098] (3) The bandwidth of the Gaussian function is adaptively determined according to equation (9) to obtain the optimal bandwidth;
[0099] (4) Perform adaptive full-rank fitting integral transformation on the original near-infrared spectral matrix (training set) according to equations (3) and (9) to obtain the transformed data matrix (training set);
[0100] (5) Construct a content prediction model based on equation (1). The transformed data matrix is then combined with partial least squares method to establish the content prediction model;
[0101] Testing phase:
[0102] (1) Input the raw near-infrared spectrum (test set or actual collected data);
[0103] (2) Perform an adaptive full-rank fitting integral transform on the original spectrum (test set or actual collected data) to obtain the transformed data;
[0104] (3) Content prediction. The transformed data is input into a regression model for quantitative analysis, and the content prediction results are output to achieve quantitative analysis.
[0105] 5. Experiments and Discussion
[0106] To verify the effectiveness of the proposed transformation method, we compared and analyzed the changes in the near-infrared spectra themselves before and after the transformation, the changes in the Pearson correlation coefficient between the spectral data matrix and the true values, and the changes in the prediction error of the analyte content during actual quantitative analysis. To verify the general applicability of the method, experiments were conducted on two significantly different datasets.
[0107] 5.1 Experimental Data and Environment
[0108] Wastewater Dataset: Using Texas Instruments' Near-Infrared Spectroscopy Module (NIRS) NIRscan TM The NanoEvaluation Module collects the transmission near-infrared spectrum of wastewater; the actual image is shown below. Figure 6 As shown, wastewater spectra were continuously collected at a municipal wastewater treatment plant for approximately one week. Simultaneously, the COD value of samples from the same source water was measured using a Shimadzu online chemical oxygen demand (COD-4200) analyzer. The measurement point was a biological treatment tank at the wastewater treatment plant, where the COD value of the sample fluctuated between approximately 15-25 mg / L. After removing data significantly affected by human error or external interference, a dataset was created. Forty samples were selected to form the training set, and ten samples formed the test set.
[0109] Maize Dataset: This publicly available dataset contains 80 near-infrared spectral data points of maize, corresponding to 80 true moisture content values, with a fluctuation range of 9-11%.
[23] A training set of 64 samples and a test set of 16 samples were selected.
[0110] Experimental environment and software: The computer used was running Windows 10, the algorithm was implemented in Python, and the modeling software was Matlab.
[0111] 5.2 Wastewater Dataset Experiment and Analysis
[0112] 5.2.1 Fitting Error
[0113] The wavelength range of the near-infrared spectrum of wastewater is 901-1700 nm, with 228 discrete points. This means that 228 Gaussian functions are required for full-rank fitting, and each discrete point wavelength corresponds to the center position of a Gaussian function. Setting the bandwidth of the Gaussian function to 10 nm, the height of the Gaussian function is obtained by solving equation (3). Figure 7 The example shows a series of Gaussian functions.
[0114] The Gaussian functions are summed to generate the fitted spectrum. Figure 8 The image shows a comparison between the fitted spectrum and the original spectrum.
[0115] from Figure 8 As can be seen, the fitted spectrum completely overlaps with the original spectrum, verifying that there is no fitting error in the full-rank fitting.
[0116] 5.2.2 Optimal Gaussian Function Bandwidth
[0117] The bandwidth of the Gaussian function is set to 4-44 nm, gradually increasing in 4 nm increments (the wavelength interval for near-infrared spectroscopy is approximately 3.5 nm, therefore the minimum bandwidth of the Gaussian function cannot be less than 3.5 nm). Figure 9 This shows a comparison between the adaptively determined optimal Gaussian function bandwidth and the near-infrared spectrum:
[0118] Figure 9 It can be seen that the width of the Gaussian function is closely related to the local waveform of the near-infrared spectrum: the near-infrared spectral peaks near 950nm, 1150nm and 1400nm are more obvious, and the corresponding Gaussian function bandwidth is smaller, which can achieve noise control; while there are no obvious spectral peaks at other positions, and the bandwidth is larger, which is conducive to effective decomposition and recombination.
[0119] 5.2.3 Comparison of prediction errors before and after transformation
[0120] Using the Gaussian function center positions and bandwidths determined in 5.2.1 and 5.2.2, curve fitting is performed on the original spectral data, and the height of each Gaussian function is calculated. Finally, the area of each Gaussian function is calculated, and the adaptive full-rank fitting integral transform result is output. Figure 10 A comparison chart of the original spectral data and the transformed results:
[0121] Depend on Figure 10 The transformation results show that the proposed method causes different changes in the original near-infrared spectrum at different waveforms: at locations where the original spectrum undergoes drastic deformation, the transformed shape is basically consistent with the original data, proving that the proposed method does not amplify the rate of change at sharp points in the original data, i.e., it does not amplify noise; at locations where the original spectrum is relatively smooth, obvious peaks appear after transformation, indicating that the information at these points is decomposed and recombined, allowing overlapping information to be displayed. Furthermore, the transformation time for each step is less than 0.1 seconds, meeting the requirements for real-time detection.
[0122] In near-infrared spectroscopy, the Pearson correlation coefficient characterizes the linear correlation between spectral energy and the true value of the analyte content at different wavelengths. If the correlation coefficient at a certain wavelength increases, it indicates that overlapping information has been revealed through transformation. The Pearson correlation coefficients between the energy value and the COD value at each wavelength before and after spectral matrix transformation are calculated, and the results are as follows: Figure 11 As shown:
[0123] Depend on Figure 11 As can be seen, the Pearson correlation coefficient changed after spectral transformation. Combined with near-infrared spectral shape analysis of wastewater, the Pearson correlation coefficients between the data and the true values in the spectrally smooth bands (1040-1120nm, 1224-1325nm, 1539-1605nm) showed a significant improvement, reaching values of 0.25 and 0.30, respectively. The maximum Pearson correlation coefficient before transformation was 0.22, while the maximum value after transformation was 0.3, representing an improvement of 36%.
[0124] This transformation method was applied to the actual quantitative analysis of COD content in wastewater. Partial least squares (PSLS) methods were used to construct prediction models for the analyte content on three wastewater datasets: the original dataset, the transformed dataset, and the transformed dataset with added multivariate scattering correction (MSC) processing. The mean square errors of the prediction models established for the three datasets are shown in Table 1.
[0125] Table 1 Prediction error of COD content in wastewater
[0126] Type of processing MSE No processing 1.4307 Proposed method 1.0657 Proposed method with MSC 1.0018
[0127] As shown in Table 1, the prediction error of the model was reduced by approximately 25% after the proposed transformation method. Furthermore, by combining the proposed transformation method with the MSC near-infrared spectroscopy preprocessing method, the mean square error of the model prediction was reduced to 1.001838, further decreasing the prediction error by 6% compared to the method without MSC preprocessing.
[0128] 5.3 Experiment and Analysis of Maize Dataset
[0129] The adaptive full-rank fitting integral transform of the near-infrared spectrum of maize was performed using the same procedure as in 5.2. Since the original spectra were different, the number of Gaussian functions was set to 700 when transforming the maize spectral data, while the other settings remained unchanged. Figure 12 The comparison between the original maize spectral data before and after transformation is shown:
[0130] Depend on Figure 12 Similar conclusions can be drawn from the spectral transformation of wastewater: the changes are significant in areas with flat waveforms, and remain essentially unchanged in areas with distinct waveforms. Each transformation takes approximately 0.1 seconds, meeting the requirements for real-time monitoring.
[0131] The results of the Pearson correlation coefficient calculation are as follows: Figure 13 As shown.
[0132] Figure 13 It can be seen that in the wavelength ranges of 1232-1453nm, 1849-1893nm, and 2022-2290nm, the Pearson correlation coefficient between the transformed data matrix and the true value is significantly improved, reaching 0.75 and 0.79 respectively. The maximum value of the Pearson correlation coefficient before the transformation was 0.65, which means that the maximum value has increased by 20%.
[0133] A regression model for predicting moisture content was established on the original near-infrared spectral dataset of maize using the same modeling method. The experimental results are shown in Table 2.
[0134] Table 2 Prediction error of corn moisture content
[0135] Type of processing MSE No processing 0.0986 Proposed method 0.0667 Proposed method with MSC 0.0662
[0136] The experimental results in Table 2 demonstrate that the proposed method reduces the prediction error by approximately 30%. Combined with multivariate scattering correction, the prediction error is further reduced by 0.7%. This experiment also verifies the general applicability of the proposed method.
[0137] 5.4 Discussion
[0138] Experimental verification based on two near-infrared spectral datasets shows that: using information decomposition and recombination based on fitting integrals can reduce the impact of peak overlap in the near-infrared spectral band on quantitative analysis, and the decomposition and recombination process does not require determining the parameters of the actual peaks in the spectrum; near-infrared spectroscopy has a wide range of applications, and the peak overlap of different detection objects is different, so the adaptive transformation method can be better applied to different data.
[0139] 6. Summary
[0140] (1) A full-rank fitting integral transform method is proposed. The fitting process does not require iteration. The unique solution of the Gaussian function parameters can be obtained by solving the system of equations, without introducing fitting error. When the number of discrete spectral points is 700, the calculation time is about 0.1s, which meets the real-time requirements. The full-rank fitting is combined with integral operation to realize the decomposition and recombination of the original information.
[0141] (2) In order to adapt the full-rank fitting integral transform to the varied near-infrared spectral waveforms, an adaptive full-rank integral fitting transform method is proposed by combining the method of adaptively determining the bandwidth of the Gaussian function. This method can simultaneously meet the requirements of improving the resolution of data in the flat region and not amplifying the noise in the region with obvious spectral peaks, effectively decompose and reorganize the original data, and reduce the influence of spectral peak overlap.
[0142] (3) The proposed transformation method was applied to the quantitative analysis of wastewater and corn data respectively. The experiment showed that the Pearson correlation coefficient between the transformed data and the true value of the content of the analyte was improved, with the maximum value increasing by more than 20%. The mean square error of the prediction of wastewater COD and corn protein content was reduced by more than 25%, indicating that the proposed method has universal applicability.
[0143] With the rapid development of near-infrared and other detection technologies, effectively extracting data information is a key research direction for various detection data, including near-infrared spectroscopy. Future work will continue to study information extraction methods to further improve data utilization. The references cited in this invention are detailed below:
[0144] [1]Z.Yang,Z.Wang,M.Lin,Nondestructive Testing of Jujube Water Based on the NTRS,
[0145] Xinjiang Agricultural Sciences 58(2021)2320-2326.
[0146] [2]H.Yu,W.Du,Z.Lang,K.Wang,J.Long,A Novel Integrated Approach toCharacterization ofPetroleum Naphtha Properties from Near-InfraredSpectroscopy,Transactions onInstrumentation and Measurement,70(2021)1-13.
[0147] [3]Y.Tang,Z.Chen,Soil pH Prediction Based on Convolution NeuralNetwork and Near InfraredSpectroscopy,Spectroscopy and Spectral Analysis,41(2021)892-897.
[0148] [4]W.Lu,Modern Near Infrared Spectroscopy Analytical Technology,ChinaPetrochemical Press,Beijing,2010.
[0149] [5]D.D.Silalahi,H.Midi,J.Arasan,M.S.Muatafa,J.P.Caliman,Robustgeneralizedmultiplicative scatter correction algorithm on pretreatment ofnear infrared spectral data,Vibrational Spectroscopy,97(2018)55-65.
[0150] [6]X.Bian,K.Wang,E.Tan,P.Diwu,F.Zhang,Y.Guo,A selective ensemblepreprocessingstrategy for near-infrared spectral quantitative analysis ofcomplex samples,Chemometrics andIntelligent Laboratory Systems,197(2020)103916.
[0151] [7]Q.Xu,L.Guo,K.Du,B.Shan,F.Zhang,A Variable Selection Method forNear-infraredSpectroscopy Based on Iterative Shrinkage Window BootstrappingSoft Shrinkage Algorithm,Journal of Instrumental Analysis,41(2022)1229-1241.
[0152] [8]K.Zheng,T.Feng,W.Zhang,X.Huang,Z.Li,D.Zhang,Y.Yao,X.Zou,Variableselection bydouble competitive adaptive reweighted sampling for calibrationtransfer of near infraredspectra,Chemometrics and Intelligent LaboratorySystems,191(2019)109-117.
[0153] [9]H.Yu,Y.Yun,W.Zhang,H.Chen,D.Liu,Q.Zhong,W.Chen,W.Chen,Three-stephybridstrategy towards efficiently selecting variables in multivariatecalibration of near-infraredspectra,Spectrochimica Acta Part A:Molecular andBiomolecular Spectroscopy,224(2020)117376.
[0154]
[10] Q.Shen,Y.Xu,H.Kang,J.Bu,W.Guo,Research Status of Decomposition ofOverlappingPeaks By Mathematical Methods at Home and Abroadm,ValueEngineering,30(2011)197-197.
[0155]
[11] M.Li,Y.Sheng,Study on Application of Gaussian Fitting Algorithmto Building Model ofSpectral Analysis,Spectroscopy and Spectral Analysis,28(2008)2352-2355.
[0156]
[12] A.Sadat,I.J.Joye,3 Chapter 3:Peak Fitting Applied to FourierTransform Infrared and RamanSpectroscopic Analysis of Proteins,Amultidimensional view on zein proteins:Structure andfunctionality in doughand bread systems,10(2022)38-61.
[0157]
[13] M.Asemani,A.R.Rabbani,Detailed FTIR spectroscopy characterizationof crude oil extractedasphaltenes:Curve resolve of overlapping bands,Journalof Petroleum Science andEngineering,185(2020)106618.
[0158]
[14] M.Fevzioglu,O.K.Ozturk,B.R.Hamaker,O.H.Campanella,Quantitativeapproach to studysecondary structure of proteins by FT-IR spectroscopy,usinga model wheat gluten system,International Journal of BiologicalMacromolecules,164(2020)2753-2760.
[0159]
[15] M.F.Wahab,T.C.O'Haver,Wavelet transforms in separation sciencefor denoising and peakoverlap detection,Journal of Separation Science,43(2020)1998-2010.
[0160]
[16] G.Yang,J.Dai,X.Liu,M.Chen,X.Wu,Spectral feature extraction basedon continuouswavelet transform and image segmentation for peak detection,Analytical Methods,12(2020)169-178.
[0161]
[17] D.Li,A.L.Hansen,C.Yuan,L.Bruschweiler-Li,R.Bruschweiler,DEEPpicker is a deepneural network for accurate deconvolution of complex two-dimensional NMR spectra,Naturecommunications,12(2021)1-13.
[0162]
[18] J.Cai,Y.Xiao,X.Li.De-noising of tobacco near infraredspectroscopy based on generalized Stransform,Acta Tabacaria Sinica,23(2017)9-14.
[0163]
[19] B.Hu,D.W.Sun,H.Pu,Q.Wei,Rapid nondestructive detection of mixedpesticides residueson fruit surface using SERS combined with self-modelingmixture analysis method,Talanta,217(2020)120998.
[0164]
[20] Y.Sun,L.Luo,Y.Liu.Analysis of Functional Group Mole Fraction ofComplex Fuels Basedon Neural Network and Infrared Spectrum,Journal ofEngineering Thermophysics,43(2022)1116-1122.
[0165]
[21] G.Wang,H.Ye,Principal component analysis and partial leastsquares,Tsinghua UniversityPress,Beijing,2012.
[0166]
[22] A.Goshtasby,W.D.Oneill,Curve fitting by a sum of Gaussians,CVGIP:Graphical Modelsand Image Processing,56(1994)281-288.
[0167]
[23] https: / / eigenvector.com / resources / data-sets / #corn-sec。
Claims
1. A method for applying near-infrared spectroscopy, the near-infrared spectroscopy being transformed near-infrared spectroscopy obtained by using a correlation and Gaussian curve fitting based adaptive near-infrared spectroscopy transformation method, and the transformed near-infrared spectroscopy being used for quantitative analysis, the method being characterized in that the implementation process of the correlation and Gaussian curve fitting based adaptive near-infrared spectroscopy transformation method is as follows: (1) determining the number of Gaussian functions according to the number n of discrete points of the original spectrum, so that each discrete wavelength corresponds to a Gaussian function, and each Gaussian function corresponds to a center wavelength; (2) adaptively determining the optimal bandwidth of each Gaussian function according to the original near-infrared spectroscopy and formula (9), so as to ensure that subsequent transformation does not amplify noise and effectively displays overlapping peak information; (3) substituting the center wavelength of step (1) and the optimal bandwidth of step (2) into the linear equation group shown in formula (5), to obtain the height A of each Gaussian function; (4) integrating the area of each Gaussian function according to the bandwidth δ and the height A of the Gaussian function, and replacing the original spectrum with the integrated data; taking any two adjacent points in the original spectrum data as the research object, the corresponding wavelengths of the two points being λ(i) and λ(j), and the corresponding energies of the two points being S(i) and S(j); assuming that the fitting parameters of two Gaussian functions are only affected by the left and right adjacent Gaussian functions, then S(i) and S(j) can be expressed in the form of the addition of three Gaussian functions after full rank fitting, as shown in formula (6); and integrating the areas of the two Gaussian functions to obtain the energies of the two discrete points after full rank fitting integral transformation: (5) the implementation process of the quantitative analysis method is as follows: the quantitative analysis is divided into modeling and detection; (6) the modeling stage: (1) inputting the original near-infrared spectroscopy matrix as the training set; the number of rows of the matrix is the number of training set samples, and the number of columns of the matrix is the number of spectral discrete points; (2) calculating the average spectrum, and determining the Gaussian function bandwidth by using the average spectrum of the original near-infrared spectroscopy matrix, so as to reduce the processing time; (3) performing the adaptive near-infrared spectroscopy transformation on the original near-infrared spectroscopy matrix in step (1) according to claim 1; (4) constructing the content prediction model according to formula (1), and establishing the content prediction model by using the transformed data matrix and the partial least squares method; U = Tβ (1) wherein, U is the score matrix of the dependent variable, T is the score matrix of the independent variable, and β is the regression coefficient matrix; and (7) the detection stage: wherein represents the fitted spectrum data, x is the wavelength, A(i), δ(i), λ(i) are the height, bandwidth and center position of the ith Gaussian function, respectively, m is the number of Gaussian functions involved in the fitting. (5) In the formula, A is a height vector of a Gaussian function, is a fitting spectrum vector, S is an original spectrum vector; n is a number of spectrum discrete points; λ(1) represents a wavelength corresponding to the spectrum discrete point number 1, λ(n) represents a wavelength corresponding to the spectrum discrete point number n, and the others are the same; represents a fitting value corresponding to the first spectrum discrete point; S(1) represents an original value of the first spectrum discrete point; A(1) represents a height of a Gaussian function corresponding to the first spectrum discrete point; and the others are the same; the matrix on the left side of formula (5) is a full rank matrix. 。 2. A method for quantitative analysis based on near infrared spectroscopy, characterized by, (1) input the actually collected data as a test set; (2) perform the adaptive near-infrared spectrum transformation on the actually collected data to obtain transformed data; (3) content prediction, input the transformed data into a regression model for quantitative analysis, output a content prediction result, and realize quantitative analysis.
3. The method according to claim 2, wherein The near-infrared spectrum-based quantitative analysis method is used for detecting the water content of the grain.
4. The method according to claim 3, wherein The near-infrared spectrum-based quantitative analysis method is used for detecting the COD content in the sewage.