Chemometrics analysis method based on partial least square method
The stoichiometric analysis through partial least squares method solves the accuracy and efficiency of air pollutant detection, realizes effective detection and control of pollutants of common types of pollutants in the air, and improves the accuracy and efficiency of analysis results.
Patent Information
- Application Number
- CN202510467931.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-25
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art cannot effectively detect and control other types of pollutants in the air except for common pollutants, and the traditional least squares method is inaccurate when the sample number is insufficient or there are multiple correlations of independent variables.
The partial least squares method is used for stoichiometric analysis, the original data is obtained through spectral analysis, preprocessing and feature extraction, a component analysis model is constructed, and the component analysis model is used for detection, and abnormal data points are judged based on structural similarity.
It improves the accuracy and efficiency of air pollutant detection, reduces the demand for sample size, can handle multifaceted complex structural models, eliminates complex collinear relationships between spectra, and ensures the correctness of analysis results.
Smart Images

Figure CN120369659A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental detection, and particularly to a chemometric analysis method, device, computing device and computer storage medium based on partial least squares method. Background Art
[0002] With the continuous development of social economy and science and technology, the whole society has increasingly realized the importance of environmental protection, and thus targeted protection has been carried out for different types such as air pollution, water pollution, and soil pollution.
[0003] For air pollution, currently, it is usually to face the main pollutant types contained in the atmosphere as a whole, judge the possible main pollution sources based on the pollutant types, and thus set targeted pollution control methods or use targeted pollution removal methods.
[0004] However, this method can usually only target common pollutants in the atmosphere, and cannot perform source detection and source treatment, so it is impossible to clearly know which pollutant types are actually generated at the pollution source. Therefore, there is a great possibility that other types of pollutants except common pollutant types will escape at the source, resulting in the area near the pollution source being affected by this type of pollution. From the perspective of overall atmospheric detection, the proportion of this type of pollution is low, and effective removal and prevention and control methods cannot be carried out. Moreover, when detecting the components of pollutants based on spectra, CLS (Constrained Least Squares) has a large demand for the number of samples, and when there is a multiple correlation between independent variables, CLS will fail. Therefore, when the number of samples is small or there is a multiple correlation between independent variables, the detection results will be inaccurate, and it is impossible to clearly know the pollutant types, and thus effective treatment and removal cannot be carried out. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a chemometric analysis method based on partial least squares method and a corresponding chemometric analysis device, computing device and computer storage medium based on partial least squares method.
[0006] According to one aspect of the present invention, there is provided a chemometric analysis method based on partial least squares method, the method comprising:
[0007] Performing spectral analysis on a detection object to obtain the original spectral data of the detection object;
[0008] Performing preprocessing on the collected original spectral data to obtain the processed spectral data after preprocessing;
[0009] Feature extraction is performed based on the preprocessed spectral data, and latent variables are constructed based on preset conditions for independent variables and dependent variables to complete the construction of the component analysis model;
[0010] The component analysis model is used to detect the components of the detection object, and a component detection result is obtained.
[0011] In the above solution, the spectral analysis of the detection object to obtain the original spectral data of the detection object further includes:
[0012] When the detection object is in a gaseous phase, the components are directly subjected to spectral analysis using an infrared spectrometer to obtain the original spectral data corresponding to the detection object;
[0013] When the detection object is in a liquid phase, the headspace sampling method is used to heat the detection object to volatilize it and inhale it into the infrared spectrometer to obtain the original spectral data corresponding to the detection object;
[0014] When the detection object is in a solid phase, the tablet pressing method is used to make a potassium bromide tablet, which is then placed in the optical chamber of the infrared spectrometer to obtain the original spectral data corresponding to the detection object.
[0015] In the above solution, the preprocessing includes at least centering and / or standardization;
[0016] The preprocessing of the collected original spectral data further includes:
[0017] The centering is to subtract the mean value corresponding to the spectral data from each variable in the spectral data so that the average value of the processed data is 0;
[0018] The standardization is to divide each variable in the spectral data by the standard deviation corresponding to the spectral data so that the variance of the processed spectral data is 1.
[0019] In the above solution, the feature extraction based on the preprocessed spectral data, constructing latent variables based on preset conditions for independent variables and dependent variables to complete the construction of the component analysis model further includes:
[0020] Based on the maximum covariance theory, the covariance between the latent variable, which is a linear combination of independent variables, and the dependent variable is maximized, and thus the first latent variable is determined according to the first independent variable; where,
[0021] Let X be an n-dimensional independent variable matrix, Y be a dependent variable vector, t1 be the first latent variable, t1 is a linear combination of X, that is, t1 = Xw1, where w1 is the n-dimensional first weight vector corresponding to the first latent variable t1; based on the maximum covariance theory, the covariance between t1 and Y is max[Cov(t1,Y)], and it can be obtained that w1 = argmax[Cov(Xw1,Y)], therefore, Furthermore, it can be known that Combined with the constraint condition ‖w1‖ = 1, the analytical solution of w1 is calculated to obtain the first latent variable t1; where Cov is the covariance, max[Cov(t1,Y)] is the maximum covariance between the first latent variable t1 and the dependent variable Y, and argmax is the function for finding the parameter of the maximum value. is the transpose of w1, and X T is the transpose of X;
[0022] Regression modeling of the dependent variable Y based on the first latent variable t1 The predicted value of the dependent variable is calculated Based on the first independent variable, the second independent variable X2 is calculated through residuals; regression modeling of the second latent variable t2 is performed using the second independent variable X2, that is And based on the maximum covariance theory to construct the second latent variable t2, that is t2 = X2w2. By calculating the analytical solution of the second weight vector w2, the second latent variable t2 is obtained; this step is repeated until the preset number of latent variables is constructed; where p1 is the linear regression coefficient of the first latent variable t1 with respect to the dependent variable Y; the preset number of latent variables is k.
[0023] Based on the preset number of latent variables, a multiple linear regression model is constructed, and the estimated values of the regression coefficients and the intercept term are determined based on the least squares method to obtain a component analysis model; where the multiple linear regression model is
[0024]
[0025] where b0 is the intercept term, and b1, b2...b k are the regression coefficients corresponding to the respective latent variables.
[0026] In the above solution, the component analysis model is used to detect the components of the detection object to obtain a component detection result, which further includes:
[0027] The component analysis model is used to detect the components of the detection object to obtain a component detection result;
[0028] Based on the result, back-calculation is performed to determine whether the preset requirements are met;
[0029] If the preset requirements are met, the component detection result is sent out; if the preset requirements are not met, the component analysis model is reconstructed.
[0030] In the above solution, the method further includes:
[0031] Extract at least two detection data points from the original spectral data and obtain the target parameters of the detection data points;
[0032] Calculate the structural similarity between detection data points based on the target parameters of the detection data points;
[0033] Determine the characteristic bands and abnormal data points in the spectral data according to the structural similarity between the detection data points, and perform data processing on the abnormal data points.
[0034] In the above solution, the method further includes:
[0035] The target parameters at least include: amplitude, frequency, and depth information;
[0036] The structural similarity between detection data points is
[0037]
[0038] where F a,b is the structural similarity function; DTW is the DTW distance between detection data points; (d, l) α is the depth information of the α-th detection data point; (d, l) β is the depth information of the β-th detection data point; |Δf| is the frequency difference between detection data points; |Δa| is the amplitude difference between detection data points.
[0039] According to another aspect of the present invention, there is provided a chemometric analysis device based on partial least squares method, including: a spectral acquisition module, a data preprocessing module, a model construction module, and a component detection module; wherein,
[0040] The spectral acquisition module is used to perform spectral analysis on the detection object to obtain the original spectral data of the detection object;
[0041] The data preprocessing module is used to preprocess the collected original spectral data to obtain the preprocessed spectral data;
[0042] The model construction module is used to perform feature extraction based on the preprocessed spectral data, and construct latent variables based on preset conditions for independent variables and dependent variables to complete the construction of the component analysis model;
[0043] The component detection module is used to detect the components of the detection object by using the component analysis model to obtain the component detection result.
[0044] According to yet another aspect of the present invention, there is provided a computing device, including: a processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus;
[0045] The memory is used to store at least one executable instruction, and the executable instruction enables the processor to perform operations corresponding to the above-mentioned chemometric analysis method based on partial least squares method.
[0046] According to another aspect of the present invention, a computer storage medium is provided, wherein the storage medium stores at least one executable instruction, wherein the executable instruction enables a processor to perform operations corresponding to the above-mentioned chemometric analysis method based on partial least squares method.
[0047] According to the technical solution provided by the present invention, spectral analysis is performed on the detection object to obtain original spectral data of the detection object; preprocessing is performed on the collected original spectral data to obtain preprocessed processed spectral data; feature extraction is performed based on the preprocessed spectral data, and latent variables are constructed based on preset conditions for independent variables and dependent variables to complete the construction of a component analysis model; and the component analysis model is used to detect the components of the detection object to obtain component detection results. By distinguishing the gas, liquid and solid phases of the detection object, the detection object is spectrally analyzed according to the corresponding method to obtain the spectral data of the detection object; the spectral data of the detection object obtained is preprocessed, and the data is prepared for data analysis based on the component analysis model through centering and standardization, so that the data analysis can be completed more accurately; through the partial least squares method, based on the maximum covariance theory, by calculating the analytical solution of the weight vector corresponding to each potential variable, the potential variables in the component analysis model are determined in turn, and a multivariate prior regression model is constructed to obtain the component analysis model. Therefore, through the partial least squares method, the model's demand for the number of samples is greatly reduced, and there is no need to analyze whether the data conforms to the normal distribution. It can handle complex structural models with multiple dimensions, and can effectively reduce the dimension of the data, eliminate the possible complex collinearity relationship between spectra, and greatly improve the accuracy and analysis efficiency of the qualitative analysis results; based on the component detection results, back calculation is performed to determine whether the results meet the requirements, and if not, the model is reconstructed to further ensure the correctness of the final analysis results, which is conducive to obtaining a reasonable treatment method after completing the component analysis and processing the components in a targeted manner. In addition, by detecting data points in the original spectral data and determining the structural similarity between data points based on their parameters, and thereby processing the abnormal data in the original spectral data, data optimization is achieved, which is more conducive to subsequent data analysis and obtains accurate component analysis results.
[0048] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings.
[0049] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings
[0050] The accompanying drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the accompanying drawings:
[0051] Figure 1 A flowchart showing a chemometric analysis method based on partial least squares according to an embodiment of the present invention is shown;
[0052] Figure 2 A flowchart showing a method for constructing a component analysis model based on partial least squares according to an embodiment of the present invention is shown;
[0053] Figure 3 A flowchart showing a result processing method based on the inverse calculation of component analysis results according to an embodiment of the present invention is shown;
[0054] Figure 4 A flowchart showing a method for preprocessing spectral data based on structural similarity judgment according to an embodiment of the present invention is shown;
[0055] Figure 5 A block diagram showing the structure of a chemometric analysis device based on partial least squares according to an embodiment of the present invention is shown;
[0056] Figure 6 A schematic diagram showing the structure of a computing device according to an embodiment of the present invention is shown. Detailed Embodiments
[0057] The following describes the preferred embodiments of the present invention with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.
[0058] Figure 1 A flowchart showing a chemometric analysis method based on partial least squares according to an embodiment of the present invention is shown. The method includes the following steps:
[0059] Step S101, perform spectral analysis on the detection object to obtain the original spectral data of the detection object.
[0060] Preferably, when the detection object is in a gaseous state, directly use an infrared spectrometer to perform spectral analysis on the components to obtain the original spectral data corresponding to the detection object;
[0061] When the object to be detected is in liquid phase, the headspace sampling method is used to heat the object to be detected to volatilize it, and then inhale it into the infrared spectrometer to obtain the original spectral data corresponding to the object to be detected;
[0062] When the object to be detected is in solid phase, the tablet pressing method is used to make a potassium bromide tablet, and then put it into the optical chamber of the infrared spectrometer to obtain the original spectral data corresponding to the object to be detected.
[0063] Step S102: Preprocess the collected original spectral data to obtain the processed spectral data after preprocessing.
[0064] Preferably, the preprocessing includes at least centering and / or standardization;
[0065] The preprocessing of the collected original spectral data further includes:
[0066] The centering is to subtract the mean value corresponding to the spectral data from each variable in the spectral data, so that the average value of the processed data is 0;
[0067] The standardization is to divide each variable in the spectral data by the standard deviation corresponding to the spectral data, so that the variance of the processed spectral data is 1.
[0068] Step S103: Extract features from the preprocessed spectral data, construct latent variables based on preset conditions for independent variables and dependent variables, and complete the construction of the component analysis model.
[0069] Preferably, the component analysis model is constructed based on the partial least squares method.
[0070] Step S104: Use the component analysis model to detect the components of the object to be detected and obtain the component detection result.
[0071] According to a chemometric analysis method based on partial least squares provided by this embodiment, spectral analysis is performed on a detection object to obtain the original spectral data of the detection object; preprocessing is performed on the collected original spectral data to obtain the processed spectral data after preprocessing; feature extraction is performed based on the spectral data after preprocessing, and latent variables are constructed based on preset conditions for independent variables and dependent variables to complete the construction of a component analysis model; the component analysis model is used to detect the components of the detection object to obtain a component detection result. Through a chemometric analysis method based on partial least squares provided by this embodiment, for the obtained spectral data of the detection object, preprocessing is performed, and through centering and standardization, data preparation is done for data analysis based on the component analysis model, enabling more accurate data analysis to be completed; a component analysis model is constructed through partial least squares, thereby greatly reducing the model's requirement for the number of samples, without the need to analyze whether the data conforms to normal distribution, capable of handling complex structural models with multiple dimensions, and can effectively reduce the dimension of the data, eliminating possible collinearity relationships between spectra, greatly improving the accuracy and analysis efficiency of qualitative analysis results; back-calculation is performed based on the component detection result to further ensure the correctness of the finally obtained analysis result, which is beneficial for obtaining a reasonable processing method after component analysis and specifically processing the components therein.
[0072] Figure 2 Fig. 4 shows a schematic flowchart of a method for constructing a component analysis model based on partial least squares according to an embodiment of the present invention;
[0073] As Figure 2 shown, the method includes the following steps:
[0074] Step S201, based on the maximum covariance theory, maximize the covariance between the latent variable, which is a linear combination of independent variables, and the dependent variable, thereby determining the first latent variable according to the first independent variable.
[0075] Preferably, let X be an n-dimensional independent variable matrix, Y be a dependent variable vector, t1 be the first latent variable, t1 be a linear combination of X, that is, t1 = Xw1, where w1 is the n-dimensional first weight vector corresponding to the first latent variable t1; based on the maximum covariance theory, the covariance between t1 and Y is max[Cov(t1,Y)], and it can be obtained that w1 = argmax[Cov(Xw1,Y)], therefore, Furthermore, it can be known that Combined with the constraint condition ‖w1‖ = 1, the analytical solution of w1 is calculated to obtain the first latent variable t1; where Cov is the covariance, max[Cov(t1,Y)] is the maximum covariance between the first latent variable t1 and the dependent variable Y, and argmax is the maximum value parameter function, is the transpose of w1, X T is the transpose of X.
[0076] Step S202: Calculate a preset number of latent variables based on the first latent variable.
[0077] Preferably, perform regression modeling on the dependent variable Y based on the first latent variable t1 Calculate the predicted value of the dependent variable Based on the first independent variable, calculate the second independent variable X2 through residuals; perform regression modeling on the second latent variable t2 using the second independent variable X2, that is And construct the second latent variable t2 based on the maximum covariance theory, that is, t2 = X2w2. By calculating the analytical solution of the second weight vector w2, obtain the second latent variable t2; repeat this step until the preset number of latent variables is constructed; where p1 is the linear regression coefficient of the first latent variable t1 on the dependent variable Y; the preset number of latent variables is k.
[0078] Step S203: Based on the preset number of latent variables, construct a multiple linear regression model, and determine the estimated values of the regression coefficients and the intercept term based on the least squares method to obtain a component analysis model.
[0079] Preferably, the multiple linear regression model is
[0080]
[0081] where b0 is the intercept term, and b1, b2...b k are the regression coefficients corresponding to the respective latent variables.
[0082] According to the above method, through partial least squares method, based on the maximum covariance theory, by calculating the analytical solutions of the weight vectors corresponding to each latent variable, the latent variables in the component analysis model can be determined in sequence, a multiple linear regression model can be constructed, and then a component analysis model can be obtained. Thus, through the partial least squares method, the demand of the model for the number of samples is greatly reduced, there is no need to analyze whether the data conforms to normal distribution, it can handle complex structural models with multiple dimensions, and it can effectively reduce the dimension of the data, eliminate the possible collinearity relationship between spectra, and greatly improve the accuracy and analysis efficiency of the qualitative analysis results.
[0083] Figure 3 shows a schematic flowchart of a result processing method for back-calculating based on component analysis results according to an embodiment of the present invention;
[0084] As Figure 3 shown, the method includes the following steps:
[0085] Step S301: Detect the components of the detection object using the component analysis model to obtain the component detection result.
[0086] Step S302: Perform back-calculation based on the result to determine whether the preset requirements are met.
[0087] Specifically, if the preset requirements are met, execute Step S303; if the preset requirements are not met, execute Step S304.
[0088] Step S303: Send out the component detection result.
[0089] Step S304: Reconstruct the component analysis model.
[0090] According to the above method, back-calculation can be performed based on the component detection result to determine whether the result meets the requirements, and when it does not meet the requirements, the model is reconstructed, so as to further ensure the correctness of the finally obtained analysis result, which is beneficial to obtaining a reasonable processing method after component analysis and processing the components therein in a targeted manner.
[0091] Figure 4 Shows a schematic flowchart of a spectral data preprocessing method based on structural similarity judgment according to an embodiment of the present invention;
[0092] As Figure 4 shown, the method includes the following steps:
[0093] Step S401: Extract at least two detection data points from the original spectral data and obtain the target parameters of the detection data points.
[0094] Step S402: Calculate the structural similarity between the detection data points based on the target parameters of the detection data points.
[0095] Preferably, the target parameters at least include: amplitude, frequency, and depth information;
[0096] The structural similarity between the detection data points is
[0097]
[0098] where F a,b is the structural similarity function; DTW is the DTW distance between the detection data points; (d, l) α is the depth information of the α-th detection data point; (d, l) β is the depth information of the β-th detection data point; |Δf| is the frequency difference between the detection data points; |Δa| is the amplitude difference between the detection data points.
[0099] Step S403: Determine the characteristic bands and abnormal data points in the spectral data according to the structural similarity between the detected data points, and perform data processing on the abnormal data points.
[0100] Preferably, the data processing may at least include: data cleaning, data replacement, and / or data filling.
[0101] According to the above method, for the detected data points in the original spectral data, the structural similarity between the data points can be determined based on their parameters, and thus the abnormal data in the original spectral data can be processed to achieve data optimization, which is more conducive to subsequent data analysis and obtain a more accurate component analysis result.
[0102] Figure 5 FIG. shows a structural block diagram of a chemometric analysis device based on partial least squares method according to an embodiment of the present invention. As Figure 5 shown, the system includes: a spectral acquisition module 501, a data preprocessing module 502, a model construction module 503, and a component detection module 504; wherein,
[0103] The spectral acquisition module 501 is configured to perform spectral analysis on a detection object to obtain the original spectral data of the detection object.
[0104] Specifically, the spectral acquisition module 501 is further configured to,
[0105] When the detection object is in a gas phase, directly use an infrared spectrometer to perform spectral analysis on the components to obtain the original spectral data corresponding to the detection object;
[0106] When the detection object is in a liquid phase, use the headspace injection method to heat the detection object to volatilize it, and inhale it into the infrared spectrometer to obtain the original spectral data corresponding to the detection object;
[0107] When the detection object is in a solid phase, use the tablet pressing method to make a potassium bromide tablet, and then put it into the optical chamber of the infrared spectrometer to obtain the original spectral data corresponding to the detection object.
[0108] The data preprocessing module 502 is configured to perform preprocessing on the acquired original spectral data to obtain the processed spectral data after preprocessing.
[0109] Preferably, the preprocessing at least includes centering and / or standardization.
[0110] Specifically, the data preprocessing module 502 is further configured to,
[0111] The centering is to subtract the mean value corresponding to the spectral data from each variable in the spectral data so that the average value of the processed data is 0;
[0112] The standardization is to divide each variable in the spectral data by the corresponding standard deviation of the spectral data, so that the variance of the processed spectral data is 1.
[0113] Preferably, the data preprocessing module 502 is further configured to
[0114] Extract at least two detection data points from the original spectral data, and obtain the target parameters of the detection data points;
[0115] Based on the target parameters of the detection data points, calculate the structural similarity between the detection data points; where
[0116] The target parameters at least include: amplitude, frequency, and depth information;
[0117] The structural similarity between the detection data points is
[0118]
[0119] where F a,b is the structural similarity function; DTW is the DTW distance between the detection data points; (d, l) α is the depth information of the α-th detection data point; (d, l) β is the depth information of the β-th detection data point; |Δf| is the frequency difference between the detection data points; |Δa| is the amplitude difference between the detection data points;
[0120] According to the structural similarity between the detection data points, determine the characteristic bands and abnormal data points in the spectral data, and perform data processing on the abnormal data points.
[0121] The model construction module 503 is configured to perform feature extraction based on the preprocessed spectral data, and construct latent variables based on preset conditions for independent variables and dependent variables to complete the construction of the component analysis model.
[0122] Preferably, the model construction module 503 is further configured to
[0123] Based on the maximum covariance theory, make the covariance between the latent variable, which is a linear combination of independent variables, and the dependent variable the largest, and thus determine the first latent variable according to the first independent variable; where
[0124] Let X be an n-dimensional independent variable matrix, Y be a dependent variable vector, t1 be the first latent variable, t1 is a linear combination of X, that is, t1 = Xw1, where w1 is the n-dimensional first weight vector corresponding to the first latent variable t1; based on the maximum covariance theory, the covariance between t1 and Y is max[Cov(t1, Y)], and it can be obtained that w1 = argmax[Cov(Xw1, Y)], so Furthermore, it can be known that Combined with the constraint condition ‖w1‖ = 1, calculate the analytical solution of w1 to obtain the first latent variable t1; where Cov is the covariance, max[Cov(t1, Y)] is the maximum covariance between the first latent variable t1 and the dependent variable Y, and argmax is the function for finding the parameter of the maximum value. is the transpose of w1, and X T is the transpose of X;
[0125] Perform regression modeling of the dependent variable Y based on the first latent variable t1 Calculate the predicted value of the dependent variable Based on the first independent variable, calculate the second independent variable X2 through residuals; use the second independent variable X2 to perform regression modeling on the second latent variable t2, that is And construct the second latent variable t2 based on the maximum covariance theory, that is t2 = X2w2. By calculating the analytical solution of the second weight vector w2, obtain the second latent variable t2; repeat this step until the preset number of latent variables is constructed; where p1 is the linear regression coefficient of the first latent variable t1 with respect to the dependent variable Y; the preset number of latent variables is k.
[0126] Based on the preset number of latent variables, construct a multiple linear regression model, and determine the estimated values of the regression coefficients and the intercept term based on the least squares method to obtain a component analysis model; where the multiple linear regression model is
[0127]
[0128] where b0 is the intercept term, and b1, b2... b k are the regression coefficients corresponding to the respective latent variables.
[0129] The component detection module 504 is used to detect the components of the detection object using the component analysis model to obtain a component detection result.
[0130] Preferably, the component detection module 504 is further used for
[0131] Detect the components of the detection object using the component analysis model to obtain a component detection result;
[0132] Perform back-calculation based on the result to determine whether the preset requirements are met;
[0133] If the preset requirements are met, send out the component detection result; if the preset requirements are not met, reconstruct the component analysis model.
[0134] A chemometric analysis device based on partial least squares according to this embodiment includes: a spectrum acquisition module, a data preprocessing module, a model construction module, and a component detection module; wherein, the spectrum acquisition module is used to perform spectrum analysis on a detection object to obtain the original spectrum data of the detection object; the data preprocessing module is used to preprocess the collected original spectrum data to obtain the processed spectrum data after preprocessing; the model construction module is used to extract features based on the spectrum data after preprocessing, construct latent variables based on preset conditions for independent variables and dependent variables, so as to complete the construction of a component analysis model; the component detection module is used to use the component analysis model to detect the components of the detection object to obtain a component detection result. Through the chemometric analysis device based on partial least squares provided by this embodiment, by distinguishing the gas, liquid, and solid phases of the detection object, the detection object is subjected to spectrum analysis according to the corresponding method to obtain the spectrum data of the detection object; for the obtained spectrum data of the detection object, preprocessing is performed, and through centering and standardization, data preparation is done for data analysis based on the component analysis model, enabling it to complete data analysis more accurately; through partial least squares, based on the maximum covariance theory, by calculating the analytical solutions of the weight vectors corresponding to each latent variable, the latent variables in the component analysis model are determined in sequence, a multiple linear regression model is constructed, and then a component analysis model is obtained. Thus, through partial least squares, the requirement of the model for the number of samples is greatly reduced, there is no need to analyze whether the data conforms to normal distribution, a complex structure model with multiple dimensions can be processed, and the data can be effectively reduced in dimension, eliminating the possible collinearity relationship between spectra, greatly improving the accuracy and analysis efficiency of the qualitative analysis result; based on the component detection result, back-calculation is performed to determine whether the result meets the requirements, and when it does not meet the requirements, the model is reconstructed, thereby further ensuring the correctness of the finally obtained analysis result, which is beneficial to obtaining a reasonable processing method after component analysis and processing the components therein specifically. In addition, for the detection data points in the original spectrum data, the structural similarity between the data points is determined based on their parameters, and thus the abnormal data in the original spectrum data is processed to achieve data optimization, which is more conducive to subsequent data analysis and obtaining accurate component analysis results.
[0135] The present invention also provides a non-volatile computer storage medium, and the computer storage medium stores at least one executable instruction, and the executable instruction can execute a chemometric analysis method based on partial least squares in any of the above method embodiments.
[0136] Figure 6 The structural schematic diagram of a computing device according to an embodiment of the present invention is shown, and the specific implementation of the computing device is not limited in the specific embodiments of the present invention.
[0137] AsFigure 6 As shown, the computing device may include: a processor 602, a communications interface 604, a memory 606, and a communication bus 608.
[0138] Among them:
[0139] The processor 602, the communications interface 604, and the memory 606 communicate with each other through the communication bus 608.
[0140] The communications interface 604 is used to communicate with network elements of other devices such as clients or other servers.
[0141] The processor 602 is used to execute the program 610, and specifically can execute the relevant steps in the above-mentioned embodiments of the chemometric analysis method based on partial least squares.
[0142] Specifically, the program 610 may include program code, and the program code includes computer operation instructions.
[0143] The processor 602 may be a central processing unit CPU, or a specific integrated circuit ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present invention. One or more processors included in the computing device may be of the same type of processor, such as one or more CPUs; or may be of different types of processors, such as one or more CPUs and one or more ASICs.
[0144] The memory 606 is used to store the program 610. The memory 606 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk memory.
[0145] The program 610 is specifically used to cause the processor 602 to execute a chemometric analysis method based on partial least squares in any of the above method embodiments. For the specific implementation of each step in the program 610, reference may be made to the corresponding steps and units in the above-mentioned embodiments of the chemometric analysis method based on partial least squares, which will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding process descriptions in the foregoing method embodiments, which will not be elaborated here.
[0146] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other apparatus. Various general-purpose systems may also be used in conjunction with the teachings based hereon. The structure required to construct such systems will be apparent from the above description. In addition, the present invention is not directed to any particular programming language. It should be appreciated that the present invention as described herein may be implemented in various programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present invention.
[0147] In the specification provided herein, numerous specific details are set forth. However, it can be understood that embodiments of the present invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0148] Similarly, it should be understood that in order to streamline this disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the claims reflect, the inventive aspects lie in less than all the features of the preceding single embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the present invention.
[0149] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) can be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0150] In addition, those skilled in the art can understand that although some of the embodiments described herein include certain features included in other embodiments rather than other features, the combination of features of different embodiments means that it is within the scope of the present invention and forms different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0151] Each component embodiment of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components according to the embodiments of the present invention. The present invention can also be implemented as a device or apparatus program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0152] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A chemometric analysis method based on partial least squares method, comprising: Performing spectral analysis on a detection object to obtain the original spectral data of the detection object; Preprocessing the collected original spectral data to obtain the processed spectral data after preprocessing; Performing feature extraction based on the spectral data after preprocessing, constructing latent variables based on preset conditions for independent variables and dependent variables, so as to complete the construction of a component analysis model; Using the component analysis model to detect the components of the detection object to obtain a component detection result.
2. The method according to claim 1, characterized in that, The performing spectral analysis on a detection object to obtain the original spectral data of the detection object further includes: When the detection object is in a gas phase, directly using an infrared spectrometer to perform spectral analysis on the components to obtain the original spectral data corresponding to the detection object; When the detection object is in a liquid phase, using the headspace injection method, heating the detection object to volatilize it, and sucking it into the infrared spectrometer to obtain the original spectral data corresponding to the detection object; When the detection object is in a solid phase, using the tablet pressing method to make a potassium bromide tablet, and then putting it into the optical chamber of the infrared spectrometer to obtain the original spectral data corresponding to the detection object.
3. The method according to claim 1, wherein The preprocessing at least includes centering and / or standardization; The preprocessing the collected original spectral data further includes: The centering is to subtract the mean value corresponding to the spectral data from each variable in the spectral data, so that the average value of the processed data is 0; The standardization is to divide each variable in the spectral data by the standard deviation corresponding to the spectral data, so that the variance of the processed spectral data is 1.
4. The method according to claim 1, wherein The performing feature extraction based on the spectral data after preprocessing, constructing latent variables based on preset conditions for independent variables and dependent variables, so as to complete the construction of a component analysis model further includes: Based on the maximum covariance theory, making the covariance between the latent variable that is a linear combination of independent variables and the dependent variable the largest, and thus determining the first latent variable according to the first independent variable; wherein, Let X be an n-dimensional independent variable matrix, Y be a dependent variable vector, and t1 be the first latent variable. t1 is a linear combination of X, i.e., t1 = Xw1, where w1 is the n-dimensional first weight vector corresponding to the first latent variable t1. Based on the maximum covariance theory, the covariance between t1 and Y is max[Cov(t1, Y)], and we can obtain w1 = agrmax[Cov(Xw1, Y)]. Therefore, Furthermore, it can be known that Combined with the constraint condition ‖w1‖ = 1, the analytical solution of w1 is calculated to obtain the first latent variable t1. Among them, Cov is the covariance, max[Cov(t1, Y)] is the maximum covariance between the first latent variable t1 and the dependent variable Y, argmax is the maximum value parameter function, is the transpose of w1, and X T is the transpose of X; Regression modeling of the dependent variable Y based on the first latent variable t1 The predicted value of the dependent variable is calculated Based on the first independent variable, the second independent variable X2 is calculated through residuals; the second latent variable t2 is regressed based on the second independent variable X2, that is And based on the maximum covariance theory to construct the second latent variable t2, that is, t2 = X2w2. By calculating the analytical solution of the second weight vector w2, the second latent variable t2 is obtained; repeat this step until the preset number of latent variables is constructed; where p1 is the linear regression coefficient of the first latent variable t1 on the dependent variable Y; the preset number of latent variables is k; Based on a preset number of latent variables, constructing a multiple linear regression model, and determining the estimated values of the regression coefficients and the intercept term based on the least squares method to obtain a component analysis model; wherein, the multiple linear regression model is Among them, b0 is the intercept term, and b1, b2... b k are the regression coefficients corresponding to the respective latent variables.
5. The method according to claim 1, wherein The using the component analysis model to detect the components of the detection object to obtain a component detection result further includes: Using the component analysis model to detect the components of the detection object to obtain a component detection result; Performing back-calculation based on the result to judge whether the preset requirements are met; If the preset requirements are met, the component detection result is sent out; if the preset requirements are not met, the component analysis model is reconstructed.
6. The method according to claim 1, characterized in that The method further includes: Extracting at least two detection data points from the original spectral data and obtaining the target parameters of the detection data points; Calculating the structural similarity between the detection data points based on the target parameters of the detection data points; Determining the characteristic bands and abnormal data points in the spectral data according to the structural similarity between the detection data points, and performing data processing on the abnormal data points.
7. The method according to claim 6, wherein The method further includes: The target parameters at least include: amplitude, frequency, and depth information; The structural similarity between the detection data points is Among them, F a,b is the structural similarity function; DTW is the DTW distance between detected data points; (d, l) α is the depth information of the α-th detected data point; (d, l) β is the depth information of the β-th detected data point; |Δf| is the frequency difference between detected data points; |Δa| is the amplitude difference between detected data points.
8. A chemometric analysis device based on partial least squares method, comprising: A spectral acquisition module, a data preprocessing module, a model construction module, and a component detection module; wherein, the spectral acquisition module is configured to perform spectral analysis on a detection object to obtain the original spectral data of the detection object; the data preprocessing module is configured to perform preprocessing on the collected original spectral data to obtain the processed spectral data after preprocessing; the model construction module is configured to perform feature extraction based on the spectral data after preprocessing, and construct latent variables based on preset conditions for independent variables and dependent variables to complete the construction of a component analysis model; the component detection module is configured to use the component analysis model to detect the components of the detection object to obtain a component detection result.
9. A computing device, comprising: A processor, a memory, a communication interface, and a communication bus, and the processor, the memory, and the communication interface complete communication with each other through the communication bus; the memory is used to store at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to a chemometric analysis method based on partial least squares as described in any one of claims 1-7.
10. A computer storage medium, wherein the storage medium stores at least one executable instruction, and the executable instruction causes the processor to perform operations corresponding to a chemometric analysis method based on partial least squares as described in any one of claims 1-7.
Citation Information
Patent Citations
Sample component determination method based on optimizing partial least squares regression model
CN104949936A
Weighted modeling local optimization method for spectral baseline correction
CN111999258A
Blueberry soluble solid detection method, device, equipment and medium
CN117668766A
Hyperspectral water quality parameter inversion model based on FOD and optimal spectral characteristics
CN117892633A
Method and system for extracting net signal in near-infrared spectrum
WO2023123329A1