Method for rapidly detecting content of betacyanin based on spectrophotometric method
Through a rapid detection method based on spectrophotometry, combined with multivariate correction, spectral feature extraction and pattern recognition technology and deep learning technology, the beet erection content is detected, and the final results are obtained through data fusion and ensemble learning, which solves the problem of spectral interference of traditional detection methods and achieves efficient and accurate beet erection content detection.
Patent Information
- Application Number
- CN202510082787.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-06-06
AI Technical Summary
Traditional beet melanin content detection methods are susceptible to spectral interference from other substances, and it is difficult to accurately extract the characteristic information of beet melanin, thereby affecting the accuracy of content measurement.
The rapid detection method based on spectrophotometry is adopted to obtain spectral data through full-spectrum scanning, pre-process and multivariate correction, and combine spectral feature extraction and pattern recognition technology and deep learning technology to detect beet red pigment content, and the final results are obtained through data fusion and ensemble learning.
It realizes rapid and accurate detection of beet melanin content, improves detection efficiency and accuracy, reduces the impact of performance fluctuations in a single method on the results, and ensures the stability and reliability of the detection results.
Smart Images

Figure CN120102483A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of betalain content detection, and in particular to a rapid betalain content detection method based on spectrophotometry. Background Art
[0002] In many industries such as food, medicine, and cosmetics, betalain is an important natural pigment, and accurate detection of its content is of vital importance. Spectrophotometry has become one of the commonly used methods for detecting betalain content due to its advantages such as widespread equipment and relatively simple operation.
[0003] In actual samples, in addition to betalain, they often contain a variety of other substances, whose spectra overlap or interfere with the spectrum of betalain in some areas, making it difficult for traditional analysis methods based on a single wavelength or simple absorbance ratio to accurately extract the characteristic information of betalain, thus affecting the accuracy of content determination. Summary of the invention
[0004] In view of the shortcomings of the prior art, the present invention provides a method for rapid detection of betalain content based on spectrophotometry, which solves the problem that traditional detection methods are easily affected by spectral interference of other substances, making it difficult to accurately extract the characteristic information of betalain, thereby affecting the accuracy of content determination.
[0005] To achieve the above objectives, the present invention is implemented by the following technical scheme: a method for rapid detection of betalain content based on spectrophotometry, comprising the following steps:
[0006] S1. Experimental operation: Prepare the sample into standard solutions of different concentrations, use a spectrophotometer to perform full spectrum scanning on the standard solution, and record the absorbance values at different wavelengths to form a spectral data matrix;
[0007] S2, preprocessing: preprocess the spectral data matrix to obtain basic data;
[0008] S3, first analysis: using multivariate correction technology to analyze the basic data, thereby obtaining first data;
[0009] S4, second analysis: using spectral feature extraction and pattern recognition technology to analyze the basic data, thereby obtaining second data;
[0010] S5, third analysis: using deep learning technology to analyze the detection data to obtain third data;
[0011] S6. Fusion and integration: The spectral data matrix is subjected to data fusion and integrated learning according to the first data, the second data, and the third data, so as to obtain the betalain content.
[0012] Preferably, the scanning range of the full spectrum scanning in S1 is 350 nm to 750 nm, and the wavelength interval is set to 1 nm.
[0013] Preferably, the pre-processing in S2 comprises the following steps:
[0014] S201, performing baseline correction on the spectral data matrix and fitting using a polynomial function;
[0015] S202, then performing Savitzky-Golay convolution smoothing processing on it;
[0016] S203, normalizing the smoothed spectral data matrix to obtain basic data.
[0017] Preferably, the polynomial function of the baseline correction is P(λ)=a 0 +α 1 λ+α 2 λ 2 +...+α k λ k Where P(λ) is the baseline fitting value at wavelength λ, λ is the wavelength, a 0 , a 1 , a 2 , ..., a k is the polynomial coefficient, k is the order of the polynomial, and the polynomial coefficient is determined by the least squares method so that ∑(A ij -P(λ) 2 Minimum, the convolution smoothing process uses a smoothing filter with a window size of w, according to the formula Calculate, where s = (w-1) / 2, c k are polynomial coefficients, and the normalization process is performed according to the formula Carry out, where A″ min , A″ max The minimum and maximum values in the spectral data matrix are obtained in turn to obtain the basic data.
[0018] Preferably, the step of analyzing the basic data using multivariate correction technology in S3 includes the following steps:
[0019] S301, taking the basic data as the independent variable matrix and the corresponding standard solution concentration vector as the dependent variable vector, and decomposing them by partial least squares method;
[0020] S302, using a leave-one-out cross-validation method to determine the number of principal components, and performing variable importance projection analysis, thereby obtaining a first model, and then inputting the spectral data matrix into the first model for content prediction, thereby obtaining first data.
[0021] Preferably, the step of analyzing the basic data by using spectral feature extraction and pattern recognition technology in S4 includes the following steps:
[0022] S401, perform principal component analysis feature extraction on the spectral data of the standard solution and the interfering substance in the basic data, select the first p principal components so that the cumulative contribution rate reaches more than 90%, and use the projection of the sample on these p principal components as a new feature vector;
[0023] S402, performing support vector machine classification on known betalain samples and interfering substances, using radial basis function kernel function, and optimizing parameters through grid search method and cross validation;
[0024] S403, constructing a support vector machine regression model for the principal component eigenvectors and concentrations of standard solutions of different concentrations, using a radial basis function kernel function, optimizing the model by minimizing the mean square error, and inputting the principal component eigenvectors of the samples into the support vector machine regression model to obtain second data.
[0025] Preferably, the step of analyzing the detection data using deep learning technology in S5 includes the following steps:
[0026] S501. Construct a convolutional neural network model with the spectral data matrix as input, including convolution layer, pooling layer, fully connected layer and output layer. In the convolution layer operation, the hyperparameters are determined by optimization methods such as random search or genetic algorithm to minimize the mean square error loss function. Where n is the number of samples, C i,pred is the betalain content value predicted by the model for the i-th sample, C i is the true betacyanin content value of the i-th sample, and the hyperparameters include convolution kernel size, step size, pooling window size, and number of neurons;
[0027] S502, training and predicting the convolutional neural network model, including dividing the standard solution spectral data and concentration data into a training set, a validation set and a test set, using a back propagation algorithm and a stochastic gradient descent optimizer to update the model parameters, and inputting the spectral data of the sample into the trained convolutional neural network model for prediction and visualization, thereby obtaining third data.
[0028] Preferably, the data fusion and ensemble learning in S6 is to determine the respective weights ω according to the performance of the first data, the second data, and the third data on the validation set. PLS ,ω PCA-SVM and ω CNN , and perform weighted average calculation, that is, Thus betalain content is obtained, and the properties include R 2 Value, MSE value.
[0029] The present invention provides a method for rapid detection of betalain content based on spectrophotometry. It has the following beneficial effects:
[0030] 1. The present invention forms a complete detection process from data collection, preprocessing, different technical analysis to final fusion integration. Each step provides better data or more accurate analysis results for subsequent steps, and together realizes rapid and accurate detection of betalain content, improves detection efficiency and accuracy, thereby solving the problem that traditional detection methods are easily affected by spectral interference of other substances, making it difficult to accurately extract the characteristic information of betalain, thereby affecting the accuracy of content determination.
[0031] 2. The present invention ensures the stability of the model under different samples and experimental conditions through standardized operations of data preprocessing and model optimization methods in various analysis techniques. Data fusion and ensemble learning determine the weights of each method based on its performance on the validation set, further reducing the impact of performance fluctuations of a single method on the final result, thereby providing reliable and consistent results. Whether it is testing samples from different batches or using them in different laboratories, the stability of detection accuracy can be guaranteed, which helps to establish unified quality inspection standards and processes.
[0032] 3. In the data collection stage, the present invention obtains rich information through full spectrum scanning, effectively reduces interference through preprocessing, handles multivariate relationships and collinearity problems through multivariate correction technology, accurately extracts features and performs classification and regression through spectral feature extraction and pattern recognition technology, and automatically learns complex features through deep learning technology. Finally, the results of the three methods are integrated to comprehensively consider their respective advantages, fully mine the effective information in the spectral data, and reduce errors caused by the limitations of a single method or data interference. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a stereogram of the present invention. DETAILED DESCRIPTION
[0034] The technical solution of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0035] Example:
[0036] Please see attached Figure 1 The embodiment of the present invention provides a method for rapid detection of betalain content based on spectrophotometry, comprising the following steps:
[0037] S1. Experimental operation: Prepare the sample into standard solutions of different concentrations, use a spectrophotometer to perform full spectrum scanning on the standard solution, and record the absorbance values at different wavelengths to form a spectral data matrix;
[0038] S2, preprocessing: preprocess the spectral data matrix to obtain basic data;
[0039] S3, first analysis: using multivariate correction technology to analyze the basic data, thereby obtaining first data;
[0040] S4, second analysis: using spectral feature extraction and pattern recognition technology to analyze the basic data, thereby obtaining second data;
[0041] S5, third analysis: using deep learning technology to analyze the detection data to obtain third data;
[0042] S6. Fusion and integration: The spectral data matrix is subjected to data fusion and integrated learning according to the first data, the second data, and the third data, so as to obtain the betalain content.
[0043] Specifically, by preparing the samples into standard solutions of different concentrations, using a spectrophotometer to perform a full spectrum scan of the standard solutions, and recording the absorbance values at different wavelengths to form a spectral data matrix, rich spectral information is obtained, covering the absorbance characteristics of betalain at various wavelengths at different concentrations, providing sufficient data basis for subsequent accurate analysis. The use of multi-concentration standard solutions ensures that when the analysis model is subsequently established, the model can adapt to the detection of betalain content levels, thereby improving the versatility and accuracy of the detection method and avoiding detection errors caused by a narrow concentration range.
[0044] By preprocessing the spectral data matrix, basic data is obtained, thereby reducing the impact of non-target factors on the spectral data, making subsequent data-based analysis more reliable and improving the accuracy of betalain content detection.
[0045] By adopting multivariate correction technology to analyze the basic data, the first data is obtained, so as to simultaneously consider the complex relationship between multiple independent variables (wavelengths) and dependent variables (concentrations), fully mine the useful information in the spectral data, and improve the prediction ability of the model.
[0046] By using spectral feature extraction and pattern recognition technology to analyze the basic data, the secondary data is obtained, thereby reducing the data dimension, retaining the main information, removing redundant information, improving the efficiency and accuracy of subsequent analysis, identifying betalain and interfering substances, and improving the prediction accuracy of betalain content.
[0047] By using deep learning technology to analyze the detection data, the third data is obtained, and complex feature representations are learned from the spectral data. Without manual feature extraction, the betalain content can be predicted with high precision. This is particularly suitable for processing large-scale, high-dimensional spectral data, which improves the accuracy and efficiency of detection.
[0048] By fusing and integrating the spectral data matrix according to the first data, the second data, and the third data, the betalain content is obtained. Therefore, according to the prediction results of the three different methods of the first data (PLS), the second data (PCA-SVM), and the third data (CNN), the advantages of each method are fully utilized, the shortcomings of each method are compensated, the overall detection performance is improved, and the error caused by model bias or data adaptability problems of a single method is reduced. The accuracy, stability and reliability of the betalain content detection results are improved, the detection results are closer to the true value, and a more reliable basis is provided for practical applications.
[0049] Through the coordinated implementation of each step, a complete detection process is formed from data collection, preprocessing, different technical analysis to final fusion integration. Each step provides better data or more accurate analysis results for the subsequent steps, and together realizes the rapid and accurate detection of betalain content, improves the detection efficiency and accuracy, and thus solves the problem that traditional detection methods are easily affected by spectral interference from other substances, making it difficult to accurately extract the characteristic information of betalain, thereby affecting the accuracy of content determination.
[0050] The scanning range of the full spectrum scan in S1 is 350nm to 750nm, and the wavelength interval is set to 1nm.
[0051] The preprocessing in S2 includes the following steps:
[0052] S201, performing baseline correction on the spectral data matrix and fitting using a polynomial function;
[0053] S202, then performing Savitzky-Golay convolution smoothing processing on it;
[0054] S203, normalizing the smoothed spectral data matrix to obtain basic data.
[0055] The polynomial function for baseline correction is P(λ)=a 0 +α 1 λ+α 2 λ 2 +...+α k λ k Where P(λ) is the baseline fitting value at wavelength λ, λ is the wavelength, a 0 , a1 , a 2 , ..., a k is the polynomial coefficient, k is the order of the polynomial, and the polynomial coefficient is determined by the least squares method so that ∑(A ij -P(λ) 2 Minimum, convolution smoothing uses a smoothing filter with a window size of w, according to the formula Calculate, where s = (w-1) / 2, c k are polynomial coefficients, and the normalization process is based on the formula Carry out, where A″ min , A″ max The minimum and maximum values in the spectral data matrix are obtained in turn to obtain the basic data.
[0056] Specifically, by performing baseline correction on the spectral data matrix and fitting it with a polynomial function, the baseline drift is corrected, allowing the spectrum to truthfully reflect the absorption characteristics of betalain, laying the foundation for accurate content measurement. By performing Savitzky-Golay convolution smoothing on it, noise is reduced and the signal-to-noise ratio is improved, highlighting the betalain absorption peak, and facilitating subsequent feature extraction and quantitative analysis, increasing detection accuracy. By normalizing the smoothed spectral data matrix, basic data is obtained, thereby unifying the data scale, eliminating the impact of measurement differences, enhancing data comparability, and using the model to reasonably assign weights to improve detection accuracy.
[0057] The analysis of basic data using multivariate correction techniques in S3 includes the following steps:
[0058] S301, taking the basic data as the independent variable matrix and the corresponding standard solution concentration vector as the dependent variable vector, and decomposing them by partial least squares method;
[0059] S302, using a leave-one-out cross-validation method to determine the number of principal components, and performing variable importance projection analysis, thereby obtaining a first model, and then inputting the spectral data matrix into the first model for content prediction, thereby obtaining first data.
[0060] Specifically, the basic data are taken as the independent variable matrix and the corresponding standard solution concentration vector as the dependent variable vector, and the partial least squares method is used for decomposition. Thus, the basic data are taken as the independent variable matrix X and the standard solution concentration vector as the dependent variable vector Y, and PLS decomposition is used to calculate the score, loading, weight vector and orthogonal decomposition, thereby integrating multivariate information dimensionality reduction, retaining concentration-related information, and laying the foundation for modeling.
[0061] The number of principal components is determined by using the leave-one-out cross-validation method, and variable importance projection analysis is performed at the same time to obtain the first model. The spectral data matrix is then input into the first model for content prediction to obtain the first data, thereby optimizing the model to improve generalization and prediction capabilities, screening key variables to simplify the model, and improving detection accuracy.
[0062] The analysis of basic data using spectral feature extraction and pattern recognition technology in S4 includes the following steps:
[0063] S401, perform principal component analysis feature extraction on the spectral data of the standard solution and the interfering substance in the basic data, select the first p principal components so that the cumulative contribution rate reaches more than 90%, and use the projection of the sample on these p principal components as a new feature vector;
[0064] S402, performing support vector machine classification on known betalain samples and interfering substances, using radial basis function kernel function, and optimizing parameters through grid search method and cross validation;
[0065] S403, constructing a support vector machine regression model for the principal component eigenvectors and concentrations of standard solutions of different concentrations, using a radial basis function kernel function, optimizing the model by minimizing the mean square error, and inputting the principal component eigenvectors of the samples into the support vector machine regression model to obtain second data.
[0066] Specifically, by performing principal component analysis and feature extraction on the spectral data of standard solutions and interfering substances in the basic data, the first p principal components are selected so that the cumulative contribution rate reaches more than 90%. The projection of the sample on these p principal components is used as a new feature vector, thereby achieving data dimensionality reduction and feature simplification, reducing data complexity while retaining the main spectral information, highlighting the betalain-related characteristics, and enhancing the efficiency and accuracy of subsequent analysis.
[0067] By performing support vector machine classification on known betalain samples and interfering substances, using radial basis function kernel function, and optimizing parameters through grid search method and cross-validation, betalain and interfering substances can be effectively distinguished, the sample category can be accurately determined, the classification accuracy and reliability can be improved, and a qualitative basis can be provided for subsequent quantitative analysis.
[0068] A support vector machine regression model is constructed by using the principal component eigenvectors and concentrations of standard solutions of different concentrations. The radial basis function kernel function is used to optimize the model by minimizing the mean square error. The principal component eigenvector of the sample is input into the support vector machine regression model to obtain the second data, thereby establishing a quantitative relationship between the eigenvector and the concentration and accurately predicting the betalain content.
[0069] The use of deep learning technology to analyze detection data in S5 includes the following steps:
[0070] S501. Construct a convolutional neural network model with the spectral data matrix as input, including convolution layer, pooling layer, fully connected layer and output layer. In the convolution layer operation, the hyperparameters are determined by optimization methods such as random search or genetic algorithm to minimize the mean square error loss function. Where n is the number of samples, C i,pred is the betalain content value predicted by the model for the i-th sample, C i is the true betacyanin content value of the i-th sample, and the hyperparameters include convolution kernel size, step size, pooling window size, and number of neurons;
[0071] S502, training and predicting the convolutional neural network model, including dividing the standard solution spectral data and concentration data into a training set, a validation set and a test set, using a back propagation algorithm and a stochastic gradient descent optimizer to update the model parameters, and inputting the spectral data of the sample into the trained convolutional neural network model for prediction and visualization, thereby obtaining third data.
[0072] Specifically, a convolutional neural network model is constructed with the spectral data matrix as input, thereby constructing a model framework that can automatically learn complex feature representations from spectral data, and optimizing hyperparameters to improve the accuracy of the model in predicting betalain content. The convolutional neural network model is trained and predicted, and the model parameters are updated using the back propagation algorithm and the stochastic gradient descent optimizer. The spectral data of the sample is input into the trained convolutional neural network model for prediction and visualization to obtain the third data. Through training and optimization, the model learns the relationship between spectral data and betalain content, and the prediction and visualization provide intuitive result display, thereby improving the prediction accuracy and reliability.
[0073] The data fusion and ensemble learning in S6 is to determine the weights ω of the first data, the second data, and the third data according to their performance on the validation set. PLS ,ω PCA-SVM and ω CNN , and perform weighted average calculation, that is, Thus, the betalain content is obtained, and the properties include R 2 Value, MSE value.
[0074] Specifically, by determining the respective weights according to the performance of the first data, the second data, and the third data on the validation set, and performing weighted average calculation, the betacyanin content is obtained, thereby integrating the results of multiple methods and giving full play to the advantages of each method, such as PLS processing multivariate relationships, PCA-SVM dealing with spectral interference, and CNN automatic learning features, so that the final result is more accurate and reliable, and the advantages of multiple methods are complementary, and the detection accuracy is improved. By considering the performance of each method on the validation set to determine the weight, the error caused by model bias or data adaptability problems of a single method is reduced, and the stability of the detection result is enhanced.
[0075] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for rapid detection of betalain content based on spectrophotometry, characterized in that: The following steps are involved: S1. Experimental operation: Prepare the sample into standard solutions of different concentrations, use a spectrophotometer to perform full spectrum scanning on the standard solution, and record the absorbance values at different wavelengths to form a spectral data matrix; S2, preprocessing: preprocess the spectral data matrix to obtain basic data; S3, first analysis: using multivariate correction technology to analyze the basic data, thereby obtaining first data; S4, second analysis: using spectral feature extraction and pattern recognition technology to analyze the basic data, thereby obtaining second data; S5, third analysis: using deep learning technology to analyze the detection data to obtain third data; S6. Fusion and integration: The spectral data matrix is subjected to data fusion and integrated learning according to the first data, the second data, and the third data, so as to obtain the betalain content.
2. The method for rapid detection of betalain content based on spectrophotometry according to claim 1, characterized in that: The scanning range of the full spectrum scanning in S1 is 350 nm to 750 nm, and the wavelength interval is set to 1 nm.
3. The method for rapid detection of betalain content based on spectrophotometry according to claim 1, characterized in that: The pre-processing in S2 comprises the following steps: S201, performing baseline correction on the spectral data matrix and fitting using a polynomial function; S202, then performing Savitzky-Golay convolution smoothing processing on it; S203, normalizing the smoothed spectral data matrix to obtain basic data.
4. The method for rapid detection of betalain content based on spectrophotometry according to claim 3, characterized in that: The polynomial function of the baseline correction is P(λ)=a0+a1λ+α2λ 2 +...+a k λ k , where P(λ) is the baseline fitting value at wavelength λ, λ is the wavelength, a0, a1, a2, ..., a k is the polynomial coefficient, k is the order of the polynomial, and the polynomial coefficient is determined by the least squares method so that ∑(A ij -P(λ) 2 Minimum, the convolution smoothing process uses a smoothing filter with a window size of w, according to the formula Calculate, where s = (w-1) / 2, c k are polynomial coefficients, and the normalization process is performed according to the formula Carry out, where A″ min , A″ max The minimum and maximum values in the spectral data matrix are obtained in turn to obtain the basic data.
5. The method for rapid detection of betalain content based on spectrophotometry according to claim 1, characterized in that: The use of multivariate correction technology to analyze basic data in S3 includes the following steps: S301, taking the basic data as the independent variable matrix and the corresponding standard solution concentration vector as the dependent variable vector, and decomposing them by partial least squares method; S302, using a leave-one-out cross-validation method to determine the number of principal components, and performing variable importance projection analysis, thereby obtaining a first model, and then inputting the spectral data matrix into the first model for content prediction, thereby obtaining first data.
6. The method for rapid detection of betalain content based on spectrophotometry according to claim 1, characterized in that: The analysis of basic data using spectral feature extraction and pattern recognition technology in S4 includes the following steps: S401, perform principal component analysis feature extraction on the spectral data of the standard solution and the interfering substance in the basic data, select the first p principal components so that the cumulative contribution rate reaches more than 90%, and use the projection of the sample on these p principal components as a new feature vector; S402, performing support vector machine classification on known betalain samples and interfering substances, using radial basis function kernel function, and optimizing parameters through grid search method and cross validation; S403, constructing a support vector machine regression model for the principal component eigenvectors and concentrations of standard solutions of different concentrations, using a radial basis function kernel function, optimizing the model by minimizing the mean square error, and inputting the principal component eigenvectors of the samples into the support vector machine regression model to obtain second data.
7. The method for rapid detection of betalain content based on spectrophotometry according to claim 1, characterized in that: The use of deep learning technology to analyze the detection data in S5 includes the following steps: S501. Construct a convolutional neural network model with the spectral data matrix as input, including convolution layer, pooling layer, fully connected layer and output layer. In the convolution layer operation, the hyperparameters are determined by optimization methods such as random search or genetic algorithm to minimize the mean square error loss function. Where n is the number of samples, C i,pred is the betalain content value predicted by the model for the i-th sample, C i is the true betacyanin content value of the i-th sample, and the hyperparameters include convolution kernel size, step size, pooling window size, and number of neurons; S502, training and predicting the convolutional neural network model, including dividing the standard solution spectral data and concentration data into a training set, a validation set and a test set, using a back propagation algorithm and a stochastic gradient descent optimizer to update the model parameters, and inputting the spectral data of the sample into the trained convolutional neural network model for prediction and visualization, thereby obtaining third data.
8. The method for rapid detection of betalain content based on spectrophotometry according to claim 1, characterized in that: The data fusion and ensemble learning in S6 is to determine the respective weights w according to the performance of the first data, the second data, and the third data on the validation set. PLS 、w PCA-SVM and w CNN , and perform weighted average calculation, that is, Thus betalain content is obtained, and the properties include R 2 Value, MSE value.
Citation Information
Cited By
Method for predicting and analyzing anthocyanin content of red onion based on CIELab color quantization
CN120369645A
Method for rapidly and nondestructively detecting quality of beet seeds based on near infrared spectrum technology
CN121324300A
Mine sample analysis method and system based on spectrum correction
CN121438113A