A method for identifying and quantifying multi-type starch adulterated sweet potato starch
By using near-infrared spectroscopy and chemometrics, a 1D-CNN model was established, which solved the problem of rapid identification and quantification of adulteration of various potato starches. This enabled efficient and accurate starch type identification and adulteration detection, thereby improving food safety and market order.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INST OF AGRI QUALITY STANDARDS & TESTING TECH RES HUBEI ACADEMY OF AGRI SCI
- Filing Date
- 2025-04-09
- Publication Date
- 2026-05-22
AI Technical Summary
Existing technologies are insufficient for the rapid and accurate identification and quantification of various types of potato starch adulteration. Traditional methods are time-consuming, complex to operate, and environmentally unfriendly, and there is a lack of effective near-infrared spectroscopy combined with chemometrics.
Near-infrared spectroscopy combined with chemometrics was employed. Spectral data of potato starch were acquired using a Frontier near-infrared Fourier transform spectrometer. Preprocessing was performed using CWT or 1st method, and a 1D-CNN model was established to classify, identify, and quantify sweet potato starch, potato starch, corn starch, and cassava starch.
It enables rapid and accurate identification of pure and non-pure sweet potato starch, improves the accuracy and efficiency of starch type identification, combats counterfeit and adulterated starch in the market, and ensures food safety.
Smart Images

Figure CN120334171B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of food testing technology, specifically relating to a method for identifying and quantifying sweet potato starch adulterated with various types of starch. Background Technology
[0002] Starch, as a core food reserve substance in plants and their derivatives, is one of the most abundant types of carbohydrates in nature. In the human diet, starch dominates in terms of calories and dietary energy, and is a key substrate for metabolic activities. Due to its abundance, low cost, non-toxicity, and biodegradability, it has received special attention from various food and non-food industries. In the food industry, it is frequently used in many dairy products, baked goods, soups and sauces, and meat products. However, because of the significant price differences between different types of starch, some unscrupulous merchants attempt to adulterate expensive starches with cheaper ones to obtain exorbitant profits. For example, sweet potato starch, extracted from sweet potatoes, has wide applications in the food, paper, and textile industries. Sweet potato starch, along with corn starch, potato starch, and cassava starch, constitutes the main edible starches in my country. The prices of different types of starch vary greatly in the market; cassava starch and corn starch are inexpensive, costing less than 2 yuan per kilogram, while sweet potato starch is often more than twice the price of cassava starch per kilogram. Therefore, many small workshops and manufacturers use tapioca starch, corn starch, potato starch, etc., to replace some of the sweet potato starch during production and processing in order to reduce production costs. Adulterated sweet potato starch and vermicelli products not only seriously disrupt market order and harm consumers' legitimate rights and interests, but also pose potential health risks. Therefore, research on the identification technology of adulterated starch and its products is of great significance for maintaining the order of the domestic and international starch market and ensuring food safety.
[0003] Different types of starch do not differ significantly in appearance and are difficult to distinguish with the naked eye. Conventional methods for starch identification mainly fall into two categories: morphological observation and physicochemical property analysis. The former utilizes methods such as scanning electron microscopy (SEM) to directly observe the morphology of starch granules, while the latter relies on the analysis of the ratio of amylose to amylopectin and the assessment of physicochemical properties such as gelatinization behavior. Traditional analytical methods have advantages in accuracy, but their disadvantages include long testing cycles, complex operations, sample damage, and environmental pollution. The limitations of current detection methods pose new challenges to the standardization of food starch detection indicators, prompting an urgent need to establish a rapid, accurate, simple, and environmentally friendly starch classification technology.
[0004] Near-infrared spectroscopy (NIR) technology, as an analytical tool, exhibits remarkable advantages such as speed, convenience, non-destructive nature, and high efficiency, providing strong support for a variety of analytical tasks. It is not only easy to operate but also allows for rapid and effective determination of the composition or properties of substances without damaging the sample, making it a highly favored technique in scientific research and industrial testing. In recent years, near-infrared spectroscopy has been widely used in food quality testing due to its ability to acquire characteristic information about food components. Chemometrics, as an important branch of spectroscopy, helps to significantly reduce spectral noise levels, enhance the precision of analysis, effectively eliminate interfering factors, and deeply uncover valuable information hidden within spectral data, ultimately ensuring the accuracy of analytical results.
[0005] Significant progress has been made in identifying and classifying samples using near-infrared spectroscopy combined with chemometrics to address food adulteration. However, rapid and accurate methods for identifying and quantifying various types of tuber starch, especially sweet potato starch, are currently lacking. Summary of the Invention
[0006] In view of the above-mentioned shortcomings of the prior art, the key point of the present invention lies in the extraction and classification of characteristic components of potato starch, aiming to establish a rapid and accurate method for identifying starch types under various types of potato starch adulteration, and providing a method for identifying and quantifying sweet potato starch adulterated with multiple types of starch.
[0007] The present invention provides a method for identifying and quantifying various types of adulterated sweet potato starch, comprising the following steps:
[0008] S1. Take four pure samples: sweet potato starch, potato starch, cassava starch, and corn starch, and prepare binary and ternary mixture samples containing sweet potato starch.
[0009] S2. Acquire the sample spectrum using a Frontier near-infrared Fourier spectrometer to obtain the raw spectral data of the sample;
[0010] S3. Randomly divide the training set and test set into a 7:3 ratio, preprocess the original spectral data of the samples using CWT or 1st method, establish a classification and quantitative prediction model for the training set using PLS or 1D-CNN method, and verify and evaluate the model in conjunction with the test set.
[0011] S4. Use the classification and quantitative prediction model established in S3 to classify, identify and quantitatively predict the content of unknown starch samples.
[0012] Preferably, in the binary mixture sample, the mass ratio of sweet potato starch is 10%-90%, with the remainder being potato starch, tapioca starch, or corn starch; in the ternary mixture sample, the mass ratio of sweet potato starch is 50%, with the remainder being any two of potato starch, tapioca starch, and corn starch, each with a mass ratio of 10%-40%. The content of each starch in the binary mixture sample can be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, or 90%. In the ternary mixture sample, the mass ratio of sweet potato starch is 50%, with the remainder being 10%, 20%, 30%, or 40% of each starch.
[0013] Preferably, step S2 is as follows: The sample spectrum is acquired using a Frontier near-infrared Fourier spectrometer, with a sample amount of 2g. The near-infrared spectrum is collected by diffuse reflectance analysis using an integrating sphere, with a near-infrared wavenumber range of 4000–10000 cm⁻¹. -1 Scan resolution 4cm -1 The number of scans was 32; 10 samples were replicated for each sample, and the spectrum of each sample was collected 3 times to obtain the original spectral data of the sample.
[0014] Preferably, step S3 is as follows: randomly divide the training set and the test set according to a 7:3 ratio, preprocess the original spectral data of the samples using the CWT method, establish a classification and quantitative prediction model for the training set using the 1D-CNN method, and verify and evaluate the model in conjunction with the test set.
[0015] Preferably, the classification prediction model is a five-category, four-category, or two-category prediction model. The five-category model includes pure sweet potato starch, potato starch, cassava starch, corn starch, and mixed starch. The four-category model includes pure sweet potato starch, potato starch, cassava starch, and corn starch. The two-category model includes pure sweet potato starch, potato starch, cassava starch, and corn starch as one category, and mixed starch as another category.
[0016] Preferably, the 1D-CNN method includes a feature extraction layer, one 2×1 convolutional kernel, a 2×1 pooling kernel, an Adam gradient descent algorithm, 500 training iterations, an initial learning rate of 0.001, an L2 regularization parameter, a learning rate reduction factor of 0.1, and a learning rate of 0.0001.
[0017] This invention uses common low-value raw starches such as cassava starch and corn starch as adulterants in sweet potato starch. Near-infrared spectroscopy is used as the detection method to establish a characteristic near-infrared spectral database of sweet potato starch, potato starch, corn starch, cassava starch, etc. Combined with chemometrics, the obtained spectra are preprocessed and multiple identification models are established to achieve the identification of pure sweet potato starch and impure sweet potato starch, and to predict the proportion of sweet potato starch in them.
[0018] The method of this invention has important reference value for the formulation of classification standards for potato starch, laying a solid foundation for the rapid and accurate identification of starch types, which is conducive to combating the counterfeiting and adulteration of starch in the market, guiding the storage and processing of sweet potato starch, helping to screen high-quality potato starch raw materials for research on potato starch products, and improving the overall value and practicality of starch-based products. Attached Figure Description
[0019] Figure 1 These are the spectra of different tuber starches; where a is the original spectrum; b is the average spectrum of different types of spectra, where Gan represents sweet potato starch, Mu represents cassava starch, Tu represents potato starch, Yu represents corn starch, Gan+Tu represents a binary mixture of sweet potato starch and potato starch, Gan+Mu represents a binary mixture of sweet potato starch and cassava starch, Gan+Yu represents a binary mixture of sweet potato starch and corn starch, and Gan+Yu+Tu represents a ternary mixture of sweet potato starch, corn starch, and potato starch; c is the spectrum treated with MSC; d is the spectrum treated with CWT; e is the spectrum treated with 1st; and f is the spectrum treated with SNV.
[0020] Figure 2 These are PCA diagrams of different types of potato starch; where a is the original spectrum; b is the average spectrum of different types of spectra; c is the spectrum processed by MSC; d is the spectrum processed by CWT; e is the spectrum processed by 1st; and f is the spectrum processed by SNV. Among these, SYS represents ternary mixed starch, EYS represents binary mixed starch, CS represents corn starch, PS represents potato starch, MS represents cassava starch, and SPS represents sweet potato starch.
[0021] Figure 3 This is a diagram showing the original spectral classification of five types of starch.
[0022] Figure 4 It is a five-class confusion matrix of different preprocessed spectra.
[0023] Figure 5 This is the classification effect of the original spectrum in the two-class case.
[0024] Figure 6 The classification results are obtained by different preprocessing spectra for two different classification scenarios.
[0025] Figure 7 This is the classification effect of pure potato starch with different spectra.
[0026] Figure 8 This is a graph showing the trend of prediction error in the training set as a function of the number of factors.
[0027] Figure 9These are the predicted results of sweet potato starch content in different spectra.
[0028] Figure 10 This is a five-class classification result based on the 1D-CNN method.
[0029] Figure 11 These are the results of two-class and four-class classification based on the 1D-CNN method.
[0030] Figure 12 The results are based on the 1D-CNN method for predicting sweet potato starch content. Detailed Implementation
[0031] The following embodiments are further illustrations of the present invention, but not limitations thereof.
[0032] Example 1
[0033] 1. Materials and Methods
[0034] 1.1 Materials and Sample Preparation
[0035] The starches used included sweet potato starch, corn starch, potato starch, and cassava starch. The sweet potato starch samples were from different batches produced on the production line of Hubei Jinyue Agricultural Products Development Co., Ltd., with a purity of 100%, totaling 10 samples. The corn starch, potato starch, and cassava starch samples were purchased in batches from the market, with the manufacturers located in Jingshan City (Hubei, China), to ensure the representativeness of the samples. The potato starch, cassava starch, and corn starch samples met the respective standards GB / T8885, GB / T 8884, and GB / T 29343. A total of 10 batches of samples were collected for each type.
[0036] Different batches of pure sweet potato starch, potato starch, cassava starch, and corn starch were pulverized and ground, then passed through a 60-mesh sieve. Three 2g portions of each sample were weighed. In addition to these samples, binary and ternary mixtures were prepared between the pure samples to increase the variability of the dataset. The sweet potato starch content in the binary mixtures ranged from 10% to 90% by mass, with 10 samples of each binary mixture and 3 replicates per sample. The mixing ratios of the adulterated sweet potato starch samples in the ternary mixtures are shown in Table 1. After mixing, 2g of each prepared sample was weighed and mixed thoroughly using a shaker.
[0037] Table 1. Explanation of Sample Quantity and Adulteration Ratio
[0038]
[0039]
[0040] 1.2 Spectral Acquisition
[0041] 2g of the above sample was placed in quartz cups. Against an air background, the near-infrared spectrum of each solid sample was collected at room temperature using diffuse reflectance analysis via an integrating sphere. Spectral acquisition and scanning were performed using a Frontier near-infrared Fourier spectrometer (near-infrared wavenumber range 4000–10000 cm⁻¹). -1 Scan resolution 4cm -1 (32 scans) Perkin Elmer, USA.
[0042] 1.3 Calculation
[0043] The spectral data matrix obtained from the scan consists of 1290 spectra (rows) and 3001 variables (columns), including 120 spectra of pure sweet potato starch, cassava starch, potato starch, and corn starch; 810 spectra of binary sweet potato mixed starch; and 360 spectra of ternary sweet potato mixed starch. The average spectra of different types of starch were used to compare their differences. Meanwhile, methods such as Standard Normal Variation (SNV), the first derivative (1st), continuous wavelet transform (CWT), and multiplicative scattering correction were used to preprocess the spectra to correct scattering effects and remove background noise from the instrument or samples.
[0044] 1.4 Modeling
[0045] Classification Model: Samples were divided into 5 categories: pure sweet potato starch, tapioca starch, potato starch, corn starch, and mixed starches. A random algorithm was used to divide the spectra of different types of starch samples into two groups: a training set (70% of the samples, 902 spectra) and a test set (30% of the samples, 388 spectra).
[0046] Quantitative model: Predicting the proportion of sweet potato starch, and the prediction results.
[0047] 1.5 Model Evaluation
[0048] Classification Model: The model is evaluated using true positive (TP), true negative (TN), false positive (FP), false negative (FN), sensitivity (Sens.), specificity (Spec.), and accuracy (Acc.). Sensitivity, specificity, and accuracy are the ratios of correctly classified target, non-target, and overall samples. Their mathematical definitions are given in Equations 1, 2, and 3, respectively.
[0049]
[0050] Quantitative model: through the coefficient of determination (R²) 2 The quality of quantitative models is evaluated using the root mean square error of prediction (RMSEP) and relative analysis error (RPD). 2 RMSEP is used to measure how well a predictive model fits observed data. RMSEP represents the difference between predicted and actual values; the smaller the RMSEP value, the higher the model's predictive accuracy. RPD is the ratio of sample standard deviation to prediction standard error. The higher the RPD value, the stronger the model's predictive ability. Generally, an RPD > 3 is considered to indicate a model with good predictive ability.
[0051] 1.6 Software Configuration
[0052] The raw data was exported from MathWorks to MATLAB R2016b on Windows 10 Enterprise Edition, with an Intel Core i7-6700K CPU and 64GB of memory.
[0053] 2. Results and Analysis
[0054] 2.1 Spectral Analysis
[0055] Figure 1 In the image, 'a' represents the original spectra of different types of starch. The original spectra show that the variation trends of different types of starch are similar, all around 8326 cm⁻¹. -1 6848cm -1 6354cm -1 5876cm -1 5738cm -1 5610cm -1 5183cm -1 4758cm -1 4384cm -1 4310cm -1 Several absorption peaks were observed. Wavenumber: 8326 cm⁻¹ -1 This is usually related to the stretching vibration of the CH bond, especially the combined frequency band of the OH stretching vibration and CH bending vibration in primary and secondary alcohols. 6848cm -1 This wavenumber may be related to the deformation vibration of the CH bond or the out-of-plane bending vibration of the CH on the benzene ring. 6354 cm⁻¹ -1 This wavenumber may be related to the deformative vibration of the NH bond or the stretching vibration of the CH bond, especially in nitrogen-containing compounds such as amino acids or amines. 5876 cm⁻¹ -1 This wavenumber may be related to the stretching vibration of the CH bond or the second overtone of the OH bond. In proteins, it may be related to the amide V band. 5738cm -1This wavenumber may be related to the stretching vibrations of the CH bond or the OH bond, especially the OH stretching vibrations in alcohols or phenols. 5610 cm⁻¹ -1 This wavenumber may be related to the stretching vibrations of CH bonds, especially the deformation vibrations of CH bonds in alkyl chains. 5183 cm⁻¹ -1 This wavenumber may be related to the stretching vibrations of the CH bond or the NH bond, and is commonly found in nitrogen-containing compounds. 4758 cm⁻¹ -1 This wavenumber may be related to the deformation vibration or stretching vibration of the CH bond, especially in long-chain alkyl groups. 4384 cm⁻¹ -1 This wavenumber may be related to the deformation vibration or stretching vibration of the CH bond, especially in saturated hydrocarbons. 4310 cm⁻¹ -1 This wavenumber may be related to the deformation vibration of CH bonds, especially the CH bonds on sp3 hybridized carbon atoms.
[0056] To compare the differences between various types of spectra, the average spectrum of each type was plotted. Figure 1 b. Comparing different types of pure tuber starch reveals that the differences in absorption peaks among them are small, with only differences in relative intensity. Potato starch has the highest relative intensity, followed by sweet potato starch, while cassava flour and corn starch have the lowest intensity, and the spectral differences between these two types of starch are relatively small. Sweet potato starch shows a larger intensity difference compared to the other three types of tuber starch. Spectral peaks provide information about the chemical structure of the samples, but this structural information is difficult to discern visually and requires processing using chemometric methods.
[0057] To eliminate noise, baseline drift, and background interference in the samples, four spectral preprocessing methods—MSC (multivariate scattering correction), CWT (continuous wavelet transform), 1st (first derivative), and SNV (standard normal variable)—were used to process the spectra. The processed spectra are shown below. Figure 1 As shown in cf. Through CWT( Figure 1 d) and 1st ( Figure 1 The spectra of samples processed by methods e) show spectral shift and reduced peak width. This is addressed by MSC ( Figure 1 c) and SNV ( Figure 1 The spectral peak drift after processing by method f) was also improved to some extent.
[0058] 2.2 PCA (Principal Component Analysis) Analysis of Different Potato Starches
[0059] A model for identifying different types of potato starch was established using the PCA method. Figure 2In the figure, 'a' represents the PCA results of the original spectra. It can be seen that for pure tuber starches, the confidence ellipses of sweet potato starch and potato starch are clearly distinguishable from those of corn starch and cassava starch. However, the confidence ellipses of cassava starch and corn starch overlap considerably, indicating that pure sweet potato starch can be 100% distinguished from other starch types. For mixtures of different starches, binary and ternary mixtures of sweet potato starch overlap significantly, and multi-component mixtures also show considerable overlap with pure sweet potato starch. This result indicates that mixtures of sweet potato starch adulterated with other starches cannot be identified by PCA.
[0060] Figure 2 In the table, 'be' represents the PCA results after different spectral preprocessing methods, including 1st, MSC, SNV, and CWT. The classification results after 1st and CWT preprocessing are slightly improved compared to the original spectral PCA results, with the confidence ellipses for sweet potato starch and potato starch showing greater distance. However, the spectral PCA results after MSC and SNV preprocessing are worse than the original spectral PCA results, with a small overlap in the confidence ellipses for potato starch and sweet potato starch. This indicates that the SNV and MSC methods not only eliminate noise interference but also remove some of the difference information between pure sweet potato starch and potato starch in the spectra.
[0061] The results above indicate that near-infrared spectroscopy combined with different processing methods can achieve non-destructive identification of pure sweet potato starch from other types of starch; however, inappropriate pretreatment methods can lead to a decrease in model accuracy. Furthermore, PCA and spectral pretreatment alone are insufficient for identifying multi-component mixed sweet potato starches. Therefore, further exploration of other classification methods is needed.
[0062] 2.3 Establishment and Optimization of Classification Model
[0063] The training and test sets were randomly divided in a 7:3 ratio for qualitative identification analysis.
[0064] 2.3.1 Identification of pure starch and mixed starch
[0065] All starch samples were categorized into five classes: pure starch and mixed starch. Class 1 was for mixed starch, and classes 2, 3, 4, and 5 were for pure sweet potato, cassava, potato, and corn starch samples, respectively. The training and test set results are shown in Table 2. Figure 3As shown in the figure, the samples are divided into five categories, but the training set can only correctly predict two categories: category 1 (mixed starch) and category 4 (potato starch). The prediction accuracy for category 1 (mixed starch) is 91.4%, and for category 4 (potato starch) it is 100%. In the test set, the prediction accuracy for category 1 (mixed starch) is 91.7%, and for category 4 (potato starch) it is 100%. Although the model classifies mixed starch and potato starch well, its accuracy in predicting sweet potato, cassava, and corn starch is poor.
[0066] The classification accuracy of different preprocessing methods (CWT, 1st, MSC, SNV) on the training and test sets for various starches (mixed starches, sweet potato starch, tapioca starch, potato starch, corn starch) is shown in the figure. Figure 4 Table 2 shows the prediction accuracy on the training and test sets, as well as the overall prediction accuracy.
[0067] Table 2 Classification results of five training and test sets after different spectral preprocessing.
[0068]
[0069] Analysis of the prediction accuracy data of different processing methods (CWT, 1st, MSC, SNV) on various starches (mixed starches, sweet potato starch, cassava starch, potato starch, and corn starch) reveals differences in performance among the methods. Overall, different methods exhibit varying predictive abilities across different starch categories. Some methods achieve 100% accuracy on specific starch type training sets, reflecting their strong learning ability for corresponding starch features. While the overall prediction accuracy values are similar but fluctuate, indicating room for optimization, all methods show improvement compared to the results from the original spectra.
[0070] The CWT method demonstrates significant advantages in mixed starch prediction, achieving an accuracy of 98.7% on the training set and 97.2% on the test set. It is extremely accurate in feature extraction and classification of sweet potato starch, potato starch, and corn starch, with 100% accuracy on both the training and test sets, and 95.0% accuracy for cassava starch on the training set. The 1st method also performs excellently in mixed starch classification, with prediction accuracy of 98.6% on the training set and 96.7% on the test set. It achieves 100% accuracy in predicting sweet potato, potato, and corn starch, and 94.7% accuracy in predicting cassava starch on both the training and test sets. The MSC and SNV methods perform similarly to the original spectra in classifying mixed starch, achieving 100% prediction accuracy for cassava starch, potato starch, and corn starch on both the training and test sets, thus improving the overall classification accuracy.
[0071] In summary, the CWT and 1st methods demonstrated high accuracy in predicting mixed starches and various single starches, outperforming other methods overall. They consistently achieved 100% accuracy, particularly for sweet potato starch. Performance varied across methods for cassava starch, while most methods achieved 100% accuracy for potato and corn starch. While the current models show high accuracy in predicting mixed starches, their ability to classify and identify different starch types simultaneously needs improvement. Further research is needed to enhance starch prediction accuracy.
[0072] All starch samples were categorized into two classes: pure starch and mixed starch samples, for modeling purposes. Class 1 represents mixed starch, and Class 2 represents pure potato starch samples. The prediction results for the original spectra training and test sets are shown below. Figure 5 The prediction accuracy for mixed starches was 92.5% in the training set and 94.1% in the test set; the accuracy for pure potato starches was 100%. The overall prediction accuracy was the proportion of correctly predicted samples to all samples. With a total of 902 and 388 samples in the training and test sets respectively, and the number of correctly predicted samples being the sum of the diagonal numbers (836 and 366 respectively), the overall prediction accuracy for the training and test sets was 92.7% and 94.3%. Compared to five-class classification, binary classification significantly improved the classification performance for both mixed and pure potato starches. Further comparison of the binary classification performance after spectral preprocessing was conducted, selecting CWT and 1st, two five-class classification methods with good results, for comparison. The classification results are shown in the figure below. Figure 6 The CWT method improved the prediction accuracy for mixed starches to 100%, and the prediction accuracy for pure potato starches to 96.6% in the training set and 97.3% in the test set, with an overall prediction accuracy of 99.7% for both the training and test sets. The classification results after processing with the 1st method showed that the prediction accuracy for mixed starches was improved to 100%, and the prediction accuracy for pure potato starches to 99.5% in the training set and 99.4% in the test set, with an overall prediction accuracy of 99.6% (898 / 902) in the training set and 99.5% (386 / 388) in the test set.
[0073] 2.3.2 Identification of pure starch
[0074] All pure starch samples were modeled and predicted in four categories: sweet potato, cassava, potato, and corn starch samples, respectively. The training and test set results are shown below. Figure 7 Analysis of the prediction accuracy data of different processing methods (RaW, CWT, 1st) on various starches (sweet potato starch, cassava starch, potato starch, corn starch) shows that there is no significant difference in performance among the methods. They can achieve 100% accuracy on the training sets of various pure starch types, reflecting their strong learning ability of pure starch characteristics.
[0075] 2.4 Establishment and Optimization of Quantitative Model
[0076] This study predicts the sweet potato starch content in tuber starch using the PLS (Partial Least Squares) method. First, the influence of the number of factors on the prediction results is examined. The trend of the prediction results with the number of factors is shown in the figure. Figure 8 As the number of factors decreases, the training set error also decreases; the larger the number of factors, the slower the rate of decrease. Therefore, 16 was chosen as the optimal number of factors. The prediction results for sweet potato starch content when the number of factors is 16 are shown below. Figure 9 The statistical results of the predictions are shown in Table 3.
[0077] Table 3. Prediction results of different proportions of sweet potato starch
[0078]
[0079] Table 3 shows the prediction results of sweet potato starch using raw spectral data and spectral data with different preprocessing methods. The linear correlation coefficients R0 of the three types of data (Raw, CWT, and 1st) in the training set are compared. 2 The three metrics are: root mean square error (RMSEP) and prediction deviation (RPD). The R values for the three types of data in the training set are... 2 The correlation coefficients were 0.9610, 0.9927, and 0.9914, respectively, and the RMSEPs were 0.059, 0.026, and 0.028, respectively. It can be seen that the correlation coefficients after spectral preprocessing are better than those of the original spectra, and the CWT method is slightly better than the 1st method. However, the RPD values for the three datasets were 5.1, 11.7, and 10.8, respectively, all greater than 3, indicating that the sweet potato content models established for all three datasets can be used for accurate quantification. The test sets R for Raw, CWT, and 1st data... 2 The RMSEP values were 0.9644, 0.9895, and 0.9861, respectively; and 0.057, 0.031, and 0.036, respectively. Comparison of the test and training sets R... 2 The results from RMSEP and the model are quite similar, indicating that the established model does not exhibit overfitting. In summary, all three datasets can be used to predict the sweet potato starch content in mixed starches. The prediction results are shown in the figure below. Figure 9 .
[0080] 2.5 Identification of Adulteration Types and Prediction of Adulteration Content in Potato Starch Based on 1D-CNN (One-Dimensional Convolutional Neural Network) Method
[0081] 2.5.1 Classification results based on the 1D-CNN method
[0082] A classification model for different types of potato starch was constructed based on the 1D-CNN method. The CNN parameters include a feature extraction layer, one 2×1 convolutional kernel, a 2×1 pooling kernel, the Adam gradient descent algorithm, 500 training iterations, an initial learning rate of 0.001, an L2 regularization parameter, a learning rate reduction factor of 0.1, and a learning rate of 0.0001. The constructed CNN convolutional neural network was used for the classification and quantification of potato starch.
[0083] The results of five-class classification of sweet potato starch using the 1D-CNN method based on raw spectral data are shown below. Figure 10 The overall prediction accuracy is shown in Table 4. The figure shows that the overall prediction accuracy for the five starch categories in the training set is 100%, while the overall accuracy in the test set is 99%. Specifically, the prediction accuracy for mixed starch (category 1), sweet potato starch (category 2), and potato starch (category 4) in the test set is 100%. However, two samples in cassava starch (category 3) and corn starch (category 5) are misclassified, resulting in prediction accuracy of 60% and 80% respectively, thus lowering the overall prediction accuracy. Analysis of the prediction accuracy data for various starches (mixed starch, sweet potato starch, cassava starch, potato starch, and corn starch) using different processing methods (CWT, 1st, MSC, SNV) shows that the prediction accuracy improved to 100% after spectral preprocessing. This indicates that spectral preprocessing can improve the classification performance of the 1D-CNN method. Furthermore, the four preprocessing methods have comparable improvement effects on CNN, suggesting that the CNN method has a higher adaptability and wider selection range for spectral preprocessing methods.
[0084] Table 4. Classification results of five training and test sets after different spectral preprocessing based on 1D-CNN method.
[0085]
[0086] 2.5.2 Quantitative Results Based on 1D-CNN
[0087] When only one type of starch is mixed in sweet potato starch, taking the mixture of sweet potato starch and potato starch as an example, the prediction results are shown in Table 5. For the original spectrum, the R values in the training and test sets... 2 The RMSEP values were 0.9818 and 0.9649, respectively, with RMSEP values of 0.058 and 0.080, and an RPD of 5.2. This indicates that the CNN-based model's predictions showed good correlation with the true values, low error, and an RPD greater than 3, meeting the requirements for accurate quantification. After CWT and 1st spectral preprocessing, the RMSEP values of the training and test sets were [data missing]. 2 All parameters showed slight improvements, while RMSEP decreased and RPD increased. This indicates that spectral preprocessing can improve the quantitative model.
[0088] Table 5. Prediction results of different proportions of sweet potato starch based on 1D-CNN
[0089]
[0090] 2.6 Comparison of Results Based on 1D-CNN and PLS
[0091] By comparing the differences in classification and quantification results between the two methods, it can be found that in the classification application of different types of adulterated sweet potato starch, the accuracy of 1D-CNN reached or approached 100% in five-class, two-class, and four-class classification tasks (99% only in five-class Raw spectral data, 100% in the rest), while PLS achieved a maximum of 96.9% in five-class classification and 99.5% in two-class classification. After spectral preprocessing, the accuracy of 1D-CNN in five-class classification improved from 99% to 100% in Raw data, while the accuracy of PLS in five-class classification improved from 91.7% to 96.9%, but still lower than the 100% accuracy of 1D-CNN. This result indicates that the classification results of PLS depend on the choice of preprocessing method, while the classification results of the 1D-CNN method do not depend on the type of spectral preprocessing. Regardless of the spectral preprocessing technique used, 100% accurate classification results can be obtained, indicating that the classification ability of CNN is far superior to that of PLS.
[0092] Quantitative results show that the RPD of both PLS and 1D-CNN quantitative models is greater than 3, indicating that the established prediction models can meet the requirements for accurate quantification. However, the RPD of the 1D-CNN quantitative model is superior to that of the PLS method. Specifically, when the optimal spectral preprocessing method CWT is selected, the RPD of the 1D-CNN method is significantly higher. 2 The PLS prediction results are 0.9885 and RMSEP is 0.047. 2 It is equivalent to 0.9895 and RMSEP is 0.031.
[0093] In summary, in classification tasks, PLS can only achieve binary classification of starch, while 1D-CNN can simultaneously predict different types of potato starch (five-category classification) when multiple types of starch are mixed, and its prediction accuracy is much higher than that of the PLS method. It can achieve better prediction accuracy than PLS combined with spectral preprocessing techniques. In quantitative tasks, 1D-CNN can also achieve prediction capabilities comparable to the PLS method. Therefore, the 1D-CNN method can simultaneously classify each type of starch and predict the sweet potato starch content in mixed starch when multiple types of mixed starch are adulterated.
[0094] Table 6 Comparison of PLS and 1D-CNN classification results
[0095]
[0096] Table 7 Comparison of PLS and 1D-CNN Quantitative Prediction Results for Sweet Potato
[0097]
Claims
1. A method for identifying and quantifying sweet potato starch adulterated with multiple types of starch, characterized in that, Includes the following steps: S1. Take four pure samples: sweet potato starch, potato starch, cassava starch, and corn starch, and prepare binary and ternary mixture samples containing sweet potato starch. S2. Sample spectra were acquired using a Frontier near-infrared Fourier spectrometer. The sample amount was 2 g. Near-infrared spectra were collected by diffuse reflectance analysis using an integrating sphere. The near-infrared wavenumber range was 4000–10000 cm⁻¹. -1 Scan resolution 4 cm -1 The number of scans was 32; 10 samples were replicated for each sample, and the spectrum of each sample was collected 3 times to obtain the raw spectral data of the sample. S3. Randomly divide the training set and test set into a 7:3 ratio, preprocess the original spectral data of the samples using CWT or 1st method, establish a classification and quantitative prediction model for the training set using 1D-CNN method, and verify and evaluate the model in combination with the test set. S4. Use the classification and quantitative prediction model established in S3 to classify, identify and quantitatively predict the content of unknown starch samples; The 1D-CNN method described above, wherein the parameters of the CNN include a feature extraction layer, 2 One convolutional kernel of size 1, and a pooling kernel of size 2.
1. The Adam gradient descent algorithm is used, with 500 training iterations, an initial learning rate of 0.001, an L2 regularization parameter, a learning rate reduction factor of 0.1, and a learning rate of 0.0001.
2. The method according to claim 1, characterized in that, The binary mixture sample contains 10%-90% sweet potato starch by mass, with the remainder being potato starch, cassava starch, or corn starch; the ternary mixture sample contains 50% sweet potato starch by mass, with the remainder being any two of potato starch, cassava starch, and corn starch, with each having a mass ratio of 10%-40%.
3. The method according to claim 1, characterized in that, Step S3 is as follows: randomly divide the training set and the test set according to a 7:3 ratio, preprocess the original spectral data of the samples using the CWT method, establish a classification and quantitative prediction model for the training set using the 1D-CNN method, and verify and evaluate the model in conjunction with the test set.
4. The method according to claim 1, characterized in that, The classification prediction model is a five-category, four-category, or two-category prediction model. The five-category model includes pure sweet potato starch, potato starch, cassava starch, corn starch, and mixed starch. The four-category model includes pure sweet potato starch, potato starch, cassava starch, and corn starch. The two-category model includes pure sweet potato starch, potato starch, cassava starch, and corn starch as one category, and mixed starch as another category.