Method for identifying and quantifying multi-type starch adulterated sweet potato starch
The classification and quantitative prediction model established by near-infrared spectroscopy and stoichiometric method solves the problem of rapidly identifying and quantifying the adulteration of various potatoes, achieving rapid and accurate starch type identification and adulteration detection, and improving food safety and market order.
Patent Information
- Application Number
- CN202510441558.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing technology is difficult to quickly and accurately identify and quantify the adulteration of various potato starch. The traditional method has a long detection cycle, complex operation and poor environmental friendliness, which affects food safety and market order.
Near-infrared spectroscopy technology combined with stoichiometric method was used to establish a characteristic near-infrared spectral database of sweet potato starch, potato starch, corn starch and tapioca starch, and a classification and quantitative prediction model was established through 1D-CNN and PLS methods to achieve the identification of starch type and the prediction of sweet potato starch proportions.
It has achieved rapid and accurate identification of starch types, cracked down on the counterfeiting and adulteration of starch in the market, guided the collection and storage and processing of potato starch, and improved the value and practicality of starch-based products.
Smart Images

Figure CN120334171A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of food detection, and particularly relates to a method for identifying and quantifying sweet potato starch adulterated with multiple types of starch. Background Art
[0002] Starch, as the core food reserve substance in plants and their derivatives, is one of the most abundant types of carbohydrates in nature. In human daily diet, starch occupies a dominant position in calories and dietary energy and is a key substrate on which the body depends for metabolic activities. Due to its abundance, low cost, non-toxicity, and biodegradability, it has attracted special attention from various food and non-food industries. In the food industry, it is often used in many dairy products, baked goods, soups, sauces, and meat products. However, since the prices of different types of starch vary greatly, some unscrupulous merchants attempt to adulterate high-priced starch with low-priced starch to obtain huge profits. For example, sweet potato starch is extracted from sweet potatoes and has a wide range of applications in multiple industries such as food, papermaking, and textiles. Sweet potato starch, together with corn starch, potato starch, cassava starch, etc., constitutes the main edible starches in China. The prices of different types of starch in the market vary greatly. Among them, cassava starch and corn starch are cheap, less than 2 yuan per catty, while sweet potato starch is often more than twice that of cassava starch per catty. Therefore, many small workshops and manufacturers use cassava starch, corn starch, potato starch, etc. to replace a part of sweet potato starch during production and processing to reduce production costs. Adulterated sweet potato starch and vermicelli products not only seriously disrupt the market order and damage the legitimate rights and interests of consumers, but also pose potential health risks. Therefore, the research on starch and its product adulteration identification technology is of great significance for maintaining the order of the domestic and international starch markets and ensuring food safety.
[0003] There are little differences in appearance among different types of starch, and it is very difficult to distinguish them with the naked eye. The conventional means for starch type identification mainly fall into morphological observation methods and physicochemical property analysis methods. The former can directly observe the starch granule morphology by methods such as scanning electron microscopy (SEM), while the latter relies on the analysis of the ratio of amylose to amylopectin and the evaluation of physicochemical properties such as gelatinization behavior characteristics. Traditional analysis methods have advantages in accuracy, but the disadvantages are long detection cycles, complex operations, damage to samples, and environmental pollution. The limitations of current detection means pose new challenges to the standardization of food starch detection indicators, which urgently prompts us to establish a rapid, accurate, simple, and environmentally friendly starch type classification technology.
[0004] As an analytical tool, near-infrared spectroscopy (NIR) exhibits remarkable characteristics of rapidity, convenience, non-destructiveness, and high efficiency, providing strong support for various analytical tasks. It not only has simple operation but also can quickly and effectively determine the composition or properties of substances without damaging the samples, making it a highly favored technical means in the fields of scientific research and industrial testing. In recent years, near-infrared spectroscopy has been widely applied to the field of food quality detection because it can obtain characteristic information in food components. Chemometrics, as an important branch in the field of spectroscopy, helps to significantly reduce the spectral noise level, enhance the fineness of analysis, effectively eliminate interference factors, and deeply explore the valuable information hidden in spectral data, ultimately achieving the goal of ensuring the accuracy of analysis results.
[0005] In the problem of food adulteration, significant achievements have been made in the identification and classification of samples using near-infrared spectroscopy technology combined with chemometric methods. However, there is currently a lack of rapid and accurate methods for the identification and quantification of various types of potato starches, especially sweet potato starch. Summary of the Invention
[0006] Aiming at the above deficiencies of the existing technology, the key point of the present invention lies in the extraction and classification of the characteristic components of potato starches, aiming to establish a rapid and accurate method for identifying the types of starches in the case of adulteration of various types of potato starches, and providing a method for identifying and quantifying sweet potato starch adulterated with multiple types of starches.
[0007] A method for identifying and quantifying sweet potato starch adulterated with multiple types of starches according to the present invention includes the following steps:
[0008] S1. Take four pure samples of sweet potato starch, potato starch, cassava starch, and corn starch, and prepare binary mixture samples and ternary mixture samples containing sweet potato starch;
[0009] S2. Use a Frontier-type near-infrared Fourier spectrometer to collect the sample spectra and obtain the original spectral data of the samples;
[0010] S3. Randomly divide the samples into a training set and a test set according to a ratio of 7:3, preprocess the original spectral data of the samples using the CWT or 1st method, establish classification and quantitative prediction models for the training set using the PLS method or 1D-CNN method, and verify and evaluate the models in combination with the test set;
[0011] S4. Use the classification and quantitative prediction models established in S3 to classify and identify unknown starch samples and predict and determine their contents.
[0012] Preferably, in the binary mixture sample, the mass ratio of sweet potato starch is 10%-90%, and the balance is potato starch, cassava starch or corn starch; in the ternary mixture sample, the mass ratio of sweet potato starch is 50%, and the balance is any two of potato starch, cassava starch and corn starch, and the mass ratio of each is 10%-40%. The content of each starch in the binary mixture sample can be 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80% or 90%. In the ternary mixture sample, the mass ratio of sweet potato starch is 50%, and the content of each of the remaining starches can be 10%, 20%, 30% or 40%.
[0013] Preferably, step S2 is as follows: collect the sample spectrum with a Frontier type near-infrared Fourier spectrometer, the sample amount is 2 g, collect the near-infrared spectrum of the sample through an integrating sphere diffuse reflection, and the near-infrared wave number range is 4,000-10,000 cm -1 , the scanning resolution is 4 cm -1 , and the number of scanning times is 32 times; 10 sample replicates are set for each sample, and the spectrum of each sample is collected 3 times to obtain the original spectrum data of the sample.
[0014] Preferably, step S3 is as follows: randomly divide the training set and the test set according to a ratio of 7:3, preprocess the original spectrum data of the sample by the CWT method, establish a classification and quantitative prediction model for the training set by the 1D-CNN method, and verify and evaluate the model in combination with the test set.
[0015] Preferably, the classification prediction model is a five-classification, four-classification or two-classification prediction model. The five-classification is pure sweet potato starch, potato starch, cassava starch, corn starch and mixed starch; the four-classification is pure sweet potato starch, potato starch, cassava starch and corn starch; the two-classification is that pure sweet potato starch, potato starch, cassava starch and corn starch are in one category, and mixed starch is in one category.
[0016] Preferably, for the 1D-CNN method, the parameters of the CNN include a feature extraction layer, 1 convolution kernel with a size of 2×1, the pooling kernel size is 2×1, the Adam gradient descent algorithm is adopted, the number of training times is 500, the initial learning rate is 0.001, the L2 regularization parameter, and the learning rate decay factor is 0.1, and the learning rate is 0.0001.
[0017] The present invention takes common low-value raw material starches such as cassava starch and corn starch as the doping objects of sweet potato starch, uses near-infrared spectroscopy as the detection means, establishes a characteristic near-infrared spectrum database of sweet potato starch, potato starch, corn starch, cassava starch, etc., preprocesses the obtained spectrum in combination with chemometrics and establishes a variety of discrimination models to realize the discrimination of pure sweet potato starch and non-pure sweet potato starch, and predicts the proportion of sweet potato starch therein.
[0018] The method of the present invention has important reference significance for formulating the classification standard of tuber starches, lays a solid foundation for quickly and accurately identifying the starch type, is conducive to cracking down on bad behaviors such as starch counterfeiting and adulteration in the market, guiding the storage and processing of sweet potato starch, helping to screen high-quality tuber starch raw materials for the research of tuber starch products, and improving the overall value and practicability of starch-based products. Description of the Drawings
[0019] Figure 1 are the spectra of different tuber starches; among them, a is the original spectrum; b is the average spectrum of different types of spectra, where the sample Gan represents sweet potato starch, Mu represents cassava starch, Tu represents potato starch, Yu represents corn starch, Gan+Tu represents the binary mixture of sweet potato starch and potato starch, Gan+Mu represents the binary mixture of sweet potato starch and cassava starch, Gan+Yu represents the binary mixture of sweet potato starch and corn starch, and Gan+Yu+Tu represents the ternary mixture of sweet potato starch, corn starch and potato starch; c is the spectrum processed by MSC; d is the spectrum processed by CWT; e is the spectrum processed by 1st; f is the spectrum processed by SNV.
[0020] Figure 2 is the PCA diagram of different types of tuber starches; among them, a is the original spectrum; b is the average spectrum of different types of spectra; c is the spectrum processed by MSC; d is the spectrum processed by CWT; e is the spectrum processed by 1st; f is the spectrum processed by SNV; where the sample SYS represents ternary mixed starch, EYS represents binary mixed starch, CS represents corn starch, PS represents potato starch, MS represents cassava starch, and SPS represents sweet potato starch.
[0021] Figure 3 is the classification effect diagram of the original spectra of five starches.
[0022] Figure 4 is the five-classification confusion matrix of different preprocessed spectra.
[0023] Figure 5 is the classification effect of the original spectra in the two-classification situation.
[0024] Figure 6 is the classification effect of different preprocessed spectra in the two-classification situation.
[0025] Figure 7 is the classification effect of different spectra of pure tuber starches.
[0026] Figure 8 is the trend chart of the change of the prediction result error of the training set with the number of factors.
[0027] Figure 9Are the predicted results of sweet potato starch content with different spectra.
[0028] Figure 10 Are the five-classification results based on the 1D-CNN method.
[0029] Figure 11 Are the two-classification and four-classification results based on the 1D-CNN method.
[0030] Figure 12 Are the predicted results of sweet potato starch content based on the 1D-CNN method. Detailed implementation manners
[0031] The following embodiments are further descriptions of the present invention rather than limitations thereof.
[0032] Embodiment 1
[0033] 1. Materials and methods
[0034] 1.1 Materials and sample preparation
[0035] The starches used include sweet potato starch, corn starch, potato starch, and cassava starch. The sweet potato starch is from different batches on the production line of Hubei Jinyue Agricultural Products Development Co., Ltd., with a purity of 100% and a total of 10 samples. The corn starch, potato starch, and cassava starch samples are purchased in batches from the market, and the producer is located in Jingshan City (Hubei, China) to ensure the representativeness of the samples. Among them, the potato starch, cassava starch, and corn starch meet the implementation standards GB / T8885, GB / T 8884, and GB / T 29343 respectively. There are 10 batches of samples for each type.
[0036] The pure sweet potato starch, potato starch, cassava starch, and corn starch of different batches are crushed and ground, and then sieved through a 60-mesh sieve after grinding. 3 portions are weighed for each sample, with 2 g for each portion. In addition to these samples, binary and ternary mixtures between pure samples are also prepared to increase the variability of the data set. Among them, the change in sweet potato starch in the binary mixed samples ranges from 10% - 90% in mass ratio, and there are 10 samples for each binary mixed sample, with 3 replicates for each sample; for the ternary mixed samples, the mixing ratio of the adulterated sweet potato starch samples is shown in Table 1. After mixing, 2 g of each prepared sample is weighed and mixed evenly using an oscillator.
[0037] Table 1 Description of the number of samples and the adulteration ratio of samples
[0038]
[0039]
[0040] 1.2 Spectrum acquisition
[0041] Take 2 g of the above-mentioned sample and place it in a quartz cup. Using air as the background, collect the near-infrared spectra of each solid sample through an integrating sphere diffuse reflection at room temperature. The acquisition and scanning of the spectra are carried out using a Frontier-type near-infrared Fourier spectrometer (near-infrared wave number range 4000 - 10000 cm -1 , the scanning resolution is 4 cm -1 , and the number of scans is 32 times) from Perkin Elmer, USA.
[0042] 1.3 Calculation
[0043] The spectral data matrix obtained by scanning consists of 1290 spectra (rows) and 3001 variables (columns), including 120 spectra of pure sweet potato starch, cassava starch, potato starch, and corn starch, 810 spectra of binary sweet potato mixed starch, and 360 spectra of ternary sweet potato mixed starch. The average spectra of different types of starches are used to compare the differences between them. At the same time, methods such as Standard Normal Variate (SNV), the first derivative (1st), continuous wavelet transform (CWT), and Multiplicative Scatter Correction are used to preprocess the spectra to correct the scattering effect in the spectral data and remove the noise background brought by the instrument or sample.
[0044] 1.4 Modeling
[0045] Classification model: The samples are divided into 5 categories, namely pure sweet potato starch, cassava starch, potato starch, corn starch, and mixed starch. The spectra of starch samples of different types are randomly divided into two groups, a training set and a test set, using a random algorithm. Among them, the training set samples account for 70%, 902 spectra, which are used to establish a prediction model, and the test set samples account for 30%, 388 spectra.
[0046] Quantitative model: Predict the proportion of sweet potato starch, and the prediction results.
[0047] 1.5 Model evaluation
[0048] Classification model: Use true positive (TP), true negative (TN), false positive (FP), false negative (FN), sensitivity (Sens.), specificity (Spec.), and accuracy (Acc.) to evaluate the model. Sensitivity, specificity, and accuracy are the ratios of correctly classifying target, non-target, and overall samples. Their mathematical definitions are given in Formulas 1, 2, and 3 respectively.
[0049]
[0050] Quantitative model: The quality of the quantitative model was evaluated by the coefficient of determination (R 2 ), root mean square error of prediction (RMSEP), and ratio of performance to deviation (RPD). R 2 was used to measure the goodness of fit of the prediction model to the observed data. RMSEP represents a measure of the difference between the predicted value and the true value. The smaller the value of RMSEP, the higher the prediction accuracy of the model. RPD is the ratio of the sample standard deviation to the standard error of prediction. The higher the RPD value, the stronger the prediction ability of the model. Generally, when RPD > 3, the model has good prediction ability.
[0051] 1.6 Software configuration
[0052] The original data was exported from MathWorks to Windows 10 Enterprise Edition of MATLAB R2016b, with an Intel Core i7-6700K CPU and 64 Gb of memory.
[0053] 2. Results and analysis
[0054] 2.1 Spectral analysis
[0055] Figure 1 In [a], the original spectra of different types of starches are shown. It can be seen from the original spectra that the changing trends of different types of starches are similar, with absorption peaks at 8326 cm -1 , 6848 cm -1 , 6354 cm -1 , 5876 cm -1 , 5738 cm -1 , 5610 cm -1 , 5183 cm -1 , 4758 cm -1 , 4384 cm -1 , 4310 cm -1 and so on. The wave number 8326 cm -1 is usually related to the stretching vibration of the C-H bond, especially to the combination band of the O-H stretching vibration and C-H bending vibration in primary and secondary alcohols. The wave number 6848 cm -1 may be related to the deformation vibration of the C-H bond or the out-of-plane bending vibration of the C-H bond on the benzene ring. The wave number 6354 cm -1 may be related to the deformation vibration of the N-H bond or the stretching vibration of the C-H bond, especially in nitrogen-containing compounds such as amino acids or amines. The wave number 5876 cm -1 may be related to the stretching vibration of the C-H bond or the second overtone of the O-H bond. In proteins, it may be related to the amide V band. The wave number 5738 cm -1This wavenumber may be related to the stretching vibration of C-H bonds or the stretching vibration of O-H bonds, especially the O-H stretching vibration in alcohols or phenolic compounds. 5610 cm -1 This wavenumber may be related to the stretching vibration of C-H bonds, especially the deformation vibration of C-H bonds in the alkyl chain. 5183 cm -1 This wavenumber may be related to the stretching vibration of C-H bonds or the stretching vibration of N-H bonds, commonly found in nitrogen-containing compounds. 4758 cm -1 This wavenumber may be related to the deformation vibration of C-H bonds or the stretching vibration of C-H bonds, especially in long-chain alkyl groups. 4384 cm -1 This wavenumber may be related to the deformation vibration of C-H bonds or the stretching vibration of C-H bonds, especially in saturated hydrocarbon compounds. 4310 cm -1 This wavenumber may be related to the deformation vibration of C-H bonds, especially the C-H bonds on sp3 hybridized carbon atoms.
[0056] To compare the differences between various spectra, the average spectra of each type of spectrum were plotted, as shown in Figure 1 b in. Comparing different types of pure potato starches, it can be seen that the differences in absorption peaks between pure potato starches are small, only the differences in relative intensities. Among them, the relative intensity of potato starch is the highest, followed by sweet potato starch, and the intensities of tapioca starch and corn starch are the lowest, and the spectral differences between these two types of starches are small. The intensity difference between sweet potato starch and the other three types of potato starches is relatively large. The spectral peaks provide chemical structure information of the samples, but it is difficult to visually distinguish these structural information, and chemometric methods need to be used for processing.
[0057] To eliminate the noise, baseline drift and background interference of the samples, four spectral preprocessing methods, namely MSC (Multiplicative Scatter Correction), CWT (Continuous Wavelet Transform), 1st (First Derivative) and SNV (Standard Normal Variate), were used to process the spectra. The processed spectra are as shown in Figure 1 c-f in. Through the sample spectra processed by two methods, CWT ( Figure 1 d in) and 1st ( Figure 1 e in), the spectral drift and peak width are reduced. Through the spectra processed by MSC ( Figure 1 c in) and SNV ( Figure 1 f in) methods, the peak drift of the spectra is also improved to a certain extent.
[0058] 2.2 PCA (Principal Component Analysis) Analysis of Different Potato Starches
[0059] A discrimination model of different potato starches was established by using the PCA method. Figure 2In this case, a is the PCA result of the original spectrum. It can be seen that for pure tuber starches, the confidence ellipses of sweet potato starch and potato starch are significantly different from those of corn starch and cassava starch. However, there is a large overlap between the confidence ellipses of cassava starch and corn starch, indicating that pure sweet potato starch can be 100% discriminated from other types of starches. For mixed starches of different types, the binary mixed sweet potato starch and the ternary mixed sweet potato starch overlap with each other to a high degree. At the same time, there is also a large degree of overlap between the multi-component mixed sweet potato starch and the pure sweet potato starch. This result shows that the mixture of sweet potato starch doped with other types of starches cannot be discriminated by the PCA method.
[0060] Figure 2 In this case, b - e are the PCA results after different spectral pre-treatments such as 1st, MSC, SNV, and CWT respectively. The classification results after 1st and CWT pre-treatments are slightly improved compared with the PCA result of the original spectrum, and the confidence ellipses of sweet potato starch and potato starch are farther apart. However, the spectral PCA results after MSC and SNV pre-treatments are worse than the PCA result of the original spectrum, and there is a small overlap between the confidence ellipses of potato starch and sweet potato starch. This indicates that the SNV and MSC methods not only eliminate noise interference but also eliminate some differential information in the spectrum related to pure sweet potato starch and potato starch.
[0061] The above results show that near-infrared spectroscopy combined with different processing methods can achieve non-destructive discrimination of pure sweet potato starch from other types of starches, but inappropriate pre-treatment methods will lead to a decrease in model accuracy. At the same time, it is impossible to discriminate multi-component mixed sweet potato starch only by PCA and spectral pre-treatment. Therefore, it is necessary to further explore other classification methods.
[0062] 2.3 Establishment and optimization of the classification model
[0063] Randomly divide the training set and the test set according to a ratio of 7:3 for qualitative discrimination analysis.
[0064] 2.3.1 Discrimination of pure starch and mixed starch
[0065] The sample categories of all starches are divided into a total of 5 types of labels for pure starch and mixed starch samples for modeling prediction. Among them, 1 type is mixed starch, and types 2, 3, 4, and 5 are pure sweet potato, cassava, potato, and corn starch samples respectively. The results of the training set and the test set are shown in Table 2. Figure 3As can be seen from the figure, the samples are divided into five categories, but the training set can only correctly predict two types of samples, namely the first type of mixed starch and the fourth type of potato starch. The prediction accuracy rate of the mixed starch (type 1) is 91.4%, and the prediction accuracy rate of the potato starch (type 4) is 100%. In the test set, the prediction accuracy rate of the mixed starch (type 1) is 91.7%, and the prediction accuracy rate of the potato starch (type 4) is 100%. Although the model has a good classification for the mixed starch and potato starch, the correct prediction effect for sweet potato starch, cassava starch and corn starch is poor.
[0066] The classification effect of the correct rate of various starches (mixed starch, sweet potato starch, cassava starch, potato starch, corn starch) in the training set and test set by using different pretreatment methods (CWT, 1st, MSC, SNV) is shown in Figure 4 , the prediction accuracy rate in the training set and test set, as well as the overall prediction accuracy rate are shown in Table 2.
[0067] Table 2 Classification effect of five types of training sets and test sets after different spectral pretreatments
[0068]
[0069] Through the analysis of the prediction accuracy rate data of different processing methods (CWT, 1st, MSC, SNV) for various starches (mixed starch, sweet potato starch, cassava starch, potato starch, corn starch), it can be seen that there are differences in the performance of each method. Generally speaking, the prediction abilities of different methods for each starch category are different. Some methods can achieve a 100% accuracy rate for the training set of specific starch types, reflecting their strong learning ability for the corresponding starch characteristics. And the overall prediction accuracy rate values are similar but fluctuate, indicating that there is room for optimization, but they are all improved compared with the results of the original spectrum.
[0070] The CWT method has obvious advantages in the prediction of mixed starch, with a training set accuracy rate of 98.7% and a test set accuracy rate of 97.2%. It is extremely accurate in the feature extraction and classification of sweet potato starch, potato starch and corn, with a 100% accuracy rate for both the training set and the test set, and a 95.0% accuracy rate for the cassava starch training set; the 1st method has excellent classification effect for mixed starch, with prediction accuracy rates of 98.6% for the training set and 96.7% for the test set respectively. The prediction accuracy rates for sweet potato, potato and corn starches are all 100%, and the prediction accuracy rates for cassava starch are 94.7% for the training set and 100% for the test set respectively; the classification effects of the MSC method and the SNV method for mixed starch are similar to those of the original spectrum, and the training set and test set prediction accuracy rates for cassava starch, potato starch and corn starch are 100%, thus improving the overall classification accuracy rate.
[0071] Overall comparison shows that the spectra processed by CWT and the 1st method have high prediction accuracy for mixed starch and various single starches. Their overall performance is better than other methods. In particular, the prediction accuracy for sweet potato starch is always 100%. For cassava starch, the performance of each method varies. For potato and corn starches, most methods can reach 100%. Since the current model has a high prediction accuracy for mixed starch, but there is room for improvement in the simultaneous classification and identification of different types of starches. If we want to improve the starch prediction accuracy, further exploration can be carried out.
[0072] The sample categories of all starches are divided into two types of labels: pure starch and mixed starch samples for modeling. Among them, type 1 is mixed starch, and type 2 represents pure tuber starch samples. The prediction results of the training set and test set of the original spectra are shown in Figure 5 . The prediction accuracy rates for mixed starch are 92.5% for the training set and 94.1% for the test set respectively; the accuracy rates for pure tuber starches are both 100%. The overall prediction accuracy rate is the proportion of the number of correctly predicted samples to all samples. That is, the total number of samples in the training set and test set are 902 and 388 respectively, and the number of correctly predicted samples is the sum of the diagonal numbers, which are 836 and 366 respectively. Therefore, the overall prediction accuracy rates of the training set and test set are 92.7% and 94.3% respectively. Compared with the five-class classification, the two-class classification significantly improves the classification effect of mixed starch and pure tuber starches. Further comparing the effects of the two-class classification after spectral preprocessing, we select the two classification methods, CWT and 1st, with good effects in the five-class classification situation for comparison. The classification effect diagrams are shown in Figure 6 . The prediction accuracy rate of the CWT method for mixed starch is increased to 100%, and the prediction accuracy rates for pure tuber starches are increased to 96.6% for the training set and 97.3% for the test set respectively. The overall prediction accuracy rates of the training set and test set are both 99.7%. The classification results after the 1st method is processed are that the prediction accuracy rate for mixed starch is increased to 100%, and the prediction accuracy rates for pure tuber starches are increased to 99.5% for the training set and 99.4% for the test set respectively. The overall prediction accuracy rate of the training set is 99.6% (i.e., 898 / 902), and the overall accuracy rate of the test set is 99.5% (i.e., 386 / 388).
[0073] 2.3.2 Identification of Pure Starch
[0074] All pure starch samples, a total of 4 types, are modeled and predicted. Among them, 1, 2, 3, and 4 are pure sweet potato, cassava, potato, and corn starch samples respectively. The results of the training set and test set are shown in Figure 7 . Through the analysis of the prediction accuracy rate data of different processing methods (RaW, CWT, 1st) for various starches (sweet potato starch, cassava starch, potato starch, corn starch), it can be seen that there is no obvious difference in the performance of each method. The training set can achieve a 100% accuracy rate for various types of pure starches, reflecting its strong learning ability for the characteristics of pure starches.
[0075] 2.4 Establishment and Optimization of Quantitative Models
[0076] Based on the PLS method (partial least squares method), the content of sweet potato starch in potato starch was predicted. First, the influence of the number of factors on the prediction results was investigated. The trend of the prediction results changing with the number of factors is shown in Figure 8 . As the number of factors decreases, the training set error decreases accordingly. The greater the number of factors, the slower the decreasing trend. Therefore, 16 was determined as the number of factors. When the number of factors is 16, the prediction results of the sweet potato starch content are shown in Figure 9 , and the prediction result statistics are shown in Table 3
[0077] Table 3 Prediction Results of Different Proportions of Sweet Potato Starch
[0078]
[0079] Table 3 shows the prediction results of sweet potato starch for the original spectral data and the spectral data with different pre-treatments. Comparing the linear correlation coefficients R 2 , root mean square error of prediction RMSEP, and RPD of the three types of data Raw, CWT, and 1st in the training set, the Rs 2 of the three types of data in the training set are 0.9610, 0.9927, and 0.9914 respectively, and the RMSEPs are 0.059, 0.026, and 0.028 respectively. It can be seen that the correlation coefficient after spectral pre-treatment is better than that of the original spectrum, and the CWT method is slightly better than the 1st method. However, the RPD values of the three types of data are 5.1, 11.7, and 10.8 respectively, all greater than 3, indicating that the sweet potato content models established by the three types of data can be used for precise quantification. The Rs 2 of the test sets of Raw, CWT, and 1st data are 0.9644, 0.9895, and 0.9861 respectively; the RMSEPs are 0.057, 0.031, and 0.036 respectively. Comparing the R 2 and RMSEP results of the test set and the training set, the two results are relatively close, indicating that the established model does not have overfitting linearity. In summary, all three types of data can be used to predict the content of sweet potato starch in mixed starch, and the prediction result diagram is shown in Figure 9 .
[0080] 2.5 Identification of Adulteration Types and Prediction of Adulteration Content of Potato Starch Based on the 1D-CNN (One-Dimensional Convolutional Neural Network) Method
[0081] 2.5.1 Classification Results Based on the 1D-CNN Method
[0082] A classification model for different potato starches was constructed based on the 1D-CNN method. The parameters of the CNN included a feature extraction layer, one convolutional kernel with a size of 2×1, a pooling kernel size of 2×1. The Adam gradient descent algorithm was used, with 500 training epochs, an initial learning rate of 0.001, an L2 regularization parameter, a learning rate decay factor of 0.1, and a learning rate of 0.0001. The constructed CNN convolutional neural network was used for the classification and quantification of potato starches.
[0083] The results of classifying sweet potato starch into five categories based on the 1D-CNN method using the original spectral data are shown in Figure 10 , and the overall prediction accuracy is statistically shown in Table 4. It can be seen from the figure that the overall prediction accuracy of the five types of starches in the training set is 100%, and the total accuracy in the test set is 99%. Among them, the prediction accuracies of the mixed starch (class 1), sweet potato starch (class 2), and potato starch (class 4) in the test set are all 100%. There are 2 misclassified samples in cassava starch (class 3) and corn starch (class 5), resulting in prediction accuracies of 60% and 80% respectively, thus reducing the overall prediction accuracy. Through the analysis of the prediction accuracy data of different processing methods (CWT, 1st, MSC, SNV) for various starches (mixed starch, sweet potato starch, cassava starch, potato starch, corn starch), it can be seen that the prediction accuracy after spectral preprocessing has increased to 100%, indicating that spectral preprocessing can improve the classification effect of the 1D-CNN method, and the improvement effects of the four preprocessing methods on the CNN are quite similar, indicating that the CNN method has a higher adaptability to spectral preprocessing methods and a wider selection range.
[0084] Table 4 Classification effects of five types of training sets and test sets after different spectral preprocessings based on the 1D-CNN method
[0085]
[0086] 2.5.2 Quantitative results based on 1D-CNN
[0087] When only one type of starch is mixed in sweet potato starch, taking the mixture of sweet potato starch and potato starch as an example, the prediction results are shown in Table 5. For the original spectrum, the R 2 in the training set and the test set are 0.9818 and 0.9649 respectively, the RMSEP are 0.058 and 0.080 respectively, and the RPD is 5.2, indicating that the prediction results of the model established based on CNN have a good correlation with the true values, low errors, and the RPD is also greater than 3, meeting the requirements for accurate quantification. After CWT and 1st spectral preprocessing, the R 2 in both the training set and the test set have increased slightly, the RMSEP has decreased, and the RPD has increased. This indicates that spectral preprocessing can improve the quantitative model.
[0088] Table 5 Prediction Results of Sweet Potato Starches with Different Proportions Based on 1D-CNN
[0089]
[0090] 2.6 Result Comparison between 1D-CNN and PLS
[0091] By comparing the differences in classification and quantification results between the two methods, it can be found that in the classification application of different types of adulterated mixed sweet potato starches, in the five-classification, two-classification, and four-classification tasks, the correct rates of 1D-CNN all reach or are close to 100% (only 99% under the five-classification Raw spectrum, and the rest are 100%), while the highest correct rate of PLS in five-classification is 96.9%, and the highest in two-classification is 99.5%. After spectral preprocessing, the five-classification correct rate of 1D-CNN increased from 99% of Raw to 100%, and the five-classification correct rate of PLS increased from 91.7% of Raw to 96.9%, but it is still lower than 100% of 1D-CNN. This result indicates that the classification result of PLS depends on the choice of preprocessing method, while the classification result of the 1D-CNN method does not depend on the type of spectral preprocessing. Regardless of which spectral preprocessing technology is used, a 100% correct classification result can be obtained, indicating that the classification ability of CNN is much higher than that of the PLS method.
[0092] The quantitative results show that the RPD of the quantitative models of both PLS and 1D-CNN methods is greater than 3, indicating that the established prediction models can all meet the requirements of accurate quantification. However, the RPD of the 1D-CNN quantitative model is better than that of the PLS method. Among them, when the selected optimal spectral preprocessing method is CWT, the prediction result R 2 of the 1D-CNN method is 0.9885, and the RMSEP is 0.047. The prediction result (R 2 is 0.9895, and the RMSEP is 0.031) of PLS is comparable to it.
[0093] In summary, in the classification task, PLS can only achieve binary classification of starches, while 1D-CNN can simultaneously predict different types of potato starches (five-classification) when there are multiple types of starch doping, and the prediction accuracy is much higher than that of the PLS method, and a higher prediction accuracy better than that of PLS combined with spectral preprocessing technology can be obtained; in the quantitative task: 1D-CNN can also obtain a prediction ability comparable to that of the PLS method. Therefore, using the 1D-CNN method can simultaneously achieve the classification of each type of starch when there is multi-type mixed starch adulteration and the prediction of the content of sweet potato starch in the mixed starch.
[0094] Table 6 Comparison of Classification Results between PLS and 1D-CNN
[0095]
[0096] Comparison of Quantitative Prediction Results of Sweet Potato by PLS and 1D-CNN
[0097]
Claims
1. A method for identifying and quantifying sweet potato starch adulterated with multiple types of starch, characterized in that, It includes the following steps: S1. Take four pure samples of sweet potato starch, potato starch, cassava starch, and corn starch, and prepare binary mixture samples and ternary mixture samples containing sweet potato starch; S2. Collect the sample spectra using a Frontier-type near-infrared Fourier spectrometer to obtain the original spectral data of the samples; S3. Randomly divide the samples into a training set and a test set at a ratio of 7:3, preprocess the original spectral data of the samples using the CWT or 1st method, establish a classification and quantitative prediction model for the training set using the PLS method or the 1D-CNN method, and verify and evaluate the model in combination with the test set; S4. Use the classification and quantitative prediction model established in S3 to classify and identify unknown starch samples and predict and determine their content quantification.
2. The method according to claim 1, wherein In the binary mixture sample, the mass ratio of sweet potato starch is 10%-90%, and the balance is potato starch, cassava starch, or corn starch; in the ternary mixture sample, the mass ratio of sweet potato starch is 50%, and the balance is any two of potato starch, cassava starch, and corn starch, with each having a mass ratio of 10%-40%.
3. The method according to claim 1, characterized in that, The step S2 is as follows: Collect the sample spectrum with a Frontier type near-infrared Fourier spectrometer. The sample amount is 2 g. Collect the near-infrared spectrum of the sample through an integrating sphere for diffuse reflection. The near-infrared wave number range is 4,000 - 10,000 cm -1 , and the scanning resolution is 4 cm -1 . The scanning times are 32 times. For each type of sample, 10 sample replicates are set, and the spectrum of each sample is collected 3 times to obtain the original spectral data of the sample.
4. The method according to claim 1, wherein The step S3 is: randomly divide the samples into a training set and a test set at a ratio of 7:3, preprocess the original spectral data of the samples using the CWT method, establish a classification and quantitative prediction model for the training set using the 1D-CNN method, and verify and evaluate the model in combination with the test set.
5. The method according to claim 1, wherein The classification prediction model is a five-classification, four-classification, or two-classification prediction model. The five-classification is pure sweet potato starch, potato starch, cassava starch, corn starch, and mixed starch; the four-classification is pure sweet potato starch, potato starch, cassava starch, and corn starch; the two-classification is that pure sweet potato starch, potato starch, cassava starch, and corn starch are in one category, and mixed starch is in one category.
6. The method according to claim 1, characterized in that, For the 1D-CNN method, the parameters of the CNN include a feature extraction layer, 1 convolutional kernel with a size of 2×1, a pooling kernel size of 2×1, the Adam gradient descent algorithm is used, the number of training times is 500, the initial learning rate is 0.001, the L2 regularization parameter, the learning rate decay factor is 0.1, and the learning rate is 0.0001.
Citation Information
Patent Citations
Method for detecting low-quality starch-doped potato starch fast
CN105044014A
Method of discriminating perchlorate content range in tea leaf based on near infrared spectroscopy
CN110308109A
Matrix difference mushroom variety identification method
CN118050460A
Method and system for identifying lotus root starch adulteration and storage medium
CN118334649A
Rapid quality evaluation method for simultaneously detecting producing area and multi-index content of scutellaria baicalensis
CN118730968A
Cited By
Method for identifying and quantifying adulterated sweet potato starch and method for training detection model
CN121877793A
Method for identifying and quantifying adulterated sweet potato starch, and detection model training method
CN121877793B