Edible starch species rapid identification method and system

By using a portable near-infrared spectrometer and a multi-model integrated classifier, combined with label consistency auditing and secondary verification mechanisms, the portability and accuracy issues of rapid identification of edible starch species have been solved, achieving efficient and accurate identification of starch species.

CN121521801APending Publication Date: 2026-02-13INSPECTION & QUARANTINE TECH CENT SHANDONG ENTRY EXIT INSPECTION & QUARANTINE BUREAU
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511982027.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-02-13

AI Technical Summary

Technical Problem

Existing technologies struggle to quickly and accurately identify different types of edible starch, especially in field settings. Traditional methods are costly, rely on laboratory equipment, and fail to meet the demands for portability and timeliness.

Method used

A portable near-infrared spectrometer was used to process spectral data with standard normal variable transformation and Savitzky-Golay first-order derivative filtering. Starch species were identified through a multi-model ensemble classifier, including a weighted soft voting strategy of K-nearest neighbors, random forest and XGBoost classifiers, combined with label consistency audit and secondary verification mechanisms.

Benefits of technology

It enables rapid, non-destructive, and highly accurate on-site identification of various common edible starches, solving the problems of traditional methods being time-consuming, costly, and dependent on laboratory environments, and providing an efficient and reliable starch species identification solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121521801A_ABST
    Figure CN121521801A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent identification, and provides an edible starch species rapid identification method and system. The method comprises the following steps: placing a powdery starch sample to be detected in a culture dish, striking off the surface, placing the culture dish on a turntable which rotates at a constant speed, and dynamically collecting original spectral data of the sample in a diffuse reflection mode by using a portable near-infrared spectrometer which is fixed on a bracket and has a vertically downward probe in a dark box environment; carrying out standard normal variable transformation and Savitzky-Golay first-order derivative filtering processing on the original spectral data in sequence; and inputting the preprocessed spectral data into a pre-trained starch species classification model, and outputting a species identification result. By constructing a standardized dynamic spectrum acquisition process, an optimized spectrum preprocessing scheme and an integrated classification model, rapid, lossless and high-accuracy field identification of various common edible starches can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to intelligent identification technology, specifically to a method and system for rapid identification of edible starch species. Background Technology

[0002] Starch, one of the most abundant polysaccharides in nature, is an indispensable thickener, stabilizer, and gelling agent in the food industry. Its functional properties, such as gelatinization temperature, peak viscosity, and retrogradation, directly depend on its source species. For example, potato starch has high viscosity and high transparency, making it suitable for sauces and frozen foods; cassava starch is known for its good freeze-thaw stability and film-forming properties, and is commonly used in candies and convenience foods; while mung bean starch, due to its high gel strength and good elasticity, is the preferred raw material for making vermicelli and rice noodles. Therefore, ensuring the authenticity of the species of starch raw materials is not only crucial for enterprises to guarantee stable product quality, but also central to protecting consumers' right to know and combating adulteration in the market.

[0003] Existing identification methods each have limitations. Chemical methods, such as the iodine colorimetric reaction, can distinguish the ratio of amylose to amylopectin, but cannot be precise at the species level. Microscopic observation methods are highly subjective and difficult to standardize. While high-performance liquid chromatography or mass spectrometry are accurate, they are costly and require complex pretreatment, making it difficult to meet the timeliness and portability requirements of rapid on-site screening. These inherent defects pose significant challenges to the source control and process supervision of the starch supply chain.

[0004] Near-infrared spectroscopy (NIR) has become a research hotspot in modern food analysis due to its advantages such as speed, non-destructive nature, reagent-free operation, and the ability to perform online / on-site analysis. However, it still faces many challenges in practical applications. The NIR spectra of different starches are highly similar in overall profile, with significant overlap of characteristic peaks, especially among species of the same family and genus, such as corn and wheat, or mung beans and peas, where spectral differences are subtle, making high-precision classification difficult. The physical properties of samples, such as particle size, packing density, and surface smoothness, introduce significant scattering effects, masking true chemical information. The robustness and generalization ability of models are insufficient. Many studies have achieved high accuracy under ideal laboratory conditions, but performance often drops significantly when the models are applied to samples from different batches and origins. Furthermore, existing research is mostly focused on benchtop instruments and laboratory environments, lacking portable, integrated solutions for practical applications.

[0005] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0006] The purpose of this application is to provide a method and system for rapid identification of edible starch species, which has the advantages of enabling rapid, non-destructive, and highly accurate on-site identification of a variety of common edible starches.

[0007] Firstly, this application provides a method for rapid identification of edible starch species, the technical solution of which is as follows: Spectral acquisition steps: Place the powdered starch sample to be tested in a petri dish and smooth the surface. Place the petri dish on a rotating turntable at a uniform speed. In a dark environment, use a portable near-infrared spectrometer fixed on a support with the probe pointing vertically downward to dynamically acquire the raw spectral data of the sample in a diffuse reflectance manner. Spectral preprocessing steps: The raw spectral data are sequentially subjected to standard normal transformation and Savitzky-Golay first-order derivative filtering; Species identification steps: Input the preprocessed spectral data into the pre-trained starch species classification model and output the species identification results.

[0008] Furthermore, the starch species classification model is trained based on a process that includes label consistency auditing; The specific training process includes: Data preparation: Near-infrared spectral training sets of various starch species were obtained, and the spectra were preprocessed by standard normal variable transformation and Savitzky-Golay first derivative filtering to obtain the preprocessed full spectrum data. Base model training: Based on the preprocessed full-spectrum data, K-nearest neighbors, random forest and XGBoost classifiers were trained as base models respectively; Label auditing and optimization: Based on the predicted probability of the training samples by the base model, the consistency of the original sample labels is audited, and labels with inconsistent confidence are screened and corrected to form an optimized training set; Integration: The base model is retrained using the optimized training set and integrated using a weighted soft voting strategy, where the voting weight of the K-nearest neighbor classifier is 40%, and the voting weights of the random forest and XGBoost classifiers are each 30%.

[0009] Furthermore, the method also includes: For the same sample to be tested, M parallel spectra are continuously acquired in the spectral acquisition step, and spectral preprocessing and species identification steps are performed respectively to obtain M initial identification results; In M initial identification results, if the frequency of occurrence of a preset species exceeds a preset threshold N, the sample to be tested is determined to belong to the preset species; where M is an integer greater than or equal to 3, and N satisfies: M > N ≥ M / 2.

[0010] Furthermore, the preset species include: corn, potato, cassava, wheat, sweet potato, mung bean, lotus root, and pea.

[0011] Furthermore, the step of inputting the preprocessed spectral data into a pre-trained starch species classification model and outputting species identification results includes: automatically triggering a secondary verification mechanism when the identification result output by the starch species classification model is corn starch or wheat starch; the secondary verification mechanism includes: From the preprocessed spectral data, extract the average absorbance value A_avg within the preset characteristic wavelength range; The average absorbance value A_avg is compared with a first preset threshold Th1 and a second preset threshold Th2, where Th1 > Th2; If A_avg ≥ Th1, it is finally confirmed as wheat starch; if A_avg ≤ Th2, it is finally confirmed as corn starch; if Th2 < A_avg < Th1, an alert for abnormal spectral characteristics is output.

[0012] Furthermore, the spectral acquisition step also includes: When the turntable rotates, the portable near-infrared spectrometer is controlled to acquire spectra in an intermittently triggered scanning mode; wherein, a single spectral acquisition is triggered when the turntable rotates to a preset equally divided angle position, and K spectra are acquired in one complete rotation cycle, where K≥4; the average of the K spectra is used as the representative spectral data of the sample.

[0013] Secondly, this application also provides a rapid identification system for edible starch species, comprising: The spectral acquisition module is used to place the powdered starch sample to be tested in a petri dish and smooth the surface. The petri dish is then placed on a rotating turntable at a uniform speed. In a dark environment, a portable near-infrared spectrometer fixed on a support with the probe pointing vertically downward is used to dynamically acquire the raw spectral data of the sample in a diffuse reflectance manner. The spectral preprocessing module is used to sequentially perform standard normal variable transformation and Savitzky-Golay first-order derivative filtering on the raw spectral data. The species identification module is used to input preprocessed spectral data into a pre-trained starch species classification model and output species identification results.

[0014] Furthermore, the system also includes a training module for training a starch species classification model based on a process that includes label consistency auditing; the specific training process includes: Data preparation: Near-infrared spectral training sets of various starch species were obtained, and the spectra were preprocessed by standard normal variable transformation and Savitzky-Golay first derivative filtering to obtain the preprocessed full spectrum data. Base model training: Based on the preprocessed full-spectrum data, K-nearest neighbors, random forest and XGBoost classifiers were trained as base models respectively; Label auditing and optimization: Based on the predicted probability of the training samples by the base model, the consistency of the original sample labels is audited, and labels with inconsistent confidence are screened and corrected to form an optimized training set; Integration: The base model is retrained using the optimized training set and integrated using a weighted soft voting strategy, where the voting weight of the K-nearest neighbor classifier is 40%, and the voting weights of the random forest and XGBoost classifiers are each 30%.

[0015] Thirdly, this application also proposes an electronic device comprising: one or more processors, and a memory for storing one or more computer programs; the computer programs are configured to be executed by the one or more processors, and the programs include steps for performing the rapid identification method for edible starch species described in the first aspect.

[0016] Fourthly, this application also proposes a storage medium storing a computer program; the program is loaded and executed by a processor to implement the steps of the rapid identification method for edible starch species as described in the first aspect.

[0017] As described above, the rapid identification method and system for edible starch species provided in this application involves placing the powdered starch sample to be tested in a petri dish, smoothing the surface, and then placing the petri dish on a rotating turntable at a uniform speed. In a dark environment, a portable near-infrared spectrometer fixed to a support with the probe pointing vertically downwards is used to dynamically acquire the raw spectral data of the sample using diffuse reflectance. The raw spectral data is then subjected to standard normal transformation and Savitzky-Golay first-order derivative filtering. The preprocessed spectral data is input into a pre-trained starch species classification model, and the species identification result is output. By constructing a standardized dynamic spectral acquisition process, optimizing the spectral preprocessing scheme, and integrating the classification model, this method solves the problems of traditional starch species identification methods being time-consuming, costly, and dependent on laboratory environments, thus failing to meet the needs of rapid on-site screening. It has the significant advantages of enabling rapid, non-destructive, and highly accurate on-site identification of various common edible starches, providing an efficient and reliable technical means. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating the steps of the rapid identification method for edible starch species disclosed in an embodiment of the present invention; Figure 2 This is a schematic diagram of starch product with a smoothed surface in a polystyrene petri dish, as disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of the rapid identification system for edible starch species disclosed in an embodiment of the present invention. Detailed Implementation

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these embodiments belong; the terminology used herein and in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit these embodiments; the terms in the specification of these embodiments and the foregoing description of the accompanying drawings include and have, and any variations thereof, and are intended to cover non-exclusive inclusion. The terms first, second, etc., in the specification of these embodiments and the foregoing drawings are used to distinguish different objects and not to describe a particular order.

[0021] The implementation details of the technical solution in this embodiment are described in detail below: Firstly, this embodiment proposes a rapid identification method for edible starch species, such as... Figure 1 As shown, the method includes: S101, Spectral acquisition steps: Place the powdered starch sample to be tested in a petri dish and smooth the surface. Place the petri dish on a rotating turntable at a uniform speed. In a dark environment, use a portable near-infrared spectrometer fixed on a support with the probe pointing vertically downward to dynamically acquire the raw spectral data of the sample in a diffuse reflectance manner.

[0022] Specifically, in this embodiment, a SupNIR1500 series portable near-infrared analyzer (Juguang Technology, Hangzhou, China) was used, with a wavelength range of 1000–1800 nm and a resolution of 1.5 nm. Spectral acquisition was performed using diffuse reflectance. An electronic balance (ME204E / 02, Mettler Toledo, Switzerland) with an accuracy of 0.1 mg was used. A motorized turntable (custom-made, adjustable speed) and polystyrene petri dishes (90 mm × 15 mm, Beijing Lanjieke Technology Co., Ltd., China) were also used. During spectral acquisition, approximately one part of the starch sample was evenly filled into the polystyrene petri dish, and the surface was leveled with a stainless steel ruler to ensure a smooth sample surface. The petri dish was placed on the motorized turntable, rotating at 5 revolutions / min. Measurements were performed using a portable NIR probe fixed to an iron stand, with the probe center vertically aligned with the center of the petri dish, approximately 1 cm from the sample surface. Ambient light was turned off in a self-made darkroom to eliminate stray light interference. After the instrument warmed up for 30 minutes, background correction was performed using a standard white board. Ten consecutive spectral acquisitions were performed on each sample, with each scan consisting of 10 iterations, to obtain an average spectrum with a high signal-to-noise ratio. For example... Figure 2 The diagram shown is a schematic of the starch product with a smoothed surface in a polystyrene petri dish in this embodiment.

[0023] Specifically, leveling the surface of the powdered starch sample in a petri dish involves pouring an appropriate amount of starch sample into a disposable polystyrene petri dish and using a stainless steel ruler or similar tool to scrape the sample surface until it is completely smooth and flush with the edge of the petri dish. The purpose of this operation is to eliminate differences in diffuse reflection light paths and scattering effects caused by uneven powder accumulation and surface roughness, ensuring that the physical state of the sample surface is consistent for each measurement, thereby reducing spectral baseline shifts and intensity fluctuations introduced by these factors.

[0024] Placing the culture dish on a uniformly rotating turntable means placing the smoothed culture dish at the center of the motorized turntable's platform. The turntable rotates at a uniform speed (e.g., 5 revolutions per minute) to ensure continuous variation of the sample area illuminated by the probe within the measurement cycle. This is equivalent to adding a spatial dimension scan to the static single-point measurement, allowing the final average spectrum to represent information about a ring-shaped region of the sample. This significantly improves the sample representativeness and spectral repeatability of a single measurement, effectively overcoming the influence of potential microscopic inhomogeneities in powder samples on the measurement results.

[0025] Furthermore, in a darkroom environment, a portable near-infrared spectrometer with the probe fixed to a stand and vertically downward was used for data acquisition. Two key operational points were considered. First, the operation took place in a darkroom environment, i.e., within a closed or light-shielded device. This was intended to completely shield the spectrometer from visible light and other stray light sources, ensuring that the signal entering the spectrometer detector came entirely from diffuse reflection of the sample from the instrument's own light source, thus eliminating ambient light noise interference. Second, the probe was fixed vertically downward, meaning it was vertically fixed using a rigid support such as a metal stand, with the center of its emission / reception window aligned with the center of the turntable, approximately 1 cm from the sample surface. This fixed probe ensured absolute consistency in the beam incident angle and receiving geometry for each measurement, a crucial geometric condition for guaranteeing spectral comparability; the vertical downward arrangement conformed to the standard optical configuration for diffuse reflectance measurements.

[0026] In practice, dynamic acquisition using diffuse reflection refers to starting the turntable and spectrometer. As the turntable rotates, near-infrared light emitted by the spectrometer's light source is perpendicularly incident on the rotating sample surface. After multiple scattering and absorption within the sample, some of the diffusely reflected light is received by the probe and converted into a spectral signal. Dynamic acquisition, unlike traditional static measurements, involves signal integration during the sample's uniform circular motion relative to the light beam. This method cleverly combines time integration and spatial averaging: within a rotation cycle, the probe collects reflected light signals from different points on the sample's circular path. The instrument software or subsequent data processing then performs an arithmetic average of the multiple spectra obtained during this period (e.g., 10 consecutive scans), ultimately outputting a raw spectral data set with a higher signal-to-noise ratio and stronger representativeness.

[0027] The spectral acquisition scheme in this application standardizes the physical state by leveling the surface, averages the spatial region by rotating the turntable, eliminates external light interference by using a darkroom environment, ensures consistent geometric conditions by vertically fixing the probe, and finally fuses multi-point, multi-spectral information into a single high-quality spectrum through dynamic acquisition. This series of operations works synergistically to fundamentally solve the industry-wide problems of poor reproducibility and insufficient representativeness in near-infrared diffuse reflectance measurements of powder samples, providing a stable, reliable, and high-quality raw data foundation for subsequent spectral preprocessing and model recognition. It is precisely this rigorous, standardized acquisition process that makes it possible to obtain stable spectral data comparable to that of desktop computers, which was previously difficult to obtain on portable devices, thus supporting the establishment of subsequent high-precision classification models.

[0028] S102, Spectral preprocessing step: Perform standard normal variable transformation and Savitzky-Golay first derivative filtering on the original spectral data in sequence.

[0029] Specifically, in this embodiment, the raw spectral data is often affected by physical scattering effects and random noise, directly impacting model performance. This study employs a two-step joint preprocessing strategy: first, SNV is used to eliminate baseline shifts and slope variations caused by differences in sample particle size and density; then, Savitzky-Golay first-order derivative filtering (window length 11, polynomial order 2) is applied to enhance spectral feature resolution and suppress low-frequency drift. It is worth noting that multivariate scattering correction (MSC) did not provide significant performance improvement in preliminary experiments and was therefore not included in the final process. This preprocessing scheme effectively improves the spectral signal-to-noise ratio, laying the foundation for subsequent modeling.

[0030] Specifically, performing a standard normal variable transformation on the raw spectral data sequentially refers to applying a mathematical processing method to each acquired raw near-infrared diffuse reflectance spectrum to eliminate spectral baseline drift and scattering effects caused by differences in the physical state of the sample. This transformation uses the absorbance values ​​at all wavelengths of each spectrum as a benchmark, calculates the mean and standard deviation of the spectral data, then subtracts the mean from the raw absorbance value at each wavelength and divides by the standard deviation. The mathematical expression is: X_SNV(λ) = [X_raw(λ) - μ] / σ, where X_raw(λ) is the raw absorbance at wavelength λ, and μ and σ are the mean and standard deviation of the absorbance at all wavelengths of the spectrum, respectively. The purpose of this operation is to ensure that the data distribution of each spectrum follows a standard normal distribution with a mean of 0 and a standard deviation of 1, thereby effectively suppressing the overall differences in diffuse reflectance intensity caused by inconsistencies in sample particle size, distribution density, and surface roughness, allowing subsequent analysis to focus more on the spectral shape information reflecting the chemical composition of the sample.

[0031] The Savitzky-Golay first-order derivative filtering of the spectrum after SNV transformation refers to calculating the first derivative of the spectrum using the Savitzky-Golay convolution smoothing method. This process requires pre-setting two key parameters: window length and polynomial order. In a preferred embodiment of this application, the window length is fixed at 11 data points (corresponding to a spectral range of approximately 16.5 nm), and the polynomial order is fixed at 2. The processing procedure is as follows: at each wavelength point in the spectrum, a local window is formed by selecting 11 data points before and after that point; within this window, a second-order polynomial is used for least-squares fitting; then, the first derivative value of the fitted polynomial at the center point is directly calculated, and this value is used as the output value after processing at that center wavelength point. By sliding the window through the entire spectrum, a first-order derivative spectrum is finally obtained.

[0032] In practical applications, the choice of a window length of 11 and a polynomial order of 2 is based on a trade-off between the spectral characteristics of starch and the noise level. A wider window (e.g., 11 points) provides sufficient smoothness to effectively suppress high-frequency random noise, while the second-order polynomial ensures fitting flexibility while avoiding overfitting. This parameter combination was determined after optimization through grid search and cross-validation, achieving the best balance between enhancing spectral resolution (highlighting subtle differences between adjacent wavelengths) and suppressing noise. The core objectives of performing first-order derivative processing are twofold: first, to eliminate any possible residual baseline drift (derivative operations are insensitive to constant and linear terms); and second, to amplify the subtle differences in the positions and shapes of specific absorption peaks among different starch species. These differences may be masked by strong absorption backgrounds in the original spectrum or in spectra processed only by SNV, but they appear as clear positive and negative peaks in the first-order derivative spectrum, greatly improving the ability to distinguish spectral features.

[0033] This application's scheme constructs a dedicated preprocessing workflow for near-infrared spectra of powdered starch by sequentially combining standard normal variable transformation (SNV) with Savitzky-Golay first-order derivative filtering (SG) with specific parameters. SNV first normalizes physical interference in terms of amplitude, while SG first-order derivative then sharpens and highlights chemical fingerprint information in terms of shape. These two steps complement each other and are indispensable. If only SNV is used, subtle peak-valley differences in the spectrum remain insignificant, making it difficult for the model to learn; if the derivative of the original spectrum is directly calculated, the baseline tilt introduced by physical scattering will be amplified into severe interference. It is precisely this specific sequence and combination of specific parameters that allows spectral data with relatively low signal-to-noise ratios obtained from portable devices to be transformed into effective inputs with distinct features and high consistency, laying a crucial data foundation for subsequent machine learning models to achieve high-precision classification. This preprocessing scheme is a creative optimization for this application scenario (portable devices, powdered starch) and is one of the key technical aspects of achieving high accuracy in the overall method.

[0034] S103, Species identification step: Input the preprocessed spectral data into the pre-trained starch species classification model and output the species identification results.

[0035] Furthermore, the preset species include: corn, potato, cassava, wheat, sweet potato, mung bean, lotus root, and pea. The starch species classification model is trained based on a process that includes label consistency auditing; the specific training process includes: Data preparation: Near-infrared spectral training sets of various starch species were obtained, and the spectra were preprocessed by standard normal variable transformation and Savitzky-Golay first derivative filtering to obtain the preprocessed full spectrum data. Base model training: Based on the preprocessed full-spectrum data, K-nearest neighbors, random forest and XGBoost classifiers were trained as base models respectively; Label auditing and optimization: Based on the predicted probability of the training samples by the base model, the consistency of the original sample labels is audited, and labels with inconsistent confidence are screened and corrected to form an optimized training set; Integration: The base model is retrained using the optimized training set and integrated using a weighted soft voting strategy, where the voting weight of the K-nearest neighbor classifier is 40%, and the voting weights of the random forest and XGBoost classifiers are each 30%.

[0036] Specifically, the label auditing and optimization step refers to a mechanism for proactively identifying and correcting potential labeling errors in training data. Its core operation involves using three trained base models—K-Nearest Neighbors, Random Forest, and XGBoost—to predict the probability distribution of each sample in the training set and obtain its classification as belonging to each species. For a given sample, if multiple base models (e.g., at least two) show highly consistent predictions (for the highest-probability category, all with probabilities greater than 0.9), and these consistent predictions differ from the sample's original label, the sample is labeled as a low-confidence label sample or a suspected mislabeled sample. Subsequently, these labeled samples undergo manual review, and their labels are confirmed or corrected based on information such as sample source records and spectral curve morphology comparisons. The set of all audited and confirmed samples (including samples that did not trigger audits and corrected samples) is called the optimized high-confidence training set. This step aims to improve quality from the data source, ensuring that the data labels used for the final integrated model training are as accurate as possible, thus laying a solid foundation for the model's high accuracy and strong generalization ability.

[0037] The weighted soft voting strategy used in the ensemble model means that during the model application phase, for a new input spectrum, the three base models (K-Nearest Neighbors, Random Forest, and XGBoost) each output a probability vector representing the probability that the input belongs to each species. Weighted soft voting is not simply voting on the predicted class labels, but rather performing a weighted average of these three probability vectors. Specifically, the calculation is: P_ensemble(Class_i) = 0.4 * P_knn(Class_i) + 0.3 * P_rf(Class_i) + 0.3 * P_xgb(Class_i), where P_ensemble(Class_i) is the final probability given by the ensemble model that the sample belongs to the i-th class, and P_knn, P_rf, and P_xgb are the probabilities given by the three base models, respectively. Finally, the ensemble model outputs the class with the highest P_ensemble value as the discrimination result. The weight allocation (40%, 30%, 30%) is the result of optimization based on a large number of cross-validation experiments, in order to balance the stability (KNN is sensitive to local features) and generalization ability (RF and XGBoost have strong global pattern learning ability) of different algorithms.

[0038] In practical applications, the entire training process forms a complete closed loop from data cleaning to heterogeneous ensemble. Data preparation and preprocessing ensure the consistency of input features; independently trained heterogeneous base models (KNN based on distance, RF based on decision tree forest, and XGBoost based on gradient boosting) learn the mapping relationship between spectra and species from different perspectives, providing diverse expert opinions; label auditing acts as a quality control post, using the consensus of these experts to reverse-verify the reliability of the original data, significantly reducing noise introduced by human annotation negligence or sample confusion; finally, weighted soft voting ensemble plays the role of an intelligent decision-making committee. It does not treat all expert opinions equally, but rather weights and synthesizes their opinions based on historical performance (reflected by weights), thus making a more robust and accurate final decision than any single expert.

[0039] This application's solution creatively solves the generalization performance bottleneck caused by label noise in training data and the limitations of a single model in spectral modeling by designing and implementing a model training process that includes label consistency auditing and optimized weighted soft voting integration. The label auditing mechanism actively cleans the training set, improving the reliability of the learning foundation; while the weighted soft voting integration of heterogeneous models effectively combines the advantages of different algorithms, enhancing the model's ability to discriminate complex, highly similar spectral patterns and its overall robustness. It is this systematic model construction method that enables this application to achieve high-precision, high-reliability rapid identification of eight starch species using spectra collected by portable devices, meeting the stringent requirements of accuracy and stability for on-site detection.

[0040] In this embodiment, five-fold hierarchical cross-validation is used to evaluate the performance of four mainstream classifiers in the full-spectrum feature space: K-Nearest Neighbors (KNN), Random Forest (RF), Partial Least Squares Discriminant Analysis (PLS-DA), and XGBoost. All models undergo hyperparameter tuning via grid search, with balanced accuracy as the evaluation metric. Finally, a weighted soft-voting ensemble model is constructed. Based on the results of multiple experimental pre-analysis, the predicted probabilities of KNN (40%), RF (30%), and XGBoost (30%) are fused to further improve generalization ability and robustness.

[0041] To evaluate the reliability of manually labeled data in a spectral dataset, a label consistency auditing method based on high-confidence predictions from multiple models was designed and implemented. The core assumption of this method is that when multiple high-performing classification models give highly consistent and confident predictions for a sample, and these predictions are inconsistent with the original labels, labeling errors may exist. Similar consensus-based prediction strategies have proven effective in cleaning mislabeled data from spectral modeling. Addressing the class imbalance problem in the dataset, this study compares and analyzes two mainstream strategies: resampling and cost-sensitive learning. Specifically, the classic Synthetic Minority Oversampling Technique (SMOTE) is used to synthesize new minority class samples to balance the training set; simultaneously, a class-weighted method is employed, assigning higher weights to the minority class in the loss function to balance the contributions of different classes to model training.

[0042] Furthermore, the method further includes: for the same sample to be tested, M parallel spectra are continuously acquired in the spectral acquisition step, and spectral preprocessing and species identification steps are performed respectively to obtain M initial identification results; in the M initial identification results, if the frequency of occurrence of a preset species exceeds a preset threshold N, the sample to be tested is determined to belong to the preset species; where M is an integer greater than or equal to 3, and N satisfies: M > N ≥ M / 2.

[0043] Specifically, in this embodiment, the continuous acquisition of M parallel spectra refers to the process of acquiring M spectral curves continuously at fixed time intervals or in sync with the turntable during a complete sample measurement, without any movement or disturbance to the sample. These M spectra originate from the reflection signals of the same sample at different spatial locations (due to the turntable rotation) within a short period. They are chemically identical, but each contains independent measurement noise and minute spatial variability information, hence the term "parallel spectra." This operation is equivalent to performing M repeated measurements on the same sample, aiming to obtain a set of observational samples representative of the sample's spectral characteristics, providing a data basis for subsequent statistical judgment.

[0044] The consensus threshold N satisfies the condition M > N ≥ M / 2, which is the core design of the majority voting rule. N ≥ M / 2 (i.e., N is greater than or equal to half of M) means that to make a definitive species determination, more than half of the parallel spectral measurements must agree. This is a reasonable and rigorous threshold based on statistics, ensuring that the final conclusion is based on the mainstream opinion (majority consensus) in the measurement results, rather than accidental consistency that may be caused by random errors. For example, when M=10, N should be at least 5, but in practice, N=8 (i.e., 80% consensus is required) is often set for higher confidence. The condition M > N ensures the practicality of the rule, avoiding overly stringent requirements such as unanimous approval (N=M), which could easily lead to an inability to make a determination due to a single accidental error. This threshold design achieves an optimal balance between the reliability of the determination and the practicality of the method.

[0045] In practical applications, the execution of the majority vote constitutes the final decision-making checkpoint for rapid on-site screening. Its workflow can be detailed as follows: The system records and displays M identification results sequentially, for example, [corn, corn, wheat, corn, corn, corn, corn, wheat, corn, corn]. Subsequently, the system counts that corn appears 8 times and wheat appears 2 times. If the preset threshold N=8, then the frequency of corn (8) equals N, and it is determined to be corn starch; if N=9, then no species meets the standard, and the system will output the conclusion that no consensus can be reached, and it is recommended to retest or send it to the laboratory for confirmation. This mechanism has a dual advantage: First, it has a strong ability to resist random interference. Even if a few measurements are misjudged due to noise or other reasons, as long as the correct results occupy the majority (≥N), the final conclusion is still correct; Second, the conclusion is clear and auditable. If the output is not A, it is a retest, and the original M results are retained for traceability, avoiding the risk of forcibly judging based on a single uncertain result.

[0046] This application's solution creatively transforms the quality control concept of averaging repeated measurements in the laboratory into an intelligent decision-making logic for consensus-building through repeated measurements, suitable for rapid on-site testing scenarios, by introducing a majority voting rule based on parallel spectroscopy. This rule directly addresses the pain point of potentially insufficient stability of single measurements by portable devices in complex on-site environments. Through simple multiple measurements and statistics, it significantly improves the reliability and robustness of single-sample identification results. It makes the entire system not just a spectral analyzer, but also an intelligent decision-making terminal with inherent fault-tolerance mechanisms and clear quality control standards. When consensus is high, it can provide a highly confident definitive answer; when consensus is low, it honestly acknowledges uncertainty and triggers subsequent processes. This greatly enhances the practical value and credibility of the method in key scenarios such as market supervision and production line acceptance.

[0047] Furthermore, the step of inputting the preprocessed spectral data into a pre-trained starch species classification model and outputting species identification results includes: automatically triggering a secondary verification mechanism when the identification result output by the starch species classification model is corn starch or wheat starch; The secondary verification mechanism includes: extracting the average absorbance value A_avg within a preset characteristic wavelength range from the preprocessed spectral data; comparing the average absorbance value A_avg with a first preset threshold Th1 and a second preset threshold Th2, where Th1 > Th2; if A_avg ≥ Th1, then it is finally confirmed as wheat starch; if A_avg ≤ Th2, then it is finally confirmed as corn starch; if Th2 < A_avg < Th1, then an alert for abnormal spectral characteristics is output.

[0048] Specifically, in this embodiment, the secondary verification mechanism includes: accurately extracting the arithmetic mean of the absorbance at all wavelength points within the preset characteristic wavelength range [1560 nm, 1640 nm] from the spectral data after standard normal variable transformation and Savitzky-Golay first-order derivative filtering, denoted as A_avg; comparing the A_avg value with a pre-calibrated first preset threshold Th1 and second preset threshold Th2, where Th1 > Th2; if A_avg ≥ Th1, then it is finally confirmed as wheat starch; if A_avg ≤ Th2, then it is finally confirmed as corn starch; if Th2 < A_avg < Th1, then an output warning indicating abnormal spectral characteristics and suggesting a review is issued.

[0049] Specifically, the automatically triggered secondary verification mechanism refers to a conditional judgment process executed by the system after obtaining the preliminary identification results of the integrated model. The logic is that this mechanism is activated only when the preliminary result is corn starch or wheat starch. This is mainly because corn starch and wheat starch originate from the same family (Poaceae), and their macroscopic physicochemical properties and microscopic molecular structures (such as the straight-chain / branched chain ratio and crystal type) are extremely similar under certain conditions, resulting in highly similar near-infrared spectra in overall shape, making them one of the most easily confused species pairs in model discrimination. Therefore, this application specifically adds an independent verification checkpoint based on different principles for this difficult pair of samples, aiming to distinguish them with higher confidence and issue early warnings for abnormal situations.

[0050] The extraction of the average absorbance value A_avg within the preset characteristic wavelength range [1560 nm, 1640 nm] is the core quantitative step in the secondary verification. The selection of the wavelength range [1560 nm, 1640 nm] is not arbitrary but based on in-depth analysis of the spectra of a large number of pure corn and wheat starches. This range lies in the strong absorption region of the OH first-order overtone and CH combination overtone in the near-infrared spectrum, and is extremely sensitive to the state of bound water in starch molecules and the chemical environment of the CH groups on the glucose ring. Studies have found that although the overall spectra of the two starches are similar, within this specific sub-range, due to subtle differences in molecular structure and hydration characteristics, the average absorbance A_avg of wheat starch is generally and consistently higher than that of corn starch. The calculation of A_avg, which involves summing the absorbance values ​​at each wavelength point within this range and dividing by the total number of wavelength points, is a simple and robust feature extraction method that can effectively characterize the overall absorption intensity of the sample in this key chemically sensitive wavelength band.

[0051] In practical applications, comparing and determining A_avg with thresholds Th1 and Th2 constitutes a clear three-stage decision rule. Thresholds Th1 and Th2 are pre-calibrated through statistical analysis of a large number of known standard spectral samples of pure corn starch and pure wheat starch: Th1 is usually set near the low quantile (e.g., the 5th percentile) of the A_avg value distribution of pure wheat starch to ensure that the A_avg value of the vast majority of wheat starch is not lower than it; Th2 is set near the high quantile (e.g., the 95th percentile) of the A_avg value distribution of pure corn starch to ensure that the A_avg value of the vast majority of corn starch is not higher than it; and Th1 must be greater than Th2 to leave a buffer zone between the typical value ranges of the two species. During the determination: if A_avg is very high (≥Th1), its spectral characteristics strongly favor the typical pattern of wheat starch, so it is confirmed as wheat starch; if A_avg is very low (≤Th2), it strongly favors the typical pattern of corn starch, so it is confirmed as corn starch; if A_avg falls in the middle buffer zone (Th2 < A_avg < Th1), it means that the spectrum of the sample is neither completely like typical corn nor completely like typical wheat in this key feature, and may be mixed starch, slightly deteriorated starch, or from a special variety. Therefore, the system does not force a classification, but outputs a warning message, prompting the operator to use a more accurate method (such as DNA analysis) for verification.

[0052] This application's solution creatively combines a complex machine learning black-box model with an interpretable white-box rule based on explicit material knowledge by designing and implementing a secondary verification mechanism for corn / wheat starch. This mechanism acts as an efficient, specialized quality inspector, specifically verifying the model's conclusions at the highest-risk discrimination point (i.e., corn-wheat). Its advantages are: first, it enhances the robustness of key discriminations by using the stability of chemical characteristic differences to compensate for potential model instability on extremely similar samples; second, it possesses anomaly detection capabilities, with the buffer setting enabling the system to automatically identify non-standard samples with atypical spectral characteristics, which is difficult for simple classification models to achieve; third, it is simple and efficient to implement, involving only the calculation of an interval average and two comparisons, resulting in extremely low computational overhead and almost no increase in overall detection time. This collaborative working mode of model-led judgment, rule verification, and anomaly warning significantly enhances the reliability and practicality of the entire rapid identification system in real-world applications.

[0053] Furthermore, in this application, the spectral acquisition step further includes: when the turntable rotates, controlling the portable near-infrared spectrometer to acquire spectra in an intermittently triggered scanning mode; wherein, a single spectral acquisition is triggered when the turntable rotates to a preset equally divided angle position, and K spectra are acquired within one complete rotation cycle, where K≥4; the K spectra are averaged and used as representative spectral data of the sample.

[0054] Specifically, the intermittently triggered scanning method refers to a discontinuous spectral acquisition control strategy that is strictly synchronized with the angular position of the turntable. Its core lies in precisely binding the time dimension of spectral acquisition with the angular dimension of the sample's spatial position. Operationally, firstly, according to a preset number of acquisitions K (e.g., K=4), the 360 ​​degrees of the turntable are divided into K equal angular intervals (e.g., each interval is 90 degrees). The system monitors the turntable angle in real time through an encoder mounted on the rotating shaft. When the turntable rotates to the starting point of each equally divided angle (e.g., 0°, 90°, 180°, 270°), the encoder sends a synchronization pulse signal, which immediately triggers the near-infrared spectrometer to perform a complete spectral scan. During rotation outside of trigger points, the spectrometer remains in a waiting state. This mode of rotating to a specific position before acquiring data is called intermittently triggered scanning.

[0055] The triggering mechanism, which activates when the turntable rotates to a preset, equally spaced angle position, is crucial for achieving uniform spatial sampling. The equally spaced angle ensures that the trigger points are evenly distributed across the entire circumference of the sample within one cycle. This means that the K spectra ultimately used for averaging originate from K different, centrally symmetrical, and evenly spaced annular regions on the sample disk. For example, when K=4, the sampling points are equivalent to sampling once each in the east, south, west, and north directions of the sample disk. This is fundamentally different from continuous sampling, which may result in random distribution or overlap of sampling points due to mismatches between rotation speed and sampling frequency. This design ensures that the spatial information of the sample is systematically and unbiasedly collected, reflecting the overall average state of the sample to the greatest extent possible, and avoiding the risk of affecting the overall representativeness due to the accidental collection of a local anomaly (such as a small agglomerate or void).

[0056] In practical applications, averaging K spectra to obtain representative spectral data is a data fusion step based on statistical principles. Arithmetic averaging K (K≥4) independent spectra obtained spatially uniformly and temporally has significant benefits: First, it effectively suppresses random noise: random factors such as spectrometer electronic noise and environmental disturbances can be significantly reduced through averaging in multiple independent measurements, thereby improving the signal-to-noise ratio of a single representative spectrum. Second, it achieves spatial homogenization: powder samples may have microscopic inhomogeneities. Averaging multiple measurements at different points in space is equivalent to spatially integrating the chemical composition of the sample, resulting in a spectrum that better represents the overall chemical properties of the batch of samples and reduces the impact of local variations. Third, it overcomes motion artifacts: in intermittent triggering mode, the turntable and probe are in a defined relative geometric relationship at the moment of each acquisition (guaranteed by the trigger angle), avoiding minor spectral distortions (similar to motion blur) that may occur due to sample movement during continuous acquisition, ensuring that each spectrum involved in averaging has higher clarity and geometric consistency.

[0057] This application's solution creatively solves the balance between the systematic nature of spatial sampling and the quality of single measurements in dynamic diffuse reflectance measurement by introducing an intermittently triggered scanning acquisition mode. Through angle-synchronized triggering, it transforms what might otherwise be random dynamic sampling into a controlled, spatially uniform quasi-static sampling sequence. This mode not only inherits the advantages of spatial averaging in dynamic measurements but also endows this averaging process with systematicity, repeatability, and optimal spatial coverage through the rules of angle equalization and trigger synchronization. The resulting single representative spectrum exhibits significantly better data quality and information reliability than simple continuous acquisition averaging or single-point static acquisition, providing an exceptionally stable and reliable input foundation for subsequent preprocessing and model recognition. This is one of the key technological innovations enabling high-precision identification using portable devices.

[0058] Secondly, this application also provides a rapid identification system for edible starch species, such as... Figure 3 As shown, it includes: The spectral acquisition module 301 is used to place the powdered starch sample to be tested in a petri dish and smooth the surface. The petri dish is placed on a rotating turntable at a uniform speed. In a dark room environment, a portable near-infrared spectrometer fixed on a support with the probe pointing vertically downward is used to dynamically acquire the original spectral data of the sample in a diffuse reflectance manner. Spectral preprocessing module 302 is used to sequentially perform standard normal variable transformation and Savitzky-Golay first derivative filtering on the original spectral data; The species identification module 303 is used to input the preprocessed spectral data into the pre-trained starch species classification model and output the species identification results.

[0059] Furthermore, the system also includes a training module for training a starch species classification model based on a process that includes label consistency auditing; the specific training process includes: Data preparation: Near-infrared spectral training sets of various starch species were obtained, and the spectra were preprocessed by standard normal variable transformation and Savitzky-Golay first derivative filtering to obtain the preprocessed full spectrum data. Base model training: Based on the preprocessed full-spectrum data, K-nearest neighbors, random forest and XGBoost classifiers were trained as base models respectively; Label auditing and optimization: Based on the predicted probability of the training samples by the base model, the consistency of the original sample labels is audited, and labels with inconsistent confidence are screened and corrected to form an optimized training set; Integration: The base model is retrained using the optimized training set and integrated using a weighted soft voting strategy, where the voting weight of the K-nearest neighbor classifier is 40%, and the voting weights of the random forest and XGBoost classifiers are each 30%.

[0060] Thirdly, this application also proposes an electronic device comprising: one or more processors, and a memory for storing one or more computer programs; the computer programs are configured to be executed by the one or more processors, and the programs include steps for performing the rapid identification method for edible starch species described in the first aspect.

[0061] Fourthly, this application also proposes a storage medium storing a computer program; the program is loaded and executed by a processor to implement the steps of the rapid identification method for edible starch species as described in the first aspect.

[0062] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for rapid identification of edible starch species, characterized in that, include: Spectral acquisition steps: Place the powdered starch sample to be tested in a petri dish and smooth the surface. Place the petri dish on a rotating turntable at a uniform speed. In a dark environment, use a portable near-infrared spectrometer fixed on a support with the probe pointing vertically downward to dynamically acquire the raw spectral data of the sample in a diffuse reflectance manner. Spectral preprocessing steps: The raw spectral data are sequentially subjected to standard normal transformation and Savitzky-Golay first-order derivative filtering; Species identification steps: Input the preprocessed spectral data into the pre-trained starch species classification model and output the species identification results.

2. The method for rapid identification of edible starch species according to claim 1, characterized in that, The starch species classification model was trained based on a process that includes label consistency auditing; The specific training process includes: Data preparation: Near-infrared spectral training sets of various starch species were obtained, and the spectra were preprocessed by standard normal variable transformation and Savitzky-Golay first derivative filtering to obtain the preprocessed full spectrum data. Base model training: Based on the preprocessed full-spectrum data, K-nearest neighbors, random forest and XGBoost classifiers were trained as base models respectively; Label auditing and optimization: Based on the predicted probability of the training samples by the base model, the consistency of the original sample labels is audited, and labels with inconsistent confidence are screened and corrected to form an optimized training set; Integration: The base model is retrained using the optimized training set and integrated using a weighted soft voting strategy, where the voting weight of the K-nearest neighbor classifier is 40%, and the voting weights of the random forest and XGBoost classifiers are each 30%.

3. The method for rapid identification of edible starch species according to claim 2, characterized in that, The method further includes: For the same sample to be tested, M parallel spectra are continuously acquired in the spectral acquisition step, and spectral preprocessing and species identification steps are performed respectively to obtain M initial identification results; In M initial identification results, if the frequency of occurrence of a preset species exceeds a preset threshold N, the sample to be tested is determined to belong to the preset species; where M is an integer greater than or equal to 3, and N satisfies: M > N ≥ M / 2.

4. The method for rapid identification of edible starch species according to claim 3, characterized in that, The preset species include: corn, potato, cassava, wheat, sweet potato, mung bean, lotus root, and pea.

5. The method for rapid identification of edible starch species according to claim 4, characterized in that, The step of inputting the preprocessed spectral data into a pre-trained starch species classification model and outputting species identification results includes: automatically triggering a secondary verification mechanism when the identification result output by the starch species classification model is corn starch or wheat starch; the secondary verification mechanism includes: From the preprocessed spectral data, extract the average absorbance value A_avg within the preset characteristic wavelength range; The average absorbance value A_avg is compared with a first preset threshold Th1 and a second preset threshold Th2, wherein Th1 > Th2; If A_avg ≥ Th1, it is finally confirmed as wheat starch; if A_avg ≤ Th2, it is finally confirmed as corn starch; if Th2 < A_avg < Th1, an alert for abnormal spectral characteristics is output.

6. The method for rapid identification of edible starch species according to claim 1, characterized in that, The spectral acquisition step also includes: When the turntable rotates, the portable near-infrared spectrometer is controlled to acquire spectra in an intermittently triggered scanning mode; wherein, a single spectral acquisition is triggered when the turntable rotates to a preset equally divided angle position, and K spectra are acquired in one complete rotation cycle, where K≥4; the average of the K spectra is used as the representative spectral data of the sample.

7. A rapid identification system for edible starch species, characterized in that, include: The spectral acquisition module is used to place the powdered starch sample to be tested in a petri dish and smooth the surface. The petri dish is then placed on a rotating turntable at a uniform speed. In a dark environment, a portable near-infrared spectrometer fixed on a support with the probe pointing vertically downward is used to dynamically acquire the raw spectral data of the sample in a diffuse reflectance manner. The spectral preprocessing module is used to sequentially perform standard normal variable transformation and Savitzky-Golay first-order derivative filtering on the raw spectral data. The species identification module is used to input preprocessed spectral data into a pre-trained starch species classification model and output species identification results.

8. The rapid identification system for edible starch species according to claim 7, characterized in that, The system also includes a training module for training a starch species classification model based on a process that includes label consistency auditing; the specific training process includes: Data preparation: Near-infrared spectral training sets of various starch species were obtained, and the spectra were preprocessed by standard normal variable transformation and Savitzky-Golay first derivative filtering to obtain the preprocessed full spectrum data. Base model training: Based on the preprocessed full-spectrum data, K-nearest neighbors, random forest and XGBoost classifiers were trained as base models respectively; Label auditing and optimization: Based on the predicted probability of the training samples by the base model, the consistency of the original sample labels is audited, and labels with inconsistent confidence are screened and corrected to form an optimized training set; Integration: The base model is retrained using the optimized training set and integrated using a weighted soft voting strategy, where the voting weight of the K-nearest neighbor classifier is 40%, and the voting weights of the random forest and XGBoost classifiers are each 30%.

9. An electronic device, the electronic device comprising: One or more processors, a memory for storing one or more computer programs; characterized in that the computer programs are configured to be executed by the one or more processors, the programs including steps for performing the rapid identification method for edible starch species as described in any one of claims 1-6.

10. A storage medium storing a computer program; characterized in that, The program is loaded and executed by a processor to implement the steps of the rapid identification method for edible starch species as described in any one of claims 1-6.