Grain multi-component decoupling prediction method, modeling method, device and storage medium
The grain multi-component decoupling prediction method based on cascade spectral features utilizes secondary and primary loop models combined with SHAP value analysis to dynamically adjust the contribution of spectral features, thus solving the problem of moisture interference on grain quality detection and achieving high-precision and robust grain component prediction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HEFEI UNIV
- Filing Date
- 2026-01-28
- Publication Date
- 2026-04-24
AI Technical Summary
In non-destructive testing of grain quality, the influence of moisture content on protein content leads to excessive reliance on moisture changes in the model, weakening the model's generalization ability and robustness, making it difficult to accurately predict multiple components of grain.
A multi-group decoupling prediction method for grain based on cascade spectral features is adopted. By constructing secondary and primary loop models, machine learning and interpretable artificial intelligence (such as SHAP value analysis) are used to screen spectral features, establish an inner and outer double closed-loop architecture, dynamically adjust the contribution of band features, suppress moisture interference signals, and enhance protein features.
It effectively suppressed moisture interference signals, improved the model's prediction accuracy and robustness, achieved high-precision decoupled prediction of multiple components of grain, and enhanced the model's generalization ability.
Smart Images

Figure CN121615512B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of non-destructive testing technology for agricultural products, specifically relating to a multi-component decoupling prediction method, modeling method, equipment, and storage medium for grains. Background Technology
[0002] In the research on non-destructive testing of grain quality based on spectral imaging, grain, as a biological organism, has a complex internal composition, and the spectral reflection signals of each component affect each other. This has long been a major bottleneck problem for spectral technology in the field of grain quality testing.
[0003] Taking the interaction between protein and water in soybeans as an example, the OH bonds in water molecules have extremely strong absorption characteristics in the near-infrared band, with absorption intensity far exceeding the intensity of characteristic signals of other components, making it one of the biggest influencing factors in spectral signal-based quality detection. Furthermore, since the moisture content in soybean seeds is usually negatively correlated with protein content, and moisture content is easily affected by environmental temperature and humidity fluctuations, the spectral signal of water not only physically masks the weak characteristic absorption of protein, but also establishes a statistical correlation with the target variable.
[0004] Therefore, the "shortcut learning" phenomenon is prone to occur in the process of spectral modeling, that is, the model relies too much on the change of sample moisture content and ignores the identification of the protein's inherent spectral signal, which seriously weakens the model's generalization ability and robustness.
[0005] Similar influences exist not only between protein and water in soybeans, but also between other components in other grains, presenting the same technical challenges. Summary of the Invention
[0006] The main objective of this invention is to provide a method, modeling method, equipment, and storage medium for predicting multi-group decomposition of grains, in order to overcome the shortcomings of the prior art.
[0007] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0008] In a first aspect, the present invention provides a grain multi-group decoupling prediction modeling method based on cascade spectral features, comprising:
[0009] A training database is provided, which includes spectral data of the target grain, calibration values of the first component content, and calibration values of the second component content.
[0010] The first band set that has a strong correlation with the content of the first component in the spectral data is extracted as the first component feature library, and the second band set that has a strong correlation with the content of the second component in the model prediction is extracted as the second component feature library.
[0011] Using the first component feature library as input and the first component content calibration value as label, a machine learning method is used to train the secondary loop model and output the prediction result of the first component.
[0012] Using the prediction results of the first component and the feature library of the second component as joint inputs, and the content calibration value of the second component as a label, the main loop model is trained using machine learning methods.
[0013] The secondary loop model and the primary loop model are trained iteratively. The combination of the obtained secondary loop model, primary loop model, and the indices of the first band set and the second band set is used as the grain multi-group decoupling prediction model. During the iteration process, the first band set and the second band set are adjusted by an automatic parameter optimization algorithm so that the feature contribution of the bands sensitive to the first component gradually decreases during the iteration process, while the feature contribution of the bands sensitive to the second component gradually increases.
[0014] Secondly, the present invention also provides a grain multi-group decoupling prediction method based on cascade spectral features, comprising:
[0015] Provide the measured spectral data of the target grain and the grain multi-group decoupling prediction model obtained by the above-mentioned grain multi-group decoupling prediction modeling method;
[0016] The component features of the measured spectral data are extracted based on the band index;
[0017] Input the features of the first component into the sub-loop model to obtain the predicted value of the first component;
[0018] The second component's features and the first component's predicted value are input into the main loop model to obtain the second component's predicted value.
[0019] Thirdly, the present invention also provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program, and the computer program is executed by the processor to perform the above-described grain multi-group decoupling prediction modeling method, or to run the grain multi-group decoupling prediction model obtained by the above-described grain multi-group decoupling prediction modeling method.
[0020] Fourthly, the present invention also provides a readable storage medium storing a computer program, wherein the computer program, when run, executes the above-described grain multi-group decoupling prediction modeling method, or the readable storage medium stores a grain multi-group decoupling prediction model obtained by the above-described grain multi-group decoupling prediction modeling method.
[0021] Compared with the prior art, the beneficial effects of the present invention include at least the following:
[0022] This invention not only corrects spectral reflectance values but also constructs a closed-loop prediction mechanism based on the state of the first component. It combines interpretable artificial intelligence to dynamically screen bands and establishes an inner and outer double closed-loop cascade architecture of the main loop and the secondary loop. This architecture can continuously adjust parameters and update the feature library, thereby systematically suppressing interference signals and strengthening the intrinsic features of the second component. The introduction of an interpretable feature screening mechanism allows for real-time evaluation of the importance of each band during model training, achieving a combination of data-driven model and spectral physical priors. This effectively avoids the problem of declining generalization performance caused by the model relying on the "shortcut features" of the first component during training.
[0023] The above description is merely an overview of the technical solution of the present invention. In order to enable those skilled in the art to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of the present invention are described below in conjunction with detailed drawings. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 The diagram illustrates the steps of a soybean protein content detection method based on feature fusion, provided in a typical embodiment of the present invention.
[0026] Figure 2 This illustrates a data graph after preprocessing with a filter and a standard normal variable, as provided in a typical embodiment of the present invention.
[0027] Figure 3 A flowchart of dynamic SHAP feature filtering provided in a typical embodiment of the present invention is shown;
[0028] Figure 4 This invention provides a typical embodiment of the invention, showing the characteristic importance and spectral comparison before and after removing the water-sensitive band.
[0029] Figure 5 This diagram illustrates the characteristic distribution of hyperspectral moisture bands in soybeans according to a typical embodiment of the present invention.
[0030] Figure 6 This diagram illustrates a SHAP-based feature interpretation diagram for predicting soybean protein after water removal, provided in a typical embodiment of the present invention.
[0031] Figure 7The diagram shows a comparison of the water-protein spectral decoupling effects provided in a typical embodiment and a comparative embodiment of the present invention. Detailed Implementation
[0032] In view of the shortcomings of the prior art, this invention has been proposed after long-term research and extensive practice. The technical solution, its implementation process and principles will be further explained below.
[0033] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0034] The purpose of this invention is to overcome the severe interference of strong moisture absorption effect on the quantitative analysis of trace components such as proteins in existing near-infrared hyperspectral technology, as well as the technical shortcomings of single models such as poor robustness and weak physical interpretability. This invention proposes a non-destructive detection method and system based on spectral decoupling using feature selection and a cascade architecture with fused interpretable learning. This method, through a dual-driven approach of mechanism and data, effectively eliminates strong interference signals under complex biological mechanisms, constructing a high-precision and generalizable grain quality detection model.
[0035] To achieve the above objectives, this invention first provides a grain multi-group decoupling prediction modeling method based on cascade spectral features, which includes the following steps:
[0036] A training database is provided, which includes spectral data of the target grain, calibration values of the first component content, and calibration values of the second component content.
[0037] The first band set that has a strong correlation with the content of the first component in the spectral data is extracted as the first component feature library, and the second band set that has a strong correlation with the content of the second component in the model prediction is extracted as the second component feature library.
[0038] Using the first component feature library as input and the first component content calibration value as label, a machine learning method is used to train the secondary loop model and output the prediction result of the first component.
[0039] Using the prediction results of the first component and the feature library of the second component as joint inputs, and the content calibration value of the second component as a label, the main loop model is trained using machine learning methods.
[0040] The secondary loop model and the primary loop model are trained iteratively. The combination of the obtained secondary loop model, primary loop model, and the indices of the first band set and the second band set is used as the grain multi-group decoupling prediction model. During the iteration process, the first band set and the second band set are adjusted by an automatic parameter optimization algorithm so that the feature contribution of the bands sensitive to the first component gradually decreases during the iteration process, while the feature contribution of the bands sensitive to the second component gradually increases.
[0041] As a typical example of the above technical solution, taking the decoupling prediction of protein and moisture in soybeans as an example, a specific embodiment of the present invention provides a grain moisture and protein decoupling prediction modeling method based on cascade spectral features, characterized by comprising the following steps:
[0042] Step 1: Use a hyperspectral imaging system to obtain spectral data of the sample from 400 to 1700 nm and the sample moisture and protein calibration values. Specifically, you can place the same type of sample in different moisture environments and obtain grain samples with different moisture content based on the placement time.
[0043] Step 2: Using the SHAP interpretable artificial intelligence method, construct spectral databases for water and protein characteristics respectively;
[0044] The method for constructing the moisture feature library is as follows: based on the SHAP value analysis, the contribution of the full-band spectrum to the moisture content is analyzed, and a set of bands with high contribution is selected.
[0045] The protein feature library is constructed by analyzing the contribution of the full-band spectrum to protein content based on SHAP values and screening out the set of bands with high contribution.
[0046] Step 3: By defining auxiliary variable prediction modeling in the secondary loop and target prediction modeling in the primary loop, a decoupled modeling system for cascaded sample moisture and protein based on spectral information is established.
[0047] Secondary loop auxiliary variable prediction modeling: Using the established moisture spectral database as feature input, machine learning methods are employed for model training, and the optimal prediction model is obtained through feedback optimization using a defined moisture spectral modeling loss function. The moisture state value can be a continuous value or a discrete gradient value, and moisture is also used as one of the input features for primary loop target modeling.
[0048] Main loop target prediction modeling: Using the spectral band data in the protein feature library and the moisture prediction results of the sub-modeling loop as common inputs, a machine learning method is used to model the protein content. By defining a protein spectral modeling loss function, feedback optimization is performed to obtain the best prediction model and predict the final protein content of the grain.
[0049] The embodiments of the present invention exemplarily propose a soybean protein prediction method based on cascade decoupling of spectral features. This method includes two levels: moisture secondary loop closure and protein main loop closure.
[0050] 1. The moisture secondary loop constructs a moisture feature library through SHAP feature analysis, and divides the moisture state of the samples into at least three, preferably five moisture gradient levels, according to the degree of moisture influence on the spectrum; for each gradient, the difference function between the absolute dry spectrum and the water-containing spectrum is fitted to establish a fitting model for the moisture-spectral difference.
[0051] 2. The protein main loop is modeled only in the bands corresponding to the protein feature library, and the water state labels output by the sub-loop are fused as the model input variables;
[0052] 3. During the training process, an automatic parameter optimization algorithm is introduced. With the error compensation function and model prediction performance as the objectives, the spectral feature subset and model hyperparameters are dynamically adjusted to form a cascade iterative update mechanism between the main and auxiliary loops, thus constructing an inner and outer double closed-loop spectral decoupled modeling architecture.
[0053] Verification has shown that the cascade closed-loop mechanism proposed in this invention can gradually reduce the feature contribution of the water-sensitive band during the iteration process, while the protein-related band is gradually strengthened, thereby realizing the active extraction of intrinsic protein features, avoiding model overfitting to water information, and improving prediction performance and robustness.
[0054] Some existing technologies have also proposed technical solutions for the decoupling prediction of moisture and protein, such as Chinese invention patent CN117191739A: a method and system for quantitative detection of near-infrared spectroscopy to correct moisture interference. However, the spectral correction method provided by this technical solution only uses the difference between absolutely dry samples and water-containing samples for one-time correction, which is an open-loop penalty processing method. It fails to establish the coupling and decoupling relationship between moisture and protein in spectral space, and is obviously inferior to the present invention in terms of model accuracy and generalization ability.
[0055] In some implementations, the extraction process of the first component feature library and the second component feature library specifically includes:
[0056] Calculate the contribution of each band in the spectral data to the content of the first component or the content of the second component by calculating the band SHAP value; sort each band based on the band SHAP value, and select the corresponding bands in the high band SHAP value interval as the first band set or the second band set.
[0057] In some implementations, the SHAP value of the band is calculated as follows:
[0058]
[0059]
[0060] in, This represents the band SHAP value corresponding to band j; This indicates the number of grain samples in the training database; Indicates sample The SHAP value of a sample in band j; S represents a feature subset that does not include band j; Represents the set of all frequency bands; Indicates the use of feature subsets The predicted value of the time feature analysis model, which is independent of the secondary loop model and the primary loop model, where X belongs to S; This represents the set of feature subset X and band j.
[0061] It should be noted that two different prediction models are involved in the implementation of this invention:
[0062] Feature analysis model: Used to initially calculate the SHAP contribution value of each band in order to screen out an initial feature library that is strongly correlated with the first and second components. This model is no longer used after the initial screening of the feature library is completed. Its structure can be the same as the following cascade prediction model, or it can be somewhat different, as long as it can perform the corresponding function.
[0063] Cascade prediction model: including the secondary loop model and the primary loop model, the final grain multi-group decoupled prediction model is constructed through iterative training, and is used to predict the component content of the actual test sample.
[0064] Regarding how to divide it, in some implementation schemes, the method specifically includes:
[0065] The spectral data is divided into a single first component sensitive band set, a single second component sensitive band set, and a two-component overlapping sensitive band set according to the SHAP value of the band. The first band set includes the single first component sensitive band set, and the second band set includes the single second component sensitive band set and the two-component overlapping sensitive band set.
[0066] Calculate the compensation variable for the difference between the single second component sensitive band set and the double component overlapping sensitive band set caused by the difference in the content of different first components;
[0067] During iterative training, the corresponding band ranges of the single second component sensitive band set and the double component overlapping sensitive band set are adjusted according to the prediction results of the first component and the difference compensation variable, and the corresponding band range of the single first component sensitive band set is adjusted according to the changes of the double component overlapping sensitive band set.
[0068] In some implementations, the prediction result of the first component is represented as multiple discrete levels, and the difference compensation variable is the inter-level compensation amount corresponding to the discrete level.
[0069] In the above embodiments, the moisture and protein content of soybeans are still taken as an example. The largest value The feature library X, composed of several wavebands, was selected to identify the feature library with the highest correlation to water content. w =(λ w1 …λ wk (excluding the portion overlapping with the protein, where λ) w1 …λ wk This represents the set of wavelength values that are highly correlated with water content but not highly correlated with proteins. It also includes the feature library X, which shows the highest protein correlation. p =(λ p 1 …λ pk (excluding the portion overlapping with water, where λ) p 1 …λ pk This represents the set of wavelength values that are highly correlated with proteins but not highly correlated with water. Let X be... w With X p The m overlapping bands are X c =(λ c 1 …λ cm ), where λ c 1 …λ cm Indicates λ w1 …λ wk With λ p 1 …λ pk The set of overlapping wavelength values.
[0070] In some implementations, the difference compensation variable is calculated as follows:
[0071]
[0072]
[0073] in:
[0074] Indicates the moisture level Below, the difference compensation variable for the sensitive band set of the second component;
[0075] Indicates the moisture level Below, the concentration of sensitive bands with overlapping dual components is the first Difference compensation variables for each band;
[0076] It is a function about wavelength obtained by fitting the spectral difference curve;
[0077] Indicates moisture level The corresponding basic compensation error;
[0078] Indicates the first The water-protein coupling coefficient of each overlapping band, with a value greater than 1, is used to characterize the intensity of interference from water changes in that overlapping band.
[0079] Specifically, for example, a hyperspectral imaging system can be used to acquire spectral data of the sample from 400 to 1700 nm. Moisture can then be divided according to its gradient. Discrete levels For the same sample, the difference in moisture content caused X p X c Spectral differences are denoted as: and Establishment by fitting the spectral difference curve. , δ l This is the basic compensation error corresponding to moisture level l.
[0080] In some implementations, the process of the automatic parameter optimization algorithm is represented as follows:
[0081]
[0082]
[0083] in, Indicates the first The single second component sensitive band set at the next iteration; Indicates the iteration round; Indicates the first The first iteration of the two-component overlapping sensitive band set Values for each band.
[0084] As a specific example, target prediction modeling for the primary and secondary loops can employ various machine learning algorithms, such as XGBoost, whose prediction function is... The sum of the functions of the regression trees is expressed as:
[0085]
[0086] in For the first The joint feature vector of each sample Let k represent the regression tree space, where k is the index of the regression tree. Indicates the first The prediction function of each regression tree is based on the input features. Output a predicted value.
[0087] A feature subset selection method based on prediction accuracy feedback is used for both the auxiliary variable prediction modeling of the secondary loop and the target prediction modeling of the primary loop. Simultaneously, the auxiliary prediction results of the secondary loop are used as new feature values input into the target modeling process of the primary loop. For X... w With X p Non-overlapping bands are retained, and m overlapping bands X are retained. c Different moisture contents can interfere with protein prediction. In order to solve the nonlinear relationship between moisture content and spectral reflectance, the characteristic spectral reflectance value of the protein is adjusted according to the following formula.
[0088] Let X w With X p The m overlapping bands are X c =(λ c 1 …λ cm )
[0089]
[0090]
[0091] X p ( )and The reflectance values of the protein and the overlapping spectral bands after iteration. This represents the reflectance values of the current protein and the overlapping spectral bands. and This is the offset value for the overlapping spectral bands of water and protein at level l. This value varies depending on the magnitude of the water gradient, thereby reducing the mutual influence between water and protein spectra and decoupling the spectral information of water and protein. The adjusted band is then re-applied to X. p Constructing a protein prediction spectral feature database X p ’ =[X p ( ), ].
[0092] In the process of selecting the feature subset, each feature subset Soli is defined, and the corresponding spectral feature X is selected. p ’Train the prediction model and calculate the objective function: f(solid) = ,in , These are the weighting coefficients. The proportion of the feature subset. Let m be the target prediction mean square error, m be the total number of bands before screening, and k be the number of bands after screening.
[0093] This invention also provides a further application of the above technical solution, namely, a grain multi-group decoupling prediction method based on cascade spectral features, which includes the following steps:
[0094] Provides the measured spectral data of the target grain and the grain multi-group decoupling prediction model obtained by the grain multi-group decoupling prediction modeling method provided in any of the above embodiments;
[0095] The component features of the measured spectral data are extracted based on the band index;
[0096] Input the features of the first component into the sub-loop model to obtain the predicted value of the first component;
[0097] The second component's features and the first component's predicted value are input into the main loop model to obtain the second component's predicted value.
[0098] This invention also provides an electronic device, characterized in that it includes a processor and a memory, wherein the memory stores a computer program, and the computer program is executed by the processor to perform the grain multi-group decoupling prediction modeling method provided in any of the above embodiments, or to run the grain multi-group decoupling prediction model obtained by running the grain multi-group decoupling prediction model provided in any of the above embodiments.
[0099] This invention also provides a readable storage medium, characterized in that the readable storage medium stores a computer program, which, when run, executes the grain multi-group decoupling prediction modeling method provided in any of the above embodiments, or the readable storage medium stores a grain multi-group decoupling prediction model obtained by the grain multi-group decoupling prediction modeling method provided in any of the above embodiments.
[0100] The technical solution of the present invention will be further described in detail below through several embodiments and in conjunction with the accompanying drawings. However, the selected embodiments are only for illustrating the present invention and do not limit the scope of the present invention.
[0101] Example 1
[0102] This embodiment illustrates a grain moisture-protein decoupling prediction modeling method based on cascaded spectral features. Its core lies in interpreting the spectral interference mechanism of moisture through interpretability analysis and constructing a prediction framework that uses cascaded spectral feature information and performs decoupling modeling through feedback of main and secondary loop information. The method specifically includes the following steps:
[0103] First, an interpretable artificial intelligence (AI) database was constructed using spectral data. A hyperspectral imaging system was used to acquire spectral data of grain samples (such as soybeans) in the 400-1700 nm range. Then, a game theory-based interpretable AI method, particularly SHAP (Shapley Additive Explanations) value analysis, was introduced to perform in-depth analysis of the entire spectrum, dividing it into a moisture-sensitive region X. w Protein-sensitive region X p X, the protein water overlap region c Preferably, specifically:
[0104] Constructing a spectral feature library related to moisture content: Further, using sample moisture content as the prediction target, train an initial prediction model (e.g., XGBoost) and calculate the SHAP value contribution of each band to moisture prediction. This contribution is quantified as the absolute mean of the SHAP values. The top k bands with the highest contribution are selected. The initial number of spectral feature bands is generally chosen from a wide range to form a water feature database. (Excluding the portion overlapping with the protein).
[0105] Constructing a protein feature library: In parallel, using sample protein content as the prediction target, SHAP value analysis is used to screen out the most critical features for protein prediction. The study found that the initial selection of spectral feature bands was also relatively wide, thus forming the protein feature library. Excluding the portion overlapping with moisture. Generally speaking... and There is a significant overlap, and the band set formed by the overlapping bands is X. c .
[0106] Construct a model of the impact of moisture differences on spectral reflectance changes: Divide the moisture values into 5 gradients according to their degree of influence on spectral changes, and use C for each gradient. l Let (l=1, 2, 3, 4, 5) be the value. Preferably, this value can be discretely categorized based on prior knowledge, such as "high," "medium-high," "medium," "medium-low," and "low." This represents the difference in X values caused by varying moisture content in the same sample. p X c Spectral differences are denoted as: ΔX p and △Xc Establishment by fitting the spectral difference curve. δ l This is the basic compensation error corresponding to level l moisture.
[0107] Secondly, there is the establishment of a cascade prediction model.
[0108] A cascaded structure is adopted, using moisture prediction as an intermediate variable to assist in the accurate prediction of final protein content.
[0109] Secondary loop-assisted prediction model: using a moisture feature database And coincident spectral database X c Constructing a moisture prediction model X w ’ As input, train a machine learning model (Such as XGBoost, Support Vector Machine). To further improve model performance, an automatic parameter optimization algorithm is added, based on... By adjusting the secondary loop moisture spectral database, a more robust moisture prediction model can be obtained. This is achieved through establishing... The root mean square error (RMSE) of prediction is iteratively updated to the optimal subset of moisture feature bands, so that the RMSE trained based on this subset is minimized.
[0110] In the training of the main loop protein content prediction, in each iteration, based on the predicted water content, according to... , To adjust protein spectral data; and simultaneously according to , The goal of automatic parameter optimization is to find an optimal subset of protein features. This can maximize the accuracy of protein content prediction. The objective function for optimization is to maximize the main loop model. Coefficient of determination on the validation set Using the above strategy, based on the top 20 protein-sensitive bands selected by SHAP values, further optimization is performed to determine the final bands to be used for deployment. One core band.
[0111] Furthermore, the main loop model employs the XGBoost algorithm, whose prediction function is the sum of regression tree functions, which can effectively handle high-dimensional features and nonlinear relationships.
[0112] Example 2
[0113] This embodiment illustrates the specific application and testing of the modeling method provided in Embodiment 1 above, as shown below:
[0114] Soybean protein content prediction based on cascade decoupling model
[0115] I. System Setup and Data Acquisition
[0116] The implementation platform for this embodiment includes: a hyperspectral imaging system (ImSpector N17E) covering the spectral range of 400-1700 nm, a standard halogen tungsten lamp light source, an integrating sphere to ensure uniform illumination, and a computer equipped with data processing software. The experimental samples consisted of 95 soybean samples from different varieties and batches to ensure representativeness. All samples were equilibrated in a constant temperature and humidity environment (25 ± 1℃) for at least 24 hours before testing to minimize instantaneous moisture fluctuations. Each sample was scanned using the hyperspectral imaging system, and the raw data was converted to relative reflectance through black-and-white correction. Each sample was scanned three times, and the average spectrum was taken as the raw spectral data for that sample. Subsequently, the protein content reference value for each sample was determined using the Kjeldahl method (referring to AOAC 992.23 standard), and its moisture content was recorded.
[0117] II. Spectral Preprocessing
[0118] The acquired raw spectral data were first subjected to a standard normal variable (SNV) transformation to eliminate baseline drift caused by sample particle size and surface scattering. Subsequently, Savitzky-Golay filtering (SG filtering) was used for smoothing and denoising. In this embodiment, the preprocessing combination (SG+SNV) proved to be the most effective and was used for all subsequent analyses, as detailed below. Figure 2 As shown, the spectral reflectance curves of soybean samples with different moisture grades in the wavelength range of 900–1700 nm are displayed. The five curves of different colors in the figure represent five moisture discrete grades, namely low moisture sample (C1), medium-low moisture sample (C2), and medium-low moisture sample (C3). 2) Medium moisture sample (C) 3) The average spectra of medium-high moisture samples (C4) and high moisture samples (C5).
[0119] III. Construction of a water and protein feature library
[0120] Training the initial prediction model:
[0121] Using spectral data from all 225 bands as input, and measured water content and protein content as output targets, two independent regression models, denoted as Model_water and Model_protein, are preferably trained using Python's XGBoost library.
[0122] 2. SHAP value analysis and feature library construction:
[0123] 1. Use TreeExplainer from the SHAP library to interpret the two models above. Further, for each model and each sample, calculate the SHAP values for each of the 225 bands.
[0124] 2. Calculate the global importance of the band: Based on the relevant formulas above, calculate the absolute mean of the SHAP value for each band j.
[0125] 3. Screening key wavebands:
[0126] To avoid the "shortcut learning" (i.e., the model incorrectly relies on water-correlated bands to obtain surface accuracy) that occurs in traditional full-spectrum modeling methods, we introduce a game theory-based interpretable artificial intelligence model—SHAP (SHapley Additive exPlanations)—to perform importance analysis on spectral bands.
[0127] The theoretical basis of SHAP originates from the Shapley value in cooperative game theory. Its core idea is to treat each spectral band as a "player" and the model's prediction result as a "payoff." By calculating the marginal contribution of each band across all possible combinations, the importance of the band to the prediction result is measured. The calculation formula is as follows:
[0128]
[0129] in, For the set of all bands, Not including bands a subset of For the first The SHAP value of a band is the average marginal contribution of that band across all possible combinations.
[0130] Two prediction models were constructed using "moisture content" and "protein content" as target variables, respectively, and the absolute mean of the SHAP value for each band was calculated. The larger the value, the greater the contribution of that band to the model.
[0131] To suppress moisture interference, the SHAP value of moisture-related bands is used as a penalty term, gradually decaying during iteration. Simultaneously, the SHAP weight of protein-related bands is increased, allowing the model to gradually establish a feature distribution prioritizing physical interpretability. This method retains the interpretability advantage of SHAP while incorporating the physical absorption patterns of moisture and proteins (OH, amide I / II, etc.) in near-infrared spectroscopy, thus achieving a two-way fusion of data-driven and spectral physical priors. This method can automatically identify and gradually eliminate moisture-spectral interference bands during training, generating a "protein intrinsic spectral feature library," a deep interpretability advantage not found in traditional full-spectrum modeling and simple spectral difference correction methods.
[0132] For the moisture feature library, select The top 20 bands (i.e.) These wavelengths, as physically verified, are precisely concentrated in the known strong absorption regions of moisture, such as 970nm, 1200nm, and 1450nm, forming a subset of moisture characteristic bands. .
[0133] Accordingly, for the protein feature library, the top 20 bands (i.e., The results showed that after removing moisture interference, the key wavelength bands significantly shifted to the characteristic absorption region of protein amide bonds in the 1400-1700 nm range, forming a subset of protein characteristics. The overlapping wavelengths are 1430nm, 1480nm, 1490nm, 1545nm, and 1560nm, forming the overlapping spectral database X. c The specific process of dynamic SHAP feature selection is as follows: Figure 3 As shown, the characteristic importance and spectral comparison before and after removing the water-sensitive band are as follows: Figure 4As shown, the left subplot a represents the peak intensity distribution of the water band, illustrating the variation of peak reflectance intensity with wavelength in the characteristic band strongly correlated with water content (i.e., the "water band") within the wavelength range of 900 nm to 1700 nm. The water band intensity significantly increases to approximately 0.140 at 1000 nm and maintains a stable high-intensity plateau in subsequent longer wavelength bands (1300 nm to 1700 nm), which intuitively reflects the widespread and strong characteristic absorption of water in the near-infrared region. The right subplot b represents the peak intensity distribution of the non-water band after removing the water band, showing the peak reflectance (or absorbance) intensity distribution of the characteristic band strongly correlated with protein content (i.e., the "non-water band") within the same wavelength range. Its intensity value is significantly lower than that of the water band, especially at 1000 nm. At nm, the intensity of the water band is about twice that of the non-water band; this comparison clearly shows that, without decoupling, the spectral signal of water completely covers the weak feature signal of the protein in terms of intensity. This is the main physical reason for the model to "shortcut learning" and rely on water information to predict proteins.
[0134] IV. Establishing and training a cascade prediction model
[0135] 1. Primary Model (Moisture State Prediction):
[0136] Extract the corresponding moisture feature subset from the spectral data of each sample. Furthermore, using this as the input feature and the measured moisture content as the output label, an XGBoost regression model is trained. .
[0137] The continuous predicted values output by this model represent the moisture state. In this embodiment, to further enhance robustness, the following is used: Based on the sample distribution, the moisture content was divided into five discrete grades: low moisture (C1), low to medium moisture (C2), medium moisture (C3), medium to high moisture (C4), and high moisture (C5). The characteristic distribution map of soybean hyperspectral moisture bands is shown below. Figure 5 As shown, where:
[0138] Figure 5 As a whole, the logic is to first perform basic statistical description of the spectral data (AC subgraph), and then show the dynamic feature selection process based on these statistics (DF subgraph).
[0139] (a) The subplot represents the average spectral statistics, which serves to show the overall statistical characteristics of the full-band spectral data. This is the starting point for all subsequent analyses.
[0140] (b) The subplot represents the low-frequency band average spectral statistics, which is specifically designed to show the statistical characteristics of the low-frequency band (usually corresponding to regions with strong absorption and broad spectral characteristics such as moisture) and is used to initially identify areas that are significantly affected by moisture.
[0141] In the figure, Max / Min represents the extreme reflectance values of all samples in the low-frequency band. Due to the strong absorption of moisture, the Max value here is expected to be relatively low.
[0142] SD measures the difference in reflectance of samples in the low-frequency band. The key point is that because moisture absorption is strong and widespread, the reflectance of samples in this region may be very similar, therefore the SD value is expected to be small. This data confirms that moisture interference is a signal of "high background, low difference".
[0143] The relative fluctuation of CV within the low-frequency band. Although SD is small, if the average value (reflectivity) is also low, the CV value is not necessarily small. This indicator helps to more accurately determine the signal stability in this region.
[0144] (c) The subplot represents the mid-frequency band average spectral statistics, which is specifically designed to display the statistical characteristics of the mid-frequency band (which may contain more refined features of components such as proteins and starch). It is used to find differential regions related to the target component.
[0145] In the figure, Max / Min represents the extreme reflectance values of all samples in the mid-frequency band. The Max value may be higher than in the low-frequency band.
[0146] SD (Discrete Scattering) is a key comparative indicator. It is expected that the SD value in the mid-frequency band will be significantly higher than that in the low-frequency band. This is because differences in protein content between samples will be reflected in the spectrum of this region, leading to increased data dispersion. An increase in SD is a data indicator of the presence of effective discriminative information.
[0147] The relative fluctuations of CV in the mid-frequency band. Combined with increased SD, the CV pattern can reveal the ratio of information to noise in this region.
[0148] (d) The subplot compares the feature selection results based on statistical intervals. Its purpose is to dynamically demonstrate the selection process. It shows how the statistical indicators (such as Max and Min) of the retained bands change when different statistical thresholds (such as setting a range for SD) are set, intuitively illustrating how the selection criteria affect the quality of the feature library.
[0149] In this figure, Max and Min are no longer the statistical values of the original data, but rather the new statistical values corresponding to the band set retained after filtering. This visually demonstrates how the filtering threshold (e.g., "only retaining bands with SD > a certain value") affects the statistical properties of the final feature library. For example, how does the average Max / Min of the retained bands change as the threshold changes? This helps researchers determine whether the filtering criteria are reasonable (e.g., after removing bands with excessively small SDs, is the range of Max / Min values for the retained bands more effective in distinguishing samples?).
[0150] (e) The subgraph represents the feature selection stability analysis, which serves to demonstrate the robustness of the method. Through multiple simulations and iterations, it is shown that the key statistical indicators remain stable under different levels of selection rigor, indicating that the selection results are reliable and not random.
[0151] The Max, Min, and SD values shown in this figure represent the statistical values of the final feature library after multiple screening experiments. If these values remain stable across different experiments (with a flat change curve), it proves that the screening method is reliable and the results are not random. For example, a stable SD value means that the data dispersion of the internal bands in each selected feature library is consistent, which provides evidence for the reproducibility of the method.
[0152] (f) The subplot represents the relationship between the screening threshold and spectral features, serving to provide a basis for automatic optimization. It quantitatively displays the relationship between the screening threshold (horizontal axis) and the quality of the final feature library (vertical axis, such as a certain statistic), enabling the computer algorithm (automatic parameter optimization) to automatically find the optimal threshold.
[0153] Typically, the horizontal axis represents the screening threshold (such as the lower limit of SD), and the vertical axis represents a certain spectral feature statistic, which is likely an average statistic of the set of retained bands, such as the average reflectance of all retained bands (related to Max / Min) or the average CV value.
[0154] This graph clearly illustrates how the quality of the feature library systematically changes when the screening threshold is adjusted. It is precisely the input required by the "automatic parameter optimization algorithm": a well-defined, optimizable functional relationship. The algorithm can automatically search for the threshold point that optimizes the performance (e.g., R²) of the final predictive model.
[0155] 2. Principal Model (Protein Content Prediction):
[0156] Accordingly, a subset of protein features is extracted from the spectral data of each sample. .Will Moisture grading labels and repeating bands output by the primary model The vectors are concatenated to form a joint feature vector.
[0157] Using this joint feature vector as input and the measured protein content as the output label, the final XGBoost regression model is trained. The mathematical expression of this model is as described above.
[0158] Figure 6 The image shows a SHAP-based feature interpretation diagram for soybean protein prediction after moisture strip removal, where:
[0159] Subgraph a is the SHAP decision graph, which illustrates the prediction decision paths for multiple samples in the test set. This graph explains how the cascade prediction model (particularly the main loop model) of this invention predicts protein content. In the figure:
[0160] The horizontal axis represents the model's output value, i.e., the predicted protein content; the vertical axis lists the 12 spectral features (bands) that contribute most to the model's prediction, sorted from top to bottom in descending order of their average impact on the prediction result (SHAP value); the multiple colored curves in the figure each represent the prediction trajectory of a test sample. The curves start from the model's expected baseline value (the average of all sample predictions) at the left end, and extend to the right as the contribution value (SHAP value) of each feature is accumulated sequentially, ending at the final predicted value of that sample.
[0161] Subplot b is a SHAP waterfall plot, which focuses on a single representative test sample and breaks down its predicted values in a "waterfall" fashion, providing the most detailed explanation of feature contributions. In the plot, the bottommost E[f( X )] = 39.139 represents the expected value (baseline) of the model's predictions for all samples. Starting from the baseline, each row corresponds to the contribution value (SHAP value) of a feature band (such as Band_1088, Band_1300, etc.) to the predicted value of that specific sample. The red bars to the right represent positive contributions (pushing the predicted value up from the baseline), and the blue bars to the left represent negative contributions (pulling the predicted value down from the baseline). The length of the bars precisely represents the magnitude of the contribution.
[0162] V. Model Performance Verification and Results
[0163] To verify the effectiveness of the invention, the 95 samples were preferably randomly divided into a modeling set (76 samples) and a prediction set (19 samples) in an 8:2 ratio. Furthermore, all models were trained and evaluated using 5-fold cross-validation.
[0164] Figure 7 The illustration shows a comparison of the water-protein spectral decoupling effects in an embodiment of the present invention. This embodiment also includes a performance comparison: the performance of the proposed cascade decoupling model (XGBoost-Cascade) is compared with that of a single XGBoost model using full-band spectroscopy.
[0165] The results are shown in Table 1 below. The cascade decoupling model achieved the best performance on the prediction set.
[0166] Table 1. Predictive performance test results of different models
[0167]
[0168] Interpretability verification: For the principal-level model Further SHAP analysis revealed that the model's decision-making is highly concentrated in the protein characteristic band of 1500-1700 nm, while the contribution of the water-sensitive band is almost zero. This confirms from an interpretability perspective that the model has successfully transitioned from "water-driven" to "protein-driven," and the decision-making mechanism is highly consistent with the physical mechanism of near-infrared spectroscopy, significantly enhancing the model's credibility and reliability.
[0169] Example 3
[0170] This embodiment illustrates the application of the above modeling method in soybean component detection, as shown below:
[0171] In practical applications, the trained primary model Principal-level model The system also deploys the band index of the feature library into the online detection system. When a new, unknown soybean sample passes through the online hyperspectral camera, the system automatically executes the following process:
[0172] Collect spectral data and perform SG+SNV preprocessing.
[0173] Based on the pre-stored water and protein feature library indexes, extract the corresponding feature subsets. , and overlapping areas .
[0174] Will enter The model yields moisture classifications. .
[0175] Will enter The model outputs the predicted protein content of the sample in real time. .
[0176] Based on the above embodiments, it is clear that the specific implementation methods provided by the present invention have at least the following beneficial effects:
[0177] 1. Fundamentally decouple spectral interference: By using the interpretability tool SHAP value, data-driven screening is combined with prior knowledge of near-infrared spectroscopy physics to systematically identify and eliminate strong interference bands caused by water, breaking the dilemma of "shortcut learning" and enabling the model to focus on the intrinsic spectral characteristics of proteins.
[0178] 2. Significantly Improved Model Accuracy and Robustness: A cascaded modeling architecture is adopted, decomposing the complex spectral-composition mapping problem into two sub-problems: "moisture sensing" and "protein prediction." As shown in the example, this strategy increases the cross-validation determination coefficient of the XGBoost model from 0.8711 to over 0.9319, effectively suppressing overfitting and enhancing the model's generalization ability under different environmental conditions.
[0179] 3. Transparent decision-making process and clear physical meaning: The entire methodological process, from the construction of the feature library (transferring key bands to the protein amide bond absorption region) to the decision of the cascade model (using water state as a clear input), has clear physical interpretability, surpassing the traditional "black box" model and providing a solid basis for the reliability of the model.
[0180] Moreover, the embodiments of this invention provide a general methodology: the cascade decoupling modeling method for spectral information proposed in this invention is not only applicable to the decoupling of soybean protein and moisture, but can also provide a generalizable analytical paradigm for the non-destructive detection of other grains with strong interfering components (such as the interference of corn oil on starch detection, the interference of moisture in wheat on gluten content, and not limited to these).
[0181] It should be understood that the above embodiments are merely illustrative of the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for multi-group decoupling prediction modeling of grain based on cascade spectral features, characterized in that, include: A training database is provided, which includes spectral data of the target grain, calibration values of the first component content and calibration values of the second component content, wherein the combination of the first component and the second component includes any one of the following: a combination of moisture and protein, a combination of oil and starch, and a combination of moisture and gluten. The first band set that has a strong correlation with the content of the first component in the spectral data is extracted as the first component feature library, and the second band set that has a strong correlation with the content of the second component in the model prediction is extracted as the second component feature library. Using the first component feature library as input and the first component content calibration value as label, a machine learning method is used to train the secondary loop model and output the prediction result of the first component. Using the prediction results of the first component and the feature library of the second component as joint inputs, and the content calibration value of the second component as a label, the main loop model is trained using machine learning methods. The secondary loop model and the primary loop model are trained iteratively. The combination of the obtained secondary loop model, primary loop model, and the indices of the first band set and the second band set is used as the grain multi-group decoupling prediction model. During the iteration process, the first band set and the second band set are adjusted by an automatic parameter optimization algorithm so that the feature contribution of the bands sensitive to the first component gradually decreases during the iteration process, while the feature contribution of the bands sensitive to the second component gradually increases.
2. The grain multi-group decoupling prediction modeling method according to claim 1, characterized in that, The extraction process of the first component feature library and the second component feature library specifically includes: Calculate the band SHAP value of the contribution of each band in the spectral data to the content of the first component or the content of the second component; Each band is sorted based on its SHAP value, and the corresponding bands in the high SHAP value range are selected as the first band set or the second band set.
3. The grain multi-group decoupling prediction modeling method according to claim 2, characterized in that, The calculation method for the SHAP value of the band is expressed as follows: ; ; in, This represents the band SHAP value corresponding to band j; This indicates the number of grain samples in the training database; Indicates sample The SHAP value of a sample in band j; S represents a feature subset that does not include band j; Represents the set of all frequency bands; Indicates the use of feature subsets The predicted value of the time feature analysis model, which is independent of the secondary loop model and the primary loop model, where X belongs to S; This represents the set of feature subset X and band j.
4. The grain multi-group decoupling prediction modeling method according to claim 2, characterized in that, Specifically, it includes: The spectral data is divided into a single first component sensitive band set, a single second component sensitive band set, and a two-component overlapping sensitive band set according to the SHAP value of the band. The first band set includes the single first component sensitive band set, and the second band set includes the single second component sensitive band set and the two-component overlapping sensitive band set. Calculate the compensation variable for the difference between the single second component sensitive band set and the double component overlapping sensitive band set caused by the difference in the content of different first components; During iterative training, the corresponding band ranges of the single second component sensitive band set and the double component overlapping sensitive band set are adjusted according to the prediction results of the first component and the difference compensation variable, and the corresponding band range of the single first component sensitive band set is adjusted according to the changes of the double component overlapping sensitive band set.
5. The grain multi-group decoupling prediction modeling method according to claim 4, characterized in that, The prediction result of the first component is represented as multiple discrete levels, and the difference compensation variable is the inter-level compensation amount corresponding to the discrete level.
6. The grain multi-group decoupling prediction modeling method according to claim 5, characterized in that, The calculation method for the difference compensation variable is expressed as follows: ; ; in: Indicates the moisture level Below, the difference compensation variable for the sensitive band set of the second component; Indicates the moisture level Below, the concentration of sensitive bands with overlapping dual components is the first Difference compensation variables for each band; It is a function about wavelength obtained by fitting the spectral difference curve; Indicates moisture level The corresponding basic compensation error; Indicates the first The water-protein coupling coefficient of each overlapping band, with a value greater than 1, is used to characterize the intensity of interference from water changes in that overlapping band.
7. The grain multi-group decoupling prediction modeling method according to claim 6, characterized in that, The process of the automatic parameter optimization algorithm is expressed as follows: ; ; in, Indicates the first The single second component sensitive band set at the next iteration; Indicates the iteration round; Indicates the first The first iteration of the two-component overlapping sensitive band set Values for each band.
8. A multi-group decoupling prediction method for grain based on cascade spectral features, characterized in that, include: Provides the spectral data of the target grain and the grain multi-group decoupling prediction model obtained by the grain multi-group decoupling prediction modeling method according to any one of claims 1-7; The component features of the measured spectral data are extracted based on the band index; Input the features of the first component into the sub-loop model to obtain the predicted value of the first component; The second component's features and the first component's predicted value are input into the main loop model to obtain the second component's predicted value.
9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, and the computer program is executed by the processor to execute the grain multi-group decoupling prediction modeling method according to any one of claims 1-7, or to run the grain multi-group decoupling prediction model obtained by the grain multi-group decoupling prediction modeling method according to any one of claims 1-7.
10. A readable storage medium, characterized in that, The readable storage medium stores a computer program, which, when run, executes the grain multi-group decoupling prediction modeling method according to any one of claims 1-7; or, the readable storage medium stores a grain multi-group decoupling prediction model obtained by the grain multi-group decoupling prediction modeling method according to any one of claims 1-7.
Citation Information
Patent Citations
Near infrared spectrum quantitative detection method and system for correcting moisture interference
CN117191739A
Model construction method and device for estimating nutrient content of wolfberry based on hyperspectrum
CN116482037A
Method and equipment for simultaneous determination of multiple components in ammonia escape and storage medium
CN119198602A