Method for detecting quality of rapeseed based on characteristic wavelength optimization and multi-model fusion

By using a method of characteristic wavelength optimization and multi-model fusion, the problems of high hardware cost, data redundancy and poor adaptability of multiple indicators in rapeseed quality testing are solved, realizing high-precision, rapid and non-destructive testing of multiple indicators of rapeseed, which is suitable for large-scale testing in the field.

CN122084567APending Publication Date: 2026-05-26HUAZHONG AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUAZHONG AGRI UNIV
Filing Date
2025-12-16
Publication Date
2026-05-26

Smart Images

  • Figure CN122084567A_ABST
    Figure CN122084567A_ABST
Patent Text Reader

Abstract

This invention discloses a rapeseed quality detection method based on feature wavelength optimization and multi-model fusion. The method involves collecting full-band near-infrared diffuse reflectance spectral data of rapeseed samples and simultaneously measuring the true values ​​of physicochemical indicators to construct an original dataset. Based on the data distribution characteristics of each physicochemical indicator, the original dataset is divided into a training set and a test set. Preprocessing algorithms are selected for each physicochemical indicator to process the full-band near-infrared diffuse reflectance spectral data. Dimensionality reduction algorithms are selected for each physicochemical indicator to extract feature wavelengths from the preprocessed spectral data. Predictive models for each physicochemical indicator are established based on the training set, and the performance of the predictive models is verified using the test set to determine the optimal algorithm combination for each physicochemical indicator. The variable importance projection algorithm is used to calculate the comprehensive contribution score of each wavelength point in the full-band spectrum to the four physicochemical indicators. A multi-indicator comprehensive threshold is set, and the core light source wavelength of the portable device is determined from the high-score region.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of non-destructive testing technology for agricultural products, specifically to a method for detecting the quality of rapeseed based on characteristic wavelength optimization and multi-model fusion. Background Technology

[0002] As an important oilseed crop, the internal quality indicators (oil content, protein, glucosinolates, and moisture content) of rapeseed directly determine its economic value and processing difficulty. Currently, the quality detection of rapeseed mainly relies on traditional laboratory chemical analysis methods. However, although chemical detection methods are highly accurate, they have obvious limitations: (1) Destructive and time-consuming. These methods are all destructive, requiring the sample to be crushed, and each test takes a long time (usually 2-4 hours), which cannot meet the needs of real-time, large-scale field testing. (2) High cost and environmental pollution. The detection process requires a large amount of chemical reagents (such as petroleum ether, concentrated sulfuric acid, etc.), which is not only costly but also generates waste liquid and pollutes the environment. (3) Complex operation. The testing personnel have high professional skills requirements, making it difficult to popularize at the purchasing site or in the field.

[0003] To address the shortcomings of traditional methods, near-infrared spectroscopy (NIRS) has been widely applied to agricultural product testing due to its rapid and non-destructive characteristics. Existing commercial near-infrared spectrometers or portable devices are usually based on miniature spectrometers (such as NIR-M-R2, NeoSpectra, etc.) to collect continuous, full-band (e.g., 900-1700 nm) spectral data. In existing technologies, researchers typically use full-band spectral data, combined with characteristic wavelength screening algorithms such as CARS (competitive adaptive reweighted sampling), UVE (uninformation variable elimination), or SPA (continuous projection algorithm), to establish regression models such as partial least squares (PLS) to predict rapeseed quality. However, the following problems still exist: (1) High hardware cost. Although devices based on miniature spectrometers are more portable than desktop computers, the core spectroscopic components are still expensive, making it difficult to promote them on a large scale as low-cost devices. (2) Data redundancy. Full-band spectra contain hundreds or even thousands of wavelength variables, with high data dimensionality, a large amount of redundant information and noise, resulting in a large computational load for the model and high performance requirements for embedded processors.

[0004] To reduce costs and hardware complexity, existing technologies are beginning to develop towards multi-channel spectral detection based on LED arrays. This technology uses only a few specific wavelengths (such as LED light sources) to replace the full-band light source and spectroscopic system. Currently, there is research on portable LED detection devices for single indicators such as rice moisture and apple sugar content. However, although multi-channel spectral technology reduces hardware costs, the following technical bottlenecks still exist in the training and construction of models for the detection of multiple indicators (oil content, protein, glucosinolates, and moisture content) of rapeseed: (1) Lack of systematic modeling methods for low-variable data. Existing spectral processing procedures (such as CARS, UVE, and other feature selection algorithms) are designed for full-band (high-dimensional) data. When the data dimension is extremely low (e.g., only 10 channels / variables), traditional feature selection algorithms are no longer applicable. (2) Poor adaptability of multiple indicators. The spectral response characteristics of rapeseed oil content, protein, glucosinolates, and moisture content are different. Existing technologies often use a single "preprocessing + regression" mode to process all indicators, which cannot take into account the prediction accuracy of different physicochemical values. For example, for multi-channel data with only 10 variables, the maximum difference projection of the covariance matrix is ​​not significant, and conventional PLS regression is not as effective as Extreme Learning Machine (ELM) on some indicators. (3) Existing models are not accurate enough. Directly transferring the modeling method of the whole band to multi-channel data often leads to a significant decrease in prediction accuracy, especially for trace components such as glucosinolates, where there is a lack of effective specific preprocessing and model combination strategies.

[0005] In summary, the existing technology lacks an optimal model selection and training method specifically for multi-band (multi-channel) spectral data that can simultaneously meet the high-precision detection of multiple physicochemical indicators of rapeseed. Summary of the Invention

[0006] This invention proposes a rapeseed quality detection method based on feature wavelength optimization and multi-model fusion, in order to solve the technical problem of optimizing core feature wavelengths and overcoming the limitations of a single model strategy under conditions of very few variable inputs.

[0007] To address the aforementioned technical problems, this invention provides a method for detecting rapeseed quality based on characteristic wavelength optimization and multi-model fusion, characterized by comprising the following steps: Step S1: Collect full-band near-infrared diffuse reflectance spectral data of rapeseed samples, and simultaneously determine the true values ​​of four physicochemical indicators of the rapeseed samples: oil content, protein, glucosinolates and moisture content, to construct the original dataset; Step S2: Based on the data distribution characteristics of each physicochemical index, select either a random partitioning method or a spatially balanced random partitioning method to divide the original dataset into a training set and a test set; Step S3: Select a preprocessing algorithm for each physicochemical index to process the full-band near-infrared diffuse reflectance spectral data; Step S4: Select a dimensionality reduction algorithm for each physicochemical index to extract feature wavelengths from the preprocessed spectral data; Step S5: Establish prediction models for each physicochemical index based on the training set, verify the performance of the prediction models using the test set, and determine the optimal algorithm combination for each physicochemical index. Step S6: Calculate the comprehensive contribution score of each wavelength point in the full-band spectrum to the four physicochemical indicators using the variable importance projection algorithm, set a multi-indicator comprehensive threshold, and determine the core light source wavelength of the portable device from the high-score region.

[0008] Preferably, in step S1, the spectral range of the full-band near-infrared diffuse reflectance spectral data is 900 nm to 1700 nm, the oil content is determined by Soxhlet extraction, the protein content is determined by Kjeldahl nitrogen determination, the glucosinolate content is determined by palladium chloride colorimetric method, and the water content is determined by direct drying method.

[0009] Preferably, in step S2, a spatial equilibrium random partitioning method is used for oil content, and a random partitioning method is used for protein, glucosinolates and water content. The spatial equilibrium random partitioning method calculates the Euclidean distance between samples based on the variable space of spectral and physicochemical values, and uses an iterative approach to select the sample with the farthest distance to add to the training set.

[0010] Preferably, in step S3, the preprocessing algorithm includes convolutional smoothing, multivariate scattering correction, and standard normal variable transformation; standard normal variable transformation is selected for oil content, glucosinolates, and water content, and multivariate scattering correction is selected for proteins.

[0011] Preferably, the standard normal variable transformation independently performs mean centering and standardization on the spectrum of each sample, and the calculation formula is as follows: ; in, For the transformed th The first sample Data points, For the original first The first sample Data points, For the first The average of all data points in a sample For the first The standard deviation of all data points in a sample.

[0012] Preferably, the dimensionality reduction algorithm includes a competitive adaptive reweighted sampling method, a non-informative variable elimination method, and a continuous projection algorithm; Competitive adaptive reweighted sampling was selected for oil content, glucosinolates and water content, and non-informative variable elimination was selected for proteins.

[0013] Preferably, the formula for calculating variable weights in the competitive adaptive reweighted sampling method is as follows: ; in, For the first The absolute values ​​of the regression coefficients of each variable, For the first The weights of each variable, This represents the number of variables remaining in each sampling.

[0014] Preferably, in step S5, the prediction model includes a partial least squares regression model and an extreme learning machine model; the coefficient of determination R is used. 2 The root mean square error (RMSE) and mean absolute error (MAE) of the prediction model are used as evaluation metrics.

[0015] Preferably, in step S6, the calculation expression for the variable importance projection value algorithm is: ; In the formula, For the first The importance projection values ​​of each feature; For the dimensions of the projection; The sum of the main components; For the first The weights of each principal component or the proportion of variance explained; For the first The variable in the first... The coefficients in each principal component; For the first The Euclidean norm of the loading vectors of the principal components.

[0016] Preferably, the core light source wavelength includes 10 discrete wavelengths, namely: 850nm, 930nm, 975nm, 1000nm, 1090nm, 1210nm, 1380nm, 1450nm, 1470nm and 1550nm, which correspond to the vibrational absorption peak positions of the CH, NH and OH functional groups inside rapeseed.

[0017] The beneficial effects of the present invention include at least the following: This invention fully considers the differences in spectral response characteristics of four physicochemical indicators: oil content, protein, glucosinolates, and moisture content. It establishes differentiated data processing and modeling workflows for each indicator, overcoming the problem of insufficient prediction accuracy for some indicators caused by the uniform modeling strategy used in existing technologies. Compared with traditional chemical analysis methods, this invention achieves rapid and non-destructive testing of rapeseed quality, eliminating the need for sample crushing and the consumption of chemical reagents such as petroleum ether and concentrated sulfuric acid. The testing process is simple and fast, avoiding environmental pollution from waste liquids, and can meet the needs of large-scale real-time testing in the field and at the purchasing stage. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of the method flow according to an embodiment of the present invention; Figure 2 This is a schematic diagram comparing the prediction accuracy of the device in an embodiment of the present invention; Figure 3 This is a schematic diagram of the external verification results of the device in an embodiment of the present invention; Figure 4 This is a schematic diagram of the Pearson analysis results in an embodiment of the present invention; Figure 5 This is a schematic diagram of the competitive adaptive reweighting analysis results according to an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating the prediction trends of different principal components in the Uve of this invention. Figure 7 This is a schematic diagram of the analysis results for eliminating non-information variables in an embodiment of the present invention; Figure 8 This is a schematic diagram of the SPA analysis results in an embodiment of the present invention; Figure 9 This is a schematic diagram of the near-infrared spectral data prediction results of a commercial instrument according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the variable importance projection analysis results in an embodiment of the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0020] This invention provides a method for rapeseed quality detection based on characteristic wavelength optimization and multi-model fusion, comprising the following steps: Step S1: Collect full-band near-infrared diffuse reflectance spectral data of rapeseed samples, and simultaneously determine the true values ​​of four physicochemical indicators of rapeseed samples: oil content, protein, glucosinolates and moisture content, to construct the original dataset.

[0021] Specifically, this invention selected 41 rapeseed varieties from major rapeseed producing areas and 24 rapeseed varieties harvested from various regions, totaling 65 samples, as experimental materials. To ensure the stability and repeatability of spectral acquisition, the experimental environment was controlled within the range of 50% to 60% relative humidity and 25 to 28 degrees Celsius.

[0022] Spectral data acquisition was performed using a standard benchtop near-infrared spectrometer, covering the full near-infrared region from 900 nm to 1700 nm. Darkroom calibration was performed before each spectral acquisition to eliminate the effects of instrument drift. Specific acquisition parameters were set as follows: 6 unsmoothed averages, 32x gain, and 0.635 ms exposure time. For each rapeseed variety, samples were divided into three categories, with 5 acquisitions per category followed by averaging, ultimately yielding 195 near-infrared diffuse reflectance spectra.

[0023] Simultaneously with the spectral acquisition, chemical analysis of four physicochemical indicators of the rapeseed samples was conducted. Oil content was determined using Soxhlet extraction, protein content using the Kjeldahl method, glucosinolate content using the palladium chloride colorimetric method, and moisture content using the direct drying method.

[0024] Following the aforementioned data collection process, a raw dataset was constructed, comprising 65 varieties, 195 full-band near-infrared spectra, and corresponding true values ​​for four physicochemical indicators. The distribution range of each physicochemical indicator in this dataset is as follows: oil content approximately 30% to 50%, protein approximately 18% to 28%, glucosinolates approximately 8 to 138 micromoles per gram, and water content approximately 4.8% to 10.3%.

[0025] Step S2: Based on the data distribution characteristics of each physicochemical index, select either a random partitioning method or a spatially balanced random partitioning method to divide the original dataset into a training set and a test set.

[0026] In the process of establishing a near-infrared spectral correction model, the reasonable partitioning of the training and test sets has a significant impact on the model's generalization ability. This invention, based on the data distribution characteristics of different physicochemical indicators, optimizes both random partitioning and spatially balanced random partitioning methods to determine the most suitable dataset partitioning strategy for modeling each indicator.

[0027] Random partitioning is the most common method for splitting datasets. Its characteristic is that it does not consider the range of sample content, data distribution, or abundance, but randomly assigns samples to the training and test sets. This method is simple and direct, and suitable for situations where the data distribution is relatively uniform. Experimental results have shown that for the three indicators of protein, glucosinolates, and water content, due to their relatively uniform data distribution or the existence of specific biologically significant intervals, random partitioning can achieve good modeling results.

[0028] The Spatial Balanced Randomization (SPXY) method is a partitioning strategy that considers the spatial distribution of data, making it particularly suitable for situations with large data variability and uneven distribution. Unlike the conventional SPXY method, this invention takes into account the clear biological thresholds of rapeseed physicochemical values ​​(e.g., a safe storage limit for moisture content). Therefore, it preserves the original data scale to avoid distortion of threshold information introduced by standardization, calculating the Euclidean distance between samples solely based on the variable space constituted by the spectrum and physicochemical values. The mathematical expression of this method is: ; In the formula, and These represent the attribute vectors of the two samples, The number of attribute variables, For the sample and The SPXY method uses an iterative approach to select samples. In each iteration, the sample furthest from all samples in the current training set is added to the training set until a predetermined training set size is reached. The remaining samples are used as the test set. This method ensures that the training set covers the main directions of change in the sample space, improving the model's predictive ability for boundary samples.

[0029] Taking into account the data characteristics of various physicochemical indicators, the dataset partitioning strategy determined in this invention is as follows: For the oil content indicator, due to its large data variation and wide distribution range, the SPXY partitioning method is adopted to improve the generalization ability of the model; for the protein, glucosinolate, and water content indicators, a random partitioning method is adopted. This differentiated partitioning strategy can better adapt to the modeling needs of different physicochemical indicators.

[0030] Step S3: Select a preprocessing algorithm for each physicochemical index to process the full-band near-infrared diffuse reflectance spectral data.

[0031] Preprocessing of near-infrared spectra is a crucial step in establishing high-quality calibration models. According to Beer-Lambert's law, the absorbance A in near-infrared spectra is proportional to the optical path length l, sample concentration c, and absorption coefficient ε, i.e., A = εlc. Theoretically, absorbance should uniquely correspond to concentration, but in actual measurements, it is subject to interference from various factors. Uneven particle size distribution in rapeseed samples can cause Mie scattering, altering the effective optical path and leading to baseline tilt and nonlinear response. Differences in sample packing density or agglomeration can create local refractive index abrupt changes, altering apparent absorbance. Moisture absorption may change the sample's dielectric constant, further amplifying scattering noise. These interferences manifest in the spectrum as baseline drift and high-frequency noise, and in severe cases, may mask the characteristic absorption peaks of the target components. Therefore, appropriate preprocessing methods are needed to suppress scattering effects, achieve noise reduction and smoothing, and enhance linear response.

[0032] Therefore, this invention employs three methods—convolutional smoothing filtering, multivariate scattering correction, and standard normal variable transformation—to preprocess spectral data, and selects the optimal preprocessing method for different physicochemical indices.

[0033] (a) Convolutional Smoothing Filter (SG) Convolutional smoothing filtering is a convolutional smoothing technique based on local polynomial fitting. It uses the least squares method to perform polynomial fitting on the data within a window to smooth noise and preserve spectral peak shape characteristics. This algorithm can significantly maintain peak width and area while effectively removing random noise. For data points with a window size of N, its polynomial fitting formula is: ; In the formula, For the smoothed data points, The raw data within the window. For polynomial coefficients, It is half the size of the window.

[0034] It is important to note that this algorithm suffers from boundary effects, which can distort the first and last k data points. In practical applications, this needs to be addressed through mirror continuation or truncation. Since convolutional smoothing is relatively weak at correcting baseline drift and scattering effects, it performs worse than multivariate scattering correction and standard normal transformation when processing rapeseed spectral data.

[0035] (ii) Multivariate Scattering Correction (MSC) Multivariate scattering correction corrects for multiplicative scattering effects caused by uneven particle distribution by assuming a linear relationship between the scattering effects of all samples and the reference spectrum. This method establishes a scattering model based on a set of reference samples (typically the average of all sample spectra) and effectively eliminates multiplicative scattering interference caused by differences in particle size by forcing all spectra to align with the scattering characteristics of the reference spectrum. Its mathematical expression is:

[0036] In the formula, The corrected spectrum The original spectrum, For the regression intercept, The regression slope, This is the error term.

[0037] Multiple scattering correction is particularly suitable for predicting protein content because the stretching and bending vibrations of CH, NH and OH bonds in the amino acid side chains and peptide bonds of proteins absorb near-infrared waves from 950 to 1200 nm, and the characteristic vibrations of C=O and NH in the amide I and amide II bands absorb near-infrared waves from 1300 to 1600 nm. The characteristic peaks at both ends of the protein are relatively independent, and multiple scattering correction can preserve these characteristic information well.

[0038] (III) Standard Normal Variable Transformation (SNV) The Standard Normal Transform (SMT) independently centers and standardizes the spectrum of each sample, eliminating additive scattering and baseline drift caused by optical path fluctuations or particle scattering. Unlike multivariate scattering correction, the SMT does not require a reference spectrum and processes each sample individually. Its calculation formula is as follows: ; in, For the transformed th The first sample Data points, For the original first The first sample Data points, For the first The average of all data points in a sample For the first The standard deviation of all data points in a sample.

[0039] Through systematic experimental verification and comparative analysis, the preprocessing strategy determined in this invention is as follows: for the three indicators of oil content, glucosinolates, and water content, the standard normal variable transformation (SNV) preprocessing method is preferred; for the protein indicator, the multivariate scattering correction (MSC) preprocessing method is preferred. This differentiated preprocessing strategy can retain the characteristic spectral information corresponding to each physicochemical indicator to the greatest extent, laying the foundation for subsequent characteristic wavelength selection and model building.

[0040] Step S4: Select a dimensionality reduction algorithm for each physicochemical index to extract feature wavelengths from the preprocessed spectral data.

[0041] In order to extract characteristic wavelengths that are highly correlated with the physicochemical properties of rapeseed from full-band spectral data, reduce data dimensionality and improve modeling efficiency, this invention first evaluates the feasibility of single-wavelength prediction through Pearson correlation coefficient analysis, and then adopts three characteristic wavelength screening methods: competitive adaptive reweighted sampling, non-information variable elimination, and continuous projection algorithm, and selects the best dimensionality reduction algorithm for different physicochemical properties.

[0042] (I) Single-wavelength correlation analysis Pearson correlation coefficients were used to quantify the linear correlation strength between single-wavelength absorbance in near-infrared spectra and target physicochemical values. Analysis showed that glucosinolates and water content exhibited the highest correlations (0.556 and 0.513) with absorbance at 1086 nm and 1091 nm, respectively; oil content and protein content showed the highest correlations (-0.434 and -0.449) with absorbance at 974 nm and 1447 nm, respectively. Due to the low single-wavelength prediction accuracy of each physicochemical value, direct single-wavelength modeling is not feasible; further multi-wavelength feature selection algorithms are needed to determine the characteristic wavelength combinations for each physicochemical value.

[0043] (ii) Competitive Adaptive Reweighted Sampling (CARS) Competitive adaptive reweighted sampling is a feature variable selection method that combines Monte Carlo sampling with partial least squares (PLS) model regression coefficients. It dynamically filters wavelength variables that significantly contribute to the model. This invention sets up 100 Monte Carlo sampling iterations and 5-fold cross-validation. Its core steps include: first, constructing an initial training set by drawing 80% of the samples with replacement from all samples; establishing a PLS regression model and calculating the regression coefficients for each wavelength variable; second, exponentially decaying wavelength elimination: in each iteration, low-importance wavelengths are gradually eliminated according to an exponential decay function; and third, adaptive reweighted sampling: the importance of spectral wavelengths is assessed based on the absolute value of the PLS regression coefficients, and feature selection is achieved through a weight allocation mechanism. High-weight wavelengths represent those that significantly contribute to the prediction of physicochemical values ​​and are therefore prioritized for retention. Low-weight wavelengths may be noise or collinear variables and are eliminated based on probability. A roulette wheel selection mechanism is used to select a subset of wavelengths based on their weighted probabilities, ensuring that important wavelengths are retained with a high probability while introducing randomness to avoid getting trapped in local optima. Iterative iteration and optimal subset selection: The above steps are repeated, and after each iteration, the partial least squares model constructed from the current wavelength subset is cross-validated, and its root mean square error (RMSECV) is calculated. Through multiple rounds of optimization and comparison, the wavelength combination that minimizes the global cross-validation error is selected as the optimal feature variable set. The formula for calculating variable weights is: ; in, For the first The absolute values ​​of the regression coefficients of each variable, For the first The weights of each variable, This represents the number of variables remaining in each sampling.

[0044] After competitive adaptive reweighted sampling, the number of characteristic wavelengths for oil content, protein, glucosinolates and water content were 56, 73, 24 and 41, respectively.

[0045] (iii) Uninformation Variable Elimination (UVE) The non-information variable elimination method is a characteristic wavelength selection method based on statistical tests. It introduces artificial noise variables to construct a noise-signal comparison model, screening wavelength variables that have significant explanatory power for target physicochemical values. Its core process includes: first, artificially generating a noise matrix, mainly containing Gaussian white noise and uniform noise, with the noise intensity set at 15% of the original spectrum's standard deviation; then combining the spectral matrix and the noise matrix to form an extended matrix; next, constructing a multiple linear regression model based on the extended matrix and the target physicochemical value vector, extracting the regression coefficients of the noise component, and calculating the maximum absolute value; if the absolute value of the regression coefficient for a certain wavelength in the original spectrum is less than this maximum value, it is determined to be invalid information and eliminated. This algorithm has advantages such as resistance to overfitting, no need to preset the number of wavelengths, and strong robustness, making it particularly suitable for predicting protein content.

[0046] (iv) Continuous Projection Algorithm (SPA) The continuous projection algorithm is a data dimensionality reduction method that eliminates multicollinearity among wavelengths and selects the most informative and uncorrelated characteristic wavelengths through vector orthogonal projection. Essentially, it is an iterative application of Gram-Schmidt orthogonalization, maximizing the information content of the subspace by progressively constructing orthogonal wavelength subspaces. However, due to the correlation characteristics of various physicochemical values ​​with the near-infrared spectrum in this invention, the continuous projection algorithm has limited improvement on the model performance.

[0047] Taking into account the effects of various dimensionality reduction algorithms, the feature wavelength screening strategy determined in this invention is as follows: for the three indicators of oil content, glucosinolates and water content, the competitive adaptive reweighted sampling method (CARS) is preferred; for the protein indicator, the uninformative variable elimination method (UVE) is preferred.

[0048] Step S5: Establish prediction models for each physicochemical index based on the training set, verify the performance of the prediction models using the test set, and determine the optimal algorithm combination for each physicochemical index.

[0049] In near-infrared spectroscopy analysis, establishing a quantitative calibration model between spectral data and target physicochemical values ​​is the core step. The main steps are: Model training: using sample spectral data with known physicochemical values, a prediction model is established through a regression algorithm; Model validation: the model's generalization ability is evaluated through cross-validation, and hyperparameters are optimized; Prediction of unknown samples: the spectrum of the sample to be tested is collected and input into the model to obtain predicted values.

[0050] This invention uses two algorithms, partial least squares regression and extreme learning machine, to build a regression model, taking into account the modeling needs of both linear and nonlinear relationships.

[0051] (a) Partial Least Squares Regression (PLS) Partial Least Squares Regression (PLS) is a bilinear factor model that extracts latent variables by simultaneously decomposing the spectral matrix and the physicochemical matrix, and establishes a linear regression relationship by maximizing the covariance between the two. The main steps of PLS ​​are as follows: standardize the predictor variable data (X) and the response variable data (Y), calculate the first PLS coefficient, i.e., find a linear combination of the predictor variable matrix that maximizes the covariance of the predictor and response variables, subtract the data corresponding to the portion explained by the PLS coefficient from X and Y, and repeat the calculation of the PLS coefficient until the termination condition is met (in this invention, the termination condition is the variance explanatory weight and the principal component count).

[0052] The termination condition of this invention is set as the proportion of variance explained and the number of principal components. Partial least squares regression is particularly suitable for situations with a large number of variables and can effectively handle multicollinearity problems in spectral data.

[0053] (ii) Extreme Learning Machine (ELM) Extreme Learning Machines (ELMs) are highly efficient single-hidden-layer feedforward neural network architectures. Their core feature is the transformation of the original nonlinear optimization problem into a system of linear equations by randomly generating and fixing the parameters of the hidden layer nodes, thus significantly improving training efficiency. ELMs are particularly suitable for high-dimensional data regression and classification tasks. The ELM model structure includes an input layer, hidden layers, and an output layer. The input layer receives spectral data, and the hidden layers contain multiple neurons, the number of which is typically chosen based on the task complexity. Commonly used activation functions include Sigmoid, ReLU, or RBF. The output layer predicts physicochemical values ​​by linearly combining the outputs of the hidden layers.

[0054] Extreme learning machines can better avoid local minima and overfitting problems, and usually outperform partial least squares regression in cases with few variables.

[0055] (III) Model Evaluation Indicators The evaluation index for the calibration model is the coefficient of determination (R²). 2 The parameters are: root mean square error of prediction (RMSE) and mean absolute error (MAE). The coefficient of determination is used to characterize the degree of variation of the explanatory variables in the model, R0. 2 The closer the value is to 1, the stronger the linear correlation between the predicted and measured values. The root mean square error (RMSE) reflects the degree of dispersion between the predicted and actual values. The smaller the RMSE, the higher the model's prediction accuracy. The mean absolute error (MAE) represents the absolute average of the prediction error. The smaller the MAE value, the more ideal the prediction effect.

[0056] (iv) Optimal combination of models for the entire spectrum Based on partial least squares regression, a prediction model is constructed. Combining the above optimization steps, the optimal technical path and model performance for each physicochemical index are shown in Table 1. Table 1. Processing procedure and results of the full-band near-infrared spectroscopy prediction model

[0057] As shown in Table 1, the oil content model, employing a combination of SPXY partitioning, SNV preprocessing, CARS feature selection, and PLS regression, achieved a coefficient of determination of 0.92 and a root mean square error of only 1.13%, demonstrating excellent predictive performance. The water content model also performed well, with a coefficient of determination of 0.94 and a root mean square error of 0.32%. The coefficients of determination for the protein and glucosinolate models were 0.85 and 0.79, respectively, both meeting the accuracy requirements for practical detection.

[0058] Step S6: Calculate the comprehensive contribution score of each wavelength point in the full-band spectrum to the four physicochemical indicators using the variable importance projection algorithm, set a comprehensive threshold for multiple indicators, and determine the core light source wavelength of the portable device from the high-score region.

[0059] To achieve low cost and portability of the detection device, the initially selected dozens of characteristic wavelengths need to be further compressed into a very small number of discrete wavelengths that can be integrated into the hardware. This invention utilizes a variable importance projection algorithm to calculate the contribution score of each wavelength point in the full-band spectrum to the four physicochemical indicators in the partial least squares model. The expression for calculating the variable importance projection value is as follows: ; In the formula, For the first The importance projection values ​​of each feature; For the dimensions of the projection; The sum of the main components; For the first The weights of each principal component or the proportion of variance explained; For the first The variable in the first... The coefficients in each principal component; For the first The Euclidean norm of the loading vectors of the principal components.

[0060] This invention establishes a multi-index comprehensive threshold strategy: the sum of the variable importance projection values ​​of any two physicochemical indicators is greater than 2, or the sum of the variable importance projection values ​​of four physicochemical indicators is greater than 3.6. Based on this threshold strategy, high-weight peak regions are identified from the full-band spectrum, initially screening out nine bands: 929, 974, 996, 1088, 1207, 1382, 1447, 1467, and 1550 nm. Considering the literature reports a strong correlation between 850 nm and water content, ten core characteristic wavelengths were finally determined: 850 nm, 930 nm, 975 nm, 1000 nm, 1090 nm, 1210 nm, 1380 nm, 1450 nm, 1470 nm, and 1550 nm.

[0061] These 10 core wavelengths have a clear correspondence with the vibrational absorption peaks of key functional groups within rapeseed: around 850 nm corresponds to the third-order harmonic absorption of the CH bond; the 930-975 nm region corresponds to the second-order harmonic absorption of the OH and NH bonds; the 1000-1090 nm region corresponds to the second-order harmonic absorption of the CH bond; around 1210 nm corresponds to the combination-frequency absorption of the CH bond; the 1380-1470 nm region corresponds to the first-order harmonic absorption of the OH and CH bonds; and around 1550 nm corresponds to the first-order harmonic absorption of the NH bond. Based on cost considerations and the 20 nm LED bandwidth characteristics, the center wavelength of some LEDs is finely adjusted by 3 to 5 nm to facilitate industrial production.

[0062] Based on the 10 core wavelengths identified above, this invention also provides a portable multi-channel spectral detection device. Using this developed portable multi-channel spectral detection device, multi-channel spectra of 61 varieties and 181 groups of rapeseed were collected to develop a physicochemical value evaluation model for the device. Multi-channel spectra of 18 randomly selected rapeseed varieties were collected to evaluate the device's effectiveness. Unlike the near-infrared spectroscopy process, the number of variables in the multi-channel spectra is only 10, making variable selection and conventional preprocessing methods inapplicable. The physicochemical value model construction process uses first derivative (1st), second derivative (2nd), maximum and minimum scaling (MMS), SNV, and MSC) preprocessing methods and PCA and None dimensionality reduction methods. Dataset partitioning (Random, SPXY) and regression algorithms (PLS, ELM) remain unchanged. In the preprocessing, SG (Smoothing Mode) can cause the window to cover too many variables when used for a small number of variable points, resulting in over-smoothing. Therefore, first and second derivatives and maximum and minimum scaling are used instead. In variable selection, CARS, LARS, and UVE are not suitable for cases with only a few variables. PCA is an analysis based on all original variables and can be used in multi-channel spectral detection. Therefore, no processing and PCA dimensionality reduction are used instead of CARS, LARS, and UVE.

[0063] The test set results of the multi-channel spectral model are shown in Table 2. The data segmentation method for each physicochemical value is the same as that for near-infrared spectroscopy. However, in the preprocessing, the four physicochemical values, MMS, 2nd, SNV, and 2nd, showed the best predictive performance. Since there are only 10 sets of variables, the maximum difference projection of the covariance matrix is ​​not significant. Only protein showed significant improvement in model performance after dimensionality reduction using PCA. Among the regression algorithms, ELM generally outperforms PLS regression in cases with fewer variables because it better avoids local minima and overfitting. Due to the relatively uniform water content distribution and the absence of local minima and local special spectra, the ELM model performed slightly worse than the PLS model.

[0064] Table 2. Multichannel spectral processing process and results

[0065] In summary, after Random+MMS+ELM modeling, the RMSE, R², and MAE for oil content prediction were 2.06%, 0.78%, and 1.67%, respectively; after SPXY+2nd+PCA+ELM treatment, the RMSE, R², and MAE for protein prediction were 1.53%, 0.72%, and 1.38%, respectively; after SPXY+SNV+ELM treatment, the RMSE, R², and MAE for glucosinolate prediction were 16.23 μmol / g, 0.66, and 13.88 μmol / g, respectively; and after SPXY+2nd+PLS treatment, the RMSE, R², and MAE for water content prediction were 0.37%, 0.77%, and 0.29%, respectively. The device's prediction accuracy for each physicochemical value meets the detection requirements. A comparison of the prediction accuracy of commercial instruments with that of this device is shown below. Figure 2 As shown.

[0066] After validating the model performance within the device, it is necessary to further determine the device's usability. External validation results were obtained by randomly selecting 18 rapeseed varieties and collecting multi-channel spectral datasets. Figure 3 As shown. The external validation RMSE, R², and MAE results for oil content, protein, glucosinolates, and water content were 2.04%, 0.69%, 1.58%; 1.52%, 0.67%, 1.25%; 18.86 μmol / g, 0.52%, 15.03 μmol / g; and 0.36%, 0.74%, 0.33%, respectively. There were slight normal losses in oil content, protein content, and water content. The accuracy of glucosinolates was significantly reduced. Data set verification revealed that the glucosinolate content in the external validation set was low, and high physicochemical values ​​existed in both the training and testing sets, leading to a positive bias in the external validation. The predicted values ​​were slightly higher than the actual values, and the R² was lower. Considering that glucosinolates are trace elements, an RMSE of 18.86 μmol / g is still sufficient to determine the content range. Overall, the portable device, while ensuring detection speed and reducing costs, achieved the actual needs of rapid field detection for all physicochemical indicators.

[0067] Example 1 This embodiment illustrates the method of the present invention through specific calculation examples. Using 41 rapeseed varieties from major rapeseed producing areas and 24 rapeseed varieties harvested from various regions (a total of 65 samples) as experimental materials, the preset data collection conditions were 50-60% relative humidity and 25-28℃ ambient temperature. The device was used to collect the spectra of the experimental materials. Before each near-infrared spectral acquisition, darkroom calibration was performed with parameters of 6 unsmoothed averages, a gain of 32x, and an exposure time of 0.635ms. For multi-channel spectroscopy, darkroom calibration was performed every 30 minutes with parameters of 5 unsmoothed averages, a gain of 4x, and an exposure time of 350ms. Spectral information from 65 varieties was collected, with each variety divided into 3 categories. Five acquisitions were performed for each category, and the average was taken, resulting in a total of 195 spectra (Set-I represents near-infrared spectra, and Set-II represents the multi-channel spectra of the invented device). In addition, the same method was used to collect multi-channel spectral information and physicochemical values ​​of 30 sets of rapeseed (Set-III) as an external validation set.

[0068] Depend on Figure 4 It was found that glucosinolates and water content showed the highest correlations (0.556 and 0.513) with absorbance at 1086 nm and 1091 nm, respectively; while oil content and protein showed the highest correlations (-0.434 and -0.449) with absorbance at 974 nm and 1447 nm, respectively. The single-wavelength prediction accuracy for each physicochemical value was low, making direct extraction of single-wavelength models impossible; further analysis is needed to determine the characteristic wavelengths of the physicochemical values.

[0069] 1. Optimal selection of characteristic wavelengths Since single-wavelength analysis has low accuracy, further optimization algorithms are needed to determine the characteristic wavelengths of various physicochemical values ​​of rapeseed. Therefore, competitive adaptive reweighting, elimination of non-information variables, and continuous projection algorithms are used to process the near-infrared spectrum of rapeseed, establish a regression model, and determine the optimal wavelengths.

[0070] 1.1 Car Analysis Cars dimensionality reduction is a feature variable selection method combining Monte Carlo sampling and PLS model regression coefficients. This invention sets 100 Monte Carlo samplings and a 5-fold cross-validation method. The wavelength iteration direction is set to forced iteration with fixed elimination. Wavelength iteration changes are correlated with the root mean square error fluctuations of oil content and protein content, as shown below. Figure 5 As shown in (a) of the diagram. In the initial iterations (0-30 rounds) of the Cars algorithm, unimportant wavelengths were removed, initially optimizing the model, and the RMSE value fluctuated and decreased. In later iterations (40-50 rounds), more wavelengths were removed, causing the model to lose some key information, and the RMSE value fluctuated and increased. The error trends of glucosinolates and water content are shown in the diagram. Figure 5As shown in (b), the Cars analysis only reduced the wavelength and had little impact on RMSE. The RMSEs of physicochemical values ​​such as oil content and protein were 1.58%, 1.35%, 17.913 μmol / g, and 4.227%, respectively, and the number of wavelengths after screening were 56, 73, 24, and 41, respectively.

[0071] 1.2 Uve Analysis Uve uses statistical information from irrelevant variables related to noise to select characteristic variables of the spectrum itself. It's a characteristic variable selection method that uses the statistical distribution of regression coefficients of the target matrix based on the noisy independent variable matrix to determine the variable's characteristics. Before adding noise, the optimal principal component of the decision kernel must be determined. Figure 6 The results of principal component analysis (PCA) from 1 to 10 are presented, using the RMSE of physicochemical values ​​as an indicator. The prediction errors for oil content, protein content, and water content all initially decreased and then increased with increasing PCA counts, with optimal PCA counts of 7, 7, 7, and 8 for oil content, protein content, and water content, respectively. The actual physicochemical values ​​of glucosinolates fluctuated significantly and did not respond positively in partial least squares regression, reaching their optimal value with a PCA count of 2.

[0072] Noise combination matrix such as Figure 7 As shown in (a), orange represents the Gaussian noise region with an intensity of 15% of the original spectrum standard deviation, and the noise center value is the mean of the original spectrum standard deviation. Figure 7 Figure (b) shows the characteristic matrix of oil content and its screening process (the other physicochemical values ​​are similar). The red horizontal line in the figure represents the maximum value Cmax of the noise matrix; wavelengths with characteristic coefficients less than Cmax are discarded, while those greater are retained. Due to the stretching vibrations of methyl (-CH3) and methylene (-CH2-), the bending vibrations of CH bonds, and the vibrations of C=C double bonds, the retention peaks of the oil content matrix appear in the 1385nm-1633nm range. The retention peaks of protein, glucosinolates, and water content are concentrated in the 1040nm-1128nm range. Figure 7 As shown in (b), (c), (d), and (e), the characteristic coefficients in the oil content characteristic matrix exhibit larger fluctuations and more pronounced peaks. This is because oil content is a broad concept, encompassing multiple components such as oleic acid, linoleic acid, triglycerides, phytosterols, and phospholipids, resulting in a more complex mapping relationship in the near-infrared spectrum. Glucosides and water content, on the other hand, are single components, exhibiting similar, gradual trends. After removing uninformative variables, the RMSE values ​​for each physicochemical value were 8.7%, 13.4%, 77.3 μmol / g, and 6.5%. The number of characteristic wavelengths were 69, 32, 23, and 9.

[0073] 1.3 Spa Analysis After preprocessing, the original matrix was optimized using Spatial wavelength sequence selection, and PLS models were established for oil content, protein, glucosinolates, and water content, respectively. The influence of the number of wavelengths in the model on the prediction error is as follows: Figure 8As shown, the error trends of each physicochemical value model can be divided into three stages. In the initial stage, as the number of feature wavelengths increases, the RMSE decreases significantly, indicating that increasing the number of feature wavelengths helps to better capture information from the data, thereby improving prediction accuracy. When the number of feature wavelengths increases to a certain level, the RMSE tends to plateau, meaning that the added wavelengths no longer significantly improve the prediction effect, and the model has captured enough information. When the number of feature wavelengths continues to increase, the RMSE shows slight fluctuations. This is because adding too many features introduces noise, leading to overfitting or instability in the model.

[0074] by Figure 8 Taking a protein model as an example, in the initial stage, increasing the number of feature wavelengths can introduce new information and improve the model's predictive ability, thus reducing RMSE. However, once a threshold is reached, adding more features introduces redundant information and no longer significantly improves prediction performance. By selecting the most effective feature wavelengths, noise and redundancy can be reduced, improving the model's generalization ability. Considering both the number of wavelengths and prediction accuracy, the number of feature wavelengths for spectral data with protein as the indicator is 12, and the selected bands account for 5.8% of the original information. The number of feature wavelengths for oil content, glucosinolates, and water content are 18, 16, and 19, respectively.

[0075] 2. Near-infrared spectral prediction model for rapeseed In the near-infrared spectral data processing, this invention employs multiple preprocessing methods (SG, MSC, SNV) and feature extraction methods (CARS, SPA, UVE) to process the spectral data, selecting the feature wavelengths most relevant to the physicochemical values ​​of rapeseed (oil content, protein, glucosinolates, moisture content). By combining dataset partitioning methods (Random, SPXY) and regression methods (PLS, ELM), the selected spectral data is used as input, and different physicochemical values ​​are used as labels to develop a regression prediction model. The near-infrared spectral prediction model processing flow and results are as follows: Figure 9 As shown.

[0076] Among the preprocessing methods, multivariate scattering correction (MSC) and standard normal variable transformation (SNV) showed better processing results. Because rapeseed samples are unevenly distributed within the sample cup, particle size and spacing cause scattering and shifting of light within the sample, leading to baseline drift and increased noise. MSC effectively corrects the multiplicative scattering effect caused by uneven particle distribution by assuming a linear relationship between the scattering effect of all samples and the reference spectrum. SNV, on the other hand, eliminates additive scattering and baseline drift caused by optical path fluctuations or particle scattering by independently centering and standardizing the spectrum of each sample. SG smoothing filtering can remove random noise from spectral data, but its correction effect on baseline drift and scattering effects is weak, making it less effective than MSC and SNV when processing rapeseed spectral data.

[0077] In feature selection, SPA selects a set of variables with the largest projection length by minimizing collinearity among variables. In this invention, the physicochemical values ​​have low correlation with near-infrared spectroscopy, and SPA treatment has a negative effect on the model. UVE, after adding noise, effectively removes irrelevant variables of proteins because the stretching and bending vibrations of CH, NH, and OH bonds in the amino acid side chains and peptide bonds absorb near-infrared waves in the 950-1200 nm range, and the characteristic vibrations of C=O and NH in the amide I and amide II bands absorb near-infrared waves in the 1300-1600 nm range. The characteristic peaks at both ends of the protein are clearly independent, and UVE performs better when there is no multi-peak convergence. CARS has the highest versatility (but stricter parameter requirements). By iteratively selecting wavelengths and weighting them based on model performance and wavelength importance, it selects the wavelengths most relevant to the prediction target, showing good effects on oil content, water content, and glucosinolates.

[0078] In data partitioning, the SPXY method considers the eigenvalues ​​of the data to find and retain the main direction of data change, and has a better partitioning effect for constant elements. Oil content and protein can be partitioned using SPXY (the protein content is slightly lower, and the results of SPXY and Random partitioning are similar, so both methods can be used). The Random method does not consider the range of sample content, data distribution and abundance, and has good results in glucosinolates (8~138μmol / g) and water content (4.8~10.3%).

[0079] Both PLS and ELM are regression algorithms, but PLS is a multivariate regression technique that can extract the most relevant information to the dependent variable from a large number of wavelengths, making it particularly suitable for situations with a large number of variables. In this invention, the number of near-infrared spectral variables is 206, and after feature selection, it is approximately 40. Multicollinearity of spectral data can be solved by principal component analysis within the PLS model, but ELM, as a single-hidden-layer feedforward neural network, performs poorly in solving multicollinearity. Therefore, the PLS model performs better in predicting near-infrared spectroscopy.

[0080] In summary, the processing procedures and results of the optimal models for each physicochemical value are shown in Table 3. After SPXY+SNV+CARS+PLS treatment, the RMSE, R2, and MAE of oil content were 1.13%, 0.92%, and 0.91%, respectively; after Random+MSC+UVE+PLS treatment, the RMSE, R2, and MAE of protein content were 1.26%, 0.85%, and 0.93%, respectively; after Random+SNV+CARS+PLS treatment, the RMSE, R2, and MAE of glucosinolates content were 12.23 μmol / g, 0.79%, and 9.7 μmol / g, respectively; after Random+SNV+CARS+PLS treatment, the RMSE, R2, and MAE of water content were 0.32%, 0.94%, and 0.25%, respectively; after each treatment, the number of characteristic wavelengths for oil content, protein content, glucosinolates content, and water content were 56, 73, 24, and 41, respectively.

[0081] Table 3 Near-infrared spectroscopy processing procedures and results

[0082] 3. Determination of the core wavelength of the device The number of light sources is crucial for the design of portable devices. Since there are many wavelengths after filtering by various algorithms, Variable Importance Projection (VIP) scores are used for further analysis. The VIP calculation formula is shown below: ; In the formula, For the first The importance projection values ​​of each feature; For the dimensions of the projection; The sum of the main components; For the first The weights of each principal component or the proportion of variance explained; For the first The variable in the first... The coefficients in each principal component; For the first The Euclidean norm of the loading vectors of the principal components.

[0083] Analysis results as follows Figure 10As shown, the wavelengths 929, 974, 996, 1088, 1207, 1382, 1447, 1467, and 1550 nm were selected as the boundary when the sum of two physicochemical values ​​(VIP) > 2 or the sum of four physicochemical values ​​(VIP) > 3.6. Furthermore, literature review indicates a strong correlation of water content at 850 nm. Combining the above nine groups, ten groups of LEDs were selected at 850, 930, 975, 1000, 1090, 1210, 1380, 1450, 1470, and 1550 nm. Based on cost and a 20 nm LED bandwidth, the center wavelength of some LEDs was fine-tuned by 3-5 nm for production purposes.

[0084] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described; only preferred embodiments of the present invention are illustrated. The descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the present invention. As long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

[0085] It should be noted that those skilled in the art can make various modifications and improvements without departing from the inventive concept, and these all fall within the scope of protection of this invention. Therefore, the scope of protection of this invention should be determined by the appended claims.

Claims

1. A method for detecting rapeseed quality based on characteristic wavelength optimization and multi-model fusion, characterized in that, Includes the following steps: Step S1: Collect full-band near-infrared diffuse reflectance spectral data of rapeseed samples, and simultaneously determine the true values ​​of four physicochemical indicators of the rapeseed samples: oil content, protein, glucosinolates and moisture content, to construct the original dataset; Step S2: Based on the data distribution characteristics of each physicochemical index, select either a random partitioning method or a spatially balanced random partitioning method to divide the original dataset into a training set and a test set; Step S3: Select a preprocessing algorithm for each physicochemical index to process the full-band near-infrared diffuse reflectance spectral data; Step S4: Select a dimensionality reduction algorithm for each physicochemical index to extract feature wavelengths from the preprocessed spectral data; Step S5: Establish prediction models for each physicochemical index based on the training set, verify the performance of the prediction models using the test set, and determine the optimal algorithm combination for each physicochemical index. Step S6: Calculate the comprehensive contribution score of each wavelength point in the full-band spectrum to the four physicochemical indicators using the variable importance projection algorithm, set a multi-indicator comprehensive threshold, and determine the core light source wavelength of the portable device from the high-score region.

2. The method according to claim 1, characterized in that, In step S1, the spectral range of the full-band near-infrared diffuse reflectance spectral data is 900 nm to 1700 nm, the oil content is determined by Soxhlet extraction, the protein content is determined by Kjeldahl nitrogen determination, the glucosinolate content is determined by palladium chloride colorimetric method, and the water content is determined by direct drying method.

3. The method according to claim 1, characterized in that, In step S2, a spatial equilibrium random partitioning method is used for oil content, and a random partitioning method is used for protein, glucosinolates and water content. The spatial equilibrium random partitioning method calculates the Euclidean distance between samples based on the variable space of spectral and physicochemical values, and uses an iterative approach to select the sample with the farthest distance to add to the training set.

4. The method according to claim 1, characterized in that, In step S3, the preprocessing algorithm includes convolution smoothing, multivariate scattering correction, and standard normal variable transformation; standard normal variable transformation is selected for oil content, glucosinolates, and water content, while multivariate scattering correction is selected for proteins.

5. The method according to claim 4, characterized in that, The standard normal variable transformation independently performs mean centering and standardization on the spectrum of each sample, and the calculation formula is as follows: ; in, For the transformed th The first sample Data points, For the original first The first sample Data points, For the first The average of all data points in a sample For the first The standard deviation of all data points in a sample.

6. The method according to claim 1, characterized in that, In step S4, the dimensionality reduction algorithm includes competitive adaptive reweighted sampling, non-informative variable elimination, and continuous projection algorithm; Competitive adaptive reweighted sampling was selected for oil content, glucosinolates and water content, and non-informative variable elimination was selected for proteins.

7. The method according to claim 6, characterized in that, The formula for calculating variable weights in the competitive adaptive reweighted sampling method is as follows: ; in, For the first The absolute values ​​of the regression coefficients of each variable, For the first The weights of each variable, This represents the number of variables remaining in each sampling.

8. The method according to claim 1, characterized in that, In step S5, the prediction model includes a partial least squares regression model and an extreme learning machine model; the coefficient of determination R is used. 2 The root mean square error (RMSE) and mean absolute error (MAE) of the prediction model are used as evaluation metrics.

9. The method according to claim 1, characterized in that, In step S6, the calculation expression for the variable importance projection value algorithm is as follows: ; In the formula, For the first The importance projection values ​​of each feature; For the dimensions of the projection; The sum of the main components; For the first The weights of each principal component or the proportion of variance explained; For the first The variable in the first... The coefficients in each principal component; For the first The Euclidean norm of the loading vectors of the principal components.

10. The method according to claim 9, characterized in that, The core light source wavelength includes 10 discrete wavelengths: 850nm, 930nm, 975nm, 1000nm, 1090nm, 1210nm, 1380nm, 1450nm, 1470nm and 1550nm. These 10 discrete wavelengths correspond to the vibrational absorption peak positions of the CH, NH and OH functional groups inside rapeseed.