Crop quality index content prediction method, system, equipment, medium and product

By combining near-infrared spectral data and conductivity data, multivariate scattering correction and standard normal variable method are used to pre-treat, and feature bands are screened using competitive adaptive reweighting sampling method to construct a ridge regression model, solving the limitations and overfitting problems of traditional crop quality index prediction methods, and achieving high-precision multi-component synchronous detection.

CN120373536APending Publication Date: 2025-07-25ARTIFICIAL INTELLIGENCE INNOVATION RES INST OF ZHEJIANG UNIV OF TECH BINJIANG DISTRICT HANGZHOU
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510444468.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The traditional crop quality index content prediction method has a long time, high cost, high operator skills requirements and limitations in real-time production environments. In addition, mid-infrared spectroscopy technology is interfered with by factors such as light scattering, baseline drift, and concentration effects in high-precision principal component prediction, which reduces the prediction accuracy.

Method used

The combination of near-infrared spectral data and conductivity data was used to pre-treat the multivariate scattering correction and standard normal variable method, and the characteristic band screening was performed using competitive adaptive reweighting sampling method, and a ridge regression model was constructed to predict the content of crop quality indexes.

Benefits of technology

It improves the accuracy and prediction efficiency of crop quality index content prediction, breaks through the limitations of the traditional single detection mode, solves the overfitting problem in multi-component synchronous detection, and improves the prediction accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373536A_ABST
    Figure CN120373536A_ABST
Patent Text Reader

Abstract

The invention discloses a crop quality index content prediction method, system and device, a medium and a product, and relates to the technical field of crop quality index detection.The method comprises the steps that near infrared spectrum data and conductivity data of a to-be-detected crop sample are obtained and preprocessed, and the conductivity of the to-be-detected crop sample is obtained; carrying out characteristic wave band screening on the preprocessed near infrared spectrum data to obtain screened near infrared spectrum data, and fusing the screened near infrared spectrum data and the preprocessed conductivity data to obtain a fused characteristic vector of the to-be-detected crop sample; respectively inputting the fusion feature vectors of the to-be-detected crop samples into the crop quality index content prediction models to obtain corresponding predicted values of the crop quality index content; the crop quality index content is free amino acid content, starch content, soluble sugar content or vitamin C content. The accuracy and the prediction efficiency of crop quality index content prediction are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of crop quality index detection, and particularly to a method, system, device, medium and product for predicting the content of crop quality indexes. Background Technique

[0002] Predicting the content of crop quality indexes is not only a technical requirement, but also a key link throughout the entire "field to table" chain. Therefore, accurately determining the content of the main components in crops is crucial for agricultural production and the food industry. Traditional methods for predicting the content of main components usually involve chemical analysis, which have disadvantages such as long covering time, high cost, and high requirements for operator skills; in addition, the application of these methods in real-time production environments has certain limitations. Currently, mid-infrared spectroscopy technology has been widely used to analyze the components of agricultural products. This technology can provide information about the sample components by measuring the mid-infrared spectra absorbed, reflected or transmitted in the sample. However, traditional mid-infrared spectroscopy technology has limitations in predicting the content of high-precision main components because spectral data may be interfered by factors such as light scattering, baseline drift, and concentration effects, reducing the accuracy of prediction.

[0003] Therefore, it is necessary to provide a method for predicting the content of crop quality indexes to solve the above problems. Summary of the Invention

[0004] The purpose of the present application is to provide a method, system, device, medium and product for predicting the content of crop quality indexes, which improves the accuracy and prediction efficiency of predicting the content of crop quality indexes.

[0005] To achieve the above purpose, the present application provides the following solutions:

[0006] In a first aspect, the present application provides a method for predicting the content of crop quality indexes, and the method for predicting the content of crop quality indexes includes:

[0007] Obtain a crop sample to be measured;

[0008] Obtain the near-infrared spectral data and conductivity data of the crop sample to be measured;

[0009] Preprocess the near-infrared spectral data and conductivity data of the crop sample to be measured respectively to obtain the preprocessed near-infrared spectral data and preprocessed conductivity data of the crop sample to be measured;

[0010] Perform feature band screening on the preprocessed near-infrared spectral data of the crop sample to be measured to obtain the screened near-infrared spectral data of the crop sample to be measured;

[0011] Fuse the screened near-infrared spectral data and preprocessed conductivity data of the crop sample to be measured to obtain the fused feature vector of the crop sample to be measured;

[0012] Input the fused feature vector of the crop sample to be measured into each crop quality index content prediction model respectively to obtain the predicted values of the corresponding crop quality index contents; the crop quality index contents are free amino acid content, starch content, soluble sugar content or vitamin C content; each of the crop quality index content prediction models is obtained by training each regression model using a training set.

[0013] Optionally, obtaining the near-infrared spectral data and conductivity data of the crop sample to be measured specifically includes:

[0014] Use a near-infrared hyperspectral camera to perform near-infrared spectral detection on the crop sample to be measured to obtain the near-infrared spectral data of the crop sample to be measured;

[0015] Use a conductivity meter to measure the conductivity of the crop sample to be measured to obtain the conductivity data of the crop sample to be measured.

[0016] Optionally, preprocess the near-infrared spectral data of the crop sample to be measured and the conductivity data of the crop sample to be measured respectively to obtain the preprocessed near-infrared spectral data and preprocessed conductivity data of the crop sample to be measured, specifically including:

[0017] Adopt the multiplicative scatter correction method to process the near-infrared spectral data of the crop sample to be measured to obtain the corrected near-infrared spectral data of the crop sample to be measured;

[0018] Use the standard normal variate method to process the corrected near-infrared spectral data of the crop sample to be measured to obtain the preprocessed near-infrared spectral data of the crop sample to be measured;

[0019] Normalize the conductivity data of the crop sample to be measured to obtain the preprocessed conductivity data of the crop sample to be measured.

[0020] Optionally, adopt the competitive adaptive reweighted sampling method to perform feature band selection on the normalized near-infrared spectral data of the crop sample to be measured.

[0021] Optionally, the training process of the crop quality index content prediction model specifically includes:

[0022] Construct a training set; the training set includes the fused feature vectors of multiple sample crop samples and the true values of the free amino acid content, starch content, soluble sugar content and vitamin C content of the sample crops;

[0023] Construct a regression model;

[0024] Taking the fusion feature vector of the sample crop samples as the input, and taking the true value of the free amino acid content, the true value of the starch content, the true value of the soluble sugar content, or the true value of the vitamin C content of the sample crops as the output, train each regression model to obtain the corresponding prediction model for the content of crop quality indicators.

[0025] Optionally, the regression model is a ridge regression model, a random forest regression model, a support vector machine regression model, or a fully connected neural network model.

[0026] In a second aspect, the present application provides a prediction system for the content of crop quality indicators. The prediction system for the content of crop quality indicators is used to implement the prediction method for the content of crop quality indicators. The prediction system for the content of crop quality indicators includes:

[0027] A sample acquisition unit for acquiring a crop sample to be measured;

[0028] A data acquisition unit for acquiring the near-infrared spectrum data and conductivity data of the crop sample to be measured;

[0029] A preprocessing unit for preprocessing the near-infrared spectrum data of the crop sample to be measured and the conductivity data of the crop sample to be measured respectively, to obtain the preprocessed near-infrared spectrum data and preprocessed conductivity data of the crop sample to be measured;

[0030] A characteristic band screening unit for screening characteristic bands from the preprocessed near-infrared spectrum data of the crop sample to be measured to obtain the screened near-infrared spectrum data of the crop sample to be measured;

[0031] A fusion unit for fusing the screened near-infrared spectrum data and the preprocessed conductivity data of the crop sample to be measured to obtain the fusion feature vector of the crop sample to be measured;

[0032] A prediction unit for inputting the fusion feature vector of the crop sample to be measured into each prediction model for the content of crop quality indicators respectively to obtain the predicted values of the corresponding content of crop quality indicators; the content of the crop quality indicators is the free amino acid content, the starch content, the soluble sugar content, or the vitamin C content; each prediction model for the content of crop quality indicators is obtained by training each regression model using a training set.

[0033] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement the prediction method for the content of crop quality indicators described in any one of the above.

[0034] Fourthly, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method for predicting the content of crop quality indicators described in any one of the above is implemented.

[0035] Fifthly, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the method for predicting the content of crop quality indicators described in any one of the above is implemented.

[0036] According to the specific embodiments provided by the present application, the present application has the following technical effects:

[0037] The present application discloses a method, a system, a device, a medium and a product for predicting the content of crop quality indicators. First, the near-infrared spectrum data of the crop sample to be measured and the conductivity data of the crop sample to be measured are preprocessed, and the characteristic bands of the preprocessed near-infrared spectrum data are screened. The screened near-infrared spectrum data and the preprocessed conductivity data are fused to obtain the fused feature vector of the crop sample to be measured. Then, the fused feature vector is respectively input into each crop quality indicator content prediction model to obtain the predicted values of the corresponding crop quality indicator content. The crop quality indicator content is the free amino acid content, the starch content, the soluble sugar content or the vitamin C content. By integrating the near-infrared spectrum and conductivity data in the crop in a multi-dimensional data manner, the limitations of the traditional single detection mode are broken through, and the accuracy and efficiency of predicting the content of agricultural product quality indicators are improved. By constructing a crop quality indicator content prediction model based on ridge regression, the overfitting problem of traditional algorithms such as PLS in multi-component synchronous detection is overcome, and the prediction accuracy is improved. Description of the Drawings

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0039] Figure 1 Schematic flowchart of the method for predicting the content of crop quality indicators provided by an embodiment of the present application;

[0040] Figure 2 Schematic diagram of the near-infrared spectrum of a young soybean sample provided by an embodiment of the present application;

[0041] Figure 3 Schematic diagram of the near-infrared spectrum characteristic bands of the screened young soybean sample obtained by the CARS algorithm provided by an embodiment of the present application;

[0042] Figure 4 Schematic diagram of the characteristic band screening results of vitamin C based on the CARS algorithm provided by an embodiment of the present application;

[0043] Figure 5 Schematic diagram of the characteristic band screening results of free amino acids based on the CARS algorithm provided by an embodiment of the present application;

[0044] Figure 6 Schematic diagram of the characteristic band screening results of soluble sugar based on the CARS algorithm provided by an embodiment of the present application;

[0045] Figure 7 Schematic diagram of the characteristic band screening results of starch based on the CARS algorithm provided by an embodiment of the present application;

[0046] Figure 8 Schematic diagram of the verification results of starch content prediction of edamame samples based on ridge regression provided by an embodiment of the present application;

[0047] Figure 9 Schematic diagram of the verification results of free amino acid content prediction of edamame samples based on ridge regression provided by an embodiment of the present application;

[0048] Figure 10 Schematic diagram of the verification results of soluble sugar content prediction of edamame samples based on ridge regression provided by an embodiment of the present application;

[0049] Figure 11 Schematic diagram of the verification results of vitamin C content prediction of edamame samples based on ridge regression provided by an embodiment of the present application;

[0050] Figure 12 Schematic diagram of the prediction results of starch content of edamame samples provided by an embodiment of the present application;

[0051] Figure 13 Schematic diagram of the prediction results of free amino acid content of edamame samples provided by an embodiment of the present application;

[0052] Figure 14 Schematic diagram of the prediction results of soluble sugar content of edamame samples provided by an embodiment of the present application;

[0053] Figure 15 Schematic diagram of the prediction results of vitamin C content of edamame samples provided by an embodiment of the present application;

[0054] Figure 16 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners

[0055] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0056] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0057] In an exemplary embodiment, as Figure 1 shown, a method for predicting the content of crop quality indicators is provided, including the following steps. Among them:

[0058] Step S1, obtain the crop sample to be measured.

[0059] Step S2, obtain the near-infrared spectrum data and conductivity data of the crop sample to be measured.

[0060] As an alternative implementation manner, step S2 specifically includes:

[0061] Step S21, use a near-infrared hyperspectral camera to perform near-infrared spectrum detection on the crop sample to be measured to obtain the near-infrared spectrum data of the crop sample to be measured.

[0062] Step S22, use a conductivity meter to measure the conductivity of the crop sample to be measured to obtain the conductivity data of the crop sample to be measured.

[0063] Step S3, preprocess the near-infrared spectrum data of the crop sample to be measured and the conductivity data of the crop sample to be measured respectively to obtain the preprocessed near-infrared spectrum data and preprocessed conductivity data of the crop sample to be measured.

[0064] Specifically, in order to eliminate the background, noise, and interference caused by instrument and position changes during the data acquisition process, it is necessary to preprocess the original near-infrared spectrum data. By applying the multiplicative scatter correction (MSC) and standard normal variate (SNV) methods to preprocess the original near-infrared spectrum data, the interference of noise and baseline drift can be eliminated.

[0065] As an alternative implementation manner, step S3 specifically includes:

[0066] Step S31: Process the near-infrared spectral data of the crop sample to be measured using the multiplicative scatter correction method to obtain the corrected near-infrared spectral data of the crop sample to be measured.

[0067] Specifically, the multiplicative scatter correction method corrects the influence of multiple scattering by constructing the ratio of the sample near-infrared spectral data (i.e., the near-infrared spectral data of the crop sample) to the reference near-infrared spectral data (i.e., the standard near-infrared spectral data without any baseline drift and any interference). Multiple scattering will cause an increase in the optical path length and a decrease in the signal intensity. The basic formula for MSC correction is as follows:

[0068] R MSC =(R sample -R min ) / (R ref -R min ) (1)

[0069] where R MSC is the corrected near-infrared spectral data of the crop sample; R sample is the sample near-infrared spectral data; R ref is the reference near-infrared spectral data, and R min is the minimum value of the near-infrared spectral data. The purpose of MSC correction is to remove the residual scattering components in the spectrum and improve the accuracy and comparability of the sample near-infrared spectral data.

[0070] Step S32: Process the corrected near-infrared spectral data of the crop sample to be measured using the standard normal variate method to obtain the preprocessed near-infrared spectral data of the crop sample to be measured.

[0071] Specifically, the standard normal variate method is a normalization technique used to eliminate the differences caused by differences in light intensity and baseline drift. The basic formula for SNV processing is as follows:

[0072] R SNV =(R MSC -μ) / σ (2)

[0073] where R SNV is the preprocessed near-infrared spectral data; μ is the average value of the corrected near-infrared spectral data of the crop sample; σ is the standard deviation of the corrected near-infrared spectral data of the crop sample. For each spectrum, the SNV method subtracts the average value from each data point in the near-infrared spectral data and then divides by the standard deviation. In this case, the overall intensity change and baseline drift in the near-infrared spectral data will be eliminated, thus highlighting the spectral features.

[0074] By comparing the original near-infrared spectral data, both MSC and SNV can effectively reduce noise, but there is little difference in terms of trends.

[0075] Step S33, normalize the conductivity data of the crop sample to be measured to obtain the preprocessed conductivity data of the crop sample to be measured.

[0076] Specifically, standardize the conductivity data of each crop sample. Since the conductivity data and the near-infrared spectroscopy data have different magnitudes, it is necessary to normalize them so as to perform model training after combining with the near-infrared spectroscopy data. The standardization formula for the conductivity data is as follows:

[0077]

[0078] Where X is the conductivity value of the crop sample, mean is the mean value of the conductivity of the crop sample, and std_dev is the standard deviation of the conductivity data of the crop sample. By this method, the conductivity data is normalized to a magnitude similar to that of the near-infrared spectroscopy data, facilitating subsequent fusion and analysis.

[0079] Step S4, perform feature band screening on the preprocessed near-infrared spectroscopy data of the crop sample to be measured to obtain the screened near-infrared spectroscopy data of the crop sample to be measured.

[0080] As an optional implementation manner, in step S4, the competitive adaptive reweighted sampling method is used to perform feature band selection on the normalized near-infrared spectroscopy data of the crop sample to be measured.

[0081] Specifically, the competitive adaptive reweighted sampling (CARS) method is used, based on Monte Carlo sampling and partial least squares (PLS) regression coefficients, for optimizing the feature bands of the near-infrared spectroscopy data. CARS is used to sample the preprocessed near-infrared spectroscopy data. In each iteration, an exponential decay function is used to randomly reselect the calibration set samples for easy selection of bands. The data is optimized by the competitive adaptive reweighted sampling method. The subset with the minimum root mean square error of cross-validation is selected as the feature band. This algorithm is implemented using the sklearn library in Python, and the feature bands are extracted from the preprocessed near-infrared spectroscopy data.

[0082] Step S5, fuse the screened near-infrared spectroscopy data and the preprocessed conductivity data of the crop sample to be measured to obtain the fused feature vector of the crop sample to be measured.

[0083] Specifically, the conductivity data is used as an additional physical feature and fused with the selected near-infrared spectral bands. The introduction of conductivity data can help improve the model's understanding of the physical and chemical properties of samples and make up for the deficiencies that may be brought by single spectral information.

[0084] Assume that n spectral band features are extracted and expressed as:

[0085] X spectral =[x1, x2,..., x n (4)

[0086] where X spectral represents the spectral feature vector extracted from the near-infrared spectral data, and each of x1, x2, x3,..., x n corresponding to x i represents the value of the i-th feature band, which is the most representative spectral band selected by the CARS algorithm.

[0087] Each crop sample corresponds to a conductivity value, and the conductivity data is denoted as d.

[0088] Taking the conductivity data as an additional feature, it is concatenated with the spectral feature vector to form a fused feature vector:

[0089] X fused =[x1, x2,..., x n , d] (5)

[0090] where X fused represents the fused feature vector obtained by concatenating the spectral feature vector X spectral and the conductivity data d.

[0091] Step S6: Input the fused feature vector of the crop sample to be measured into the crop quality index content prediction model to obtain the predicted value of the corresponding crop quality index content. The crop quality index content includes: free amino acid content, starch content, soluble sugar content, and vitamin C content. The crop quality index content prediction model is obtained by training a regression model using a training set.

[0092] Among them, taking the fused feature vector X fused as the input and inputting it into the crop quality index content prediction model, the predicted value y of the corresponding crop quality index content is output:

[0093] y = f(X fuesd ) (6)

[0094] where f is the regression model and y is the prediction result of the crop quality index content prediction model (i.e., the predicted value of the crop quality index content).

[0095] As an alternative implementation, in step S6, the training process of the crop quality index content prediction model specifically includes:

[0096] Step S61, constructing a training set.

[0097] Specifically, the training set includes: the fusion feature vectors of multiple sample crop samples and the true values (i.e., labels) of the sample crop quality index contents. Among them, the labels include: the true values of the free amino acid content, starch content, soluble sugar content, and vitamin C content of the sample crops. Among them, the fusion feature vectors of the sample crop samples are obtained by fusing the screened near-infrared spectral data and the preprocessed conductivity data of the sample crop samples.

[0098] The true values of the sample crop quality index contents are obtained by detecting the sample crop samples through experiments, specifically including:

[0099] 1) Dissolve the sample crop sample in hydrochloric acid, perform derivatization treatment, and then separate and detect it by high performance liquid chromatography to determine the true value of the free amino acid content of the sample crop sample.

[0100] 2) After extracting the sample crop sample, add amylase for hydrolysis, and detect the generated glucose by colorimetry or high performance liquid chromatography to obtain the true value of the starch content of the sample crop sample.

[0101] 3) Vitamin C has strong reducibility. Add oxidized 2,6-dichlorophenol indophenol dye to the sample crop sample. The vitamin C in the sample crop sample undergoes a quantitative oxidation-reduction reaction with the oxidized 2,6-dichlorophenol indophenol dye. When the solution color fades from blue to colorless, it is the titration end point, and the true value of the vitamin C content of the sample crop sample is calculated by the volume of the consumed dye.

[0102] 4) Soluble sugar dehydrates under the action of concentrated sulfuric acid to generate furfural substances, which combine with phenol to form an orange-yellow compound. Place the sample crop sample in concentrated sulfuric acid and perform colorimetric determination at a wavelength of 490 nm to determine the true value of the soluble sugar content of the sample crop sample.

[0103] Specifically, randomly divide the screened near-infrared spectral data and the preprocessed conductivity data into a training set and a test set according to a ratio of 7:3. A total of 235 training samples and 101 test samples are used. The conductivity data and the spectral data are kept consistent during the division process to ensure the integrity of the modal information during training and testing.

[0104] Step S62, constructing a regression model.

[0105] As an alternative implementation, in step S62, the regression model is a ridge regression model, a random forest regression model, a support vector machine regression model, or a fully connected neural network model.

[0106] Preferably, the regression model is a ridge regression model.

[0107] In order to accurately predict the contents of free amino acids, starch, vitamin C, and soluble sugars in crops through near-infrared spectral information and conductivity data, this application selects a variety of mainstream regression models for systematic experiments. The aim is to construct a mathematical mapping relationship between the input (near-infrared spectral data and conductivity data) and the output (contents of free amino acids, starch, vitamin C, and soluble sugars), so as to achieve efficient prediction of the crops to be measured. The regression models used include a ridge regression model, a random forest regression model, a support vector machine regression model, and a fully connected neural network model. Among them, as an extension of linear regression, ridge regression effectively solves the problem of multicollinearity by imposing L2 regularization on the regression coefficients, and it performs particularly well especially when the data samples are relatively limited. In contrast, when the sample size is small, random forests, neural networks, and support vector machines are prone to overfitting due to their high model complexity, and their response ability is weak when dealing with high correlations between variables. Among them, considering that there is a significant correlation between the characteristic bands and the number of data samples is limited, this application preferentially uses the ridge regression model as the final regression model.

[0108] Step S63: Using the fusion feature vector of the sample crop samples as the input and the true values of the free amino acid content, starch content, soluble sugar content, or vitamin C content of the sample crops as the output, train each regression model to obtain the corresponding prediction model for the crop quality index content.

[0109] Model architecture:

[0110] Taking the ridge regression model as an example, in order to ensure the best performance of the model, the key hyperparameters of the ridge regression model are systematically optimized. Ridge regression is used as the main prediction model to make full use of its advantages in dealing with multicollinearity and capturing nonlinear relationships. Ridge regression effectively reduces the variance of the model parameters and improves the generalization ability of the model by introducing L2 regularization.

[0111] Selection of the regularization parameter (α) of the ridge regression model:

[0112] The cross-validation method is adopted to select the optimal regularization parameter within a predetermined range through grid search to balance the bias and variance of the model. The regularization parameter α set for the prediction task of the crop quality index is set to 0.05.

[0113] Model training:

[0114] Based on the optimized hyperparameters, the ridge regression model and the random forest model are trained on the training set respectively:

[0115] Training of the ridge regression model: Using the fusion vector as the input feature and the content of the crop quality index as the target variable (i.e., the model input), the model parameters are fitted by minimizing the loss function with an L2 regularization term to obtain the prediction model for the content of the crop quality index. To adapt to the input requirements of the ridge regression model, the training set and its labels are subjected to necessary shape conversion processing, and the regularization parameter α is set to 0.1 to achieve the best balance between model complexity and generalization ability.

[0116] Model validation and evaluation:

[0117] The trained ridge regression model is validated and its performance is evaluated through cross-validation and an independent validation set. Metrics such as the coefficient of determination (R 2 ) and the root mean square error (RMSE) are used to comprehensively measure the prediction accuracy and stability of the model, ensuring its generalization ability on different data sets.

[0118] The training set is used to train the regression model, and the test set is used to test the performance of the trained regression model. Based on the coefficient of determination R 2 and RMSE between the predicted values and the true values of the test set as metrics to determine the effectiveness of the model. Among them:

[0119]

[0120] Among them, represents the predicted value of the content of the i-th crop quality index; y i represents the true value of the content of the i-th crop quality index, and m is the number of crop quality index contents.

[0121] The prediction accuracy of the starch content in this application is about 97%, and the prediction accuracy of the free amino acid, vitamin C, and soluble sugar contents is about 98%.

[0122] Beneficial effects:

[0123] 1) Innovatively integrating the frequency doubling differences (chemical properties) of chemical bonds such as C-H, C-O, and H-O in organic molecules in the near-infrared spectrum in the 780-2500mm band with the conductivity (physical property) for multi-dimensional data integration, breaking through the limitations of traditional single detection modes and improving the accuracy of predicting the content of agricultural product quality indicators.

[0124] 2) Constructing a multi-task joint prediction model based on ridge regression, overcoming the overfitting problem of traditional algorithms such as PLS in multi-component synchronous detection.

[0125] Example 1:

[0126] Edamame (i.e., vegetable soybeans) is an important agricultural product with high nutritional value and is an important source for human food, feed, and industrial applications. Taking edamame as an example, the method for predicting the content of crop quality indicators in this application will be further elaborated below.

[0127] Step S1, obtain edamame samples.

[0128] The edamame materials used in this example were provided by the vegetable soybean research group of a certain vegetable research institute. This population contains 336 materials, including different varieties from multiple regions. They are planted in March every year, harvested in June of the same year, and field phenotypic traits are identified. According to the completely randomized block design, deep ditches and high ridges are used, the ridge width is 0.9 - 1.1 m, 2 rows are planted, the row spacing is 40 - 45 cm, the hole spacing is 25 - 30 cm, and 3 - 5 seeds are sown in each hole. The field fertilizer and water management are kept consistent with production.

[0129] For each material, several fresh seeds of edamame at a relatively consistent commercial maturity stage are taken, frozen and preserved in liquid nitrogen, then dried for 96 h to remove moisture using a freeze dryer (LABCONCO, USA), and then ground into fine powder using a JXFSTPRP-24L full-automatic sample rapid grinder of Shanghai Jingxin brand. The 336 core germplasm populations of edamame are numbered and used as edamame samples.

[0130] Step S2, obtain the near-infrared spectroscopy data and conductivity data of the edamame samples.

[0131] Step S21, collection of near-infrared spectroscopy data.

[0132] The instrument used for collecting near-infrared spectroscopy data is a near-infrared hyperspectral camera, model: FS-15. The wavelength range of the collected near-infrared is between 900 nm and 1700 nm. The freeze-dried powder of edamame is placed at the sampling port of the spectrometer for spectral data collection. After running the spectrometer, near-infrared spectroscopy data in the wavelength range of 900 nm - 1700 nm is obtained.

[0133] The edamame samples are loaded into sample cups with a diameter of 5 cm and a height of about 1.5 cm. The thickness and compaction degree of the edamame samples in the sample cups are kept consistent to minimize the measurement error caused by uneven sample loading. Each edamame sample is scanned three times, and the average absorption spectrum is calculated as the spectral value. The near-infrared spectroscopy data of all edamame samples are collected and stored.

[0134] Visualize the above-stored near-infrared spectroscopy data to obtain a near-infrared spectroscopy schematic diagram as Figure 2As shown, in this embodiment, within the near-infrared spectrum range of 336 core germplasm populations of green soybeans, the absorption peak of starch is around 6700 cm -1 (1492 nm); meanwhile, for free amino acids, their absorption peaks are concentrated around 4700 cm -1 (2127 nm) and 5100 cm -1 (1960 nm); for soluble sugars, their absorption peaks are mainly concentrated around 1100 nm, and for vitamin C, its absorption peaks are mainly concentrated around 1350 nm.

[0135] Step S22: Measure the conductivity data of the green soybean samples.

[0136] Use a precision conductivity meter to measure the conductivity of the pressed green soybean samples. The conductivity meter uses the four-probe method, applying current through the external electrodes and measuring the voltage change in the green soybean samples through the internal electrodes. This can reduce the influence of contact resistance on the measurement results and improve the measurement accuracy. Record the conductivity data of each green soybean sample. To ensure the stability of the results, each green soybean sample is measured at least three times, and the average value is taken as the final conductivity value. The conductivity data can reflect the internal physical properties of the samples, especially those related to conductivity and ion movement. Combining with the spectral data, conductivity can provide additional sample information, which helps to improve the prediction accuracy of the contents of free amino acids, starch, vitamin C, and soluble sugars.

[0137] Step S3: Preprocess the near-infrared spectral data and the conductivity data of the green soybean samples respectively to obtain the preprocessed near-infrared spectral data and the preprocessed conductivity data of the green soybean samples.

[0138] Preprocess the near-infrared spectral data of the green soybean samples by the MSC method and the SNV method to eliminate the interference of noise and baseline drift. MSC can eliminate the spectral differences caused by different scattering levels during the near-infrared spectrum measurement (the near-infrared spectra are different due to different measurement positions, light during measurement, etc.), and enhance the correlation between the near-infrared spectrum and the data. This method needs to use the "ideal near-infrared spectrum" (the change of the near-infrared spectrum and the content of components in the sample satisfy a direct linear relationship) to correct the spectral differences caused by different scattering levels (i.e., correct the baseline translation and offset phenomena of the near-infrared spectral data). Since it is difficult to obtain the "ideal near-infrared spectrum", the average value of all near-infrared spectral data is used to replace the "ideal near-infrared spectrum" in this algorithm. The specific results are as Figure 2 shown.

[0139] Perform standardization processing on the conductivity data to obtain the preprocessed conductivity data.

[0140] Step S4: Screen the preprocessed near-infrared spectral data of the edamame samples to obtain the screened near-infrared spectral data of the edamame samples.

[0141] In this embodiment, CARS is used to extract significant characteristic bands. Figure 3 The distribution of these characteristic bands is shown. From Figure 3 it can be observed that the extracted characteristic bands are mainly concentrated between 900 nm and 1700 cm -1 Compared with the original near-infrared spectral data, the number of bands is significantly reduced. Applying the CARS algorithm, highly correlated band number results for different crop quality indicators are obtained. The band number results are as Figures 4 - 7 shown. It can be seen that for soluble sugar, the number of highly correlated bands is 58; for free amino acids, the number of highly correlated bands is 114; for vitamin C, the number of highly correlated bands is 29; for starch, the number of highly correlated bands is 26.

[0142] In addition, from Figure 3 the extracted characteristic bands, the distribution of the characteristic bands obtained based on the CARS algorithm is relatively uniform. The CARS algorithm selects and extracts the most representative features through a competitive process without explicitly favoring specific bands.

[0143] Step S5: Fuse the screened near-infrared spectral data and the preprocessed conductivity data of the edamame samples to obtain the fused feature vectors of the edamame samples.

[0144] Step S6: Input the fused feature vectors of the edamame samples into the prediction model for the content of crop quality indicators to obtain the predicted values of the corresponding edamame quality indicator contents. The edamame quality indicator contents are: free amino acid content, starch content, vitamin C content, and soluble sugar content. The training set includes the fused feature vectors of multiple sample edamame samples and the true values of the free amino acid content, starch content, vitamin C content, and soluble sugar content of the sample edamame.

[0145] Among them, when obtaining the true values of the sample edamame quality indicator contents, specific instruments are used to measure the true values of the free amino acid content, starch content, vitamin C content, and soluble sugar content of the edamame samples.

[0146] Each numbered sample was divided into three parts for processing and testing. Take out the first part of the edamame sample and dissolve it in a hydrochloric acid solution of a certain concentration to ensure that the sample is completely dissolved and evenly distributed. Perform derivatization on the dissolved sample. Derivatization is to make free amino acids more easily detectable and separable, usually by using appropriate reagents for chemical reactions. Inject the derivatized sample into a high-performance liquid chromatograph for separation and detection. HPLC can accurately separate different free amino acids and measure their contents through a detector, and finally obtain the true value of the free amino acid content in the sample; take out the second part of the edamame sample and extract the starch in the edamame sample by an appropriate extraction method (such as extraction with water or buffer solution). Add amylase to the extracted sample for enzymatic hydrolysis. Amylase will decompose starch into glucose for subsequent detection. Use colorimetry or high-performance liquid chromatography (HPLC) to detect the glucose content generated after enzymatic hydrolysis; colorimetry measures the change in absorbance of the sample solution through a colorimeter to calculate the glucose content, while HPLC separates and detects glucose molecules to obtain a more accurate true value of the starch content; add oxidized 2,6-dichloroindophenol dye to the sample edamame sample, and calculate the true value of the vitamin C content of the sample crop sample by calculating the volume of the consumed dye; place the sample crop sample in concentrated sulfuric acid and measure the true value of the soluble sugar content of the sample crop sample by colorimetry at a wavelength of 490 nm.

[0147] The ridge regression model was used as the main regression model to make full use of its advantages in dealing with multicollinearity and capturing nonlinear relationships. The regularization parameter (α) set for the prediction task of the important quality indicators of edamame was set to 0.05.

[0148] Step S7, based on the coefficient of determination R of the predicted value and the actual value of the test set 2 and RMSE as indicators to determine the effectiveness of the crop quality indicator content prediction model.

[0149] Tables 1, 2, 3, and 4 are the verification results of different crop quality indicator content prediction models obtained by modeling based on different regression models. Figures 8 - 11 It is a schematic diagram of the performance results of different crop quality indicator content prediction models obtained by modeling based on the ridge regression model on the test set. It can be seen that the predicted values of starch content, free amino acid content, soluble sugar, and vitamin C obtained by the crop quality indicator content prediction model of this application have a good fitting effect with the observed data (test set); it can be seen that on the performance of the test set, the ridge regression model has obvious advantages. Figures 12 - 15 It is a schematic diagram of the prediction results of different crop quality indicator contents.

[0150] Table 1 Verification Results Table of Starch Content Prediction Model

[0151] Random Forest Support Vector Machine Ridge Regression <![CDATA[R 2 =(CARS)]]> 0.944 0.767 0.973 RMSE(CARS) 0.272 0.645 0.223

[0152] Table 2 Verification Results Table of Free Amino Acid Content Prediction Model

[0153] Random Forest Support Vector Machine Ridge Regression <![CDATA[R 2 (CARS)]]> 0.943 0.561 0.977 RMSE(CARS) 0.276 0.533 0.225

[0154] Table 3 Verification Results Table of Soluble Sugar Content Prediction Model

[0155] Random Forest Support Vector Machine Ridge Regression <![CDATA[R 2 (CARS)]]> 0.953 0.51 0.969 RMSE(CARS) 0.251 0.432 0.315

[0156] Table 4 Verification Results Table of Vitamin C Content Prediction Model

[0157] Random Forest Support Vector Machine Ridge Regression <![CDATA[R 2 (CARS)]]> 0.941 0.561 0.983 RMSE(CARS) 0.28 0.533 0.217

[0158] Based on the same inventive concept, the embodiments of the present application further provide a crop quality index content prediction system for implementing the crop quality index content prediction method involved above. The implementation solutions provided by this system to solve problems are similar to those recorded in the above method. Therefore, the specific limitations in one or more embodiments of the crop quality index content prediction system provided below can refer to the limitations on the crop quality index content prediction method in the above text, and will not be elaborated here.

[0159] In an exemplary embodiment, a crop quality index content prediction system is provided, including:

[0160] A sample acquisition unit for acquiring crop samples to be measured.

[0161] A data acquisition unit for acquiring near-infrared spectrum data and conductivity data of the crop samples to be measured. A preprocessing unit for preprocessing the near-infrared spectrum data and the conductivity data of the crop samples to be measured respectively, to obtain the preprocessed near-infrared spectrum data and the preprocessed conductivity data of the crop samples to be measured.

[0162] A characteristic band screening unit for screening characteristic bands from the preprocessed near-infrared spectrum data of the crop samples to be measured, to obtain the screened near-infrared spectrum data of the crop samples to be measured.

[0163] A fusion unit for fusing the screened near-infrared spectrum data and the preprocessed conductivity data of the crop samples to be measured, to obtain the fused feature vector of the crop samples to be measured.

[0164] A prediction unit for respectively inputting the fusion feature vectors of the crop samples to be measured into each crop quality index content prediction model to obtain the predicted values of the corresponding crop quality index contents; the crop quality index contents are free amino acid content, starch content, soluble sugar content or vitamin C content; each of the crop quality index content prediction models is obtained by training each regression model using a training set.

[0165] In an exemplary embodiment, a computer device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement a method for predicting crop quality index contents.

[0166] In an exemplary embodiment, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements a method for predicting crop quality index contents.

[0167] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements a method for predicting crop quality index contents.

[0168] In an exemplary embodiment, a computer device is provided. The computer device can be a server or a terminal, and its internal structure diagram can be as Figure 16 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used for exchanging information between the processor and external devices. The communication interface of the computer device is used for communicating with external terminals through a network connection. When the computer program is executed by the processor, it implements a method for predicting crop quality index contents.

[0169] Those skilled in the art can understand that Figure 16 the structure shown in

[0170] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.

[0171] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.

[0172] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.

[0173] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.

[0174] In this text, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.

Claims

1. A method for predicting the content of crop quality indicators, characterized in that, The method for predicting the content of crop quality indicators includes: Obtain a crop sample to be measured; Obtain the near-infrared spectral data and conductivity data of the crop sample to be measured; Preprocess the near-infrared spectral data and conductivity data of the crop sample to be measured respectively to obtain the preprocessed near-infrared spectral data and preprocessed conductivity data of the crop sample to be measured; Screen the characteristic bands of the preprocessed near-infrared spectral data of the crop sample to be measured to obtain the screened near-infrared spectral data of the crop sample to be measured; Fuse the screened near-infrared spectral data and preprocessed conductivity data of the crop sample to be measured to obtain the fused feature vector of the crop sample to be measured; Input the fused feature vector of the crop sample to be measured into each crop quality indicator content prediction model respectively to obtain the predicted values of the corresponding crop quality indicator contents; the crop quality indicator contents are free amino acid content, starch content, soluble sugar content or vitamin C content; each of the crop quality indicator content prediction models is obtained by training each regression model using a training set.

2. The method for predicting the content of crop quality indicators according to claim 1, wherein Obtaining the near-infrared spectral data and conductivity data of the crop sample to be measured specifically includes: Use a near-infrared hyperspectral camera to perform near-infrared spectral detection on the crop sample to be measured to obtain the near-infrared spectral data of the crop sample to be measured; Use a conductivity meter to measure the conductivity of the crop sample to be measured to obtain the conductivity data of the crop sample to be measured.

3. The method for predicting the content of crop quality indicators according to claim 1, characterized in that, Preprocessing the near-infrared spectral data and conductivity data of the crop sample to be measured respectively to obtain the preprocessed near-infrared spectral data and preprocessed conductivity data of the crop sample to be measured specifically includes: Use the multiplicative scatter correction method to process the near-infrared spectral data of the crop sample to be measured to obtain the corrected near-infrared spectral data of the crop sample to be measured; Use the standard normal variate method to process the corrected near-infrared spectral data of the crop sample to be measured to obtain the preprocessed near-infrared spectral data of the crop sample to be measured; Normalize the conductivity data of the crop sample to be measured to obtain the preprocessed conductivity data of the crop sample to be measured.

4. The method for predicting the content of crop quality indexes according to claim 3, wherein, Use the competitive adaptive reweighted sampling method to select characteristic bands from the normalized near-infrared spectral data of the crop sample to be measured.

5. The method for predicting the content of crop quality indicators according to claim 1, characterized in that, The training process of the crop quality indicator content prediction model specifically includes: Construct a training set; the training set includes the fused feature vectors of multiple sample crop samples and the true values of the free amino acid content, starch content, soluble sugar content and vitamin C content of the sample crops; Construct a regression model; Use the fused feature vector of the sample crop sample as the input and the true value of the free amino acid content, starch content, soluble sugar content or vitamin C content of the sample crop as the output to train each regression model to obtain the corresponding crop quality indicator content prediction model.

6. The method for predicting the content of crop quality indicators according to claim 5, characterized in that, The regression model is a ridge regression model, a random forest regression model, a support vector machine regression model, or a fully connected neural network model.

7. A prediction system for the content of crop quality indicators, characterized in that, The crop quality index content prediction system is used to implement the crop quality index content prediction method according to any one of claims 1-6. The crop quality index content prediction system includes: A sample acquisition unit for acquiring crop samples to be measured; A data acquisition unit for acquiring the near-infrared spectrum data and conductivity data of the crop samples to be measured; A preprocessing unit for preprocessing the near-infrared spectrum data of the crop samples to be measured and the conductivity data of the crop samples to be measured respectively, to obtain the preprocessed near-infrared spectrum data and preprocessed conductivity data of the crop samples to be measured; A characteristic band screening unit for screening the characteristic bands of the preprocessed near-infrared spectrum data of the crop samples to be measured, to obtain the screened near-infrared spectrum data of the crop samples to be measured; A fusion unit for fusing the screened near-infrared spectrum data and the preprocessed conductivity data of the crop samples to be measured, to obtain the fusion feature vector of the crop samples to be measured; A prediction unit for inputting the fusion feature vector of the crop samples to be measured into each crop quality index content prediction model respectively, to obtain the predicted values of the corresponding crop quality index contents; the crop quality index contents are free amino acid content, starch content, soluble sugar content, or vitamin C content; each of the crop quality index content prediction models is obtained by training each regression model using a training set.

8. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the crop quality index content prediction method according to any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the crop quality index content prediction method according to any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the crop quality index content prediction method according to any one of claims 1-6.

Citation Information

Cited By

  • Characteristic spectrum analysis-based identification method for flavonoid-containing compounds of elm branches and leaves

    CN121528379A