A rapeseed lysine content detection method, system, device and medium
By constructing a target lysine content prediction model through spectral data processing, the problems of low detection efficiency and low accuracy of traditional chemical methods are solved, enabling rapid, non-destructive, and accurate detection of lysine content in rapeseed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN INST OF TECH
- Filing Date
- 2023-07-12
- Publication Date
- 2026-07-24
AI Technical Summary
Traditional chemical methods for detecting lysine content in rapeseed are cumbersome, inefficient, and inaccurate, failing to meet the need for rapid and non-destructive testing.
A target lysine content prediction model based on spectral data was adopted. By removing outliers and screening wavelength points, the target lysine content prediction model was constructed, and spectral data was obtained using near-infrared spectroscopy for prediction.
It enables rapid, non-destructive, and accurate detection of lysine content in rapeseed, and is simple to operate, low in cost, and highly efficient, avoiding the use of chemical reagents.
Smart Images

Figure CN117110231B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of amino acid detection technology, specifically to a method, system, equipment, and medium for detecting lysine content in rapeseed. Background Technology
[0002] Rapeseed is an important oil crop in my country, distributed in the Yangtze River Basin, Qinghai, Inner Mongolia and other regions. At present, rapeseed is the most important source of edible vegetable oil in my country, and its planting area and total output account for about 30% of the world's rapeseed planting area and total output.
[0003] Amino acid composition is one of the important indicators determining the nutritional value of rapeseed. Lysine is one of the essential amino acids for humans and mammals. Lysine has positive nutritional significance in promoting human growth and development, enhancing immunity, antiviral activity, promoting lipid oxidation, and alleviating anxiety. It can also promote the absorption of certain nutrients and work synergistically with other nutrients to better exert their physiological functions. Therefore, how to quickly identify and determine the lysine content in rapeseed has become an important issue in the field of food safety.
[0004] Currently, traditional chemical methods are commonly used for oil quality testing. However, these methods often require chemical reagents, which are cumbersome, time-consuming, and generally costly. They also fail to meet the need for rapid and non-destructive on-site testing. Summary of the Invention
[0005] The technical problem this invention aims to solve is that traditional chemical methods for detecting lysine content in rapeseed are cumbersome, inefficient, and have low accuracy. To address this problem, this invention provides a method, system, equipment, and medium for detecting lysine content in rapeseed.
[0006] The technical solution of the present invention to solve the above-mentioned technical problems is as follows:
[0007] A method for detecting lysine content in rapeseed includes the following steps:
[0008] Obtain the target spectral data corresponding to the rapeseed to be tested;
[0009] The target spectral data is input into a pre-trained target lysine content prediction model, and the target lysine content prediction model outputs the lysine content in the rapeseed to be detected.
[0010] The target lysine content prediction model is obtained through the following steps:
[0011] Step 1: Obtain a sample set for rapeseed, and divide the sample set into an original training set and an original test set; wherein, the sample set includes the lysine content and corresponding original spectral data of multiple rapeseed samples;
[0012] Step 2: Remove outliers from the original training set and the original test set to obtain a first training set and a first test set; wherein the first training set includes multiple first training data and the first test set includes multiple first test data.
[0013] Step 3: Filter wavelength points according to the relevant parameters corresponding to each first training data to obtain the second training set; Filter wavelength points according to the relevant parameters corresponding to each first test data to obtain the second test set;
[0014] Step 4: Determine the target lysine content prediction model based on the second training set, the second test set, and the pre-constructed initial lysine content prediction model.
[0015] To address the aforementioned technical problems, the present invention also provides a rapeseed lysine content detection system, comprising:
[0016] The target data acquisition module is used to acquire the target spectral data corresponding to the rapeseed to be tested; the content detection module is used to input the target spectral data into a pre-trained target lysine content prediction model, and output the lysine content in the rapeseed to be tested through the target lysine content prediction model.
[0017] A prediction model determination module is used to determine the prediction model for the target lysine content;
[0018] The prediction model determination module includes:
[0019] The sample data acquisition module is used to acquire a sample set for rapeseed, and divide the sample set into an original training set and an original test set; wherein, the sample set includes the lysine content and corresponding original spectral data of multiple rapeseed samples;
[0020] An outlier removal module is used to remove outliers from the original training set and the original test set respectively to obtain a first training set and a first test set; wherein, the first training set includes multiple first training data and the first test set includes multiple first test data;
[0021] The wavelength point filtering module is used to filter wavelength points according to the relevant parameters corresponding to each first training data to obtain a second training set; and to filter wavelength points according to the relevant parameters corresponding to each first test data to obtain a second test set.
[0022] The model determination module is used to determine the target lysine content prediction model based on the second training set, the second test set, and the pre-constructed initial lysine content prediction model.
[0023] To address the aforementioned technical problems, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the rapeseed lysine content detection method as described above.
[0024] To address the aforementioned technical problems, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the method for detecting lysine content in rapeseed as described above.
[0025] The beneficial effects of this invention are as follows: Based on relevant parameters, wavelength points are screened in the spectral data after outlier removal to overcome the problem of low prediction accuracy caused by multicollinearity among wavelengths in the spectral data. This allows the trained target lysine content prediction model to better utilize wavelength points that have a positive effect on lysine content prediction and reduce the negative impact of other redundant wavelength points on lysine content prediction, thereby improving the prediction performance of the target lysine content prediction model. This invention provides a green, non-destructive, and rapid detection method that uses a trained target lysine content prediction model to predict the lysine content in rapeseed to be tested. It requires no chemical reagents, is environmentally friendly, and has the advantages of simple operation, low detection cost, high detection efficiency, and high detection accuracy. Attached Figure Description
[0026] Figure 1 This is a schematic flowchart of the method for detecting lysine content in rapeseed in this invention.
[0027] Figure 2 This is a schematic diagram of the process for determining the target lysine content prediction model in this invention. Detailed Implementation
[0028] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.
[0029] Example 1
[0030] This embodiment provides a method for detecting the lysine content in rapeseed, such as... Figure 1 and Figure 2 As shown, it includes the following steps:
[0031] S1, Obtain the target spectral data corresponding to the rapeseed to be detected;
[0032] S2, input the target spectral data into the pre-trained target lysine content prediction model, and output the lysine content in the rapeseed to be detected through the target lysine content prediction model;
[0033] The target lysine content prediction model is obtained through the following steps:
[0034] Step 1: Obtain a sample set for rapeseed, and divide the sample set into an original training set and an original test set; wherein, the sample set includes the lysine content and corresponding original spectral data of multiple rapeseed samples;
[0035] Step 2: Remove outliers from the original training set and the original test set to obtain a first training set and a first test set; wherein the first training set includes multiple first training data and the first test set includes multiple first test data.
[0036] Step 3: Filter wavelength points according to the relevant parameters corresponding to each first training data to obtain the second training set; Filter wavelength points according to the relevant parameters corresponding to each first test data to obtain the second test set;
[0037] Step 4: Determine the target lysine content prediction model based on the second training set, the second test set, and the pre-constructed initial lysine content prediction model.
[0038] This invention overcomes the problem of low prediction accuracy caused by multicollinearity among wavelengths in spectral data by screening wavelength points after outlier removal. This allows the trained target lysine content prediction model to better utilize wavelength points that have a positive effect on lysine content prediction and reduce the negative impact of other redundant wavelength points on lysine content prediction, thereby improving the prediction performance of the target lysine content prediction model. For rapeseed to be tested, the target spectral data corresponding to the rapeseed to be tested is obtained, and the target lysine content prediction model is used to predict the lysine content in the rapeseed to be tested based on the target spectral data. This method requires no chemical reagents, is environmentally friendly, and has the advantages of simple operation, low detection cost, high detection efficiency, and high detection accuracy.
[0039] In this embodiment, in step 1, the sample set is obtained through the following steps:
[0040] Multiple rapeseed samples were obtained;
[0041] For each rapeseed sample, the rapeseed sample was dried and cooled to obtain a rapeseed sample, and the lysine content in the rapeseed sample was determined.
[0042] For each rapeseed sample, near-infrared spectral scanning is performed on the rapeseed sample to obtain the original spectral data corresponding to the rapeseed sample;
[0043] The sample set is obtained based on the original spectral data.
[0044] In this embodiment, the rapeseed sample and the rapeseed to be tested are scanned using a near-infrared spectrometer. For acquiring the sample set, this embodiment uses near-infrared spectroscopy to obtain rich information such as the intensity and position differences of the infrared absorption peaks of the rapeseed. Subsequently, by extracting the required feature information, the content of the substance (i.e., lysine) is predicted. This method has the advantages of fast analysis speed, high analysis efficiency, and low analysis cost, facilitating online analysis and enabling widespread application in product quality testing. It avoids the problems of traditional chemical methods, which require chemical reagents, are cumbersome to operate, time-consuming, and cannot meet the needs of rapid and non-destructive on-site testing.
[0045] In this embodiment, for each rapeseed sample, the near-infrared spectral scanning of the rapeseed sample to obtain the corresponding raw spectral data includes the following steps:
[0046] The rapeseed sample is scanned by near-infrared spectroscopy according to each preset scanning position to obtain multiple near-infrared spectra of the rapeseed sample. Each near-infrared spectra has multiple wavelength points. The horizontal axis of the near-infrared spectra is wavelength, the vertical axis is absorbance, and each wavelength point is the absorbance value corresponding to a wavelength value.
[0047] The average spectrum is obtained by averaging the various near-infrared spectra.
[0048] Based on the average spectrum, the original spectral data corresponding to the rapeseed sample is determined. The original spectral data is a set obtained by sorting each wavelength point in the average spectrum according to its corresponding wavelength value from smallest to largest.
[0049] In this embodiment, for each rapeseed sample, the process of drying and cooling the rapeseed sample includes: placing the rapeseed sample in an oven to dry it, and obtaining the rapeseed sample after it has cooled down; the preset scanning locations are 6.
[0050] In this embodiment of the invention, the original spectral data corresponding to the rapeseed sample is determined based on the average spectral diagram, which better reflects the actual component information in the rapeseed sample and facilitates the improvement of subsequent training results.
[0051] The original training set includes multiple first original spectral data, each of which is one original spectral data; the original test set includes multiple second original spectral data, each of which is one original spectral data.
[0052] In this embodiment, step 2, which involves removing outliers from the original training set to obtain the first training set, includes the following steps:
[0053] Step 2.1: For each of the first original spectral data, input the first original spectral data into the pre-established partial least squares model (hereinafter referred to as PLS model), and output the squared prediction error (hereinafter referred to as Q statistic) and Hotelling statistic (hereinafter referred to as T2 statistic) corresponding to the first original spectral data through the PLS model.
[0054] Step 2.2: For each of the first original spectral data, determine whether the rapeseed sample corresponding to the first original spectral data is an outlier sample based on the preset error threshold, the preset statistical threshold, the Q statistic and T2 statistic corresponding to the first original spectral data.
[0055] Step 2.3: Remove the first original spectral data of each corresponding rapeseed sample that is an outlier from the original training set to obtain the updated training set;
[0056] Step 2.4: Adjust the model parameters of the PLS model, use the updated training set as the original training set, and repeat steps 2.1 to 2.3 until all rapeseed samples corresponding to the first original spectral data in the original training set are no longer outliers, thus obtaining the first training set.
[0057] In this embodiment, 10-fold cross-validation is used to obtain a reliable and stable PLS model. The PLS model is used to predict the lysine content in rapeseed samples corresponding to the spectral data input into the PLS model.
[0058] In this embodiment, for the original training set, the Q-statistics and T2 statistics corresponding to each of the first original spectral data are plotted on a pre-established two-dimensional graph using MATLAB software, and outlier samples are removed using MATLAB software. Specifically, for each of the first original spectral data, if the Q-statistic corresponding to the first original spectral data is greater than or equal to the preset error threshold and / or the T2 statistic corresponding to the first original spectral data is greater than or equal to the preset statistical threshold, then the rapeseed sample corresponding to the first original spectral data is an outlier sample; if the Q-statistic corresponding to the first original spectral data is less than the preset error threshold and the T2 statistic corresponding to the first original spectral data is less than the preset statistical threshold, then the rapeseed sample corresponding to the first original spectral data is not an outlier sample.
[0059] In this embodiment, each first training data is a first original spectral data, and each first test data is a second original spectral data. The method of obtaining the first test set based on the original test set is similar to the method of obtaining the first training set based on the original training set, and will not be described in detail here.
[0060] During data processing, outlier observations (i.e., some original spectral data in the sample set) are often encountered. To address this, this invention calculates the Q-statistic and T2-statistic using the PLS model, and then defines a standard (i.e., a method for determining whether a rapeseed sample is an outlier based on a preset error threshold and a preset statistic threshold). Based on this standard, it determines whether each piece of original spectral data is an outlier, thereby identifying whether the corresponding rapeseed sample is an outlier and thus removing outliers. This invention ensures the prediction accuracy of the PLS model by detecting and filtering out observations that the PLS model cannot interpret well (specifically, original spectral data with large Q-statistic values) and observations far from the center of normal observations (specifically, original spectral data with large T2-statistic values).
[0061] In this embodiment, the relevant parameters include correlation coefficient, mean, and standard deviation, and the correlation coefficient includes a first correlation coefficient and a second correlation coefficient;
[0062] Step 3, which involves filtering wavelength points based on the relevant parameters corresponding to each of the first training data to obtain the second training set, includes the following steps:
[0063] Each of the first training data is mapped into a matrix to obtain a two-dimensional spectral matrix; wherein each row of the two-dimensional spectral matrix corresponds to one of the first training data, and each element of the two-dimensional spectral matrix is a wavelength point of the first training data;
[0064] For each column of data in the two-dimensional spectral matrix, the correlation coefficient between the data in that column and the data in the first column of the two-dimensional spectral matrix is taken as the first correlation coefficient corresponding to that column of data.
[0065] The column data corresponding to the first correlation coefficient that is less than the preset minimum correlation threshold are deleted from the two-dimensional spectral matrix to obtain the first spectral matrix;
[0066] For each column of data in the first spectral matrix, multiple second correlation coefficients are determined based on the correlation coefficient between the data in that column and each column in the first spectral matrix.
[0067] For each column of data in the first spectral matrix, the mean and standard deviation of the data in that column are determined based on the corresponding second correlation coefficients; wherein the mean is the average of the second correlation coefficients, and the standard deviation is the root mean square deviation of the second correlation coefficients.
[0068] Delete column data with a mean greater than a preset mean and / or a standard deviation greater than a preset standard deviation from the first spectral matrix to obtain the second spectral matrix;
[0069] For each row of data in the second spectral matrix, a second training data is formed based on the elements of that row;
[0070] The second training set is obtained based on each of the second training data.
[0071] In this embodiment, the second training set includes multiple second training data, and the second test set includes multiple second test data. The method of obtaining the second training set based on the first training set is similar to the method of obtaining the second test set based on the first test set, and will not be described in detail here.
[0072] In this embodiment, the correlation coefficient corresponding to each column of data in the matrix is calculated using the Pearson correlation coefficient; for each column of data in the first spectral matrix, the correlation coefficient between the data in that column and each column in the first spectral matrix is the multiple second correlation coefficients corresponding to that column of data.
[0073] In this embodiment, based on the correlation coefficient, a two-stage spectral screening is performed on the first training set and the first test set respectively, so that the final target lysine content prediction model can efficiently extract spectral information, thereby improving the prediction accuracy of the model. Taking the first training set as an example, the first stage of spectral screening is first performed: by calculating the first correlation coefficient between the data in each column of the two-dimensional spectral matrix and the physicochemical values (i.e., the data in the first column of the two-dimensional spectral matrix), the screening result of the first stage (i.e., the first spectral matrix) determined according to the preset minimum correlation threshold and each of the first correlation coefficients ensures that the selected column data and the physicochemical values have a high correlation. Then, based on the screening result of the first stage, the second stage of spectral screening is performed: for wavelengths, if the correlation coefficient between a wavelength and other wavelengths is large, it means that there is multicollinearity between the wavelength and other wavelengths. Therefore, based on the first spectral matrix, according to the preset mean threshold, the preset standard deviation threshold, and the mean and standard deviation calculated based on the correlation coefficient, the wavelengths in the second stage of screening result (i.e., the second spectral matrix) have low linear correlation, overcoming the problem of low prediction accuracy caused by multicollinearity between wavelengths in spectral data.
[0074] In this embodiment, step 4 includes the following steps:
[0075] Step 4.1: Train the pre-constructed initial lysine content prediction model using the second training set to obtain the intermediate lysine content prediction model;
[0076] Step 4.2: For each of the second test data, input the second test data into the intermediate lysine content prediction model to obtain the prediction result corresponding to the second test data;
[0077] Step 4.3: For each second test data, determine the prediction error corresponding to the second test data based on the label and prediction result; wherein, the label corresponding to the second test data is the lysine content in the rapeseed sample corresponding to the second test data;
[0078] Step 4.4: If each prediction error is within the prediction error range, the intermediate lysine content prediction model is determined as the target lysine content prediction model; if there is a corresponding prediction error that is not within the prediction error range, the intermediate lysine content prediction model is parameter-tuned to obtain an optimized lysine content prediction model, and the optimized lysine content prediction model is used as the initial lysine content prediction model, returning to step 4.1.
[0079] In this embodiment, the initial lysine content prediction model adopts an existing one-dimensional convolutional neural network. The initial lysine content prediction model includes an input layer, a convolutional layer, a pooling layer, a first fully connected layer, a second fully connected layer, and an output layer connected in sequence. The output layer has two neurons, which correspond to the two categories of high lysine content and low lysine content. In order to reduce the complexity of the model, the initial lysine content prediction model uses only one convolutional layer, which includes eight convolutional kernels.
[0080] For each of the second training data, the second training data and the corresponding label (i.e., the lysine content in the rapeseed sample corresponding to the second training data) are input into the initial lysine content prediction model. The convolutional layer performs convolution operations through convolution kernels with pre-set sizes and strides to output multiple feature maps. The pooling layer uses max pooling to downsample the output of the convolutional layer to extract its local features. The first fully connected layer and the second fully connected layer map the features to the sample space for classification (dividing them into the high lysine content category or the low lysine content category). The output layer normalizes its input through the Softmax function and outputs the prediction result corresponding to the second training data (i.e., the predicted lysine content in the rapeseed sample corresponding to the second training data).
[0081] In this embodiment, before training the initial lysine content prediction model, its weights are pre-initialized. For each piece of second training data, after inputting the second training data into the initial lysine content prediction model, forward propagation is performed to obtain the prediction result output by the initial lysine content prediction model. By calculating the loss function value between the prediction result output by forward propagation and the label corresponding to the second training data, the loss function value is backpropagated to adjust the weights of the initial lysine content prediction model using a method that minimizes the error, thereby obtaining the intermediate lysine content prediction model. During model training, continuously adjusting the model's weights improves the efficiency of determining the lysine content prediction model.
[0082] Preferably, regularization terms and random deactivation are added to both the first and second fully connected layers in the initial lysine content prediction model to avoid overfitting as much as possible; an activation layer is set between the convolutional layer and the pooling layer in the initial lysine content prediction model, and the activation layer uses the ReLU function as the activation function to avoid the gradient vanishing problem.
[0083] Preferably, the method further includes:
[0084] Each of the first training data in the first training set is preprocessed to obtain multiple first preprocessed data.
[0085] Each of the first test data in the first test set is preprocessed to obtain multiple second preprocessed data.
[0086] Step 3 is as follows:
[0087] Wavelength points are selected based on the relevant parameters corresponding to each of the first preprocessed data to obtain the second training set; wavelength points are selected based on the relevant parameters corresponding to each of the second preprocessed data to obtain the second test set.
[0088] In this embodiment, the preprocessing method can be at least one of convolutional smoothing, baseline correction, multivariate scattering correction, standard normal transformation, and normalization. Since absorbance data (i.e., the original spectral data) has both absorption and scattering characteristics, this method enhances the spectral data features by preprocessing each of the first training data and each of the second training data, thereby improving the efficiency and accuracy of obtaining the target lysine content prediction model. For example, standard normal transformation is applied to the target data (i.e., the first training data and the first test data) to reduce spectral errors caused by scattering between the rapeseed samples; baseline correction is applied to the target data to eliminate baseline drift and improve spectral resolution; and normalization is applied to the target data to enhance spectral absorption characteristics. For the same target data, when multiple preprocessing methods are used, the preprocessed data are superimposed based on the complementarity of each method to enhance spectral data features, resulting in a target lysine content prediction model with strong generalization ability.
[0089] In this embodiment, the method of obtaining the second training set based on the first training set is similar to the method of obtaining the second training set based on each of the first preprocessed data and the method of obtaining the second test set based on each of the second preprocessed data, and will not be described in detail here.
[0090] Example 2
[0091] Based on the same principle as the rapeseed lysine content detection method described in Embodiment 1 above, this embodiment provides a rapeseed lysine content detection system, comprising:
[0092] The target data acquisition module is used to acquire the target spectral data corresponding to the rapeseed to be tested;
[0093] The content detection module is used to input the target spectral data into a pre-trained target lysine content prediction model, and output the lysine content in the rapeseed to be detected through the target lysine content prediction model;
[0094] A prediction model determination module is used to determine the prediction model for the target lysine content;
[0095] The prediction model determination module includes:
[0096] The sample data acquisition module is used to acquire a sample set for rapeseed, and divide the sample set into an original training set and an original test set; wherein, the sample set includes the lysine content and corresponding original spectral data of multiple rapeseed samples;
[0097] An outlier removal module is used to remove outliers from the original training set and the original test set respectively to obtain a first training set and a first test set; wherein, the first training set includes multiple first training data and the first test set includes multiple first test data;
[0098] The wavelength point filtering module is used to filter wavelength points according to the relevant parameters corresponding to each first training data to obtain a second training set; and to filter wavelength points according to the relevant parameters corresponding to each first test data to obtain a second test set.
[0099] The model determination module is used to determine the target lysine content prediction model based on the second training set, the second test set, and the pre-constructed initial lysine content prediction model.
[0100] Specifically, the sample data acquisition module is used to acquire a sample set of rapeseed, and is used for:
[0101] Multiple rapeseed samples were obtained;
[0102] For each rapeseed sample, the rapeseed sample was dried and cooled to obtain a rapeseed sample, and the lysine content in the rapeseed sample was determined.
[0103] For each rapeseed sample, near-infrared spectral scanning is performed on the rapeseed sample to obtain the original spectral data corresponding to the rapeseed sample;
[0104] The sample set is obtained based on the original spectral data.
[0105] Specifically, for each rapeseed sample, the sample data acquisition module performs near-infrared spectral scanning on the rapeseed sample to obtain the original spectral data corresponding to the rapeseed sample, and is used for:
[0106] The rapeseed sample is scanned by near-infrared spectroscopy according to each preset scanning location to obtain multiple near-infrared spectra of the rapeseed sample. Each near-infrared spectra has multiple wavelength points. The horizontal axis of the near-infrared spectra is wavelength, and the vertical axis is absorbance.
[0107] The average spectrum is obtained by averaging the various near-infrared spectra.
[0108] Based on the average spectrum, the original spectral data corresponding to the rapeseed sample is determined. The original spectral data is a set obtained by sorting each wavelength point in the average spectrum according to its corresponding wavelength value from smallest to largest.
[0109] The original training set includes multiple first original spectral data;
[0110] The outlier removal module is used to remove outliers from the original training set to obtain the first training set, including:
[0111] The parameter value acquisition unit is used to input the first original spectral data into a pre-established partial least squares model for each first original spectral data, and output the squared prediction error and Hotelling statistic corresponding to the first original spectral data through the partial least squares model.
[0112] The outlier sample determination unit is used to determine whether the rapeseed sample corresponding to each of the first original spectral data is an outlier sample based on a preset error threshold, a preset statistic threshold, the squared prediction error corresponding to the first original spectral data, and the Hotling statistic.
[0113] The training set update unit is used to remove the first original spectral data of each corresponding rapeseed sample that is an outlier from the original training set to obtain an updated training set.
[0114] The original training set determination unit is used to adjust the model parameters of the partial least squares model, and use the updated training set as the original training set. The parameter value acquisition unit, the outlier sample determination unit, and the training set update unit are executed sequentially until the rapeseed samples corresponding to each first original spectral data in the original training set are no longer outliers, thus obtaining the first training set.
[0115] The relevant parameters include correlation coefficient, mean, and standard deviation, and the correlation coefficient includes a first correlation coefficient and a second correlation coefficient.
[0116] The wavelength point filtering module is used to filter wavelength points according to the relevant parameters corresponding to each first training data set. When obtaining the second training set, it is specifically used for:
[0117] Each of the first training data is mapped into a matrix to obtain a two-dimensional spectral matrix; wherein each row of the two-dimensional spectral matrix corresponds to one of the first training data, and each element of the two-dimensional spectral matrix is a wavelength point of the first training data;
[0118] For each column of data in the two-dimensional spectral matrix, the correlation coefficient between the data in that column and the data in the first column of the two-dimensional spectral matrix is taken as the first correlation coefficient corresponding to that column of data.
[0119] The column data corresponding to the first correlation coefficient that is less than the preset minimum correlation threshold are deleted from the two-dimensional spectral matrix to obtain the first spectral matrix;
[0120] For each column of data in the first spectral matrix, multiple second correlation coefficients are determined based on the correlation coefficient between the data in that column and each column in the first spectral matrix.
[0121] For each column of data in the first spectral matrix, the mean and standard deviation of the data in that column are determined based on the corresponding second correlation coefficients; wherein the mean is the average of the second correlation coefficients, and the standard deviation is the root mean square deviation of the second correlation coefficients.
[0122] Delete column data with a mean greater than a preset mean and / or a standard deviation greater than a preset standard deviation from the first spectral matrix to obtain the second spectral matrix;
[0123] For each row of data in the second spectral matrix, a second training data is formed based on the elements of that row;
[0124] The second training set is obtained based on each of the second training data.
[0125] The model determination module includes:
[0126] The model training unit is used to train the pre-constructed initial lysine content prediction model using the second training set to obtain the intermediate lysine content prediction model.
[0127] The model prediction unit is used to input the second test data into the intermediate lysine content prediction model for each second test data to obtain the prediction result corresponding to the second test data.
[0128] The prediction error determination unit is used to determine the prediction error corresponding to each second test data according to the label corresponding to the second test data and the prediction result; wherein, the label corresponding to the second test data is the lysine content in the rapeseed sample corresponding to the second test data;
[0129] The model determination unit is used to determine the intermediate lysine content prediction model as the target lysine content prediction model if each prediction error is within the prediction error range; if there is a corresponding prediction error that is not within the prediction error range, the intermediate lysine content prediction model is parameter-tuned to obtain an optimized lysine content prediction model, and the optimized lysine content prediction model is used as the initial lysine content prediction model. The model training unit, the model prediction unit, the prediction error determination unit, and the model determination unit are executed sequentially.
[0130] Preferably, the system further includes a data preprocessing module, the data preprocessing module being used for:
[0131] Each of the first training data in the first training set is preprocessed to obtain multiple first preprocessed data.
[0132] Each of the first test data in the first test set is preprocessed to obtain multiple second preprocessed data.
[0133] The wavelength point filtering module is used to filter wavelength points according to the relevant parameters corresponding to each of the first preprocessed data to obtain a second training set; and to filter wavelength points according to the relevant parameters corresponding to each of the second preprocessed data to obtain a second test set.
[0134] Example 3
[0135] To address the aforementioned technical problems, this embodiment provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the rapeseed lysine content detection method as described in Embodiment 1.
[0136] Example 4
[0137] To address the aforementioned technical problems, this embodiment provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the rapeseed lysine content detection method as described in Embodiment 1.
[0138] In the description of this invention, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0139] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," and "some examples" indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0140] Furthermore, the functional units in the embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0141] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for detecting lysine content in rapeseed, characterized in that, Includes the following steps: Obtain the target spectral data corresponding to the rapeseed to be tested; The target spectral data is input into a pre-trained target lysine content prediction model, and the target lysine content prediction model outputs the lysine content in the rapeseed to be detected. The target lysine content prediction model is a one-dimensional convolutional neural network model, which is obtained through the following steps: Step 1: Obtain a sample set for rapeseed, and divide the sample set into an original training set and an original test set; wherein, the sample set includes the lysine content and corresponding original spectral data of multiple rapeseed samples; Step 2: Remove outliers from the original training set and the original test set to obtain a first training set and a first test set; wherein the first training set includes multiple first training data and the first test set includes multiple first test data. Step 3: Filter wavelength points according to the relevant parameters corresponding to each first training data to obtain the second training set; Filter wavelength points according to the relevant parameters corresponding to each first test data to obtain the second test set; The step of filtering wavelength points based on the relevant parameters corresponding to each of the first training data to obtain the second training set specifically includes: Each of the first training data is mapped into a matrix to obtain a two-dimensional spectral matrix; wherein each row of the two-dimensional spectral matrix corresponds to one of the first training data, and each element of the two-dimensional spectral matrix is a wavelength point of the first training data; For each column of data in the two-dimensional spectral matrix, the correlation coefficient between the data in that column and the data in the first column of the two-dimensional spectral matrix is taken as the first correlation coefficient corresponding to that column of data; the column of data corresponding to the first correlation coefficient that is less than a preset minimum correlation threshold is deleted from the two-dimensional spectral matrix to obtain the first spectral matrix. For each column of data in the first spectral matrix, based on the correlation coefficient between the data in that column and each column in the first spectral matrix, multiple second correlation coefficients are determined for that column of data; based on each of the second correlation coefficients, the mean and standard deviation of that column of data are determined; wherein, the mean is the average of each of the second correlation coefficients, and the standard deviation is the root mean square deviation of each of the second correlation coefficients; columns with a mean greater than a preset mean and / or a standard deviation greater than a preset standard deviation are deleted from the first spectral matrix to obtain the second spectral matrix; For each row of data in the second spectral matrix, a second training data is formed based on the elements in that row; based on each of the second training data, the second training set is obtained. Step 4: Determine the target lysine content prediction model based on the second training set, the second test set, and the pre-constructed initial one-dimensional convolutional neural network model; The second test set includes multiple second test data sets.
2. The method according to claim 1, characterized in that, In step 1, the sample set is obtained through the following steps: Multiple rapeseed samples were obtained; For each rapeseed sample, the rapeseed sample was dried and cooled to obtain a rapeseed sample, and the lysine content in the rapeseed sample was determined. For each rapeseed sample, near-infrared spectral scanning is performed on the rapeseed sample to obtain the original spectral data corresponding to the rapeseed sample; The sample set is obtained based on the original spectral data.
3. The method according to claim 2, characterized in that, For each rapeseed sample, performing near-infrared spectral scanning on the rapeseed sample to obtain the corresponding raw spectral data includes the following steps: The rapeseed sample is scanned by near-infrared spectroscopy according to each preset scanning location to obtain multiple near-infrared spectra of the rapeseed sample. Each near-infrared spectra has multiple wavelength points. The horizontal axis of the near-infrared spectra is wavelength, and the vertical axis is absorbance. The average spectrum is obtained by averaging the various near-infrared spectra. Based on the average spectrum, the original spectral data corresponding to the rapeseed sample is determined. The original spectral data is a set obtained by sorting each wavelength point in the average spectrum according to its corresponding wavelength value from smallest to largest.
4. The method according to claim 1, characterized in that, The original training set includes multiple first original spectral data; Step 2 involves removing outliers from the original training set to obtain the first training set, and includes the following steps: Step 2.1: For each of the first original spectral data, input the first original spectral data into the pre-established partial least squares model, and output the squared prediction error and Hotelling statistic corresponding to the first original spectral data through the partial least squares model; Step 2.2: For each of the first original spectral data, determine whether the rapeseed sample corresponding to the first original spectral data is an outlier sample based on the preset error threshold, the preset statistic threshold, the squared prediction error corresponding to the first original spectral data, and the Hotelling statistic. Step 2.3: Remove the first original spectral data of each corresponding rapeseed sample that is an outlier from the original training set to obtain the updated training set; Step 2.4: Adjust the model parameters of the partial least squares model, use the updated training set as the original training set, and repeat steps 2.1 to 2.3 until all rapeseed samples corresponding to the first original spectral data in the original training set are no longer outliers, thus obtaining the first training set.
5. The method according to claim 1, characterized in that, Step 4 includes the following steps: Step 4.1: Train the initial one-dimensional convolutional neural network model using the second training set to obtain the intermediate lysine content prediction model; Step 4.2: For each of the second test data, input the second test data into the intermediate lysine content prediction model to obtain the prediction result corresponding to the second test data; Step 4.3: For each second test data, determine the prediction error corresponding to the second test data based on the label and prediction result; wherein, the label corresponding to the second test data is the lysine content in the rapeseed sample corresponding to the second test data; Step 4.4: If each prediction error is within the prediction error range, the intermediate lysine content prediction model is determined as the target lysine content prediction model; if there is a corresponding prediction error that is not within the prediction error range, the intermediate lysine content prediction model is parameter-tuned to obtain an optimized lysine content prediction model, and the optimized lysine content prediction model is used as the initial one-dimensional convolutional neural network model, returning to step 4.
1.
6. The method according to claim 1, characterized in that, The method further includes: Each of the first training data in the first training set is preprocessed to obtain multiple first preprocessed data. Each of the first test data in the first test set is preprocessed to obtain multiple second preprocessed data. Step 3 is as follows: Wavelength points are selected based on the relevant parameters corresponding to each of the first preprocessed data to obtain the second training set; wavelength points are selected based on the relevant parameters corresponding to each of the second preprocessed data to obtain the second test set.
7. A rapeseed lysine content detection system, characterized in that, include: The target data acquisition module is used to acquire the target spectral data corresponding to the rapeseed to be tested; The content detection module is used to input the target spectral data into a pre-trained target lysine content prediction model, and output the lysine content in the rapeseed to be detected through the target lysine content prediction model; A prediction model determination module is used to determine the prediction model for the target lysine content; The prediction model determination module includes: The sample data acquisition module is used to acquire a sample set for rapeseed, and divide the sample set into an original training set and an original test set; wherein, the sample set includes the lysine content and corresponding original spectral data of multiple rapeseed samples; An outlier removal module is used to remove outliers from the original training set and the original test set respectively to obtain a first training set and a first test set; wherein, the first training set includes multiple first training data and the first test set includes multiple first test data; The wavelength point filtering module is used to filter wavelength points according to the relevant parameters corresponding to each first training data to obtain a second training set; and to filter wavelength points according to the relevant parameters corresponding to each first test data to obtain a second test set. The wavelength point filtering module is used to filter wavelength points according to the relevant parameters corresponding to each first training data to obtain a second training set, specifically including: Each of the first training data is mapped into a matrix to obtain a two-dimensional spectral matrix; wherein each row of the two-dimensional spectral matrix corresponds to one of the first training data, and each element of the two-dimensional spectral matrix is a wavelength point of the first training data; For each column of data in the two-dimensional spectral matrix, the correlation coefficient between the data in that column and the data in the first column of the two-dimensional spectral matrix is taken as the first correlation coefficient corresponding to that column of data; the column of data corresponding to the first correlation coefficient that is less than a preset minimum correlation threshold is deleted from the two-dimensional spectral matrix to obtain the first spectral matrix. For each column of data in the first spectral matrix, based on the correlation coefficient between the data in that column and each column in the first spectral matrix, multiple second correlation coefficients are determined for that column of data; based on each of the second correlation coefficients, the mean and standard deviation of that column of data are determined; wherein, the mean is the average of each of the second correlation coefficients, and the standard deviation is the root mean square deviation of each of the second correlation coefficients; columns with a mean greater than a preset mean and / or a standard deviation greater than a preset standard deviation are deleted from the first spectral matrix to obtain the second spectral matrix; For each row of data in the second spectral matrix, a second training data is formed based on the elements in that row; based on each of the second training data, the second training set is obtained. The model determination module is used to determine the target lysine content prediction model based on the second training set, the second test set, and the pre-constructed initial lysine content prediction model.
8. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for detecting lysine content in rapeseed as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method for detecting lysine content in rapeseed as described in any one of claims 1 to 6.