Construction method and application of rice seed single-particle activity detection model
Through hyperspectral imaging technology and multi-level data fusion method, a single-particle vitality detection model of rice seeds was constructed, solving the problems of complex, time-consuming and low accuracy of existing detection methods, and achieving fast and accurate rice seed vitality detection.
Patent Information
- Application Number
- CN202411861329.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-05-23
AI Technical Summary
The existing rice seed vitality detection methods are complex, time-consuming and low-precision, making it difficult to efficiently identify the vitality of a large number of rice seeds.
The rice seed single-particle vitality detection model is constructed by combining hyperspectral imaging technology with multi-level data fusion method. The model includes steps such as data preprocessing, feature variable selection, machine learning modeling and model optimization.
The rapid, accurate and non-damaging samples of rice seed vitality detection is achieved, which improves the detection efficiency and accuracy, and can simultaneously efficiently identify the vitality of a large number of rice seeds.
Smart Images

Figure CN120030867A_ABST
Abstract
Description
Technical Field
[0001] The invention relates to the technical field of crop authenticity identification, and in particular to a method for constructing a rice seed single-grain vitality detection model and its application. Background Art
[0002] Rice is one of the most important food crops in the world, feeding nearly half of the world's population. As a major food crop, rice production directly affects global food security and economic stability. Rice seed quality is the basis for determining the efficiency of rice production, and seed vigor is one of the core indicators of seed quality. Seed vigor testing is to determine the vitality, germination rate and growth potential of seeds under adverse conditions through a series of physiological, biochemical and morphological evaluation methods. Seeds with high vigor can maintain a high germination rate and growth rate under various environmental pressures, ensuring the advantage of early crop growth. This is of great significance for agricultural producers to make effective decisions under uncertain climate and environmental conditions.
[0003] Seed vigor testing can screen out seeds with high vigor, thereby improving germination and emergence rates and ensuring field planting density and uniformity. Using seeds with high vigor can reduce seed waste and the need for replanting, reducing production costs. In addition, seeds with high vigor are less dependent on fertilizers and pesticides during planting, further reducing agricultural inputs.
[0004] High-vigor seeds usually have stronger stress resistance and can better cope with adverse environmental conditions such as drought, low temperature and pests and diseases. This is of great significance for agricultural production in the context of climate change. Seed vigor testing provides key data and reference basis for breeding work, helping breeding experts to select new rice varieties with more vitality and adaptability, and promote agricultural scientific and technological progress. High-quality seeds can not only increase yields, but also improve the quality and market value of rice, thereby enhancing the competitiveness of agricultural products and increasing farmers' income. Through scientific seed vigor testing, we can optimize planting structures and management strategies, reduce the use of pesticides and fertilizers, reduce environmental pollution, and promote the sustainable development of agriculture.
[0005] However, traditional rice seed vitality detection technologies such as standard germination test, rapid aging test, conductivity measurement and TTC staining usually have disadvantages such as long detection cycle, damage to samples or the need to use chemical reagents. Especially when there are a large number of samples, the workload is huge. Generally, sampling detection methods are used to detect individual samples to evaluate the overall vitality of the samples, which is inaccurate. Therefore, it is urgent to develop new analytical technologies that are accurate, do not damage samples, and can efficiently identify the vitality of a large number of rice seeds at the same time.
[0006] Hyperspectral imaging technology is a method that combines imaging and spectral analysis. It can collect images and spectral data in the visible to near-infrared range, can detect a large number of seeds in a short time, does not need to destroy the seeds, and is suitable for large-scale seed screening. At the same time, it can be combined with automated equipment to realize the automation and standardization of seed viability detection, reduce human errors, and improve detection efficiency. At present, some scholars have applied hyperspectral technology to the detection of rice seed viability. The document "Determination of viability and vigor of naturally-aged rice seeds using hyperspectral imaging with machine learning" published in May 2022 discloses a method for viability detection based on NIR-HSI data using principal component analysis (PCA) full spectrum + characteristic wavelength to construct a convolutional neural network (CNN) and conventional machine learning methods (support vector machine (SVM) and logistic regression (LR)). The document "Research on rice seed variety and vigor grade detection based on hyperspectral imaging technology" published in April 2022 discloses a method for viability detection of 5 rice seed varieties based on Vis / NIR-HSI data.
[0007] Since HSI technology can detect data in the visible to near-infrared range, the hyperspectral data in the visible range (Vis-HSI) mainly reflects the physical characteristics of the detected object, such as appearance, color, and morphology, while the hyperspectral data in the near-infrared range (NIR-HSI) focuses on reflecting the information of the object's hydrogen-containing groups (CH, NH, OH), that is, the organic chemical components. However, most previous studies only used a single data source (Vis-HSI or NIR-HSI) for vitality detection, and few combined the two types of data for analysis. Data fusion is a method or technical framework that combines data from multiple sources to achieve the goal of detecting vitality. Various mathematical methods give play to the complementary advantages of data from different sources to improve the detection or discrimination effect. Therefore, the fusion of rice Vis-HSI and NIR-HSI data is a possible means to achieve high-precision vitality detection. Data fusion methods are generally divided into three types of fusion methods: low-level (data layer), middle-level (feature layer) and high-level (decision layer), and each of the three methods has its own advantages. In order to meet the needs of the seed and grain industry to accurately identify the vitality of a large number of rice seeds, it is necessary to further develop the fusion methods of rice Vis-HSI and NIR-HSI data to find a detection model for accurate and rapid detection of single-grain rice vitality. Summary of the invention
[0008] The technical problem to be solved by the present invention is that the single-grain rice seed vitality detection method in the prior art is complicated, time-consuming, and has low precision.
[0009] The present invention solves the above technical problems through the following technical means:
[0010] The method for constructing a rice seed single grain vitality detection model comprises the following steps:
[0011] S1. Low-level fusion: The Vis-HSI average spectrum and the NIR-HSI average spectrum of each calibration set sample are concatenated end to end to obtain a low-level fusion spectrum, and then a machine learning discriminant algorithm is used to construct a low-level fusion model of the low-level fusion spectrum and the corresponding vitality category label;
[0012] S2. Middle-level fusion: The same variable selection algorithm is used to select feature variables for the Vis-HSI average spectrum and NIR-HSI average spectrum of each calibration set sample, and then the two sets of spectra after feature selection are spliced end to end to form a middle-level fusion spectrum. Then, a machine learning discriminant algorithm is used to construct a middle-level fusion model of single-grain rice vitality for the middle-level fusion spectrum and the corresponding vitality category label;
[0013] S3. Model optimization: In the middle-level fusion, at least N variable selection algorithms are used to screen feature variables, and in steps S1 and S2, multiple machine learning discrimination algorithms are used for modeling processing. Finally, all the processed models of steps S1 and S2 are arranged in descending order of recognition performance;
[0014] S4. High-level fusion: Use the top x best models selected in step S3 to predict the validation set samples in turn, and then horizontally splice the prediction results and category labels in order as the true values. Use the machine learning algorithm to build a discriminant model for the above prediction values and true values, that is, use the prediction results of each model test set and the corresponding actual labels as training data for training, and count the model training results.
[0015] S5. Determine the final model: obtain different training results by changing the value of x in step S4 and the machine learning algorithm used, and select the model with the best result as the best rice vitality discrimination model.
[0016] Furthermore, the method for obtaining the Vis-HSI average spectrum and the NIR-HSI average spectrum in step S1 is:
[0017] Hyperspectral HSI data acquisition: All rice grains are placed on a counting plate in order to ensure that each grain is separated. Two hyperspectral cameras with detection ranges in the near-infrared range and visible light range are used to collect data on the samples to obtain near-infrared hyperspectral NIR-HSI data and visible hyperspectral Vis-HSI data of the rice on the entire counting plate.
[0018] Sample HSI data segmentation: extract the region of interest of each seed, and then perform padding operation on the rice sample highlight data: that is, place the sample in the center of the canvas of the same size, and set the background of the non-sample area to 0, so that the hyperspectral data of each seed is segmented. Since the collected HSI data includes two groups, NIR-HSI and Vis-HSI, they need to be segmented separately to obtain the NIR-HSI data and Vis-HSI data of each rice grain;
[0019] Averaging of single-grain HSI data: The average spectrum is extracted from the NIR-HSI data cube of the segmented single-grain seed. The calculation process is as follows: the dimension of the NIR-HSI data cube is W×W×C1, where C1 is the data in the near-infrared spectrum dimension, including n wavelength points. The average values of the sample image areas corresponding to the 1st, 2nd, 3rd, …, nth wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the NIR-HSI data of the seed sample is obtained; similarly, the average spectrum is extracted from the Vis-HSI data cube of the segmented single-grain seed: the dimension of the Vis-HSI data cube is W×W×C2, where C2 is the data in the visible spectrum dimension, including m wavelength points. The average values of the sample image areas corresponding to the 1st, 2nd, 3rd, …, mth wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the Vis-HSI data of the seed sample is obtained.
[0020] Furthermore, in step S1, before the Vis-HSI average spectrum and the NIR-HSI average spectrum of each calibration set sample are spliced head to tail, the average NIR-HSI and Vis-HSI spectra of the single rice grains need to be preprocessed, and the preprocessing method is Z-score standardization: the calculation formula is as follows:
[0021]
[0022] in, is the jth spectral data of the ith variety, N is the number of varieties, μ (i) is the mean, σ (j) is the standard deviation.
[0023] Furthermore, the characteristic variable selection algorithm used in steps S1 and S2 is a competitive adaptive reweighting algorithm, a minimum angle regression algorithm, a continuous projection algorithm, and an unsupervised learning characteristic wavelength selection algorithm based on spectral clustering combined with Laplace scoring method.
[0024] Furthermore, the machine learning modeling algorithms used in step 1, step 2, and step 4 are K nearest neighbor algorithm, linear discriminant analysis algorithm, and extreme gradient boosting tree algorithm.
[0025] The present invention also provides a rice seed single grain vitality detection model construction system, comprising:
[0026] Low-level fusion module: concatenate the Vis-HSI average spectrum and the NIR-HSI average spectrum of each calibration set sample to obtain a low-level fusion spectrum, and then use a machine learning discriminant algorithm to construct a low-level fusion model of the low-level fusion spectrum and the corresponding vitality category label;
[0027] Middle-level fusion module: The same variable selection algorithm is used to select feature variables for the Vis-HSI average spectrum and NIR-HSI average spectrum of each calibration set sample, and then the two sets of spectra after feature selection are spliced end to end to form a middle-level fusion spectrum. Then, a machine learning discriminant algorithm is used to construct a middle-level fusion model of single-grain rice vitality for the middle-level fusion spectrum and the corresponding vitality category label;
[0028] Model optimization module: In the middle-level fusion, at least N variable selection algorithms are used to screen feature variables, and in the low-level fusion module and the middle-level fusion module, multiple machine learning discrimination algorithms are used for modeling processing. Finally, all the processed models of the low-level fusion module and the middle-level fusion module are arranged in descending order of recognition performance;
[0029] High-level fusion module: Use the top x best models selected in the step model optimization module to predict the validation set samples in turn, and then horizontally splice the prediction results and category labels in order as the true values. Use the machine learning algorithm to build a discriminant model for the above prediction values and true values, that is, use the prediction results of each model test set and the corresponding actual labels as training data for training, and count the model training results.
[0030] Determine the final model module: obtain different training results by changing the value of x in the high-level fusion module and the machine learning algorithm used, and select the model with the best result as the optimal rice vitality discrimination model.
[0031] Furthermore, the method for obtaining the Vis-HSI average spectrum and the NIR-HSI average spectrum in the low-level fusion module is:
[0032] Hyperspectral HSI data acquisition: All rice grains are placed on a counting plate in order to ensure that each grain is separated. Two hyperspectral cameras with detection ranges in the near-infrared range and visible light range are used to collect data on the samples to obtain near-infrared hyperspectral NIR-HSI data and visible hyperspectral Vis-HSI data of the rice on the entire counting plate.
[0033] Sample HSI data segmentation: extract the region of interest of each seed, and then perform padding operation on the rice sample highlight data: that is, place the sample in the center of the canvas of the same size, and set the background of the non-sample area to 0, so that the hyperspectral data of each seed is segmented. Since the collected HSI data includes two groups, NIR-HSI and Vis-HSI, they need to be segmented separately to obtain the NIR-HSI data and Vis-HSI data of each rice grain;
[0034] Averaging of single-grain HSI data: The average spectrum is extracted from the NIR-HSI data cube of the segmented single-grain seed. The calculation process is as follows: the dimension of the NIR-HSI data cube is W×W×C1, where C1 is the data in the near-infrared spectrum dimension, including n wavelength points. The average values of the sample image areas corresponding to the 1st, 2nd, 3rd, …, nth wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the NIR-HSI data of the seed sample is obtained; similarly, the average spectrum is extracted from the Vis-HSI data cube of the segmented single-grain seed: the dimension of the Vis-HSI data cube is W×W×C2, where C2 is the data in the visible spectrum dimension, including m wavelength points. The average values of the sample image areas corresponding to the 1st, 2nd, 3rd, …, mth wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the Vis-HSI data of the seed sample is obtained.
[0035] Furthermore, in the low-level fusion module, before concatenating the Vis-HSI average spectrum and the NIR-HSI average spectrum of each calibration set sample, the average NIR-HSI and Vis-HSI spectra of a single grain of rice need to be preprocessed by Z-score standardization: the calculation formula is as follows:
[0036]
[0037] in, is the jth spectral data of the ith variety, N is the number of varieties, μ (i) is the mean, σ (j) is the standard deviation.
[0038] Furthermore, the feature variable selection algorithms used in the low-level fusion module and the middle-level fusion module are competitive adaptive reweighting algorithm, minimum angle regression algorithm, continuous projection algorithm, and unsupervised learning feature wavelength selection algorithm based on spectral clustering combined with Laplace scoring method.
[0039] The present invention also provides a method for predicting the vitality of an unknown single rice grain using the above rice seed single grain vitality detection model.
[0040] The advantages of the present invention are:
[0041] The hyperspectral multi-level data fusion method designed by the present invention can achieve more ideal results in rice seed vitality detection compared with traditional data fusion methods and traditional hyperspectral modeling methods. The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some of the technical features therein by equivalents; and these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 is a flow chart of a model building method in an embodiment of the present invention;
[0043] Figure 2 This is a diagram of modeling results of high-level fusion of different parameters in the case of the present invention. DETAILED DESCRIPTION
[0044] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in combination with the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] This embodiment provides a method for constructing a rice seed single grain vitality detection model based on hyperspectral and multi-level data fusion. Figure 1 As shown, the following steps are included:
[0046] Step 1. Sample collection: Collect no less than 400 rice grains and 200 rice grains as a calibration sample set for modeling and a validation set for improving the model, respectively.
[0047] Step 2, hyperspectral (HSI) data acquisition: All rice grains are placed on a grain counting plate in order to ensure that each grain is separated, and two hyperspectral cameras with detection ranges in the near-infrared (800nm-2500nm) range and visible light (380nm-800nm) range are used to collect data on the samples to obtain near-infrared hyperspectral (NIR-HSI) data and visible hyperspectral (Vis-HSI) data of the rice on the entire grain counting plate.
[0048] Step 3, sample HSI data segmentation: Since several hyperspectral images of rice grains are placed together, it is necessary to use image processing technology including but not limited to the Otsu method (OSTU) of threshold segmentation combined with the connected domain method to extract the region of interest (ROI) of each seed, and then perform padding operation on the rice sample highlight data: that is, place the sample in the center of a canvas of the same size (W×W, W is a pixel), and set the background of the non-sample area to 0, so that the hyperspectral data of each seed is segmented. Since the collected HSI data include NIR-HSI and Vis-HSI groups, they need to be segmented separately to obtain the NIR-HSI data and Vis-HSI data of each rice grain.
[0049] Step 4, average of single-grain HSI data: extract the average spectrum of the NIR-HSI data cube of the segmented single-grain seed, and the calculation process is as follows: the dimension of the NIR-HSI data cube is W×W×C1, where C1 is the data in the near-infrared spectrum dimension, including n wavelength points, and the average values of the sample image areas corresponding to the 1st, 2nd, 3rd, ..., nth wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the NIR-HSI data of the seed sample is obtained; similarly, for the Vis-HSI data cube of the segmented single-grain seed, extract the average spectrum: the dimension of the Vis-HSI data cube is W×W×C2, where C2 is the data in the visible spectrum dimension, including m wavelength points, and the average values of the sample image areas corresponding to the 1st, 2nd, 3rd, ..., mth wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the Vis-HSI data of the seed sample is obtained.
[0050] Step 5, single rice seed germination: using the method provided in the Chinese national standard GB / T 5520-2011 Grain and Oil Inspection Seed Germination Test, each seed was germinated. After 7 days, the vitality of the seeds was evaluated based on whether they germinated. If they germinated, they were marked as vigorous and their category label was assigned a value of 1. If they did not germinate, they were marked as inactive and their category label was assigned a value of 0.
[0051] Step 6: Spectral preprocessing: Use appropriate preprocessing methods to preprocess the average NIR-HSI and Vis-HSI spectra of single rice grains.
[0052] Step 7: Multi-level fusion model construction: The multi-level fusion method is used to screen the optimal rice vitality prediction model, which includes 5 steps:
[0053] 1) Low-level fusion: The Vis-HSI average spectrum and the NIR-HSI average spectrum of each calibration set sample are concatenated end to end to obtain a low-level fusion spectrum, and then a suitable machine learning discriminant algorithm is used to construct a low-level fusion model of the above spectra and the corresponding vitality category labels.
[0054] 2) Middle-level fusion: The same variable selection algorithm is used to screen the feature variables for the Vis-HSI average spectrum and NIR-HSI average spectrum of each calibration set sample, and then the two sets of spectra after feature selection are spliced end to end to form a middle-level fusion spectrum. Then, a suitable machine learning discriminant algorithm is used to construct a middle-level fusion model of single-grain rice vitality for the above spectra and the corresponding vitality category labels.
[0055] 3) Model optimization: In step 2), at least 3 variable selection algorithms are used to screen feature variables, and in steps 1) and 2), at least 3 machine learning discrimination algorithms are used for modeling. Finally, all processed models of 1) and 2) are arranged in descending order of recognition performance. The recognition performance of the model is evaluated by accuracy and recall. The higher these two indicators are, the better the model recognition performance is.
[0056] 4) High-level fusion: Use the first x (x≤the total number of models in step 3)) best models selected in 3) to predict the samples of the validation set in turn, and then horizontally splice the prediction results in order, and then horizontally splice their category labels in order as the true values. Use the appropriate machine learning algorithm to build a discriminant model for the above predicted values and true values, that is, use the prediction results of each model test set and the corresponding actual labels as training data for training, and count the model training results.
[0057] 5) Determine the final model: By changing the value of x in 4) and the machine learning algorithm used, different training results are obtained, and the model with the best result is selected as the optimal rice vigor discrimination model.
[0058] S8. Prediction: For unknown single-grain rice samples, use the methods described in steps S2, S3, and S4 to collect HSI data, segment the grain HSI data, and calculate the average spectrum of each grain. Use the method described in step S6 to perform spectral preprocessing. Finally, use the high-level fusion model constructed in step S7 to perform prediction to obtain the vitality discrimination result.
[0059] In this embodiment, the preprocessing method used in step S6 is Z-score standardization: this method calculates the mean and standard deviation of each column of the data set, and standardizes the data based on the mean and standard deviation of the original data. New data = (original data - mean) / standard deviation. The mean and standard deviation calculation formulas are as follows:
[0060]
[0061] in, is the jth spectral data of the ith variety, N is the number of varieties, μ (i) is the mean, σ (j) is the standard deviation.
[0062] In this embodiment, the feature variable selection algorithms used in step 1) and step 2) in step S7 are competitive adaptive reweighted sampling (CARS), least angle regression (LAR), successive projections algorithm (SPA), and unsupervised learning feature wavelength selection algorithm (SPE) based on spectral cluster-Laplace score. The principles of these algorithms are as follows:
[0063] (1) CARS: The competitive adaptive reweighting algorithm is an effective algorithm for characteristic wavelength selection. Its core idea is to dynamically adjust the weight of each wavelength in each round of iteration through an adaptive re-sampling method, and then continuously optimize, update and screen the candidate wavelength subset. Specifically, it includes the following steps: ① Initially randomly select some different bands as the initial candidate set of wavelength subsets. ② Use the candidate wavelength subset to build a partial least squares model (PLS)
[49] , and use the evaluation index PRESS to evaluate the current model performance as a baseline. ③ Re-sample and update the weights. Randomly sample all wavelengths and calculate the evaluation index corresponding to the model of the sampled subset. If it is higher than the baseline model, increase the weight of the corresponding wavelength, otherwise reduce the weight of the wavelength. ④ Update the subset. According to the weight after the subset update, select several wavelength combinations with the highest weight to form a new candidate wavelength subset. ⑤ Repeat steps ② to ④ until the termination condition is met, such as the maximum number of iterations.
[0064] (2) LAR: Least angle regression algorithm is an effective method for feature wavelength selection and belongs to the embedded feature selection method. The core idea of the LAR algorithm is to reduce the coefficients of all feature wavelengths, select the feature wavelength most relevant to the current residual in each iteration, and use the least angle regression to update the coefficient of the feature wavelength, so as to finally achieve the purpose of feature wavelength selection and coefficient estimation. It mainly includes the following steps: ① Initialize the coefficients of all feature wavelengths to 0, and the residual vector is equal to the target variable vector. ② Perform correlation calculation, calculate the correlation between each feature and the residual vector, and take the absolute value of the correlation coefficient. ③ Select the feature wavelength and update the coefficient, select the feature most relevant to the reference, and use the least angle regression to calculate the increase in the feature coefficient, and then update the coefficients of all feature wavelengths with non-zero coefficients. ④ Recalculate the residual vector using the updated coefficient value. ⑤ Repeat steps ② to ④ until the termination condition is met, that is, all features are selected or the maximum number of features is reached, and the features corresponding to the non-zero coefficients are selected as the final wavelength candidate set.
[0065] (3) SPA: The core idea of the continuous projection algorithm is to project all wavelengths onto the target variable, and gradually select the optimal characteristic wavelength subset by minimizing the evaluation index of the model error. The general steps are as follows: ① Initialization. Arrange all wavelengths in descending order according to the sum of the squares of the projection values of the target variable, and then select the top few wavelengths as the initial wavelength candidate set. ② Use the current wavelength subset candidate set to build a multivariate linear regression model and calculate the evaluation index PRESS as the baseline. ③ Projection and error calculation. Add each unselected wavelength in all wavelengths to the current wavelength candidate set in order, build a new model, and calculate the evaluation index of the new model. ④ Subset update. Find the unselected wavelength that minimizes the evaluation index from ③ and add it to the wavelength candidate set. ⑤ Repeat steps ② to ④ until the termination condition is met.
[0066] (4) SPE: The unsupervised learning feature wavelength selection algorithm based on spectral clustering combined with Laplace score is referred to as SPE. Its core principle is to perform spectral clustering on feature wavelengths, cluster the feature wavelength set into several feature wavelength subsets of different categories, and then select the best classification situation based on the clustering index silhouette coefficient; finally, under this classification situation, the Laplace score of each wavelength in each category relative to its category is calculated, and the feature wavelengths with the highest scores in each category are selected because they can represent their category to the greatest extent; after completing the above steps, the screened feature wavelength set can be obtained.
[0067] The machine learning modeling algorithms used in step 1), step 2), and step 4) in step S7 are K-Nearest Neighbor (KNN), Linear Discriminant Analysis (LDA), and eXtreme Gradient Boosting (XGB). The principles and steps of these algorithms are as follows:
[0068] (1) K-nearest neighbor algorithm (KNN): The core idea of KNN is to take the vitality categories of the K nearest neighbor training samples of a rice seed sample with unknown vitality level according to the distance or similarity between it and the rice seed with known vitality level, and predict the sample with unknown vitality level through mechanisms such as voting. The main steps are as follows: ① Select an appropriate distance metric. This study selects the commonly used Mahalanobis distance to calculate the distance between the sample of unknown vitality category and all known vitality training samples. ② Sort by the distance calculated in ① and select the K neighbors of known vitality samples closest to the unknown vitality sample. ③ Count the number of samples of each vitality level among the K neighbors, and classify the seed samples of unknown vitality level into the vitality level category with the largest number. The principle of KNN is simple and direct. It does not require estimation of model parameters or prior knowledge. It has good interpretability and can handle the classification task in this study well.
[0069] (2) Linear Discriminant Analysis (LDA): It is a commonly used supervised learning method that is used in dimensionality reduction, pattern recognition and classification tasks. The core idea of LDA is to maximize the inter-class scatter matrix while retaining the intra-class scatter matrix, and then find the most ideal discriminant direction, so that rice seed samples with different vitality can obtain the maximum degree of separation in this direction. The main steps are as follows: ① Calculate the mean vector of rice seed samples of different vitality categories. ② Calculate the scatter matrix; the intra-class scatter matrix represents the degree of dispersion of rice samples with the same vitality, and the inter-class scatter matrix describes the degree of dispersion of the mean vector of rice samples with different vitality levels. ③ Solve the generalized eigenvalue problem max(Sb*w) / (Sw*w), where Sb is the inter-class scatter matrix and Sw is the intra-class scatter matrix. The obtained w is the optimal discriminant direction vector. ④ Project the rice seed samples to the optimal discriminant direction, and then complete the classification task based on the different positions on the projected discriminant direction. LDA has the advantages of simplicity, high efficiency, and clear principles. It works best when the data conforms to the Gaussian distribution assumption, and the discriminant direction also has good interpretability.
[0070] (3) Extreme Gradient Boosting Tree (XGB). XGB is a model that has been widely used. It is not only used in some industries, but also in many top-ranked solutions in competitions on competition platforms such as Kaggle and Tianchi. The core idea of XGB is to learn multiple decision trees in an incremental additive way, then use the second-order expansion to approximate the loss function, and finally use Newton's method to iteratively optimize each tree. The general steps are as follows: ① Initialize a regression tree model, and the vitality prediction for all rice samples is the same. ② Calculate the residual. For each rice seed sample prepared for training, calculate the residual between its true label and prediction. ③ Fit this regression tree using all the residuals. ④ Calculate the weight of the regression tree. Use the regression tree obtained in ③ as a classifier. By approximating the second-order Taylor expansion of the loss function, use Newton's method to calculate the optimal weight. ⑤ Update the model. Multiply the calculated regression tree by the weight coefficient obtained in ④ and add it to the existing model. ⑥ Repeat steps ② to ⑤ until the termination condition is met. XGB uses many optimization measures including but not limited to cache optimization and parallel computing, so that XGB can have excellent computing efficiency on data of a certain scale. It also uses L1L2 regularization and subsampling to control overfitting.
[0071] Case: A method for rapid determination of rice vitality, the steps are as follows:
[0072] 1. Sample collection: Seed samples of 6 rice varieties were used, including 800 grains of Changbai 21, 400 grains of Jihong No. 6, 200 grains of Jiudao 86, 200 grains of Jiyuanxiang No. 1, 200 grains of Songfeng 696 and 200 grains of Tongyu 335. The above seed samples were provided by Sinochem Modern Agriculture Co., Ltd., totaling 2,000 grains, and placed under different storage conditions. To ensure the accuracy of the data, abnormal grains with mechanical damage caused by the transportation process in these samples were manually removed. The average size of the sample grains was about 0.8 cm long and 0.35 cm wide. In view of the fact that this embodiment uses a 200-hole counting plate for hyperspectral data acquisition and a 96-hole germination plate for germination experiments, appropriate adjustments were made to achieve accurate data correspondence: the last 8 seeds of each 200-hole counting plate will be discarded, so that each plate of seed samples can completely correspond to two 96-hole germination plates. Finally, after deducting a small amount of loss during the experiment, 1,878 valid seed samples were retained for subsequent analysis.
[0073] 2. HSI data acquisition: All rice grains were placed on a counting plate in order, and data were collected on the samples using the Specim FX hyperspectral acquisition platform (equipped with a NIR-HSI camera with a range of 900nm-1700nnm and a Vis-NIR-HSI camera with a range of 400nm-1000nm) to obtain near-infrared hyperspectral (NIR-HSI) data and visible hyperspectral (Vis-HSI) data of the rice on the entire counting plate.
[0074] 3. Sample HSI data segmentation: The Otsu method (OSTU) combined with the connected domain method is used to extract the region of interest (ROI) of each seed, and then the padding operation is performed on the rice sample highlight data: that is, the sample is placed in the center of the canvas of the same size (32×32, 32 pixels), and the background of the non-sample area is set to 0, so that the hyperspectral data of each seed is segmented. Since the collected HSI data include NIR-HSI and Vis-HSI groups, they need to be segmented separately to obtain the NIR-HSI data and Vis-HSI data of each rice grain.
[0075] 4. Average of single rice grain HSI data: The average spectrum is extracted from the segmented NIR-HSI data cube of the single rice grain. The calculation process is as follows: The dimension of the NIR-HSI data cube is 32×32×C1, where C1 is the data of the near-infrared spectrum dimension, including 224 wavelength points. The average values of the sample image areas corresponding to the 1st, 2nd, 3rd, …, 224th wavelength points of the spectrum dimension are calculated respectively as the absorbance value of the current wavelength point, until the average absorbance value of all wavelength points is calculated, and the NIR-HSI value of the seed sample is obtained. The average spectrum of the Vis-HSI data is extracted from the segmented single seed Vis-HSI data cube: the dimension of the Vis-HSI data cube is 32×32×C2, where C2 is the data in the visible spectrum dimension, containing 448 wavelength points. The average values of the sample image areas corresponding to the 1st, 2nd, 3rd, …, 224th wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the Vis-HSI data of the seed sample is obtained.
[0076] 5. Germination experiment: Put several seeds on the plate into a 96-well plate with a section of the bottom cut off, add water and culture in the culture room, conduct a germination experiment, change the water every day, and after 7 days, mark whether the seeds are viable based on whether they germinate. If they germinate, they are marked as viable, otherwise they are marked as viable.
[0077] According to the germination experiment, there were 1,271 rice seed samples that had germinated, accounting for 66% of the total samples; there were 643 rice seed samples that had not germinated, accounting for 34% of the total samples.
[0078] According to the hyperspectral data obtained above, the germination results were used as labels, and the above samples were randomly classified into training set, validation set, and test set in a ratio of 8:1:1. The number of germinated and ungerminated rice seeds in the training set was 989 and 513, the number of validation sets was 123 and 65, and the number of test sets was the same as the validation set, which was 123 and 65. The training set was used to build the model, and the validation set and the test set were used to verify the model performance. If the three data sets are better, the model performance is better. The training set and the validation set are used in the optimization of the multi-level fusion model to evaluate the recognition performance of the low-level and middle-level fusion models, while the test set is used in the optimization of the multi-level fusion model to evaluate the recognition performance of the high-level fusion model.
[0079] 6. Construct a multi-level fusion model of rice vitality: All NIR-HSI and Vis-HSI data are Z-score standardized, and then the preprocessed data and the real data labels are combined and substituted into the low-level fusion and high-level fusion training set data for five-fold cross-validation. The prediction results of multiple models on the test set are then passed to the high-level fusion part for further screening and secondary training, and then verified on the validation set.
[0080] In the low-level fusion, the Vis-HSI average spectrum and the NIR-HSI average spectrum of each calibration set sample are concatenated end to end to obtain the low-level fusion spectrum. Then, the KNN, LDA and XGB algorithms are used to obtain the low-level fusion model of the above spectra and the corresponding vitality category labels. A total of 3 groups of low-level fusion models for comparison are obtained.
[0081] In the middle-level fusion, the Vis-HSI average spectrum and NIR-HSI average spectrum of each calibration set sample were firstly screened for feature variables using the same variable selection algorithm (CARS, LAR, SPA and SPE algorithms respectively), and then the two sets of spectra after feature selection were concatenated end to end to form a middle-level fusion spectrum. Then, the KNN, LDA and XGB algorithms were used to construct the middle-level fusion model of single-grain rice vitality for the above spectra and the corresponding vitality category labels. A total of 12 sets of middle-level fusion models for comparison were obtained.
[0082] The predicted values and true values of the test sets of the above 15 groups of models were counted to evaluate the recognition performance (precision and recall) of the models. The results are shown in Table 1.
[0083] Table 1 Prediction performance evaluation of low-level fusion and mid-level fusion models
[0084]
[0085] Table 1 shows the training and verification results of the multi-level spectral data fusion model. The second column indicates the specific fusion method, while the third column indicates the machine learning algorithm used for modeling. For example, the CARS row in the middle-level fusion refers to the splicing and fusion of the CARS wavelength selection results of the NIR-HSI data and the CARS wavelength selection results of the VIS-HSI spectral data as the characteristic wavelength, and then further adopts the subsequent KNN / LDA / XGB algorithm to model the result.
[0086] In Table 1, the best middle-layer fusion model is the LDA model based on CARS wavelength selection, with an accuracy of 89.9% and 86.7% on the calibration set and validation set; the best low-layer fusion model is the XGB model, with an accuracy of 86.2% and 86.7% on the calibration set and validation set.
[0087] After listing the results of the above low-level and middle-level fusion models, high-level fusion is performed: take the first N (N≤15) best models in Table 1, predict the test set samples in turn, and then horizontally splice the prediction results in order, and then horizontally splice their category labels in order as the true values. Use the KNN / LDA / XGB algorithm to build a discriminant model for the above predicted values and true values, that is, use the predicted results of the test set of each model and the corresponding actual labels as training data for training, and count the model training results.
[0088] Figure 2 It is the result of high-level fusion (the number of combined models ranges from 3 to 12). The horizontal axis represents the number of the top N models with the best performance used, and the vertical axis represents the prediction accuracy of the fused model on the validation set.
[0089] Depend on Figure 2 It can be seen that there are five groups of models in the high-level fusion method that have achieved the best accuracy of 88.3% in the test set, namely 1) KNN model based on the prediction results of the top 3 best low-level and middle-level fusion models, 2) KNN model based on the prediction results of the top 4 best low-level and middle-level fusion models, 3) XGB model based on the prediction results of the top 3 best low-level and middle-level fusion models, 4) XGB model based on the prediction results of the top 4 best low-level and middle-level fusion models, and 5) XGB model based on the prediction results of the top 5 best low-level and middle-level fusion models. The results of the above five high-level fusion models are higher than the best prediction results of low-level fusion and middle-level fusion, so these models are considered to be the best models.
[0090] Based on the above steps, taking model 1) in the above high-level fusion model as an example, the modeling process of a complete multi-level model is described as follows: using CARS to screen the characteristic variables of the NIR-HSI average spectrum and the Vis-HSI average spectrum of a single grain of rice, then splicing the variables together, and using LDA to build a discriminant model; similarly, using SPA to perform the above variable screening step, and then using LDA to build a discriminant model, and using LARS to perform the above variable screening step, and then using LDA to build a discriminant model. Use the above three models (the above models are the top three groups of models with the best performance in Table 1) to predict the vitality of the test set, splice the prediction results head to tail, and use the KNN algorithm to build a discriminant model of these predicted values and true values (obtained through germination experiments).
[0091] When predicting a new rice sample, the HSI data is collected in the same manner as in the embodiment, the grain HSI data is segmented and the average spectrum of each grain is calculated, the spectrum is preprocessed using Z-scores standardization, and finally the above-mentioned multi-level fusion model is used for prediction to obtain the vitality discrimination result.
[0092] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a rice seed single-grain vitality detection model, characterized in that: The following steps are involved: S1. Low-level fusion: The Vis-HSI average spectrum and the NIR-HSI average spectrum of each calibration set sample are concatenated end to end to obtain a low-level fusion spectrum, and then a machine learning discriminant algorithm is used to construct a low-level fusion model of the low-level fusion spectrum and the corresponding vitality category label; S2. Middle-level fusion: The same variable selection algorithm is used to select feature variables for the Vis-HSI average spectrum and NIR-HSI average spectrum of each calibration set sample, and then the two sets of spectra after feature selection are spliced end to end to form a middle-level fusion spectrum. Then, a machine learning discriminant algorithm is used to construct a middle-level fusion model of single-grain rice vitality for the middle-level fusion spectrum and the corresponding vitality category label; S3. Model optimization: In the middle-level fusion, at least N variable selection algorithms are used to screen feature variables, and in steps S1 and S2, multiple machine learning discrimination algorithms are used for modeling processing. Finally, all the processed models of steps S1 and S2 are arranged in descending order of recognition performance; S4. High-level fusion: Use the top x best models selected in step S3 to predict the validation set samples in turn, and then horizontally splice the prediction results and category labels in order as the true values. Use the machine learning algorithm to build a discriminant model for the above prediction values and true values, that is, use the prediction results of each model test set and the corresponding actual labels as training data for training, and count the model training results. S5. Determine the final model: obtain different training results by changing the value of x in step S4 and the machine learning algorithm used, and select the model with the best result as the best rice vitality discrimination model.
2. The method for constructing a rice seed single grain vitality detection model according to claim 1, characterized in that: The method for obtaining the Vis-HSI average spectrum and the NIR-HSI average spectrum in step S1 is: Hyperspectral HSI data acquisition: All rice grains are placed on a counting plate in order to ensure that each grain is separated. Two hyperspectral cameras with detection ranges in the near-infrared range and visible light range are used to collect data on the samples to obtain near-infrared hyperspectral NIR-HSI data and visible hyperspectral Vis-HSI data of the rice on the entire counting plate. Sample HSI data segmentation: extract the region of interest of each seed, and then perform padding operation on the rice sample highlight data: that is, place the sample in the center of the canvas of the same size, and set the background of the non-sample area to 0, so that the hyperspectral data of each seed is segmented. Since the collected HSI data includes two groups, NIR-HSI and Vis-HSI, they need to be segmented separately to obtain the NIR-HSI data and Vis-HSI data of each rice grain; Averaging of single-grain HSI data: The average spectrum is extracted from the NIR-HSI data cube of the segmented single-grain seed. The calculation process is as follows: the dimension of the NIR-HSI data cube is W×W×C1, where C1 is the data in the near-infrared spectrum dimension, including n wavelength points. The average values of the sample image areas corresponding to the 1st, 2nd, 3rd, …, nth wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the NIR-HSI data of the seed sample is obtained; similarly, the average spectrum is extracted from the Vis-HSI data cube of the segmented single-grain seed: the dimension of the Vis-HSI data cube is W×W×C2, where C2 is the data in the visible spectrum dimension, including m wavelength points. The average values of the sample image areas corresponding to the 1st, 2nd, 3rd, …, mth wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the Vis-HSI data of the seed sample is obtained.
3. The method for constructing a rice seed single grain vitality detection model according to claim 2, characterized in that: In step S1, before the Vis-HSI average spectrum and the NIR-HSI average spectrum of each calibration set sample are spliced head to tail, the average NIR-HSI and Vis-HSI spectra of the single rice grains need to be preprocessed. The preprocessing method is Z-score standardization: the calculation formula is as follows: in, is the jth spectral data of the ith variety, N is the number of varieties, μ (i) is the mean, σ (j) is the standard deviation.
4. The method for constructing a rice seed single grain vitality detection model according to any one of claims 1 to 3, characterized in that: The characteristic variable selection algorithms used in the steps S1 and S2 are competitive adaptive reweighting algorithm, minimum angle regression algorithm, continuous projection algorithm, and unsupervised learning characteristic wavelength selection algorithm based on spectral clustering combined with Laplace scoring method.
5. The method for constructing a rice seed single grain vitality detection model according to any one of claims 1 to 3, characterized in that: The machine learning modeling algorithms used in step 1, step 2, and step 4 are K nearest neighbor algorithm, linear discriminant analysis algorithm, and extreme gradient boosting tree algorithm.
6. A rice seed single grain vitality detection model construction system, characterized in that: include: Low-level fusion module: concatenate the Vis-HSI average spectrum and the NIR-HSI average spectrum of each calibration set sample to obtain a low-level fusion spectrum, and then use a machine learning discriminant algorithm to construct a low-level fusion model of the low-level fusion spectrum and the corresponding vitality category label; Middle-level fusion module: The same variable selection algorithm is used to select feature variables for the Vis-HSI average spectrum and NIR-HSI average spectrum of each calibration set sample, and then the two sets of spectra after feature selection are spliced end to end to form a middle-level fusion spectrum. Then, a machine learning discriminant algorithm is used to construct a middle-level fusion model of single-grain rice vitality for the middle-level fusion spectrum and the corresponding vitality category label; Model optimization module: In the middle-level fusion, at least N variable selection algorithms are used to screen feature variables, and in the low-level fusion module and the middle-level fusion module, multiple machine learning discrimination algorithms are used for modeling processing. Finally, all the processed models of the low-level fusion module and the middle-level fusion module are arranged in descending order of recognition performance; High-level fusion module: Use the top x best models selected in the step model optimization module to predict the validation set samples in turn, and then horizontally splice the prediction results and category labels in order as the true values. Use the machine learning algorithm to build a discriminant model for the above prediction values and true values, that is, use the prediction results of each model test set and the corresponding actual labels as training data for training, and count the model training results. Determine the final model module: obtain different training results by changing the value of x in the high-level fusion module and the machine learning algorithm used, and select the model with the best result as the optimal rice vitality discrimination model.
7. The system for constructing a rice seed single-grain vitality detection model according to claim 6, characterized in that: The method for obtaining the Vis-HSI average spectrum and the NIR-HSI average spectrum in the low-level fusion module is: Hyperspectral HSI data acquisition: All rice grains are placed on a counting plate in order to ensure that each grain is separated. Two hyperspectral cameras with detection ranges in the near-infrared range and visible light range are used to collect data on the samples to obtain near-infrared hyperspectral NIR-HSI data and visible hyperspectral Vis-HSI data of the rice on the entire counting plate. Sample HSI data segmentation: extract the region of interest of each seed, and then perform padding operation on the rice sample highlight data: that is, place the sample in the center of the canvas of the same size, and set the background of the non-sample area to 0, so that the hyperspectral data of each seed is segmented. Since the collected HSI data includes two groups, NIR-HSI and Vis-HSI, they need to be segmented separately to obtain the NIR-HSI data and Vis-HSI data of each rice grain; Averaging of single-grain HSI data: The average spectrum is extracted from the NIR-HSI data cube of the segmented single-grain seed. The calculation process is as follows: the dimension of the NIR-HSI data cube is W×W×C1, where C1 is the data in the near-infrared spectrum dimension, including n wavelength points. The average values of the sample image areas corresponding to the 1st, 2nd, 3rd, …, nth wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the NIR-HSI data of the seed sample is obtained; similarly, the average spectrum is extracted from the Vis-HSI data cube of the segmented single-grain seed: the dimension of the Vis-HSI data cube is W×W×C2, where C2 is the data in the visible spectrum dimension, including m wavelength points. The average values of the sample image areas corresponding to the 1st, 2nd, 3rd, …, mth wavelength points in the spectrum dimension are calculated respectively as the absorbance values of the current wavelength point, until the absorbance average values of all wavelength points are calculated, and the average spectrum of the Vis-HSI data of the seed sample is obtained.
8. The system for constructing a rice seed single-grain vitality detection model according to claim 7, characterized in that: In the low-level fusion module, before the Vis-HSI average spectrum and the NIR-HSI average spectrum of each calibration set sample are concatenated head to tail, the average NIR-HSI and Vis-HSI spectra of a single rice grain need to be preprocessed. The preprocessing method is Z-score standardization: the calculation formula is as follows: in, is the jth spectral data of the ith variety, N is the number of varieties, μ (i) is the mean, σ (j) is the standard deviation.
9. The system for constructing a rice seed single grain vitality detection model according to any one of claims 6 to 8, characterized in that: The feature variable selection algorithms used in the low-level fusion module and the middle-level fusion module are competitive adaptive reweighting algorithm, minimum angle regression algorithm, continuous projection algorithm, and unsupervised learning feature wavelength selection algorithm based on spectral clustering combined with Laplace scoring method.
10. Use the rice seed single grain vitality detection model according to any one of claims 1 to 5 to predict the vitality of an unknown single grain of rice.
Citation Information
Cited By
Multi-modal detection method, system, equipment and medium for viability of rice germplasm resources
CN120632782A