Lettuce seed variety identification method based on multispectral imaging technology and machine learning algorithm
The morphological and spectral characteristic data of lettuce seeds were obtained through multi-spectral imaging technology, and the LDA model was constructed using machine learning algorithms, which solved the problem of low identification accuracy of lettuce varieties in the existing technology, and achieved efficient and accurate identification of lettuce seed varieties.
Patent Information
- Application Number
- CN202411855269.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-17
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art has problems such as complex operation, strong destructiveness and low accuracy in the identification of lettuce varieties. Especially when processing lettuce seeds, it is difficult to effectively distinguish different varieties.
Multispectral imaging technology combined with machine learning algorithms is used to obtain the morphological characteristics and spectral characteristic data of lettuce seeds, and a linear discriminant analysis (LDA) model is constructed to achieve high-precision identification of lettuce seed varieties.
It realizes rapid, lossless and accurate identification of lettuce seed varieties, improves identification accuracy, saves time and cost, and has important application value.
Smart Images

Figure CN120198700A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of seed identification, and particularly relates to a method for identifying lettuce seed varieties based on multispectral imaging technology and machine learning algorithms. Background Art
[0002] Lettuce (Lactuca sativa L.) is an annual plant belonging to the Compositae family and is a vegetable crop widely cultivated globally. According to statistics in 2013, the global lettuce production reached 24.9 million tons, with more than half of the production coming from China (Boriss and Brunke., 2005). Lettuce germplasm resources and varieties are rich and diverse. According to the edible parts, it can be divided into stem lettuce and leaf lettuce. According to morphological characteristics, leaf lettuce can be further divided into upright lettuce, loose-leaf lettuce, and head lettuce (Lebeda et al., 2006). With the gradual increase in lettuce consumption and the continuous application of new breeding technologies, the number of lettuce varieties available for cultivation globally is also increasing. For example, the number of newly added lettuce varieties in the European region usually exceeds 100 each year (Van Treuren et al., 2008). However, due to the diversity of different agricultural regional environments, there are strict requirements for different lettuce varieties. The same lettuce variety has significantly different economic yields and qualities in different planting environments. Therefore, identifying and selecting suitable varieties is crucial for ensuring lettuce yield and quality (Al-Shammari and Saud, 2014).
[0003] Current methods for crop variety identification mainly include morphological identification, biochemical identification, and genetic identification. Morphological identification is a relatively traditional method for variety identification. It distinguishes different varieties by observing the external morphological characteristics of crops, such as plant height, leaf shape, flower color, seed shape and color, etc. (Asati et al., 2023). Biochemical identification analyzes the isozyme bands in crop tissues through protein electrophoresis technology to achieve variety identification (Cooke, 2020). With the development of genetics-related technologies, such as the emergence of molecular marker technology and genomics, the application of protein electrophoresis and isozyme analysis in variety identification has gradually been replaced by more efficient and accurate molecular biology-related technologies (Sinha et al., 2023). For example, in the current assessment of lettuce variety genetic diversity, molecular marker technologies that have been adopted include random amplified polymorphic DNA (RAPD) (Sharma et al., 2018), amplified fragment length polymorphism ( et al., 2021), Target Region Amplification Polymorphism (TRAP) (Hu et al., 2005), Simple Sequence Repeat (SSR) (Rui et al., 2020), Single Nucleotide Polymorphism (SNPs) (Kwon et al., 2013), Selective Amplification of Microsatellite Polymorphic Loci (SAMPL) (Hassan, 2024), Genome-Wide Association Study (GWAS) (Nair., 2024), etc. In addition, molecular marker technology has also been successfully applied to lettuce variety identification and genotype analysis (Van de Wiel et al., 1999; Hong et al., 2013; Rauscher and Simko, 2013). Nevertheless, due to the relatively limited genetic diversity of lettuce, DNA markers suitable for lettuce variety identification must have a high information content (Simko, 2009; Simko et al., 2009). In addition, due to factors such as limitations in detection sites, insufficient stability, and high requirements for experimental skills, genetic identification methods such as molecular markers still have certain limitations in lettuce variety identification and diversity analysis (Sakiyama, 2014).
[0004] In view of this, developing a crop variety identification method with simple operation, high efficiency and non-destructiveness is the common goal pursued by breeders. In recent years, with the development of technologies such as spectral technology, machine learning and artificial intelligence, technologies such as multispectral imaging, hyperspectral imaging, infrared spectroscopy, magnetic resonance imaging (MRI) have gradually been used to identify germplasm resources, crop varieties, seed vigor and other quality characteristics of crops (ElMasry et al., 2019). For example, hyperspectral imaging provides electromagnetic spectrum information in the ultraviolet to near-infrared band (380-2500 nm), and can capture the spatial distribution and image characteristics of samples, providing richer information for germplasm analysis (Ambrose et al., 2016). Some researchers used near-infrared hyperspectral images combined with deep learning to identify the vigor of rice seeds harvested in different years, and finally the identification accuracy of KN4 rice seeds using the self-built convolutional neural network reached 99.5018% (Yang et al., 2020); Infrared spectroscopy, as a sensing technology, can quickly, non-destructively and efficiently detect the chemical components on the surface and inside of the seeds to be tested (Sakiyama et al., 2014). Some studies used near-infrared (Near Infrared, NIR) non-destructive technology combined with machine learning and spectral dimensionality reduction methods to distinguish the vigor of individual wheat seeds. Finally, all 8 joint models constructed could well distinguish the vigor of individual wheat seeds at 3 different levels, with quite good performance. Among them, the PCA-ELM and SPA-RF algorithms can be used as two optimal classification models (Fan et al., 2020); In addition, some studies used near-infrared spectroscopy to analyze the damage and vigor of soybean and corn grains. Finally, it was found that in the two partial data collected, for non-aged seeds and artificially accelerated aged seeds, non-heat-treated seeds and heat-damaged seeds, the identification accuracy was higher than 98% (Wang et al., 2020). Although these spectral technologies have been widely used, they still have certain limitations. For example, the detection results of near-infrared spectroscopy will be affected by the coincidence of the defect position and the spectral acquisition area; while the penetration power of hyperspectral imaging technology is poor, and it is difficult to judge the deep internal situation of the sample (Walsh et al., 2020).
[0005] Multispectral imaging technology (MSI), as a non-destructive identification technology emerging in recent years, can also combine spectral information with computer vision information to quickly provide information on a series of morphological or spectral characteristics of the sample to be tested (ElMasry et al., 2019). Different from infrared spectroscopy and hyperspectral imaging, this technology uses laser-type illumination with LEDs (light-emitting diodes), ranging from the ultraviolet radiation band (≥200 nm) to the near-infrared range (≤1000 nm), and uses discrete filters (usually 3 to 20 discontinuous bands) to form discrete spectral bands. When the light irradiates the target object, the CCD monochrome sensor reflects the signal and converts it into a spectral image, which contains the specific physical and chemical characteristics of the sample (BOELT et al., 2018; ELMASRY et al., 2019), and the scanning time is also faster than magnetic resonance imaging. At the same time, with the continuous development of deep learning technology and machine learning technology, multispectral imaging technology has also been more widely applied. For example, random forest and artificial neural network models constructed using multispectral videometer Lab data and selected nutrient composition data calculated various nutritional characteristics of black rice with accuracies of 85.35% and 99.9% (Buenafe et al., 2022); the use of a linear discriminant analysis (LDA) classifier and an SVM model combined with multispectral data can achieve the identification of different cultivated varieties of oat seeds (Fu et al., 2023); the stacked ensemble learning SEL model combined with MSI technology can distinguish single seeds of three alfalfa species (JIA, et al., 2022); the use of orthogonal partial least squares discriminant analysis (OPLS-DA) combined with MSI analysis can achieve the classification of coffee bean seeds (Mihailova et al., 2022), sugar beet seeds and tomato cultivated varieties (Salimi and Boelt, 2019; Shrestha et al., 2015), and systematically evaluate the quality of soybean seeds ( et al., 2023).
[0006] There is currently no report on the research of a method for identifying lettuce varieties based on multispectral imaging technology and machine learning algorithms. Summary of the Invention
[0007] Aiming at the above deficiencies of the prior art, the present invention for the first time constructs a method for identifying lettuce seed varieties based on multispectral imaging technology and machine learning algorithms. It is a model constructed using the morphological and spectral characteristic data of lettuce seeds obtained based on multispectral imaging technology. Among them, the average classification accuracy of the test set of the LDA (linear discriminant analysis) model reaches 92.7%, and the accuracy of the prediction set of the LDA model is even as high as 93%. This model has a very high recognition accuracy and shows good model stability.
[0008] A method for identifying lettuce seed varieties based on multispectral imaging technology and machine learning algorithms of the present invention includes the following steps:
[0009] S1. Select representative lettuce seeds of different varieties for constructing a model;
[0010] S2. Collect multispectral image information of lettuce seeds of different varieties;
[0011] S3. Obtain morphological feature data of lettuce seeds of different varieties under multispectral images;
[0012] S4. Obtain spectral feature data of lettuce seeds of different varieties under multispectral images;
[0013] S5. Construct an LDA model by using the obtained morphological feature data and spectral feature data of lettuce seeds of different varieties;
[0014] S6. According to the morphological feature data and spectral feature data of the lettuce seeds of the variety to be identified collected in steps S2 - S4, substitute them into the LDA model constructed in step S5 to conduct identification of lettuce seed varieties.
[0015] Preferably, the different varieties for constructing the model in step S1 are Guasi Hong, Kexing No. 5 Lettuce, Siqing Lettuce, Nencui Cabbage Lettuce, Ruili Cabbage Lettuce, Cabbage Lettuce, Italian Lettuce, Cream Lettuce, Ice Leaf Lettuce, American Big Fast - growing, Big Fast - growing, Four - season Oyster Plant, Peacock Lettuce, Roman Lettuce, Qingxiang Hongyou Oyster Plant.
[0016] Preferably, in step S2, a Videometer Lab4 multispectral imaging system is used to obtain monochromatic images of lettuce seeds under irradiation at 19 specific wavelengths from 365nm to 970nm, so as to obtain multispectral image information.
[0017] More preferably, the 19 specific wavelengths from 365nm to 970nm are 365nm, 405nm, 430nm, 450nm, 470nm, 490nm, 515nm, 540nm, 570nm, 590nm, 630nm, 645nm, 660nm, 690nm, 780nm, 850nm, 880nm, 940nm and 970nm wavelengths.
[0018] More preferably, the Videometer Lab4 multispectral imaging system is calibrated before each use, and geometric correction and optical frequency measurement are continuously performed using white, black, and white dotted line plates. The image size is 2192×2192 pixels, the resolution is 40μm / pixel, and 32 bits / pixel.
[0019] Preferably, the morphological features in the step S3 include area, length, width, compact circle, compact ellipse, BetaShape_a, BetaShape_b, vertical skew, CIELab L*, CIELab a*, CIELab b*, saturation, and hue.
[0020] Preferably, the spectral feature data in the step S4 is the average spectral reflectance of each variety of lettuce seeds irradiated at wavelengths of 365nm, 405nm, 430nm, 450nm, 470nm, 490nm, 515nm, 540nm, 570nm, 590nm, 630nm, 645nm, 660nm, 690nm, 780nm, 850nm, 880nm, 940nm, and 970nm respectively.
[0021] Preferably, in the step S5, before data analysis, the data is randomly split into 70% training set data and 30% test set data using the validation set method. Then, an LDA model is constructed based on the combined data of the morphological feature data and spectral feature data of different varieties of lettuce seeds. The parameter setting and result analysis visualization of the model are completed through the R language software version 4.1.3 and the online data analysis platform SPSSPRO version 1.1.26.
[0022] Although the average values of the 15 seed morphological feature indicators adopted in the present invention show significant differences among species, due to the variation or overlap among the data of 300 seeds selected for each variety, it is difficult to directly distinguish 15 lettuce varieties by a single seed morphological feature or spectral feature. This may be due to the very similar spectral-texture-morphological features of the seeds among different lettuce varieties (Concepcion II et al., 2020). In addition, some lettuce varieties have the same seed origin, and their cultivation climates and seed treatment methods are also very similar, which results in smaller differences among similar seeds in terms of seed quality at harvest, increasing the difficulty of variety classification (Dornbos, 2020).
[0023] In the present invention, we first attempt to apply MSI technology to the classification and identification of different lettuce varieties to reduce seed screening errors. Then, by further combining this technology with machine learning algorithms, we can finally accurately, non-destructively, and quickly identify seeds of different lettuce varieties.
[0024] Both principal component analysis and nCDA analysis showed some differences among the population data of seeds from different lettuce varieties. Among them, the results of nCDA were better than those of principal component analysis, which might be because nCDA is a supervised dimensionality reduction technique while principal component analysis is unsupervised (Jia et al., 2023). The results in this invention also indicated that neither PCA nor nCDA could effectively distinguish the 15 varieties completely, especially those samples near the boundary. We speculated that this might be because the morphological and spectral characteristic data means of seeds with the same seed coat color were very close, resulting in a larger within-group variance. The lettuce seeds of different varieties surrounded their respective clustering centers, but due to the scattered distribution of sample points and multiple overlaps between sample points, it affected the analysis and unification of the between-group differences of samples by PCA and nCDA, leading to an unsatisfactory final result. In addition, these two data dimensionality reduction analysis methods themselves would also miss many key classification information, thus affecting the classification effect of varieties. For example, PCA was mainly to optimize the variance value of data in the projection subspace while minimizing the dimension (Hamouda et al., 2020). However, in this process, PCA might miss data very important for the classification task, especially when dealing with low-energy directions (Song et al., 2018). Therefore, it was necessary to construct a corresponding quantitative discrimination model to achieve a better classification effect.
[0025] The overall classification capabilities of the BP, SVM, RF, and LDA models constructed based on the seed morphological characteristics and spectral characteristic data of 15 lettuce varieties in this invention were significantly better than those of the models constructed with single data. Among them, the LDA and BP models had the best classification effects, both reaching an accuracy rate of over 90%. And through the final model pre-test experiment, it was found that the overall effect of the LDA discriminant analysis model was the best.
[0026] The beneficial effects of this invention are as follows:
[0027] The method proposed in this invention is a lettuce variety identification method using multispectral imaging technology and machine learning algorithms. The model constructed with the obtained lettuce seed morphological characteristics and spectral characteristic data obtained the LDA model and verified that the prediction accuracy rate of this model was relatively high.
[0028] The non-destructive detection of lettuce and the model established in this invention solved the problem of lettuce variety identification at the seed stage, saved the problem of using vegetative bodies as detection materials required by many variety identification methods, and saved time, labor, and cost for lettuce production. In addition, this technology may be applicable to the evaluation of lettuce seed quality such as seed vigor and pest and disease conditions, and has important application value. Description of the Drawings
[0029] Figure 1It is a schematic diagram of the internal structure (right) and the VideometerLab multispectral imaging system (left).
[0030] Figure 2 They are morphological diagrams of the vegetative organs during the edible period of 15 lettuce varieties.
[0031] Figure 3 They are images of the seeds of 15 lettuce varieties under an RGB light source.
[0032] Figure 4 They are box plots of the differences in the shape characteristics of the seeds of 15 lettuce varieties; (A) circular compactness, (B) elliptical compactness, (C) width-to-length ratio, (D) BetaShpe_a, (E) BetaShpe_b, (F) vertical skewness. Different lowercase letters in the figure represent significant differences among the varieties in the same shape (P<0.05).
[0033] Figure 5 They are box plots of the differences in the binary characteristics of the seeds of 15 lettuce varieties; (A) width, (B) area, (C) length. Different lowercase letters in the figure represent significant differences among the varieties in the same shape (P<0.05).
[0034] Figure 6 They are box plots of the differences in the color characteristics of the seeds of 15 lettuce varieties; (A) CIELab L*, (B) CIELab A*, (C) CIELab B*, (D) saturation, (E) hue. Different lowercase letters in the figure represent significant differences among the varieties in the same shape (P<0.05).
[0035] Figure 7 They are line graphs of the average light reflectance of the seeds of 15 lettuce varieties at 19 wavelengths (nm). Different colors represent different lettuce varieties, and the error bars represent ± standard deviation.
[0036] Figure 8 They are two-dimensional scatter plots of PCA and OPLS-DA based on the multispectral data of 15 lettuce varieties; two-dimensional plane graphs of the first two principal components (PCs) of the dataset based on 15 lettuce varieties: (A) morphological combined with spectral characteristics, (B) morphological characteristics, (C) spectral characteristics.
[0037] Figure 9 They are two-dimensional score scatter plots of nCDA based on the multispectral data of 15 lettuce varieties; based on 15 lettuce varieties: (A) morphological characteristics combined with spectral characteristics, (B) spectral characteristics, (C) morphological characteristics.
[0038] Figure 10RGB, nCDA, and PCA images of seeds of 15 lettuce varieties; based on multispectral data of 19 wavelengths, nCDA and PCA image conversions were performed on the multispectral images of all lettuce varieties, and different nCDA and PCA values were colored from blue to red from low to high (in the legend, red corresponds to 2.00, green corresponds to 0.00, and blue corresponds to -2.00).
[0039] Figure 11 Results of the RF, BP, LDA, and SVM models based on the morphological feature data of seeds of different lettuce varieties in the test set; (A) Radar chart for performance evaluation of the four models based on the indicators of accuracy, F1, sensitivity, and precision. The range of 0.5 - 1 in the figure represents the magnitude range of the four evaluation indicator values. (B) Confusion matrix of the BP model. (C) Confusion matrix of the LDA model. (D) Confusion matrix of the RF model. (E) Confusion matrix of the SVM model.
[0040] Figure 12 Analysis results of the RF, BP, LDA, and SVM classification models based on the morphological feature data of seeds of different lettuce varieties in the training set; (A) Radar chart for performance evaluation of the four models based on the indicators of accuracy, F1, sensitivity, and precision. The range of 0.5 - 1 in the figure represents the magnitude range of the four evaluation indicator values. (B) Confusion matrix of the BP model. (C) Confusion matrix of the LDA model. (D) Confusion matrix of the RF model. (E) Confusion matrix of the SVM model.
[0041] Figure 13 Results of the RF, BP, LDA, and SVM models based on the spectral feature data of seeds of different lettuce varieties in the test set; (A) Radar chart for performance evaluation of the four models based on the indicators of accuracy, F1, sensitivity, and precision. The range of 0.3 - 1 in the figure represents the magnitude range of the four evaluation indicator values. (B) Confusion matrix of the BP model. (C) Confusion matrix of the LDA model. (D) Confusion matrix of the RF model. (E) Confusion matrix of the SVM model.
[0042] Figure 14 Analysis results of the RF, BP, LDA, and SVM models based on the spectral feature data of seeds of different lettuce varieties in the training set; (A) Radar chart for performance evaluation of the four models based on the indicators of accuracy, F1, sensitivity, and precision. The range of 0.5 - 1 in the figure represents the magnitude range of the four evaluation indicator values. (B) Confusion matrix of the BP model. (C) Confusion matrix of the LDA model. (D) Confusion matrix of the RF model. (E) Confusion matrix of the SVM model.
[0043] Figure 15Analysis results of the test sets of RF, BP, LDA, and SVM classification models for the seed morphological characteristics combined with spectral characteristic data of different lettuce varieties; (A) Radar chart for evaluating the performance of the four models based on the indicators of accuracy, F1, sensitivity, and precision. The range of 0.5 - 1 in the figure represents the magnitude range of the four evaluation index values, (B) Confusion matrix of the BP model, (C) Confusion matrix of the LDA model, (D) Confusion matrix of the RF model, (E) Confusion matrix of the SVM model.
[0044] Figure 16 Analysis results of the training sets of RF, BP, LDA, and SVM classification models based on the seed morphological characteristics combined with spectral characteristic data of different lettuce varieties; (A) Radar chart for evaluating the performance of the four models based on the indicators of accuracy, F1, sensitivity, and precision. The range of 0.5 - 1 in the figure represents the magnitude range of the four evaluation index values, (B) Confusion matrix of the BP model, (C) Confusion matrix of the LDA model, (D) Confusion matrix of the RF model, (E) Confusion matrix of the SVM model.
[0045] Figure 17 Correlation analysis between spectral characteristics and morphological characteristics among seeds of different lettuce varieties. Pearson correlation analysis is performed on spectral data, and Mantel test is performed between spectral and morphological characteristic data. In the Pearson correlation analysis of spectral data, red indicates positive correlation, blue indicates negative correlation, gray indicates no correlation, and the darker the color, the stronger the correlation. The size of the circle represents the significance of the correlation; while in the Mantel test of spectral and morphological characteristics, orange indicates significant difference and green indicates no significant difference; Mantel's r represents the correlation between two distance matrices. Generally, if Mantel's r is positive, it indicates a positive correlation between the two distance matrices; if Mantel's r is negative, it indicates a negative correlation between the two distance matrices.
[0046] Figure 18 Radar chart for evaluating the performance and prediction confusion matrix of the seed morphological characteristics combined with spectral characteristic data of different batches of lettuce varieties on the constructed LDA and BP models; (A) Radar chart for evaluating the performance of the LDA model based on the indicators of accuracy, F1, sensitivity, and precision. The range of 0.2 - 1 in the figure represents the magnitude range of the four evaluation index values, (B) Prediction confusion matrix of the LDA model, (C) Radar chart for evaluating the performance of the BP model based on the indicators of accuracy, F1, sensitivity, and precision. The range of 0.2 - 0.5 in the figure represents the magnitude range of the four evaluation index values, (D) Prediction confusion matrix of the BP model.
[0047] Figure 19 Two-dimensional scatter score chart of the LDA model based on the combined data of morphological and spectral characteristics of seeds of 15 lettuce varieties in different batches. Detailed implementation mode
[0048] The following examples are further illustrations of the present invention rather than limitations thereof.
[0049] Example 1
[0050] 1. Materials and methods
[0051] 1.1 Select representative samples
[0052] Seeds of 15 common lettuce varieties in the domestic market are included, namely: Hanging Silk Red (H.S.Red), Sinovac No.5 Lettuce (Sinovac), Four Season Green Lettuce (F.S.G.L), Tender Crisp Head Lettuce (T.C.H.L), Ruili Head Lettuce (R.H.L), Cabbage Lettuce (C.L), Italian Lettuce (Itlian.L), Butterhead Lettuce (B.L), Iceberg Lettuce (Iceberg.L), American Fast-Growing (A.F.G), Fast-Growing (F.G), Four Season Mustard Greens (F.S.M.G), Peacock Lettuce (P.L), Cos Lettuce (Cos.L), Aromatic Red Mustard (A.Red.M).
[0053] Seeds of all lettuce varieties were uniformly purchased in May 2024. Among them, the seed batches for the training group and the test group used to build the model were produced in the second half of 2023, and the seed batches for the model prediction group were produced in the first half of 2024. The germination rates of the above lettuce seeds of each variety were tested to be above 85%, and the seeds to be tested were stored in a 4°C refrigerator for later use.
[0054] 1.2 Multispectral image acquisition of different lettuce seeds
[0055] In this example, a Videometer Lab4 (Videometer, Denmark) multispectral imaging system is used to obtain seed multispectral image information. As Figure 1As shown, the sample needs to be placed under a hollow sphere with a high-definition camera configured at the top. When acquiring images, the sphere will descend to cover the sample below, thus creating uniform illumination conditions with minimal shadow and specular reflection. The sample is irradiated at 19 specific wavelengths (from 365 nm (UV) to 970 nm (NIR)) by 19 flash diodes (LEDs) placed in the middle, and monochromatic images are acquired. This instrument needs to be calibrated before each use, that is, geometric correction and optical frequency measurement are continuously performed using white, black, and white dotted line plates. The image size is 2192×2192 pixels, with a resolution of 40 μm / pixel and 32 bits / pixel. At the same time, we selected seeds of healthy lettuce varieties to obtain multi-spectral image information. A total of 300 seeds of each variety were used. Each time, 100 seeds (repeated three times) were placed in a disposable petri dish. Among them, 210 seeds were randomly selected as the training set, and the remaining 90 seeds were used for the independent test set. The seeds should be placed evenly in the petri dish without touching each other, ensuring the accuracy of extracting the characteristic values of individual seeds in the later stage.
[0056] 1.3 Multi-spectral image analysis
[0057] In the presented spectral images, in addition to the main body of the seeds, there will be some impurities and backgrounds. These irrelevant elements need to be identified and separated in the Blob Toolbox, and then the morphological and spectral information of each seed (blobs) is extracted. The morphological characteristic indicators selected in this embodiment include area, length, width, compact circle, compact ellipse, BetaShape_a, BetaShape_b, vertical skew, CIELab L*, CIELab a*, CIELab b*, saturation, and hue. The detailed content of each feature is shown in (Table 1). The extracted seed spectral characteristics represent the average intensity of the reflected light at each single wavelength calculated from all seed pixels in the image. Subsequently, the data is classified and sorted for further analysis.
[0058] Table 1 Statistical table of morphological characteristic information extracted based on the Videometer Lab multi-spectral imaging system
[0059]
[0060]
[0061] 1.4 Statistical analysis of multi-spectral imaging technology data
[0062] Using principal components analysis (PCA), an unsupervised algorithm commonly used for data dimensionality reduction that can reduce the dimensionality of data without the need for labels, the multispectral data of lettuce seeds of different varieties was analyzed. PCA transforms the original data set into a new set of uncorrelated variables, called principal components (PCs), through linear transformation, where PC1 represents the highest variance, PC2 represents the second highest variance, and so on, thereby extracting important information and data frameworks from high-dimensional data. In addition, we also used the Normalized Canonical Discriminant Analysis (nCDA) built into Videometer Lab software version 3.14, a statistical analysis method also used for data dimensionality reduction and classification. It maximizes the differences between groups and reduces the differences within groups, and is used to explore the spectral data structure of lettuce seeds of different varieties (Bartoli et al., 2022). Among them, PCA uses the stats package (3.5.0) in R (4.1.3) software for data analysis, and the "ggplot2" package (3.3.3) to achieve plotting. nCDA uses the candisc package in R (4.1.3) software for two-dimensional analysis, and the "ggplot2" package (3.3.3) to achieve plotting, and uses the Transformation Builder function built into VideometerLab v3.14 software to perform image transformation on the nCDA and PCA results of all lettuce seeds.
[0063] Based on the multispectral data extracted from lettuce seeds, algorithms such as Support Vector Machine (SVM), Random Forest (RF), Back Propagation (BP) neural network algorithm, and Linear Discriminant Analysis (LDA) were used to attempt to distinguish different lettuce varieties. SVM is a kernel method that divides high-dimensional data by finding the optimal hyperplane and has been successfully and widely applied to multivariate function prediction and non-linear classification, and SVM classifiers have obtained high-accuracy results in most applications, especially for face recognition and disease recognition applications (Abdullah et al., 2021). RF, as an ensemble learning algorithm, is a classification method recently attributed to remote sensing in the field of meta-classifiers. As a meta-classifier, it uses decision trees as the basic classifiers and combines multiple decision trees to create diverse and accurate models. Nowadays, it is widely applied to machine learning tasks in many fields, such as classification, regression, and feature selection. This algorithm reduces overfitting by randomly selecting feature subsets and data samples for each decision tree, thereby reducing the generalization error existing in the model and reducing the computational cost (Rodriguez-Galiano et al., 2012; Lundberg et al., 2018). BP, as a widely applied neural network algorithm, is a multi-layer feedforward network trained by error backpropagation. As a supervised learning algorithm, during the process of optimizing the loss function or objective function, this algorithm solves the gradients of the parameters participating in the operation and then uses gradient descent to update the weights until the desired output effect is achieved (Jia et al., 2023). LDA is also a common supervised machine learning algorithm. Different from PCA, its central idea is simply "the within-class variance is minimized and the between-class variance is maximized after projection", so as to achieve the maximum discriminant ability. Therefore, it is mainly used for the classification and prediction of samples to be tested (Yu et al., 2021). In the embodiments, the parameter settings and result analysis visualization of all machine learning classification models were completed through R (4.1.3) software and the online data analysis platform SPSSPRO (version 1.1.26) https: / / www.spsspro.com . Completed.
[0064] Before data analysis, the validation set method was used to randomly split the data into 70% training data and retain 30% test data. At the same time, in order to avoid the problem of model overfitting, the ten-fold cross-validation and random search methods were used to optimize the relevant parameters of the BP, SVM, and RF models to ensure their stability and effectiveness to the greatest extent. In addition, accuracy, precision, recall, and F1 were used as model evaluation indicators. The calculation equation is as follows:
[0065] Among them, TP, FP, TN and FN represent true positive, false positive, true negative and false negative, respectively (Jia et al, 2023).
[0066] Finally, the combined correlation diagrams of morphological indexes and spectral indexes, and spectral indexes and spectral indexes of different varieties of lettuce seeds were produced using the linkET package in R (4.1.3) software.
[0067] 2. Results
[0068] 2.1 Lettuce varieties and their seed morphological characteristics
[0069] We collected morphological images of the vegetative organs of 15 lettuce varieties at their best edible stage ( Figure 2 ). According to morphological characteristics and growth habits, they can be divided into the following types: HSRed, Sinovac and FSGL are stem lettuces, and the rest are leaf lettuces. Among leaf lettuces, TCHL, RHL, CL, Itlian.L, BL, Iceberg.L, AFG and FG are loose-leaf lettuces, and FSMG, PL, Cos.L and A.Red.M are upright lettuces. The 15 lettuce varieties present different morphological characteristics. The seeds (seed coats) of the 15 lettuce varieties appear in two colors under RGB light source, among which CL, TCHL, Itlian.L, FSMG, BL, Iceberg.L, Cos.L have white seed coats, and the other seeds have black seed coats ( Figure 3 ). Therefore, it is difficult to directly distinguish different varieties of seeds with the naked eye under RGB light source.
[0070] 2.2 Morphological characteristics of lettuce seeds of different varieties under multi-spectral
[0071] Fourteen morphological characteristics (Table 1) were screened from the MSI images of seeds of 15 lettuce varieties and compared among the varieties. The results showed that all measured characteristics within the 95% confidence interval were significantly different among groups. Among them, in terms of shape characteristics, the values of BetaShape_a and BetaShape_b of iceberg lettuce were the highest among groups, while the values of length-width ratio, circularity, and ellipticity were the lowest among groups; the value of BetaShape_b of Kexing No. 5 was the lowest among groups, and the values of length-width ratio, circularity, and ellipticity were the highest among groups; the value of vertical skewness was the highest among groups for Sisiju Youmai; at the same time, among the 15 lettuce varieties, except for the value of ellipticity being relatively close within the group, the distributions of other characteristics within the group were relatively loose( Figure 4 ). There were also significant differences in binary characteristics among the 15 lettuce varieties. Among them, the area and width of Kexing No. 5 seeds were the largest among groups, while the area, length, and width of American Dashusheng were significantly lower than those of most other varieties, which was also consistent with the actual size of the seeds( Figure 5 ). In addition, there were differences in color characteristics among the 15 lettuce varieties. The group mean values of CIELab L and hue value of butterhead lettuce; CIELab A and saturation of Guasihong; CIELab B and saturation of Sisijing lettuce were the highest among groups and significantly higher than those of other varieties. It is worth mentioning that the CIELab L value of all lettuce seeds with white seed coats was very significantly higher than that of black seed coat seeds, probably because white seed coat lettuce seeds reflect light sources stronger, thus increasing the seed brightness( Figure 6 ).
[0072] 2.3 Spectral characteristics of lettuce seeds of different varieties
[0073] The average reflection spectra (365 - 970 nm) among different lettuce varieties are as Figure 7As shown, the overall curve is relatively smooth, and the average spectral reflectance of all lettuce varieties increases with the increase of the irradiation wavelength. The change trends are similar, and the reflectance differences among different varieties at the same wavelength are significant (Table 2). Generally speaking, the average reflectance of the 7 lettuce varieties with white seed coats at each wavelength band is significantly higher than that of all lettuce with dark seed coats. Among them, the reflectance of romaine lettuce and iceberg lettuce in the ultraviolet and far-infrared wavelength bands of 365 - 970 nm has always been the highest among groups. The average reflectance of ice plant lettuce has always been between that of white-seed-coated and dark-seed-coated lettuces when it is at 400 nm - 970 nm. In the NIR (780 - 970 nm), we found that the spectral reflectance of peacock lettuce is significantly lower than that of other varieties (P<0.05). At the same time, in the spectral range of 365 nm - 780 nm, the differences among dark-seed-coated lettuce seeds are not as significant as those of white-seed-coated lettuce seeds. Therefore, from the spectral data results, it can be seen that multi-spectral data has a certain effect on the identification of seeds with different seed coat colors, especially for white-seed-coated seeds with excellent results, but the classification effect on dark-seed-coated seeds is not ideal.
[0074] Table 2 Average spectral reflectance of seeds of 15 lettuce varieties in 19 wavelength bands
[0075]
[0076]
[0077] Continued Table 2
[0078]
[0079]
[0080] Continued Table 2
[0081]
[0082] Further, principal component analysis and nCDA analysis were performed on the multi-spectral information of lettuce seeds of different varieties. First, the outlier and standardization processing were carried out on the seed spectral data to ensure the accuracy of the analysis. The results of the principal component analysis combining morphological features and spectral features showed that the first two principal components could explain 92% and 5.44% of the original variance among seeds respectively. Among them, 15 kinds of lettuce seeds were divided into 3 categories. The within-group differences of most seeds were larger than the between-group differences, so there were still overlapping phenomena for many varieties (A in Figure 8 ). For morphological features, the first two principal components could explain 99.95% of the variation among varieties, which were 96.11% and 3.84% respectively. Among them, peacock lettuce and Dashusheng were somewhat separated from other varieties and did not completely overlap ( Figure 8B) in it. For spectral features, the variance interpretation rates of the first two principal components are 96.81% and 2.83% respectively. However, it is still impossible to effectively distinguish 15 kinds of lettuce seeds only by using spectral feature data. Nevertheless, it has a good effect on distinguishing seeds with different morphological sizes ( Figure 8 C) in it. Therefore, whether it is the PCA score plot of morphological and spectral data or the PCA score plot of combined data, it is impossible to completely separate different lettuce seeds effectively.
[0083] As a most commonly used unsupervised exploratory multivariate data analysis technique, PCA performs principal component analysis on the morphological features and spectral data extracted from all seeds. PCA mainly focuses on maximizing the variance of variables rather than the differences between groups. Therefore, its analysis effect on samples with insignificant differences between groups is poor (ElMasry et al, 2019; Hu et al., 2020). In this regard, we also adopted normalized canonical discriminant analysis (nCDA), a supervised data analysis method for data dimensionality reduction and classification, to explore the classification of combined features, spectral features, and morphological features of different lettuce seeds. The results showed that for morphological feature data, Can1 and Can2 of nCDA analysis explained 93.26% and 2.62% of the variation respectively. All varieties were classified into two parts. Among them, ice leaf lettuce, romaine lettuce, Guasi red, Kexing No. 5, Dasisheng, and Ruili heading lettuce did not show complete overlap, while the overlap rate was relatively high among the remaining varieties ( Figure 9 C) in it. For spectral feature data, Can1 and Can2 of nCDA explained 90.39% of the variation. Among them, American Dasisheng, Guasi red, butterhead lettuce, romaine lettuce, and ice leaf lettuce did not show complete overlap, and the overlap rate was relatively high among other varieties ( Figure 9 B) in it; finally, for the combined data of morphological features and spectral features, Can1 and Can2 of nCDA analysis explained 83.52% of the variation, and all varieties were classified into four categories. Among them, Siqing, Kexing No. 5, peacock lettuce, Dasisheng, and Ruili heading lettuce were in one category; Italian lettuce, tender and crispy heading lettuce, and Siji youmaicai were in one category; Guasi red, Qingxiang youmaicai, and American Dasisheng were in one category; heading lettuce, romaine lettuce, butterhead lettuce, and ice leaf lettuce were in one category, and there was different degrees of overlap among the varieties in each part ( Figure 9 A) in it. Generally speaking, the effect of nCDA analysis is better than that of PCA.
[0084] To more intuitively display the classification results, we randomly selected 100 seeds from each lettuce variety and performed nCDA and PCA visualization on the seed multispectral images. From Figure 10It can be seen that the RGB images captured by multispectral can observe the color and size changes of different seed categories. However, these images fail to depict other differences in the seeds of different lettuce varieties. Under the RGB light source, lettuce seeds are divided into two types, black and white, due to the different colors of their seed coats. The results of nCDA imaging show that the nCDA values of Ruili heading lettuce, Dasusheng, Guasihong, Qingxiangyoumaicai, Kongque lettuce, American Dasusheng, and Kexing No. 5 are relatively low, so the images show more blue areas. While the nCDA values of Nencuijieqiu lettuce, Italian lettuce, Sijiyoumaicai, Iceberg lettuce, Heading lettuce, Roman lettuce, and Cream lettuce are higher, and the images present different degrees of red areas. Among them, Roman lettuce, Iceberg lettuce, and Italian lettuce have more red areas. The PCA results are opposite to those of nCDA, that is, the seeds of the eight lettuce varieties that show more blue areas in nCDA all show more red areas under PCA. Among them, Kexing No. 5, Dasusheng, and Sijiqing lettuce have the most red areas; while the seeds that show more red areas in nCDA all turn blue in PCA, and among them, Iceberg lettuce, Sijiyoumaicai, and Nencuijieqiu lettuce have deeper blue colors. Generally speaking, in the nCDA analysis, the differences among lettuce seeds of the same seed coat color show different effects due to the seed varieties, but the differences among seeds of different seed coat colors are very significant. The PCA results are similar to those of nCDA, but the difference among lettuce seeds of the same seed coat color is slightly worse than that of nCDA.
[0085] Therefore, both nCDA and PCA analyses can distinguish lettuce varieties with different seed coat colors and seed sizes. However, for lettuce seeds with very similar sizes and basically the same seed coat colors, both nCDA and PCA analyses may not be able to make an intuitive assessment. To find a more effective classification method, we further established BP, LDA, RF, and SVM classification models using the morphological or spectral feature data extracted from the seeds of 15 lettuce varieties (n = 4500 seeds) obtained from multispectral imaging. 2.4 Identification models based on the morphological features, spectral feature data, and combined data of different lettuce variety seeds
[0086] When using the morphological feature data of 15 lettuce varieties, the indicators of the four models all decreased to varying degrees. The classification accuracies of RF, BP, LDA, and SVM on the test set are 0.592, 0.683, 0.633, and 0.654 respectively, and the F1 values are 0.591, 0.683, 0.631, and 0.648 respectively. The sensitivity range of the four models is from 0.591 to 0.683, and the precision rate reaches the highest value of 0.687 in the BP model ( Figure 11in A). The differences in the effects among the four models are small, and the model effects are poor. The results of the confusion matrix of the test group show that the recognition rates of the four models for crisp-heading lettuce, Italian lettuce, heading lettuce, and Ruili heading lettuce are only about 10%-50%, and the recognition error rates are relatively high; for other varieties, the recognition accuracies of the four models are between 60%-80%( Figure 11 in B-E). On the training set data of the morphological characteristics of 15 lettuce varieties, all performance indicators of the four models are higher than those of the test group, between 0.6-0.7( Figure 12 in A). The results of the confusion matrix of the training group show that the misrecognition rates of the four models for crisp-heading lettuce, Italian lettuce, heading lettuce, and Ruili heading lettuce are still about 50% on average, and the overall recognition rate is slightly higher than that of the test group( Figure 12 in B-E).
[0087] When building models based on the spectral characteristic data of 15 lettuce varieties, the classification accuracies of BP and LDA on the test set are 0.616 and 0.737 respectively, while the accuracies of RF and SVM are only 0.321 and 0.294. The F1 is the highest in the LDA model, which is 0.738, but only 0.268 in SVM. The sensitivity ranges of the four models are from 0.294 to 0.737, and the highest value of the precision rate is in the LDA model, reaching 0.744( Figure 13 in A). It can be found from the four evaluation indicators that the spectral characteristic data is weaker in the effects of RF and SVM than in the two models constructed by the morphological characteristic data. At the same time, further analyzing the results of the confusion matrix of the test group of the four models of RF, BP, LDA, and SVM, the recognition rate of RF for butterhead lettuce reaches about 70%, and the classification recognition accuracies of other varieties are relatively low. The results of SVM are similar to those of RF. Except that the recognition rates of butterhead lettuce and ice plant lettuce reach about 60%, there are also problems with low recognition rates for other varieties, and even the recognition accuracy rates for Guasi red and all-the-year-round lettuce are 0. BP and LDA have difficulty distinguishing between Italian lettuce and heading lettuce, and the correct discrimination rates of other varieties are basically between 60%-80%( Figure 13 in B-E); for the training group data, the four model indicators of RF are all improved compared with the test group, approaching about 0.5, but generally speaking, the discrimination effect is still not good. The model indicators of SVM become even lower, and it is speculated that there may be problems such as underfitting or other parameters in the model( Figure 14 in A), and the recognition rates of the training group confusion matrix results for each lettuce variety are basically the same as those of the test group, and no significantly different model results are seen( Figure 14 in B-E).
[0088] By analyzing the combined data of morphological features and spectral features, we can see that the classification accuracies of RF, BP, LDA, and SVM on the test set are 0.727, 0.896, 0.927, and 0.736 respectively, and the F1 values are 0.722, 0.896, 0.927, and 0.734 respectively. The sensitivity range of the four models is from 0.727 to 0.927, and the precision rate is the highest in the LDA model, reaching 0.928 ( Figure 15 A in Figure 15 ). Therefore, the overall performance of the BP and LDA models is the best. Further comprehensive analysis of the confusion matrices of the four models of RF, BP, LDA, and SVM shows that ( Figure 16 B - E in
[0089] ) there are varying degrees of problems with their accuracy in distinguishing tender and crispy head lettuce from Italian lettuce, Dashusheng lettuce, and peacock lettuce. At the same time, the correct discrimination rates of RF for head lettuce and Ruili head lettuce are only 46.67% and 43.33% respectively, while those of SVM reach 51.1% and 67.78% respectively. The recognition rate of LDA for Ruili head lettuce is only 72.2%, and the recognition error rate for other lettuce varieties is about 15%. For BP, except for Dashusheng, Kexing No. 5, Roman lettuce, and ice leaf lettuce, the recognition rates for other varieties are between 85% and 100%. On the other hand, the correct discrimination rates of RF for head lettuce and Ruili head lettuce on the training set data reach 61% and 49% respectively, and the results of LDA, SVM, and BP are similar to those on the test set ( Figure 16 A in
[0089] ), and the classification effect of the LDA and BP models is relatively excellent overall.To explore whether there is a correlation between the morphological characteristics and spectral characteristic data of lettuce seeds of different varieties, Pearson analysis was first performed among 14 morphological characteristic indexes, and Mantel tests were conducted between 14 morphological characteristic indexes and 19 wavelengths. The results showed that there were positive and negative correlation relationships to varying degrees among the 14 morphological characteristic indexes. Among them, among the binary characteristics, area was extremely significantly positively correlated with length (P<0.01) and extremely significantly positively correlated with width (P<0.01); among the shape characteristics, width-to-length ratio was extremely significantly positively correlated with elliptical compactness, elliptical compactness was extremely significantly positively correlated with circular compactness, and BetaShpe_a was extremely significantly positively correlated with BetaShpe_b (P<0.01); while among the color characteristics, CIELab L* was extremely significantly negatively correlated with CIELab a* and saturation (P<0.01), and hue was extremely significantly negatively correlated with CIELab b* and saturation (P<0.01). At the same time, except for the lack of correlation between shape characteristics and color characteristics, between binary characteristics (such as area), and some indexes between color characteristics, there were positive or negative correlations to varying degrees among most indexes. And through Mantel tests, the first three wavelengths in the ultraviolet and visible light bands and the first wavelength in the far-infrared light band were all significantly positively correlated with binary characteristics. At the same time, the last wavelength in the ultraviolet, visible, and far-infrared light bands was all significantly positively correlated with hue, and there was no significant correlation between other wavelengths and the 14 morphological characteristic indexes ( Figure 17 ).
[0090] 2.5 Variety identification of lettuce seeds of different batches based on LDA and BP models
[0091] Seeds of the same variety in different batches often have certain differences due to the influence of cultivation and storage environments. At the same time, to further determine the stability of the constructed model, we used lettuce seeds of each variety in different batches (a total of 750 seeds, 50 seeds for each experimental variety) to conduct a prediction experiment on the model. From Figure 15 A in it can be seen that based on the combined data of the morphological characteristics and spectral characteristics of lettuce seeds of each variety, the recognition accuracies of the LDA and BP models are both above 90%, belonging to relatively excellent classification models. Therefore, we planned to collect the combined data of the morphological characteristics and spectral characteristics of 15 lettuce variety seeds in different batches and then input them into the previously constructed LDA and BP models, and conduct a variety classification prediction experiment. The results found that the classification prediction effect of the LDA model on lettuce variety seeds of different batches was still very excellent, and all aspects of evaluation indexes of the model reached above 0.93. Except for head lettuce and romaine lettuce, the prediction accuracy of the LDA model for other lettuce varieties was basically above 90%, and even the prediction accuracy for ice leaf lettuce and Qingxiang oil wheat lettuce reached 100%( Figure 18) To make the results clearer, we plotted a two-dimensional scatter score plot of the LDA model based on the combined data of this batch of lettuce seeds ( Figure 19 ). For the combined data of morphological features and spectral features, LD1 and LD2 of LDA explained 91.3% of the variation. Except for the relatively high classification overlap rates among several varieties such as Sijiqing, American Big Fast Growth, Guasihong, Romaine Lettuce, Italian Lettuce, and Head Lettuce, other varieties basically did not show complete overlap, and the classification effect of the overall varieties was better than that of the previous nCDA and PCA analyses. However, the classification effect of the BP model on the lettuce seeds of different varieties in different batches became worse. It can be seen from the confusion matrix that the evaluation indicators of the BP model were very low, and the precision was even only 0.09, and the model predicted most of the seeds as Peacock Lettuce ( Figure 18 ). In summary, through the model prediction experiment, it can be seen that the LDA model based on the combined data of morphological features and spectral features of lettuce seeds has excellent classification ability in the training group, test group, and prediction group, and the stability of the model is relatively high.
Claims
1. A method for identifying lettuce seed varieties based on multispectral imaging technology and machine learning algorithm, characterized in that: The following steps are involved: S1. Select representative lettuce seeds of different varieties for model construction; S2. Collect multispectral image information of different varieties of lettuce seeds; S3. Obtaining morphological feature data of different varieties of lettuce seeds under multispectral images; S4. Obtaining spectral feature data of multispectral images of lettuce seeds of different varieties; S5. constructing an LDA model using the acquired morphological characteristic data and spectral characteristic data of lettuce seeds of different varieties; S6. According to steps S2-S4, the morphological feature data and spectral feature data of the lettuce seeds of the variety to be identified are collected, and the data are substituted into the LDA model constructed in step S5 to identify the lettuce seed variety.
2. The method according to claim 1, characterized in that The different varieties used to construct the model in step S1 are red silk lettuce, Kexing No. 5 lettuce, four-season green lettuce, tender and crisp head lettuce, Ruili head lettuce, head lettuce, Italian lettuce, cream lettuce, ice leaf lettuce, American fast-growing, fast-growing, four-season oil lettuce, peacock lettuce, Roman lettuce, and fragrant red oil lettuce.
3. The method according to claim 1, characterized in that The step S2 is to use the Videometer Lab4 multispectral imaging system to obtain a monochrome image of the lettuce seeds under 19 specific wavelengths of 365nm-970nm, thereby obtaining multispectral image information.
4. The method according to claim 3, characterized in that The 19 specific wavelengths of 365nm-970nm are 365nm, 405nm, 430nm, 450nm, 470nm, 490nm, 515nm, 540nm, 570nm, 590nm, 630nm, 645nm, 660nm, 690nm, 780nm, 850nm, 880nm, 940nm and 970nm.
5. The method according to claim 3, characterized in that: The Videometer Lab4 multispectral imaging system was calibrated before each use, and white, black, and white dotted line plates were used continuously for geometric correction and optical frequency measurement. The image size was 2192×2192 pixels, the resolution was 40 μm / pixel, and 32 bits / pixel.
6. The method according to claim 1, characterized in that The morphological features in step S3 include area, length, width, compact circle, compact ellipse, BetaShape_a, BetaShape_b, vertical skew, CIELab L*, CIELab a*, CIELab b*, saturation and hue.
7. The method according to claim 1, characterized in that The spectral characteristic data in step S4 are the average spectral reflectance of each variety of lettuce seeds under irradiation at wavelengths of 365nm, 405nm, 430nm, 450nm, 470nm, 490nm, 515nm, 540nm, 570nm, 590nm, 630nm, 645nm, 660nm, 690nm, 780nm, 850nm, 880nm, 940nm, and 970nm, respectively.
8. The method according to claim 1, characterized in that The step S5 is to randomly divide the data into 70% training set data and 30% test set data using the validation set method before data analysis, and then construct an LDA model based on the combined data of morphological feature data and spectral feature data of different varieties of lettuce seeds. The model parameter setting and result analysis visualization are completed through R language software version 4.1.3 and online data analysis platform SPSSPRO version 1.1.26.