Chrysanthemum salt tolerance high-throughput identification method
By using RGB image processing and machine learning algorithms, a chrysanthemum salt tolerance prediction model was established, which solved the problems of time-consuming, labor-intensive and subjective bias in traditional chrysanthemum salt tolerance assessment, and realized high-throughput, non-destructive evaluation of chrysanthemum salt tolerance.
Patent Information
- Application Number
- CN202511690768.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2026-03-10
AI Technical Summary
Traditional methods for assessing salt tolerance in chrysanthemums are time-consuming, labor-intensive, have low throughput, are subject to subjective bias, and require destructive sampling, making it difficult to achieve high-throughput, non-destructive, and objective salt tolerance evaluation.
Using RGB image processing technology, the salt tolerance D value is calculated by dimensionality reduction through principal component analysis. Combined with morphological and textural feature parameters, a salt tolerance prediction model is established using the random forest algorithm to realize the automated classification of salt tolerance levels of chrysanthemum germplasm resources.
It enables rapid and accurate assessment of salt tolerance in chrysanthemums, reduces the input of human and material resources, avoids the influence of subjective human factors, and is suitable for high-throughput screening of chrysanthemums and their closely related species.
Smart Images

Figure CN121640200A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of flower crop germplasm identification and evaluation and breeding, and particularly relates to a high-throughput identification method for salt tolerance of chrysanthemum. BACKGROUND
[0002] Chrysanthemum morifolium Ramat. plays an important role in the world flower industry, and chrysanthemum is sensitive to salt stress. Salt stress in cultivation often causes chrysanthemum plants to be dwarf, leaf yellowing and wilting, flowering process to be blocked, and even whole plant death. In production, soil secondary salinization has become a problem restricting chrysanthemum production. Screening salt-tolerant germplasm and breeding salt-tolerant varieties are effective ways to solve this problem. At present, traditional chrysanthemum salt tolerance evaluation mainly relies on artificial observation, such as leaf wilting degree, mortality, and measurement of stress index of salt-related traits such as plant height, aboveground and underground fresh weight, and aboveground and underground dry weight. The comprehensive evaluation and salt tolerance grade are divided by using the membership function method or the principal component analysis method. Among them, the biomass comprehensively reflects the performance of plants in aspects such as configuration, light absorption and carbon assimilation, and is an important indicator for measuring plant growth conditions and environmental adaptability. However, the traditional phenotype acquisition method has the disadvantages of time-consuming and laborious, low throughput, subjective bias, and the need for destructive sampling. Therefore, it is urgent to develop a simple, non-destructive, high-throughput and objective chrysanthemum salt tolerance phenotype evaluation technology to screen salt-tolerant excellent germplasm. SUMMARY
[0003] The purpose of the present application is to provide a high-throughput identification method for salt tolerance of chrysanthemum, which is used for quickly and accurately evaluating the salt stress response ability of chrysanthemum materials and estimating the aboveground biomass, and solves the problems of time-consuming and laborious, low throughput, subjective bias and the need for destructive sampling in the traditional phenotype acquisition method.
[0004] The technical scheme of the present application is a high-throughput identification method for salt tolerance of chrysanthemum, comprising the following steps: (1) Selecting multiple chrysanthemum germplasm resources, setting different concentrations of sodium chloride solution for salt stress treatment, and calculating a comprehensive salt tolerance D value by principal component analysis method from multiple salt-related traits; (2) After the salt stress phenotype appears, collecting RGB images of four side angles and one overhead angle of each plant, and performing image color calibration and standardized naming; (3) Segmenting the image in the excess green and HSV color space to obtain a plant binary image, and extracting two types of digital feature parameters of morphology and texture, and analyzing the reliability of the feature parameters by using a linear regression model; (4) The representative features which are significantly responsive to salt stress and genetically stable are screened from the digital feature parameters by independent sample t-test, variability analysis, feature importance analysis and generalized heritability analysis; (5) The screened representative features are divided into multiple feature subsets, and a random forest algorithm is used to train a salt tolerance prediction model with the representative features as independent variables and the salt tolerance D value as the dependent variable; (6) Based on the salt tolerance prediction model, the importance of each feature is evaluated, and the optimal feature combination with the smallest prediction error is selected to automatically divide the chrysanthemum germplasm resources into salt tolerance grades.
[0005] Further, in step (1), the salt stress treatment specifically includes: setting a control group, a 100mM NaCl low concentration treatment group and a 200mM NaCl high concentration treatment group; the calculation of the salt tolerance D value includes: standardizing the original index matrix of 22 salt tolerance related traits, obtaining the eigenvalue and eigenvector of the correlation coefficient matrix, calculating the principal component score and performing forward normalization processing, and finally taking the contribution rate of each principal component as the weight to calculate the weighted average value to obtain the D value.
[0006] Further, in step (3), the morphological feature parameters include total projection area, convex hull area, perimeter, green projection area, plant height, plant width, green projection area ratio, height-width ratio, side view total projection area and edge area ratio, and perimeter area ratio; the texture feature parameters include histogram texture features, gray level co-occurrence matrix features and gray level-gradient co-occurrence matrix features extracted based on G, H and V three color components.
[0007] Further, in step (3), the determination coefficient R² and its significance are analyzed by using a linear regression model.
[0008] Further, in step (4), the specific standards for screening are as follows: first, through independent sample t-test, the feature parameters with no significant difference P>0.05 between the treatment group and the control group are removed; then, variability analysis is performed, and the feature parameters with a coefficient of variation CV average value less than 2% are removed; then, feature importance analysis is performed, and the feature parameters with a normalized importance average value less than 0.3 are removed by using a multilayer perceptron model combined with a SHAP method; finally, generalized heritability analysis is performed, and the feature parameters with a heritability H²<0.6 are removed, and the finally retained feature parameters are taken as the representative features which are significantly responsive to salt stress and genetically stable.
[0009] Further, in step (5), the multiple feature subsets include: an all-feature subset, a morphological feature subset, a histogram texture feature subset, a gray level co-occurrence matrix feature subset and a gray level-gradient co-occurrence matrix feature subset.
[0010] Further, in step (6), the optimal feature combination is determined by ten-fold cross-validation, wherein the optimal feature combination comprises 7 feature parameters for low-concentration salt stress treatment, and the optimal feature combination comprises 17 feature parameters for high-concentration salt stress treatment.
[0011] Further, the specific method of automatic division of salt tolerance grades is as follows: the mean value of the predicted D value of each resource under different salt treatments is calculated, and according to the standard deviation distance between the mean value and the overall mean value, the resource is divided into five grades of extremely salt-tolerant I, relatively salt-tolerant II, salt-tolerant III, non-salt-tolerant IV and extremely non-salt-tolerant V.
[0012] Beneficial effects: Compared with the prior art, the present application has the following remarkable advantages: the present application is suitable for all chrysanthemum and its related species plants, and can be used as a reference for other potted crops; the feature parameters extracted by image can replace artificial statistical indicators to quantify the salt tolerance of chrysanthemum and divide the salt tolerance grades, avoiding the influence of human subjective factors on the evaluation of chrysanthemum salt tolerance, and having important theoretical and practical significance for chrysanthemum salt tolerance breeding; the prediction model of chrysanthemum aboveground biomass is established by using machine learning algorithm, which can realize non-destructive dynamic measurement of plants and greatly reduce the labor and material inputs. BRIEF DESCRIPTION OF DRAWINGS
[0013] Figure 1 is the experimental process of the present application.
[0014] Figure 2 is the feature parameter reliability analysis of the present application. Figure 2 in (a) and (b) are the correlation analysis of the plant height and fresh weight collected manually and the plant height and total projection area extracted from the RGB image, Figure 2 in (c) is the high correlation between the feature parameters and the phenotypic characteristics.
[0015] Figure 3 is the filtering and determination of the salt tolerance related feature parameters of the present application. Figure 3 in (a) is the t-test of the treatment group and the control group at the side view and top view stages, *** P<0.001, ** P<0.01, * P<0.05, and ns represents no significant difference, Figure 3 in (b) is the MLP mean value of all feature parameters, Figure 3 in (c) is the coefficient of variation (CV) of all feature parameters, Figure 3 in (d) is the general heritability (H2) of 103 feature parameters.
[0016] Figure 4 is the feature parameter selection responding to salt stress of the present application; and is the principal component analysis of 5 types of feature parameters at different salt concentration treatments of side view and top view angles.
[0017] Figure 5The present invention utilizes the random forest algorithm to predict the optimal prediction model for D values under different salt concentrations and ranks the importance of image feature parameters; wherein (a) and (b) respectively use cross-validation to screen important feature parameters, with the optimal number of feature parameters for low salt concentration treatment being 7 and the optimal number of feature parameters for high salt concentration treatment being 17; (c) and (d) respectively use the screened feature parameters to predict the D values for different salt concentration treatments using the random forest algorithm.
[0018] Figure 6 The present invention provides the optimal model for predicting aboveground biomass under different salt concentrations using the random forest algorithm and the ranking of the importance of image feature parameters; wherein (a) is the optimal model for predicting fresh weight under different salt treatments using all feature parameter datasets, and (b) is the optimal model for predicting dry weight under different salt treatments using all feature parameter datasets. Detailed Implementation
[0019] The technical solution of the present invention will be further described below with reference to the accompanying drawings.
[0020] like Figure 1 As shown, this embodiment of the invention provides a high-throughput phenotypic identification method for salt tolerance in chrysanthemums based on RGB images, comprising the following steps: Step 1: Traditional salt tolerance evaluation and analysis, including the following steps: Step 11: Materials and Field Trial Design Twenty-four chrysanthemum and closely related plant resources from different sources and without direct kinship were selected as experimental materials (Table 1). All materials are preserved at the Chrysanthemum Germplasm Resource Conservation Center of Nanjing Agricultural University. During the vegetative growth season of chrysanthemums, robust and uniformly growing cuttings were collected and propagated in plug trays (V vermiculite:V perlite = 1:1). After normal management and good rooting, they were transplanted into black flowerpots (top diameter × height = 15 × 12 cm) with pure vermiculite as the cultivation environment. After transplanting, all cuttings were placed in plastic turnover boxes (length × width × height = 65 cm × 43 cm × 16 cm) for 4 days to allow them to acclimate. A fixed amount of water (5 L) was applied to each plastic turnover box. The experiment included a control group (CK) and a salt stress treatment group (LS: NaCl 100 mM, HS: NaCl 200 mM). A completely randomized block design was used, with 3 blocks and 5 replicates per block. During the treatment period, the control group was watered with 1 / 4 Hogland nutrient solution as usual, while the salt stress treatment group was watered with the corresponding NaCl solution in addition to the same amount of 1 / 4 Hogland nutrient solution. The treatment was carried out every 5 days.
[0021] Step 12: Materials and Field Trial Design Immediately after transplanting, the height of each plant was investigated, which was the height before treatment. After the appearance of salt stress phenotypes, the mild, moderate and severe salt stress phenotypes were scored according to the leaf color and leaf morphology of the plants, and the yellow leaf rate and plant survival rate were calculated. The score value, i.e. the wilting index (WI) of the plants, was the sum of the leaf color score and the leaf morphology score after salt treatment. The yellow leaf rate was the ratio of the number of damaged leaves (including yellow leaves and dead leaves) in the treatment group to the total number of intact leaves of the plant. Since the control group plants grew healthily, the above indicators were only statistically analyzed for the treatment group plants.
[0022] After the third scoring, the trait indicators of each plant were investigated, including plant height, growth rate, aboveground fresh weight, underground fresh weight, aboveground dry weight, underground dry weight, fresh weight root-shoot ratio, dry weight root-shoot ratio, and 22 other salt tolerance-related traits. In addition to the yellow leaf rate, survival rate and wilting index, the salt tolerance coefficient (%) of each traditional trait was calculated according to the formula: treatment group / control group x 100%.
[0023] The D value of the above indicators was calculated by principal component analysis, and the salt tolerance D value (D) was used as the standard for comprehensive evaluation of chrysanthemum salt tolerance. The calculation formula of the principal component analysis D value is: after standardizing the original index matrix, the eigenvalue λj and the corresponding eigenvector Lj of the correlation coefficient matrix R are obtained, and the principal component Fj = ∑(Lj x Z) is obtained, where Z is the standardized index value. If the principal component is positively correlated with salt tolerance, its standardized score Yj = (Fj - Fmin) / (Fmax - Fmin); if it is negatively correlated, Yj = 1 - (Fj - Fmin) / (Fmax - Fmin). Fmin and Fmax are the minimum and maximum scores of the principal component in all resources. Finally, the contribution rate wj = λj / ∑λ of each principal component is used as the weight, and the weighted average value D = ∑(wj x Yj) of the standardized scores of all principal components of each resource is calculated as the standard for comprehensive evaluation of chrysanthemum salt tolerance; the larger the value, the stronger the salt tolerance. The larger the D value, the better the salt tolerance of the resource.
[0024] Step 13: Traditional salt tolerance evaluation According to the average D value obtained from the two salt concentrations (Table 1), and referring to the grading method of Yan et al. (2020), all materials were divided into 5 salt tolerance grades: among them, 16 materials were extremely salt-tolerant (I), 21 materials were relatively salt-tolerant (II), 132 materials were salt-tolerant (III), 7 materials were not salt-tolerant (IV), and 28 materials were extremely not salt-tolerant (V).
[0025] Table 1 Salt tolerance information of 204 test materials
[0026]
[0027]
[0028]
[0029]
[0030] Step 2: RGB image feature parameter screening and prediction model establishment, including the following steps: Step 21: RGB image acquisition and feature parameter extraction After the appearance of salt stress phenotype, all potted plants were imaged in a 80 cm x 80 cm x 80 cm photographic studio with fixed light source in the room. A Canon EOS 90D camera was used for shooting, and an A2-sized white paper was used as the background, with a black circle with a radius of 1.5 cm marked on each corner, and a 1 cm x 1 cm black square as a ruler. The X-RITE mini 24 color card for color calibration and the two-dimensional code for picture naming were fixed. The two-dimensional code was generated by an online two-dimensional code generator (http: / / www.qinms.com / webapp / qrcode / MultiQrcode.aspx). According to the size of the plants, the appropriate shooting distance was selected and fixed to meet the size requirements of the field of view. In this test, the imaging distance of the overhead camera and the side-view camera was 60 cm and 100 cm, respectively, with a resolution of 6960 x 4640. The overhead camera collected the top phenotype data of the plants, and the side-view camera collected the phenotype data at 0°, 90°, 180° and 270° angles, i.e. 5 RGB images were collected for each plant.
[0031] Image color calibration was performed using ColorChecker Camera Calibration (X-RITE, USA); due to the large number of images collected, we used the pyzbar package of python to decode the two-dimensional code and achieve image renaming. MATLAB R2022b (MathWorks Inc., USA) was used to complete the excess green (E x G) and HSV color space image segmentation and feature parameter extraction Figure 1 Through image processing analysis, a total of 103 feature parameters based on RGB images were extracted, including 10 morphological feature parameters and 93 texture feature parameters. Texture features were calculated under green (G), hue (H) and value (I) components, respectively, and were denoted as G texture, H texture and V texture (Table 2).
[0032] Table 2 Detailed definition of traditional traits and image feature parameters
[0033] Step 22: Reliability Analysis of Characteristic Parameters To verify the reliability of non-destructive digital feature parameters of RGB images, linear regression models were constructed using the `lm` function of the `stats` package in R language, respectively, to correlate manually measured plant height, aboveground fresh weight, and corresponding digital feature parameters. The adjusted coefficient of determination (adj.R²) and its significance were then labeled using the `ggpmisc` package. Figure 2 a and b). The results showed that under low salt stress, the artificial plant height was strongly linearly correlated with the plant height in the RGB images (LS, R 2 =0.91, P <0.001; CK, R 2 =0.84, P <0.001). The fresh weight of the aboveground parts was also significantly correlated with TPA (LS, R 2 =0.25, P <0.001; CK, R 2 =0.64, P <0.001), the correlation was lower than that of plant height.
[0034] The Pearson correlation coefficients between 22 manually measured indicators and 103 numerical feature parameters were calculated using the `corr.test` function of the `psych` package. After FDR correction, the correlation coefficients were expressed as | R 2 |>0.6 and P<0.001 are used as thresholds to construct edge files, which are then imported into Cytoscape v3.10.0 to generate an index-parameter correlation network. Figure 2 c), the network contains 8×63 highly correlated node pairs, showing a strong correlation between 8 manually measured indicators and 63 numerical characteristic parameter indicators. R 2 >0.6, P <0.001).
[0035] The above analysis shows that the measurement results based on RGB images have a strong positive correlation with the results of manual measurement, indicating that the non-destructive feature parameters obtained based on RGB images can, to some extent, replace traditional destructive measurements.
[0036] Step 3: Filtering effective feature parameters, including the following steps: After obtaining 103 feature parameters, the data was filtered using NI LabVIEW 2023 software. Then, the following steps were followed to select the feature parameters responding to salt stress: Step 31: Significance Analysis The `t.test` function from the `stats` package in R software, supplemented by dplyr for data processing, was used to perform independent sampling on all characteristic parameters of the control group, low-concentration salt treatment group, and high-concentration salt treatment group, along with their corresponding lateral and top-view angles. t We conducted tests to assess the effects of salt stress treatment on characteristic parameters, setting the confidence level at 95% to identify characteristic parameters that showed significant differences under salt stress.
[0037] The results show that ( Figure 3 (a) Of the 103 characteristic parameters, 16 parameters showed no significant difference between different salt treatments and the control. P >0.05), including PAR, G_ME, G_SD, G_ST, G_TM, G_CON_ME, G_COR_ME, G_ASM_SD, G_CON_SD, G_IDM_SD, G_GGCM1, G_GGCM4, G_GGCM6, G_GGCM8, G_GGCM14, and G_GGCM15. After removing these characteristic parameters, the remaining parameters showed statistically significant differences under salt stress treatment, indicating that they can effectively reflect the effects of salt stress on Asteraceae plants.
[0038] Step 32: Variability Analysis IBM SPSS version 27 was used to calculate the coefficient of variation (CV) of all characteristic parameters of the control group, low-concentration salt treatment group, and high-concentration salt treatment group and their corresponding side-view and top-view angles to measure the dispersion and discriminative ability of the parameters among different genotypes. Characteristic parameters with a mean CV of less than 2% were removed to ensure that the characteristic parameters have sufficient variability among samples.
[0039] The results showed that ( Figure 3 c) The CVs of different parameters ranged from 0.14% to 263.05%, with an average of 27.16%, indicating significant differences in properties among materials. To ensure the validity of the parameters, 21 characteristic parameters with CVs less than 2% were removed, including G-ME, G-ET, H-UF, V-UF, G-ENT_ME, H-ASM_ME, H-IDM_ME, V-ASM_ME, V-IDM_ME, G-GGCM6, G-GGCM11, G-GGCM13, G-GGCM14, H-GGCM1, H-GGCM4, H-GGCM5, H-GGCM15, V-GGCM1, V-GGCM4, V-GGCM5, and V-GGCM15.
[0040] Step 33: Feature Importance Analysis A multilayer perceptron (MLP) model was constructed using IBM SPSS version 27. All feature parameters of the control group, low-concentration salt treatment group, and high-concentration salt treatment group, along with their corresponding side-view and top-view angles, were trained and run five times to reduce the influence of random initial weights. Subsequently, the DALEX and SHAP methods were combined to extract feature importance values, and the average value of the repeated results was calculated. Feature parameters with a normalized importance average value less than 0.3 were removed, thus retaining features that significantly contribute to the salt stress response.
[0041] The results show that ( Figure 3 (b) The importance of the normalized features ranged from 0.225 to 0.660, indicating that different parameters contributed differently to distinguishing between the control, low-salt, and high-salt treatments. Based on the set threshold, 13 parameters with an average importance less than 0.3 were removed, including PW, TBR, V_ME, V_ST, V_TM, H_ENT_ME, V_ENT_ME, V_CON_ME, V_GGCM2, V_GGCM6, V_GGCM8, V_GGCM13, and V_GGCM15. The retained parameters showed high predictive value in the MLP model and effectively reflected the differences in phenotypic responses under salt stress.
[0042] Step 34: Heritability Analysis The generalized heritability of each characteristic parameter was calculated using the lmer function in the 'lme4' package of R software. H² Estimate to measure the genetic stability of trait parameters across different genotypes and select H² Characteristic parameters ≥0.6 are used for subsequent research.
[0043] The results showed that ( Figure 3 d) The measured characteristic parameters H² The range is from 0.544 to 0.996, with most parameters... H² A value greater than 0.7 indicates that the parameter is strongly controlled by genetic factors and has good genetic stability. Only H_ASM_SD and H_IDM_SD... H ² If the value is less than 0.6, it is rejected.
[0044] By conducting parallel screening of all feature parameters through the above four types of analysis, feature parameters that meet the requirements for salt stress response research in terms of statistical significance, variability, model importance, and genetic stability are finally obtained, providing a technical basis for the accurate identification of salt stress response features and related functional analysis.
[0045] Step 4: Principal component analysis of characteristic parameters in response to salt stress Principal component analysis (PCA) was performed on the 56 valid feature parameters obtained to reflect their salt stress response at the population level. PCA was performed using the `prcomp` function from the R package `factoextra`, and the results were visualized using the `fviz_pca_ind` function with a 90% confidence interval. The feature parameters were categorized into five classes: all feature parameters, morphological feature parameters, histogram texture features, gray-level co-occurrence matrix (GLCM), and gray-level-gradient co-occurrence matrix (GGCM).
[0046] Principal component analysis results show that the first principal component (Dim1) carries 64.8–87.6% and 50.4–89.1% of the total variance in side-view and top-view angles, respectively, concentrating most of the original feature information. Figure 4 Among the features, histogram texture features showed the highest loading on Dim1 from both side views (side view ≈ 0.89, top view ≈ 0.88), indicating that subtle changes in leaf grayscale distribution are the most discriminative source of information in the current image. PCA score distribution showed that normal treatment (CK) and high salt stress (HS) were most clearly separated on the Dim1 axis, with low salt stress (LS) falling in between, forming a continuous gradient from CK to LS to HS. From the side view, all features, histogram texture, and GGCM trends were consistent, while GLCM scatter points were more concentrated at the negative end of Dim1, indicating the clearest boundary between normal control and salt treatment. From the top view, morphological scatter points were more clustered, providing better differentiation between normal and salt-treated plants compared to the other four categories. In summary, side-view GLCM and top-view morphological traits provided the best classification boundaries from different perspectives. These results suggest that texture feature parameters may also be a core potential indicator reflecting the salt stress response of Asteraceae plants.
[0047] Step 5: Confirmation of the optimal model for D-value prediction, including the following steps: Step 51: Dataset Classification and Model Building The filtered feature parameters were divided into five datasets: all feature parameters (ALL), morphological feature parameters (mor-i), histogram texture features, gray-level co-occurrence matrix (GLCM), and gray-level-gradient co-occurrence matrix (GGCM). A random forest (RF) model was used to predict low-salt stress (LS) and high-salt stress (HS). Since the D value is calculated based on the stress index of traditional traits, the stress index (SS / CK) of the feature parameters was used for prediction. 70% of the feature parameters and D values obtained from RGB images were used as the training set, and 30% as the test set. The feature parameters obtained from RGB images were used as independent variables, and the D value as the dependent variable. A five-fold cross-validation method was used to train the random forest machine learning model. The random forest model was implemented in the 'randomForest' R package, and the five-fold cross-validation was implemented in the 'Caret' R package. The root mean square error (RMSE) was used. RMSE Mean absolute error ( MAE ) and coefficient of determination ( R 2 To evaluate the accuracy of model predictions, RMSE and MAE The lower the value, R 2 The higher the value, the better the model's prediction performance.
[0048] The predicted D-values under low salt stress (LS) and high salt stress (HS) conditions were validated. The results show that, under LS conditions, the determination coefficients of each model ( R² The values are all in the range of 0.67–0.74, indicating a moderate level of prediction accuracy. The ALL dataset, which includes all feature parameters, performs best. R² =0.737). Under the HS condition, the predictive performance of each model is significantly improved, and the performance of each dataset is... R² All values exceeded 0.90, indicating that the proposed method can achieve high prediction accuracy under high salt concentration stress.
[0049] Further comparison of performance across different datasets revealed that the ALL dataset consistently demonstrated the highest prediction performance under both LS and HS conditions (respectively...). R² =0.737 and R² =0.912). This result shows that integrating all feature parameters can maximize the retention of salt stress-related information, thereby significantly improving the robustness and generalization ability of the model.
[0050] Table 3 Comparison of the performance of predicting D values based on the random forest model
[0051] Note: The best result is shown in bold.
[0052] Step 52: Selection of characteristic parameters and salt tolerance classification Based on the random forest model, D-value prediction models for low-concentration and high-concentration salt stress were constructed using all feature parameter datasets. The feature parameters were then selected based on the cross-validation curves after performing five repeated 10-fold cross-validation.
[0053] When treating with low-concentration salts, the cross-validation error is minimized when n = 7. Figure 5 a) This indicates that retaining the first 7 feature parameters yields ideal regression results. Based on the importance of the variables, the first 7 feature parameters H.COR_ME, H.COR_SD, GPAR, H.GGCM10, H.CON_SD, H.GGCM8, and H.CON_ME are selected. The machine learning model is then used to predict the D value again. Under the random forest model, the prediction accuracy is 0.737, which is quite high. Figure 5 c).
[0054] When treating with high-concentration salts, the cross-validation error is minimized when n = 17. Figure 5 (b) This indicates that retaining the first 17 feature parameters yields ideal regression results. Based on the importance of the variables, the first 17 feature parameters H.SD, H.CON_SD, H.GGCM8, H.CON_ME, GPAR, H.TM, H.ST, H.COR_ME, H.GGCM14, GPA, H.GGCM10, H.ME, H.ENT_SD, H.COR_SD, H.GGCM6, V.COR_ME, and V.SD are selected. The machine learning model is then used to predict the D value. Under the random forest model, the prediction accuracy is 0.911, which is quite high. Figure 5 d). The results show that the prediction accuracy using a small number of feature parameters is almost equivalent to that using all feature parameters, indicating that these selected feature parameters are effective in salt stress response.
[0055] By predicting the D-values for low-salt (LS) and high-salt (HS) treatments separately and calculating their mean D-values, the overall salt tolerance level of different chrysanthemum varieties under salt stress can be effectively characterized. Taking 16 resource classifications of extremely salt-tolerant varieties as an example (Table 4), the predicted D-values are highly consistent with the actual observed values, with the deviation generally controlled within ±0.05. Only a few varieties have slightly higher deviations, but the overall deviations remain within an acceptable range, and the predicted classification results are consistent with the actual classifications. This indicates that the method of predicting and averaging D-values under multiple environmental conditions can improve the stability and reliability of salt tolerance evaluation results while eliminating random errors under a single salt concentration.
[0056] Table 4. Extremely salt-tolerant resources selected by the prediction model built based on the selected feature parameters.
[0057] Step 6: Machine learning model predicts aboveground biomass Similar to predicting the D value, the filtered feature parameters were divided into five datasets: all feature parameters (ALL), morphological feature parameters (mor-i), histogram texture features, gray-level co-occurrence matrix (GLCM), and gray-level-gradient co-occurrence matrix (GGCM). A random forest model was used to predict aboveground biomass under control (CK), low-salt treatment (LS), and high-salt treatment (HS) (Table 5), with the remaining steps the same as in 5.1.
[0058] The results show that the ALL dataset performs best in fresh weight prediction, achieving high accuracy under all three processing conditions. Figure 6 ),in R² The root mean square errors were 0.825 (CK), 0.846 (LS), and 0.875 (HS), respectively. RMSE ) and mean absolute error ( MAE The dry weight prediction results were also at a low level, indicating that the model can stably capture the trend of fresh weight changes under salt stress. The dry weight prediction results were generally lower than the fresh weight, but still maintained high explanatory power. In the ALL dataset, the three treatments... R² The results were 0.823 (CK), 0.766 (LS), and 0.738 (HS), respectively, indicating that the random forest model has moderate to high reliability in dry weight prediction. Further comparison of different feature datasets revealed that the prediction accuracy of the GGCM and GLCM feature sets was close to that of the ALL dataset under certain conditions, indicating that texture features and higher-order co-occurrence matrix parameters have strong explanatory power in the representation of aboveground biomass. Overall, the random forest model can effectively integrate multidimensional image feature data to achieve non-destructive and accurate prediction of aboveground biomass, providing technical support for high-throughput monitoring of plant growth dynamics under salt stress.
[0059] Table 5. Effects of Random Forest Model Prediction on Aboveground Biomass under Normal Conditions
[0060] Note: The best result is shown in bold.
Claims
1. A high-throughput method for identifying salt tolerance of chrysanthemum, characterized in that, The method comprises the following steps: (1) selecting multiple chrysanthemum germplasm resources, setting different concentrations of sodium chloride solution for salt stress treatment, and calculating a comprehensive salt tolerance D value by principal component analysis method for multiple salt tolerance related traits; (2) after the salt stress phenotype appears, collecting the RGB images of four side angles and one overhead angle of each plant, and performing image color calibration and standardized naming; (3) segmenting the images in the excess green and HSV color space to obtain plant binary images, and extracting two types of digital feature parameters of morphology and texture, and analyzing the reliability of the feature parameters by using a linear regression model; (4) screening representative features from the digital feature parameters which are significantly responsive to salt stress and genetically stable by using independent sample t-test, variability analysis, feature importance analysis and generalized heritability analysis; (5) dividing the screened representative features into multiple feature subsets, using the random forest algorithm, taking the representative features as independent variables and the salt tolerance D value as dependent variable, and training a salt tolerance prediction model; (6) based on the salt tolerance prediction model, the importance of each feature is evaluated, the optimal feature combination with the smallest prediction error is screened out, and the salt tolerance grade of chrysanthemum germplasm resources is automatically divided.
2. The method according to claim 1, wherein, In step (1), the salt stress treatment specifically includes: setting a control group, a 100mM NaCl low concentration treatment group and a 200mM NaCl high concentration treatment group; the calculation of the salt tolerance D value includes: standardizing the original index matrix of 22 salt tolerance related traits, obtaining the eigenvalue and eigenvector of the correlation coefficient matrix, calculating the scores of each principal component and performing forward normalization processing, and finally calculating the weighted average value to obtain the D value by taking the contribution rate of each principal component as the weight.
3. The method according to claim 1, wherein, In step (3), the morphological feature parameters include total projection area, convex hull area, perimeter, green projection area, plant height, plant width, green projection area ratio, height-width ratio, side view total projection area and edge area ratio, and perimeter area ratio; the texture feature parameters include histogram texture features, gray level co-occurrence matrix features and gray level-gradient co-occurrence matrix features based on G, H and V three color components.
4. The method according to claim 1, wherein, In step (3), the linear regression model is used to analyze the coefficient of determination R² and its significance.
5. The method according to claim 1, wherein, In step (4), the specific standards for screening are as follows: first, the independent sample t-test is used to remove the feature parameters with no significant difference P > 0.05 between the treatment group and the control group; then, the variability analysis is performed to remove the feature parameters with a coefficient of variation CV average value less than 2%; then, the feature importance analysis is performed, and the multilayer perceptron model combined with the SHAP method is used to remove the feature parameters with a normalized importance average value less than 0.3; finally, the generalized heritability analysis is performed to remove the feature parameters with a heritability H² < 0.6, and the finally retained feature parameters are taken as the representative features which are significantly responsive to salt stress and genetically stable.
6. The method according to claim 1, wherein, In step (5), the multiple feature subsets include: an all-feature subset, a morphology feature subset, a histogram texture feature subset, a gray level co-occurrence matrix feature subset and a gray level-gradient co-occurrence matrix feature subset.
7. The method according to claim 1, wherein, In step (6), the optimal feature combination is determined by ten-fold cross-validation, wherein the optimal feature combination contains 7 feature parameters for low-concentration salt stress treatment, and contains 17 feature parameters for high-concentration salt stress treatment.
8. The method according to claim 1, wherein, The specific method of automatic division of salt tolerance grades is as follows: the mean value of the predicted D value of each resource under different salt treatments is calculated, and according to the standard deviation distance between the mean value and the overall mean value, it is divided into five grades of extremely salt-tolerant I, relatively salt-tolerant II, salt-tolerant III, non-salt-tolerant IV and extremely non-salt-tolerant V.