Chrysanthemum flooding tolerance efficient identification grading and biomass estimation method

Through RGB image processing and machine learning algorithms, a prediction model of chrysanthemum water resistance and biomass is established, which solves the problems of time-consuming and labor-intensive water resistance identification and biomass measurement in the existing technology, and achieves rapid and accurate water resistance identification and biomass prediction.

CN120070880APending Publication Date: 2025-05-30NANJING AGRICULTURAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411620428.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The existing chrysanthemum waterproof identification and evaluation techniques and biomass measurement methods have problems such as time-consuming and labor-intensive, low flux, large subjective deviations and the need for destructive sampling, making it difficult to achieve rapid and effective waterproof identification and lossless biomass prediction.

Method used

RGB image processing technology is adopted to collect RGB image data of chrysanthemum plants through top-view and side-view cameras, and the morphological and texture feature parameters of the plant are extracted using image segmentation algorithms, and a prediction model of chrysanthemum water resistance and biomass is established in combination with machine learning algorithms (such as random forests and gradient-enhanced trees).

Benefits of technology

Efficient identification and grading of chrysanthemum water resistance and biomass estimation have been achieved, the influence of human subjective factors is avoided, manpower and material investment is reduced, and excellent waterlogging resistance can be screened out quickly and accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070880A_ABST
    Figure CN120070880A_ABST
Patent Text Reader

Abstract

The invention discloses a chrysanthemum waterlogging tolerance efficient identification grading and biomass estimation method, which comprises the following steps: (1) selecting a plurality of representative chrysanthemum and related species plant resources from different sources for waterlogging treatment to obtain a waterlogging tolerance membership function value of each resource; (2) acquiring RGB images of the plant in a moderate waterflooding period and a heavy waterflooding period, acquiring image data of the plant at four side view angles and one overlook angle by using a camera, and performing image color calibration and batch naming; (3) processing the image by using an image segmentation algorithm to obtain a plant part binary image, calculating plant morphological characteristics of plant height, width and projection area, and extracting plant texture characteristics based on image gray and color information; (4) screening the obtained characteristic parameters according to phenotype significance t test and heritability to obtain representative image characteristic parameters responding to waterlogging stress; and (5) dividing the obtained image feature parameters into five types of data sets, establishing an optimal prediction model of waterlogging tolerance membership function values and biomass by adopting two machine learning algorithms of a random forest and a gradient boosting tree, and automatically dividing the waterlogging tolerance grade of the chrysanthemum. The method can provide reference for identification and evaluation of plant waterlogging tolerance or other abiotic stress resistance and biomass prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of identification, evaluation and breeding of germplasm resources of flower crops, and particularly relates to a method for efficient identification and grading of waterlogging tolerance and biomass estimation of chrysanthemums. Background Art

[0002] Chrysanthemum morifolium Ramat. occupies an important position in the world flower industry. However, chrysanthemums are shallow-rooted crops and are intolerant to waterlogging. Waterlogging stress seriously hinders the healthy development of the chrysanthemum industry. Screening waterlogging-tolerant germplasms and cultivating waterlogging-tolerant varieties are effective ways to solve this problem. At present, the identification of waterlogging tolerance of chrysanthemums mainly involves scoring based on the damage degree of the above-ground part and counting the yellow leaf rate, and measuring the stress indices of waterlogging-related traits such as plant height, fresh weight of the above-ground and underground parts, and dry weight of the above-ground and underground parts, and using the membership function value method for comprehensive evaluation and waterlogging tolerance level classification. Among them, biomass comprehensively reflects the performance of plants in many aspects such as architecture, light absorption, and carbon assimilation, and is an important indicator for measuring plant growth status and environmental adaptability. However, traditional phenotypic acquisition methods are time-consuming, laborious, low-throughput, subject to subjective biases, and require destructive sampling. Therefore, there is an urgent need for a rapid and effective method for identifying and evaluating waterlogging tolerance of chrysanthemums and non-destructive biomass prediction to screen excellent waterlogging-tolerant germplasms with high throughput.

[0003] With the establishment of non-destructive plant phenotypic acquisition methods and large-scale automated high-throughput phenotypic acquisition platforms, phenomics technology has been applied to plant basic research and crop breeding. Common phenotypic acquisition techniques include RGB imaging, thermal infrared imaging, chlorophyll fluorescence, and multi-spectral / hyperspectral imaging, etc. Among them, RGB imaging has received extensive attention in plant phenotypic research due to its advantages such as low cost, simple operation, and convenient maintenance. In chrysanthemums, RGB image processing technology is mainly applied in the fields of variety identification and classification, quality evaluation, and flowering period monitoring, but its application in the identification and evaluation of abiotic stresses such as waterlogging and biomass prediction is still blank. Summary of the Invention

[0004] Object of the Invention: The object of the present invention is to provide a method for efficient identification and grading of waterlogging tolerance and biomass estimation of chrysanthemums, to solve the deficiencies in the existing chrysanthemum waterlogging tolerance identification and evaluation techniques and biomass measurement methods, and to predict the above-ground biomass.

[0005] Technical Solution: The method for efficient identification and grading of waterlogging tolerance and biomass estimation of chrysanthemums according to the present invention includes the following steps:

[0006] (1) Select multiple representative plant resources of Chrysanthemum and its related genera from different sources and perform flooding treatment to obtain the membership function values of waterlogging tolerance for each resource;

[0007] (2) Collect RGB images of the plants during the moderate and severe flooding periods; use an overhead camera to collect the top image data of the plants, and a side-view camera to collect the side-view image data at four angles of 0°, 90°, 180°, and 270° of the plants, that is, a total of 5 RGB images are collected for each plant; and perform image color calibration and batch naming;

[0008] (3) Use an image segmentation algorithm to process the images to obtain the binary map of the plant part; calculate 10 morphological feature parameters such as plant height, width, and projected area of the plant; a total of 93 texture feature parameters are collected according to the gray level and color information of the plant part in the image;

[0009] (4) According to the phenotypic significance t-test and heritability, screen the obtained 103 feature parameters to obtain the representative image feature parameters that respond to waterlogging stress;

[0010] (5) Divide the obtained 96 image feature parameters into 5 categories of data sets; train two machine learning models of random forest and gradient boosting tree respectively, and select the optimal prediction model;

[0011] (6) Calculate the importance of each feature parameter for the constructed optimal MFVW prediction model, sort the feature parameters, and select the number of feature parameters with the smallest error.

[0012] Furthermore, in step (1), the flooding treatment is as follows: The potting simulation flooding method is used to identify the waterlogging tolerance at the seedling stage at the 10-12 leaf stage. Set the treatment group and the control group, and use the membership function method to conduct a traditional comprehensive evaluation of the waterlogging tolerance of chrysanthemums at the two periods of moderate flooding T1 and severe flooding T2.

[0013] Furthermore, the two flooding periods are the 23rd and 39th days of the flooding treatment respectively; the phenotypes collected during flooding include score value, dead leaf rate, plant height, aboveground fresh weight, and aboveground dry weight; the calculation formula of the membership function value is: If the index data is positively correlated with waterlogging tolerance, then X i =(X - X min ) / (X max - X min ); If the index data is negatively correlated with waterlogging tolerance, then X i = 1 - (X - X min ) / (X max - X min ), where X is the measured value of a certain index of a certain resource, and X min and X max are the minimum and maximum values in the population; finally, calculate the average membership function value MFVW of all the indexes of each resource as the standard for comprehensively evaluating the waterlogging tolerance of chrysanthemums. The larger this value is, the stronger its resistance.

[0014] Further, in step (2), the RGB image color calibration is performed using Adobe Lightroom Classic and the software ColorChecker Camera Calibration (X-RITE, USA) provided by X-RITE; according to the QR code in the RGB image, each picture is named using python in the format of "number_CK / T_F1 / C1_1", where the QR code is generated by a QR code generator.

[0015] Further, in step (3), the image segmentation is performed in the excess green E×G and HSV color spaces, and the intersection of the E×G and HSV segmentations is the final binary image of the plants.

[0016] Further, in step (3), the 10 morphological feature parameters include: total projected area TPA, convex hull area CHA, perimeter P, green projected area GPA, plant height H, plant width W, green projected area ratio GPAR, height-width ratio HWR, ratio of total projected area to edge area in side view TBR, and perimeter-area ratio PAR; the texture feature parameters mainly have three types, namely histogram texture features, gray-level co-occurrence matrix, and gray-level gradient co-occurrence matrix.

[0017] Further, the texture features are extracted under the three components of G, H, and V respectively.

[0018] Further, in step (4), the feature parameter screening is as follows: First, the independent samples t-test is used to select the image feature parameters with significant differences between the treatment group and the control group, and the confidence interval is 95%; then the broad-sense heritability (H 2 ) of the feature parameters is calculated, and the feature parameters with higher heritability (H 2 ≥0.2) are selected for the next analysis.

[0019] Further, in step (5), the 5 types of data sets include: all image feature parameters ALL, morphological feature parameters mor-i, histogram texture features texture, gray-level co-occurrence matrix GLCM, and gray-level gradient co-occurrence matrix GGCM.

[0020] Further, in step (5), the feature parameters obtained from the RGB images and the manually collected data are divided into training sets and test sets. Taking the feature parameters obtained from the RGB images as independent variables and the manually collected phenotypic data as dependent variables, the random forest and gradient boosting tree models are trained respectively.

[0021] The present invention also provides an application of the above method in the efficient identification and grading of chrysanthemum waterlogging tolerance, which is characterized in that after step (6) of claim 1, the following steps are further included:

[0022] (7) Using the 13 selected characteristic parameters GPA, GPAR, H_SD, H_ST, H_ME, H_ENT_ME, H_COR_ME, V_ENT_ME, V_COR_SD, V_COR_ME, H_GGCM6, H_GGCM8, and H_GGCM14 to construct a random forest prediction model again to predict the waterlogging tolerance membership function value MFVW of each resource;

[0023] (8) Classify the chrysanthemum resources according to the predicted MFVW: MFVW ≥ mean + 1.64 × standard deviation is extremely waterlogging-tolerant (Ⅰ), mean + 1.64 × standard deviation > MFVW ≥ mean + standard deviation is relatively waterlogging-tolerant (Ⅱ), mean + standard deviation > MFVW ≥ mean - standard deviation is waterlogging-tolerant (Ⅲ), mean - standard deviation > MFVW ≥ mean - 1.64 × standard deviation is not waterlogging-tolerant (Ⅳ), and MFVW < mean - 1.64 × standard deviation is extremely not waterlogging-tolerant (Ⅴ).

[0024] Beneficial effects: Compared with the prior art, the present invention has the following remarkable advantages: It is applicable to all chrysanthemum and its related plant species and can be used for reference by other potted crops; The characteristic parameters extracted from RGB images can replace manual statistical indicators to quantify the waterlogging tolerance of chrysanthemums and divide the waterlogging tolerance levels, avoiding the influence of subjective factors on the evaluation of chrysanthemum waterlogging tolerance, which has important theoretical and practical significance for chrysanthemum waterlogging tolerance breeding; A prediction model of the above-ground biomass of chrysanthemums is established using machine learning algorithms, which can achieve non-destructive dynamic measurement of plants and greatly reduce the input of human and material resources. Brief Description of the Drawings

[0025] Figure 1 is the experimental process of the method for identifying and grading the waterlogging tolerance and estimating the biomass of chrysanthemums based on RGB image processing of the present invention; Figure 1 in (a), 180 resources are subjected to flooding treatment, Figure 1 in (b), lateral and top-view RGB images are collected at two flooding periods of moderate stress (T1) and severe stress (T2), Figure 1 in (c), preprocessing such as color correction and renaming is performed on the collected RGB images, Figure 1 in (d), the binary image of the plant and the three components G, H, and V are extracted through an image segmentation algorithm, Figure 1 in (e), the finally extracted 10 morphological characteristic parameters and 93 texture characteristic parameters.

[0026] Figure 2 is the reliability analysis of the characteristic parameters of the present invention; Figure 2 in (a) and (b) are the correlation analyses of the plant height and fresh weight collected manually and the plant height and total projected area extracted from RGB images, respectively, Figure 2In (c), hierarchical clustering was performed on 180 resources based on morphological characteristic parameters.

[0027] Figure 3 Filtration and determination of waterlogging tolerance-related characteristic parameters of the present invention; Figure 3 In (a), t-tests were performed on the treatment group and the control group in two stages, ***P < 0.001, **P < 0.01, *P < 0.05, ns indicates no significant difference. Figure 3 In (b), the coefficient of variation (CV) of all characteristic parameters Figure 3 In (c), the broad-sense heritability (H 2 )

[0028] Figure 4 Selection of characteristic parameters for early response to waterlogging stress of the present invention; (a) is the principal component analysis of five categories of characteristic parameters during moderate and severe flooding periods, and (b) is the phenotypic distribution of H_SD, H_ENT_SD, and H_GGCM101 of normal and treated plants of 180 resources during two flooding periods.

[0029] Figure 5 The optimal models for predicting MFVW and above-ground biomass using the random forest algorithm and the importance ranking of image characteristic parameters in the present invention; among them, (a) are the optimal models for predicting fresh weight using the morphological characteristic parameter dataset, dry weight using all characteristic parameter datasets, and MFVW using the GLCM dataset respectively, (b) is the screening of important characteristic parameters using cross-validation, and the number of optimal characteristic parameters is 13, and (c) is the prediction of MFVW using the 13 screened characteristic parameters using the random forest and gradient boosting tree algorithms respectively. Detailed implementation manners

[0030] The technical solution of the present invention will be further described below with reference to the accompanying drawings.

[0031] The embodiment of the present invention provides a method for efficient identification and grading of chrysanthemum waterlogging tolerance and biomass estimation, including the following steps:

[0032] I. Traditional waterlogging tolerance evaluation and analysis

[0033] 1. Materials and field experiment design

[0034] A total of 180 plant resources of Chrysanthemum and its related genera from different sources without direct genetic relationship were selected as experimental materials (Table 1), including 116 wild resources and 64 cultivated varieties. All materials are preserved in the "China Chrysanthemum Germplasm Resource Conservation Center" of Nanjing Agricultural University, and those skilled in the art can obtain the above germplasm from the "China Chrysanthemum Germplasm Resource Conservation Center". Cuttings with consistent growth and no pests and diseases were planted in the greenhouse of Nanjing Hushu Experimental Base (119.12°N, 31.80°E) in September 2022. The experiment adopted a completely randomized block design, with a treatment group (WS) and a control group (WW). The treatment group had 3 biological replicates, and the control group had 2 biological replicates. Each replicate had 5 plants. After about 8 - 10 days of slow seedling growth and when the plants reached the 10 - 12 leaf stage, the plants were placed in plastic turnover boxes (65 cm long × 43 cm wide × 16 cm high). Tap water was injected into the turnover boxes of the treatment group to make the water surface 3 cm higher than the soil surface. During the treatment period, water was replenished regularly to ensure the water level, and the control group plants were managed normally. Figure 1 a).

[0035] 2 Materials and field experiment design

[0036] After 23 days (T1, moderate stress) and 39 days (T2, severe stress) of waterlogging treatment, the plants were scored according to the leaf color and leaf morphology, and the dead leaf rate was calculated. The score value was the sum of the leaf color score and the leaf morphology score after waterlogging treatment. The dead leaf rate was the proportion of the number of damaged leaves of the treatment group plants to the total number of intact leaves of the plant. Since the control group plants grew healthily, the above indicators were only statistically analyzed for the treatment group plants.

[0037] After each scoring, the aboveground biomass information of the treated materials was collected manually from the soil surface. The plant height (cm) of the materials was measured with a ruler; the aboveground fresh weight (g) of the materials was obtained by weighing. The fresh aboveground tissues were blanched at 105°C for 20 minutes and dried to a constant weight (≥72 hours) at 80°C to measure the aboveground dry weight (g) to obtain the dry weight data. Except for the score value and the dead leaf rate, the waterlogging tolerance coefficient (WRC, %) of each traditional trait was calculated according to the formula: treatment group / control group × 100.

[0038] The membership function method was used to calculate the membership function values of the above indicators, and the membership function value of waterlogging tolerance (MFVW) was used as the standard for comprehensively evaluating the waterlogging tolerance of chrysanthemums. Among them, the larger the MFVW value, the better the waterlogging tolerance of the resource. The scoring standard and MFVW calculation were mainly based on the waterlogging tolerance evaluation system established by Su et al. (2016).

[0039] 3 Traditional waterlogging tolerance evaluation

[0040] According to the average values of MFVW in two periods (Table 1), and referring to the classification method of Yan et al. (2020), all materials were divided into 5 waterlogging tolerance levels: among them, there were 10 materials with extreme waterlogging tolerance (Ⅰ), 19 materials with relatively waterlogging tolerance (Ⅱ), 124 materials with waterlogging tolerance (Ⅲ), 19 materials with waterlogging intolerance (Ⅳ), and 8 materials with extremely waterlogging intolerance (Ⅴ).

[0041] Table 1 Information on 180 tested materials and their waterlogging tolerance

[0042]

[0043]

[0044]

[0045]

[0046]

[0047] II RGB Image Acquisition and Feature Parameter Extraction

[0048] 1 RGB Image Collection

[0049] After each scoring, all potted plants were imaged in a 80 cm × 80 cm × 80 cm photography studio with a fixed light source indoors ( Figure 1 b). A Canon EOS 90D camera was used for shooting. The background was a white paper of A2 size, with black circles with a radius of 1.5 cm marked at each of the four corners, and a black square of 1 cm × 1 cm size as a scale. The X-RITE Mini 24-color card for color calibration and the QR code for picture naming were fixed. The QR code was generated by an online QR code generator (http: / / www.qinms.com / webapp / qrcode / MultiQrcode.aspx). According to the size of the plant, the appropriate shooting distance was selected and fixed to meet the field of view size requirements. In this experiment, the imaging distances of the top-view camera and the side-view camera were 60 cm and 100 cm respectively, and the resolution was 6960 × 4640. The top-view camera collected the phenotypic data of the top of the plant, and the side-view camera collected the phenotypic data of the plant at four angles of 0°, 90°, 180°, and 270°, that is, a total of 5 RGB images were collected for each plant.

[0050] 2 Extraction of Digital Feature Parameters Based on RGB Images

[0051] Image color calibration was performed using ColorChecker Camera Calibration (X-RITE, USA); due to the excessive number of collected images, we used the pyzbar package in Python to decode QR codes and achieve image renaming. Figure 1 c). Image segmentation and feature parameter extraction in the excessive green (E×G) and HSV color spaces were completed using MATLAB R2022b (MathWorks Inc., USA). Figure 1 d). Through image processing and analysis, a total of 103 feature parameters based on RGB images were extracted, including 10 morphological feature parameters and 93 texture feature parameters (Table 2). The texture features were calculated separately under three components: green (G), hue (H), and value (I), denoted as G texture, H texture, and V texture respectively (Table 2; Figure 1 e).

[0052] Table 2 Detailed definitions of traditional traits and image feature parameters

[0053]

[0054]

[0055] 3 Reliability analysis of feature parameters

[0056] To evaluate the reliability of the image-based feature parameter data, correlation analysis was performed between the manually measured plant height and aboveground fresh weight and the corresponding feature parameters. Figure 2 a and b). At the T2 stage of the stress treatment, the plant height measured manually was highly positively correlated with the plant height obtained from RGB images (WS, R 2 = 0.95, P < 0.001; WW, R 2 = 0.93, P < 0.001)). The fresh weight was also linearly related to TPA (WS, R 2 = 0.69, P < 0.001; WW, R 2 = 0.49, P < 0.001), but weaker than the plant height. The measurement results based on RGB images had a strong positive correlation with the manual measurement results, indicating that the non-destructive feature parameters obtained from RGB images could better reflect the true values of traditional indicators.

[0057] In addition, according to the calculation of MFVW, morphological feature parameters (TPA, CHA, P, W, GAP, PAR, and GPAR) were selected, and their waterlogging tolerance coefficients (treatment group / control group × 100) were calculated. Hierarchical clustering was performed using the 'ggtree' package in R software, and 180 materials could be divided into 5 subgroups. Figure 3c), which is relatively consistent with the five waterlogging tolerance levels of traditional waterlogging tolerance evaluation, indicating that the morphological characteristic parameters can quantify the differences in plant waterlogging response.

[0058] 4 Screening of effective characteristic parameters

[0059] After obtaining 103 characteristic parameters, the characteristic parameters responding to flooding stress were screened according to the following steps: (1) Independent-samples t-tests were performed using IBMSPSS version 27 (IBM, Armonk, USA) to select characteristic parameters with significant differences between the treatment group and the control group, with a confidence interval of 95%; (2) The generalized heritability (H 2 ) of the characteristic parameters was calculated using the lmer function in the 'lme4' package of R software, and characteristic parameters with higher heritability (H 2 ≥0.2) were selected for further analysis.

[0060] The results showed ( Figure 2 ), whether viewed from the side or top, most characteristic parameters had significant differences. Six non-significant characteristic parameters, namely HWR, G_UF, G_ET, G_ENT_SD, G_GGCM3, and G_GGCM11, were excluded based on the P-value. The characteristic parameters of different materials were significantly different, with the coefficient of variation ranging from 0.16% to 513.63%, and the average coefficient of variation was 34.04%; the H 2 of these characteristic parameters ranged from 0.050 to 0.997, and the H 2 of most characteristic parameters was greater than 0.5. In summary, the waterlogging tolerance of Asteraceae plants is mainly controlled by genetic factors, and the reliability of their characteristic parameters is relatively high. We believe that the phenotypic variation of characteristic parameters with H 2 lower than 0.2 is mainly caused by environmental factors, so H_COR_SD was excluded in further analysis.

[0061] 5 Analysis of characteristic parameters responding to waterlogging stress

[0062] Principal component analysis was performed on the obtained effective characteristic parameters to reflect their waterlogging stress response at the population level. Principal component analysis was performed using the prcomp function in the R package factoextra, and the results of principal component analysis were visualized using the fviz_pca_ind function, with a confidence interval of 90%. The characteristic parameters were divided into 5 categories, namely all characteristic parameters, morphological characteristic parameters, histogram texture features, gray-level co-occurrence matrix (GLCM), and gray-level-gradient co-occurrence matrix (GGCM). Among these 5 categories, only Dim1 explained 59.9 - 74.3% and 64.8 - 81.6% of the phenotypic variation at T1 and T2, respectively ( Figure 4a). Compared with the other four types, GLCM can better distinguish between the control group and the treated group of plants at the early stress stage of T1, while GGCM shows better discrimination ability at the T2 stage. These results indicate that the texture feature parameters may also be potential indicators reflecting the waterlogging response of Compositae plants. Three texture features, namely H_SD, H_ENT_ME, and H_GGCM8, were selected from the histogram texture, GCLM, and GGCM respectively for further analysis. From Figure 4 As can be seen from b, with the prolongation of waterlogging stress time, the discrimination ability of the control group and the treated group for individuals is significantly enhanced, which can be used as an index for identifying waterlogging-sensitive and waterlogging-tolerant chrysanthemum resources.

[0063] 6MFVW Prediction Optimal Model Confirmation

[0064] 6.1 Dataset Classification and Model Establishment

[0065] The selected characteristic parameters were divided into 5 types of datasets, namely all characteristic parameters (ALL), morphological characteristic parameters (mor-i), histogram texture features (texture), gray-level co-occurrence matrix (GLCM), and gray-level-gradient co-occurrence matrix (GGCM). The random forest model (RF) and the gradient boosting tree model (GBT) were used to predict MFVW during the severe treatment period. Since MFVW is calculated based on the stress index of traditional traits, the stress index (WS / WW×100) of the characteristic parameters was used for prediction. 70% of the characteristic parameters and MFVW obtained from RGB images were used as the training set, and 30% as the test set. Taking the characteristic parameters obtained from RGB images as the independent variable and MFVW as the dependent variable, two machine learning models, random forest and gradient boosting tree, were trained respectively using the five-fold cross-validation method. The random forest and gradient boosting tree models were implemented in the 'randomForest' and 'xgboost' R packages respectively, and the five-fold cross-validation was implemented in the 'Caret' R package. The root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R 2 ) were used to evaluate the accuracy of the model prediction. The lower the values of RMSE and MAE, and the higher the R 2 , the better the model prediction effect.

[0066] Regarding the prediction results of MFVW in the T2 period (Table 3), except for the GGCM dataset under the gradient boosting tree model, the R 2 of each model is above 0.8, showing a good prediction effect. For the random forest model, the prediction accuracy of the GCLM dataset is the highest (R 2 = 0.858), while the prediction accuracy of the histogram texture dataset under the gradient boosting tree model is the best (R 2= 0.841). The average values of the evaluation metrics for different datasets under the same model were calculated respectively, and it was found that the prediction accuracy of the random forest model (R 2 = 0.834) for MFVW was better than that of the gradient boosting tree model (R 2 = 0.810).

[0067] Table 3 Comparison of the prediction effects of MFVW based on random forest and gradient boosting tree models

[0068]

[0069] Note: The optimal results are shown in bold.

[0070] 6.2 Feature parameter selection and waterlogging tolerance classification

[0071] According to the optimal prediction model - the random forest model, an MFVW prediction model was constructed using all feature parameter datasets. By performing ten-fold cross-validation repeated 5 times, the feature parameters were selected or discarded based on the cross-validation curve. When n = 13, the cross-validation error was the smallest ( Figure 5 a), indicating that an ideal regression result could be obtained by retaining the first 13 feature parameters. According to the importance of the variables, the first 13 feature parameters GPA, GPAR, H_SD, H_ST, H_ME, H_ENT_ME, H_COR_ME, V_ENT_ME, V_COR_SD, V_COR_ME, H_GGCM6, H_GGCM8, and H_GGCM14 were selected to predict MFVW again using machine learning models. Under the random forest and gradient boosting tree models, the prediction accuracies were 0.825 and 0.803 respectively, with quite high prediction accuracies ( Figure 5 b, c). The results show that the prediction accuracy using a small number of feature parameters is almost the same as that using all feature parameters, indicating that these selected feature parameters are effective in waterlogging response.

[0072] The chrysanthemum resources were classified according to the predicted MFVW: MFVW ≥ mean + 1.64 × standard deviation is extremely waterlogging tolerant (Ⅰ), mean + 1.64 × standard deviation > MFVW ≥ mean + standard deviation is relatively waterlogging tolerant (Ⅱ), mean + standard deviation > MFVW ≥ mean - standard deviation is waterlogging tolerant (Ⅲ), mean - standard deviation > MFVW ≥ mean - 1.64 × standard deviation is not waterlogging tolerant (Ⅳ), and MFVW < mean - 1.64 × standard deviation is extremely not waterlogging tolerant (Ⅴ). Twelve extremely waterlogging tolerant resources are shown in Table 4.

[0073] Table 4 Extremely waterlogging tolerant resources screened by the prediction model constructed based on the selected 13 feature parameters

[0074]

[0075] Among them, the waterlogging tolerance grades of 8 resources, such as Aster tataricus, Chrysanthemum nankingense × Tanacetum vulgare, Artemisia lancea, Dendranthema zawadskii var. latilobum × Ajania przewalskii, Artemisia leucophylla, Artemisia vulgaris, Dendranthema indicum × Crossostephium chinense, and Artemisia viridis, all belong to extreme waterlogging tolerance and can be used for subsequent chrysanthemum waterlogging tolerance breeding and molecular mechanism research.

[0076] III. Machine learning model for predicting aboveground biomass

[0077] Similar to the prediction of MFVW, the screened characteristic parameters were divided into 5 categories of data sets, namely all characteristic parameters (ALL), morphological characteristic parameters (mor-i), histogram texture features (texture), gray-level co-occurrence matrix (GLCM), and gray-level gradient co-occurrence matrix (GGCM). The random forest model and gradient boosting tree model were used to predict the aboveground biomass obtained under different treatment conditions (WS and WW) and treatment periods (T1 and T2). The remaining steps were the same as in 6.1.

[0078] For aboveground biomass, the prediction effect of the two models at T2 was better than that at T1, and the prediction effect of the treatment group was better than that of the control group (Table 4). Generally speaking, the prediction accuracy of fresh weight was the highest using the morphological characteristic parameter data set, and the prediction accuracy of dry weight was the highest using all characteristic parameter data sets. By calculating the average value of each evaluation index of the same model at different periods and treatment conditions, it was found that the prediction accuracy of the random forest model for MFVW, fresh weight, and dry weight was better than that of the gradient boosting tree model. Therefore, the random forest model was the best model for prediction.

[0079] Table 5 Effects of predicting aboveground biomass under treatment and normal conditions based on random forest and gradient boosting tree models

[0080]

[0081] Note: The optimal results are shown in bold.

Claims

1. A method for efficiently identifying and grading the waterlogging tolerance of chrysanthemum and estimating its biomass, characterized in that: The following steps are involved: (1) Select multiple representative plant resources of the genus Chrysanthemum and its closely related species from different sources, perform flooding treatment, and obtain the membership function value of waterlogging resistance of each resource; (2) During the moderate and severe flooding periods, the side-view camera was used to collect image data of the plants at four angles: 0°, 90°, 180°, and 270°, and the top-view camera was used to collect image data of the top of the plants. That is, a total of 5 RGB images were collected for each plant, and image color calibration and batch naming were performed; (3) The image was processed using an image segmentation algorithm to obtain a binary image of the plant; 10 morphological characteristic parameters such as plant height, width, and projection area were extracted; the texture characteristics of the plant were extracted based on the grayscale and color information of the plant in the image, and a total of 93 texture characteristic parameters were obtained; (4) Based on the phenotypic significance t-test and heritability, the 103 characteristic parameters obtained were screened to obtain representative image characteristic parameters that responded to waterlogging stress; (5) The 96 image feature parameters obtained were divided into five data sets; two machine learning models, random forest and gradient boosting tree, were trained respectively, and the optimal prediction model was selected; (6) The importance of each feature parameter of the constructed optimal MFVW prediction model is calculated, and the feature parameters are sorted to select the number of feature parameters with the smallest error.

2. The method for efficiently identifying and grading waterlogging tolerance and estimating biomass of chrysanthemum according to claim 1, characterized in that: In step (1), the flooding treatment is as follows: a potted simulated flooding method is used to identify the waterlogging resistance of seedlings at the 10-12 leaf age, a treatment group and a control group are set up, and a traditional comprehensive evaluation of the waterlogging resistance of chrysanthemum is performed by the membership function method under moderate flooding T1 and severe flooding T2 conditions.

3. A method for efficiently identifying and grading waterlogging tolerance and estimating biomass of chrysanthemum according to claim 2, characterized in that: The two flooding periods were the 23rd and 39th days of flooding treatment respectively; the phenotypes collected during flooding included score value, dead leaf rate, plant height, aboveground fresh weight, and aboveground dry weight; the calculation formula of the membership function value was: if the index data was positively correlated with waterlogging tolerance, then X i =(XX min ) / (X max -X min ); if the index data is negatively correlated with waterlogging resistance, then X i =1-(XX min ) / (X max -X min ) where X is the measured value of a certain indicator of a certain resource, X min and X max are the minimum and maximum values ​​in the population; finally, the average value of the membership function values ​​of all indicators of each resource is calculated as the standard for comprehensive evaluation of chrysanthemum waterlogging resistance. The larger the value, the stronger its resistance.

4. A method for efficiently identifying and grading waterlogging tolerance and estimating biomass of chrysanthemum according to claim 1, characterized in that: In step (2), Python is used to name each image according to the QR code in the RGB image, in the format of "number_CK / T_F1 / C1_1". The QR code is generated by a QR code generator.

5. A method for efficiently identifying and grading waterlogging tolerance and estimating biomass of chrysanthemum according to claim 1, characterized in that: In step (3), image segmentation is performed in the excess green E×G and HSV color space, and the intersection of the E×G and HSV segmentations is the final plant binary image.

6. A method for efficiently identifying and grading waterlogging tolerance and estimating biomass of chrysanthemum according to claim 1, characterized in that: In step (3), the 10 morphological feature parameters include: total projected area TPA, convex hull area CHA, perimeter P, green projected area GPA, plant height H, plant width W, green projected area ratio GPAR, aspect ratio HWR, side view total projected area to edge area ratio TBR and perimeter area ratio PAR; texture features are extracted from the three components of G, H and V, and there are three main types, namely histogram texture features, grayscale co-occurrence matrix and grayscale-gradient co-occurrence matrix.

7. A method for efficiently identifying and grading waterlogging tolerance and estimating biomass of chrysanthemum according to claim 1, characterized in that: In step (4), the feature parameter screening is as follows: first, an independent sample t-test is used to select image feature parameters with significant differences between the treatment group and the control group, with a confidence interval of 95%; then the broad heritability (H 2 ), select the characteristic parameter with higher heritability (H 2 ≥0.2) for further analysis.

8. The method for efficiently identifying and grading waterlogging tolerance and estimating biomass of chrysanthemum according to claim 1, characterized in that: In step (5), the five types of data sets include: all image feature parameters ALL, morphological feature parameters mor-i, histogram texture features texture, gray-level co-occurrence matrix GLCM and gray-gradient co-occurrence matrix GGCM.

9. The method for efficiently identifying and grading waterlogging tolerance and estimating biomass of chrysanthemum according to claim 1, characterized in that: In step (5), the feature parameters obtained through RGB images and the manually collected data are divided into a training set and a test set, and the feature parameters obtained based on the RGB images are used as independent variables and the manually collected phenotypic data are used as dependent variables to train the random forest and gradient boosting tree models respectively.

10. An application of the method according to claims 1 to 9 in the efficient identification and classification of waterlogging resistance of chrysanthemum, characterized in that: After step (6), the following steps are also included: (7) The random forest prediction model was constructed again using the 13 selected characteristic parameters GPA, GPAR, H_SD, H_ST, H_ME, H_ENT_ME, H_COR_ME, V_ENT_ME, V_COR_SD, V_COR_ME, H_GGCM6, H_GGCM8, and H_GGCM14 to predict the waterlogging resistance membership function value MFVW of each resource; (8) Chrysanthemum resources were classified according to the predicted MFVW: MFVW ≥ mean + 1.64 * standard deviation, extremely resistant to waterlogging (I); mean + 1.64 * standard deviation > MFVW ≥ mean + standard deviation, relatively resistant to waterlogging (II); mean + standard deviation > MFVW ≥ mean - standard deviation, resistant to waterlogging (III); mean - standard deviation > MFVW ≥ mean - 1.64 * standard deviation, intolerant to waterlogging (IV); MFVW < mean - 1.64 * standard deviation, extremely intolerant to waterlogging (V).