A method for efficiently predicting cassava single plant tuber fresh weight based on whole genome
The linear regression model constructed using whole-genome selection and machine learning algorithms solves the problems of long evaluation cycle and high cost of single-plant fresh weight in traditional cassava breeding, enabling early and accurate screening of high-yield potential materials and improving breeding efficiency and prediction accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-04
- Publication Date
- 2026-07-21
AI Technical Summary
Traditional cassava breeding involves long evaluation cycles, high costs, and low screening efficiency for single-plant fresh weight, and is greatly affected by environmental factors. Modern molecular breeding methods struggle to explain the cumulative effects of multiple genes, impacting the accuracy and stability of predictions.
A predictive model was constructed using whole-genome selection combined with machine learning algorithms. SNPs that were significantly associated with the fresh weight of individual tubers were selected from the cassava training population. The fresh weight of individual tubers was predicted at the seedling stage using a linear regression model. The predictive model was constructed by combining machine learning algorithms such as lasso regression and elastic net regression and 100 cross-validations were performed to select the linear regression model.
It achieves high-precision prediction of single-plant fresh weight of potatoes during the seedling stage, shortens the breeding cycle, reduces costs, improves screening efficiency, has low prediction error, explains about 81.6% of phenotypic variation, and has good generalization ability.
Smart Images

Figure CN122435983A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of plant molecular breeding technology. Specifically, it relates to a method for efficiently predicting the fresh tuber weight of a single cassava plant based on whole genome, and in particular, a method for early prediction of the fresh tuber weight of a single cassava plant by combining machine learning algorithms with whole genome SNP data. Background Technology
[0002] Cassava (Manihot esculenta) is a globally important root vegetable crop. Its tubers are rich in starch and are a major energy source for hundreds of millions of people in tropical and subtropical regions. The fresh weight of cassava tubers not only determines the yield per unit area but is also an important economic trait for comprehensively evaluating the production performance of cultivated materials, and is one of the core objectives of cassava genetic improvement breeding.
[0003] Traditional methods for improving the fresh weight of cassava tubers mainly rely on field breeding trials and phenotypic selection. However, the fresh weight of individual tubers can only be accurately measured at maturity, resulting in long evaluation cycles, high field trial costs, and low screening efficiency. In addition, this trait is controlled by multiple genes and is easily affected by environmental factors (such as precipitation, temperature, and soil nutrients), making traditional selection methods highly susceptible to environmental noise interference.
[0004] In modern molecular breeding techniques, marker-assisted selection (MAS) based on single molecular markers has limited effectiveness in handling complex traits and struggles to explain the cumulative effects of multiple genes. While genomic selection (GS) can estimate individual breeding values using whole-genome marker information, how to combine it with efficient algorithms to further improve the accuracy and stability of predicting the complex quantitative trait of cassava tuber fresh weight remains a pressing issue for current breeding technologies. Summary of the Invention
[0005] The purpose of this invention is to provide a method for predicting the fresh weight of cassava tubers from individual plants based on whole-genome selection, aiming to solve the technical problems of long evaluation cycle, high cost and low screening efficiency of individual tuber fresh weight in traditional cassava breeding, and to achieve early and accurate screening of high-yield potential materials.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] In a first aspect, the present invention provides a method for efficiently predicting the fresh weight of cassava tubers per plant based on whole-genome sequencing, comprising the following steps:
[0008] (1) Obtain whole genome SNP data of the cassava sample to be tested;
[0009] (2) Substitute the whole genome SNP data into the prediction model to predict the fresh weight of the cassava plant to be tested;
[0010] The prediction model is constructed based on SNP sites that are significantly correlated with the fresh weight of individual cassava plants, selected from the cassava training population. The model formula is as follows: ,in This represents the predicted fresh weight of a single potato plant for the i-th individual. This represents the population mean. Indicates the total number of SNP sites. This represents the genotype coding value of the i-th individual at the j-th SNP locus. This represents the effect size of the j-th SNP locus. The random error term is represented; the genotype coding value is 0, 1 or 2, representing the homozygous reference type, heterozygous type and homozygous variant type, respectively.
[0011] Furthermore, the SNP loci significantly associated with the fresh weight of individual cassava plants were obtained through the following screening method: genome-wide association analysis was performed on the cassava training population, and the association significance between each SNP locus and the fresh weight phenotype of individual cassava plants was calculated using a linear mixture model, and SNP loci with a p-value less than 0.05 were screened; the prediction model was a linear regression model selected after 100 repeated cross-validations, which included machine learning models such as lasso regression, elastic net regression, cross-validation elastic net regression, kernel ridge regression, partial least squares regression, support vector regression, multinomial kernel support vector regression, random forest regression, and linear regression.
[0012] Furthermore, the cassava sample to be tested in step (1) is a cassava seedling sample, and the collection site is the tender leaf tissue.
[0013] Furthermore, the method predicts the fresh weight of individual cassava plants during the cassava seedling stage.
[0014] Furthermore, the phenotypic data of the fresh weight of individual tubers in the cassava training population were determined in the following manner: during the normal maturity period or the preset harvest period, cassava plants with uniform growth in the field were harvested one by one, all tubers of each plant were dug up completely, the above-ground parts were removed and the soil and impurities attached to the surface of the tubers were cleaned, all tubers of the same plant were collected, and their fresh weight was weighed using an electronic balance or platform scale, and the obtained weight was recorded as the phenotypic data of the fresh weight of individual tubers of that plant.
[0015] Furthermore, the method is applied to the early screening of cassava hybrid breeding populations.
[0016] On the other hand, the present invention also provides a cassava breeding and screening method, which uses the method to predict the fresh weight of a single cassava plant in the cassava sample to be tested, and selects high-yielding cassava varieties based on the prediction results.
[0017] The beneficial effects of this invention are as follows:
[0018] (1) Shorten the breeding cycle: It can predict the fresh weight of a single potato at maturity based on genotype data during the seedling stage, without waiting for the plant to mature and be harvested, which significantly shortens the breeding cycle.
[0019] (2) Reduced costs: Reduced the scale of field planting and field management costs of non-potential materials, and reduced the human and material input for phenotyping.
[0020] (3) High prediction accuracy: By integrating whole genome SNP information with machine learning algorithms, the model can explain about 81.6% of phenotypic variation, with low prediction error and good generalization ability.
[0021] (4) High screening efficiency: It realizes the batch and standardized prediction and screening of breeding materials, and improves the efficiency of breeding high-yield cassava varieties. Attached Figure Description
[0022] Figure 1 A flowchart provided for embodiments of the present invention.
[0023] Figure 2 This is a graph showing the results of a genome-wide association study (GWAS) based on 377 cassava samples.
[0024] Figure 3 A comparative diagram of whole-genome selection models for fresh weight of cassava tubers per plant.
[0025] Figure 4 This is a diagram illustrating the various evaluation indicators of the Linear model.
[0026] Figure 5 Scatter plot showing the correlation between actual and predicted values of the whole genome selection Linear model for fresh weight of cassava tubers per plant. Detailed Implementation
[0027] To enable those skilled in the art to better understand the technical solutions of this invention, the present application will be further described in detail below with reference to embodiments.
[0028] Figure 1 A flowchart of the present invention has been shown, and the following detailed description is provided in conjunction with specific embodiments.
[0029] Example 1: Determination of fresh weight phenotypic data of cassava single plant tubers
[0030] The 377 cassava samples used in the experiment were preserved at the experimental base of the Institute of Tropical Biotechnology, Chinese Academy of Tropical Agricultural Sciences (Wenchang). Cassava germplasm resources from different geographical origins and genetic backgrounds were numbered and recorded. At the normal maturity stage or the predetermined harvest time, cassava plants with uniform growth and no obvious pests or diseases were harvested individually in the field. All tubers of each plant were completely dug up, the above-ground parts were removed, and the soil and impurities attached to the tuber surface were cleaned to avoid damaging the tuber tissue. Subsequently, all tubers from the same plant were collected, and their fresh weight was measured using an electronic balance or platform scale. The measured weight was recorded as the fresh weight of the tubers per plant. Two biological replicates were set up for each tested material, and the average value of the results from each replicate was used as the phenotypic data of the fresh weight of the tubers per plant for subsequent statistical analysis and genome-wide selection model construction.
[0031] Example 2: Genome-wide association analysis identifies significant loci for fresh weight of cassava tubers in individual plants.
[0032] 1. Genotyping data determination
[0033] Paired-end sequencing was performed on the 377 cassava samples using GBS (genotyping by sequencing) genome sequencing technology and the Illumina HiSeq sequencing platform. Low-quality sequences and sequences with adapter contamination were removed using Fastp software to obtain high-quality next-generation sequencing data. Using VCFtools software, filtering was performed based on integrity >0.8 and minimum allele frequency (MAF) ≥0.05, resulting in 1,507,463 high-quality SNP loci from the 377 cassava samples for subsequent analysis.
[0034] 2. Genome-wide association analysis
[0035] Based on 1,507,463 SNP loci obtained through filtering, genome-wide association studies (GWAS) were performed on each phenotype using GEMMA software and a linear mixture model (LMM) to identify SNP loci significantly associated with the phenotype. The LMM model effectively considers population structure and kinship, reducing the occurrence of false positives. The analysis results show that ( Figure 2 ), 66,041 loci were significantly associated with the phenotype (P < 0.05), and then genotype files of 377 samples were generated based on VCF, and 5,295 common loci were selected as genotype loci for the genome-wide selection model.
[0036] Example 3: Optimal Genome-Wide Selection Prediction Model
[0037] 1. Preparation of phenotypic and genotypic files
[0038] Genotypic data of selected SNP loci and phenotypic data of 377 cassava tuber fresh weights were used as input data. Basic quality control was performed using VCFtools to screen for biallelic genes with MAF > 0.05. Then, PLINK was used to remove redundant loci based on linkage disequilibrium (LD) to reduce feature dimensionality. Next, missing values were imputed using Tassel and Beagle. Finally, the genotypes were converted to numerical codes (0, 1, 2) and cleaned to generate a genotype feature matrix file that could be directly used for model construction.
[0039] 2. Prediction Model Construction
[0040] Lasso Regression (LassoRegression), Elastic Net Regression, Elastic Net with Cross-Validation (ElasticNetCV), Kernel Ridge Regression, Partial Least Squares Regression (PLS Regression), Support Vector Regression (SVR), Support Vector Regression with Polynomial Kernel (SVR_poly), Random Forest Regressor, and Linear Regression were selected for predictive model building and performance evaluation. Nine GS prediction models were constructed using cross-validation, and all were trained and tested with 100 repeated random sampling to ensure the stability and generalization ability of the results. In each repetition, 80% of the data was randomly sampled as the training set and 20% as the test set to predict the fresh weight of cassava tubers per plant.
[0041] 3. Evaluation of the prediction model
[0042] The coefficient of determination (R²) measures the goodness of fit of the model; a higher value indicates more accurate predictions. The root mean square error (RMSE), the square root of the MSE, more directly reflects the magnitude of the error and quantifies the predictive ability of the models. Based on these indicators, the performance of the nine models in predicting the fresh weight of cassava tubers per plant was ranked and evaluated. The prediction results of different models under multiple cross-validation conditions were compared and analyzed. The results show that there are significant differences in the predictive performance of the models. The Linear model has the most concentrated R² and RMSE values, indicating higher predictive accuracy and stability. Figure 3 ).
[0043] Linear model ;in, This represents the phenotypic value of the i-th individual; This represents the population mean; Indicates the total number of SNP sites; This represents the genotype coding value of the i-th individual at the j-th SNP locus (usually coded as 0, 1, or 2, representing homozygous reference, heterozygous, and homozygous variant, respectively). This represents the effect size of the j-th SNP site; This represents the random error term, which is usually assumed to have a mean of 0 and a variance of 0. The model follows a normal distribution. In the test population, the predictive coefficient of determination (R²) for the fresh weight of individual cassava plants reached as high as 0.816, indicating that the model can explain approximately 81.6% of the phenotypic variation; its prediction error was low, with an RMSE of 0.623, demonstrating high predictive accuracy and good generalization ability. Figure 4 Further analysis of the fit between the actual and predicted values revealed a significant positive correlation between them, and the data distribution was concentrated, with no obvious systematic bias. Figure 5 Based on the comprehensive analysis of model ranking results and prediction performance, the Linear model has a significant advantage in predicting the fresh weight of cassava tubers per plant. It can be used as the core prediction model for the fresh weight of tubers per plant in subsequent breeding populations, for early screening of high-yield potential materials, and to support whole-genome selection breeding practices.
[0044] Example 4: A method for predicting the fresh weight of cassava tubers per plant during the seedling stage
[0045] Young leaf samples were collected from the cassava seedlings (approximately 4-6 true leaves). Genomic DNA was extracted using the conventional CTAB method, and SNP genotyping was performed to obtain the corresponding whole-genome SNP data. After quality control and encoding of the genotype data, it was input into the Linear prediction model selected in Example 3 to calculate the predicted fresh weight of a single cassava plant.
[0046] High-potential materials are screened based on the predicted fresh weight of individual potatoes and given priority for the next round of breeding programs and multi-environment validation trials. Compared with the traditional screening method that requires large-scale phenotypic testing after maturity, the method of this invention can complete early prediction and initial screening at the seedling stage, significantly reducing the field planting and testing costs of low-potential materials and improving breeding screening efficiency and resource utilization.
[0047] In summary, the method of the present invention can complete early prediction and preliminary screening during the seedling stage, significantly reducing the field planting and testing costs of low-potential materials, and improving breeding screening efficiency and resource utilization.
Claims
1. A method for efficiently predicting the fresh weight of cassava tubers per plant based on whole genome, characterized in that, Includes the following steps: (1) Obtain whole genome SNP data of the cassava sample to be tested; (2) Substitute the whole genome SNP data into the prediction model to predict the fresh weight of the cassava plant to be tested; The prediction model is constructed based on SNP sites that are significantly correlated with the fresh weight of individual cassava plants, selected from the cassava training population. The model formula is as follows: ,in This represents the predicted fresh weight of a single potato plant for the i-th individual. This represents the population mean. Indicates the total number of SNP sites. This represents the genotype coding value of the i-th individual at the j-th SNP locus. This represents the effect size of the j-th SNP locus. The random error term is represented; the genotype coding value is 0, 1 or 2, representing the homozygous reference type, heterozygous type and homozygous variant type, respectively.
2. The method according to claim 1, characterized in that, The SNP loci significantly associated with the fresh weight of individual cassava plants were obtained through the following screening method: genome-wide association analysis was performed on the cassava training population, and the association significance between each SNP locus and the fresh weight phenotype of individual cassava plants was calculated using a linear mixture model, and SNP loci with a p-value less than 0.05 were screened; the prediction model was a linear regression model selected after 100 repeated cross-validations, which included a machine learning model including lasso regression, elastic net regression, cross-validation elastic net regression, kernel ridge regression, partial least squares regression, support vector regression, multinomial kernel support vector regression, random forest regression, and linear regression.
3. The method according to claim 1, characterized in that, The cassava sample to be tested in step (1) is a cassava seedling sample, and the collection site is the tender leaf tissue.
4. The method according to claim 1, characterized in that, The method is used to predict the fresh weight of a single cassava plant during the seedling stage.
5. The method according to claim 1, characterized in that, The phenotypic data of single-plant fresh weight of cassava in the cassava training population were determined in the following way: During the normal maturity period or the preset harvest period of cassava, cassava plants with uniform growth in the field were harvested one by one. All tubers of a single plant were dug up completely, the above-ground parts were removed, and the soil and impurities attached to the surface of the tubers were cleaned. All tubers of the same plant were collected, and their fresh weight was weighed using an electronic balance or platform scale. The obtained weight was recorded as the phenotypic data of single-plant fresh weight of cassava for that plant.
6. The method according to any one of claims 1-5, characterized in that, The method is applied to the early screening of cassava hybrid breeding populations.
7. A method for cassava breeding and screening, characterized in that, Using the method described in any one of claims 1-6, the fresh weight of a single cassava plant in the cassava sample to be tested is predicted, and high-yielding cassava varieties are selected based on the prediction results.