Grape fruit longitudinal diameter whole genome selective breeding method
Through the whole genome selection breeding method, using whole genome association analysis and machine learning models, the problem of difficult to accurately control the longitudinal diameter of fruit in grape breeding was solved, and efficient and low-cost fruit longitudinal diameter prediction and breeding were achieved.
Patent Information
- Application Number
- CN202510745798.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-09-16
AI Technical Summary
Existing grape breeding methods make it difficult to accurately control the longitudinal diameter of the fruit, resulting in long breeding cycles, high costs and insufficient variety diversity. In particular, traditional hybrid breeding offspring are highly random and the prediction accuracy of molecular marker-assisted technology is low.
The whole genome selection breeding method is adopted. Through whole genome association analysis and machine learning models, the whole genome data of grape samples are used to predict the longitudinal diameter of the fruit, and the Ridge regression model is combined to achieve high-precision prediction.
It significantly improves the accuracy of fruit longitudinal diameter prediction and breeding efficiency, reduces breeding costs, shortens the breeding cycle, and can achieve precise regulation of fruit longitudinal diameter at the seedling stage.
Smart Images

Figure CN120656536A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of breeding and screening, and in particular to a method for whole-genome selection breeding of grape fruit longitudinal diameter. Background Art
[0002] Berry diameter refers to the diameter of a grape berry along its longitudinal axis (i.e., its maximum length), and is a key indicator of grape size and shape. Berry diameter is not only closely related to grape variety characteristics, cultivation environment, and management practices, but also has a significant impact on both winemaking and fresh-eating wines. In winemaking, berry diameter directly reflects berry maturity and flavor intensity. Grapes with longer diameters typically contain more skin and flesh, resulting in a higher skin-to-flesh ratio. This helps extract tannins and increase pigment concentration, thereby enhancing the wine's structure and color richness. Especially for red wines, a larger diameter contributes to a richer body and more complex taste. However, excessively large berries can lead to uneven sugar accumulation, while a low skin-to-flesh ratio can result in a thin wine with a lack of depth. Therefore, controlling berry diameter in wine grape management can help optimize grape quality and ensure the final wine possesses the desired taste and structure. In fresh-eating grape production, berry diameter is a key appearance indicator. Grapes with larger berries have a fuller, more attractive shape and are therefore generally preferred by consumers. Larger berries generally mean more juice and a better eating experience, resulting in a juicier, sweeter taste and a richer flesh, which enhances the chewiness and texture of fresh-eating grapes. However, excessively enlarged berries can lead to thickened skin and loose flesh, affecting taste and storage. They can also disrupt the sugar-acid balance due to uneven distribution of water and nutrients during growth. Therefore, in the cultivation of fresh-eating grapes, proper control of berry diameter is crucial for ensuring the taste, appearance, and storage properties of the fruit.
[0003] Grape breeding methods primarily fall into two categories: traditional methods and modern technologies. Traditional methods include hybridization, bud mutation selection, and seed selection. Hybridization can integrate the superior traits of different parental lines to produce new varieties with superior overall performance, such as disease resistance, stress tolerance, or high-quality varieties. However, hybrid offspring are prone to trait segregation, requiring a long and complex selection and breeding process, and are limited by the gene pools of the parents. Bud mutation selection leverages natural plant variation to uncover new genetic resources, but the frequency of variation is low and the screening workload is substantial. Seed selection involves sowing seeds to select seedlings with superior traits, but this can result in a long breeding cycle and severe trait segregation in offspring. Modern technologies have introduced methods such as marker-assisted selection, gene editing, and whole-genome selection. Marker-assisted selection utilizes molecular markers to rapidly and accurately detect plant genetic information, improving breeding efficiency and accuracy. However, this method is technically challenging and expensive. Gene editing can modify or knock out specific genes, overcoming the limitations of traditional breeding methods, but its effectiveness in improving quantitative traits is limited. Whole-genome selection breeding, an emerging breeding method, integrates high-throughput sequencing technology with bioinformatics. By constructing predictive models to evaluate the genetic information of individual whole genomes, this method enables rapid and accurate early selection. This method fully utilizes whole-genome genetic information to significantly improve breeding efficiency, shorten breeding cycles, and reduce breeding costs. With the continuous development and improvement of biotechnology, whole-genome selection breeding will play an even greater role in grape breeding, strongly promoting the innovative development of the grape industry.
[0004] In summary, given the significant impact of fruit diameter on the grape industry, using whole-genome selection breeding to accurately predict fruit diameter will significantly accelerate the selection and breeding process of new grape varieties, achieve precise regulation and customized cultivation of fruit diameter, thereby effectively promoting the innovative development of the grape industry and meeting the demand for high-quality grapes in the winemaking and fresh food markets. Summary of the Invention
[0005] Traditional grape breeding relies primarily on crossbreeding based on subjective experience. However, grapes have a complex genetic background. Fruit diameter, a typical quantitative trait, is regulated by multiple genes. This results in a high degree of randomness in the offspring of hybrid breeding. During the breeding process, obtaining individuals that meet expectations is extremely difficult, and precise improvement of traits like fruit diameter is even more difficult. Furthermore, grapes are perennial crops, requiring several years of juvenile growth from seedling cultivation to flowering and fruiting, significantly reducing breeding efficiency. Furthermore, fruit volume is highly susceptible to growing conditions and environmental factors, and accurate assessment often requires years of continuous observation after yields stabilize. While molecular marker-assisted techniques can predict traits to a certain extent, they are often based on a single or limited number of markers. Complex quantitative traits like fruit diameter are the result of the combined effects of multiple genes and cannot be fully influenced by a single gene or single variant locus. By ignoring the complex interactions between genes, predictions fail to accurately reflect the actual grape fruit diameter, significantly reducing the accuracy and reliability of the predictions. These limitations greatly prolong the conventional grape breeding cycle, increase breeding costs, and reduce breeding results, which is also a major reason for the lack of variety diversity in the current grape market. The present invention provides a grape fruit longitudinal diameter whole genome selection breeding method, comprising the following steps:
[0006] Step 1: Count the longitudinal diameters of the berries of 323 grape samples and collect leaf samples for whole-genome next-generation sequencing to obtain next-generation sequencing data;
[0007] Use Fastp software to perform quality control on the second-generation sequencing data to obtain high-quality second-generation sequencing data;
[0008] Step 2: Use GTX software to map high-quality second-generation sequencing data to the PNT2T genome for variant detection and obtain a variant site dataset;
[0009] Step 3: Divide the 323 samples into a training set and a test set, with 259 samples in the training set and 64 samples in the test set;
[0010] Step 4: Perform genome-wide association analysis based on the variant site dataset and phenotypic data of the training set;
[0011] Step 5: Sort the variant site dataset according to the p-value of the genome-wide association analysis results, and divide the variant sites according to the gradient;
[0012] The dimensionality of the variant site dataset is reduced according to the linkage disequilibrium (LD) relationship, and a representative site in each LD region is selected as input data;
[0013] Step 6: Import the associated variant sites and phenotypic data from the training set in Step 5 into five classic regression models to train models for predicting grape longitudinal diameter, and select the model with the highest accuracy;
[0014] Step 7: Import the variant site datasets in the 64 test sets into the selected regression model for prediction; evaluate the accuracy of the model.
[0015] Furthermore, the quality control process in step 1 includes excluding low-quality data and data with adapter contamination in the second-generation sequencing data.
[0016] Furthermore, in the variant site dataset in step 2, Beagle software is used to fill in missing values in the variant site dataset.
[0017] Furthermore, the genome-wide association analysis process in step 4 includes analyzing the correlation between the variant site dataset and the trait data using a mixed linear model of GEMMA and a Wald P-value test.
[0018] Furthermore, in step 5, the order of sorting the p values is from small to large.
[0019] Furthermore, in step five, the gradient partitioning of the variant sites is performed by taking the first 100, 500, 1000, 5000, 100000, 500000, and 1000000 sites respectively.
[0020] Furthermore, the regression model in step 6 includes ridge regression, lasso regression, linear elastic network regression, linear support regression and support vector regression-polynomial kernel.
[0021] Compared with the prior art, the technical solution provided by the present invention has the following beneficial effects:
[0022] The proposed genomic selection prediction method for grape berry longitudinal diameter size, based on 5,000 variable dimensionality reduction loci and a Ridge regression model, achieved a linear correlation coefficient of R = 0.96 between model predictions and actual phenotypic values in 64 test samples, demonstrating that the model can replace manual testing. This method significantly reduces the screening cost of grape hybrid progeny and improves the prediction of berry longitudinal diameter, possessing important application value in breeding. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 This is a flowchart of the grape fruit longitudinal diameter genome prediction process of the present invention;
[0024] Figure 2 The results of the genome-wide association analysis of the longitudinal diameter of 259 grape berries of the present invention;
[0025] Figure 3 Comparison of the prediction results of full gene selection based on machine learning in the present invention;
[0026] Figure 4 . is the correlation between the predicted value and the actual measured value under the optimal model of the present invention. DETAILED DESCRIPTION
[0027] The present invention will be described in further detail below with reference to the accompanying drawings and specific embodiments. The embodiments of the present invention are provided for purposes of illustration and description and are not intended to be exhaustive or to limit the invention to the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. The embodiments are chosen and described to better illustrate the principles of the invention and its practical application, and to enable those skilled in the art to understand the invention and design various embodiments with various modifications suitable for specific applications.
[0028] Example, 1) Sample data collection
[0029] The longitudinal diameter of berries from 323 grape samples was counted, and leaf samples were collected for whole-genome next-generation sequencing. Sequencing data were quality-controlled using Fastp software to exclude low-quality data and data with adapter contamination, generating high-quality next-generation sequencing data. High-quality next-generation sequencing data were mapped to the PNT2T genome using GTX software (v. 2.1.11, http: / / www.gtxlab.com / ) for variant detection, generating a dataset of variant loci. Missing values in the VCF files were then filled using Beagle software.
[0030] 2) Sample division and genome-wide association analysis
[0031] The 323 samples were divided into a training set and a test set for subsequent machine learning model training and model evaluation. The training set contained 259 samples and the test set contained 64 samples.
[0032] Based on the variant sites and phenotypic data of 259 samples in the training set, genome-wide association analysis was performed, and the correlation between variant data and trait data was analyzed using the GEMMA mixed linear model (-lmm) and Wald P value test ( Figure 2Next, the variants were ranked by p-value from smallest to largest in the genome-wide association analysis, and the top 100, 500, 1,000, 5,000, 10,000, 50,000, 100,000, 500,000, and 1,000,000 variants were selected using a gradient partitioning algorithm. These variants were most linearly correlated with grape berry longitudinal diameter. Dimensionality reduction was then performed on the variant dataset based on linkage disequilibrium (LD) relationships, with a representative locus in each LD region selected as input data.
[0033] 3) Model training and evaluation
[0034] The associated variant sites and phenotypic data in the above training set were imported into 5 classic regression models to train the model for predicting the longitudinal diameter of grapes, and the model with the highest accuracy was selected ( Figure 3 ). These models include ridge regression (ridge), lasso regression (Lasso), linear elastic network regression (ElasticNet), linear support regression (SVR_linear) and support vector regression-polynomial kernel (SVR_poly). Then, the high-quality variant site data in the 64 test sets were imported into the screened models with high prediction accuracy for further prediction. The predicted values were compared and the best prediction model was selected. The results showed that when the number of variant sites was 5000, the redundant Ridge regression model had the highest prediction accuracy, and there was a significant linear relationship between the actual value and the predicted value (p < 2.2e-16), and the correlation coefficient R = 0.96. This shows that the prediction effect of the model has reached the level of manual detection, which can replace tedious experiments and manual screening, and can predict the longitudinal diameter of the fruit of hybrid offspring at the seedling stage. ( Figure 4 ).
[0035] Obviously, the embodiments described are only some of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field and related fields without making creative efforts should fall within the scope of protection of the present invention. Structures, devices, and operating methods not specifically described and explained in the present invention shall be implemented in accordance with conventional means in the field unless otherwise specified or limited.
Claims
1. A method for whole-genome selection breeding of grape fruit longitudinal diameter, characterized in that: The following steps are involved: Step 1: Count the longitudinal diameters of the berries of 323 grape samples and collect leaf samples for whole-genome next-generation sequencing to obtain next-generation sequencing data; Use Fastp software to perform quality control on the second-generation sequencing data to obtain high-quality second-generation sequencing data; Step 2: Use GTX software to map high-quality second-generation sequencing data to the PNT2T genome for variant detection and obtain a variant site dataset; Step 3: Divide the 323 samples into a training set and a test set, with 259 samples in the training set and 64 samples in the test set; Step 4: Perform genome-wide association analysis based on the variant site dataset and phenotypic data of the training set; Step 5: Sort the variant site dataset according to the p-value of the genome-wide association analysis results, and divide the variant sites according to the gradient; The dimensionality of the variant site dataset is reduced according to the linkage disequilibrium (LD) relationship, and a representative site in each LD region is selected as input data; Step 6: Import the associated variant sites and phenotypic data from the training set in Step 5 into five classic regression models to train models for predicting grape longitudinal diameter, and select the model with the highest accuracy; Step 7: Import the variant site datasets in the 64 test sets into the selected regression model for prediction; evaluate the accuracy of the model.
2. The method for whole genome selection breeding of grape fruit longitudinal diameter according to claim 1, characterized in that: The quality control process in step 1 includes excluding low-quality data and data with connector contamination in the second-generation sequencing data.
3. The method for whole genome selection breeding of grape fruit longitudinal diameter according to claim 1, characterized in that: In the step 2, Beagle software is used to fill in the missing values of the variant site dataset.
4. The method for whole-genome selection and breeding of grape fruit longitudinal diameter according to claim 1, characterized in that: The genome-wide association analysis process in step 4 includes using a mixed linear model of GEMMA and a Wald P-value test to analyze the correlation between the variant site dataset and the trait data.
5. The method for whole-genome selection and breeding of grape fruit longitudinal diameter according to claim 1, characterized in that: In step 5, the order of sorting the p values is from small to large.
6. The grape berry longitudinal diameter whole genome selection breeding method according to claim 1, characterized in that: In the step 5, the gradient division variation sites are respectively taken as the first 100, 500, 1000, 5000, 10000, 50000, 100000, 500000, and 1000000 sites.
7. The method for whole-genome selection and breeding of grape berry longitudinal diameter according to claim 1, characterized in that: The regression models in step six include ridge regression, lasso regression, linear elastic network regression, linear support regression and support vector regression-polynomial kernel.