Grape fruit volume whole genome selective breeding method

Through the whole genome selection breeding method, combined with high-throughput sequencing and bioinformatics, a prediction model was constructed to solve the problem of fruit volume control in traditional grape breeding, achieve early screening and precise breeding, and improve breeding efficiency and fruit volume control effects.

CN120613006APending Publication Date: 2025-09-09AGRICULTURAL GENOMICS INSTITUTE AT SHENZHEN CHINESE ACADEMY OF AGRICULTURAL SCIENCES (SHENZHEN BRANCH GUANGDONG LABORATORY FOR LINGNAN MODERN AGRICULTURE)
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510751666.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Traditional grape breeding methods make it difficult to accurately control fruit volume, resulting in low breeding efficiency, high costs, and difficulty in meeting the market's demand for diverse grape varieties.

Method used

A whole-genome selection breeding method is used, combined with high-throughput sequencing and bioinformatics, to construct a prediction model, and machine learning technology is used to screen out individuals with genetic potential for large fruits, achieving early and accurate screening.

Benefits of technology

Significantly shorten the breeding cycle, improve breeding efficiency, reduce costs, achieve precise control of fruit volume, and meet the market demand for diverse grape varieties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120613006A_ABST
    Figure CN120613006A_ABST
Patent Text Reader

Abstract

The invention relates to a grape fruit volume whole genome selective breeding method. The method comprises the following steps: preparing data; the method comprises the following steps: collecting a sample of a grape germplasm resource, and carrying out fruit volume measurement on the sample; performing whole genome sequencing on the samples, and excluding unqualified samples with low quality and joint pollution in the samples to obtain high-quality second-generation sequencing data; and mapping the high-quality next-generation sequencing data to a high-quality reference genome PN40024 for variation typing, performing variation calling in GTX software, filtering and screening to obtain effective variation sites, and obtaining variation sites. The invention relates to the technical field of plant genetic breeding. According to the grape fruit volume whole genome selective breeding method, the limitation of a traditional breeding method in the aspect of increasing the grape fruit volume is overcome, wide application prospects and important popularization value are brought to the field of grape breeding, and sustainable development and progress of the grape industry are expected to be promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of plant genetic breeding, in particular to a method for whole-genome selection breeding of grape fruit volume. Background Art

[0002] Fruit volume, a key physical characteristic of grape berries, is a crucial indicator of berry size. For table grape varieties, berry volume directly influences consumer perception and purchasing intention: larger berries are often more popular because they are not only more attractive in appearance but also have a higher edible percentage, providing a better eating experience. Furthermore, larger berries that meet consumer taste and flavor expectations better meet consumer expectations, enhancing the grape's market competitiveness and added value. However, larger berry volume is not always better. For wine grape varieties, berry volume is also a crucial factor: only moderate berry volume can produce richer, higher-quality wines. Overly large berries hinder the release of flavor compounds during fermentation and may indicate excessive water content, which can dilute the juice's flavor and reduce the wine's alcohol and flavor concentration. On the other hand, undersized berries may result in a high seed content and a low pulp content, affecting the ratio and content of tannins and sugars, ultimately compromising the wine's quality and taste. Therefore, grape breeders need to fully consider the important factor of fruit volume and use scientific breeding methods to better meet the market's different demands for fresh grapes and wine grapes, so as to promote the sustainable development and prosperity of the grape industry.

[0003] Common grape breeding methods include hybridization, bud mutation selection, and seed selection, each with its own advantages and disadvantages. Hybridization combines the superior traits of different parental lines to create hybrid offspring with improved overall performance, facilitating the development of new varieties with disease resistance, stress tolerance, or superior quality. However, hybrid offspring can exhibit trait segregation, requiring extensive selection and cultivation, a complex and time-consuming process that is also limited by the gene pools of the parents. Bud mutation selection leverages natural plant variation to discover new genetic resources, but the frequency of variation is low and the screening effort is arduous. Seed selection involves sowing seeds to obtain seedlings and selecting individuals with superior traits, but this requires a longer breeding cycle and results in greater trait segregation in offspring. In addition to these traditional breeding methods, modern breeding techniques such as marker-assisted selection and gene editing are also increasingly being applied to grape breeding. Marker-assisted selection utilizes molecular markers to rapidly and accurately detect and analyze genetic information in plants, improving breeding efficiency and accuracy. However, this method requires high technical requirements and is relatively expensive. Gene editing breeding can modify or knock out specific genes in a targeted manner to create individuals with new traits, breaking through the limitations of traditional breeding. However, the technology is difficult and is not very effective for quantitative traits.

[0004] Whole-genome selection breeding, an emerging breeding method, combines high-throughput sequencing technology with bioinformatics methods. By constructing predictive models to evaluate the genetic information of individual grapes across the entire genome, it enables early, rapid, and accurate selection. Whole-genome selection breeding fully utilizes the genetic information of the entire genome, improving breeding efficiency and accuracy, shortening the breeding cycle, and reducing breeding costs. With the continuous development and improvement of biotechnology, whole-genome selection breeding is expected to play an increasingly important role in grape breeding, providing strong support for the sustainable development and innovation of the grape industry.

[0005] Traditional grape breeding methods rely primarily on empirical hybridization. However, due to the complex genetic background of grapes, and the fact that fruit volume is a quantitative trait controlled by multiple genes, traditional methods struggle to precisely dissect the interactions between these genes. Consequently, this approach is subject to significant randomness, making it difficult to select individuals that meet expectations in subsequent generations, and even more challenging to achieve precise trait improvement. Furthermore, as a perennial crop, grapes undergo several years of juvenile growth from seedling to fruiting, resulting in low breeding efficiency. Furthermore, fruit volume is susceptible to growing conditions and environmental factors, and accurate assessment often requires years of continuous observation after yield stabilization. These limitations significantly prolong conventional grape breeding cycles, increase breeding costs, and reduce breeding effectiveness, contributing to the current lack of variety diversity in the grape market. Therefore, there is an urgent need to explore and innovate breeding methods and techniques to improve both efficiency and accuracy. Summary of the Invention

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A grape fruit volume whole genome selection breeding method comprises the following steps:

[0007] Step 1: Data preparation;

[0008] collecting samples of grape germplasm resources and measuring the fruit volume of the samples;

[0009] Perform whole genome sequencing on the samples, exclude low-quality samples and unqualified samples with adapter contamination, and obtain high-quality second-generation sequencing data;

[0010] The high-quality second-generation sequencing data were mapped to the high-quality reference genome PN40024 for variant typing, and GTX software was used to perform variant calling and filtering to obtain effective variant sites, thereby obtaining variant sites;

[0011] Based on the variant sites and phenotypic data, a genome-wide association analysis is performed to screen out high-quality variant sites;

[0012] The variant sites are sorted from small to large according to the p-value of the whole genome association analysis results, that is, they are sorted from strong to weak according to the phenotypic association.

[0013] Step 2: Model construction and evaluation;

[0014] All samples were randomly divided into training and test sets, with 80% of the samples used as training sets and 20% of the samples used as test sets, and this was repeated 100 times for subsequent analysis;

[0015] Use the data from the training set to build the model;

[0016] During the model construction process, the genotype data of the sample is used as the feature variable. The genotype data is specifically the genotype of the variant site, and different numbers of variant sites are selected according to the gradient;

[0017] The model building method used five models, including linear elastic network regression, lasso regression, ridge regression, linear support regression, and support vector regression-polynomial kernel model;

[0018] Each trained model is evaluated using the test set data. Missing values ​​at valid variant sites in the test set samples must be filled with training set data. The accuracy of the prediction is analyzed by comparing the predicted fruit volume with the actual sample volume.

[0019] A comprehensive comparison was made of 100 models constructed with different sample divisions, different numbers of variant sites, and different model construction methods, and the model with the best prediction accuracy was selected as the whole genome selection model.

[0020] Preferably, the first 100, 500, 1000, 5000, 10000, 50000, 100000, 500000, and 1000000 mutation sites are selected for model construction.

[0021] Preferably, the phenotypic data of the sample is used as the target variable in the model building process.

[0022] The present invention provides a whole-genome selection breeding method for grape fruit size, which has the following beneficial effects:

[0023] (1) This whole-genome selection breeding method for grape berry size integrates machine learning and genomic prediction technologies to construct a highly efficient prediction model. This model enables early screening of breeding populations, selecting individuals with genetic potential for large fruit by accurately predicting fruit size, significantly shortening the breeding cycle and significantly improving breeding efficiency. This method not only overcomes the limitations of traditional breeding methods in achieving increased grape berry size, but also offers broad application prospects and significant promotional value in the field of grape breeding, and is expected to promote the sustainable development and progress of the grape industry.

[0024] (2) This whole-genome selection breeding method for grape berry volume demonstrates the feasibility of machine learning for phenotypic prediction. Based on a dimensionality-reduced dataset of 5,000 variant sites and a Rigde regression model, the present invention can efficiently and accurately predict grape berry volume. In the test set, the model's predicted values ​​showed a significant linear correlation with the actual phenotypic values ​​(p < 2.2e-16), with a Pearson correlation coefficient of R = 0.94, reaching a level that can replace manual screening. This method is particularly effective for early fruit volume prediction and screening of seedlings in the juvenile stage (vegetative reproductive stage).

[0025] (3) This genome-wide selection breeding method for grape berry size, based on 5,000 variable dimensionality reduction loci and a Ridge regression model, achieved a linear correlation coefficient of R = 0.94 between model predictions and actual phenotypic values ​​in 64 test samples, demonstrating that the model can replace manual testing. This method significantly reduces the cost of screening grape hybrid progeny and improves the prediction of berry size, possessing important breeding application value. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 This is a flowchart of the whole genome prediction of grape fruit volume of the present invention;

[0027] Figure 2 This is a genome-wide association analysis diagram of grape fruit volume of the present invention;

[0028] Figure 3 This is a comparison chart of the prediction effects of the models ElasticNet, Lasso, Ridge, SVR-linear and SVR-ploy of the present invention. DETAILED DESCRIPTION

[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0030] See also Figure 1-3 , the present invention provides a technical solution:

[0031] The whole genome selection breeding method for grape berry volume includes the following steps:

[0032] Step 1: Data preparation;

[0033] collecting samples of grape germplasm resources and measuring the fruit volume of the samples;

[0034] Perform whole genome sequencing on the samples, exclude low-quality samples and unqualified samples with adapter contamination, and obtain high-quality second-generation sequencing data;

[0035] The high-quality second-generation sequencing data were mapped to the high-quality reference genome PN40024 for variant typing, and GTX software was used to perform variant calling and filtering to obtain effective variant sites, thereby obtaining variant sites;

[0036] Based on the variant sites and phenotypic data, a genome-wide association analysis is performed to screen out high-quality variant sites;

[0037] The variant sites are sorted from small to large according to the p-value of the whole genome association analysis results, that is, they are sorted from strong to weak according to the phenotypic correlation.

[0038] Step 2: Model construction and evaluation;

[0039] All samples were randomly divided into training and test sets, with 80% of the samples used as training sets and 20% of the samples used as test sets, and this was repeated 100 times for subsequent analysis;

[0040] Use the data from the training set to build the model;

[0041] During the model construction process, the genotype data of the sample is used as the feature variable. The genotype data is specifically the genotype of the variant site, and different numbers of variant sites are selected according to the gradient;

[0042] The model building method used five models, including linear elastic network regression, lasso regression, ridge regression, linear support regression, and support vector regression-polynomial kernel model;

[0043] Each trained model is evaluated using the test set data. Missing values ​​at valid variant sites in the test set samples must be filled with training set data. The accuracy of the prediction is analyzed by comparing the predicted fruit volume with the actual sample volume.

[0044] A comprehensive comparison was made of 100 models constructed with different sample divisions, different numbers of variant sites, and different model construction methods, and the model with the best prediction accuracy was selected as the whole genome selection model.

[0045] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.

[0046] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A method for whole-genome selection breeding of grape berry volume, characterized in that: The following steps are involved: Step 1: Data preparation; collecting samples of grape germplasm resources and measuring the fruit volume of the samples; Perform whole genome sequencing on the samples, exclude unqualified samples of low quality and with adapter contamination, and obtain next-generation sequencing data; Mapping the second-generation sequencing data to the reference genome PN40024 for variant typing to obtain variant sites; Based on the variant sites and phenotype data, perform genome-wide association analysis to screen high-quality variant sites; Sort the variant sites according to the p-value of the genome-wide association analysis results; Step 2: Model construction and evaluation; All samples are randomly divided into training set and test set, 80% of the samples are used as training set and 20% of the samples are used as test set; Use the data from the training set to build the model; During the model construction process, the genotype data of the sample is used as the feature variable. The genotype data is specifically the genotype of the variant site, and different numbers of variant sites are selected according to the gradient; The model building method used five models, including linear elastic network regression, lasso regression, ridge regression, linear support regression, and support vector regression-polynomial kernel model; The constructed models were comprehensively compared and the model with the best prediction accuracy was selected as the whole genome selection model.

2. The grape berry volume whole genome selection breeding method according to claim 1, characterized in that: The first 100, 500, 1000, 5000, 10000, 50000, 100000, 500000, and 1000000 mutation sites were selected for model construction.

3. The grape berry volume whole genome selection breeding method according to claim 1, characterized in that: The phenotypic data of the samples are used as target variables in the model construction process.