A method for breeding high-oleic acid and high-yield peanuts based on genome-wide association analysis and genome-wide selection
Through genome-wide association analysis and genome-wide selection methods, TagSNP was developed and reference groups were constructed to detect and predict peanut hybrid offspring, which solved the problem of difficulty in polymerizing high oleic acid and high yield in the existing technology, and achieved efficient breeding of new peanut varieties.
Patent Information
- Application Number
- CN202410307529.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-03-18
AI Technical Summary
The prior art is difficult to polymerize new peanut varieties with high oleic acid and high yields, resulting in ineffective breeding.
Using genome-wide association analysis and genome-wide selection methods, TagSNP is developed by obtaining sites that control oleic acid traits, and using genome-wide selection to construct reference groups, low-generation detection and high-generation prediction of hybrid offspring are carried out to effectively polymerize high oleic acid and high-yield traits.
It has achieved excellent traits of rapid polymerization, improved breeding efficiency, saved time and cost, and selected new varieties of high oleic acid and high yield peanuts.
Smart Images

Figure CN118272562B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the field of plant genetic breeding, and in particular relates to a breeding method for high-oleic acid and high-yield peanuts based on whole genome association analysis and whole genome selection. Background Art
[0002] Peanuts are important oil crops and cash crops in my country. High-oleic peanuts have many advantages, such as good taste and low price, comparable to olive oil; not easy to be oxidized, peanuts and their products have a long shelf life; reduce harmful cholesterol, maintain beneficial cholesterol, and prevent the occurrence of cardiovascular and cerebrovascular diseases (Smith GD, Song F, Sheldon T A. Cholesterol lowering and mortality: the importance of considering initial level of risk [J]. BMJ, 1993, 306 (6889): 1367-1373.). In addition, the consumption of peanut oil is on the rise, so it is particularly important to cultivate new varieties of high-yield, high-oleic peanuts, which has become the main breeding goal of peanut breeders.
[0003] Genome-wide association study (GWAS) was initially used in humans and animals, and was first used in plant genetic research in 2001 (Thornsberry JM, Goodman MM, Doebley J, Kresovich S, Nielsen D, Buckler ES. Dwarf8 polymorphisms associate with variation in flowering time [J]. Nature genetics, 2001, 28 (3)). It sequences individuals in a crop population, detects genetic variation polymorphic sites SNP (Single Nucleotide Polymorphism) across the entire genome, combines the detected genotype data with phenotypic trait data, conducts population-level statistical analysis, discovers sites that control important agronomic traits, and provides scientific and technological support for the selection and breeding of new peanut varieties. Whole genome selection (GS) uses high-density SNP markers covering the entire genome and combines them with phenotypes to estimate the breeding values of individual crops. It assumes that at least one of these markers is in linkage disequilibrium with all quantitative trait loci that control all traits. The genotype and phenotypic information of the population is collected for modeling and construction of a reference population to predict the phenotypic values of only genotyped individuals (Grevenhof IEV, Werf JHV D. Design of reference populations forgenomic selection in crossbreeding programs[J]. Genetics Selection Evolution, 2015, 47(1):14).
[0004] The breeding of new high-oleic acid and high-yield peanut varieties requires the simultaneous aggregation of oleic acid traits and yield traits, but both are quantitative traits. Conventional breeding experience and the naked eye cannot effectively aggregate excellent traits. With the development of molecular biotechnology, we can easily obtain the genotype data of breeding materials, and use GWAS and GS technology to screen offspring materials. The low-generation detection and screening of excellent traits will greatly and effectively accelerate the breeding process of high-yield and high-quality peanuts. Summary of the invention
[0005] The purpose of the present invention is to provide a method for breeding high-oleic acid and high-yield peanuts by combining whole genome association analysis and whole genome selection, and to quickly aggregate excellent traits to improve breeding efficiency. Specifically, the present invention uses whole genome association analysis to obtain sites that control oleic acid traits, develops TagSNPs, uses whole genome selection to construct a reference group, and performs low-generation detection and high-generation prediction on hybrid offspring, respectively, effectively aggregates high oleic acid and high-yield traits, saves time and cost, and selects new varieties of high-oleic acid and high-yield peanuts.
[0006] The present invention adopts the following technical solutions:
[0007] The breeding method of high-oleic acid and high-yield peanuts based on whole genome association analysis and whole genome selection includes the following steps:
[0008] Step 1: Use genome-wide association analysis to obtain loci that control peanut oleic acid traits, develop TagSNPs, and verify them. Specifically include:
[0009] Step 1.1: Select peanut materials (more than 100 copies) to form a natural population, conduct multi-year multi-point field trials, examine the oleic acid content and single-plant productivity of each material after maturity, harvest and dry, remove outliers, and use a mixed linear model to correct the oleic acid content and single-plant productivity to obtain the respective best linear unbiased estimates as the corrected phenotypic values, thereby obtaining the phenotypic data of the population materials.
[0010] Step 1.2: Use the second-generation resequencing technology to perform resequencing of each material in the population to a depth of no less than 10×, and perform polymorphic variation site detection (call SNP) to obtain the genotype data of the population material.
[0011] Step 1.3: Use the phenotypic data (oleic acid content) and genotypic data of the population materials to conduct genome-wide association analysis, explore the sites that control the oleic acid content of peanut, and develop a set of TagSNPs.
[0012] Step 1.4: Select more than 20 peanut materials (validation group) with resequencing data to verify TagSNP, with an accuracy of no less than 0.9.
[0013] Step 2: Use whole genome selection to construct a reference population, including:
[0014] Step 2.1: Use the GBLUP model to calculate the genetic parameters and breeding values of the productivity of each material and construct a reference population.
[0015] Step 2.2: Select more than 20 peanut validation groups to validate the reference group, with an accuracy of no less than 0.4.
[0016] Step 3: Conduct low-generation detection and high-generation prediction on the hybrid offspring to effectively aggregate high oleic acid and high-yield traits, including:
[0017] Step 3.1: Breeding of high-oleic acid and high-yield peanut varieties: configure hybrid combinations (no less than 20), select the offspring using the pedigree method, sample individual plants in the F2 generation at the seedling stage, extract DNA for preservation, and screen individual plants with more than 20 full fruits at harvest. Use high-oleic acid TagSNP to perform molecular marker-assisted selection on individual plants with more than 20 full fruits to obtain multiple high-oleic acid plants with equivalent yields.
[0018] Step 3.2: The individual plants with low yield (less than 20 full fruits) were eliminated from the continuous self-pollination progeny population until the F5 generation seedling stage, when individual plant sampling was performed and DNA was extracted and preserved. When harvesting, high-yield individual plants (more than 20 full fruits) were retained for resequencing (10×), and the reference population constructed by whole genome selection was used to predict the yield of high-yield plants to obtain high-oleic acid high-yield individual plants.
[0019] Step 3.3: After the F6 generation plants are harvested, they can participate in yield tests to breed new high-oleic acid and high-yield peanut varieties.
[0020] The beneficial effects of the present invention are:
[0021] 1. The present invention performed deep sequencing on 169 materials to obtain genotype data of high-density variable sites, and multi-point field identification obtained reliable phenotypic data, which can improve the accuracy of GWAS and GS.
[0022] 2. The present invention uses GWAS results to obtain a group of SNP sites that control oleic acid content.
[0023] 3. The present invention combines Block analysis and group stepwise regression methods to develop three TagSNPs, with a verification accuracy rate of up to 0.9571.
[0024] 4. The present invention uses GS to model yield traits. In order to improve the accuracy of the model, 20 varieties are randomly selected from the population, and their phenotypic data are removed for modeling. Then, the model is used to predict the variables of the 20 varieties, and the correlation coefficient is 0.4120, which is highly accurate. According to statistics, compared with conventional selection, 66.7% of the promoted lines can be retained when the planting scale is reduced by 50%, and the accuracy is improved by 33.33%.
[0025] 5. The present invention combines the developed TagSNP with high-precision GS and applies it to the breeding process of new high-oleic acid and high-yield peanut varieties.
[0026] 6. The oleic acid trait is controlled by the major effect gene and can be screened in the early generation. The present invention uses the discovered TagSNP to detect F2 generation plants with high yields and screen high oleic acid materials, thereby improving the accuracy and efficiency of breeding.
[0027] 7. The yield trait is controlled by multiple micro-effect genes and is easily affected by the environment. The present invention uses the GS reference group to batch predict the yield of high-oleic acid materials in high-generations such as F5, eliminating the need to conduct multi-year multi-point yield tests, saving costs and improving breeding efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] Figure 1 These are histograms of phenotypic data for two traits; A: histogram of oleic acid content; B: histogram of single plant productivity.
[0029] Figure 2 is the density distribution of SNPs on peanut chromosomes.
[0030] Figure 3 Manhattan plot of oleic acid content.
[0031] Figure 4 This is the QQ plot of oleic acid content.
[0032] Figure 5 Block linkage map of significant loci.
[0033] Figure 6 Select the promotion rate statistic for the genome.
[0034] Figure 7 Correlation between predicted breeding values and true values.
[0035] Figure 8 This is a flow chart for breeding high oleic acid and high-yield peanut varieties. DETAILED DESCRIPTION
[0036] The present invention is described in more detail below through specific implementation modes to facilitate understanding of the technical solution of the present invention, but is not intended to limit the protection scope of the present invention.
[0037] 1. Materials
[0038] In this embodiment, the field planting population of peanuts includes 169 varieties or strains (Table 1), among which the Kainong series peanut varieties, strains, Qiulehua 177, and Kaixuan 016 are provided by Kaifeng Academy of Agriculture and Forestry Sciences, the Jihua series peanut varieties are provided by Hebei Academy of Agriculture and Forestry Sciences, the Zhonghua series peanut varieties are provided by the Oil Crops Research Institute of the Chinese Academy of Agricultural Sciences, the Huayu series peanut varieties are provided by Shandong Peanut Institute, and G168 (AT1-1) is introduced from Georgia, USA. 20 materials were randomly selected from the 169 materials as the validation population for whole genome association analysis and whole genome selection (Table 2), and the remaining 149 materials were subjected to GWAS and GS analysis.
[0039] Table 1. Information on 169 peanut populations
[0040]
[0041]
[0042] Table 2. Material information for association analysis and genome-wide selection accuracy verification
[0043] serial number Variety (line) name serial number Variety (line) name G1 Kai Nong 30 G81 1530G-OG-2-1N-0 G8 1225-4 G84 0317-26 G16 1443-0G-7G-7-0 G91 0317-5 G25 1483-0G-42G-1-OA G94 1359-0-(0)-6-0 G57 1327-0-4-2-0 G126 0317-25 G64 1333-0-(0)-4-0 G146 Kainong92 G70 Kainong 99 (1255-1) G150 0937-1 G73 0972-1 G165 0979-3 G76 0977-1 G186 Jihua16 G79 0974-0-69-3-1-1-0 G190 Kainong 306 (0317-60)
[0044] 2. Test methods and results
[0045] 1.1 Phenotypic data
[0046] 1.1.1 Field trial design and phenotypic investigation
[0047] The 169 materials were phenotypically identified at multiple sites over many years, with a total of four environments, namely the experimental field of Kaifeng Academy of Agriculture and Forestry Sciences in 2019 (E1), Zhaomiao Village, Hudian Township, Pingqiao District, Xinyang in 2019 (E2), the experimental field of Kaifeng Academy of Agriculture and Forestry Sciences in 2020 (E3), and the experimental field of Kaifeng Academy of Agriculture and Forestry Sciences in 2021 (E4). All experiments adopted a randomized block arrangement experimental design with three replications, a field sowing hole spacing of 20 cm, a row spacing of 40 cm, and an area of 13.34 m in each plot. 2 (6.67m×2m), and each replicate interval was 20 materials, with a control variety of Kainong 69. The experimental fields were all sandy loam with medium fertility and convenient drainage and irrigation. During the growth period of peanuts, field management and harvesting were carried out in a timely manner. At harvest, the productivity of individual plants of 169 materials was investigated, and 10 plants were harvested from each plot. The productivity of individual plants was measured after drying; the oleic acid content of 169 materials was investigated after drying, and the oleic acid content of each sample was tested three times using a near-infrared analyzer Perton DA7250.
[0048] 1.1.2 Phenotypic data processing
[0049] Microsoft Excel 2010 was used to organize and calculate the phenotypic data and delete outliers. The mixed linear model of Genstat 18th Edition software was used to calculate the best linear unbiased estimate for the three replicate test data of two traits in each environment. The corrected phenotypic data of oleic acid content and single plant productivity were used for data histogram analysis. We found that Figure 1 As shown in the figure, the oleic acid content trait has two peaks, indicating that the trait is controlled by the major effect gene and is suitable for early generation selection assisted by molecular markers. The single plant productivity phenotypic data completely conforms to the normal distribution, which indirectly reflects that the yield trait is a quantitative trait controlled by multiple micro-effect loci. Combining GS for high-generation line selection would be a more reasonable breeding scheme.
[0050] 1.2 Genotype data
[0051] About 100 mg of young leaf tissue was collected at the seedling stage, and DNA from 169 samples was extracted using a plant genomic DNA kit. DNA integrity, quality, and concentration were assessed using agarose gel electrophoresis and NanoDrop. TM The 2000 sequencing platform performed 10× resequencing on 169 materials, performed quality control on clean data, used the peanut cultivar Kaixuan 016 (completed denovo assembly) as the reference genome for variant site detection (calling SNPs), retained high-quality SNPs, and filtered the following criteria: sequencing depth dp>=3, SNP site missing rate Miss<=0.2, minor allele frequency Maf>=0.05. A total of 608,809 SNPs were obtained on 20 chromosomes, with an average density of 242.56 / Mb ( Figure 2 ), the depth and breadth of the high quality of the sequencing samples have enabled the genotype data to reach a high level.
[0052] The method for sequencing and assembling the reference genome 016 is as follows:
[0053] (1) Third-generation technology: Use the PacBioSequel II platform for third-generation sequencing, and the sequencing depth must be no less than 100×.
[0054] (2) Second-generation Illumina data: Second-generation sequencing was performed using the Illumina nova-seq PE150 platform, with a sequencing depth of no less than 100×, Q20 ≥ 90%, and Q30 ≥ 85%.
[0055] (3) Hi-C data: Select four-base enzyme or six-base enzyme according to species information to construct Hi-C library, and the sequencing depth is required to be no less than 100×, Q20 ≥ 85%, and Q30 ≥ 80%.
[0056] (4) Transcriptome data: The second-generation transcriptome sequencing of Kaixuan 016 was completed using the Illumina nova-seq PE150 platform for genome-assisted annotation. The data volume was required to be no less than 6G / sample, with Q20 ≥ 85% and Q30 ≥ 80%.
[0057] Sequencing assembly results:
[0058] (1) The sequencing data of the third-generation sequencing of Kaixuan 016 was 297.92G, with a depth of 109.77×, and a total of 549.80G sequencing data combined with the second-generation sequencing.
[0059] (2) Survey analysis was performed using kmer17 software: the genome size was 2,703.87 Mbp, which was 2,686.33 Mbp after correction, the heterozygosity rate was 0.13%, and the repetitive sequence ratio was 84.15%.
[0060] (3) The peanut genome was sequenced and denovo assembled. The assembly results are as follows: the total length of the contig is 2.53 Gbp, and the length of the contig N50 is 11.48 Mbp; the total length of the scaffold is 2.53 Gbp, and the length of the scaffold N50 is 11.48 Mbp.
[0061] (4) Use Hi-C data to mount chromosomes to obtain chromosome-level genomes.
[0062] (5) The assembly quality was evaluated by consistency, sequence integrity, EST sequence, RNA sequence, CEGMA, and BUSCO.
[0063] The alignment rate of all small fragment reads to the genome was about 99.65%, and the coverage rate was about 99.80%, indicating that the reads and the assembled genome were very consistent; 99.2% of the 1614 orthologous single-copy genes were assembled into complete single-copy genes, indicating that the assembly results were relatively complete; 241 genes were assembled from 248 CEGs (Core Eukaryotic Genes), accounting for 97.18%, indicating that the assembly results were relatively complete.
[0064] 1.3 TagSNP mining
[0065] 1.3.1 Significant SNP site detection
[0066] The genotype and phenotype data of 149 materials were used for association analysis, and potential candidate SNPs were screened out based on the significance of the association (P-value). Figure 3 and Figure 4They are Manhattan plot and QQplot of oleic acid content trait, respectively. From the figure, we can see that the significant sites are mainly distributed on chromosomes 9 and 19. After statistics, a total of 32 significant SNP sites were obtained, and the phenotypic variance explained rate (PVE, Phenotypic variance explained) analysis was performed on the associated sites (Table 3). Table 3 shows that the PVE of the significant sites is between 17% and 26%. Since the significant sites are mainly concentrated on chromosomes 9 and 19, there is linkage between the sites. In order to simplify the number of sites and reduce the screening cost, we need to develop TagSNP to simplify the molecular marker process. The data of these sites were extracted from 169 samples, and the plink software was used to perform Block analysis, and 3 Blocks were obtained ( Figure 5 ), using 20 validation population samples, grouped stepwise regression analysis was performed on 32 loci in 3 blocks, and the best combination was selected for the molecular marker-assisted locus set. A representative TagSNP was extracted from each block to replace the block, and a total of 3 TagSNPs were obtained, namely chr9_114322963, chr9_113845844, and chr19_154509990.
[0067] Steps to develop TagSNP:
[0068] (1) Extract data from 32 sites in 169 samples and use plink software to perform block analysis to obtain 3 blocks ( Figure 5 ), extract 3 TagSNPs to replace these 3 Blocks.
[0069] *9_114195794 9_114241585 9_114255210 9_114261850 9_114322963 9_114323009 9_114325078 9_114378317;
[0070] *9_113815844 9_113858071 9_113858669 9_113874096 9_113898323 9_113899092 9_113918792 9_113940661 9_113951897 9_113962231 9_1139633339_113985040 9_114001925 9_114002898 9_114035070 9_114046758 9_114091718 9_114100766 9_114167704 9_114183334;
[0071] *19_153409182 19_153491111 19_153598939 19_154509990.
[0072] (2) Using the 20 candidate sample groups and combining the block division in the reference group, the 32 SNPs were grouped and stepwise regressed, and the three best SNPs were comprehensively selected as TagSNPs.
[0073] The three TagSNPs are: "9_114322963", "9_113845844", and "19_154509990".
[0074] Table 3. Significant SNP loci and PVE of oleic acid content
[0075] SNP CHROM POS REF ALT MLM MAF PVE 19_153409182 19 153409182 A G 2.02E-11 0.3265 0.262256832 9_114001925 9 114001925 A G 2.70E-10 0.3844 0.236387112 19_153491111 19 153491111 T C 6.41E-10 0.3241 0.227595113 19_153598939 19 153598939 C G 7.55E-10 0.3163 0.22592299 9_113898323 9 113898323 G T 1.15E-09 0.3836 0.221613614 9_113962231 9 113962231 C T 2.80E-09 0.3725 0.212384579 9_114325078 9 114325078 T C 3.40E-09 0.4828 0.210362653 9_114241585 9 114241585 G T 5.48E-09 0.3725 0.205372638 9_113985040 9 113985040 A G 5.94E-09 0.3784 0.204532012 9_113951897 9 113951897 C T 6.10E-09 0.3801 0.204251537 9_113815844 9 113855324 A G 6.55E-09 0.3767 0.203504213 9_113918792 9 113918792 T C 7.34E-09 0.3854 0.202307517 9_114100766 9 114100766 C T 7.35E-09 0.3767 0.202288706 9_114046758 9 114046758 T C 7.53E-09 0.3811 0.202043154 9_114195794 9 114195794 A G 8.49E-09 0.3885 0.200777679 9_113963333 9 113963333 T C 9.24E-09 0.3801 0.199883676 19_154509990 19 154509990 T C 1.20E-08 0.4088 0.197164479 9_114255210 9 114255210 A T 2.08E-08 0.3818 0.191288142 9_113858669 9 113858669 T G 2.25E-08 0.369 0.190442046 9_113874096 9 113874096 G A 2.35E-08 0.364 0.190000148 9_113940661 9 113940661 T A 2.45E-08 0.3741 0.189540539 9_114323009 9 114323009 G A 4.15E-08 0.387 0.183909625 9_114261850 9 114261850 G A 4.17E-08 0.3818 0.183854052 9_114091718 9 114091718 A C 4.34E-08 0.3664 0.183430486 9_114183334 9 114183334 G T 4.96E-08 0.3828 0.181993539 9_114002898 9 114002898 A G 5.31E-08 0.375 0.181253438 9_113858071 9 113858071 G A 5.55E-08 0.3741 0.180778339 9_114167704 9 114167704 C T 6.39E-08 0.3889 0.179265694 9_114035070 9 114035070 G A 6.43E-08 0.3699 0.179195597 9_114378317 9 114378317 C T 7.24E-08 0.3953 0.17791279 9_113899092 9 113899092 G T 7.50E-08 0.3733 0.177540011 9_114322963 9 114322963 A G 8.28E-08 0.4764 0.176466001
[0076] 1.3.2 TagSNP site accuracy verification
[0077] The three SNPs were coded as 0-1-2, with the major allele homozygous coded as 0, the heterozygous site coded as 1, and the minor allele homozygous coded as 2 as the X variable. The field data analysis corrected data value was used as the Y variable. The R language was used for regression analysis and the harmonic R was calculated. 2 It is 0.9571, with a high degree of explanation, indicating that the above three sites can be used as high oleic acid TagSNPs.
[0078] 1.4 Genome-wide selection model
[0079] Based on the single-plant productivity data under four environments, the best linear unbiased estimate was calculated, and the phenotypic and genotypic data of 169 samples were obtained. 169 genotypic data and 149 phenotypic data were used as reference groups. The phenotypic data of 20 genotypes in the genotypic data were deleted and used as candidate groups. The performance of phenotypic traits of the 20 candidate groups was predicted, and the accuracy was evaluated with the actual phenotypes to verify the GS model.
[0080] 1.4.1 Constructing the G matrix
[0081] For quality control of genotype data, loci with missing values greater than 10% and loci with minor allele frequency (MAF) less than 0.01 were deleted, leaving 581,275 loci. The G matrix was constructed using 169 genotype data, of which 149 individuals with phenotypic data were used as the reference group, and the other 20 (Table 2) individuals with phenotypic data set as missing were used as candidate groups.
[0082]
[0083] p iis the minor allele frequency at site i, Z is the design matrix of SNP markers, and Z' is the transposed matrix of Z.
[0084] 1.4.2 Calculation of Genomic Estimated Breeding Value (GEBV) using the GBLUP model
[0085] The G matrix constructed using genotype data was combined with phenotypic data and the restricted maximum likelihood method was used to iterate and calculate the genetic parameters and BLUP breeding value (GEBV) of the productivity of each variety (Table 4). Based on the evaluated genetic parameters, the heritability of the traits was calculated (Table 5).
[0086]
[0087] X is the matrix structure of fixed factors, Z is the matrix structure of random factors, Y is the matrix structure of observations, G -1 is the inverse matrix of kinship G, is the effect value of the fixed factor (BLUE), is the effect size of the random factor (BLUP), and k is the ratio of the residual variance component to the additive variance component.
[0088] Table 4. Breeding values of 169 peanut materials
[0089]
[0090]
[0091]
[0092] When GS modeling, not only the GEBV breeding value is considered, but also the additive variance component, the residual variance component and the heritability are also important indicators. Many studies have shown that the heritability of yield is around 0.1 to 0.3. The results of this GS analysis showed that the additive variance component was 2.6275, the residual variance component was 11.234, and the heritability was 0.1896. The heritability value was within the normal range, indicating that there was no abnormality in the GS data and model operation and the results were credible.
[0093] Table 5. GS data model variation analysis
[0094] Mutation Type Mutation results Vg 2.6275 Ve 11.234 <![CDATA[h 2 ]]> 0.1896
[0095] 1.4.3 GS precision improvement statistics
[0096] The measured data of 20 samples showed that the promotion rate of hybrid combinations selected by the GS model increased by 33.33%. Through this GS model, 83.3% of the promoted varieties can be retained when the test scale is reduced by 20%, 75.0% of the promoted varieties can be retained when the test scale is reduced by 30%, 66.7% of the promoted varieties can be retained when the test scale is reduced by 40%, and 66.7% of the promoted varieties can be retained when the test scale is reduced by 50%. Figure 6 And as shown in Table 6.
[0097] Table 6. GS promotion rate statistics
[0098] Total number of varieties Select Scale Includes the number of upgraded varieties Planting scale Retention rate of promoted varieties 20 100% 12 100% 100.0% 20 90% 11 90% 91.7% 20 80% 10 80% 83.3% 20 70% 9 70% 75.0% 20 60% 8 60% 66.7% 20 50% 8 50% 66.7% 20 40% 7 40% 58.3% 20 30% 5 30% 41.7% 20 20% 4 20% 33.3% 20 10% 3 10% 25.0% 20 5% 1 5% 8.3%
[0099] 1.4.4 Accuracy Evaluation
[0100] The correlation between the predicted breeding values (GEBV values) and the true values of the candidate group of 20 materials was compared to evaluate the accuracy of the GS model. Figure 7 As shown, the scatter points are basically distributed around both sides of the straight line, and the correlation coefficient between GEBV and the true value is 0.4120, and the result can be used normally.
[0101] 1.5 Breeding of high-oleic acid and high-yield peanut varieties based on GWAS and GS
[0102] According to the goal of high oleic acid and high yield peanut breeding, hybrid combinations are configured. New varieties are selected by pedigree method, and the developed TagSNP and GS are applied to the breeding process ( Figure 8 ). Generally, high-yield materials are used as female parents, high-oleic acid materials are used as male parents, and about 20 hybrid combinations are configured. After the true and false hybrids of F1 are identified, 50 true hybrids can be obtained for each combination, and F2 is obtained by self-pollination. All individual plants are sampled and extracted and DNA is retained. When harvesting, the individual plants with low yield (less than 20 full fruits) are removed, and the remaining individual plants are screened by oleic acid molecular markers for molecular marker-assisted selection, so as to retain the high-oleic acid (oleic acid content greater than 75%) individual plants with equivalent yield. The continuous self-pollination offspring population only needs to eliminate the individual plants with low yield (less than 20 full fruits) until all the F5 individual plants are sampled and extracted and their DNA is retained. When harvesting, the high-yield individual plants (more than 20 full fruits) are retained and the corresponding sample DNA is resequenced. The GS reference group is used to predict the individual plant productivity. The individual plants with higher individual plant productivity (in this embodiment, peanuts are planted in single seeds with a hole distance of 20 cm. Under such planting conditions, individual plant productivity greater than 40g is considered a high-yield individual plant with higher individual plant productivity) are planted in rows the next year, all are harvested and removed of weeds, and then directly participate in the yield test next year, without the need for multi-year multi-point yield tests.
[0103] The embodiments described above are only preferred embodiments of the present invention and are not intended to limit the scope of implementation of the present invention. Therefore, any equivalent changes or modifications made according to the structures, features and principles described in the patent scope of the present invention should be included in the scope of the patent application of the present invention.
Claims
1. A method for breeding high-oleic acid and high-yield peanuts based on whole genome association analysis and whole genome selection, characterized in that: The following steps are involved: Step 1: Use genome-wide association analysis to obtain loci controlling peanut oleic acid traits, develop TagSNPs, and verify them; Step 2: Use whole genome selection to construct a reference population; Step 2 includes: Step 2.1: Use the GBLUP model to calculate the genetic parameters and breeding values of the productivity of each material and construct a reference group; Step 2.2: Select peanut materials with resequencing data as the validation group to validate the reference group, with an accuracy of no less than 0.4; Step 3: Conduct low-generation detection and high-generation prediction on the hybrid offspring, and aggregate high oleic acid and high-yield traits; Step 3 includes: Step 3.1: Breeding of high-oleic acid and high-yield peanut varieties: configure hybrid combinations, select offspring using the pedigree method, sample individual plants at the F2 generation seedling stage, extract DNA for preservation, select individual plants with more than 20 full fruits at harvest, and use high-oleic acid TagSNP to perform molecular marker-assisted selection on individual plants with more than 20 full fruits to obtain multiple high-oleic acid individual plants with equivalent yields; Step 3.2: Eliminate the individual plants with less than 20 full fruits from the continuous self-pollination population until the F5 generation seedling stage, extract and preserve the individual plants, and retain the individual plants with more than 20 full fruits for resequencing at harvest. Use the reference population constructed by whole genome selection to predict the yield of the individual plants with more than 20 full fruits, and obtain high-oleic acid and high-yield individual plants; Step 3.3: After the F6 generation plants are harvested, they will participate in yield tests to breed new high-oleic acid and high-yield peanut varieties.
2. The breeding method for high-oleic acid and high-yield peanuts based on whole genome association analysis and whole genome selection according to claim 1, characterized in that: Step 1 specifically includes the following: Step 1.1: Select peanut materials to form a natural population, conduct multi-year multi-point field trials, harvest and dry them at maturity, examine the oleic acid content and single-plant productivity of each material, remove outliers, use mixed linear models to correct oleic acid content and single-plant productivity to obtain the best linear unbiased estimate of each as the corrected phenotypic value, and obtain the phenotypic data of the population material; Step 1.2: Use the second-generation resequencing technology to perform resequencing of each material in the population at a depth of no less than 10×, and perform polymorphic variation site detection to obtain the genotype data of the population material; Step 1.3: Use the oleic acid content and genotypic data in the population material to conduct genome-wide association analysis, explore the loci that control the oleic acid content of peanuts, and develop TagSNPs; Step 1.4: Select peanut materials with resequencing data as the validation group to verify TagSNP, with an accuracy of no less than 0.
9.
3. The breeding method for high-oleic acid and high-yield peanuts based on whole genome association analysis and whole genome selection according to claim 2, characterized in that: In step 1.1, there are more than 100 peanut materials in the natural population.
4. The breeding method for high-oleic acid and high-yield peanuts based on whole genome association analysis and whole genome selection according to claim 2, characterized in that: In step 1.4, more than 20 peanut materials with resequencing data were selected as the validation group.
5. The breeding method for high-oleic acid and high-yield peanuts based on whole genome association analysis and whole genome selection according to claim 1, characterized in that: In step 2.2, more than 20 peanut materials with resequencing data were selected as the validation group.
6. The breeding method for high-oleic acid and high-yield peanuts based on whole genome association analysis and whole genome selection according to claim 1, characterized in that: In step 3.1, the number of hybridization combinations configured is no less than 20.
7. The method for breeding high-oleic acid and high-yield peanuts based on whole genome association analysis and whole genome selection according to claim 1, characterized in that: In step 3.2, the resequencing depth was 10×.
Citation Information
Patent Citations
High-yield large-fruit peanut hybridization combination selection method based on whole genome selection
CN115443907A
PARMS molecular marker obviously associated with arachidonic acid and linoleic acid and application of PARMS molecular marker
CN116445653A
Molecular marker of QTL (Quantitative Trait Loci) qLINOL.17 related to high oil and high oleic acid characters of peanuts, primer combination and application
CN118086564A