A method for constructing an optimal whole-genome selection system for Brassica napus
By constructing the best genome selection system for cabbage-type rapeseed, the problem of insufficient breeding prediction accuracy in the existing technology is solved, the whole genome selection method is optimized, and the breeding efficiency and accuracy are improved.
Patent Information
- Application Number
- CN202410898932.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2044-07-05
AI Technical Summary
A genome-wide selection system suitable for cabbage-type rapeseed has not been established in the prior art, which has affected the prediction accuracy and efficiency of its breeding.
A genome selection system for cabbage-type rapeseed is constructed. Through the analysis of the genotype data of the two populations, a comprehensive comparison of different whole-gene selection models, marker density, population size and proportion, non-additive effects, trait-specific SNPs and harmful mutations was established to establish a database of harmful mutations that are not restricted by computing equipment, and the whole-genome selection method was optimized.
It improves the prediction accuracy of cabbage-type rapeseed breeding, provides a reference for the best genome selection system, and promotes the breeding practice of cabbage-type rapeseed.
Smart Images

Figure CN119007802B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of Brassica napus variety breeding and molecular breeding technology, and specifically to a method for constructing an optimal genome-wide selection system for Brassica napus. Background Art
[0002] With the rapid development of genomics, as of 2023, 3517 genomes of 1575 plants have been sequenced. Among them, the number of genomes sequenced from 2021 to 2023 reached 2373, and the number of re-sequencings is countless (Xie et al. 2024). Brassica napus, one of the most important oil crops in the world, had its genome sequenced in 2014 (Chalhoub et al. 2014). In recent years, many Brassica napus populations have been re-sequenced, accompanied by the investigation of different phenotypes. The main purpose is to conduct GWAS by combining phenotypic and genotypic data to identify candidate genes for target traits (Tang et al. 2021; Lu et al. 2019). However, the breeding of excellent varieties is still an important task in breeding. For some traits, such as the unsaturated fatty acid content in Brassica napus seeds, the measurement is relatively complex. Through genome-wide selection (GS), phenotypic information can be quickly obtained. At the same time, the genome of a rapeseed germplasm can almost represent all its phenotypic information. Previous studies have carried out genome-wide selection on flowering time and disease resistance, etc., and confirmed its availability in Brassica napus (Fikere et al. 2020; Hohenlohe et al. 2015). Studies in other species have shown that the prediction accuracy of genome-wide selection is affected by many factors (Liu et al. 2018; Spindel et al. 2015), including statistical models, marker numbers, population sizes, the ratio of training populations to breeding populations, non-additive effects, harmful mutations, etc. These influencing factors play different roles in different species and different populations. Currently, multi-factor evaluations of genome-wide selection have been carried out in many species, including rice, maize, and soybean. However, an optimal genome-wide selection system has not been established in Brassica napus yet. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for constructing an optimal genome-wide selection system for Brassica napus. Based on 10 phenotypic traits of two populations, the prediction accuracies under different models, marker numbers, marker types, training population selection, and the application of non-additive effects and harmful mutations are respectively statistically analyzed. The main purpose is to evaluate the prediction ability of genome-wide selection for 10 phenotypic traits of Brassica napus, evaluate the influence of different factors on the prediction accuracy, and comprehensively compare to provide suggestions for subsequent breeding practices of Brassica napus; to solve the problems raised in the background art.
[0004] To achieve the above object, the present invention provides the following technical solutions: A method for constructing an optimal whole-genome selection system for Brassica napus, comprising at least the following steps:
[0005] S1: Determine plant materials and phenotypic data. A total of two populations are used, which are respectively called the 60K population and the WGR population;
[0006] S2: Analyze the genotype data of the two populations;
[0007] S3: Practice different whole-genome selection models. Four common GS models are selected for comparison. The four common GS models include genomic best linear unbiased prediction (GBLUP) that directly estimates genomic estimated breeding value using the genomic relationship matrix, ridge regression best linear unbiased prediction (RR-BLUP) that further calculates the GEBV of individuals in the breeding population after estimating the marker effect values based on the training population, Bayesian B (BayesB), and semi-parametric method reproducing kernel Hilbert space (RKHS);
[0008] S4: Practice whole-genome selection with different marker densities;
[0009] S5: Practice whole-genome selection with different population sizes and proportions;
[0010] S6: Practice whole-genome selection including non-additive effects;
[0011] S7: Practice whole-genome selection applying trait-specific SNPs;
[0012] S8: Construct a corresponding database by designing a method for constructing a harmful mutation database that is not restricted by computing devices, and practice whole-genome selection applying harmful mutations.
[0013] Further, the S1 at least includes the following steps:
[0014] The 60K population and the WGR population respectively contain 520 and 604 germplasms;
[0015] The germplasms are planted. Each germplasm is planted in two rows, with 10 plants in each row and a plant spacing of 20 cm;
[0016] Corresponding to the 6 six seed quality traits of the 60K population, namely glucosinolate, oleic acid, linoleic acid, linolenic acid, protein content, and oil content, they are measured by near-infrared method in the planting year;
[0017] Record 3 growth period traits. The growth period traits include the initial flowering stage, the final flowering stage, and the flowering period. The lengths of the initial flowering stage, the final flowering stage, and the flowering period are respectively counted in three time periods, and the BLUE value is calculated using the R package "lme 4 " for GS analysis;
[0018] One yield trait, silique length, was measured in the WGR population. Silique length was measured in two years, and 10 siliques were investigated for each germplasm.
[0019] Furthermore, the S2 at least includes the following steps: For the SNP data of the two populations in S1, according to the minor allele frequency being greater than 0.05, the missing proportion of the markers being less than 0.2, and the missing proportion of the samples being less than 0.3, use the PLINK software to filter the markers. Finally, 31,725 and 385,853 high-quality markers are obtained in the 60K and WGR populations respectively, and use the VCFtools tool to convert the genotype format.
[0020] Furthermore, the S3 at least includes the following steps:
[0021] GBLUP, RKHS, and BayesB are completed through the R package "BWGS" in R 4.3.0, where the GBLUP model is repeated 100 times, and RKHS and BayesB are repeated 30 times, and the models are evaluated through five-fold cross-validation;
[0022] RR-BLUP is completed through the R package "rrBLUP", using four-fifths of the population size as the training population, and the remaining one-fifth as the breeding population, and the prediction accuracy is evaluated by repeating 100 times.
[0023] Furthermore, the S4 at least includes the following steps:
[0024] Due to the difference in the time required for different models, the GBLUP model is selected in the R package "BWGS" to calculate the prediction accuracy of 10 phenotypic traits of Brassica napus at different marker densities;
[0025] Use the parameters "geno.reduct.method = "RMR", reduct.marker.size" to randomly select 100, 500, 1000, 5000, and 10,000 markers for calculation, and repeat 100 times under five-fold cross-validation.
[0026] Furthermore, the S5 at least includes the following steps:
[0027] Set 6 different population sizes according to 604 germplasms of the WGR population, and set 5 different population sizes for 520 germplasms of the 60K population. Except for the maximum population sample numbers of the two populations, the remaining population sizes decrease by 100 samples in turn;
[0028] The prediction accuracies of 10 phenotypic traits in Brassica napus were compared using the RR-BLUP model at different population sizes. Four-fifths of the set population size was selected as the training population, and one-fifth of the set population size was selected from the remaining samples of the total population as the breeding population;
[0029] This process was completed using the R package "rrBLUP" and repeated 100 times;
[0030] Meanwhile, seven different ratios of the training population to the breeding population, namely 4:1, 3:1, 2:1, 1:1, 1:2, 1:3, and 1:4, were set and also completed using the RR-BLUP model, repeated 100 times.
[0031] Furthermore, step S6 at least includes the following steps:
[0032] Calculate the additive, dominant, and epistatic effects of the markers using the R package "sommer";
[0033] Meanwhile, implement GS including additive, additive + dominant, additive + dominant + epistatic effects using the RKHS model in the R package "BGLR". Take four-fifths of the total population as the training population and the remaining one-fifth as the breeding population, and repeat 100 times.
[0034] Furthermore, step S7 at least includes the following steps:
[0035] Based on the high-quality SNP loci of the two populations in S1, perform principal component analysis on the genotype data using the PLINK software and calculate the kinship between samples in Tassel 5;
[0036] Finally, perform GWAS using the mixed linear model in Tassel 5 according to the results of principal component analysis, the phenotypic values of 10 traits, the genotypes of samples, and the kinship between samples;
[0037] According to the GWAS results, sort the markers, extract markers with different trait thresholds based on P values less than 0.1, 0.01, and 0.001 using the PLINK software, and implement genome-wide selection of trait-specific SNPs in the R package "BWGS", repeated 100 times.
[0038] A method for constructing a harmful mutation database that is not restricted by computing devices, at least including the following steps:
[0039] To eliminate the requirements for the computing device hardware, the genome file, protein file, and protein annotation file of the Brassica napus 'Darmor-bzh' reference genome were split according to chromosome information;
[0040] Apply the SIFT software in the Ubuntu system to construct a harmful mutation database for each chromosome. After the construction of the harmful mutation databases for all chromosomes is completed, merge them;
[0041] According to the obtained Brassica napus harmful mutation database, annotate the SNP data of the 60K population and the WGR population using the SIFT software, and identify SNPs with scores less than 0.05 as harmful mutations;
[0042] Apply the PLINK software to extract SNPs with scores less than 0.05. After converting the SNP data format using the VCFtools tool, implement genome-wide selection of harmful mutations in Brassica napus using the GBLUP model in the R package "BWGS", and repeat 100 times.
[0043] A method for estimating the prediction accuracy of genome-wide selection of phenotypic traits in Brassica napus:
[0044] Apply GWAS based on the mixed linear model in common GWAS software to determine the P-value corresponding to a certain phenotypic trait of Brassica napus for each marker;
[0045] Sort the P-values through Excel, and at the same time extract the markers with P-values less than 0.01, and determine the number of these markers;
[0046] Divide the number of these markers by the total number of markers. When this value is greater than 0.019, the predictability of this trait is high.
[0047] Compared with the prior art, the beneficial effects of the present invention are:
[0048] Based on 10 traits of Brassica napus, the present invention comprehensively compares the prediction accuracies under different genome-wide selection models, marker densities, population sizes and ratios, as well as the applications of non-additive effects, trait-specific SNPs and harmful mutations, and determines an optimal genome-wide selection system for Brassica napus. In addition, the present invention establishes a method for constructing a harmful mutation database that is not restricted by computing devices and a method for estimating the prediction accuracy of genome-wide selection of phenotypic traits in Brassica napus. Therefore, the present invention can provide a reference for the establishment of an optimal genome-wide selection system for Brassica napus, and at the same time can promote the application of genome-wide selection in Brassica napus. Description of the Drawings
[0049] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0050] Figure 1 Prediction accuracy of 10 phenotypic traits of Brassica napus under 4 models;
[0051] Figure 2 Prediction accuracy of 10 phenotypic traits of Brassica napus under different marker densities;
[0052] Figure 3 Prediction accuracy of 10 phenotypic traits of Brassica napus under different population sizes;
[0053] Figure 4 Prediction accuracy of 10 phenotypic traits of Brassica napus under different population ratios;
[0054] Figure 5 Prediction accuracy of 10 phenotypic traits of Brassica napus when applying additive and non-additive effects;
[0055] Figure 6 Prediction accuracy of 10 phenotypic traits of Brassica napus when applying trait-specific SNPs and harmful mutations;
[0056] Figure 7 Correlation between the number of markers and prediction accuracy of 10 phenotypic traits of Brassica napus under different P values. Specific implementation manners
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.
[0058] Example 1:
[0059] A method for constructing an optimal whole-genome selection system for Brassica napus, at least including the following steps:
[0060] S1: Determine plant materials and phenotypic data. A total of two populations are used, which are respectively called the 60K population and the WGR population;
[0061] S2: Analyze the genotype data of the two populations;
[0062] S3: Practice different whole-genome selection models. Four common GS models are selected for comparison. The four common GS models include GBLUP that directly estimates genomic estimated breeding values using the genomic relationship matrix, RR-BLUP and BayesB that calculate the GEBV of individuals in the breeding population after estimating the effect values of markers based on the training population, and the semi-parametric method RKHS;
[0063] S4: Practice whole-genome selection with different marker densities;
[0064] S5: Genome-wide selection practices for different population sizes and proportions;
[0065] S6: Genome-wide selection practices including non-additive effects;
[0066] S7: Genome-wide selection practices applying trait-specific SNPs;
[0067] S8: Constructing a corresponding database by designing a method for constructing a harmful mutation database that is not restricted by computing devices, and genome-wide selection practices applying harmful mutations.
[0068] S1 includes at least the following steps:
[0069] The 60K population and the WGR population contain 520 and 604 germplasms respectively;
[0070] Plant the population germplasms. Each germplasm is planted in two rows, with 10 plants in each row, and the plant spacing is 20 cm. In this example, it is planted at the Chongqing Rapeseed Engineering and Technology Research Center (latitude 29°45′, longitude 106°22′, altitude 238.57 m);
[0071] For the 6 grain quality traits corresponding to the 60K population, namely glucosinolate (GSL), oleic acid (18:1), linoleic acid (18:2), linolenic acid (18:3), protein content (SPC), and oil content (SOC), they are measured by near-infrared method in the planting year (in this example, in 2015);
[0072] Record 3 growth period traits. The growth period traits include the initial flowering date (DIF), the final flowering date (DFF), and the flowering period (FP). The initial flowering date, the final flowering date, and the flowering period length are statistically analyzed in three time periods (in this example, 2017, 2018, 2019), and the BLUE value is calculated using the R package "lme 4 " for GS analysis;
[0073] One yield trait, silique length (SL), is measured in the WGR population. The silique length is measured in two years (in this example, 2015, 2017), and 10 siliques are investigated for each germplasm.
[0074] S2 includes at least the following steps: For the SNP data of the two populations in S1, according to the minor allele frequency being greater than 0.05, the missing proportion of the markers being less than 0.2, and the missing proportion of the samples being less than 0.3, use the PLINK software to filter the markers. Finally, 31,725 and 385,853 high-quality markers are obtained in the 60K and WGR populations respectively, and use the VCFtools tool to convert the genotype format.
[0075] S3 includes at least the following steps:
[0076] GBLUP, RKHS, and BayesB were completed using the R package "BWGS" in R 4.3.0. For the GBLUP model, it was repeated 100 times, and for RKHS and BayesB, they were repeated 30 times. The models were evaluated through five-fold cross-validation.
[0077] RR-BLUP was completed using the R package "rrBLUP". Four-fifths of the population size was used as the training population, and the remaining one-fifth was the breeding population. The prediction accuracy was evaluated by repeating 100 times.
[0078] S4 includes at least the following steps:
[0079] Due to the difference in time required for different models, the GBLUP model was selected in the R package "BWGS" to calculate the prediction accuracy of 10 phenotypic traits of Brassica napus at different marker densities.
[0080] Using the parameters "geno.reduct.method = "RMR", reduct.marker.size", 100, 500, 1000, 5000, and 10000 markers were randomly selected for calculation and repeated 100 times under five-fold cross-validation.
[0081] S5 includes at least the following steps:
[0082] Six different population sizes were set according to 604 germplasms of the WGR population, and five different population sizes were set for the 520 germplasms of the 60K population. Except for the maximum population sample sizes of the two populations, the remaining population sizes decreased by 100 samples in sequence.
[0083] The prediction accuracy of 10 phenotypic traits of Brassica napus at different population sizes was compared through the RR-BLUP model. Four-fifths of the set population size was selected as the training population, and one-fifth of the set population size was selected from the remaining samples of the total population as the breeding population.
[0084] This process was completed using the R package "rrBLUP" and repeated 100 times.
[0085] Meanwhile, seven different ratios of training population to breeding population, namely 4:1, 3:1, 2:1, 1:1, 1:2, 1:3, and 1:4, were set and also completed through the RR-BLUP model, repeated 100 times.
[0086] S6 includes at least the following steps:
[0087] The additive, dominant, and epistatic effects of the markers were calculated using the R package "sommer".
[0088] At the same time, the RKHS model in the R package "BGLR" was used to implement GS including additive, additive-dominant, additive-dominant and epistatic effects. Four-fifths of the total population was taken as the training population, and the remaining one-fifth was taken as the breeding population, and this was repeated 100 times.
[0089] S7 at least includes the following steps:
[0090] Based on the high-quality SNP loci of the two populations in S1, the genotype data was subjected to principal component analysis using the PLINK software, and the kinship between samples was calculated in Tassel 5.
[0091] Finally, a mixed linear model was applied in Tassel 5 for GWAS based on the results of principal component analysis, the phenotypic values of 10 traits, the genotypes of samples, and the kinship between samples.
[0092] According to the GWAS results, the markers were ranked, and markers with different trait thresholds were extracted by the PLINK software according to P values less than 0.1, 0.01, and 0.001. The genomic selection of trait-specific SNPs was implemented using the GBLUP model in the R package "BWGS", and this was repeated 100 times.
[0093] Example two:
[0094] Based on S8 in the above example, this example proposes a method for constructing a harmful mutation database that is not restricted by computing devices, which at least includes the following steps:
[0095] To eliminate the requirements for the computing device hardware, the genomic file, protein file, and protein annotation file of the Brassica napus 'Darmor-bzh' reference genome were split according to chromosome information.
[0096] The SIFT software was used in the Ubuntu system to construct a harmful mutation database for each chromosome. After the harmful mutation databases for all chromosomes were constructed, they were merged.
[0097] According to the obtained Brassica napus harmful mutation database, the SNP data of the 60K population and the WGR population were annotated using the SIFT software, and SNPs with scores less than 0.05 were identified as harmful mutations.
[0098] The PLINK software was used to extract SNPs with scores less than 0.05. After converting the SNP data format using the VCFtools tool, the genomic selection of Brassica napus harmful mutations was implemented using the GBLUP model in the R package "BWGS", and this was repeated 100 times.
[0099] Based on the above example, the following statistical analysis is proposed:
[0100] The differences in the GS prediction accuracy among different GS models, different marker densities, different population sizes and ratios, and the application of trait-specific SNPs and harmful mutations were compared by the Tukey method in one-way ANOVA, and P<0.05 was considered significantly different. The comparison between the GS prediction accuracy of non-additive effects and additive effects was conducted by two-tailed Student's T-test. All the above analyses were implemented in SPSS 24. The correlation analysis between the number of markers obtained at different P values and their GS prediction accuracy for 10 phenotypic traits of Brassica napus was achieved by Origin 2024.
[0101] In the present invention, the prediction accuracies of six quality traits of Brassica napus seeds, including glucosinolate (GSL), oleic acid (18:1), linoleic acid (18:2), linolenic acid (18:3), seed protein content (SPC), and seed oil content (SOC), three growth period traits, including the initial flowering date (DIF), the final flowering date (DFF), and the flowering period (FP), and one yield trait, silique length (SL), were determined under the GBLUP, RR-BLUP, BayesB, and RKHS models. The prediction accuracies of each trait under each model (Table 1), the optimal model for each trait ( Figure 1 ), and the comprehensive optimal model BayesB were provided.
[0102] Figure 1 Note: C18:1: oleic acid, C18:2: linoleic acid, C18:3: linolenic acid, GSL: glucosinolate, SOC: seed oil content, SPC: seed protein content, DFF: final flowering date, DIF: initial flowering date, FP: flowering period, SL: silique length. Multiple comparisons were performed by Tukey's test in one-way ANOVA, and different lowercase letters indicate significant differences among different models, P<0.05. The following Figures 2 - 7 are all the same.
[0103] The present invention provides the prediction accuracies of 10 traits of Brassica napus at five marker densities of 10,000, 5,000, 1,000, 500, and 100 SNPs ( Figure 2 ), and provides an optimal marker density suitable for genome-wide selection of Brassica napus, which is 5 kb SNPs.
[0104] Figure 2 Note: Five marker densities of 10,000, 5,000, 1,000, 500, and 100 were applied.
[0105] The present invention provides the prediction accuracies of 10 traits of Brassica napus at six population sizes of 604, 520 / 500, 400, 300, 200, and 100 samples ( Figure 3) and provided an optimal population size suitable for genome-wide selection in Brassica napus, which is 400 samples.
[0106] Figure 3 Note: For 6 phenotypic traits corresponding to the 60K population, 5 population sizes were set. For 4 phenotypic traits corresponding to the WGR population, 6 population sizes were set. Error bars indicate standard deviation.
[0107] The present invention provided the prediction accuracies of 10 traits in Brassica napus in the training population and breeding population at seven ratios of 4:1, 3:1, 2:1, 1:1, 1:2, 1:3, and 1:4 ( Figure 4 ) and provided an optimal ratio of the training population to the breeding population suitable for genome-wide selection in Brassica napus, which is 3:1.
[0108] Figure 4 Note: A total of 7 ratios of the training population to the breeding population were set, including 4:1, 3:1, 2:1, 1:1, 1:2, 1:3, and 1:4. Error bars indicate standard error.
[0109] The present invention provided the prediction accuracies of 10 traits in Brassica napus in the training population and breeding population under additive effects, additive effects and dominant effects, and additive effects, dominant effects and epistatic effects ( Figure 5 ) and determined that the application of non-additive effects varies with traits, which seems unnecessary for traits with high prediction accuracy and is effective for traits with low prediction accuracy.
[0110] Figure 5 Note: A: Additive, AD: Additive and dominant, ADE: Additive, dominant and epistatic. A two-tailed Student's T-test was used to analyze the difference in GS prediction accuracy with or without non-additive effects. *, P<0.05, **, P<0.01.
[0111] The present invention provided the prediction accuracies of 10 traits in Brassica napus under trait-specific SNPs with P values less than 0.001, 0.01, and 0.1 obtained by applying GWAS ( Figure 6 ) and provided an optimal selection of trait-specific SNPs with P values less than 0.1.
[0112] Figure 6 Note: BD: Harmful mutations in Brassica napus, BM5000: 5000 markers randomly selected in a population with 10 phenotypic values, M5000: 5000 random markers in the 60K and WGR populations. P<0.1, P<0.01, and P<0.001 respectively represent the markers filtered according to the P values less than 0.1, 0.01, and 0.001 obtained by GWAS.
[0113] The present invention provides the gains in the prediction accuracy of 10 traits of Brassica napus under trait-specific SNPs with P-values less than 0.001, 0.01, and 0.1 obtained by GWAS (Table 2). For traits with relatively low prediction accuracy, when GS is performed using SNPs filtered by P-values, the gain effect is more obvious. For example, for C182, C183, SPC, and FP, the average gain exceeds 90%. The P-values at which different traits reach the maximum gain are different. C181, C182, SPC, and DIF have the maximum gain when the P-value is 0.001; C183, GSL, SOC, and FP have the maximum gain when the P-value is 0.01; while DFF and SL have the maximum gain when the P-value is 0.1.
[0114] Example 3:
[0115] By comparing the correlation between the number of markers obtained for 10 phenotypic traits of Brassica napus at different P-values and their GS prediction accuracy, it was found that the number of markers with P-values less than 0.01 was significantly positively correlated with the GS prediction accuracy ( Figure 7 ).
[0116] Based on this, the present invention provides a method for estimating the prediction accuracy of genome-wide selection of phenotypic traits in Brassica napus. This method determines the P-values of each marker corresponding to a certain phenotypic trait of Brassica napus by applying GWAS based on the mixed linear model in common GWAS software, sorts the P-values through Excel, and simultaneously extracts the markers with P-values less than 0.01 to determine the number of such markers. Divide the number of such markers by the total number of markers. If this value is greater than 0.019, the predictability of this trait is high. In the attempt of the present invention, the prediction accuracies of 4 traits with the ratio of the number of markers with P-values less than 0.01 obtained by GWAS to the total number of markers greater than 0.019 are 0.612, 0.813, 0.762, and 0.456 respectively (Table 3), and the average value reaches 0.661, which is 29% higher than the average value of the prediction accuracies of 10 phenotypic traits.
[0117] Figure 7 Note: PA: Prediction accuracy, MN0.001, MN0.01, and MN0.1 respectively represent the number of markers filtered according to P-values less than 0.001, 0.01, and 0.1. The numbers in the circles represent the Pearson correlation coefficients between traits.
[0118] The present invention provides a method for constructing a harmful mutation database of Brassica napus, as well as a method for annotating SNP harmful mutation information in Brassica napus resequencing data. At the same time, the present invention provides the prediction accuracy of 10 traits of Brassica napus in applying harmful mutations for genome-wide selection ( Figure 6 ), and determines that the application of harmful mutations varies depending on the trait, and seems unnecessary for traits with high prediction accuracy, while is effective for traits with low prediction accuracy.
[0119] As an effective breeding strategy, genome-wide selection (GS) has been widely applied in plants. It is of great significance in the practical application of GS to evaluate the influencing factors of its prediction accuracy and determine the optimal combinations that can increase the prediction accuracy. In this invention, the prediction accuracies of 10 phenotypic traits in Brassica napus were compared under different GS models, marker densities, population sizes and ratios, as well as the application of non-additive effects, trait-specific SNPs, and harmful mutations. Finally, a model was provided. First, the overall prediction ability for a trait can be evaluated based on the number of markers with P values less than 0.01 obtained from GWAS. Further, a training population of 400 samples or three times the size of the breeding population can be selected, and GS can be carried out in the BayesB model using the markers with P values less than 0.1 obtained from GWAS. The application of non-additive effects and harmful mutations varies with traits and seems unnecessary for traits with high prediction accuracy, but is effective for traits with low prediction accuracy.
[0120] Table 1 Prediction accuracies of 10 phenotypic traits in Brassica napus under 4 GS models
[0121]
[0122]
[0123] Table 2 Improvements in prediction accuracy of 10 phenotypic traits in Brassica napus under different P value selections
[0124]
[0125] Table 3 Number of markers obtained for 10 phenotypic traits in Brassica napus under different P value selections
[0126]
[0127] Note: The left side shows the number of markers corresponding to the WGR population scaled down proportionally to the 60K population. PA represents the prediction accuracy. P<0.01 / All is the ratio of the number of markers with P values less than 0.01 obtained from GWAS to the total number of markers.
[0128] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be encompassed by the present invention. Any reference signs in the claims should not be regarded as limiting the claims involved.
Claims
1. A method for constructing an optimal whole-genome selection system for Brassica napus, characterized in that: At least include the following steps: S1: Determine plant materials and phenotypic data. A total of two populations are used, which are respectively called the 60K population and the WGR population; S2: Analyze the genotype data of the two populations; S3: Practice different genomic selection models. Four commonly used GS models are selected for comparison. The four commonly used GS models include genomic best linear unbiased prediction (GBLUP) that directly estimates genomic estimated breeding value using the genomic relationship matrix, ridge regression best linear unbiased prediction (RR-BLUP) that further calculates the GEBV of individuals in the breeding population after estimating the effect values of markers based on the training population, Bayesian B (BayesB), and semi-parametric method reproducing kernel Hilbert space (RKHS); S4: Practice genomic selection with different marker densities; S5: Practice genomic selection with different population sizes and proportions; S6: Practice genomic selection including non-additive effects; S7: Practice genomic selection applying trait-specific SNPs; S8: Construct a corresponding database by designing a method for constructing a harmful mutation database that is not restricted by computing devices, and practice genomic selection applying harmful mutations; The S1 at least includes the following steps: The 60K population and the WGR population respectively contain 520 and 604 germplasms; The germplasms are planted. Each germplasm is planted in two rows, with 10 plants in each row and a plant spacing of 20 cm; Corresponding to 6 six grain quality traits of the 60K population, namely glucosinolate, oleic acid, linoleic acid, linolenic acid, protein content, and oil content, they are measured by near-infrared method in the planting year; Record three growth stage traits, where the growth stage traits include the initial flowering stage, the final flowering stage, and the flowering period. The initial flowering stage, the final flowering stage, and the length of the flowering period are respectively statistically analyzed in three time periods, and the BLUE value is calculated using the R package "lme 4 ” for GS analysis; One yield trait, silique length, is measured in the WGR population. The silique length is measured in two years, and 10 siliques are investigated for each germplasm; The S2 at least includes the following steps: For the SNP data of the two populations in S1, according to the minor allele frequency being greater than 0.05, the missing proportion of markers being less than 0.2, and the missing proportion of samples being less than 0.3, use the PLINK software to filter the markers. Finally, 31,725 and 385,853 high-quality markers are obtained in the 60K and WGR populations respectively, and use the VCFtools tool to convert the genotype format; The S3 at least includes the following steps: GBLUP, RKHS, and BayesB are completed through the R package "BWGS" in R 4.3.
0. Among them, the GBLUP model is repeated 100 times, and RKHS and BayesB are repeated 30 times, and the model is evaluated through five-fold cross-validation; RR-BLUP is completed through the R package "rrBLUP". Four-fifths of the population size is used as the training population, and the remaining one-fifth is the breeding population, and the prediction accuracy is evaluated by repeating 100 times; The S4 at least includes the following steps: Due to the difference in time required for different models, the GBLUP model is selected in the R package "BWGS" to calculate the prediction accuracy of 10 phenotypic traits of Brassica napus at different marker densities; Randomly select 100, 500, 1000, 5000, and 10000 markers for calculation using the parameters "geno.reduct.method = "RMR", reduct.marker.size", and repeat 100 times under five-fold cross-validation; The S5 at least includes the following steps: Set 6 different population sizes according to 604 germplasms of the WGR population, and set 5 different population sizes for 520 germplasms of the 60K population. Except for the maximum population samples of the two populations, the remaining population sizes decrease by 100 samples in sequence; Compare the prediction accuracies of 10 phenotypic traits of Brassica napus under different population sizes through the RR-BLUP model. Select four-fifths of the set population size as the training population, and select one-fifth of the set population size from the remaining samples of the total population as the breeding population; This process is completed through the R package "rrBLUP" and repeated 100 times; Meanwhile, set 7 different ratios of the training population to the breeding population, namely 4:1, 3:1, 2:1, 1:1, 1:2, 1:3, 1:
4. Similarly, it is completed through the RR-BLUP model and repeated 100 times; The S6 at least includes the following steps: Calculate the additive, dominant, and epistatic effects of markers through the R package "sommer"; Meanwhile, apply the RKHS model in the R package "BGLR" to achieve GS including additive, additive plus dominant, additive plus dominant and epistatic effects. Take four-fifths of the total population as the training population, and the remaining one-fifth as the breeding population, and repeat 100 times; The S7 at least includes the following steps: Based on the high-quality SNP loci of the two populations in S1, use the PLINK software to perform principal component analysis on the genotype data, and calculate the kinship between samples in Tassel 5; Finally, perform GWAS using the mixed linear model in Tassel 5 according to the principal component analysis results, the phenotypic values of 10 traits, the genotypes of samples, and the kinship between samples; According to the GWAS results, sort the markers, and extract markers with different trait thresholds according to P values less than 0.1, 0.01, and 0.001 through the PLINK software. Apply the GBLUP model in the R package "BWGS" to achieve genome-wide selection of trait-specific SNPs, and repeat 100 times.
2. A method for constructing a harmful mutation database that is not restricted by computing devices, which is used for the method for constructing an optimal whole-genome selection system for Brassica napus described in claim 1 above, and is characterized in that: At least includes the following steps: To eliminate the requirements for the hardware of the computing device, the genome file, protein file, and protein annotation file of the Brassica napus 'Darmor-bzh' reference genome were split according to chromosome information; Apply the SIFT software in the Ubuntu system to construct a harmful mutation database for each chromosome. After the harmful mutation databases of all chromosomes are constructed, merge them; According to the obtained harmful mutation database of Brassica napus, annotate the SNP data of the 60K population and the WGR population using the SIFT software, and identify SNPs with scores less than 0.05 as harmful mutations; SNPs with scores less than 0.05 were extracted using PLINK software. After converting the SNP data format with the VCFtools tool, the GBLUP model was applied in the R package "BWGS" to achieve genome-wide selection of harmful mutations in Brassica napus, and this was repeated 100 times.
3. A method for estimating the prediction accuracy of genome-wide selection of phenotypic traits in Brassica napus, which is used for the construction method of an optimal genome-wide selection system for Brassica napus described in any one of the above claims 1, and is characterized in that: In common GWAS software, GWAS based on a mixed linear model was used to determine the P-value of each marker corresponding to a certain phenotypic trait of Brassica napus; The P-values were sorted using Excel, and at the same time, the markers with P-values less than 0.01 were extracted and their quantity was determined; The quantity of this marker was divided by the total quantity of markers. When this value is greater than 0.019, the predictability of this trait is high.
Citation Information
Patent Citations
Method for predicting hybrid vigor of brassica napus based on whole genome selection
CN116844641A