A method for predicting heterosis in Brassica napus based on whole genome selection

Through the whole genome selection method, we screened the significant points of inter-group differences and association analysis of Brassica napus, constructed a multivariate stepwise regression model, solved the difficult problems of parent selection and hybrid combination prediction of Brassica napus, and improved the efficiency and accuracy of hybrid vigor breeding.

CN116844641BActive Publication Date: 2025-10-03SOUTHWEST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202310630848.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-31
Publication Date
2025-10-03
Estimated Expiration
2043-05-31

AI Technical Summary

Technical Problem

Existing technologies are difficult to guide the selection of parents and prediction of hybrid combinations of Brassica napus in a simple and effective manner, which affects the efficiency of hybrid vigor breeding.

Method used

Based on the whole genome selection method, by obtaining parental materials, determining the phenotype and genotype of the traits, screening the significant points of inter-group differences and association analysis, constructing a multivariate stepwise regression model, evaluating the accuracy of the prediction model, and determining the optimal model.

Benefits of technology

It improves the accuracy and efficiency of hybrid vigor prediction of Brassica napus, simplifies the parent selection process, and promotes the breeding of high-yield rapeseed varieties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116844641B_ABST
    Figure CN116844641B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for predicting hybrid vigor of Brassica napus based on whole genome selection, the method comprising: obtaining parent materials, configuring a plurality of hybrid combinations according to the parent materials; determining the trait phenotypes of each parent material and its hybrid combination; performing genotyping to obtain genotype data; screening out significant points of intergroup difference and significant points of association analysis based on the trait phenotypes and the genotype data; constructing multiple prediction models based on the trait phenotypes, genotype data, significant points of intergroup difference and significant points of association analysis; performing multivariate linear analysis based on the significant points of intergroup difference using a stepwise regression method to obtain a multivariate stepwise regression model; and evaluating the prediction accuracy of the multiple prediction models and the multivariate stepwise regression model based on cross-validation to determine the optimal model for the trait. The present invention constructs a simple and effective prediction index system to guide parent selection and predict hybrid combinations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent agricultural data processing, and in particular to a method for predicting hybrid vigor of Brassica napus based on whole genome selection. Background Art

[0002] Brassica napus L., a herbaceous plant in the genus Brassica in the family Cruciferae, is a major oilseed and cash crop. Breeding high-yield varieties using heterosis is a breeding area of ​​great interest to breeders. In heterosis breeding, the number of hybrid combinations available far exceeds the number of combinations breeders can actually select. Therefore, developing a simple and effective prediction index system to guide parent selection and predict hybrid combinations has important theoretical and practical significance.

[0003] Heterosis refers to the phenomenon in which the offspring of two parental lines exhibit traits that exceed those of either parent (Shang Lianguang, 2017). Darwin discovered heterosis in 1876, and Shull described it in more detail in 1914 (Darwin, 1876; Shull, 1914). Heterosis is ubiquitous in plant breeding and widely used in agricultural production. It has become a primary means of increasing yields of major crops such as corn, rice, rapeseed, and sorghum, and is crucial to agricultural production, making a significant contribution to grain production in my country and around the world (Whaley, 1944; Shao et al., 2019).

[0004] The development of hybrid vigor-utilizing systems has been a major driver of crop yield increases and hybrid variety development, garnering significant attention from breeders and researchers. Heterotic vigor breeding is widely used in various crops, but research into the genetic mechanisms of heterosis lags far behind its application. Consequently, researchers have proposed various hypotheses to explain its genetic mechanisms, including the dominance hypothesis, the overdominance hypothesis, the epistasis hypothesis, the genetic balance hypothesis, the DNA methylation hypothesis, and the genome-cytoplasmic gene interaction model hypothesis.

[0005] Rapeseed (Brassica genus, Cruciferae) is a widely cultivated oilseed crop worldwide and the fastest-growing field crop in terms of hybrid vigor research and utilization. my country is a major producer of rapeseed, with the largest rapeseed planting area in the world. In 2019, the rapeseed planting area remained stable at 658 hectares. 2, with a total output of 13.48 million tons (Gan Guoyu, 2022). Within the Brassica genus, Brassica rapa (AA, 2n = 20), Brassica nigra (BB, 2n = 16), and Brassica oleracea (CC, 2n = 18) are the three diploid basic species. Three allotetraploids were formed through natural hybridization and doubling of these basic species: Brassica napus (AACC, 2n = 38), Brassica juncea (AABB, 2n = 36), and Brassica carinata (BBCC, 2n = 34). The genomic relationship between these three species is known as the "Yu-type triangle" (UN, 1935).

[0006] Brassica napus L. (AACC, 2n=38) is an allotetraploid derived from the natural hybridization of Brassica rapa (2n=20) and Brassica oleracea (2n=18) approximately 7,500 years ago. Since its introduction to my country from Europe and Japan in the 1930s, it has been widely cultivated in the Yangtze River Basin and has become the primary cultivated variety in my country (Liu Houli, 1993). The breeding and development of Brassica napus in my country was developed based on the introduction of rapeseed resources from abroad. After Chinese breeders introduced European Brassica napus from Japan, it was widely planted in Sichuan, the middle and lower reaches of the Yangtze River, and Yunnan and Guizhou provinces in 1953 due to its high yield and strong disease resistance. Since then, a large number of Brassica napus resources have been introduced, such as the spring mid-ripening and mid-late maturing Brassica napus varieties introduced from Canada, and the high-yielding and cold-resistant Brassica napus varieties introduced from Sweden, which have greatly enriched my country's germplasm resources (Tian Hongyun, 2016).

[0007] Brassica napus is a relatively young allotetraploid species with a short cultivation history and a narrow genetic base. Therefore, developing effective methods to create new hybrids is crucial (Becker et al., 1995). Brassica napus can be divided into three ecotypes: spring, semi-winter, and winter. Numerous studies have shown that heterosis is prevalent among Brassica napus varieties. By hybridizing these three ecotypes or conducting interspecific hybridization, superior exogenous genes can be introduced, thereby expanding the genetic base and increasing genetic diversity (Grant et al., 1985). Quijada et al. (2006) hybridized winter and spring rapeseed and used the F1 generation to map QTLs for traits such as seed yield, flowering time, and plant height. They found that gene introduction from winter rapeseed increased per-plant yield in spring rapeseed. Radoev et al. (2008) constructed a population of 250 DH lines by hybridizing a conventional rapeseed variety with a synthetic rapeseed, and used this DH line population to hybridize with conventional rapeseed parents. By examining their yield-related traits, they found that they had dominant and over-dominant effects on seed yield.

[0008] Breeding high-yielding Brassica napus varieties using hybrid vigor is a breeding direction closely followed by breeders. However, breeding high-yielding varieties often requires years of multi-site mating design and data analysis and statistics, so using simple and efficient methods to guide parent selection is crucial. Summary of the Invention

[0009] In view of the shortcomings of the prior art described above, the purpose of the present invention is to provide a method for predicting heterosis in Brassica napus based on whole genome selection, and to construct a simple and effective prediction index system to guide parent selection and predict hybrid combinations.

[0010] Other features and advantages of the present application will become apparent from the following detailed description, or may be learned in part by practice of the present application.

[0011] According to one aspect of the present application, a method for predicting heterosis of Brassica napus based on whole genome selection is provided, the method comprising:

[0012] Obtaining parental materials and configuring several hybrid combinations based on the parental materials;

[0013] Based on the parental materials and the hybrid combination, determining the phenotype of each parental material and the hybrid combination thereof;

[0014] Performing genotyping based on the parental materials to obtain genotyping data;

[0015] Based on the trait phenotype and the genotype data, screen out the sites with significant differences between groups and the sites with significant association analysis;

[0016] Constructing multiple prediction models based on the phenotypic and genotypic data, significant points of intergroup differences, and significant points of association analysis;

[0017] Based on the significant differences between the groups, a multivariate linear analysis was performed using the stepwise regression method to obtain a multivariate stepwise regression model;

[0018] Based on cross-validation, the prediction accuracy of multiple prediction models and the multivariate stepwise regression model was evaluated to determine the optimal model for the trait.

[0019] Furthermore, the trait phenotypes of each parent material and their hybrid combination include intermediate parent advantage, super parent advantage, correlation of various traits, combining ability and heritability.

[0020] Furthermore, significant points of association analysis were screened out by the following method:

[0021] Perform principal component analysis and kinship analysis on the filtered and filled genotype data to obtain the PCA matrix and K matrix;

[0022] The PCA matrix and K matrix were used as covariates to establish a mixed linear model of population phenotypes and genes for association analysis:

[0023] y=Xβ+Zμ+ε

[0024] Where y is the observed phenotypic value, β is the fixed effect vector including genetic markers and population structure, μ is the random additive genetic effect vector of individuals / lines, X and Z are the design matrices with unknown fixed and random effects, respectively, and ε is the residual effect. μ and ε obey and Normal distribution, K is the genomic kinship matrix, is the additive genetic variance, is the residual variance;

[0025] The correction factor multiple test is used to determine the significance level P of the point. The significance level P reflects the degree of association between the marker and the phenotypic variation. The smaller the P value, the higher the degree of association between the marker and the trait variation.

[0026] Set different significance thresholds to filter out significant points in association analysis.

[0027] Furthermore, the points with significant differences between the groups were screened out by the following method:

[0028] Calculating the types of marker types present in the hybrid combination based on the marker types of the parents, wherein the marker types of the parents include one or a combination of a reference genome type, a mutant type, and a heterozygous type;

[0029] The hybrid combinations are grouped according to marker types, the phenotypic mean differences among the groups are statistically analyzed, and the significant difference points among the groups are determined according to the phenotypic mean differences.

[0030] Furthermore, the hybrid combination has 6 marker types, namely M P1 , M P2 , M F1 , M F2 , M BC1 and M BC2 ,

[0031] Analyze the differences between groups according to the marker type to see whether there are calculable additive effects and dominant effects, and establish the inter-group (0, 1) coefficient matrix Ka, Kd and the inter-group standardized effect matrix A based on the significance of the differences. 6*1 and D 10*1 , the additive effect a and dominant effect d of the significant site are calculated according to the following formula:

[0032] a=Ka*A / ∑ka(1,i)

[0033] d=2Kd*D / ∑kd(1,i)

[0034] The intergroup differences that can be calculated for additive effects include 6 groups: (M F2 , M BC1 )、(M F2 , M BC2 )、(M F1 , M BC1 )、(M F1 , M BC2 )、(M BC1 , M BC2 ) and (M P1 , M P2 );

[0035] There are 10 groups with calculable differences between groups in terms of dominant effect: (M F1 , M F2 )、(M BC1 , M BC2 )、(M p1 , M F2 )、(M p1 , M F1 )、(M p1 , M BC1 )、(M p1 , M BC2 )、(M p2 , M F2 )、(M p2 , M F1 )、(M p2 , M BC1 ) and (Mp2 , M BC2 ).

[0036] Furthermore, based on the significant differences between the groups, a multivariate linear analysis was performed using a stepwise regression method to obtain a multivariate stepwise regression model, which specifically includes:

[0037] The effect matrix E is established based on the additive effect a and dominant effect d of the significant differences between the groups. n*2 =(a i , d i ), the marker type matrix M of the hybrid combination is calculated based on the marker types of the hybrid parents at the significant sites 6*n , calculated using the following formula:

[0038] Ef=E·K·M

[0039] The main diagonal elements of Ef are Ef (i,i) The labeling effect of n significant sites that make up a certain material;

[0040] The effect X of the marker loci related to the trait and the phenotypic value of the trait were used for stepwise regression analysis, and a multivariate stepwise regression model was established:

[0041] y=Xβ+ε

[0042] where y is the phenotypic observation, β is the fixed effect vector, X is the marker effect matrix, ε is the residual effect, and ε obeys normal distribution, is the residual variance.

[0043] Furthermore, the cross-validation-based evaluation of the prediction accuracy of multiple prediction models and the multivariate stepwise regression model to determine the optimal model for the trait specifically includes:

[0044] Multiple prediction models were evaluated using Pearson correlation coefficients between true phenotypic values ​​and estimated breeding values;

[0045] The multivariate stepwise regression model was evaluated by the coefficient of determination and mean absolute error.

[0046] According to one aspect of the present application, a device for predicting heterosis of Brassica napus based on whole genome selection is provided, the device comprising:

[0047] an acquisition module, configured to acquire parental materials and configure a plurality of hybrid combinations according to the parental materials;

[0048] a determination module configured to determine the trait phenotypes of each parent material and the hybrid combination thereof based on the parent material and the hybrid combination;

[0049] A typing module is configured to perform genotyping based on the parental material to obtain genotyping data;

[0050] A screening module is configured to screen out significant points of inter-group differences and significant points of association analysis based on the trait phenotype and the genotype data;

[0051] A construction module is configured to construct multiple prediction models based on the trait phenotype, genotype data, significant points of inter-group differences, and significant points of association analysis;

[0052] An analysis module is configured to perform a multivariate linear analysis based on the significant difference points between the groups using a stepwise regression method to obtain a multivariate stepwise regression model;

[0053] The evaluation module is configured to evaluate the prediction accuracy of multiple prediction models and the multivariate stepwise regression model based on cross-validation, and determine the optimal model for the trait.

[0054] Furthermore, the trait phenotypes of each parent material and their hybrid combination include intermediate parent advantage, super parent advantage, correlation of various traits, combining ability and heritability.

[0055] Furthermore, the screening module is further configured to:

[0056] Perform principal component analysis and kinship analysis on the filtered and filled genotype data to obtain the PCA matrix and K matrix;

[0057] The PCA matrix and K matrix were used as covariates to establish a mixed linear model of population phenotypes and genes for association analysis:

[0058] y=Xβ+Zμ+ε

[0059] Where y is the phenotypic observation, β is the fixed effect vector, including genetic markers and population structure, μ is the random additive genetic effect vector of individuals / lines, X and Z are the design matrices with unknown fixed and random effects, respectively, and ε is the residual effect. μ and ε obey μ~N(0, and Normal distribution, K is the genomic kinship matrix, is the additive genetic variance, is the residual variance;

[0060] The correction factor multiple test is used to determine the significance level P of the point. The significance level P reflects the degree of association between the marker and the phenotypic variation. The smaller the P value, the higher the degree of association between the marker and the trait variation.

[0061] Set different significance thresholds to filter out significant points in association analysis.

[0062] Furthermore, the screening module is further configured to:

[0063] Calculating the types of marker types present in the hybrid combination based on the marker types of the parents, wherein the marker types of the parents include one or a combination of a reference genome type, a mutant type, and a heterozygous type;

[0064] The hybrid combinations are grouped according to marker types, the phenotypic mean differences among the groups are statistically analyzed, and the significant difference points among the groups are determined according to the phenotypic mean differences.

[0065] Furthermore, the hybrid combination has 6 marker types, namely M P1 , M P2 , M F1 , M F2 , M BC1 and M BC2 , the screening module is further configured to

[0066] Analyze the differences between groups according to the marker type to see whether there are calculable additive effects and dominant effects, and establish the inter-group (0, 1) coefficient matrix Ka, Kd and the inter-group standardized effect matrix A based on the significance of the differences. 6*1 and D 10 *1. Calculate the additive effect a and dominant effect d of the significant locus according to the following formula:

[0067] a=Ka*A / ∑ka(1,i)

[0068] d=2Kd*D / ∑kd(1,i)

[0069] The intergroup differences that can be calculated for additive effects include 6 groups: (M F2 , M BC1 )、(M F2 , M BC2 )、(M F1 , M BC1 )、(M F1 , M BC2 )、(M BC1 , M BC2 ) and (M P1 , M P2 );

[0070] There are 10 groups with calculable differences between groups in terms of dominant effect: (M F1 , M F2 )、(M BC1 , M BC2 )、(M p1 , M F2 )、(M p1 , M F1 )、(M p1 , MBC1 )、(M p1 , M BC2 )、(M p2 , M F2 )、(M p2 , M F1 )、(M p2 , M BC1 ) and (M p2 , M BC2 ).

[0071] Furthermore, the analysis module is further configured to:

[0072] The effect matrix E is established based on the additive effect a and dominant effect d of the significant differences between the groups. n*2 =(a i , d i ), the marker type matrix M of the hybrid combination is calculated based on the marker types of the hybrid parents at the significant sites 6*n , calculated using the following formula:

[0073] Ef=E·K·M

[0074] The main diagonal elements of Ef are Ef (i,i) The labeling effect of n significant sites that make up a certain material;

[0075] The effect X of the marker loci related to the trait and the phenotypic value of the trait were used for stepwise regression analysis, and a multivariate stepwise regression model was established:

[0076] y=Xβ+ε

[0077] where y is the phenotypic observation, β is the fixed effect vector, X is the marker effect matrix, ε is the residual effect, and ε obeys normal distribution, is the residual variance.

[0078] Furthermore, the evaluation module is further configured to:

[0079] Multiple prediction models were evaluated using Pearson correlation coefficients between true phenotypic values ​​and estimated breeding values;

[0080] The multivariate stepwise regression model was evaluated by the coefficient of determination and mean absolute error.

[0081] According to one aspect of the present application, an electronic device is provided, comprising: a controller; and a memory for storing one or more programs, wherein when the one or more programs are executed by the controller, the controller implements the method described above.

[0082] According to one aspect of the present application, a non-transitory computer-readable storage medium storing instructions is provided. When the instructions are executed by a processor, the method described above is performed.

[0083] Compared with the prior art, the present invention has the following beneficial effects:

[0084] (1) Thirty-five representative parental materials were selected from the core germplasm resource bank of Brassica napus, and 306 hybrid combinations were configured. The phenotypes of each parent and their F1 hybrids were measured, and the phenotypic data were preliminarily analyzed.

[0085] (2) The genotypes of 35 parental materials were analyzed using the Brassica napus 60K SNP chip.

[0086] SNP data were preliminarily filtered, quality controlled, and filled.

[0087] (3) Use hybrid phenotypic and genotypic data to establish a genome-wide selection model and explore the factors that affect the accuracy of genome-wide prediction.

[0088] (4) A GBLUP model based on GWAS and intergroup differential selection and a multivariate stepwise regression model based on intergroup differential selection markers were established, and cross-validation was used to explore the prediction accuracy of the models.

[0089] (5) Evaluate the various prediction models established and select the best prediction model for the trait. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] Figure 1 The figure is a flow chart of a method for predicting heterosis of Brassica napus based on whole genome selection according to an embodiment of the present invention.

[0091] Figure 2 Schematic diagram of the principle of fold cross-validation according to an embodiment of the present invention, a is the entire experimental population; b is the test population; c is the correlation calculation between the true phenotypic value of the test population and the estimated breeding value.

[0092] Figure 3 It is a phenotypic frequency distribution histogram of each trait according to an embodiment of the present invention, wherein: TSW: thousand-grain weight, g; SS: number of grains per silique, grains; SY: yield per plant, g; PH: plant height, cm; MIL: effective length of the main sequence, cm; BN: number of effective branches at one time, pieces; BH: starting point of an effective branch at one time, cm; MIS: number of effective siliques in the main sequence, pieces; SL: silique length, cm.

[0093] Figure 4Schematic diagram comparing agronomic traits of hybrid F1 and parents according to an embodiment of the present invention, Note: TSW: 1000-grain weight, g; SS: number of grains per silique, grains; SY: yield per plant, g; PH: plant height, cm; MIL: effective length of main sequence, cm; BN: number of effective primary branches, pieces; BH: starting point of an effective primary branch, cm; MIS: number of effective primary siliques, pieces; SL: silique length, cm; P1: female parent; P2: male parent; ***: P < 0.001; **: P < 0.01; *: P < 0.05.

[0094] Figure 5 1 is a distribution diagram of the heterotic values ​​of the parent and the super-parent in the single plant yield according to an embodiment of the present invention.

[0095] Figure 6 This is a heat map of the correlation between individual trait phenotypes according to an embodiment of the present invention. Note: TSW: 1000-grain weight; SS: number of grains per silique; SY: yield per plant; PH: plant height; MIL: effective length of the main sequence; BN: number of effective branches at one time; BH: starting point of an effective branch at one time; MIS: number of effective siliques in the main sequence; SL: silique length; ***: P < 0.001; **: P < 0.01; *: P < 0.05.

[0096] Figure 7 These are the DNA test results of some parental materials according to an embodiment of the present invention.

[0097] Figure 8 4 is a SNP marker density distribution diagram according to an embodiment of the present invention.

[0098] Figure 9 Schematic diagram of the prediction accuracy of GBLUP and GBLUP_D according to an embodiment of the present invention. Note: TSW: thousand-grain weight; SS: number of grains per silique; SY: yield per plant; PH: plant height; MIL: effective length of main sequence; BN: number of effective branches at one time; BH: starting point of an effective branch at one time; MIS: number of effective siliques in the main sequence; SL: silique length.

[0099] Figure 10 The prediction accuracy change trend of the GBLUP model with different marker quantities for each trait according to the embodiments of the present invention. Note: ai: prediction accuracy change trend of the GBLUP model with different marker quantities for 1000-grain weight, number of grains per silique, yield per plant, plant height, effective length of main sequence, number of effective branches at one time, starting point of effective branch at one time, number of effective siliques in main sequence, and silique length.

[0100] Figure 11Schematic diagram of the effect of population size on prediction accuracy according to an embodiment of the present invention; Note: TSW: thousand-grain weight; SS: number of grains per silique; SY: yield per plant; PH: plant height; MIL: effective length of main sequence; BN: number of effective branches at one time; BH: starting point of an effective branch at one time; MIS: number of effective siliques in the main sequence; SL: silique length.

[0101] Figure 12 GWAS principal component analysis and kinship analysis according to an embodiment of the present invention, (a) PCA scree plot; (b) kinship matrix.

[0102] Figure 13 : This is a whole-genome association analysis diagram for single-plant yield according to an embodiment of the present invention. Note: a is the QQ plot under the single-plant yield GWAS analysis; b is the Manhattan plot under the single-plant yield GWAS analysis; in the Manhattan plot, the blue horizontal line represents the significance threshold of -log10(1 / 23425)=4.37, and the sites exceeding the threshold are significantly associated sites.

[0103] Figure 14 Schematic diagram of the prediction accuracy of the GWAS association marker model at different threshold levels according to an embodiment of the present invention. Note: TSW: thousand-grain weight; SS: number of grains per silique; SY: yield per plant; PH: plant height; MIL: effective length of the main sequence; BN: number of effective branches at one time; BH: starting point of an effective branch at one time; MIS: number of effective siliques in the main sequence; SL: silique length. DETAILED DESCRIPTION

[0104] The following examples are merely intended to better illustrate the present invention, but the present invention is not limited to the examples set forth herein. Therefore, non-essential improvements and adjustments to the embodiments made by those skilled in the art based on the above-described invention and applied to other embodiments are still within the scope of the present invention.

[0105] The present invention will now be further described with reference to the accompanying drawings.

[0106] The embodiment of the present invention provides a method for predicting heterosis of Brassica napus based on whole genome selection, such as Figure 1 As shown, the method includes:

[0107] Step 1: Obtain parental materials and configure several hybrid combinations based on the parental materials;

[0108] Step 2, based on the parental materials and the hybrid combination, determining the phenotype of each parental material and the hybrid combination;

[0109] Step 3, performing genotyping based on the parental materials to obtain genotyping data;

[0110] Step 4, based on the trait phenotype and the genotype data, screening out the sites with significant differences between the groups and the sites with significant association analysis;

[0111] Step 5, constructing multiple prediction models based on the trait phenotype, genotype data, significant points of inter-group differences, and significant points of association analysis;

[0112] Step 6: Based on the significant differences between the groups, a multivariate linear analysis is performed using a stepwise regression method to obtain a multivariate stepwise regression model;

[0113] Step 7: Based on cross-validation, evaluate the prediction accuracy of multiple prediction models and the multivariate stepwise regression model to determine the optimal model for the trait.

[0114] It should be noted that the order of steps listed above is only an example, and the method does not necessarily need to be implemented in the order of steps listed above during actual implementation.

[0115] The method mainly includes the following six invention points:

[0116] 1. Phenotypic Genetic Analysis of Rapeseed

[0117] Statistical and genetic analyses of the phenotypic data of the parents and F1 hybrids revealed that the phenotypic data for each trait followed a normal distribution, with coefficients of variation ranging from 4.61% to 23.06%. The F1 hybrids exhibited significant mid-parent heterosis (MPH) and over-parent heterosis (OPH), with narrow-sense and broad-sense heritabilities ranging from 0.156 to 0.667 and 0.312 to 0.684, respectively.

[0118] 2. Genome-wide selection prediction model

[0119] Nine hybrid traits were analyzed using different whole-genome selection models. Five-fold cross-validation was used to assess the model's predictive accuracy. Results showed no significant differences in the predictive accuracy of different models for the same trait. Significant differences were observed between the different trait prediction models, which were highly correlated with the trait's heritability, with a correlation coefficient of 0.988. The effects of marker density and population size on whole-genome selection (GS) prediction accuracy were investigated using the GBLUP model. Results showed that both marker density and population size influenced GS prediction accuracy, initially increasing and then stabilizing with increasing marker density and population size. For 1000-grain weight, number of grains per silique, yield per plant, and silique length, the number of markers required to achieve optimal prediction was within 1000. For primary branch origin and plant height, the number of markers required to achieve optimal prediction was within 5000. For all traits, except for primary branch number, GS prediction accuracy reached its peak when the population size reached 250.

[0120] 3. GBLUP model based on significant sites in association analysis

[0121] A GWAS analysis was conducted using 35 parents and 306 hybrids. GBLUP prediction models were developed by screening for significant association loci at six threshold levels. Compared to the GS model, the prediction accuracy of the GBLUP models improved to varying degrees, indicating that selecting for loci associated with traits can enhance model prediction. The optimal models for different traits had different marker screening thresholds, and the prediction correlation coefficients of the optimal models ranged from 0.391 to 0.737. Prediction accuracy followed the order of magnitude of the GS model and was also influenced by the heritability of the trait.

[0122] 4. GBLUP model based on intergroup differential selection markers

[0123] The research materials were grouped using marker type as a fixed factor, and 15,789 specific loci associated with nine traits were screened at a significance level of 0.0001. The largest number of specific loci, 9,190, was associated with 1,000-grain weight, while the smallest number, 155, was associated with the number of primary effective branches. Based on the order of absolute magnitude of the specific loci's effect, 13 marker number levels were set to establish a GBLUP prediction model. The model achieved optimal prediction performance with approximately 300 specific loci. Compared with the GBLUP model established using genome-wide markers, the GBLUP model established using specific loci for each trait showed varying degrees of improvement, ranging from 0.0005 to 0.1372.

[0124] 5. Multivariate linear model based on differential selection markers between groups

[0125] Multivariate linear models were constructed using stepwise regression analysis based on the marker effect values ​​at specific loci and the phenotypic values ​​of the traits. The coefficients of determination between the predicted and phenotypic values ​​for each trait ranged from 0.351 to 0.966. The models were tested using 30-fold five-fold cross-validation, with coefficients of determination ranging from 0.247 to 0.913. The number of marker loci involved in each prediction model ranged from 9 to 92, with the 1000-grain weight prediction model having the most markers and the primary effective branch number prediction model having the fewest. Eleven markers were common to all nine traits.

[0126] 6. Prediction Model Evaluation

[0127] Using 30 replicates of five-fold cross-validation, a comparative analysis of various prediction models in the study revealed that prediction models based on trait-related molecular markers had significantly higher correlation coefficients than genome-wide selection models. The multivariate linear model based on markers expressing intergroup differences in selection performed best for eight traits: 1000-grain weight, number of grains per silique, yield per plant, plant height, starting point of a primary effective branch, number of primary effective branches, effective length of the main sequence, and number of effective siliques in the main sequence. The GBLUP model based on GWAS-significant markers performed best for silique length.

[0128] Given that the overall process of the method is known, the following embodiment will combine specific experimental data, apply the method proposed in this embodiment based on the experimental data, and analyze the results to fully illustrate the feasibility and progress of the present invention.

[0129] 1) Test materials:

[0130] The research materials used in this example are data generated from materials previously prepared and planted by the applicant. Thirty-five parental materials of Brassica napus were selected from core Brassica napus germplasm resources domestically and internationally. The names, providing institutions, and numbers of the parental materials are shown in Table 1. In March 2016, 306 F1 hybrids were generated from these 35 parental materials according to the NCII (15×20) and by adding some hybrid combinations. These 306 F1 hybrids and the 35 parental materials were planted at the Xiema Experimental Planting Base of the Rapeseed Engineering Technology Research Center in Beibei District, Chongqing, during the 2016-2017 season.

[0131] Table 1. Test materials and sources

[0132] Parent number Material name Material provider P1 Sichuan Oil 20 Sichuan Academy of Agricultural Sciences P2 Hunan Oil No. 13 Hunan Agricultural University P3 WX10329 Hunan Agricultural University P4 Nca Foreign resources P5 Ningyou No. 1 Huazhong Agricultural University P6 09-P64-1 Huazhong Agricultural University P7 10-Chong29 Huazhong Agricultural University P8 Huashuang No. 5 Huazhong Agricultural University P9 Huashuang No. 4 Huazhong Agricultural University P10 A-972 Huazhong Agricultural University P11 Shed A963 Huazhong Agricultural University P12 Ningyou No. 18 Jiangsu Academy of Agricultural Sciences P13 Yangyou No. 6 Jiangsu Academy of Agricultural Sciences P14 Yangyou No. 5 Jiangsu Academy of Agricultural Sciences P15 Zhejiang Oil 21 Zhejiang Academy of Agricultural Sciences P16 Wanyou 29 Anhui Academy of Agricultural Sciences P17 D3 Qinghai University P18 GY270 Shaanxi Hybrid Rapeseed Research Center P19 03LF1 Gansu Agricultural University P20 Nakaee Chousen Foreign resources P21 Youyan No. 2 Guizhou Rapeseed Research Institute P22 CPC 589 Oil Crops Research Institute, Chinese Academy of Agricultural Sciences P23 Zhejiang Double No. 8 Oil Crops Research Institute, Chinese Academy of Agricultural Sciences P24 Double 11 Oil Crops Research Institute, Chinese Academy of Agricultural Sciences P25 CPC 821 Oil Crops Research Institute, Chinese Academy of Agricultural Sciences P26 Shi Lifeng Foreign resources P27 Double No. 9 Oil Crops Research Institute, Chinese Academy of Agricultural Sciences P28 Cubs root Foreign resources P29 Victory rapeseed Foreign resources P30 cresor Foreign resources P31 Spoon Leaf Green Jiangsu Academy of Agricultural Sciences P32 Ningyou No. 10 Jiangsu Academy of Agricultural Sciences P33 Duoyou No. 1 Jiangsu Academy of Agricultural Sciences P34 Guangde 8104 Jiangsu Academy of Agricultural Sciences P35 Chuyou No. 1 Jiangsu Academy of Agricultural Sciences

[0133] 2) Planting and phenotypic identification of materials:

[0134] The experimental materials were planted in September 2016 using a randomized block design with two replications and three rows, with a row spacing of 0.4 m, a row length of 2.4 m, and a plant spacing of 0.24 m, with 10 plants per row. Field management followed conventional production practices. At maturity, six plants were randomly sampled from the middle row for testing, and the average phenotypic value was used for post-sequential correlation analysis. Nine traits were evaluated: thousand seed weight (TSW), seeds per siliques (SS), seeds yield per plant (SY), plant height (PH), effective length of main inflorescence (MIL), branch number (BN), branch height (BH), effective siliques of main inflorescence (MIS), and silique length (SL).

[0135] 3) Phenotypic data analysis of materials

[0136] 3.1) Phenotypic data statistics

[0137] After testing 35 parents and 306 hybrid F1, they were preliminarily sorted using Excel software to calculate the average value of each material and the average value of the block group, and basic statistical quantities such as standard deviation, kurtosis, skewness, and coefficient of variation were calculated. The data were visualized using the R program ggplot2 package (Version: 3.35).

[0138] 3.2) Calculation of heterosis-related parameters

[0139] Mid-parent heterosis (MPH) refers to the ratio of the yield or a quantitative trait of the F1 hybrid to the average value of the same trait of the parents (P1 and P2). Its formula is:

[0140]

[0141] Over-parent heterosis (OPH) refers to the ratio of the yield or numerical value of a quantitative trait of the F1 hybrid to the difference of the same trait of the high-parent (HP). Its formula is:

[0142]

[0143] 3.3) Correlation analysis of various traits

[0144] The Pearson correlation coefficients of the nine phenotypic traits of 35 parents and 306 F1 hybrids were calculated using the cor function in the R package, and the correlation analysis results were visualized using the corrplot package (Version: 0.84).

[0145] 3.4) Calculation of combining ability and heritability

[0146] When analyzing the combining ability and heritability of traits, a mixed linear model was established using the Sommer package (Version: 4.0.5) of the R (Version: 4.0.4) computer software. The model formula is as follows:

[0147] y=u+GCA+SCA+e

[0148] Where y is the phenotypic value in the environment, u is the population mean, GCA is the general combining ability of the parents, SCA is the specific combining ability of the parents, and e is the residual effect. In this model, except for the trait itself, which is considered a fixed effect, all other effects are considered random effects and are assumed to follow a normal distribution.

[0149] The broad heritability of the trait (H 2 ) refers to the percentage of genetic variance to phenotypic variance. The formula for calculating broad-sense heritability is as follows:

[0150]

[0151] The narrow sense heritability of the trait (h 2 ) refers to the percentage of additive effect variance to phenotypic variance. The formula for calculating narrow-sense heritability is as follows:

[0152]

[0153] in, is the general combining ability variance, i.e., the additive effect variance; is the variance of specific combining ability; is the error variance.

[0154] 4) Genotype identification of materials

[0155] At the rapeseed seedling stage, five plants from each parent were selected and ten discs of uniform size were punched out using a microporator. These discs, totaling 10, were then stored at -20°C until further use. Genomic DNA was extracted using the EasyPure Plant Genetics DNA kit (see the kit instructions for detailed procedures). DNA samples were electrophoresed on 1% agarose gels using a Biometra electrophoresis instrument and electrophoresis tank to ensure genomic DNA integrity and concentration. Genomic DNA was also analyzed using a Thermo Scientific NanoDrop 2000 spectrophotometer. Samples suitable for genotyping must have an A260 / 230 ratio between 1.8 and 2.2, an A260 / 280 ratio between 1.8 and 2.0, an average concentration of at least 50 ng / μl, and a total volume of at least 200 ng. DNA that met these requirements was fragmented and denatured, then hybridized to a microarray, coated, and imaged. Genotyping was performed using high-resolution images obtained using Illumina's Genome Studio software.

[0156] 5) Genome-wide selection prediction model

[0157] 5.1) Prediction Model

[0158] 5.1.1) GBLUP and GBLUP_D

[0159] GBLUP uses molecular markers to construct a matrix of individual kinship relationships for optimal linear unbiased estimation. The models used in this experiment are the additive effect GBLUP model and the additive dominant effect GBLUP model (GBLUP_D). The mixed linear model formula represented by the GBLUP model is as follows:

[0160] y=Xβ+Za+ε

[0161] Where y is the phenotypic value, which is the average of two replicates of the material in this experiment. β is the fixed effect vector, X is the matrix corresponding to all fixed effects; a is the additive effect vector and obeys Normal distribution, A = ZZ′ / 2∑p i (1-p i ), p i is the minimum allele frequency of the ith site; Z is the matrix corresponding to the random effect, ε is the random error effect vector, and obeys normal distribution.

[0162] The mixed linear model formula represented by the GBLUP_D model used in this experiment is as follows:

[0163] y=Xβ+Za+Zd+ε

[0164] Among them, a and d are additive effect vector and dominant effect vector respectively, and obey Normal distribution and Normal distribution. Where A=ZZ′ / 2∑p i (1-p i ), D = ZZ′ / ∑2p i q i (1-p i q i ), p i and q i is the frequency of the ith allele (Su et al., 2012; et al.; 2014).

[0165] The GBLUP model was constructed using the “mmer” function in the sommer package (Version: 4.1.3) of the R language programming software. The G matrix was constructed using the “A.mat” function, and the SNP genotyping was coded as -1, 0, and 1, that is, the low-frequency genotype was coded as -1, the heterozygous genotype was coded as 0, and the high-frequency genotype was coded as 1 (Wang Qi, 2020). The D matrix was constructed using the “D.mat” function, and the SNP genotyping was coded as 0 and 1, that is, the low-frequency genotype was coded as 0, and the heterozygous genotype was coded as 1.

[0166] 5.1.2) RRBLUP model

[0167] The RRBLUP model is a commonly used model in the indirect method, which can estimate the effects of all SNP markers. This model assumes that the marker effects are random effects and follow a normal distribution. It uses a mixed linear model to estimate the effect value of each marker, and then sums the effects of each marker to obtain the individual estimated breeding value (Bao Jingjing, 2020). The formula is as follows:

[0168]

[0169] Where y is the observed value, b is the fixed effect vector, X is the correlation matrix related to b, m is the number of markers for model establishment, and Z i is the genotyping code of the i-th SNP (-1, 0, 1), corresponding to the three marker types of the SNP (0 / 0, 1 / 0, 1 / 1); g i is the effect value of the i-th site; e is the random error effect, which obeys the normal distribution is the random genetic variance, and I is the identity matrix. The RRBLUP model established in this experiment was constructed using the mixd.solve function in the rrBLUP package (Version: 4.6.1).

[0170] 5.1.3) Bayesian Model

[0171] The basic formula of the Bayesian model is the same as that of RRBLUP. The Bayesian method assumes that the SNP effect and its variance follow a certain prior distribution. Different Bayesian models have different prior distribution assumptions. Among them, Bayes A assumes that all markers have an effect and g i Normal distribution Follows the scaled inverse chi-square distribution x -2 (v, S); Bayes B assumes that most markers have no effect, and that only a small number of markers have a large effect. i Normal distribution and Follows the scaled inverse chi-square distribution x -2 (v, S). The Bayes C method treats π as an unknown parameter, assuming it follows a uniform distribution U(0, 1), and assumes that the effect variances of effective SNPs are different. Bayesian LASSO (BL) assumes that the effect variances of marker loci follow a normal distribution with a double exponential distribution, also known as the Laplace distribution.

[0172] The GBLUP model was constructed using the BGLR function in the BGLR package (Version: 1.0.8) of the R language programming software, where the nIter parameter was set to 15000 and the burn in parameter was set to 5000.

[0173] 5.2) Prediction Accuracy

[0174] In genome-wide selection (GS) studies, the primary evaluation metrics for whole-genome selection are prediction accuracy and predictive power. Prediction accuracy is the Pearson correlation coefficient between the true breeding value (TBV) and the estimated breeding value (GEBV), or r(TBV:GEBV); while predictive power is the Pearson correlation coefficient between the actual phenotypic value and the predicted breeding value. Because TBV is typically unknown, prediction accuracy in this experiment was assessed using the Pearson correlation coefficient between the actual phenotypic value and the predicted breeding value.

[0175] 5.3) Research on factors affecting prediction accuracy

[0176] In order to explore the effects of marker density, population size, and heritability on the prediction accuracy of the GBLUP model, this experiment used the control variable method. When exploring the effect of marker density on the prediction accuracy of the GBLUP model, this experiment randomly selected 100, 200, 300, 400, 500, 750, 1000, 1500, 2000, 5000, 10000, 20000, and 23425 markers from 23425 SNP markers to reconstruct the G matrix for modeling, and used 30 five-fold cross-validations to evaluate the accuracy of the prediction model; when exploring the effect of population size on the prediction accuracy of the GBLUP model, this experiment randomly selected 150, 200, 250, 300, and 306 individuals from 306 hybrids to reconstruct the G matrix for modeling, and used 30 five-fold cross-validations to evaluate the accuracy of the prediction model; when exploring the effect of heritability on the prediction accuracy of the model, the average of the prediction accuracy of 30 five-fold cross-validations of the seven models was used to perform Pearson correlation analysis with the broad-sense heritability of the phenotypic data.

[0177] 5.4) Analysis of variance of prediction accuracy

[0178] In order to test whether there are significant differences among different influencing factors (model, marker density), ANOVA was performed on the prediction accuracy of different categories, and the formula is as follows:

[0179] y=μ+M i +ε

[0180] Where y is the prediction accuracy of 30 cross-validations, μ is the overall mean, and M i is the different treatments, and ε is the error.

[0181] 6) Cross-validation

[0182] In order to select a suitable model, the cross-validation (CV) method is usually used to evaluate the model. Figure 2 As shown, cross validation is to divide the data set into two parts, use one part of the data for training and building a prediction model, and the other part of the data for testing to evaluate the model. This experiment used 5-fold cross validation, and each model was cross-validated 30 times. The cross validation of this experiment was completed based on the caret package of R (R version 4.0.4) software. For the whole genome selection prediction model, the model is mainly evaluated by the Pearson correlation coefficient (r) between the true phenotypic value and the estimated breeding value. For the stepwise regression multivariate linear model, the determination coefficient (R 2 ) and mean absolute error (MAE) to evaluate the model.

[0183] 7) Molecular marker screening

[0184] 7.1) Screening of molecular marker loci based on whole-genome association analysis

[0185] This example uses TASSEL (Trait Analysis by Association, Evolution and Linkage) software (Version: 5.2.78) to perform association analysis on 340 experimental materials such as parents and hybrid F1. TASSEL has two built-in models commonly used for association analysis, GLM and MLM. In this experiment, the MLM in the software was used for analysis. First, principal component analysis (PCA) and kinship analysis were performed on the filtered and filled genotype data. Principal component analysis is based on genomic similarity to stratify the population (Zhao, 2018); kinship analysis was performed using the "scaled IBS" method in the software. The obtained PCA matrix and K matrix were used as covariates to establish a mixed linear model of population phenotype and genotype for association analysis. The formula is as follows:

[0186] y=Xβ+Zμ+ε

[0187] Where y is the observed phenotypic value, which in this study is the average of two replicates of the material, β is the fixed effect vector, including genetic markers and population structure (Q), μ is the random additive genetic effect vector of individuals / lines, X and Z are the design matrices with unknown fixed and random effects, respectively, and ε is the residual effect. μ and ε obey and Normal distribution, K is the genomic kinship matrix, is the additive genetic variance, is the residual variance.

[0188] The Bonferroni multiple test was used to determine the significance level (P) of the SNP, which reflects the degree of association between the marker and the phenotypic variation. The smaller the P value, the higher the degree of association between the marker and the phenotypic variation of the trait. In this study, different significance thresholds were set to discover significantly associated SNP sites for predictive analysis.

[0189] 7.2) Screening of molecular marker loci based on variance analysis

[0190] The SNP marker types of the parents are digitally recorded as 0 / 0 (reference genome type), 1 / 1 (mutant type), and 1 / 0 (heterozygous type). Based on the three marker types of the parents, it is inferred that the hybrid F1 may have 6 marker types, namely 1 / 1 type, that is, the two parental markers at this site are homozygous 1 / 1, named M P1 ; 1 / 0 type, that is, the two parental markers at this site are homozygous 1 / 1 and 0 / 0, named MF1 ; 0 / 0 type, that is, the two parents are homozygous 0 / 0 at this site, named M P2 ; 1 / 1, 1 / 0 and 0 / 0 mixed type, that is, both parental markers at this site are heterozygous 1 / 0, named M F2 ; 1 / 1 and 1 / 0 mixed type, that is, the two parental markers at this site are homozygous 1 / 1 and heterozygous 1 / 0, named M BC1 ; 0 / 0 and 1 / 0 mixed type, that is, the two parental markers at this site are homozygous 0 / 0 and heterozygous 1 / 0, named M BC2 ; According to the six marker types of hybrids M P1 、M P2 、M F1 、M F2 、M BC1 and M BC2 The loci were grouped and statistically analyzed for differences in phenotypic means between groups. If there was a significant difference in the means between groups at a certain significance level, the locus was considered a significant trait association locus.

[0191] According to the principles of genetics, the differences between groups were analyzed by marker type to determine whether there were calculable additive effects and dominant effects, and based on the significance of the differences, the inter-group (0, 1) coefficient matrices Ka and Kd and the inter-group standardized effect matrix A were established. 6*1 and D 10*1 The additive effect (a) and dominant effect (d) of the significant loci were calculated according to the following formula:

[0192] a=Ka*A / ∑ka(1,i)

[0193] d=2Kd*D / ∑kd(1,i)

[0194] The intergroup differences that can be calculated for additive effects include 6 groups: (M F2 , M BC1 )、(M F2 , M BC2 )、(M F1 , M BC1 )、(M F1 , M BC2 )、(M BC1 , M BC2 ) and (M P1 , M P2 ) were similar. The differences between groups with calculable dominant effects were 10 groups (M F1 , M F2 )、(M BC1 , M BC2 )、(M p1 , M F2 )、(M p1 , M F1 )、(Mp1 , M BC1 )、(M p1 , M BC2 )、(M p2 , M F2 )、(M p2 , M F1 )、(M p2 , M BC1 ) and (M p2 , M BC2 ).

[0195] 8) Multivariate stepwise regression prediction model

[0196] In this example, there are only six possible marker types M for all sites. P1 , M P2 , M F1 , M F2 , M BC1 and M BC2 , same as 7.2), the additive and dominant effects of different marker types are fixed, thus establishing the coefficient matrix K 2*6 . According to the previously estimated additive effect (a) and dominant effect (d) of the significant marker sites, the effect matrix E is established. n*2 =(a i , d i ), the marker matrix M of hybrid F1 is calculated based on the marker types of the hybrid parents at the significant sites 6*n , calculated using the following formula:

[0197] Ef=E·K·M

[0198] Then the main diagonal element Ef of Ef (i,i) The labeling effect of n significant sites that make up a certain material.

[0199] The effect X of the marker loci associated with the trait and the phenotypic value of the trait were used for stepwise regression analysis, and a multivariate linear prediction model was established, the formula of which is as follows:

[0200] y=Xβ+ε

[0201] Where y is the phenotypic observation, which is the average of two replicates in this study, β is the fixed effect vector, X is the marker effect matrix, and ε is the residual effect. normal distribution, is the residual variance.

[0202] Finally, the experimental results and analysis are as follows:

[0203] 1. Phenotypic data analysis

[0204] 1.1 Descriptive statistical analysis of traits of parents and hybrids

[0205] The phenotypic mean values ​​of each trait data of 35 parents and 306 hybrid F1 were preliminarily analyzed. The phenotypic data analysis results and frequency distribution histograms of each trait are shown in Table 2 and Figure 3 As shown, the phenotypic values ​​of all traits exhibited a normal distribution. The absolute values ​​of the skewness coefficients for each trait ranged from 0.04 to 0.79, with the smallest absolute value for primary ordinal length and the largest for silique length. The absolute values ​​of the kurtosis coefficients ranged from 0.00 to 1.94, with the smallest absolute value for number of grains per silique and the largest absolute value for silique length. Yield per plant had the highest coefficient of variation, at 24.94%, while plant height had the lowest coefficient of variation, at 5.19%. The order of coefficients of variation for the nine traits was plant height < primary ordinal length < number of primary effective branches < number of primary effective siliques < starting point of primary effective branches < number of grains per silique < silique length < 1000-grain weight < yield per plant.

[0206] Table 2 Descriptive statistical analysis of phenotypic data of each trait

[0207]

[0208] Note: 1000-grain weight, g; number of grains per silique, grains; yield per plant, g; plant height, cm; effective length of main sequence, cm; number of effective primary branches, pieces; starting point of effective primary branch of BH, cm; number of effective siliques in main sequence, pieces; silique length, cm.

[0209] 1.2 Analysis of heterosis for various hybrid traits

[0210] Combined with Table 3, the comparative analysis of the phenotypic means of the nine traits of the hybrid and the parents revealed that the average phenotypic values ​​of the nine traits of the hybrid F1 were significantly higher than those of the parents, and the distribution range was also larger than that of the parents ( Figure 4 ). The phenotypic data of 306 hybrid F1 were subjected to hybrid vigor analysis. The average values ​​of neutral parent advantage for all hybrid traits were positive, ranging from 3.76% to 44.74%. Among them, the neutral parent advantage for single plant yield was the largest, and the neutral parent advantage for 1000-grain weight was the smallest. Except for 1000-grain weight, the starting point of the first effective branch, and the number of first effective branches, the super-parent advantage for the other 6 traits was also positive. The super-parent advantage for all traits ranged from -4.67% to 26.92%, among which the super-parent advantage for single plant yield was the largest, and the super-parent advantage for 1000-grain weight was the smallest. Single plant yield showed obvious neutral parent advantage and super-parent advantage, with more than 88% of the combinations showing positive neutral parent advantage and more than 74% of the combinations showing positive super-parent advantage ( Figure 5 ).

[0211] Table 3 Intermediate heterosis and super heterosis of each trait

[0212]

[0213] 1.3 Analysis of variance of various traits of parents and hybrids

[0214] A variance analysis was performed on the phenotypes of nine traits in 340 experimental materials, including parents and F1 hybrids, using plot averages. As shown in Tables 4 and 5, the results showed that, with the exception of the number of effective siliques in the main sequence, which did not reach a significant level between the parental hybrids, all other traits showed extremely significant differences. A variance analysis was performed on the phenotypes of nine traits in 306 hybrids using plot averages, and the results were similar to those for the parents and progeny materials. With the exception of the number of effective siliques in the main sequence, which did not reach a significant level, all other traits showed extremely significant differences.

[0215] Table 4 Analysis of variance of 9 traits of parental hybrids

[0216]

[0217] Note:***:P<0.001;**:P<0.01;*:P<0.05. Note:***:0.001;**:0.01;*:0.05.

[0218] Table 5 Analysis of variance of nine hybrid traits

[0219]

[0220] Note:***:P<0.001;**:P<0.01;*:P<0.05. Note:***:0.001;**:0.01;*:0.05.

[0221] 1.4. Trait-phenotype correlation analysis

[0222] The correlation analysis of the phenotypic means between the nine traits of the parental hybrids was performed. Figure 6 As shown, for example, per-plant yield was positively correlated with the number of siliques per plant, plant height, effective length of the primary sequence, starting point of the first effective branch, and number of effective siliques in the primary sequence. Yield per plant showed extremely significant positive correlations with the number of siliques per plant, plant height, and effective length of the primary sequence (P < 0.01), and relatively significant positive correlations with the starting point of the first effective branch and number of effective siliques in the primary sequence (P < 0.05). Yield per plant did not show significant or extremely significant negative correlations with other traits. However, 1000-grain weight showed extremely significant negative correlations with the number of siliques per plant, starting point of the first effective branch, and effective length of the primary sequence, and relatively significant negative correlations with the starting point of the first effective branch and number of effective branches. Among the correlation tests among the traits, plant height had the highest positive correlation with the starting point of the first effective branch, with a correlation coefficient of 0.71. The starting point of the first effective branch had an extremely significant negative correlation with effective length of the primary sequence, with a correlation coefficient of -0.25.

[0223] 1.5 Combining ability and heritability analysis

[0224] Phenotypic combining ability and heritability analysis of nine traits in 306 F1 hybrids (Table 6) revealed that the variance of the general combining ability ranged from 50.00% to 97.18%, with the highest variance for 1000-grain weight and the lowest for number of primary effective branches. The variance of the specific combining ability ranged from 2.82% to 50%, with the highest variance for number of primary effective branches and the lowest for 1000-grain weight. The variance of the general combining ability for all nine traits was above 50%, indicating that additive effects predominantly affected these traits in the first-generation hybrid population. The large variances in the general combining ability for 1000-grain weight and number of grains per silique suggest that these two traits are strongly influenced by additive effects and are thus stably inherited. The general combining ability for number of primary effective branches was 50%, indicating that this trait is influenced not only by additive but also by non-additive effects and is susceptible to environmental factors, resulting in poor genetic stability. The narrow-sense heritability and broad-sense heritability of each trait ranged from 0.156 to 0.667 and 0.312 to 0.684, respectively. The largest narrow-sense heritability and broad-sense heritability were for 1000-grain weight, while the smallest were for the number of primary effective branches.

[0225] Table 6 Genetic parameters related to each trait

[0226] Note: TSW: thousand-seed weight; SS: number of seeds per silique; SY: yield per plant; PH: plant height; MIL: effective length of main sequence; BN: one

[0227]

[0228] BH: the number of secondary effective branches; MIS: the number of effective siliques in the main sequence; SL: the length of silique.

[0229] 2. Genotype data analysis

[0230] 2.1 DNA extraction results

[0231] The genomic DNA of 35 parents was extracted using a plant genome kit. The integrity and purity of the DNA were detected by 1% agarose gel electrophoresis. The concentration of the genomic DNA was detected by spectrophotometer. The main DNA bands were clear, and the DNA quality and concentration met the requirements of the SNP chip test ( Figure 7 ).

[0232] Genotyping of 35 parents was performed using the Illumina Brassica napus 60K SNP array, yielding a total of 52,157 markers. After preliminary filtering to exclude low-quality sites and sites with no polymorphism between the parents, 34,103 usable SNP marker data were obtained. For GWAS and GS prediction, these 34,103 markers were further processed. Data quality control was first performed using Plink software (zzz.bwh.harvard.edu / plink), with the minimum allele frequency (MAF) set to 0.01 and the missing value rate set to 0.05. Missing values ​​for SNP markers were then filled using Beagle software, resulting in a total of 23,425 SNP markers for subsequent GWAS and GS model construction.

[0233] 2.2. Marker distribution and marker density analysis

[0234] Using Darmor-bzh as the reference genome, the marker density was calculated based on the physical positions of the SNP markers obtained after filtering and a heat map was drawn (Table 7 and Figure 8 SNP markers cover the entire rapeseed genome, with the number of markers per chromosome and marker density varying at different locations on the same chromosome. The number of markers on a chromosome is correlated with its length, with the longer chromosomes C03 and C04 having the highest marker counts, at 1808 and 2353, respectively. The average marker density range for chromosomes is 0.015 to 0.073 Mb. This high marker density ensures that more genes are in linkage disequilibrium with SNP loci, which is fundamental to the accuracy of GWAS and GS.

[0235] Table 7. Distribution of SNP markers on the genome

[0236]

[0237] 3. Whole-genome prediction

[0238] In whole-genome selective breeding research, the most widely used algorithms include RRBLUP, GBLUP, Bayes A, Bayes B, Bayes C, and Bayesian LASSO. This study also attempts to use these algorithms to explore the prediction accuracy under different conditions in order to find the best prediction method.

[0239] 3.1 Prediction Accuracy of Different Prediction Models

[0240] Using 23,425 SNP markers, the effects of different GS models on the prediction accuracy of 9 traits were explored through 5-fold cross validation (Table 8). There was no significant difference in the prediction accuracy of each prediction model, and the prediction accuracy of the Bayesian method prediction model was slightly higher than that of the RRBLUP and GBLUP models. The prediction accuracy of different traits varied significantly, with r values ​​ranging from 0.193 to 0.757. The thousand-grain weight and the number of grains per silique were higher, while the number of effective branches at one time and the number of effective siliques in the main sequence were lower. For example, for single plant yield, the prediction accuracy of RRBLUP, GBLUP, Bayes A, Bayes B, Bayes C, and Bayesian LASSO models were 0.333, 0.337, 0.348, 0.349, 0.346, and 0.341, respectively. The dominant effect was added to the GBLUP model to construct the GBLUP_D model, and it was found that this model did not significantly improve the prediction accuracy ( Figure 9 A variance analysis of the 30 prediction accuracies of the seven models revealed no significant differences between the models (Table 9). Given the lack of significant differences between the models, the GBLUP model possesses an absolute advantage in computational speed. Therefore, this study primarily utilized the sommer package to establish the GBLUP model when exploring the impact of other factors on model prediction accuracy.

[0241] Table 8 Prediction accuracy and standard deviation of six genome-wide selection models for different traits Note: r is the Pearson correlation coefficient between the true phenotypic value and the estimated breeding value; sd is the standard deviation of 30 Pearson correlation coefficients.

[0242]

[0243] Table 9 Analysis of variance of prediction accuracy of 7 GS models for each trait

[0244]

[0245] Note:***:P<0.001;**:P<0.01;*:P<0.05. Note:***:0.001;**:0.01;*:0.05.

[0246] 3.2 Impact of Marker Density on Prediction Accuracy

[0247] In this study, 100, 200, 300, 400, 500, 750, 1000, 1500, 2000, 5000, 10000, 20000, and 23425 marker sites were randomly selected to establish the GBLUP model to explore the effect of different marker densities on the accuracy of GS prediction. The results showed that in GS prediction, as the marker density increased, the prediction accuracy of GS also increased significantly. When the marker density increased to a certain number, the increase trend of GS prediction accuracy slowed down ( Figure 10 ). When the number of markers for 1000-grain weight, number of grains per silique, yield per plant, and silique length is within 1000 SNPs, the r value shows an inflection point, and the prediction accuracy basically tends to be stable. Thereafter, with the increase in the number of markers, the improvement in prediction accuracy is not obvious. This shows that the kinship matrix established with 1000 SNPs can meet the requirements for the establishment of the whole-genome selection GBLUP model. The number of markers required to achieve marker stability for the starting point of the effective branch and plant height of the GBLUP model is about 5000. The prediction accuracy of the effective length of the main sequence, the number of effective siliques of the main sequence, and the number of effective branches of the first time is low, and the marker density has no obvious effect on the prediction accuracy of the GBLUP model. This may be related to the low heritability of these three traits. Therefore, for traits with low heritability, the prediction accuracy cannot be improved simply by increasing the number of markers. Variance analysis was performed on the prediction accuracy of 30 five-fold cross validations with different marker densities. There were extremely significant differences between different molecular marker densities and prediction accuracy for 1000-grain weight, number of grains per silique, yield per plant, plant height, starting point of a primary effective branch, number of effective siliques in the main sequence, and silique length. Increasing the number of markers at lower marker densities significantly improved the prediction accuracy (Table 10).

[0248] Table 10 Analysis of variance on GS prediction accuracy at different marker densities

[0249]

[0250] Note:***:P<0.001;**:P<0.01;*:P<0.05. Note:***:0.001;**:0.01;*:0.05.

[0251] 3.3 Impact of Population Size on Prediction Accuracy

[0252] This study used six populations of different sizes to analyze the accuracy of GS prediction. The results showed that ( Figure 11 ), GS prediction accuracy slowly increases with population size. When the population reaches a certain size, GS achieves optimal prediction for each trait. For traits with high heritability, smaller populations achieve optimal prediction accuracy. When the number of individuals in a population is smaller than a certain threshold, genome-wide prediction using five-fold cross-validation is insufficient for some traits with low heritability. With the exception of the number of primary effective branches, genome-wide selection achieves optimal prediction accuracy for all traits when the population size reaches 250.

[0253] 3.4 Impact of Heritability on Prediction Accuracy

[0254] Pearson correlation analysis was performed on the average values ​​of the prediction accuracy of 30 five-fold cross-validations of the seven models and the broad-sense heritability of the phenotypic data. It was found that there was a significant positive correlation between them. Traits with high heritability tended to have higher prediction accuracy. The correlation between prediction accuracy and heritability reached 0.988 (Table 11), indicating that the heritability of the trait greatly affects the prediction accuracy.

[0255] Table 11 Correlation coefficients between trait heritability and prediction accuracy

[0256]

[0257] Note: TSW: thousand-grain weight; SS: number of grains per silique; SY: yield per plant; PH: plant height; MIL: effective length of main sequence; BN: number of primary effective branches; BH: starting point of primary effective branch; MIS: number of effective siliques in the main sequence; SL: silique length.

[0258] 4. Prediction model of significant marker sites based on association analysis

[0259] The prediction results based on whole-genome selection are not very ideal, possibly because many marker sites unrelated to the target trait gene affect the results. In this study, we set out to study the prediction model using significantly associated sites.

[0260] 4.1. Screening of significant association loci by genome-wide association analysis

[0261] In this study, TASSEL software was used to perform principal component analysis and kinship analysis on 340 experimental materials including parents and hybrid F1, and the data were visualized using R software ggplot2 package and heatmap package ( Figure 12 As can be seen from the figure, the most optimal number of principal components is 5. Using the PCA+K mixed linear model for GWAS analysis, a total of 282 loci associated with yield traits were detected at the significance threshold of -log10(1 / 23425) = 4.37 (1 level) (Table 12). Of these, 201 SNP markers were associated with 1000-grain weight, the largest number. Individual markers explained 5.539% to 7.744% of the phenotypic variation. These loci were located on chromosomes A01, A03, A07, A08, C02, C03, C04, C05, C07, and C09, with the largest number of SNP loci on chromosomes C04 and C07. This is due to the influence of LD, as loci surrounding a strongly associated locus also exhibit a continuous shift in association from high to low, confirming the locus's reliability. No SNP loci were detected with the number of grains per silique at the significance threshold.

[0262] Table 12 Statistics of SNP sites significantly associated with each trait

[0263]

[0264]

[0265] Note: Significant threshold value of -log10(1 / 23425) = 4.37.

[0266] According to the observed and expected values ​​of -log10(P) of all SNPs, the qqman package (0.1.8) of R language was used to draw Quantile-Quantile scatter plots and Manhattan plots ( Figure 13 As can be seen from the figure, the observed and expected P values ​​for SNPs associated with per-plant yield are well aligned. At a significance threshold of -log10(1 / 23425) = 4.37, significant SNPs associated with per-plant yield were identified, located on chromosomes A01, A03, C03, and C08. Results for other traits were similar.

[0267] 4.2 GBLUP model based on significant sites in association analysis

[0268] Based on the results of GWAS analyses of various traits, six threshold levels (0.1, 1, 10, 100, 1000, and all levels) (23,425 marker loci) were determined using the Bonferroni multiple-test method based on the -log10(P) value. The number of markers screened at different threshold levels increased with increasing -log10(P) values ​​(Table 13). A new molecular marker matrix was constructed using the significantly associated molecular markers at different threshold levels to calculate the genomic relationship G matrix and construct a GBLUP prediction model. The results showed that for each trait, the prediction accuracy for 1000-grain weight increased with increasing threshold levels and the number of associated markers. For other traits, the prediction accuracy initially increased and then decreased. This initial increase and then decrease is hypothesized to be due to the fact that as the threshold level increases, loci associated with the target trait are gradually identified, and the number of false-positive markers increases after reaching a certain critical value. The model prediction accuracy was highest when the markers in the model were closest to the true loci, but the marker screening threshold for the 1000-grain weight prediction model did not reach the critical value. Among these traits, the threshold levels corresponding to the models with the highest prediction accuracy were different for different traits. Except for 1000-grain weight, the best prediction models based on significantly associated loci were significantly more accurate than the GS models that included all marker loci ( Figure 14). This indicates that screening for significant trait association loci is very necessary to improve the model prediction effect.

[0269] Table 13 Number of markers at different threshold levels for each trait

[0270]

[0271]

[0272] Note: n represents the number of significant loci in GWAS above the threshold, and r represents different levels of prediction accuracy.

[0273] Comparing the prediction accuracy of 30 cross-validations between the model established by GWAS screening of significant associated sites and randomly selecting the same number of marker sites, the variance analysis results showed that there was a very significant difference between the two. Screening of significant associated markers can significantly improve the prediction accuracy of the model (Table 14)

[0274] Table 14 GBLUP model variance analysis of GWAS-screened significant association markers and randomly selected markers

[0275]

[0276] Note:***:P<0.001;**:P<0.01;*:P<0.05. Note:***:0.001;**:0.01;*:0.05.

[0277] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention, and they should all be included in the scope of the claims and description of the present invention.

Claims

1. A method for predicting heterosis in Brassica napus based on whole genome selection, characterized in that: The method comprises: Obtaining parental materials and configuring several hybrid combinations based on the parental materials; Based on the parental materials and the hybrid combination, determining the phenotype of each parental material and the hybrid combination thereof; Performing genotyping based on the parental materials to obtain genotyping data; Based on the trait phenotype and the genotype data, screen out the sites with significant differences between groups and the sites with significant association analysis; Constructing multiple prediction models based on the phenotypic and genotypic data, significant points of intergroup differences, and significant points of association analysis; Based on the significant differences between the groups, a multivariate linear analysis was performed using the stepwise regression method to obtain a multivariate stepwise regression model; Based on cross-validation, the prediction accuracy of multiple prediction models and the multivariate stepwise regression model is evaluated to determine the optimal model for the trait; The significant differences between the groups were screened out by the following method: Calculating the types of marker types present in the hybrid combination based on the marker types of the parents, wherein the marker types of the parents include one or a combination of a reference genome type, a mutant type, and a heterozygous type; The hybrid combinations are grouped according to marker types, phenotypic mean differences among the groups are statistically analyzed, and significant points of difference among the groups are determined according to the phenotypic mean differences; There are 6 marker types in the hybrid combination, namely M P1 , M P2 , M F1 , M F2 , M BC1 and M BC2 , Analyze the differences between groups according to the marker type to see whether there are calculable additive effects and dominant effects, and establish the inter-group (0, 1) coefficient matrix Ka, Kd and the inter-group standardized effect matrix A based on the significance of the differences. 6*1 and D 10*1 , the additive effect a and dominant effect d of the significant site are calculated according to the following formula: a=Ka*A / ∑ka(1,i) d=2Kd*D / ∑kd(1,i) The intergroup differences that can be calculated for additive effects include 6 groups: (M F2 , M BC1 )、(M F2 , M BC2 )、(M F1 , M BC1 )、(M F1 , M BC2 )、(M BC1 , M BC2 ) and (M P1 , M P2 ); There are 10 groups with calculable differences between groups in terms of dominant effect: (M F1 , M F2 )、(M BC1 , M BC2 )、(M p1 , M F2 )、(M p1 , M F1 )、(M p1 , M BC1 )、(M p1 , M BC2 )、(M p2 , M F2 )、(M p2 , M F1 )、(M p2 , M BC1 ) and (M p2 , M BC2 ); The multivariate linear analysis was performed using the stepwise regression method to obtain a multivariate stepwise regression model, which specifically includes: The effect matrix E is established based on the additive effect a and dominant effect d of the significant differences between the groups. n*2 =(a i , d i ), the marker type matrix M of the hybrid combination is calculated based on the marker types of the hybrid parents at the significant sites 6*n , calculated using the following formula: Ef=E·K·M The main diagonal elements of Ef are Ef (i,i) The labeling effect of n significant sites that make up a certain material; The effect X of the marker loci related to the trait and the phenotypic value of the trait were used for stepwise regression analysis, and a multivariate stepwise regression model was established: y=Xβ+ε where y is the phenotypic observation, β is the fixed effect vector, X is the marker effect matrix, ε is the residual effect, and ε obeys normal distribution, is the residual variance.

2. The method according to claim 1, characterized in that The trait phenotypes of the various parental materials and their hybrid combinations include intermediate parental advantage, super parental advantage, correlation between various traits, combining ability and heritability.

3. The method according to claim 1, characterized in that The significant points of association analysis were screened out by the following method: Perform principal component analysis and kinship analysis on the filtered and filled genotype data to obtain the PCA matrix and K matrix; The PCA matrix and K matrix were used as covariates to establish a mixed linear model of population phenotypes and genes for association analysis: y=Xβ+Zμ+ε Where y is the observed phenotypic value, β is the fixed effect vector including genetic markers and population structure, μ is the random additive genetic effect vector of individuals / lines, X and Z are the design matrices with unknown fixed and random effects, respectively, and ε is the residual effect. μ and ε obey and Normal distribution, K is the genomic kinship matrix, is the additive genetic variance, is the residual variance; The correction factor multiple test is used to determine the significance level P of the point. The significance level P reflects the degree of association between the marker and the phenotypic variation. The smaller the P value, the higher the degree of association between the marker and the trait variation. Set different significance thresholds to filter out significant points in association analysis.

4. The method according to claim 1, wherein The method of evaluating the prediction accuracy of multiple prediction models and the multivariate stepwise regression model based on cross-validation to determine the optimal model for the trait specifically includes: Multiple prediction models were evaluated using Pearson correlation coefficients between true phenotypic values ​​and estimated breeding values; The multivariate stepwise regression model was evaluated by the coefficient of determination and mean absolute error.

5. A device for predicting heterosis of Brassica napus based on whole genome selection, for implementing the method according to any one of claims 1 to 4, characterized in that: The device comprises: an acquisition module, configured to acquire parental materials and configure a plurality of hybrid combinations according to the parental materials; a determination module configured to determine the trait phenotypes of each parent material and the hybrid combination thereof based on the parent material and the hybrid combination; A typing module is configured to perform genotyping based on the parental material to obtain genotyping data; A screening module is configured to screen out significant points of inter-group differences and significant points of association analysis based on the trait phenotype and the genotype data; A construction module is configured to construct multiple prediction models based on the trait phenotype, genotype data, significant points of inter-group differences, and significant points of association analysis; An analysis module is configured to perform a multivariate linear analysis based on the significant difference points between the groups using a stepwise regression method to obtain a multivariate stepwise regression model; The evaluation module is configured to evaluate the prediction accuracy of multiple prediction models and the multivariate stepwise regression model based on cross-validation, and determine the optimal model for the trait.

6. An electronic device comprising: Controller; A memory for storing one or more programs, which, when executed by the controller, causes the controller to implement the method according to any one of claims 1 to 4. 7 . A non-transitory computer-readable storage medium storing instructions, which, when executed by a processor, executes the method according to claim 1 .