A method for dividing maize heterosis group based on marker dominance effect weight
By using a marker-based dominant effect weighting method, the dominant effect of maize inbred lines was identified using SNP chips. Genetic distances were calculated and cluster analysis was performed, which solved the problem of low efficiency in the existing classification of maize heterosis groups. This method achieved rapid and accurate group classification and improved breeding efficiency.
Patent Information
- Application Number
- CN202211385428.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-11-07
AI Technical Summary
Existing methods for classifying maize heterotic groups suffer from problems such as neutral molecular markers, time-consuming and labor-intensive testing of combining ability in mating populations, and limited ability of isozyme methods to distinguish the diversity of breeding materials, resulting in low breeding efficiency.
A marker-based dominant effect weighting method was used to identify the dominant effect of maize inbred lines using SNP chips, calculate genetic distances, establish a fingerprint database, and perform cluster analysis to quickly and accurately classify maize heterotic groups.
It enables rapid and accurate classification of maize inbred lines, improves breeding efficiency, saves time, manpower and financial resources, provides a basis for breeding decisions, and reduces the blindness of combination.
Smart Images

Figure CN115910205B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of corn breeding, and particularly relates to a corn heterosis group division method based on marker dominance effect weight. BACKGROUND
[0002] Corn is one of the most successful crops in heterosis utilization, and corn breeding mainly utilizes heterosis, so the division of heterosis groups is an important content of corn breeding. In the improvement, expansion and innovation of corn germplasm, the principle of heterosis groups must be followed to improve the breeding efficiency.
[0003] At present, the commonly used heterosis group division methods in China include the methods according to morphology, yield combining ability, geographical origin and pedigree, isozyme and molecular marker. Molecular marker is a genetic marker based on the variation of nucleotide sequences in genetic material between individuals, which is not affected by environmental conditions and developmental stages, and has been successfully applied to the division of inbred line groups.
[0004] In actual breeding, breeders mainly determine the heterosis group to which an inbred line belongs according to the special combining ability of the hybrid formed between the inbred line and a standard test variety. There are the following problems in the division of heterosis groups: (1) molecular markers are neutral in the corn genome, and the contribution rate of a single molecular marker to the heterosis effect value is greatly different in breeding, and the division result of the genetic distance between markers often has deviation; (2) the accurate detection of combining ability of the test and mating population needs to be carried out in multiple points, multiple environments and multiple repetitions, and the process of combining ability identification is time-consuming and labor-intensive; (3) the methods such as isozyme are limited by genetic loci, and the diversity of breeding materials is limited. It is of great significance for future corn variety breeding to establish a more efficient and accurate heterosis group division method. SUMMARY
[0005] The purpose of the embodiments of the application is to provide a corn heterosis group division method based on marker dominance effect weight, aiming at solving the problems raised in the third part of the background.
[0006] The embodiments of the application are implemented in this way, a corn heterosis group division method based on marker dominance effect weight, the method comprises:
[0007] Take 3-4 copies of each representative inbred line of the corn heterosis group, and form a hybrid population by double cross of all representative corn inbred lines;
[0008] The genotype of the representative corn inbred line parent is identified by using an SNP chip, the contribution value of a single SNP dominance effect to the heterosis of the hybrid population is obtained as the dominance effect weight of the SNP;
[0009] The hybridization genetic distance between the representative inbred lines of each heterotic group of corn is calculated, and a fingerprint database of the hybridization genetic distance between the representative inbred lines of each heterotic group of corn is established.
[0010] SNP marker information of the to-be-tested corn inbred line is acquired, the hybridization genetic distance between the to-be-tested inbred line and the representative inbred line of the heterotic group is calculated, and the representative inbred line of the heterotic group is subjected to cluster analysis in combination with the hybridization genetic distance, so as to determine whether the minimum value of the hybridization genetic distance is within the threshold range of the representative inbred line of the target group, and then the to-be-tested inbred line is classified into a group.
[0011] Preferably, in the step of selecting 3-4 representative inbred lines of each heterotic group of corn, the hybrid phenotypes of the representative inbred lines include the following contents:
[0012] 3-4 representative inbred lines of each heterotic group of corn are selected, and the representative inbred lines of each heterotic group of corn are Reid-B73, Reid-PH6WC, Reid-tie7922, Tangsipingtou-Huangzaosu, Tangsipingtou-Chang7-2, Tangsipingtou-Zheng22, Lvdahonggu-Dan340, Lvdahonggu-Zheng22, Lancaster-Mo17, Lancaster-PH4CV, Lancaster-NK764, PA-Ye478, PA-Zheng58, PB-Qi319 and PB-Dan599.
[0013] The inbred lines are subjected to diallel cross to form a plurality of hybrid populations, and 15 inbred lines are subjected to partial diallel cross to form 105 hybrids, the hybrids are planted in Xinjiang Shihui, Wujiaku and Yili for two repetitions of yield, growth period, plant height and ear position, and BLUP calculation is performed on each trait for subsequent grouping.
[0014] Preferably, in the step of acquiring the contribution value of the single SNP dominant effect to the hybridization phenotype of the hybrid population as the weight of the dominant effect of the SNP, the content of calculating the weight of the dominant effect of the single SNP marker is as follows:
[0015] The representative inbred lines of corn are subjected to genotype identification by using a SNP chip, 15 representative inbred lines are subjected to genotype identification by using a Maize 2K breeding chip, SNP markers are subjected to quality control, and the requirements are that the frequency of micro-effect alleles is greater than 5%, the deletion rate of the site is less than 20%, and the heterozygosity rate is less than 20%.
[0016] Preferably, the content of acquiring the contribution value of the single SNP dominant effect to the hybridization phenotype of the hybrid population is as follows:
[0017] The dominant effect of the single SNP is estimated by using the Bayesian method of whole genome selection, and the prediction model is as follows:
[0018] Y = μ + Z A a+Z D d+e①,
[0019] Where Y represents the hybrid trait, Z A The design matrix for additive effects of SNP markers is a = (a1, ... a2) / a3. l ), where a is the random matrix of additive effects, and Z D The design matrix for the SNP marker dominance effect is d = (d1, ... d2). l ), where d is the random matrix of the dominance effect, l is the number of SNP labels, and d u It is the dominant effect of a single SNP, where u = 1, 2, ..., l, and the hybrid trait is one of the yield trait, growth period trait, and plant type trait.
[0020] The preferred formula for calculating the dominance effect weight of a single SNP marker u is as follows:
[0021]
[0022] in It is the average of the absolute values of the dominant effects of all SNPs.
[0023] Preferably, the steps for establishing a fingerprint database of heterosis genetic distances among representative inbred lines of each maize dominant group specifically include:
[0024] The genetic distance of heterosis between heterotic parents is calculated by combining SNP markers with weighted dominance effects. The formula is as follows:
[0025]
[0026] Where X and Y are the genotypes of the parents, X uj and Y uj n represents the frequency of the j-th allele at the u-th locus. u Let w represent the number of alleles at the u-th locus, l represent the number of loci, and w represent the number of alleles at the u-th locus. u The dominant effect weight represents the value of the u-th locus. It is noteworthy that a positive dominant effect increases heterosis genetic distance, while a negative dominant effect decreases it. The database stores heterosis genetic distance information for 15 representative inbred lines and performs cluster analysis on these 15 representative maize inbred lines to classify the corresponding heterosis groups for each representative inbred line.
[0027] Preferably, the step of calculating the genetic distance of heterosis between the inbred line to be tested and the representative inbred lines of the heterotic group specifically includes:
[0028] The to-be-tested inbred lines are identified by using a Maize 2K breeding chip, and quality control is performed, and the screened SNPs are used for subsequent calculation;
[0029] The heterosis genetic distance between the to-be-tested inbred lines and each representative inbred line of the dominant group is calculated;
[0030] The heterosis genetic distance between the to-be-tested inbred lines and each representative inbred line of the dominant group is combined for cluster analysis;
[0031] The threshold value is defined: after the heterosis genetic distance between each two representative inbred lines is calculated, the average value u and the standard deviation sigma of all heterosis genetic distances are calculated, and the value of u+sigma is taken as the threshold value corresponding to the cluster group;
[0032] It is judged whether the minimum value of the heterosis genetic distance of the to-be-tested inbred line is less than the threshold value of the corresponding cluster group, if yes, the to-be-tested inbred line is formally divided into the cluster group where the minimum genetic distance representative inbred line is located; if the minimum genetic distance is not within the threshold value range of the calculated cluster group, the to-be-tested inbred line is added to a new type of heterosis group.
[0033] The application provides an application of a corn inbred line cluster identification method in crop variety breeding, and has the beneficial effects that the genetic relationship between a to-be-tested inbred line and a representative inbred line is judged based on the combination of genomic SNP marker information and the dominant effect between inbred lines, so that a technical route for quickly and accurately classifying corn inbred lines is provided. The application method has a wide application field, and is not limited to species in terms of application field. The application method can quickly and effectively identify the cluster of crop inbred lines, and then select the corresponding heterosis mode, so that the hybrid combination is selected in a targeted manner, thereby greatly improving the breeding efficiency, saving time, manpower, material resources and financial resources, providing a decision basis for breeders, reducing the blindness of combination, and providing reference and guidance for crop breeders. BRIEF DESCRIPTION OF DRAWINGS
[0034] Figure 1 A flowchart of a corn heterosis group classification method based on marker dominant effect weight provided by the embodiment of the application;
[0035] Figure 2 A distribution diagram of 18 representative corn inbred lines high-quality SNPs on a genome provided by the embodiment of the application;
[0036] Figure 3 A cluster comparison diagram of heterosis genetic distance and molecular marker genetic distance of 18 representative corn inbred lines (left Rogers's genetic distance cluster, right heterosis genetic distance cluster) provided by the embodiment of the application;
[0037] Figure 4 The results of heterosis genetic distance and molecular marker genetic distance classification of 18 representative maize inbred lines and 2 maize inbred lines to be tested provided in the embodiments of the present invention are shown in the figure (Rogers's genetic distance cluster on the left and heterosis genetic distance cluster on the right). Detailed Implementation
[0038] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0039] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.
[0040] Genomic selection utilizes molecular markers evenly distributed across the genome. By identifying the genotypes and analyzing the phenotypes of a training population, and estimating the effect value of each marker based on the correlation between genotype and phenotype, a genetic model is established to predict phenotypes using marker genotypes. The dominance hypothesis of heterosis posits that the more dominant effects an organism possesses, the more advantageous it is for adaptive survival. Hybrids concentrate dominant genes from both parents, thus exhibiting heterosis. Models established by genomic selection, including the dominant genetic effects of individual markers, can be combined with genetic distances between markers for more accurate calculations of kinship between different inbred lines.
[0041] The parents and lineages of the main maize varieties cultivated in China are divided into six major heterotic groups. The representative inbred lines of these groups are Reid (B73, PH6WC, Tie 7922), Tangsi Pingtou (Huangzao Si, Chang 7-2, Zheng 22), Lüda Honggu (Dan 340, Zheng 22), Lancaster (Mo17, PH4CV, NK764), PA (Ye 478, Zheng 58), and PB (Qi 319, Dan 599). Fifteen representative inbred lines of these groups are selected as the representative inbred lines of each heterotic group.
[0042] Example 1
[0043] like Figure 1 The diagram shows a flowchart of a method for classifying maize heterosis groups based on the dominant effect weight of markers, provided by an embodiment of the present invention. The specific implementation of the present invention is illustrated by a simple case of identifying maize inbred line groups based on the dominant effect weight of SNP markers.
[0044] (1) Take 3-4 copies of each representative inbred line of the main maize heterotic group, perform double cross between all representative maize inbred lines to form hybrid population, and identify hybrid phenotypes such as yield in multiple years, multiple sites and multiple repetitions;
[0045] The experimental objects selected in this case are Luda Honggu, Lancaster, PA, Reid, Tang Sipingtou, and PB heterotic groups. The molecular marker data of 20 maize inbred lines Maize 55K are selected from 6 maize heterotic groups. 18 inbred lines are randomly selected as representative heterotic group control samples (pedigree source is shown in Table 1 background colorless), and 2 inbred lines are selected as test samples (background gray, PH4CV, Huang C). The maize inbred line group rapid identification method based on hybrid heterosis genetic distance is verified, and the specific implementation process is as follows:
[0046] Table 1 shows 20 representative maize inbred lines of 6 heterotic groups
[0047]
[0048]
[0049] 18 representative maize inbred lines were used to form 153 maize hybrids by partial double cross in Hainan in winter 2016. The 153 maize hybrids were planted in Xinjiang Shihui City and Xinjiang Wujiaqu City in 2017-2018. The field test used an alphalatin square design, with two repetitions, 2 rows of planting, row length 3m, row width 50+60cm, and plant spacing 23cm. The yield traits of each hybrid and its parent inbred line were investigated at the harvest stage. After BLUP of the data of each test site, it was used for subsequent analysis.
[0050] (2) The representative maize inbred line parent is genotyped by using the SNP chip to obtain the contribution value of the single SNP dominant effect to the heterosis of the hybrid population, which is used as the dominant effect weight of the SNP
[0051] The 18 representative maize inbred lines are genotyped by using Maize 55K chip. Initially, 50812 SNPs markers evenly distributed on 10 chromosomes of maize are obtained. Then, the Plink is used to delete the SNPs sites with deletion rate greater than 20%, heterozygosity greater than 20%, and micro-effect allele frequency less than 5%. The Plink is used for LD filtering, and finally 20555 high-quality SNPs are obtained for genotype analysis. The SNP density distribution on the whole genome level of maize inbred lines is shown in Figure 2 .
[0052] The Bayesian C model of the "BGLR" package of R language is used, 10-fold cross-validation method is adopted, and the prediction model Y=μ+ZA a+Z D d+e①Perform analysis, extract the yield contribution rate d of all single SNP markers obtained by model analysis u .
[0053] (3) Calculate the heterosis genetic distance between the representative 18 corn inbreds based on the weight of all SNP markers and the molecular marker genetic distance based on Rogers's method, respectively, and establish a fingerprint database of the heterosis genetic distance of the representative inbreds of each heterosis group in corn
[0054] According to the formula , the heterosis genetic distance between inbred parents is calculated, where X and Y are the genotypes of the parents, X uj and Y uj represent the frequency of the jth allele of the uth locus, n u represents the number of alleles of the uth locus, l represents the number of loci, and w u represents the dominance effect weight of the uth locus. Calculate the heterosis genetic distance between all representative inbreds, and store the heterosis genetic distance information of the 18 representative inbreds in the database. Cluster analysis is performed on the 18 representative corn inbreds using the hclust function of the "cluster" package based on R language, and the corresponding heterosis groups of each representative inbred are divided.
[0055] According to the formula , the molecular marker genetic distance between inbred parents is calculated, where X and Y are the genotypes of the parents, X uj and Y uj represent the frequency of the jth allele of the uth locus, n u represents the number of alleles of the uth locus. Calculate the molecular marker genetic distance between all representative inbreds, and store the molecular marker genetic distance information of the 18 representative inbreds in the database. Cluster analysis is performed on the 18 representative corn inbreds using the hclust function of the "cluster" package based on R language, and the corresponding heterosis groups of each representative inbred are divided. Figure 3 (Left Rogers's genetic distance clustering, right heterosis genetic distance clustering) Compare the clustering results of the heterosis groups by the two genetic distance calculation methods, and the clustering results of the heterosis genetic distance are more close to the pedigree information results.
[0056] (4) Obtain the SNP marker information of the two corn inbreds to be tested, calculate the heterosis genetic distance between the inbreds to be tested and the representative inbreds of the heterosis group, and determine whether the minimum value of the heterosis genetic distance is within the threshold range of the target group representative inbred, and then divide the class of the inbred to be tested.
[0057] Obtaining the genotyping results of the to-be-tested inbred lines Huang C and PH4CV, and performing high-quality filtering and screening on the SNP markers, the method being the same as step (2).
[0058] According to step (3), the heterotic genetic distances between the two to-be-tested corn inbred lines and the 18 representative inbred lines and the molecular marker genetic distances according to Rogers's method are calculated. In combination with the heterotic genetic distance information of the 18 representative inbred lines stored in the database, the hclust function of the "cluster" package based on R language is used to perform cluster analysis on the 20 corn inbred lines, and the corresponding heterotic groups of the representative inbred lines are divided.
[0059] The average heterotic genetic distance u of the 18 representative corn inbred lines is 0.138, the standard deviation σ is 0.034, and the threshold value corresponding to the cluster group is 0.172. The threshold value 0.296 is used to draw a line on the heterotic cluster result, and it can be seen from Figure 4 that the heterotic genetic distance of the inbred line Huang C is less than the threshold line range of the Sipingtou group, and therefore the inbred line is classified into the Sipingtou group; the heterotic genetic distance of the inbred line PH4CV is less than the threshold line range of the Lancaster group, and therefore the inbred line is classified into the Lancaster group.
[0060] The above merely describes preferred embodiments of the present application but should not be used to restrict the present application, and any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for classifying maize heterotic groups based on the weighting of marker dominance effects, characterized in that, The method includes: Take 3-4 representative inbred lines from each maize heterotic group, and perform diallel crosses of all representative maize inbred lines to form a hybrid population; Genotyping of representative maize inbred lines was performed using SNP chips to obtain the contribution value of the dominant effect of a single SNP to heterosis in the hybrid population, which was used as the dominant effect weight of that SNP. Calculate the heterosis genetic distance among representative maize inbred lines based on the dominant effect of all SNP marker weights, and establish a fingerprint database of heterosis genetic distances among representative inbred lines of each maize dominant group; Obtain SNP marker information of maize inbred lines to be tested, calculate heterosis genetic distance between the inbred lines to be tested and the representative inbred lines of the heterosis group, and perform cluster analysis based on the heterosis genetic distance of the representative inbred lines to determine whether the minimum heterosis genetic distance is within the threshold range of the representative inbred lines of the target group, and then classify the inbred lines to be tested into groups. The contribution of a single SNP dominance effect to the hybrid phenotype of a hybrid population is obtained, as follows: The dominance effect of a single SNP was estimated using Bayesian methods of genome-wide selection. The prediction model is as follows: Y=μ+Z A a+Z D d+e①, Where Y represents the hybrid trait, Z A The design matrix for additive effects of SNP markers is a = (a1, ... a2) / a3. l ), where a is the random matrix of additive effects, and Z D The design matrix for the SNP marker dominance effect is d = (d1, ... d2). l ), where d is the random matrix of the dominance effect, l is the number of SNP labels, and d u It is the dominant effect of a single SNP, where u = 1, 2, ..., l, and the hybrid trait is one of the yield trait, growth period trait, and plant type trait; The formula for calculating the dominance weight of a single SNP marker u is as follows: in It is the average of the absolute values of the dominant effects of all SNPs; The steps for establishing a fingerprint database of heterosis genetic distances among representative inbred lines of various maize dominant groups specifically include: The genetic distance of heterosis between heterotic parents is calculated by combining SNP markers with weighted dominance effects. The formula is as follows: Where X and Y are the genotypes of the parents, X uj and Y uj n represents the frequency of the j-th allele at the u-th locus. u Let w represent the number of alleles at the u-th locus, l represent the number of loci, and w represent the number of alleles at the u-th locus. u This represents the dominance weight of the u-th locus.
2. The method for classifying maize heterotic groups based on marker dominance effect weights according to claim 1, characterized in that, In the step of obtaining 3-4 representative inbred lines from each maize heterotic group, the identification of the hybrid phenotype of the representative inbred lines includes the following: Three to four representative inbred lines from each maize heterotic group were selected. The representative inbred lines from each heterotic group were Reid-B73, Reid-PH6WC, Reid-Tie7922, Tangsipingtou-Huangzaosi, Tangsipingtou-Chang7-2, Tangsipingtou-Zheng22, Lüdahonggu-Dan340, Lüdahonggu-Zheng22, Lancaster-Mo17, Lancaster-PH4CV, Lancaster-NK764, PA-Ye478, PA-Zheng58, PB-Qi319, and PB-Dan599. Diallel crosses of inbred lines were used to form a multi-hybrid population. Partial diallel crosses of 15 inbred lines resulted in 105 hybrids. These hybrids were planted in Shihezi, Wujiaqu, and Yili in Xinjiang for double-replicated identification of yield, growth period, plant height, and ear position. Optimal linear unbiased prediction calculations were performed for each trait for subsequent genotyping.
3. The method for classifying maize heterotic groups based on marker dominance effect weights according to claim 1, characterized in that, The representative maize inbred line parents were genotyped using SNP chips to obtain the contribution value of the dominant effect of a single SNP to the heterosis of the hybrid population, which was used as the dominant effect weight of the SNP. The calculation of the dominant effect weight of a single SNP marker is as follows: Genotyping of representative maize inbred lines was performed using SNP microarrays, and genotyping of 15 representative inbred lines was performed using Maize2K breeding microarrays. Quality control was performed on SNP markers, requiring minor allele frequencies >5%, locus deletion rates <20%, and heterozygosity <20%.
4. The method for classifying maize heterotic groups based on marker dominance effect weights according to claim 3, characterized in that, The steps for calculating the genetic distance of heterosis between the inbred line to be tested and the representative inbred lines of the heterotic group specifically include: Genotyping of the inbred lines to be tested was performed using the Maize 2K breeding chip, and quality control was conducted. The selected SNPs were used for subsequent calculations. Calculate the genetic distance of heterosis between the inbred line to be tested and the representative inbred lines of each dominant group; Cluster analysis was performed by combining the genetic distance of heterosis among the inbred lines to be tested and the representative inbred lines of each dominant group; Define threshold: After calculating the heterosis genetic distance between each pair of representative inbred lines, calculate the mean u and standard deviation σ of all heterosis genetic distances, and use the value of u+σ as the threshold corresponding to this cluster. Determine whether the minimum genetic distance of heterosis of the inbred line to be tested is less than the threshold of the corresponding group. If it is, then the inbred line to be tested is officially classified into the group of the representative inbred line with the minimum genetic distance. If the minimum genetic distance is not within the calculated group threshold range, then the inbred line to be tested is added to the new type of heterosis group.
Citation Information
Patent Citations
Crop variety heterosis pattern identification method based on molecular marker technology
CN108416189A
Molecular marker technology-based crop inbred line class group identification method
CN108427866A