Method for predicting character phenotypic correlation based on genetic similarity calculated by corn character high-throughput genetic loci
By using a high-throughput genetic locus calculation method, published genetic locus data are mapped onto a unified genome sequence to calculate the similarity and significance of gene sets of genetic loci among traits. This solves the problems of high cost, long cycle and small data volume in phenotypic data collection, and enables rapid and accurate prediction of phenotypic correlations among traits.
Patent Information
- Application Number
- CN202512038808.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies for studying the correlation between phenotypic traits in organisms suffer from problems such as high cost of phenotypic data collection, long cycle time, low reproducibility and availability, small data volume and weak data interpretation ability, making it difficult to effectively utilize genetic similarity to predict phenotypic correlation.
Using high-throughput genetic locus computation methods, published genetic locus data are mapped onto a unified genome sequence to calculate the similarity and significance of gene sets at genetic loci among traits, and to predict phenotypic correlations.
It enables low-cost and rapid prediction of phenotypic correlations, provides clues for elucidating the genetic basis among traits, and improves the availability and accuracy of data.
Smart Images

Figure CN121459949A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of genomics and bioinformatics, and specifically discloses a method for predicting phenotypic correlations of traits based on genetic similarity calculated from high-throughput genetic loci of maize traits. Background Technology
[0002] Phenotypic correlations among traits refer to the statistical associations that exist between the phenotypic values of different traits in an organism. Information on phenotypic correlations has significant theoretical research and practical application value. For example, in basic biological research, it helps identify the genetic basis regulating phenotypic formation and discover genes with multiple effects; and in breeding applications, it helps optimize breeding strategies and enhance selective breeding for specific traits or one or more traits.
[0003] Currently, phenotypic correlation information is mainly obtained through phenotypic measurements, including measuring the phenotypic values of target traits in the target species and performing correlation analysis (e.g., Pearson correlation coefficient analysis, Spearman rank correlation coefficient analysis) on the phenotypic values among target traits, thereby identifying statistically significant phenotypic correlations. However, methods relying on phenotypic measurements have at least four shortcomings: (1) Phenotypic data collection is costly and time-consuming. For example, for crops, it requires inputs such as land, planting, and measurement, and certain traits can only be measured at a certain growth stage (such as crop flowering and maturity). (2) The repeatability and availability of phenotypic data is low because phenotypic data of different traits need to come from the same biological individual to form paired data in order to obtain phenotypic correlation through correlation analysis of phenotypic measurements. If the phenotypic data of two target traits cannot be paired, the experiment needs to be redesigned and carried out, which further increases the experimental cost and cycle time. (3) The small amount of phenotypic data makes it difficult to discover phenotypic correlations between certain traits; (4) The data interpretation ability is weak. Although the phenotypic measurement method can estimate the phenotypic correlation, the genetic basis behind the correlation is unknown, and additional experiments need to be designed to analyze it.
[0004] Currently, high-throughput sequencing and GWAS methods have become routine techniques for uncovering the genetic regulatory basis of traits in organisms. Studies using these methods have published a vast amount of data on genetic loci regulating trait formation. Combined with the large amount of data accumulated from QTL mapping studies, this has resulted in a massive dataset of genetic loci for biological traits. Genetic similarity between traits refers to the degree of similarity among the genetic bases regulating different traits. The genetic basis of phenotypic correlation stems from genetic similarity; that is, the more shared genetic bases between traits, the higher the phenotypic correlation tends to be. Therefore, phenotypic correlation can be predicted based on genetic similarity between traits.
[0005] Accordingly, this invention discloses a method for predicting phenotypic correlations based on genetic similarity calculation of maize traits using high-throughput genetic loci. This method obtains genetic locus information in the target species that regulates or is associated with the formation of the target trait, maps the genetic locus information of each target trait onto a unified genome sequence in the target species, extracts annotation genes within the genetic locus regions of each target trait, forms a genetic locus gene set for each target trait, calculates the similarity of the genetic locus gene sets among the target traits to represent the genetic similarity between traits, calculates the significance of this similarity, and finally uses the genetic similarity between traits and its significance to predict phenotypic correlations.
[0006] This invention overcomes the above-mentioned shortcomings of phenotypic measurement methods, specifically: (1) Existing research has identified and published a large number of genetic loci (including QTL and QTN) that regulate the formation of biological traits, which can be used to calculate genetic similarity, thereby overcoming the shortcomings of high cost and long cycle of phenotypic data collection; (2) Genetic similarity analysis is the similarity between two sets of genetic loci of a trait. It does not require paired data and can overcome the shortcomings of low repeatability and availability of phenotypic data. (3) A large number of genetic loci associated with biological traits were identified by high-throughput sequencing technology and GWAS method. The large amount of data can overcome the shortcomings of the difficulty in discovering phenotypic correlations between traits. (4) Genetic loci shared among traits identified by genetic similarity methods can directly provide clues for the analysis of the genetic basis of phenotypic correlation.
[0007] In summary, this invention utilizes genomic genetic big data and bioinformatics methods to provide a method for predicting phenotypic correlations of traits based on genetic similarity calculation using high-throughput genetic loci. This method can overcome the shortcomings of phenotypic measurement methods and provide methodological support for detecting phenotypic correlations between traits in fields such as biological gene function research and breeding applications. Summary of the Invention
[0008] The purpose of this invention is to address the shortcomings of phenotypic measurement in estimating phenotypic similarity by providing a method for predicting phenotypic correlation based on genetic similarity calculated using high-throughput genetic loci, thus offering theoretical guidance for fields such as biological gene function research and breeding applications. The method of this invention includes the following four steps:
[0009] Step 1: Collection of genetic locus data: In the target species, by searching for information in published literature or by QTL or GWAS analysis, the QTL regions that regulate the target trait A or the QTN loci associated with the target trait A are obtained, which constitute the genetic locus information of the target trait A in the target species.
[0010] The QTL information includes at least the trait regulated by the QTL, the linkage group or chromosome number, the physical location of the left boundary, and the physical location of the right boundary; the QTN information includes at least the trait associated with the QTN, the linkage group or chromosome number, and the physical location on the linkage group or chromosome.
[0011] Further, following the same method described in step 1, the genetic locus information of the target trait B in the target species is searched and obtained.
[0012] By repeating step 1, genetic locus information for two or more target traits in the target species can be obtained.
[0013] Step 2, standardization of genetic locus data: The physical location of the same genetic locus is inconsistent across different versions of the maize reference genome. Therefore, genetic locus data from different sources need to be uniformly mapped to the same version of the maize reference genome. For each QTL locus of target trait A in the target species, the linkage group information and the left and right boundaries of the linkage interval of that QTL are uniformly mapped to the same version of the whole-genome sequencing reference sequence of the target species. For each QTN locus of target trait A in the target species, the chromosomal location information and the physical location information of genetic variations within that QTN locus are uniformly mapped to the same version of the whole-genome sequencing reference sequence of the target species. Furthermore, the chromosome number and physical location of each QTN locus are expanded to a certain upstream and downstream physical location interval (e.g., 10kb). The mapped QTL intervals and the expanded QTN intervals form the mapped genetic locus information for target trait A.
[0014] Further, download the genome annotation GTF or GFF file corresponding to the reference genome version of the target species mapped above. Based on the physical location information of the genes labeled as gene features (“genes”) contained in the GTF or GFF file, use random sampling without replacement to extract all gene features that overlap with the genetic locus intervals mapped to the target trait A, resulting in a gene set; for redundant genes in this set, only one is retained, resulting in the genetic locus gene set of the target trait A.
[0015] Further, following the same method described in step 2, the QTL intervals and QTN loci of trait B are organized to obtain the genetic locus set of the target trait B.
[0016] By repeating steps 1 and 2, a set of genes for two or more target traits in the target species can be obtained.
[0017] Step 3: Calculate genetic similarity: Calculating genetic similarity among target traits includes calculating the similarity of gene sets at genetic loci between the target traits and calculating the significance of this similarity. The similarity of gene sets at genetic loci between the target traits and the significance of this similarity can respectively represent the genetic similarity between the target traits and the significance of this similarity.
[0018] The Jaccard coefficient is used to calculate the genetic similarity J value, which represents the similarity of gene sets at genetic loci between target traits (e.g., target trait A and target trait B). The formula is as follows:
[0019] In the formula, A is the set of genetic loci of the target trait A, B is the set of genetic loci of the target trait B, A∩B is the intersection of the genetic loci of the target trait A and the genetic loci of the target trait B, A∪B is the union of the genetic loci of the target trait A and the genetic loci of the target trait B, and Num() is the number of genes in the set; J is the genetic similarity between the target trait A and the target trait B. The calculated J value range is [0,1]. The larger the value, the greater the genetic similarity between the two traits. Furthermore, J=0 represents that the genetic basis between the two traits is completely different, and J=1 represents that the genetic basis between the two traits is completely the same.
[0020] The bootstrap method is used to calculate the significance value of gene set similarity at genetic loci between target traits (e.g., target trait A and target trait B). The calculation formula is:
[0021] In the formula, J real J represents the true value of genetic similarity among target traits. i is the simulated value of the genetic similarity between traits obtained from the i-th simulation, N is the number of simulations in the bootstrap method, and P is the significance of the genetic similarity between the target traits.
[0022] The specific method for calculating the P-value is as follows: For each simulation of target trait A and target trait B, in the i-th simulation, from the gene features included in the genome annotation file used in step 2 of the target species, a random sampling method without replacement is used to extract gene features with the same number of elements as the gene set of the genetic locus of target trait A to form the simulated genetic locus gene set of target trait A; simultaneously, using the same method, gene features with the same number of elements as the gene set of the genetic locus of target trait B are extracted to form the simulated genetic locus gene set of target trait B; the J-value between the simulated genetic locus gene sets of target trait A and target trait B is calculated. i Value; After N simulations, N simulated J values are obtained. The simulated J values between target trait A and target trait B that are greater than J are calculated. real The number of values is divided by the number of simulations N to obtain the significance P-value of the genetic similarity between target trait A and target trait B; the range of the P-value is [0,1], and the smaller the value, the more significant the genetic similarity between the representative traits.
[0023] By repeating steps 1 to 3, the genetic similarity between any two target traits in the target species and the significance of that similarity can be obtained.
[0024] Step 4, Phenotypic Correlation Prediction: Predicting phenotypic correlation between two target traits (e.g., target trait A and target trait B) based on their genetic similarity values and significance values includes: If there is a significant genetic correlation between target trait A and target trait B, then it is predicted that there is a significant phenotypic correlation between target trait A and target trait B; conversely, if there is a small or insignificant genetic correlation between target trait A and target trait B, then it is predicted that there is no significant phenotypic correlation between target trait A and target trait B.
[0025] By repeating steps 1 to 4, the phenotypic correlation between any two target traits in the target species can be predicted. Attached Figure Description
[0026] Figure 1 This is a flowchart illustrating a method for predicting phenotypic correlations of traits based on high-throughput genetic loci calculation of maize traits, provided in an embodiment of the present invention. Figure 2 This describes the distribution of genetic loci on the genome provided in this embodiment of the invention. Figure 3 Venn diagram of the gene sets of genetic loci for target trait A and target trait B provided in this embodiment of the invention; Figure 4 This is a scatter plot of the phenotypic values between plant height and ear height for 347 maize backbone inbred lines provided in this embodiment of the invention. Detailed Implementation
[0027] The following examples are used to describe in detail the specific implementation of the present invention, but are not intended to limit the scope of the present invention.
[0028] Example 1: Using genetic locus similarity to measure the phenotypic correlation between maize plant height and ear height In maize, plant height refers to the vertical distance from the base of the plant (close to the ground) to the highest point of the tassel after pollination, while ear height refers to the vertical distance from the base of the plant (close to the ground) to the main female ear after pollination. Plant height and ear height phenotypes are important agronomic traits describing maize plant morphology and significantly influence lodging resistance, density tolerance, and yield potential. Under normal circumstances, plant height and ear height phenotypes in maize show a strong correlation; higher plant height is often accompanied by higher ear height. This relationship has important guiding significance for breeding practices. To clarify the use and role of the method in this invention for predicting phenotypic correlation using genetic similarity between traits, this invention collects QTN genetic locus data on maize plant height and ear height, performs data standardization, calculates genetic similarity, and then measures the phenotypic correlation between maize plant height and ear height.
[0029] For detailed operating procedures, please refer to [link / reference]. Figure 1 .
[0030] Step 1: Collection of genetic locus data: To obtain QTN genetic loci associated with maize plant height or ear height traits in GWAS studies.
[0031] Currently, many studies have been published analyzing the genetic basis of maize plant height and ear height traits. In 2020, Wang et al., through GWAS analysis of genotype and phenotypic data from 350 maize backbone inbred lines, identified 13 QTN loci significantly associated with maize plant height and 17 QTN loci significantly associated with maize ear height (Wang et al., 2020). In this embodiment of the invention, these QTN loci are used as genetic locus data for plant height and ear height traits.
[0032] Table 1. Genetic locus data for maize plant height and ear height.
[0033] Figure 2 This illustrates the distribution of genetic loci on the genome in embodiments of the present invention. Figure 2 The study showed 13 QTN loci (PH_1 to PH_13) that were significantly associated with maize plant height and 17 QTN loci (EH_1 to EH_17) that were significantly associated with maize ear height.
[0034] Step 2, standardization of genetic locus data: This is used to map the plant height and ear height QTN genetic loci in the embodiments of the present invention onto the latest version of the maize reference genome sequence, and to extract the annotation genes in the upstream and downstream regions of the mapped QTN loci, respectively forming the gene sets of plant height and ear height genetic loci.
[0035] The physical location of the same genetic locus varies across different versions of the maize reference genome, necessitating the mapping of genetic locus data from different sources to the same version of the maize reference genome. Furthermore, high-quality assembled whole-genome sequences not only facilitate the physical location of genetic loci but also improve the completeness and accuracy of genome feature annotation. The maize plant height and ear height genetic locus data published by Wang et al. were based on the maize reference genome version 3 (RefGen_v3). Therefore, this embodiment of the invention selects the latest maize whole-genome sequencing reference sequence (NAM-5.0) with high-quality assembly. For each QTN with plant height trait, the chromosome number and physical location on the 3rd edition reference genome (e.g., QTN numbered PH_1 is located at base 102841618 on chromosome 3 of the 3rd edition reference genome) are first converted to the chromosome number and physical location on the 5th edition reference genome (e.g., QTN numbered PH_1 is converted to base 90115066 on chromosome 3 of the 5th edition reference genome), and then expanded into a physical location interval of 10kb upstream and downstream (e.g., QTN numbered PH_1 is expanded to the interval from base 90105066 to base 90125066 on chromosome 3).
[0036] Download the genome feature annotation GFF file (Zm-B73-REFERENCE-NAM-5.0_Zm00001eb.1.gff3.gz) corresponding to the latest version of the maize whole-genome sequencing reference sequence from the Maize Genetics and Genomes Database (MaizeGDB). Based on the physical location information of the genes (“gene”) in the GFF file, extract the gene numbers for all genes within the expanded genetic locus's physical location interval. Merge the extracted genes from the expanded physical location interval corresponding to each QTN locus for the plant height trait into a set. For this set, retain only one redundant gene, forming the QTN genetic locus gene set for the plant height trait, which includes 13 genes.
[0037] Following the same method, the ear height trait QTN was processed to obtain the gene set of the QTN genetic locus for the ear height trait, which includes 26 genes.
[0038] Step 3: Calculate genetic similarity: This is used to calculate the similarity and significance of gene sets between genetic loci for plant height and ear height in embodiments of the present invention. The specific implementation method consists of the following two steps:
[0039] Step 3.1 Calculate the similarity of gene sets at genetic loci: The Jaccard coefficient is used to calculate the similarity between the gene sets of QTN genetic loci for plant height and ear height, representing the genetic similarity between the two. The specific calculation formula is as follows:
[0040] In the formula, A is the set of genes for the plant height trait QTN, B is the set of genes for the ear height trait QTN, A∩B is the intersection of the sets of genes for the plant height and ear height trait QTN, A∪B is the union of the sets of genes for the plant height and ear height trait QTN, Num() is the number of genes in the set, and J is the genetic similarity between plant height and ear height. The calculated J value should be in the interval [0,1]. The larger the value, the greater the genetic similarity between plant height and ear height. Furthermore, if J=0, it means that the genetic basis of the traits of plant height and ear height is completely dissimilar, and if J=1, it means that the genetic basis of the traits of plant height and ear height is completely identical.
[0041] Figure 3 Venn diagram of the gene sets of genetic loci for target trait A and target trait B provided in the embodiments of the present invention. Figure 3 In this study, target trait A is plant height, and target trait B is ear height. The gene sets for plant height and ear height share 3 genes, resulting in a Num(A∩B) value of 3. Each gene set for plant height and ear height has 10 and 23 unique genes, respectively. Adding the 3 shared genes, the union of the gene sets for plant height and ear height is 36 genes, resulting in a Num(A∪B) value of 36. In this embodiment, the J value is calculated to be 0.0833, indicating a genetic similarity of 0.0833 between maize plant height and ear height.
[0042] Step 3.2 Calculate the significance of genetic similarity:
[0043] The bootstrap method is used to calculate the significance value of gene set similarity at genetic loci between target traits (e.g., target trait A and target trait B). The calculation formula is:
[0044] In the formula, J real J represents the true value of genetic similarity between maize plant height and ear height traits (0.0833). iis the simulated value of genetic similarity between traits obtained from the i-th simulation, N is the number of simulations performed in the bootstrap method; P is the significance of genetic similarity between plant height and ear height traits.
[0045] In a single simulation using the bootstrap method, among the features annotated as genes in the genome feature annotation GFF file included in step 2 of this embodiment, a random sampling method without replacement is used to select the same number of genes (13) as the gene set of the QTN genetic loci for the plant height trait as the genetic locus gene set for the plant height trait. At the same time, a random sampling method is used to select the same number of genes (26) as the gene set of the QTN genetic loci for the ear height trait as the genetic locus gene set for the ear height trait. The Jaccard coefficient between the genetic locus gene set for the plant height trait simulation and the genetic locus gene set for the ear height trait simulation is calculated using the method in step 3.1 of this embodiment.
[0046] One million simulations were performed to obtain the J values for one million simulated genetic loci between plant height and ear height traits. The number of simulated J values greater than the actual J values was calculated, and this ratio was divided by the number of simulations (one million). The resulting ratio represents the significance (P-value) of the genetic similarity between plant height and ear height traits. This significance value should be in the range of [0,1], with a smaller value indicating a higher degree of significance in the genetic similarity between the traits. In this embodiment of the invention, no simulated value was greater than the actual genetic similarity between plant height and ear height, i.e., the P-value was 0, indicating that the genetic similarity between plant height and ear height was extremely significant.
[0047] Step 4, Phenotypic Correlation Prediction: Calculations showed that the genetic locus similarity between plant height and ear height was 0.0833, with a p-value of 0, indicating significant similarity. Therefore, a significant phenotypic correlation is predicted between plant height and ear height.
[0048] In their 2020 study, Wang et al. published plant height and ear height phenotypic values for 350 backbone inbred lines. Among them, 347 backbone inbred lines had paired plant height and ear height phenotypic values, which could be used to calculate the actual correlation between phenotypes and to verify the reliability of the prediction results.
[0049] Figure 4 This is a scatter plot of the phenotypic values between plant height and ear height for 347 maize backbone inbred lines provided in this embodiment of the invention. Figure 4 It can be seen that there is a significant phenotypic correlation between plant height and ear height among these 347 maize backbone inbred lines.
[0050] Calculations showed that the Pearson correlation coefficient between plant height and ear height for these 347 backbone inbred lines was 0.7284, with a P value less than 0.0001, indicating a significant phenotypic correlation between plant height and ear height.
[0051] The results show that both the genetic similarity-based prediction method and the phenotypic measurement method yielded a significant phenotypic correlation between maize plant height and ear height, indicating the feasibility and accuracy of the phenotypic correlation prediction method based on genetic similarity between traits.
[0052] The vast amount of genetic locus data on organismal traits provides a rich data resource for the method of this invention, helping to overcome the shortcomings of phenotypic measurement methods, such as high cost, long cycle, and small data volume. Furthermore, it can be repeatedly used to calculate the genetic similarity between different traits, providing clues to the genetic basis of phenotypic correlations while predicting them. The method of this invention is highly operable, not limited by species or traits, and can provide important phenotypic correlation prediction information for basic research such as gene function studies and genetic basis analysis, as well as practical applications such as variety improvement and breeding.
[0053] References
[0054] Wang B, Lin Z, Li X, Zhao Y, Zhao B, Wu G, Ma X, Wang H, Xie Y, Li Q,Song G, Kong D, Zheng Z, Wei H, Shen R, Wu H, Chen C, Meng Z, Wang T, Li Y,Li Genet.2020, 52(6):565-571.
Claims
1. A method for predicting phenotypic correlations of traits based on genetic similarity calculated from high-throughput genetic loci in maize, characterized in that, It includes the following four steps: Step 1: Collection of genetic locus data; Step 2: Standardize genetic locus data; Step 3: Calculate genetic similarity; Step 4: Phenotypic correlation prediction.
2. The method according to claim 1, characterized in that, Step 1, the collection of genetic locus data, involves obtaining genetic loci on the genome that regulate the formation of target traits in the target species, including but not limited to QTL regions that regulate target traits and QTN loci associated with target traits, which constitute the genetic locus information of each target trait under study in the target species.
3. The method according to claim 1, characterized in that, The genetic locus data standardization described in step 2 involves first selecting a reference genome sequence of the same version as the target species; then mapping the linkage group information and genetic location information of the left and right boundaries of linkage intervals in the QTL loci of the target trait, as well as the chromosomal location information and physical location information of genetic variations in the QTN loci, onto the selected whole genome sequencing sequence of the same version in the target species. Then, according to the General Transfer Format or General Feature Format genome annotation file, i.e., GTF or GFF genome annotation file, the annotated genes in the QTL region and the annotated genes in a certain region upstream and downstream of the QTN locus are taken to form the standardized genetic locus gene set of each target trait in the target species.
4. The method according to claim 1, characterized in that, The genetic similarity calculation in step 3 includes calculating the similarity of gene sets of genetic loci among the target traits to represent the genetic similarity among the target traits, and calculating the significance of the genetic similarity.
5. The similarity of gene sets at genetic loci among the target traits as described in claim 4 is calculated using the Jaccard coefficient method, and the calculation formula is as follows: In the formula, J represents the similarity of the gene sets of genetic loci, i.e., the genetic similarity value between target traits. A is the gene set of genetic loci of target trait A, B is the gene set of genetic loci of target trait B, A∩B is the intersection of the gene sets of genetic loci of target trait A and target trait B, A∪B is the union of the gene sets of genetic loci of target trait A and target trait B, and Num() is the number of genes in the set. The calculated J value range is [0,1]. The larger the value, the greater the genetic similarity between the two traits. Furthermore, J=0 represents that the genetic basis of the two traits is completely different, and J=1 represents that the genetic basis of the two traits is completely the same.
6. The calculation of the significance of the genetic similarity as described in claim 4 uses the bootstrap method, and the calculation formula is as follows: In the formula, the P-value represents the significance of genetic similarity among target traits, and J... real J represents the true value of genetic similarity among target traits. i is the simulated value of genetic similarity between traits obtained from the i-th simulation, and N is the number of simulations in the bootstrap method; The specific method for calculating the P-value is as follows: For N simulations of target trait A and target trait B, in the i-th simulation, from the gene features included in the genome annotation file used for the target species in step 2 of claim 1, a random sampling method without replacement is used to extract gene features with the same number of elements as the gene set of the genetic locus of target trait A to form the simulated genetic locus gene set of target trait A; simultaneously, using the same method, gene features with the same number of elements as the gene set of the genetic locus of target trait B are extracted to form the simulated genetic locus gene set of target trait B; the J-value between the simulated genetic locus gene sets of target trait A and target trait B is calculated. i Value; After N simulations, N simulated J values are obtained. The simulated J values between target trait A and target trait B that are greater than J are calculated. real The number of values is divided by the number of simulations N to obtain the significance P-value of the genetic similarity between target trait A and target trait B; the range of the P-value is [0,1], and the smaller the value, the more significant the genetic similarity between the representative traits.
7. The method according to claim 1, characterized in that, The phenotypic correlation prediction in step 4 is based on the genetic similarity J value between the two target traits obtained using the method of claim 5 and the significance P value obtained using the method of claim 6. The larger the J value and the higher the significance level indicated by the P value, the greater and more significant the phenotypic correlation between the two target traits can be predicted.
8. The method according to claim 1, characterized in that, The methods described in steps 1 to 4 are not limited to corn and can be used for other organisms with high-throughput genetic locus information.