Method for selecting high-efficiency molecular marker for enhancing selection efficiency for target trait and usage thereof
A new algorithm for selecting SNP markers based on variance-weighted genotypes addresses the limitations of existing techniques, enabling efficient selection of molecular markers for target traits and enhancing crop breeding processes.
Patent Information
- Application Number
- PCT/KR2024/096491
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-11
- Filing Date
- 2024-11-13
- Publication Date
- 2025-05-22
AI Technical Summary
Existing molecular marker discovery techniques, such as QTL Mapping and Genome-Wide Association Study (GWAS), are limited in processing large numbers of SNPs from Next Generation Sequencing (NGS) and are time-consuming in selecting molecular markers with high selection efficiency for target traits.
A new algorithm is developed to select SNP markers with high heritability and selection efficiency by assigning weights based on variance in genotype distributions at trait value extremes and the median, using NGS data from crossbreeding populations.
The algorithm significantly shortens the molecular marker development process, allowing for the effective selection of SNP markers with the highest selection efficiency for target traits, improving crop breeding and genetic research efficiency.
Smart Images

Figure KR2024096491_22052025_PF_FP_ABST
Abstract
Description
High-efficiency molecular marker selection method for enhancing selection efficiency for target traits and its use
[0001] The present invention relates to a method for selecting a high-efficiency molecular marker to enhance selection efficiency for a target trait and its use.
[0002] Molecular markers are analytical tools that detect diverse mutations within the genome and are utilized in diverse fields, including biology, genetic research, and trait analysis. Molecular markers can be used to study the relationship between genes and traits and to quickly select individuals with desired traits without morphological selection. As the scope of molecular marker application expands, the rapid and accurate identification of trait-associated genes or molecular markers has become a key tool in crop breeding. It is also essential for improving the selection efficiency of quantitative traits such as crop productivity, quality, resistance, and adaptability. However, the search for molecular markers with high individual selection efficiency remains very limited.
[0003] To date, conventional molecular marker discovery techniques, such as QTL Mapping and Genome-Wide Association Study (GWAS), have been primarily used to explore trait-associated molecular markers. QTL mapping uses crossbreeding populations to identify chromosomal locations or molecular markers associated with traits. Representative methods for selecting trait-associated markers include Joinmap and QTL cartographer. However, these programs are limited to a few thousand molecular markers, making them incapable of processing the hundreds of thousands of SNPs recently acquired through Next-generation sequencing (NGS). GWAS, a method for identifying molecular markers with high trait associations through statistical analysis, is a widely used technique for exploring genetic associations with an unlimited number of molecular markers. However, this method is complex and time-consuming, requiring the selection process for highly selective markers among candidate molecular markers.
[0004] Effective selection and utilization of molecular markers is a key element in crop breeding and individual selection, leading to an increasing demand for molecular markers capable of detecting a wider range of traits. The development of high-throughput sequencing devices, known as next-generation sequencing (NGS), has enabled rapid, comprehensive genome-wide analysis of structural mutations. However, the sheer number of identified structural mutations, ranging from hundreds of thousands to millions, has significantly increased the complexity of molecular marker selection. Therefore, there is a need to develop a molecular marker selection algorithm that can maximize selection efficiency with a minimal number of molecular markers for target traits from the large-scale structural mutations (SNPs, Indels) present throughout the genome, as identified through NGS and other methods.
[0005] Meanwhile, Korean Patent Publication No. 2603207 discloses a 'method for discovering SNP markers associated with phenotypes using statistical regularization and selection probability', but does not describe the 'method for selecting high-efficiency molecular markers for enhancing selection efficiency for target traits and use thereof' of the present invention, which sorts SNPs according to trait values within a population and assigns different weights to the genotypes of individuals distributed at both ends of the trait values and the genotypes of individuals distributed in the middle to select SNP markers highly associated with traits at both ends.
[0006] The present invention was derived from the above-mentioned needs, and the inventors developed a novel algorithm for selecting SNP markers with the highest selection efficiency for a target trait compared to the existing quantitative trait locus (QTL) mapping or GWAS. Specifically, by utilizing NGS technology to obtain genome sequence data from F2 crossbred populations or collected genetic resources, a large number of SNPs existing within the population were searched. Next, by measuring the trait value for the target trait and assigning weights to the SNPs according to the measured trait value, an algorithm was developed to select SNP markers with high heritability and excellent selection efficiency for individuals related to the target trait. Afterwards, experimental verification was performed using the SNP-based KASP primer set selected through the above algorithm. In conclusion, the present invention was completed by confirming that when SNP markers selected through the new algorithm of the present invention are used rather than SNP markers selected through existing QTL mapping or GWAS, individuals for the target trait can be selected more effectively, and that among numerous SNP markers explored through QTL mapping or GWAS, SNP markers closely linked to the target trait and with improved selection efficiency can be selected.
[0007] In order to solve the above problem, the present invention comprises: (1) a step of calculating a score (Sj) of each locus using [Formula 1] by differently assigning weights according to variance based on the genotypes of individuals distributed at both ends of the aligned trait values and the genotypes of individuals having the median value of the target trait for each of at least 10 reference SNPs; and
[0008] [Formula 1]
[0009]
[0010] (2) A method for determining, detecting, selecting or isolating a target trait-associated single nucleotide polymorphism (SNP) marker, including a step of selecting a top Q locus among markers (positions) between 1 and L using [Formula 2] for the score (Sj) of each locus in step (1) above, is provided.
[0011] [Formula 2]
[0012]
[0013] In addition, the present invention provides a method for producing a plant or seed population, comprising: a step of selecting a plant or seed population according to a method of determining, detecting, selecting, or isolating the target trait-associated SNP marker; and a step of crossing or selfing one or more selected plants or plants grown from the selected seeds.
[0014] In addition, the present invention provides a method for introducing or increasing expression of a target trait-related SNP in a subject, comprising a step of determining, detecting, selecting or isolating a target trait-related SNP marker according to the method for determining, detecting, selecting or isolating the target trait-related SNP marker, and a step of expressing or increasing expression of the target trait-related SNP in the subject.
[0015] In addition, the present invention provides a method for producing a plant having a target trait, comprising a step of determining, detecting, selecting or isolating a target trait-related SNP marker according to a method for determining, detecting, selecting or isolating a target trait-related SNP marker, and a step of expressing or increasing the expression of the target trait-related SNP in a subject.
[0016] In addition, the present invention provides a plant having a target trait produced by a method for introducing or increasing the expression of a SNP related to a target trait in the subject or a method for producing a plant having a target trait.
[0017] Conventional molecular marker discovery technologies, such as QTL mapping or GWAS, have been used to discover molecular markers associated with target traits. However, they have limitations in processing the hundreds of thousands of SNPs recently acquired through NSG, and the process of selecting molecular markers with high selection efficiency among candidate molecular markers is complex and time-consuming. Using the algorithm of the present invention, the molecular marker development process can be significantly shortened, enabling the effective selection of molecular markers with the highest selection efficiency for target traits quickly and effectively. Therefore, it can be utilized in various fields such as trait research, genetic research, and crop breeding.
[0018] Figure 1 shows the results of QTL mapping for ripening traits in pepper fruits. Red arrow: SNP with the highest LOD (Logarithm of the odds) on chromosome 10.
[0019] Figure 2 shows the results of a GWAS using the GAPIT (Genome association and prediction integrated tool) program on ripening traits in pepper fruits. Red boxes: SNPs showing high LOD on chromosomes 10 and 6.
[0020] Figure 3 shows the results of KASP analysis performed on parental lines, their F2 population, and fast-maturing individuals corresponding to the comparison group using a primer set for KASP based on the candidate SNP marker 'C10674K' on chromosome 10 associated with ripening traits of pepper fruits.
[0021] Figure 4 shows the results of KASP analysis performed on parental lines, their F2 population, and fast-maturing individuals corresponding to the comparison group using a primer set for KASP based on the candidate SNP marker 'C06233B' on chromosome 6 associated with ripening traits of pepper fruits.
[0022] Figure 5 is a distribution of SNP marker selection values associated with pepper fruit ripeness analyzed using the algorithm of the present invention. Orange indicates SNP markers highly correlated with pepper fruit coloring traits selected using the algorithm of the present invention, and green indicates the portion of the SNP markers that passed the best fit line and exhibited the greatest variation.
[0023] Figure 6 shows the results of KASP analysis performed on the parent lines, their F2 population, and fast-maturing individuals corresponding to the comparison group using a primer set for KASP based on the SNP marker 'C09222M' associated with the ripeness of pepper fruits selected through the algorithm of the present invention.
[0024] Figure 7 shows the results of QTL mapping for traits related to resistance to chlorosis. Red arrow: SNP with the highest LOD on chromosome 7.
[0025] Figure 8 shows the results of GWAS performed on traits related to resistance to chlorosis using the GAPIT and TASEEL (Trait Analysis by aSSociation) programs. Red box: SNP showing the highest LOD on chromosome 7.
[0026] Figure 9 is a distribution diagram of SNP marker selection values associated with resistance to yellowing disease analyzed using the algorithm of the present invention.
[0027] Figure 10 is a chromosomal distribution diagram of SNP markers associated with resistance to yellowing disease selected using the algorithm of the present invention.
[0028] Figure 11 shows the results of KASP analysis performed on parent lines (mother line, resistant (S); father line, susceptible (R)) and recombinant inbred lines (RILs) using a primer set for KASP based on the SNP marker 'R0712812M' associated with resistance to yellowing disease selected through the algorithm of the present invention.
[0029] Figure 12 shows the PCR results confirming the presence or absence of Pun1, a pungency-related gene, in the chili pepper genetic resource. Lane 1: C25, lane 2: C26, lane 3: C27, lane 4: C28, lane 5: C29.
[0030] Figure 13 shows the results of a GWAS using the GAPIT program for traits related to the pungency of chili peppers. Red boxes: SNPs showing high LOD on chromosomes 2, 9, and 10.
[0031] Figure 14 is a distribution diagram of SNP marker selection values associated with the spiciness of peppers analyzed using the algorithm of the present invention.
[0032] Figure 15 is a chromosomal distribution diagram of SNP markers associated with chili pepper spiciness selected using the algorithm of the present invention.
[0033] In order to achieve the purpose of the present invention, the present invention comprises: (1) a step of calculating a score (Sj) of each locus using [Formula 1] by differently assigning weights according to variance based on the genotypes of individuals distributed at both ends of the aligned trait values and the genotypes of individuals having the median value of the target trait for each of at least 10 reference SNPs; and
[0034] [Formula 1]
[0035]
[0036] (2) A method for determining, detecting, selecting or isolating a target trait-associated single nucleotide polymorphism (SNP) marker, including a step of selecting a top Q locus among markers (positions) between 1 and L using [Formula 2] for the score (Sj) of each locus in step (1) above, is provided.
[0037] [Formula 2]
[0038]
[0039] In a method according to one embodiment of the present invention, Q may be an integer from 1 to 100, preferably an integer from 1 to 10, but is not limited thereto.
[0040] In a method according to one embodiment of the present invention, the method may further include a step of analyzing the genome sequence information of each individual or offspring group obtained by mating the paternal and maternal lines to obtain a reference SNP.
[0041] In a method according to one embodiment of the present invention, the genotype of the individual can be collected by analyzing the genome sequence information of each individual or offspring group obtained by mating the paternal and maternal lines to collect SNPs.
[0042] In a method according to one embodiment of the present invention, the method may additionally include, but is not limited to, a step of measuring a trait value for a target trait for each individual or offspring group obtained by crossing the paternal and maternal lines, and listing SNPs in order of the trait values.
[0043] In a method according to one embodiment of the present invention, the trait value is measured by a target trait for each individual or offspring group obtained by crossing the paternal and maternal lines, and SNPs can be listed in the order of the trait value.
[0044] In a method according to one embodiment of the present invention, the SNP of step (1) may be selected as a SNP having a minor allele frequency (MAF) of 5% or more, a missing data of 30% or less, and a homozygous type among polymorphic SNPs between parental lines, but is not limited thereto.
[0045] In a method according to one embodiment of the present invention, the target trait-associated SNP marker may be selected by repeatedly performing steps (1) and (2) on individuals obtained by sequentially applying a significance probability (p-value) range of 5% or less, 10% or less, and 15% or less to both ends of the aligned trait values of step (1), but is not limited thereto.
[0046] In a method according to one embodiment of the present invention, the target trait-associated SNP marker can be selected from a subject. The subject may include one or more selected from the group consisting of seeds, plants, animals, bacteria, and insects, and preferably includes plants or seeds, but is not limited thereto.
[0047] Additionally, in a method according to one embodiment of the present invention, the target trait-associated SNP marker can be selected from a sample. The sample may include one or more selected from the group consisting of seeds, plants, animals, bacteria, and insects, and preferably includes plants or seeds, but is not limited thereto.
[0048] In a method according to one embodiment of the present invention, the target trait may be, but is not limited to, a trait including ripening of pepper fruits, a trait including resistance to radish chlorosis, or a trait including a pungency gene (Pun1) of pepper. In addition, the target trait may be, but is not limited to, a trait including a resistance gene R1, R8, or Rpi-amr1 against a Phytophthora pathogen. In addition, the target trait may be, but is not limited to, a trait including one or more selected from the group consisting of herbicide resistance, disease resistance, insect or pest resistance, fungal disease resistance, virus resistance, nematode resistance, bacterial disease resistance, fatty acid biosynthesis, starch biosynthesis, increased grain yield, increased oil, improved nutrition, increased growth rate, fruit ripening, increased yield, improved stress tolerance, improved environmental or chemical tolerance, morphological characteristic changes, improved digestibility, industrial enzyme production, improved flavor, nitrogen fixation, and any combination thereof.
[0049] In a method according to one embodiment of the present invention, the method can show better target trait selection efficiency compared to a GWAS (Genome-Wide Association Study) program and / or QTL mapping.
[0050] In a method according to one embodiment of the present invention, the method specifically comprises:
[0051] (1) A step of obtaining SNP (single nucleotide polymorphism) by analyzing the genome sequence information of each individual or offspring group obtained by mating the paternal and maternal lines;
[0052] (2) A step of measuring the trait value for the target trait for each individual or offspring group obtained by crossing the paternal and maternal lines, and listing SNPs in the order of the trait value;
[0053] (3) A step of calculating the score (Sj) of each locus using [Formula 1] below by differently assigning weights according to variance based on the genotypes of individuals distributed at both ends of the above-mentioned sorted trait values and the genotypes of individuals having the median value of the target trait; and
[0054] [Formula 1]
[0055]
[0056] (4) The step of selecting the top 10 loci among markers (positions) between 1 and L using [Formula 3] for the score (Sj) of each locus in step (3) may be included, but is not limited thereto.
[0057] [Formula 3]
[0058]
[0059] In a method according to one embodiment of the present invention, the sequence information analysis of step (1) may be performed through next-generation sequencing (NGS), but is not limited thereto.
[0060] In a method according to one embodiment of the present invention, the target trait of step (2) may be, but is not limited to, an agricultural trait of a crop (e.g., plant height, fruit color, maturity period, taste, etc.) or a trait such as disease resistance.
[0061] In a method according to one embodiment of the present invention, the trait values in step (2) may be sorted in ascending or descending order, but are not limited thereto.
[0062] In a method according to one embodiment of the present invention, the SNP of step (1) may be selected as a SNP having a minor allele frequency (MAF) of 5% or more, a missing data of 30% or less, and a homozygous SNP among polymorphic SNPs between parental lines, but is not limited thereto.
[0063] In addition, step (2) above includes a process of aligning the SNPs of each individual according to the trait value. The reason why this process is necessary is because it is difficult to determine the correlation between the trait and the sample by looking only at the SNP results without the trait value. The measurement method for the trait value for the target trait varies depending on the type of target trait. In the case of crops, it can be performed according to the National Institute of Agrobiodiversity's Characteristics Survey Standards (https: / / www.seed.go.kr / seed / 192 / subview.do), but is not particularly limited thereto.
[0064] In a method according to one embodiment of the present invention, the two ends of step (3) refer to individuals with opposite characteristics for a specific trait. For example, when fruit ripeness is set as a target trait and the trait values of the offspring obtained by crossing the mother and father plants are measured and sorted in descending or ascending order, there are early-ripening and late-ripening samples with opposite characteristics in the two ends. Accordingly, in the present invention, weights are differently assigned according to variance based on the genotypes of the individuals distributed in the two ends and the genotypes of the individuals with the median value of the trait, so that high values are assigned to SNPs that can effectively select individuals in the two ends.
[0065] The target trait-associated SNP marker selected through the method according to one embodiment of the present invention is selected by repeatedly performing steps (3) and (4) on the obtained two-terminal individuals by sequentially applying the ranges of significance probability (p-value) of 5% or less, 10% or less, and 15% or less to the two-terminals of the aligned trait values of step (3).
[0066] Specifically, when conducting genetic analysis, more than 96 samples are used for a specific population to ensure that various groups are evenly represented in the entire population, and most analyses are performed in multiples of 96 samples (192, 288, …).
[0067] Accordingly, in the present invention, the first calculation was performed with 5% as the standard for statistical significance in 96 samples (approximately 100 samples) that analyze the group. This is because as the number of samples increases (the larger the group), it follows a normal distribution, so the calculation was performed based on samples with trait values with a significance probability (p-value) of 5% or less. However, if we proceed only with the standard of p-value 0.05 (5%) as above, the samples in the middle distribution (samples with middle trait values) are not taken into consideration, so the results will be biased toward the samples at both ends, and the curse of dimensionality (a phenomenon in which judgment is impossible due to numerous candidate markers when selecting a marker that can explain one trait with a small number of samples) may occur.
[0068] For this reason, in the step (3) above, in the middle of the sorted trait values, different weights are given to the genotypes of individuals having trait values with a significance probability (p-value) of 5% or less and the genotypes of individuals distributed in the middle, and SNPs are selected using the above [Formula 1] and the above [Formula 3] of the step (4) above, and further, in the step (3), in the middle of the sorted trait values, different weights are given to the genotypes of individuals having trait values with a significance probability of 10% or less and the genotypes of individuals distributed in the middle, and SNPs are selected through the same process as above, and further, in the step (3), in the middle of the sorted trait values, different weights are given to the genotypes of individuals having trait values with a significance probability of 15% or less and the genotypes of individuals distributed in the middle, and SNPs are selected through the same process as above. When steps (3) and (4) were repeated by applying different significance probabilities to both ends of the trait value as described above, if SNPs in the same LD (Linkage Disequilibrium) block were repeatedly selected, it was judged that the SNPs selected through this process were more likely to be associated with the actual trait value.
[0069] LD is analyzed by calculating recombination using SNP markers that move with genes, identifying related information, and predicting the relative distances between genes. LD blocks are regions of the genome that are so short or densely packed that crossover rarely occurs across generations. Therefore, the genetic information within these regions is nearly identical and is largely preserved across generations.
[0070] The present invention also comprises a step of selecting a population of plants or seeds according to a method of determining, detecting, selecting or isolating the target trait-associated SNP marker; and
[0071] A method for producing a plant or seed population is provided, comprising the step of crossbreeding or self-fertilizing one or more selected plants, or plants grown from selected seeds.
[0072] In a method for producing a plant or seed population according to one embodiment of the present invention, the selected population may have a target trait of improved ripening of pepper fruits, a target trait of improved resistance to radish chlorosis, or a target trait of improved pungency gene (Pun1) of pepper, but is not limited thereto. In addition, the selected population may have a target trait of improved resistance genes R1, R8, or Rpi-amr1 against Phytophthora pathogens, but is not limited thereto.
[0073] The present invention also provides a method for determining, detecting, selecting or isolating a target trait-associated SNP marker, comprising: a step of determining, detecting, selecting or isolating a target trait-associated SNP marker; and
[0074] A method for introducing or increasing expression of a target trait-related SNP in a subject is provided, comprising a step of expressing or increasing expression of the target trait-related SNP in the subject.
[0075] The present invention also provides a method for determining, detecting, selecting or isolating a target trait-associated SNP marker, comprising: a step of determining, detecting, selecting or isolating a target trait-associated SNP marker; and
[0076] A method for producing a plant having a target trait is provided, comprising a step of expressing or increasing the expression of a SNP related to the target trait in a target organism.
[0077] In a method for introducing or increasing the expression of a SNP related to a target trait in a subject according to one embodiment of the present invention or a method for producing a plant having a target trait, the subject is a plant, and the method may further include, but is not limited to, crossing the plant with another plant of the same species to produce a progeny plant having the target trait.
[0078] The present invention also provides a plant having a target trait produced by a method for introducing or increasing the expression of a target trait-related SNP in a subject according to one embodiment of the present invention or a method for producing a plant having a target trait.
[0079]
[0080] Hereinafter, the present invention will be described in detail by way of examples. However, the following examples are merely illustrative of the present invention and the scope of the present invention is not limited thereto.
[0081]
[0082] Example 1. Selection of markers related to ripening of pepper fruits.
[0083] In the first case, we conducted research on ripening traits in pepper fruit, developing molecular markers related to coloration using an F2 population of pepper. In this case, the present invention excluded markers commonly identified as Type I errors through QTL mapping or GWAS, and discovered highly efficient molecular markers corresponding to Type II errors.
[0084]
[0085] 1-1. Analysis materials
[0086] After breeding F2 populations using parental lines that showed differences in fruit ripening time, the ripening time of the parental lines and F2 populations was measured. Young leaves from each sample were collected, genomic DNA was extracted, and used as GBS (Genotyping-by-sequencing) samples.
[0087]
[0088] 1-2. Bioinformatics Analysis and SNP Statistical Analysis
[0089] A total of 98 lines including the parent lines were sampled, young leaves of each sample were collected, genomic DNA was extracted, digested with the restriction enzyme ApeKI, and a GBS library was created, and the base sequence was produced using Illumina's Hiseq X10. Using the produced base sequence, SNPs were selected through the bioinformatics pipeline in Seeders. SNPs that satisfied the conditions of minor allele frequency (MAF) of 5% or more and missing data of 30% or less were selected, and the chromosomal location in the reference genome (https: / solgenomics.net / organism / Capsicum_annuum / genome) and the presence of polymorphic SNPs between the parents were checked, and if necessary, additional filtering processes were performed and used in the next analysis.
[0090] Summary of SNP filter process Filter step Filter item SNP matrix loci 1 Total SNP matrix 1 2,550,758 2 MAF (minor allele frequency) >5% * 1 4,376,3513MAF >20%, Chromosome level left 4,153,7063Missing data <30% * 2116,7294 MAF >5%, Missing data <30% 14,0575 Parental polymorphic SNP homo-type locus * 3 9,9816 Chromosome level left 9,6987 MAF >20%, missing data <30%, parental polymorphic SNP homo-type left 8,3678 SNP selection for linkage map creation * 4 2,250
[0091] *1: Select SNPs with a MAF greater than 5% of the entire sample for the corresponding locus. *2: Select SNPs with missing data less than 30% of the entire sample for the corresponding locus.
[0092] *3: Select homotype SNPs among polymorphic SNPs between the two parental samples.
[0093] *4: SNPs are selected based on the maximum left limit of the Joinmap program input.
[0094]
[0095] 1-3. Selection of molecular markers related to pigmentation through QTL mapping
[0096] Using the Joinmap program for the 2,250 SNP loci finally selected, 2,242 SNP loci were used to create a linkage map containing 12 linkage groups. Using the QTL cartographer program, the QTL loci showing the highest LOD (Logarithm of the odds) in linkage group 10 were identified (Fig. 1), and in addition, QTL loci showing a standard LOD value of 4.3 or higher were detected on chromosomes 4 and 9.
[0097]
[0098] 1-4. Molecular marker selection through GWAS using the GAPIT (Genome association and prediction integrated tool) program
[0099] Due to limitations in currently developed programs, the number of SNP markers actually used to create association maps is limited to 2,000-3,000. However, GWAS has the advantage of overcoming marker limitations and allowing association analysis with a sufficient number of SNPs. In the present invention, we used 8,367 refined SNP loci and GAPIT (applying eight models), the most commonly used GWAS program, to conduct an association analysis on pepper fruit ripeness.
[0100] As a result, the most significant signal was confirmed at the C10674K marker on chromosome 10 in all GWAS models, followed by the C06233B marker on chromosome 6. Accordingly, the C10674K and C06233B markers were selected as candidate markers (Fig. 2).
[0101]
[0102] 1-5. Candidate molecular marker screening using a primer set for KASP
[0103] To measure the selection effect of individuals using the candidate SNP markers selected through the above QTL mapping and GWAS using GAPIT, a primer set for KASP based on candidate SNPs was created and the selection efficiency was confirmed targeting individuals for which trait information was already known.
[0104] First, using the primer set for KASP based on the C10674K marker selected on chromosome 10, which commonly showed the highest signal through QTL mapping and GWAS, KASP analysis was performed on parental lines with different fruit ripening times, their F2 populations, and fast-maturing individuals (control group). As a result, the genotype of the late-maturing parental line was B, and the genotype of the relatively fast-maturing parental line was A (Fig. 3). However, the control group (fast-maturing individuals) showed genotype B, which did not match the genotype of the parental line.
[0105] In addition, the median value of the trait in the F2 population of Fig. 3 is 86, so if the trait value is less than 86, it should have base A, and if it is 86 or more, it should have base B. However, as shown in Fig. 3, among the F2 samples with a trait value less than 86, only 1 out of 7 samples had base A, and among the F2 samples with a trait value of 86 or more, only 2 out of 5 samples had base B, showing a tendency not to be suitable for use as a marker for the F2 population. Through this, it was found that the C10674K marker selected on chromosome 10 is not suitable for selecting pepper individuals with different ripening times.
[0106] In addition, KASP analysis was performed on the same samples as above using a primer set for KASP based on the C06233B marker selected on chromosome 6 by GWAS. As a result, the genotype of the late-maturing parent was B, and the genotype of the relatively early-maturing parent was A. However, the comparison group (early-maturing individuals) showed genotype B, confirming that it did not match the genotype of the parent (Fig. 4). This shows that the C06233B marker is also not suitable for selecting pepper individuals with different maturation times.
[0107]
[0108] 1-6. Development of a new algorithm for selecting SNP markers related to pepper fruit coloration.
[0109] Through the above results, it was found that the selection efficiency of SNP markers selected through QTL mapping and GWAS was poor, and a new algorithm was developed to secure markers with enhanced selection efficiency.
[0110] In the first step, SNPs were obtained by analyzing the sequence information of the offspring population or collected resources obtained by mating the paternal and maternal lines.
[0111] In the second step, the trait values of each sample for the target trait were measured for the offspring population or collection resource obtained by crossing the paternal and maternal lines, sorted in ascending or descending order, and then the SNPs of each individual were listed in order of the trait values.
[0112] In the third step, the genotypes of individuals distributed in the two ends showing different trait characteristics from the above-mentioned sorted trait values and the genotypes of individuals with the median trait value were given different weights according to the variance, so that SNPs that effectively select individuals in the two ends were given higher values (Formula 1). In other words, the score (Sj) for each loci was calculated by giving weights for the additive effect.
[0113] [Formula 1]
[0114]
[0115] In the fourth step, the top 10 positions among markers (positions) between 1 and L for Sj were selected (Formula 3).
[0116] [Formula 3]
[0117]
[0118] By applying the coloring-related traits of pepper fruits to the above algorithm, 10 markers belonging to the upper group on the graph were selected and marked in orange, as shown in Fig. 5, and information on the top 5 markers among the 10 selected markers is summarized in Table 2 below.
[0119] Information on the top 5 markers selected using the algorithm of the present invention Rank Marker name Chromosome Selection index Best fit line 1C09222M9 No. 2465 Pass 2C09238A9 No. 2389 Pass 3C09238B9 No. 2367 Pass 4C09216A9 No. 2308 Pass 5C09216B9 No. 2306 Pass
[0120]
[0121] As shown in Table 2 above, the SNP markers selected through the algorithm of the present invention were confirmed to have the highest selection efficiency on chromosome 9, which is different from those identified through existing QTL mapping and GWAS. Molecular markers passing through the best fit line can be candidate markers that can be used as selection markers, and these were mainly clustered on chromosome 9. Among them, the marker with the highest value was named C09222M. The markers selected in the upper group were found to be clustered in the LD block on the chromosome, showing results very consistent with existing theory.
[0122] Examining the QTL mapping results, chromosome 9 showed the second highest signal, followed by chromosome 10. Although the signals were on the same chromosome, the physical locations of the SNPs in the genome showed a difference of several Mbp. However, the C10674K marker on chromosome 10 and the C06233B marker on chromosome 6, which were selected through QTL mapping or GWAS, showed very low values when analyzed using the algorithm of the present invention and were therefore excluded from the candidate group.
[0123]
[0124] 1-7. Verification of the C09222M marker selected through the algorithm of the present invention
[0125] Using a primer set for KASP based on the C09222M marker selected through the algorithm of the present invention, KASP analysis was performed on parent lines with different ripening periods of pepper fruits, their F2 populations, and individuals with early ripening periods (control group).
[0126] As a result, the genotype of the late-maturing parent was B, and the genotype of the relatively early-maturing parent was A, showing good polymorphism in the parental line, and the comparison group also showed genotype A, confirming that it was consistent with the genotype of the parent (Fig. 6).
[0127] In addition, the median value of the trait in the F2 population of Fig. 6 is 86, so if the trait value is less than 86, it must have base A, and if it is 86 or more, it must have base B. As shown in Fig. 6, among the F2 samples with trait values less than 86, 3 out of 6 samples had base A, and among the F2 samples with trait values greater than 86, 5 out of 6 samples had base B, indicating that they are suitable for use as markers for the F2 population.
[0128] Through this, it was found that the C09222M marker selected through the algorithm of the present invention can effectively select pepper plants with different ripening times.
[0129]
[0130] Example 2. Selection of markers related to resistance to yellowing disease
[0131] As a second case, we conducted a study on traits related to resistance to chlorosis of radish, and developed molecular markers associated with chlorosis resistance using parental lines (mother, resistant; father, susceptible) and recombinant inbred lines (RILs) of radish.
[0132]
[0133] 2-1. Analysis materials
[0134] A RIL population was bred using parental lines showing susceptibility or resistance to chlorosis. 170 RIL lines were secured and subjected to disease inoculation tests and genotyping-by-sequencing (GBS) genotyping. Young leaves from each sample were collected, genomic DNA was extracted, and used as GBS (Genotyping-by-sequencing) samples.
[0135]
[0136] 2-2. Bioinformatics Analysis and SNP Statistical Analysis
[0137] Genomic DNA was extracted from young leaves of each sample for the parental lines (2 samples) and RIL 170 lines (170 samples), digested with the restriction enzyme ApeKI, and a GBS library was constructed, which was sequenced using Illumina's Hiseq X10. Using the sequences produced, SNPs were selected through the bioinformatics pipeline within Seeders. SNPs that satisfied the conditions of MAF of 5% or more and missing data of 30% or less were selected, and the chromosomal location within the reference genome and the presence of polymorphic SNPs between the parents were checked, and if necessary, additional filtering processes were performed to use them in the next analysis.
[0138] Summary of SNP filter process Filter step Filter item No. of SNPs 1 Total SNP 7 24,590 2 MAF (minor allele frequency) >5% *1 225,1243Missing data <30% *2 171,0194Missing data <30% & MAF >5%48,0415SNP in chromosome *3 45,8427 Parental Sample Polymorphic Homotype Locus *488,8428MAF (minor allele frequency) >20%69,3379Missing data <10%20,52610Missing data <10% & MAF >20%17,34211Random selection1,576
[0139] *1: Select SNPs with MAF greater than 5% from the entire sample of the corresponding left.
[0140] *2: Select SNPs with less than 30% missing data in the entire sample on the left.
[0141] *3: Select SNP loci secured at the chromosome level of the standard genome.
[0142] *4: Select homotype SNPs among polymorphic SNPs between the two parental samples.
[0143]
[0144] 2-3. Selection of molecular markers related to resistance to yellowing disease through QTL mapping
[0145] Using the Joinmap program, 1,537 SNP loci were used to create an association map comprising 9 linkage groups, utilizing the 1,576 SNP loci finally selected. The QTL cartographer program was used to identify the QTL loci showing the highest LOD in association group 7 (Fig. 7).
[0146]
[0147] 2-4. Molecular marker selection through GWAS using GAPIT and TASEEL (Trait Analysis by aSSociation) programs
[0148] We conducted an association analysis of chlorosis resistance using 45,842 single nucleotide polymorphisms (SNPs) using the most commonly used GWAS programs, GAPIT (using the MLMM model) and TASSEL. Both programs revealed the most significant signal on chromosome 7 (Fig. 8).
[0149]
[0150] 2-5. Selection of markers related to resistance to yellowing disease using the algorithm of the present invention
[0151] Using the algorithm of the present invention to select SNP markers, we confirmed the presence of the R0712812M marker, which has a high selection efficiency, on chromosome 7. Molecular markers passing the best fit line can be candidate markers for selection (Figure 9), and all of these were clustered on chromosome 7 (Figure 10). Information on the top five markers among the ten markers selected using the algorithm of the present invention is summarized in Table 4 below.
[0152] Information on the top 5 markers selected using the algorithm of the present invention Rank Marker name Chromosome Selection index Best fit line 1 R0712812M 7th 23567 Pass 2 R0712812A 7th 23561 Pass 3 R0712812B 7th 23543 Pass 4 R071289 4 A 7th 23538 Pass 5 R0712812C 7th 23537 Pass
[0153]
[0154] The marker candidates selected using the algorithm of the present invention were identical to those selected through QTL mapping and GWAS, and were clustered in LD blocks, demonstrating strong agreement with existing theory. Therefore, the use of the algorithm of the present invention validated that markers selected through QTL mapping or GWAS exhibit superior selection efficiency.
[0155]
[0156] 2-6. Verification of the R0712812M marker selected through the algorithm of the present invention
[0157] Using a primer set for KASP based on the R0712812M marker selected through the algorithm of the present invention, KASP analysis was performed on the parental lines and RIL populations showing susceptibility or resistance to chlorosis.
[0158] As a result, it was confirmed that the resistant parental and RIL 3 lines and the susceptible parental and RIL 2 lines were accurately distinguished (Fig. 11). Through this, it was confirmed that the R0712812M marker selected through the algorithm of the present invention can effectively select radish plants resistant to chlorosis.
[0159]
[0160] Example 3. Selection of a marker related to the spiciness of red peppers.
[0161] In the third case, we verified whether the location of the previously reported Pun1 gene associated with the pungency of pepper could be accurately located using the resequencing results of 111 pepper genetic resources. The pepper genetic resources used this time were not F2 populations composed of the offspring of specific parents as in Examples 1 and 2, but independently bred lines with a more complex genome composition. We tested whether the algorithm of the present invention could be applied to such an independent genetic population. The target trait was Pun1 (located on chromosome 2, Indel marker), a previously identified pungency-related gene. After confirming the presence or absence of the gene through PCR, the PCR results were recognized as phenotypes, and whether the selected molecular marker was appropriately selected by applying the algorithm of the present invention.
[0162]
[0163] 3-1. Analysis materials
[0164] 111 lines of common pepper type and paprika type used as breeding lines were used as samples.
[0165]
[0166] 3-2. Bioinformatics Analysis and SNP Statistical Analysis
[0167] Using the Illumina platform (NovaSeq 6000), we generated genome sequences with 20x coverage for 111 strains of common pepper and paprika types. The generated genome sequences were preprocessed and mapped using the BWA program, and single nucleotide polymorphisms (SNPs) were selected using the Deepvariant program. Using the selected SNPs, we created an SNP matrix for all 111 strains, and refined the SNPs based on MAF and missing data, securing a total of 11,242,804 SNPs (Table 5).
[0168] Summary of SNP filter process Filter step Filter item SNP matrix loci 1 Total SNP matrix 4 3,671,923 2 MAF (minor allele frequency) >20% *1 12,970,3433Missing data <5% *2 37,109,4234MAF >5% and Missing data <30%11,242,804
[0169] *1: Select SNPs with MAF greater than 20% from the entire sample of the corresponding left.
[0170] *2: Select SNPs with less than 5% missing data in the entire sample on the left.
[0171]
[0172] 3-3. Trait investigation
[0173] The presence or absence of the Pun1 gene associated with pungency was used as a target trait. To confirm the presence or absence of the Pun1 gene in 111 pepper genetic resources, the PCR results using the Pun1 gene-specific PCR primers disclosed in the paper of Hai Thi Hong Truong et al. (Horticulture Environment Biotechnology, 2009, 50(4), 358-365) were organized and used as phenotypic values. The results of NCBI BLAST using the reference genome sequence information (https: / solgenomics.net / organism / Capsicum_annuum / genome) confirmed that it was located at 150 Mb on chromosome 2 (Fig. 12), which was used as the basis for judging the accuracy of future analyses.
[0174]
[0175] 3-4. Molecular marker selection through GWAS using the GAPIT program
[0176] Using 11,242,804 single nucleotide polymorphisms (SNPs), we conducted an association analysis on chili pepper spiciness using GAPIT (applied to eight models), the most widely used GWAS program. As shown in Figure 13, most models showed significant marker signals not only on chromosome 2 but also commonly on chromosomes 9 and 10, prompting further experiments.
[0177] Results showed significant differences depending on the model used. For example, the BLINK model showed the strongest signal on chromosome 2, while the FarmCPU model showed the strongest signal on chromosome 10. Most models generated a large number of signals, making it difficult to distinguish between noise and true signals. In situations where the precise location of the gene in question is unknown, this phenomenon requires significant time and effort for additional verification experiments.
[0178]
[0179] 3-5. Selection of markers related to the spiciness of red pepper using the algorithm of the present invention
[0180] The SNP markers selected using the algorithm of the present invention showed the strongest signal on chromosome 2 among the various chromosomes identified through GWAS. Consistent with the results confirmed by gel electrophoresis using the UniversalPun1 primer, the highest value was observed in the 150 Mb region of chromosome 2. Following chromosome 2, chromosome 10 showed the next highest value.
[0181] Molecular markers that pass the best-fit line can be used as candidate markers for selection (Fig. 14), and these were primarily clustered on chromosome 2 (Fig. 15). The clustering of the selected markers within LD blocks demonstrates strong agreement with existing theory, demonstrating that the algorithm of the present invention can select markers with superior selection efficiency among those selected through GWAS.
[0182] Information on the top 5 markers among the 10 markers selected through the algorithm of the present invention is summarized in Table 6 below.
[0183] Information on the top 5 markers selected using the algorithm of the present invention Rank Marker name Chromosome Selection index Best fit line 1C02200A2 No. 11.14 Pass 2C02200B2 No. 11.09 Pass 3C02200C2 No. 11.09 Pass 4C02200D2 No. 11.07 Pass 5C02200E2 No. 11.07 Pass
[0184]
[0185] The examples of Examples 1 to 3 above demonstrate that the algorithm of the present invention provides excellent efficiency in the process of selecting markers associated with various traits, and can successfully contribute to discovering and utilizing excellent candidate markers compared to existing methods.
Claims
1. (1) For each of at least 10 reference SNPs, a weight is assigned differently according to the variance based on the genotypes of the individuals distributed at both ends of the aligned trait values and the genotypes of the individuals with the median value of the target trait, and the score (S) of each locus is calculated using [Formula 1]. j ) for calculating; and [Formula 1] (2) The score (S) of each locus in step (1) above j ) using [Formula 2] to select the upper Q locus among markers (positions) between 1 and L; A method for determining, detecting, selecting or isolating a single nucleotide polymorphism (SNP) marker associated with a target trait. [Formula 2] 2. A method according to claim 1, wherein Q is an integer from 1 to 10.
3. A method according to claim 1 or 2, characterized in that the method further comprises a step of obtaining a reference SNP by analyzing the genome sequence information of each individual or offspring group obtained by mating the paternal and maternal lines.
4. A method according to any one of claims 1 to 3, characterized in that the genotype of the individual is collected by analyzing the genome sequence information of each individual or offspring group obtained by mating the paternal and maternal lines and collecting SNPs.
5. A method according to any one of claims 1 to 4, characterized in that the method further comprises a step of measuring a trait value for a target trait for each individual or group of offspring obtained by crossing the paternal and maternal lines, and listing SNPs in order of the trait values.
6. A method according to any one of claims 1 to 5, characterized in that the trait value is measured by a target trait for each individual or offspring group obtained by crossing the paternal and maternal lines, and SNPs are listed in order of the trait value.
7. A method according to any one of claims 1 to 6, wherein the SNP in step (1) has a minor allele frequency (MAF) of 5% or more, a missing data of 30% or less, and is selected as a homozygous SNP among polymorphic SNPs between parental lines.
8. A method according to any one of claims 1 to 7, characterized in that the target trait-associated SNP marker is selected by repeatedly performing steps (1) and (2) on individuals obtained by sequentially applying a significance probability (p-value) range of 5% or less, 10% or less, and 15% or less to both ends of the aligned trait values of step (1).
9. A method according to any one of claims 1 to 8, characterized in that the target trait-associated SNP marker is selected from a subject.
10. A method according to claim 9, characterized in that the target comprises at least one selected from the group consisting of seeds, plants, animals, bacteria, and insects.
11. A method according to claim 9, characterized in that the target object is a plant.
12. A method according to claim 9, characterized in that the target object is a seed.
13. A method according to any one of claims 1 to 12, characterized in that the target trait-associated SNP marker is selected from a sample.
14. A method according to claim 13, characterized in that the sample comprises at least one selected from the group consisting of seeds, plants, animals, bacteria and insects.
15. A method according to claim 13, characterized in that the sample is a plant or a seed.
16. A method according to any one of claims 1 to 15, wherein the target trait has a characteristic including ripening of pepper fruit.
17. A method according to any one of claims 1 to 16, characterized in that the target trait has a characteristic including resistance to yellowing disease.
18. A method according to any one of claims 1 to 17, characterized in that the target trait has a characteristic including a spicy taste gene (Pun1) of a red pepper.
19. A method according to any one of claims 1 to 18, characterized in that the target trait has a characteristic including a resistance gene R1 against a Phytophthora pathogen.
20. A method according to any one of claims 1 to 19, characterized in that the target trait has a characteristic including a resistance gene R8 against a Phytophthora pathogen.
21. A method according to any one of claims 1 to 20, characterized in that the target trait has a characteristic including a resistance gene Rpi-amr1 against a Phytophthora pathogen.
22. A method according to any one of claims 1 to 21, wherein the target trait comprises at least one selected from the group consisting of herbicide resistance, disease resistance, insect or pest resistance, fungal disease resistance, virus resistance, nematode resistance, bacterial disease resistance, fatty acid biosynthesis, starch biosynthesis, increased grain yield, increased oil, improved nutrition, increased growth rate, fruit ripening, increased yield, improved stress tolerance, improved environmental or chemical tolerance, changes in morphological characteristics, improved digestibility, industrial enzyme production, improved flavor, nitrogen fixation, and any combination thereof.
23. A method according to any one of claims 1 to 22, characterized in that the method shows better target trait selection efficiency compared to a Genome-Wide Association Study (GWAS) program and / or QTL mapping.
24. In the first paragraph, the method is characterized by comprising the following steps: (1) A step of obtaining SNP (single nucleotide polymorphism) by analyzing the genome sequence information of each individual or offspring group obtained by mating the paternal and maternal lines; (2) A step of measuring the trait value for the target trait for each individual or offspring group obtained by crossing the paternal and maternal lines, and listing the SNPs in order of the trait value; (3) A step of calculating the score (Sj) of each locus using the following [Formula 1] by differently assigning weights according to the variance based on the genotypes of the individuals distributed at both ends of the above-mentioned sorted trait values and the genotypes of the individuals having the median value of the target trait; and [Formula 1] (4) A step of selecting the top 10 loci among markers (positions) between 1 and L using [Formula 3] for the score (Sj) of each locus in step (3). [Formula 3] 25. A method in accordance with claim 24, wherein the SNP in step (1) is selected as a homozygous SNP among polymorphic SNPs between parental lines, having a minor allele frequency (MAF) of 5% or more, a missing data of 30% or less, and a polymorphic SNP between parental lines.
26. In the 24th paragraph, the target trait-associated SNP marker is a method characterized in that steps (3) and (4) are repeatedly performed on individuals obtained by sequentially applying a range of significance probabilities (p-values) of 5% or less, 10% or less, and 15% or less to both ends of the aligned trait values of step (3).
27. A method for producing a population of plants or seeds, comprising: selecting a population of plants or seeds according to any one of the methods of claims 1 to 26; and crossbreeding or selfing one or more of the selected plants, or plants grown from the selected seeds.
28. A method according to claim 27, characterized in that the selected population has a target trait having improved ripening of pepper fruits.
29. A method according to claim 27, characterized in that the selected population has a target trait having enhanced resistance to wilt disease.
30. A method according to claim 27, characterized in that the selected population has a target trait having an improved spicy taste gene (Pun1) of a pepper.
31. A method according to claim 27, characterized in that the selected population has a target trait having an improved resistance gene R1 against Phytophthora pathogens.
32. A method according to claim 27, characterized in that the selected population has a target trait having an improved resistance gene R8 against Phytophthora pathogens.
33. A method according to claim 27, wherein the selected population has a target trait having an improved resistance gene Rpi-amr1 against Phytophthora pathogens.
34. A step of determining, detecting, selecting or separating a SNP marker related to a target trait according to any one of the methods of clauses 1 to 26, and A method for introducing or increasing expression of a target trait-related SNP in a subject, comprising a step of expressing or increasing expression of the target trait-related SNP in the subject.
35. A step of determining, detecting, selecting or separating a SNP marker related to a target trait according to any one of the methods of clauses 1 to 26, and A method for producing a plant having a target trait, comprising a step of expressing or increasing the expression of a SNP related to the target trait in a target plant.
36. A method according to claim 34 or 35, wherein the subject is a plant, and the method further comprises crossing the plant with another plant of the same species to produce a progeny plant having the target trait.
37. A plant having a target trait produced by the method of any one of claims 34 to 36.
Citation Information
Patent Citations
Method for selecting and utilizing tag-SNP for discriminating haplotype in gene unit
KR1020180046592A
An Intelligent integrated apparatus for assembling and transferring of semiconductor light emitting device
KR1020230163764A