SNP marker for discriminating high and low protein content in rice seeds and the use thereof

A machine learning-based model using SNP markers accurately predicts rice protein content, overcoming the inefficiencies of existing methods by achieving rapid and precise protein content determination in rice seeds, thereby supporting breeding goals.

KR1020260112864APending Publication Date: 2026-07-21REPUBLIC OF KOREA (MANAGEMENT RURAL DEV ADMINISTRATION)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
KR · KR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-01-10
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing methods for determining rice protein content are time-consuming and costly, and existing genetic markers are difficult to utilize in breeding due to the complex regulation of protein content by various genetic mechanisms.

Method used

A machine learning-based prediction model using 26 SNP markers associated with high and low protein content in rice is developed, allowing for rapid and accurate prediction of protein content in rice seeds through genomic DNA analysis and machine learning algorithms such as SVM, KNN, FR, C5, and PLS.

Benefits of technology

The model achieves an accuracy of over 80% in predicting protein content, enabling early-stage determination and supporting data-driven breeding decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure PAT00005_ABST
    Figure PAT00005_ABST
Patent Text Reader

Abstract

The present invention relates to a Single Nucleotide Polymorphism (SNP) marker for determining the degree of protein content of rice seeds and a method for determining the degree of protein content of rice seeds using said marker. Specifically, the present invention relates to an SNP marker for determining protein content and a method for determining protein content using said marker, which constructs a machine learning-based rice protein content prediction model by utilizing SNP markers associated with high and low protein content, thereby enabling the prediction of rice seeds with high protein content in the early stages of growth.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] The present invention relates to an SNP marker for determining the degree of protein content of rice seeds and a method for determining the degree of protein content of rice seeds using said marker. Specifically, the present invention relates to an SNP marker for determining protein content and a method for determining protein content using said marker, which constructs a machine learning-based rice protein content prediction model by utilizing SNP markers associated with high and low protein content, thereby enabling the prediction of rice seeds with high protein content in the early stages of growth. Background Technology

[0002] Rice (Oryza sativa L.) is one of the world's most important food crops and is the primary food crop in Korea. Since 2006, imported rice for cooking has been distributed and the rice market has diversified, leading to increased consumer interest in taste and a rising demand for high-quality rice. As a result, the percentage of whole grains, moisture content, and protein content are considered important criteria for high-quality rice.

[0003] Moisture content is determined by the harvest time, drying method, and drying time; therefore, harvesting at the appropriate time and drying at 45–50°C or lower is recommended. It is known that excessive drying after harvest increases cracked grains, reduces germination rates, causes denaturation of proteins and starches, and lowers palatability.

[0004] Protein is the second most abundant component in rice, after starch. The protein content of rice is a quantitative characteristic with low heritability that is heavily influenced by the environment. As a factor that significantly impacts nutritional value and food processing characteristics, it is directly linked to rice quality, which in turn is directly related to consumer demand. Therefore, it is crucial to accurately predict and control the protein content of rice.

[0005] The protein content can only be measured after the rice has been processed into white rice through sorting, drying, hulling, and polishing processes, so the existing method of measuring this requires a significant amount of time and cost to determine the protein content of the rice.

[0006] Meanwhile, GWAS analysis of 529 rice lines has identified variants and genes (flo5, etc.) associated with protein content according to various rice ecotypes (see reference [Chen et al., Molecular Breeding, 2023]; also see Table 1 below).

[0007] Chr Lead SNP Num Population Yea Known genes / QTLs (bp) 1 sf0102103441 8 All, Ind_All, Jap_All, TrJ 2014, 2016 1 sf0102342328 5 All, Jap_All, TeJ 2014, 2015, 2016 1 sf0103093920 2 All 2014, 2016 1 sf0113186602 2 IndI 2015, 2016 Sar1a (-98.48) 1 sf0139892657 2 TeJ, Aus 2015 1-19 2 sf0206873340 4 All, Ind_All 2014, 2015 2-5 2 sf0208274542 1 IndI 2015 GluB6 (-128.21), 2-7 2 sf0212141249 2 All, Jap_All 2014 2-9 2 sf0219198685 1 All 2015 OsTudor-SN (93.15) 3 sf0311636192 3 All, Jap_All 2014 3 sf0316272964 1 IndII 2015 Susy2 (-33.18) 3 sf0319366452 2 All, Ind_All 2014 4 sf0400950425 3 All, Ind_All, Jap_All 2014, 2015, 2016 4 sf0401250213 3 All, Jap_All 2015, 2016 4 sf0401819925 5 All 2014, 2015, 2016 4 sf0403146346 4 All, Jap_All 2014, 2016 4 sf0404599105 2 All, TeJ 2014 4 sf0423552592 3 All, Ind_All 2015, 2016 5 sf0508556305 2 All, Jap_All 2014 5 sf0516166384 2 All, IndII 2015 5 sf0524443891 1 Ind_All 2014 Glb1 (-133.85) 6 sf0609401331 3 All, TeJ 2016 6 sf0618816229 5 All, Jap_All 2014, 2015, 2016 7 sf0706087386 1 All 2016 Rc (24.5), 7-4 7 sf0706126055 4 All, Ind_All 2015, 2016 7-4 7 sf0707834002 1 Ind_All 2016 OsAGPL4 (-155.28) 7 sf0708340365 2 All, Jap_All 2014 qPC-7, 7-4 7 sf0709202668 3 All, Ind_All, Jap_All 2014, 2015, 2016 7-4 7 sf0710047261 3 All 2014, 2016 7-4 7 sf0712761453 2 All, IndI 2014, 2016 7-4 7 sf0712867464 1 IndI 2014 GBSSII (-53.72) 7 sf0714417157 2 IndII 2015, 2016 7 sf0714859990 2 IndI 2014, 2015 7 sf0716523661 2 Jap_All, TeJ 2014 7 sf0719598121 4 All, TeJ 2014, 2016 7-9 7 sf0729064931 2 All 2014, 2016 8 sf0800465605 3 All, Ind_All 2015, 2016 8 sf0805366362 1 All 2014 Flo5 (14.26) 8 sf0817958573 3 IndII, IndI 2014, 2015, 2016 8-9 8 sf0827747654 2 Ind_All, IndII 2014, 2016 8 sf0828049587 2 All, Jap_All 2015 9 sf0911061640 2 Ind_All, IndII 2014 11 sf1102396805 2 All, Jap_All 2014

[0008] However, despite these findings, it was difficult to actually utilize them in breeding because protein content is regulated by various mechanisms depending on the genetic resources of rice.

[0009] Therefore, it has become necessary to apply machine learning to identify protein content-associated variant markers applicable to major domestic rice lines, including Tongilbyeo, and to make rapid and accurate predictions. There is also a need to build a machine learning-based prediction model capable of accurately and rapidly determining the protein content of rice at an early stage, such as in seeds or during the early growth stage, and to utilize it for breeding. Prior art literature

[0010] Registered Patent No. 10-1000889 (Published December 14, 2010) The problem to be solved

[0011] Accordingly, the inventors of the present invention sought to provide a means to predict the degree of protein content by constructing a machine learning-based rice protein content prediction model.

[0012] Accordingly, the objective of the present invention is to provide a molecular marker that can be used to predict the protein content of rice, and to provide a method for determining the protein content of rice using the same. means of solving the problem

[0013] To achieve the above objective, the present invention

[0014] 1) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 1 is A or G;

[0015] 2) A single nucleotide polymorphism in which the 151st base of the sequence indicated by SEQ ID NO. 2 is T or A;

[0016] 3) A single nucleotide polymorphism in which the 151st base of the sequence indicated by SEQ ID NO. 3 is T or C;

[0017] 4) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 4 is G or A;

[0018] 5) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 5 is T or C;

[0019] 6) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 6 is A or C;

[0020] 7) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 7 is G or A;

[0021] 8) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 8 is C or T;

[0022] 9) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 9 is C or T;

[0023] 10) A single nucleotide polymorphism in which the 151st base of the sequence represented by sequence number 10 is C or G;

[0024] 11) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 11 is T or G;

[0025] 12) A single nucleotide polymorphism in which the 151st base of the sequence represented by sequence number 12 is C or T;

[0026] 13) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 13 is T or A;

[0027] 14) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 14 is G or A;

[0028] 15) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 15 is C or G;

[0029] 16) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 16 is G or A;

[0030] 17) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 17 is T or C;

[0031] 18) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 18 is G or A;

[0032] 19) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 19 is C or G;

[0033] 20) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 20 is G or A;

[0034] 21) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 21 is G or A;

[0035] 22) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 22 is A or G;

[0036] 23) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 23 is T or C;

[0037] 24) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 24 is T or C;

[0038] 25) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 25 is C or T;

[0039] 26) A single nucleotide polymorphism in which the 151st base of the sequence denoted by SEQ ID NO. 26 is A or T

[0040] A composition for determining high or low protein content in rice seeds is provided, comprising a preparation capable of detecting [the protein content].

[0041] In addition, the present invention provides a kit for determining high or low protein content in rice seeds, comprising the above composition.

[0042] In addition, as a method for determining the high or low protein content of rice seeds,

[0043] i) Step of isolating genomic DNA from rice;

[0044] ii) a step of determining all the genotypes of 1) to 26) using the isolated genomic DNA as a template; and

[0045] iii) A step of determining whether the protein content of rice seeds is high or low using a model machine-trained with a machine learning algorithm selected from the group consisting of SVM, KNN, FR, C5, and PLS, based on the genotype information determined in step ii) above.

[0046] Provides a method including Effects of the invention

[0047] It is possible to determine the protein content of rice in seeds and during the early stages of growth using the genetic marker according to the present invention, and to support accurate data-based decision-making according to the breeding goal regarding the protein content of rice. Brief explanation of the drawing

[0048] Figure 1 shows a variant search pipeline in genotype analysis. Figure 2 shows the phenotypic distribution of protein content in rice. Figure 3 shows a Manhattan plot of markers associated with the protein content of rice. Figure 4 shows the results of developing a protein content prediction model through machine learning (26 SNPs). Figure 5 shows the results of inputting genotype information for 26 selected markers into a prediction model using machine learning (Control - bottom 30% low-dose group; Case - top 30% high-dose group). Specific details for implementing the invention

[0049] The present invention will be described in detail below.

[0051] The present invention discovered SNP markers associated with rice protein content through genome-wide association study (GWAS) analysis. Specifically, by searching for common genotypes among individuals classified by phenotype, genotypes of genetic loci associated with protein content were obtained from among a number of SNPs obtained throughout the genome.

[0052] Among the obtained genotypes, 26 variants capable of predicting protein content were selected through machine learning, and a model for predicting the protein content of rice was constructed based on this, and it was confirmed that the accuracy of the model was over 80%.

[0053] In the present invention, a total of 26 SNP markers were identified, and flanking sequences of 150 bp before and after these SNP markers are presented in SEQ ID NOs 1 to 26.

[0054] Using the SNP marker provided in the present invention, it can be determined that the protein content of rice seeds is higher or lower than the average protein content of rice. Specifically, using the SNP marker provided in the present invention, rice seeds determined to have a high protein content can be determined to have a protein content of 6.5% or more, preferably 6.6% or more, more preferably 6.9% or more, and rice seeds determined to have a low protein content can be determined to have a protein content of less than 6.5%.

[0055] 1) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 1 is A or G;

[0056] 2) A single nucleotide polymorphism in which the 151st base of the sequence indicated by SEQ ID NO. 2 is T or A;

[0057] 3) A single nucleotide polymorphism in which the 151st base of the sequence indicated by SEQ ID NO. 3 is T or C;

[0058] 4) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 4 is G or A;

[0059] 5) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 5 is T or C;

[0060] 6) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 6 is A or C;

[0061] 7) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 7 is G or A;

[0062] 8) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 8 is C or T;

[0063] 9) A single nucleotide polymorphism in which the 151st base of the sequence indicated by sequence number 9 is C or T;

[0064] 10) A single nucleotide polymorphism in which the 151st base of the sequence represented by sequence number 10 is C or G;

[0065] 11) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 11 is T or G;

[0066] 12) A single nucleotide polymorphism in which the 151st base of the sequence represented by sequence number 12 is C or T;

[0067] 13) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 13 is T or A;

[0068] 14) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 14 is G or A;

[0069] 15) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 15 is C or G;

[0070] 16) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 16 is G or A;

[0071] 17) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 17 is T or C;

[0072] 18) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 18 is G or A;

[0073] 19) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 19 is C or G;

[0074] 20) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 20 is G or A;

[0075] 21) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 21 is G or A;

[0076] 22) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 22 is A or G;

[0077] 23) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 23 is T or C;

[0078] 24) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 24 is T or C;

[0079] 25) A single nucleotide polymorphism in which the 151st base of the sequence denoted by sequence number 25 is C or T;

[0080] 26) A single nucleotide polymorphism in which the 151st base of the sequence denoted by SEQ ID NO. 26 is A or T

[0081] A composition for determining high or low protein content in rice seeds is provided, comprising a preparation capable of detecting [the protein content].

[0082] In the present invention, the formulation capable of detecting the single nucleotide polymorphisms of 1) to 26) may be a primer pair or a probe capable of detecting each of 1) to 26).

[0083] In addition, the present invention provides a kit for determining high or low protein content in rice seeds, comprising the above composition.

[0085] In the present invention, the term "nucleotide" refers to a deoxyribonucleotide or ribonucleotide existing in a single-stranded or double-stranded form, and includes analogs of natural nucleotides unless specifically otherwise noted.

[0086] In the present invention, the term “marker” refers to a nucleotide sequence used as a reference point when identifying genetically unrelated loci. This term also applies to nucleic acid sequences complementary to the marker sequence, such as nucleic acids used as a set of primers to amplify the marker sequence. The location of a molecular marker on a gene map is referred to as a genetic locus.

[0087] In the present invention, the term “SNP (single nucleotide polymorphism)” refers to a polymorphic region in which two or more alleles exist at a single genetic locus, and only a single nucleotide differs. That is, it refers to a case in which some nucleotides in the entire genome of a population differ from chromosome to chromosome. Typically, it is known that SNPs exist about once every 300 to 1,000 nucleotides, but are not limited thereto.

[0088] In the present invention, the term "primer" refers to a short nucleic acid sequence having a free 3' hydroxyl group at the free 3' end, capable of forming base pairs with a complementary template, and functioning as a starting point for copying the template. The primer can initiate DNA synthesis at an appropriate temperature in the presence of a suitable buffer solution, a reagent for polymerization, four different nucleotide triphosphates, and DNA polymerase or reverse transcriptase.

[0089] In the present invention, the term "complementary" means sufficiently complementary to the extent that a primer selectively hybridizes to a target nucleic acid sequence under specific annealing or hybridization conditions, and encompasses both substantially complementary and perfectly complementary, preferably meaning perfectly complementary.

[0090] In the present invention, the oligonucleotide used as a primer may also include a nucleotide analogue, for example, a phosphorothioate, an alkylphosphorothioate, or a peptide nucleic acid, or may include an intercalating agent.

[0091] In the present invention, the term "probe" refers to a hybridization probe comprising a natural or modified monomer or a linear oligomer having a bond, comprising a deoxyribonucleotide and a ribonucleotide, capable of sequence-specifically binding to the complementary strand of a nucleic acid. The probe of the present invention is an allele-specific probe in which a polymorphic site exists in a nucleic acid fragment derived from two members of the same species, so that it hybridizes to a DNA fragment derived from one member but not to a fragment derived from the other member. Preferably, the probe may be a single strand, more preferably a deoxyribonucleotide, for maximum efficiency in hybridization, but is not limited thereto.

[0092] In the present invention, the reagent for performing the amplification reaction may include, but is not limited to, DNA polymerase, dNTPs, and a buffer. The dNTPs include dATP, dCTP, dGTP, and dTTP, and the DNA polymerase may be a heat-resistant DNA polymerase, such as commercially available polymerases like Taq DNA polymerase and Tth DNA polymerase. Additionally, the kit of the present invention may further include a user manual describing optimal reaction conditions.

[0093] As a method for determining the high or low protein content of rice seeds,

[0094] i) Step of isolating genomic DNA from rice;

[0095] ii) a step of determining all the genotypes of 1) to 26) using the isolated genomic DNA as a template; and

[0096] iii) A step of determining whether the protein content of rice seeds is high or low using a model machine-trained with a machine learning algorithm selected from the group consisting of SVM, KNN, FR, C5, and PLS, based on the genotype information determined in step ii) above.

[0097] Provides a method including

[0098] The method of the present invention includes the step of isolating genomic DNA from rice.

[0099] In the present invention, genomic DNA extraction from a sample can be performed using a commonly used phenol / chloroform extraction method, SDS extraction method, CTAB separation method, or a commercially available DNA extraction kit.

[0100] In the present invention, any method known to those skilled in the art can be used to determine the genotype of 1) to 26). For example, it can be performed by known methods such as sequencing analysis, sequencing analysis using an automated DNA sequencer, pyrosequencing, hybridization by chip array, PCR-RELP (restriction fragment length polymorphism), PCR-SSCP (single strand conformation polymorphism), PCR-SSO (specific sequence oligonucleotide), ASO (allele specific oligonucleotide) hybridization combining PCR-SSO and dot hybridization, TaqMan-PCR, MALDI-TOF / MS, RCA (rolling circle amplification), HRM (high resolution melting), primer extension, Southern blot hybridization, dot hybridization, etc.

[0101] Meanwhile, machine learning analytics is the fastest-growing field in computer science, helping to derive complex correlations hidden within large-scale datasets. It utilizes a diverse set of algorithms that build models using existing data and facilitate pattern recognition and classification leading to predictions. Machine learning algorithms can be broadly divided into supervised learning and unsupervised learning; supervised learning is a method that predicts the classification of new objects by training on well-classified objects, while unsupervised learning is a method that classifies provided objects without a training process (Ang et al., Comput . Biol . Bioinform Such machine learning analysis is being utilized in various biological fields, such as coding region recognition, biomarker identification, and disease gene discovery (Han et al., Biomed Res. Int . 8496165. 2016).

[0102] In the present invention, 213 initial association markers related to protein content were selected through Genome-Wide Association Study (GWAS) analysis. Based on the information of the SNP markers of the 213 variant markers, 26 markers were selected by additionally selecting those with high priority using a machine learning feature selection method, by setting different weights according to the importance of the markers in distinguishing phenotypes.

[0103] In order to develop a genotype-based phenotype prediction algorithm through machine learning, at least five machine learning algorithms were applied, specifically machine learning algorithms selected from a group consisting of SVM, KNN, FR, C5, and PLS. In the machine learning model constructed in this way, the protein content of rice seeds can be determined quickly and accurately by examining only the genotypes of the 26 gene markers provided in the present invention.

[0105] The present invention will be explained in detail below through examples. However, the following examples are merely illustrative of the present invention, and the scope of the present invention is not limited to the following examples.

[0107] Examples 1. GWAS SNPs associated with rice protein content through analysis Marker excavation

[0108] By searching for common genotypes among individuals classified by phenotype, genotypes of loci associated with protein content were obtained from a number of SNPs found across the genome. The study included 85 superior domestic rice varieties, of which 55 were Japonica and 30 were Tongil rice (see Table 2 below).

[0109] No. Variety Name (Korean) Variety Name (English) System name Year of training biovar 01 Chamdongjin Chamdongjin Jeonju623 2020 japonica 02 Five generations Odae Suweon303 1980 japonica 03 Samgwang Samgwang Suweon474 2003 Tongil-type 04 Miryang 23 Milyang23 Milyang23 1976 Tongil-type 05 sign Giho Suweon306 1980 japonica 06 cult Yechan Iksan583 2017 japonica 07 Cheongcheong rice Cheongcheongbyeo Milyang46 1979 Tongil-type 08 Mimyeon Mimyeon Milyang260 2012 Tongil-type 09 Sangnam field rice Sangnambatbyeo Milyang93 1988 japonica 10 Yongmun rice Yongmunbyeo Suweon332 1985 Tongil-type 11 Hangangchal No. 1 Hangangchal 1 Milyang167 2006 Tongil-type 12 Book Review Seopyeong Gyehwa22 2003 japonica 13 Nokwoo Nokwoo Suweon560 2014 Tongil-type 14 Wooden puppet Mogwoo Suweon519 2009 Tongil-type 15 Aromi Aromi Milyang302 2016 japonica 16 Jangseong rice Jangseongbyeo Iri362 1985 Tongil-type 17 Nokyang Nokyang Suweon490 2006 Tongil-type 18 Samgang rice Samgangbyeo Milyang55 1982 Tongil-type 19 progress Jinbo Yeongdeog45 2009 japonica 20 Healthy red rice Geonganghongmi Milyang234 2010 japonica 21 Hyangmibyeo No. 1 Hyangmibyeo 1 Suweon393 1993 Tongil-type 22 Joeun Black Rice Joeunheukmi Cheolwoen80 2011 japonica 23 Dasan No. 1 DASAN1HO Suweon499 2006 Tongil-type 24 Nam-il Namil Suweon472 2002 japonica 25 Joan Joan Suweon478 2003 japonica 26 Borami Borami Milyang211 2007 japonica 27 black pearl Heukjinjubyeo Suweon415 1997 japonica 28 South Wind Rice Nampungbyeo Suweon294 1981 Tongil-type 29 Hangang glutinous rice Hangangchalbyeo Suweon290 1979 Tongil-type 30 Saemimyeon Saemimyeon Milyang278 2014 Tongil-type 31 Jewel Black Glue Boseogheugchal Suweon512 2008 japonica 32 MY299BK MY299BK Milyang299 2016 japonica 33 Hangaru Hangaru Suweon594 2016 japonica 34 Black rice Heugnambyeo Iri427 1997 japonica 35 Taebaek rice Taebaekbyeo Suweon287 1979 Tongil-type 36 Cheongdam Cheongdam Suweon498 2006 japonica 37 Yeongdeok rice Yeongdeogbyeo Yeongdeog3 1985 japonica 38 Snow White Seolbaek Cheolweon76 2010 japonica 39 Gangbaek Gangbaek Iksan478 2006 japonica 40 Shepherding Mogyang Suweon525 2010 Tongil-type 41 Hwaseong rice Hwaseongbyeo Suweon330 1985 japonica 42 Cheongwoo Cheongwoo Suweon585 2016 Tongil-type 43 CW92MR CW92MR Cheolweon92 2016 japonica 44 Hanareumchal Hanareumchal Milayng288 2015 Tongil-type 45 Jinbuolbyeo Jinbuolbyeo Jinbu11 1991 japonica 46 Miu Miwoo Suweon597 2017 Tongil-type 47 Jinmi rice Jinmibyeo Suweon349 1989 japonica 48 signature Seomyeong Gyehaw30 2009 japonica 49 Nong-an rice Nonganbyeo Suweon392 1993 japonica 50 Beautiful rice Areumbyeo Milyang160 1999 Tongil-type 51 Enemy Headquarters Reconnaissance Jeogjinjuchal Suweon524 2010 japonica 52 Dacheong Dacheong Iksan495 2008 japonica 53 Cheongbaekchal Cheongbaekchal Cheolweon77 2010 japonica 54 Hong Jin-ju HONGJINJU Suweon501 2006 japonica 55 Joryeong rice Joryeongbyeo Milyang107 1992 japonica 56 Youngwoo Yeongwoo Suweon573 2015 Tongil-type 57 Baekjinju No. 1 Baegjinju1ho Suweon491 2005 japonica 58 All-rounder Palbangmi Suweon543 2012 Tongil-type 59 Anmi Anmi Suweon523 2010 japonica 60 Geonyangmi Geonyangmi Suweon533 2011 japonica 61 Baekokchal Baegokchal Milyang225 2008 japonica 62 Andabye Andabyeo Suweon431 1998 Tongil-type 63 Jinseolchal Jinseolchal Jinbu50 2012 japonica 64 Dasan No. 2 Dasan2 Suweon518 2009 Tongil-type 65 Big-eyed Black Chal 1 Nunkeunheugchal1ho Milyang282 2014 japonica 66 MY298BB MY298BB Milyang298 2016 japonica 67 Daedeolbyeo No. 1 Daeripbyeo1 Suweon391 1993 japonica 68 Seonhyang Black Rice Seonhyangheukmi Suweon532 2011 japonica 69 Unilchal unilchal Unbong52 2014 japonica 70 Hanareum No. 4 Hanareum4 Milyang295 2016 Tongil-type 71 Ilpum rice Ilpumbyeo Suweon355 1990 japonica 72 Jungwon rice Jungwonbyeo Suweon325 1984 Tongil-type 73 Manmi Manmi Milyang162 2002 japonica 74 Jo Gwang Jogwang Milyang213 2007 japonica 75 Cheongun Cheongun Suweon537 2012 japonica 76 New Odae No. 1 Saeodae1ho Cheolwon 103 2022 japonica 77 True delicacy Chamjinmi Jeonju654 2022 japonica 78 Full to the brim Hangadeuk Suweon639 2022 japonica 79 Dangchan Jinmi Dangchanjinmi Jeonju632 2022 japonica 80 Amisal Amissal Jeonju653 2022 Tongil-type 81 Dapyeong Dapyeong Milyang 363 2022 japonica 82 Amimyeon Amimyeon Milyang 355 2022 Tongil-type 83 Jin Ok-chal Jinokchal Milyang 366 2022 japonica 84 Chamnuri Chamnuri Jeonju 636 2021 japonica 85 Chohong Chohong Milyang 356 2021 japonica

[0111] Rice genotyping (genomic variation by resource) analysis was performed using a pipeline based on GATK (Genome Analysis Toolkit, Broad Institute), which is the most widely used tool globally and is effectively the de facto standard. The pipeline operates in six stages, and the sequence is as follows (see Fig. 1):

[0112] i. Use Trimmomatic to remove low-quality regions and adapter sequences from the NGS sequences.

[0113] ii. The nucleotide sequence obtained from step i. is mapped to the rice standard genome IRGSP-1.0 using BWA.

[0114] iii. Analyze and organize the mapping results using the Picard program.

[0115] iv. Genomic variant information is extracted using HaplotypeCaller (GATK).

[0116] v. Correct the extraction conditions of the mutation information using the extracted mutation information.

[0117] vi. Re-extract genomic variant information using HaplotypeCaller (GATK) under corrected conditions.

[0118] Characteristically, the accuracy of the information was enhanced through two stages of genomic variant information extraction, and the increase in analysis time caused by the addition of the variant information extraction stage was overcome by using a supercomputer.

[0119] Genotypes were analyzed for 85 domestically bred rice varieties and 20 domestically collected weed rice varieties in Table 2, and the results of the analysis are shown in Table 3 (variation search results).

[0120] chromosome minimum variation Maximum variation Variation statistics Sample transition number Sample transition number average Standard deviation Chr01 samgwang 1,482 YW1428 174,246 39,737 32,265 Chr02 giho 1,550 YW1428 122,805 35,397 26,768 Chr03 giho 1,168 YW1428 140,698 32,494 25,830 Chr04 chamdongjin 1,218 YW1428 205,660 47,995 36,059 Chr05 samgwang 864 YW1428 113,793 36,114 26,976 Chr06 giho 748 YW1428 205,739 39,830 31,540 Chr07 samgwang 845 YW1428 163,085 38,837 31,173 Chr08 giho 830 YW1428 145,810 43,514 30,587 Chr09 chamdongjin 593 YW1428 123,714 30,637 23,676 Chr10 giho 740 YW1428 117,459 39,098 27,731 Chr11 giho 1,166 YW1428 171,826 46,261 32,476 Chr12 odae 943 YW3624 146,809 38,879 28,549 Total giho 13,381 YW1428 1,780,672 468,792 349,080

[0121] According to the analysis results, the least variation was generally found in Gihobyeo, followed by Samgwang and Chamdongjin. On the other hand, Geojeaengmi 2 (YW1428) and Bukjejuaengmi 1 (YW3624) showed a high amount of variation.

[0122] By chromosome, the variation was found to be least in the order of chromosomes 9, 10, 6, 8, 7, 5, 12, 11, 3, 4, 1, and 2.

[0123] Meanwhile, regarding the protein content of rice, the overall average was 6.6% and the average for Japonica was 6.5%, which was found to be slightly lower than that of Tongil rice at 6.9% (see Fig. 2).

[0124] Based on the protein content distribution shown in Figure 2, the group was divided into a high-content group (top 30%) and a low-content group (bottom 30%), and GWAS analysis was performed to select 213 associated markers that met the criterion of -log10P ≥ 6 (see Figure 3).

[0126] Examples 2. Development of a Protein Content Prediction Model Using Machine Learning

[0127] A training set for machine learning was determined using 213 protein content-associated variant markers selected through GWAS analysis in Example 1 above.

[0128] Among the 213 selected variant markers, 26 markers were selected by using a machine learning feature selection method to further select those with high priority based on the importance of the markers in distinguishing phenotypes.

[0129] To develop a genotype-based phenotype prediction algorithm using machine learning, individuals used for marker selection were randomly sampled and divided into 70% (training set) and 30% (validation set), accuracy was improved through at least 100 iterations, and at least 5 machine learning algorithms were applied.

[0130] Since the sample size is too small, there is a possibility that the accuracy of the model may be misleading due to coincidence, so the test was repeated with 1 to 5 seeds.

[0131] TP (True Positive), FP (False Positive), TN (True Negative), and FN (False Negative) were measured using a 30% validation set, and sensitivity, specificity, and accuracy were predicted. The results are shown in Figure 4.

[0132] As shown in Figure 4, most models showed an accuracy of over 80%, except for the C5 model, and in particular, when the PLS model was applied, it showed a classification accuracy of 83.6%.

[0133] Therefore, it was determined that even by investigating only the genotypes of 26 SNPs, it would be possible to predict the desired phenotype, that is, whether the rice seeds contain a high amount of protein, through a machine learning model.

[0135] Examples 3. Utilization of Machine Learning Prediction Models

[0136] As a result of reviewing the feasibility of designing a fluidigm chip for the application of a machine learning prediction model, it was confirmed that all 26 markers were suitable.

[0137] For the Fluidigm chip design, 150 bp of adjacent sequences before and after the SNP are provided, which are presented in Table 4 (SNP adjacent sequence information) below.

[0138] Marker ID Flanking sequence chr01:27390946 GGTGAAGTGGGGAGTGGGGAGTGGTGAAGTGGTTCAGTCGCAGAGTCGCTCAAGGGCCAAGGCGGTTTGAAGGATGTCTAGTTCACCTCTTGAATCATCTATCAGCACTAACTAGCTTTTTCATAACGAAATTCGAAGTGATGGCGTTCT[ A / G ]CATGGTAATACTCCCTCCGTCCAAAAAAAAAAACTCAACCTACTACTTCGATGGGATATAATCTAGTACAACAAATCTAGATAGGAGTACTAGGTTAAGTTTTTTTGGGACGAAGGAAGTATGCGGTAACTTTACATCCACCGTAGATCT(서열번호 1) chr01:39134375 TTTGCAAAATGAATCCTTTACTGTTAACACAAAATCGCACACTATCACGAATCCATATCTATGAAATGAATATCGGAGTGATCCAACCAAATTCTCTAAGCAACACCAAAGGATGTTATTACTAGGGAAAACGGAACAGTTGTACTGATC[ T / A ]GAGAAAAAAAAAAGGTTTCTCCTGTTCCTCGAACAGTTGTTCAACCCCCAATAGAGAAAAATAAAGAAACCCCATCTTGATCGGTTCTTTGCCTCTGTTCTACGAGCATACAATCCAATATGAGTAAAACCATAAAAATCATGCCGATTA(서열번호 2) chr02:23722718 GGTAAAGTGTAAAATTGATGAATAAGTGATATTGTCTAACTTACCACTATTGCTGTACTTCAAATATTCAATGAACTTTTAAGGATGTATTTACAATTAACGAAATATGGTTGCAGATCTGCCTGCGTTTTAGCTCAGCTAAAGTATTTC[ T / C ]GACTTCTGAGAGGTAATCTTAATCAGGAAGAGGGGGGGGGGGAGGGGGGGAGGGGGGGAGGGGGAAGCAGCATACCTGATGATAAGCAATCTTCTTATTTTCCTTTTTCCCCTCTTTCTCCAGTCTGTTTGACCAATGTGTTGGAATCGA(서열번호 3) chr04:29776964 TAAAGAACTATGATTCTCTATCTTTACCATTCTCGATGTTACCGTCATTCGGTGGCTTAGTAGGTGATGCATCATGTGAAGAACTACCATTCTCATGGTCAGGGTCTTTGGATATCTCAGCAGGCACGGCAGCCCCCTCACATGGCAATG[ G / A ]CATCGAAGTCAAGTGCAAGTAACTACGTAGGAATGAATCATGCGCTCTTGCAGGCGGGTGAGAGACTCTGTTGAACCAACGTATCACCGTGAGATCGTCCCCCTCAACGATCAGCTTTTTAACTGGTTCATGGTACTCAAGCATGACCTT(서열번호 4) chr04:31539521 TCTTATAATATGAAACGGAGGGAGTAATTATTATGGAGAAACAGATAGGTTATGTTTCCAGTGCCAATTCATTAATAATTTAGGATCAGAATATATGTCATGCATGTTCATATAAGTTATTGGGGTTTACTCCAGACTTCATATCCTGCG[ T / C ]ACAACTTAGCTTGCTTTCTACATGCCTTTGACATCTTCTGGGGAAAAAGGATCTGCTAAACACTGTATCATTTATCTGGTCATAGCCGGCCAACAACACTGCAGGATCGAGGTAGGGACAGTGATGAAGAAGACTTTCGAGATAGGGACT(서열번호 5) chr04:31882216 CCAAAGAGAATGAATTGGAAAGGAACTGCCCTAACTCCTCCCCACAAATATCGACTTTACATAGATGCAAGCTTGTCAAGCTTATGTTGCAACCAAGAGTCACCGTAGGATGAAAAGCACAGGAGGAGAGGAAAAGTGACTGAATTGTCC[ A / C ]TCCCTCTGCCTTATTGGATAGAACGGAACATGGGAAGTTGTATTCTGTCCGTTTTTTCAATGAAAGACATCTCCAAGGAGAGTTCTTCGATCCCAGGTTTGATAACGGTTACAAACCATTTGTCAAGAACAGCTGCACTTATATTCGGAC(서열번호 6) chr04:670596 GACGAGCTAAAGCAATAACTGCAAAGCCCGCCAGCTCTGATCAGCCCAGCACCAAGGAGTGAACTGCTACTATACCTGGCAGCTTCGCCAGTGGCAGTCAGTGCTGCTCTCGTCCAAGAAACGGATTCAGGCCAGAAGCCGGTCTATTTC[ G / A ]TCTCCGAAGCATTGCAAGGAGCGAAGACAAGATATGTTGAAATGGAAAAGCTCGCTTATGCCCTGGTGATGGCTTCACGCAAGTTCAAGCACTACTTTCAGGCCCACAAAGTCATAGTGCCATCGTAGTACCCTTTGGACGAAATACTCC(서열번호 7) chr04:674276 GCCGCCATCCTCCCAAAACCCCTTCCGAAAGGCCCTTAACGCAGCTCCGGAGCACTGCGTGTTGCTAGGCCATTCTTCCTTCGCGAAATCAGGAAAATTCCCGATGTGGTCGCAGCTGTGCGCGTGCAAGAGGCTGAGAACGAATCCTGC[ C / T ]GAAACACGAGCACAACAGTCCCCGTAGGCGCCCGCCACTTCGACAACTGAGCCGGCCGCCTCCTGCGTCCATTCAGAGAAATCGAGTGTAGTCCCGCCTTCGCCCGGAGCACCCTCCGCCCGCGCCCTAAGGTCATTAAGGGCCAGCCGG(서열번호 8) chr04:691278 ATTTCAGGAGCTTGGATCGGACACTTGCATGATAGTACATGAAGAAGTACTAGGTCCCAAACCAACACAGAAAGAATAGAATAATCTGGAAAAGAATTGAAGAAACAAACTCGTTAGTTTTGAATTTGAGTTTGATTTGAAATATTTGGT[ C / T ]GAATCCAAATCGTGCTTACGCTTGTAGCTAAAGGCATGCGACAGATGGAAGTCTAGTACATCAATCCAACCACGCTTAGAAATGTTTATGGTGCATGGCTTTGGAATAGTCAATTACCGCAAAAAATGGGACATAATTCCCAAACACGAG(서열번호 9) chr04:697800 TTATAATTCTGACAAATGTCATGATTCTTAAATAGAGAACTACATATTTTACAATTACTAATTTTGAATTTAGAGGTATTAATCTAGCCAACTGTTTAGTCCTTATATATATACTTACAAAGTAGTACCATACACTTACCACGGACTTTC[ C / G ]TATCATATAGTTAACGATCAAATTTGGTAAATCATATGAACCAGGTTCGACATAGGAATCAAAGTGGATTCGGCGAAGGATTGTACTCAAAGAACAGAGTGCGTATCGGCTGCGAGGGCATCGGCTGAGGATCGCATCGGCTAAGATGGA(서열번호 10) chr05:16878762 ACATACGCAAAGCGCAAAAAGAAAACTTTACAGGTTCAAGGCCTTTGGCCCGAACATGTTAGAAGGTAAGCCACCTGCCGAGATGTTTTGCGCCCGCCAAGACCCATGTTTTGGCCTCCTCCTTAATTTTGGCTATCAGACTAGCCACCG[ T / G ]GAGCTCTTGGTGCTAGAAGACTCTGTGGTTCCTTTCGTTCCAAATCTCCCAAGCGACTAATAGCATTATAGTCCACAAGGCCTTCTTTGGCACTTCATTGGTACATGCGGTGGCCTCCCACCACTTGAGTAGTGTCGAGGCTTGTTGCTA(서열번호 11) chr05:24095918 TAAGCAATCGGCTATATGTCCAATGTAGATAATGATATAAAGGCAATCGGCTGATGATGATGTAATAAAATAATAATATAATCCAGTATAAACCAATCGGCTAGCATTGATATGATAAAAAAGCACTAATCCGATAGTTAAAGCATACAT[ C / T ]GGCTGGAGGTCTGATGTCATGAAATCCACAAGATTAGATTAAACAGTGAAACATTTGTTGTCATCGGCTAAATCCAACTTATATGTATATGCAATCCTTATGAGCCGATGCAACGTCCAGATAACTCACCAGCTGAAACCCTTATTGGCA(서열번호 12) chr05:29327030 TCCGGTTTCTATCCTCTTGTTGGGCCTTGCGCTCAATGCCAGGCGGCCCAGTTAAACTACACGCATGACAGCCCATCACAGTTTGGACATCCAGATGGGCCAAATTGGAATATAGCGATCCATTCGGCCCATCATATTGTAAGGCCCATC[ T / A ]TATGTATACTGAACAGAAGGGAAAGCCCACGACGCACCGGTACTGAGCACCGCAGCCACGGCGCCGCCGGCTGCTCGCCGGCGAAGGGAGCAGCCACGGCGCCGCCCACCACCGCCTCGAAGGCCACCTACTGCACTCGTCGAGGTCTTG(서열번호 13) chr05:29512997 ATTCTCTTTAGCTTACAAAGCAATGAGACTGCTTTTATAACCTTTTATAGACATAACATTGGCTTGTCAAATGGTTCAAATTGAATGTACTGTAAAATTATGGATCACTAGTCCCACACAATGCTATCCAGTCGTTGTCTAGCTGATTCC[ G / A ]TGTAGCATGGAAGGTCCTCATGGCCACCAGCTGCTCCTTGATAAAGGAAGATTGAGACTTCAAGAGTTAATCCTTGCCAAGCCAATGTTATGTCTACGGAGTCACTGCTCTCTGTGTCTATTGAGCCTCTACGCTGATGAGAGGGTCAAC(서열번호 14) chr05:29682235 AAAGATAGTCCTCCAGTACTAGCCTAGTAGTTGTTCTTGTTGGGCGTGGCGGACGACCAGTCAGTGCATCGGTCACCGGTGCATGAGAATGTACTGTTACTATGACATGAAGCAGCAGCAGCCACAGCAGTGCTGCGTGCCGATTCATGC[ C / G ]AGTACTAGACAAGAGAGGGACATCACAAGGTATGCAGAAGCCGGAAACTATTTTCATCCCTGGAGGGGATATTCCCTCGTTGTATGCATGTCACTTAAATGATTATGAAAAAAAATTAAAAAATTTGAGAAGATGTATTAAGATGTGATA(서열번호 15) chr08:15824488 TCCTGTTACCACAAAAATACTAGTTCCTAGGTAGGCCCCGCGCCAACTAAAAGGTACCGAATTTTCAATAAAAACCACCATCTTTCTCTGACACCAGTATAGGTGCATGTTCTTTTTAAAAACCAACACCTTTTTGATGATAAAGGTGCC[ G / A ]GTTCTTTGGATTTTGGCGGGCAGATGGGCTCGTGGGTGGGGAAAGGCGCTATAGGTACCGATTTTACCTACTCTGGCACCTATAGTGCTGACAGCGAATTGATGTTTTGCAGTAGTGAGATGTTGTGTGTATATGGGAGGTAGGATCAGG(서열번호 16) chr08:8510533 TCCTCTCTTTTGTCCTCTTATTCATATCATTAGTCTTTTGCCTAAAATAGTCTATCCATGTTCAGTTATAGAGTATAAATTTATGCTTATTCTATAGCTTGCTCAACCATATAAATGTTCGCCTGTTTTACATGGCAAAATTGTTTAGCA[ T / C ]TTTCATGGTACTCCCTCCGTCCCACAATATAAGGGATTTTCAGTTTTTGCTTGTAACGTTTGACCACTCGTCTTATTCAAATTTTTTTTGCAAATATAAAAAATGAAAAGTTGTGCTTAAAGTACTATAGATAATAAAGTAAGTCACAAA(서열번호 17) chr08:8653144 GTGTGGCATCGGCAAGCGGTGTGCAGGTGGCGGCATCGACCGGCGACAGTCTAGGCGGCGTCGATAAGCGGCGGCGGAGATGCACGGGCTACCTAGAGGCAACAAAGGCAGGGTGGCACCTACAAGGTGTTCGACGAAATAACGACGAGG[ G / A ]ATGTGGTGGCGTGGAGCTGAGTAATCTAGCTAAACACTTTCTATACCCACCCAGCTTCTATAGGAGTCATTTCTCTTTGGAGTTAGAGTTTAGAGTTGAGAATTGAAATTTGAAGCTCTACCAAACGGGGTCCATGTATGTATACTTCCA(서열번호 18) chr11:3551290 TCCCTTCCTGAATGGCTTGAGAATCTTAAAAAGTTACAAAAGTTAACCTTGCAAAACAAAAATTTCCCTTCCTGAATTGCAGTTATTCCGTCATCTGTTTCGAATCTGTCTCAACTAGCAGTTCTCGGACTATATTCTAACAAGTTGGAA[ C / G ]GGCACATACCAAGCTTCGTAAACCTGCAAATGCTTCAACTACTAGTCATTTCCAGTAATAATCTTCATGTCAGCATACCAAAAGAGATATTTAGTATTCCATCAATAATTGCAATTGATTTATCTTTCAACAACCTAGATGGGCAACTTC(서열번호 19) chr11:3649182 AGTTTGTGTGTGGCATCGGATAGAATTGAAATTTATTTTTTCTTGCGTATGGAAATTCAAAATGAATGTGATCGATCATGGACTAGATAGAGAAATTCAAACTGCCACTGGTTTCTCATGTCGGCAGGGATGAAACAGGAGAACACCAGC[ G / A ]GTGTCTGTGTGCTGTAGATAGAGATAGACCCAAATACAAGTTTCAGTAATCTGGAGGCATGTACACTTGTGCAGTCGATTCAGATGCCTATAATTAGGCAACTCATGTATGCATCGAGTAGTTGAAGAATACTAGCTCAAATTGTGTGTG(서열번호 20) chr11:4090841 CAATGGCGAGTCTTTTACTACTCCCTTCGTTTCACAATGTAAGTCATTCTAGCATTTCTCACATTCATATGGATTTTAATGAATCTAGATAGATATATATGTCTAGATTTATTAATATCGATATGAATGTGGAAAATGCTAGAATGACTT[ G / A ]TATTGTGAAACGGAGGGAGTAGTAAGCAAAACTTTGTTTCTTTCTCATTTCTTCATGTAAATTGTAGTGCTAATGAGATGCTCATCTGTTGGCTAGAACGAGGGAGTGTGTTTCTTATTTAACTTTTCGACAGTGTGTTACTGGTTTTAT(서열번호 21) chr11:5462708 TAGTTGCATCATAAGTTAAACGAAGACATACAATCATTCAATCATTACGAACTCTGAAATTAGGTGATGGCAACAGTTAGATTAATTGTTAGTATTATTTTCCAGTCATAAACCACTGTTCTGAATTTGTTAAGAGGGAAACACTGTGGT[ A / G ]TATATTGTATGAAAATCTGATTTTAATTTGCTATTTTATATATTTTTGCCTGAGTATGCTTCATTATTCAATACAGTCTAGTGAGACGCTTTCAGTCTATCTCGAGTATGTTTCTGGGGGCTCTATCCATAAGTTGCTTCAAGAGTATGG(서열번호 22) chr11:9040443 GTGCGGTATCTCAGTGATTAGAAGGTAATAGCATCTTTATATTTTGAGTTTGTACTTTCTTGAGACTAACTGTTATACCACTTGGCACTCAAGGGAGGTTGGAAGGTTCTTTTTTTTTTACCACTCTGCTGTCATGTATGTCAGTGATAC[ T / C ]GTTATGTTGCTTCTTTGCTTTGCCTGATAGCTTGTGCAGAATGATATAAATGGAGCACCATGGCCATTTACAAATGCTCTACTAGAAAATATTGGAAACTTTGATTCAGTTTCCTTCAGTTGCTTTACGGGCTTCACATACCCAATACTG(서열번호 23) chr12:10244769 ATTAGTTTTATGTACCAAAATAAAAGTTCAAATATTAATCTGTCATCCTATCTATTTATCTATTTGTTTTTTTATATCGTGTCTTGGCTAGGATTCAAACTCAAGACCTATCTTTCATGCGTGCGCTCCTTTACCAACTCACCTACGCAT[ T / C ]GTTAGAAAATGTATATCGTTCCATTTGAGTCTACCTAACTGAGATTTGAATCATAGATTTGAATATCTAAAAGATTTCAAAAGAAAAAGTTATCACCTATAAAATTGTAGATCTCGTCGAGAGCTATAATTTCAATATAAAATTTGTCTC(서열번호 24) chr12:10271635 CGCCGCTCCGACCCAAGCCCGCCACCGCCGCCGCCGCTCCAGCCCTCATCCGCCGTCATCGCCGCCGCCGTTCAAACGCAGCAAAGGGGGAGAGGAAGGGAATGGAAGGATGGTTGTGGGCCCACGTGTAAGTGGGTCCCACTTTTTTTT[ C / T ]GATTGACAGATGGGTCCGTATATATTTTTTTAAATTCTAATGCCATGTCAGCGCCACGTTGGACGATGACCAGGTCAACACTGACACGTAAGCGCCACGTCAGCGAAACCGCCCTTCAAAACCGCCGAGGGAGTCAAATTGCACCAGTTT(서열번호 25) chr12:10866960 GTTTTCTTTTTCTTTTGGCTCAGGTAACTTATCTCGACGTTTAACAGAGGGATACTGAGTAGTTTTGACGACATTTGCAGCAGGGTCTGTATGTTCTAATCCATTGCAACCATTTGTCAAAAGTTCTCTCCTTCGACATGAAGCACTTCG[ A / T ]AGAGACTCTGCAGACAAGAAGTCTCTTGTATCAAAGTACATCACTTCATCTTCATCAGTTTCAACATCTGCCACTTGACTTGCAGTATCTGGATCAGACTCGCTTGCACTTCCTCCTGATAAAACTGAATAAAAATCTGCCAATAAAGCA(서열번호 26)

[0140] In addition, the results of obtaining phenotypes related to the protein content of rice seeds for various varieties by inputting genotype information for 26 markers indicated by the selected sequence numbers 1 to 26 into a prediction model through machine learning are shown in Fig. 5 (Control - bottom 30% low content group; Case - top 30% high content group).

[0141] Meanwhile, among the two bases in the SNP sites of the 26 markers indicated by sequence numbers 1 to 26, the first base is the reference gene and the second base is the alternative gene, and the bases represented by the SNPs of the 26 markers according to the variety are shown in FIG. 5.

[0142] For example, as shown in FIG. 5, even if a rice seed belongs to the high protein content group, the base of the SNP site of the marker indicated by SEQ ID NO. 1 may be A (e.g., Nunkeun Heukchal No. 1) or G (e.g., Dapyeong), and depending on the variety, the corresponding marker may not exist (e.g., Giho). Therefore, the base of the SNP marker of each marker provided in the present invention does not determine whether the protein content of the rice seed is high or low, and the protein content of the rice seed can be predicted by inputting the genotype of each of the 26 markers into a prediction model to derive a comprehensive result.

[0144] [References]

[0145] Chen, Pingli, et al. “The genetic basis of grain protein content in rice by genome-wide association analysis.”Molecular Breeding43.1 (2023): 1.

[0146] Kang, Min-Jeong, et al. “Identification of transcriptome-wide, nut weight-associated SNPs in Castanea crenata.”Scientific reports9.1 (2019): 13161.

[0147] Yu, Go-Eun, et al. “Machine learning, transcriptome, and genotyping chip analyzes provide insights into SNP markers identifying flower color in Platycodon grandiflorus.”Scientific Reports11.1 (2021): 8019.

Claims

Claim 1 1) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 1 is A or G; 2) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 2 is T or A; 3) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 3 is T or C; 4) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 4 is G or A; 5) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 5 is T or C; 6) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 6 is A or C; 7) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 7 is G or A; 8) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 8 is C or T; 9) A single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 9 is C or T Single nucleotide polymorphism; 10) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 10 is C or G; 11) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 11 is T or G; 12) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 12 is C or T; 13) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 13 is T or A; 14) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 14 is G or A; 15) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 15 is C or G; 16) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 16 is G or A; 17) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 17 is T or C Single nucleotide polymorphism; 18) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 18 is G or A; 19) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 19 is C or G; 20) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 20 is G or A; 21) Single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 21 is G or A;A composition for determining high or low protein content in rice seeds, comprising a preparation capable of detecting: 22) a single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 22 is A or G; 23) a single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 23 is T or C; 24) a single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 24 is T or C; 25) a single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 25 is C or T; and 26) a single nucleotide polymorphism in which the 151st base of the sequence represented by SEQ ID NO. 26 is A or T.; Claim 2 A composition according to claim 1, wherein the formulation capable of detecting the single nucleotide polymorphisms of 1) to 26) is a primer pair or probe capable of detecting each of 1) to 26). Claim 3 A kit for determining high or low protein content in rice seeds, comprising the composition of claim 1 or 2. Claim 4 A method for determining whether the protein content of a rice seed is high or low, comprising: i) a step of isolating genomic DNA from rice; ii) a step of determining all genotypes of 1) to 26) of claim 1 using the isolated genomic DNA as a template; and iii) a step of determining whether the protein content of the rice seed is high or low using a model machine-trained with a machine learning algorithm selected from the group consisting of SVM, KNN, FR, C5, and PLS, based on the genotype information determined in step ii).