Rice thousand seed weight prediction model, SNP molecular marker related to rice thousand seed weight and construction method and application of SNP molecular marker

By constructing a rice thousand-grain weight prediction model and screening relevant SNP molecular markers, the problem of improving rice thousand-grain weight was solved, rice yield was increased, and the effectiveness of molecular breeding was realized.

CN120853671APending Publication Date: 2025-10-28HUAZHONG AGRI UNIV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510695743.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies cannot efficiently utilize molecular marker technology to improve the thousand-grain weight of rice, thus affecting rice yield.

Method used

A rice thousand-grain weight prediction model was constructed. SNP molecular markers were screened using Spearman correlation coefficient, and a regression model was constructed using Lasso regression. SNP molecular markers with high absolute values ​​of weight coefficients were screened, and SNP molecular marker combinations related to rice thousand-grain weight were constructed by combining SNP loci at specific chromosomal locations.

Benefits of technology

It has enabled efficient prediction and improvement of the thousand-grain weight of rice, increased rice yield, and has important application value in molecular breeding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120853671A_ABST
    Figure CN120853671A_ABST
Patent Text Reader

Abstract

The invention provides a rice thousand seed weight prediction model, an SNP molecular marker related to rice thousand seed weight and a construction method and application of the SNP molecular marker. The construction method of the rice thousand seed weight prediction model comprises the following steps: S1, acquiring hybrid rice sample data, and constructing a data set according to thousand seed weight information of rice and SNP molecular markers; s2, based on the training set data, using a Spearman correlation coefficient to preliminarily screen SNP molecules related to the thousand seed weight of rice; s3, using Lasso regression to further screen SNP molecules related to the rice thousand seed weight, and fitting model parameters to construct a regression model, namely a rice thousand seed weight prediction model; and S4, verifying the model on the verification set. And on the basis of the obtained model, screening out the SNP molecular marker of which the absolute value of the weight coefficient is higher than a certain threshold value as the SNP molecular marker related to the thousand seed weight of rice. The screened SNP molecular marker can be used in rice yield trait molecular assisted breeding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of rice molecular breeding technology, specifically to a rice thousand-grain weight prediction model, SNP molecular markers related to rice thousand-grain weight, their construction methods, and applications. Background Art

[0002] Rice is one of the world's most important food crops, with over 60% of the population relying on it as a staple food. Further increasing rice yield is crucial for ensuring global food security. Rice yield is determined by the number of effective panicles per plant, the number of grains per panicle, and the thousand-grain weight, with thousand-grain weight being a key factor directly affecting yield. Grain shape, including grain length, width, length-to-width ratio, and thickness, directly influences thousand-grain weight. Therefore, cloning grain shape-related genes and identifying superior alleles are important pathways to improve grain shape and thousand-grain weight, thereby increasing rice yield.

[0003] Molecular markers and molecular identification are important auxiliary tools in rice molecular breeding. Molecular marker technology has advantages such as high specificity and ease of use. It is mainly applied to improve rice yield, rice quality, and resistance to insects and diseases. By using molecular marker-assisted breeding technology, high-yielding QTLs from wild rice were transferred into superior late-season rice restorer lines, thereby increasing rice yield.

[0004] In recent years, with the continuous development of molecular biology techniques, SNP (single nucleotide polymorphism) molecular markers, as high-precision and high-density genetic markers, have been widely used in plant gene mapping, variety identification, and genetic improvement. In rice research, identifying SNP molecular markers related to thousand-grain weight can provide important theoretical basis and technical support for molecular breeding of rice. Summary of the Invention

[0005] To address the problems existing in the background art, this invention provides a rice thousand-grain weight prediction model, SNP molecular markers related to rice thousand-grain weight, their construction methods, and applications. The technical solution of the present invention to solve the above-mentioned technical problems is as follows: In a first aspect, the present invention provides a method for constructing a rice thousand-grain weight prediction model, comprising the following steps: S1. Obtain hybrid rice sample data from the database, and construct a dataset based on the thousand-grain weight information and SNP molecular markers of rice, including a training set and a validation set; S2. Based on the training set data, Spearman correlation coefficient was used to preliminarily screen SNP molecular markers related to the thousand-grain weight of rice; S3. Use Lasso regression to further screen SNP molecular markers related to the thousand-grain weight of rice and fit the model parameters to construct a regression model, which is the rice thousand-grain weight prediction model. S4. Validate the model obtained in step S3 on the validation set.

[0006] According to the above scheme, in step S2, the correlation between SNP sites and the thousand-grain weight of rice is calculated using the Spearman correlation coefficient to obtain the correlation and significance of the SNP molecular markers with the thousand-grain weight. Based on the significance and FDR, SNP sites as candidate features are selected.

[0007] According to the above scheme, in step S3, a 5x cross-validation strategy is adopted.

[0008] Secondly, the present invention provides a rice thousand-grain weight prediction model, which is constructed by the above-mentioned method for constructing the rice thousand-grain weight prediction model.

[0009] Thirdly, this invention provides a method for screening SNP molecular markers related to the thousand-grain weight of rice, which, based on the above-mentioned method for constructing a thousand-grain weight prediction model for rice, further includes the following steps: S5. Based on the obtained rice thousand-grain weight prediction model, SNP molecular markers with absolute values ​​of weight coefficients higher than a certain threshold are selected as SNP molecular markers related to rice thousand-grain weight.

[0010] Fourthly, this invention provides SNP molecular markers associated with the thousand-grain weight of rice, obtained by the screening method for the aforementioned SNP molecular markers related to the thousand-grain weight of rice, and selected from SNPs at the following chromosomal locations:

[0011] The rice genome reference version is: Nipponbare Reference Genome Os-Nipponbare-Reference-IRGSP-1.0.

[0012] According to the above scheme, the polymorphism of the SNP molecular marker is as follows:

[0013] Fifthly, the present invention provides an SNP molecular marker composition associated with the thousand-grain weight of rice, which is any of the following combinations: (I) At least two of the following SNP sites: 3_22238811, 2_25132698, 7_24300032, 1_5426089, 1_32235551, 6_23599240, 3_29315748, 5_5457754, 11_25650738, 3_30362482; (II) At least two of the following SNP sites: 3_591226, 5_27909347, 10_18384112, 11_28774264, 4_27376521, 12_1192083, 2_11134239, 11_2643669, 5_22996493, 4_31361948.

[0014] According to the above scheme, it can be any of the following combinations: (I) At least two of the following SNP sites: 3_22238811, 2_25132698, 7_24300032, 1_5426089, 1_32235551, 6_23599240, 3_29315748, 5_5457754, 11_25650738, and 3_30362482; (II) At least two of the following SNP sites: 3_591226, 5_27909347, 10_18384112, 11_28774264, 4_27376521, 12_1192083, 2_11134239, 11_2643669, 5_22996493, and 4_31361948.

[0015] In a sixth aspect, the present invention provides any of the following applications of the above-mentioned SNP molecular markers associated with thousand-grain weight of rice and the SNP molecular marker compositions associated with thousand-grain weight of rice: (I) Predicting the thousand-grain weight of rice; (II) Identification and screening of high-yield rice varieties and strains; (III) Molecular-assisted breeding of high-yield rice; (IV) Improve rice germplasm resources.

[0016] The beneficial effects of this invention are as follows: This invention constructs a rice thousand-grain weight prediction model through Lasso regression, and mines SNP molecular markers and SNP marker combinations that are significantly associated with rice thousand-grain weight based on the model. These SNP markers are significantly associated with rice thousand-grain weight and can be used for molecular-assisted breeding of rice yield traits. This is of great significance for improving rice yield and enhancing the overall economic benefits of rice through molecular breeding. Attached Figure Description

[0017] Figure 1 This is a scatter plot of the predicted and actual thousand-grain weights of the input validation set data after the Lasso regression model was successfully fitted in Embodiment 1 of the present invention. Figure 2This is an identification diagram of molecular markers on the third pair of chromosomes of rice in Example 2 of the present invention that are related to the thousand-grain weight of rice. The red dots are the selected molecular marker 3_22238811. Figure 3 The dataset in Example 1 of this invention has been cleaned, and the correlation values ​​and distribution diagrams of molecular marker sites related to the thousand-grain weight trait in the whole rice genome are shown. Figure 4 Box plot of molecular marker 3_22238811 and rice thousand-grain weight in Example 2 of this invention; Figure 5 This is the SNP molecular marker density map corresponding to the dataset in Embodiment 1 of the present invention; Figure 6 Box plot of molecular marker 2_25132698 and rice thousand-grain weight in Example 2 of this invention; Figure 7 Box plot of molecular marker 3_591226 and rice thousand-grain weight in Example 2 of this invention; Figure 8 Box plot of molecular marker 5_27909347 and rice thousand-grain weight in Example 2 of this invention; Figure 9 Box plot of molecular marker 10_18384112 and rice thousand-grain weight in Example 2 of this invention; Figure 10 Box plot of molecular marker 11_28774264 and rice thousand-grain weight in Example 2 of this invention; Figure 11 Box plot of molecular marker 7_24300032 and rice thousand-grain weight in Example 2 of this invention; Figure 12 Box plot of molecular marker 4_27376521 and rice thousand-grain weight in Example 2 of this invention; Figure 13 Box plot of molecular marker 1_5426089 and rice thousand-grain weight in Example 2 of this invention; Figure 14 Box plot of molecular marker 12_1192083 and rice thousand-grain weight in Example 2 of this invention; Figure 15 Box plot of four SNP molecular marker combinations (3_22238811, 7_24300032, 1_5426089 and 1_32235551) and rice thousand-grain weight in Example 2 of the present invention; Figure 16 This is a box plot of two SNP molecular marker combinations (3_22238811 and 7_24300032) and the thousand-grain weight of rice in Example 2 of the present invention. Detailed Implementation

[0018] The principles and features of the present invention are described below with reference to the accompanying drawings and specific embodiments. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0019] Example 1: Construction of a Rice Thousand-Grain Weight Prediction Model I. Downloading and Cleaning Data Hybrid rice sample data were downloaded from the public database GropGS-Hub (https: / / iagr.genomics.cn / CropGS). This data includes samples from three hybrid combinations: indica rice × indica rice, japonica rice × japonica rice, and indica rice × japonica rice, totaling 1495 hybrid samples and 1,651,507 SNP molecular markers. The density map of SNP sites in the genome is shown below. Figure 5 As shown. Then, Plink (v1.9) was used to convert the original VCF file into a Plink binary Bed file, filtering out SNP molecular markers with genotype deletion rates higher than 5% and minor allele frequencies (MAF) lower than 5%. Further linkage disequilibrium (LD) filtering was performed, with a step size of 1, a window span of 1000kb, and r... 2 <0.5, SNP molecular markers with excessively high linkage were screened out, and a total of 99236 SNP molecular markers were obtained after screening.

[0020] II. Preliminary screening of candidate features Initial data screening involved standardizing 99,236 SNP molecular marker data according to their variation. Then, for each SNP molecular marker, the Spearman correlation coefficient was used to calculate its correlation with the trait, thus obtaining the correlation (R) and significance (p) between the molecular marker and thousand-grain weight. Molecular markers with a significance p-value < 0.05 and FDR < 0.05 were selected as candidate features, resulting in 1109 candidate molecular markers. The distribution of molecular markers on chromosomes and their correlation p-values ​​obtained through the significance p-value < 0.05 screening are shown below. Figure 3 As shown.

[0021] III. Construction of Lasso Regression Model Lasso regression was used for feature selection and parameter fitting of candidate features. Lasso regression introduces L1 regularization (parameter λ), causing the coefficients of some genes to become zero, thus achieving further gene selection. To select the optimal regularization parameter λ, a 5-fold cross-validation strategy was adopted in the dataset, using Lasso regression to further filter and select hyperparameters from the 1109 molecular markers obtained above. Specifically, the training set was divided into 5 subsets. In each round, the model was trained on 4 subsets and then tested on the remaining subset. This process was repeated 5 times, and the results of the 5 cross-validations were averaged. During each cross-validation training process, the Lasso model found the optimal model by adjusting the regularization parameter λ. This parameter controls the degree of penalty on the coefficients; an excessively large λ will cause more coefficients to be compressed to zero, achieving feature selection. Through 5-fold cross-validation, the regularization parameter λ that resulted in the best performance of the model on the test set was obtained. Next, using all the samples and the optimal parameter λ, a regression model is built on the training set and saved, which is the rice thousand-grain weight prediction model.

[0022] Meanwhile, the coefficients of the minimum residual positions were obtained. These coefficients correspond to the selected genes and indicate the degree of contribution of these genes to the final model. A total of 165 SNP molecular markers were finally screened out.

[0023] IV. Evaluating Model Performance Using the rice thousand-grain weight prediction model obtained from the training set, the thousand-grain weight of rice was calculated on both the training and independent validation sets. A scatter plot was created comparing the predicted and actual thousand-grain weights, and the Pearson correlation coefficient, corresponding p-value, RMSE (root mean square error between predicted and actual thousand-grain weights), and R² were used to calculate the predicted and actual thousand-grain weights. 2 The fit was evaluated. The results of evaluating the predicted thousand-grain weight and the actual thousand-grain weight of the above model on the validation set are as follows: Figure 1 The Pearson correlation coefficient between the predicted and actual thousand-grain weights was 0.798 (P-value = 4.16e-65), the root mean square error between the predicted and actual thousand-grain weights was 3.89, and the R-squared value between the predicted and actual thousand-grain weights was... 2 A value of 0.62 indicates that the model has good predictive performance on the validation set.

[0024] Example 2: Screening of SNP molecular markers related to thousand-grain weight in rice The SNP sites in the rice thousand-grain weight prediction model obtained in Example 1 were ranked according to the absolute value of their weight coefficients. The top 20 SNP molecular markers were selected as SNP molecular markers related to rice thousand-grain weight, as shown in Table 1 below. Table 1

[0025] The rice genome reference version is: Nipponbare Reference Genome Os-Nipponbare-Reference-IRGSP-1.0.

[0026] The Spearman correlation coefficients and p-values ​​of the SNP molecular markers related to the thousand-grain weight of rice obtained above are shown in Table 2.

[0027] Table 2

[0028] The top-ranked molecular marker, Chr3_22238811, is located at position 22238811 of the third nucleotide sequence of the rice reference genome version, Nipponbare Os-Nipponbare-Reference-IRGSP-1.0. Figure 2 As shown in Table 3, the data corresponding to the corresponding sample data at this site are classified and organized. This SNP molecular marker exhibits G / A polymorphism. The thousand-grain weight of rice corresponding to this polymorphism is shown in Table 3. The thousand-grain weight is lowest when the site is G / G, moderate when it is G / A or A / G, and highest when it is A / A. The box plots corresponding to these phenotypic values ​​are shown in Table 3. Figure 4 As shown in the figure. ANOVA (or Kruskal-Wallis H test if each group of data is not normally distributed) was used to test whether there was a significant difference in the thousand-grain weight of rice among the three different genotypes. The results showed that the P-value was 3.106e-27, indicating that this locus is indeed an important marker locus associated with the thousand-grain weight of rice.

[0029] Table 3

[0030] Molecular markers ranked 2nd to 10th (2_25132698, 3_591226, 5_27909347, 10_18384112, 11_28774264, 7_24300032, 4_27376521, 1_5426089, 12_1192083) were selected. Based on the sample data corresponding to the loci, the base polymorphism of this SNP molecular marker was classified and organized in the obtained dataset. The box plots corresponding to the phenotypic values ​​are shown below. Figure 6-14As shown. It is proven that the nine SNP loci, 2_25132698, 3_591226, 5_27909347, 10_18384112, 11_28774264, 7_24300032, 4_27376521, 1_5426089, and 12_1192083, are also SNP loci associated with the thousand-grain weight of rice.

[0031] SNP molecular markers positively correlated with thousand-grain weight of rice include: 3_22238811, 2_25132698, 7_24300032, 1_5426089, 1_32235551, 6_23599240, 3_29315748, 5_5457754, 11_25650738, and 3_30362482; The SNP molecular markers associated with the thousand-grain weight of rice include: 3_591226, 5_27909347, 10_18384112, 11_28774264, 4_27376521, 12_1192083, 2_11134239, 11_2643669, 5_22996493, and 4_31361948.

[0032] Combining at least two of the SNP molecular markers positively correlated with rice thousand-grain weight yields an SNP molecular marker composition that can be used to predict rice thousand-grain weight. For example: The following four SNP sites, 3_22238811, 7_24300032, 1_5426089, and 1_32235551, which are positively correlated with thousand-grain weight, were integrated. The corresponding data in the dataset were then categorized and organized to identify the base polymorphisms of these four SNP molecular markers. The results are as follows: Figure 15 As shown in the figure, "allunalt" corresponds to no variation at any of the four loci, "mixed" corresponds to single-locus variation at all four loci, and "allpure" corresponds to complete variation at all four loci. Compared to the change in thousand-grain weight trait value caused by single-locus variation, the result of multiple loci working together is significantly better than single-locus variation. In practical molecular breeding applications, variations at multiple loci can be combined to directionally improve rice yield.

[0033] The two SNP sites 3_22238811 and 7_24300032 were integrated, and the corresponding data from the dataset obtained in Example 2 were classified and organized to analyze the base polymorphisms of these two SNP molecular markers. The results are as follows: Figure 16 As shown, the combination of two SNP sites is more effective than the change in thousand-grain weight trait value produced by single point variation.

[0034] Alternatively, at least two of the SNP molecular markers associated with the thousand-grain weight of rice can be combined to obtain an SNP molecular marker composition for the prediction of the thousand-grain weight of rice.

[0035] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for constructing a rice thousand-grain weight prediction model, characterized in that, Includes the following steps: S1. Obtain hybrid rice sample data from the database, and construct a dataset based on the thousand-grain weight information and SNP molecular markers of rice, including a training set and a validation set; S2. Based on the training set data, Spearman correlation coefficient was used to preliminarily screen SNP molecular markers related to the thousand-grain weight of rice; S3. Use Lasso regression to further screen SNP molecular markers related to the thousand-grain weight of rice and fit the model parameters to construct a regression model, which is the rice thousand-grain weight prediction model. S4. Validate the model obtained in step S3 on the validation set.

2. The method for constructing the rice thousand-grain weight prediction model according to claim 1, characterized in that, In step S2, the correlation between SNP sites and thousand-grain weight of rice is calculated using Spearman correlation coefficient to obtain the correlation and significance of the SNP molecular markers with thousand-grain weight. Based on the significance and FDR, SNP sites are selected as candidate features.

3. The method for constructing the rice thousand-grain weight prediction model according to claim 1, characterized in that, In step S3, a 5x cross-validation strategy is adopted.

4. A rice thousand-grain weight prediction model, characterized in that, It is constructed by the method described in any one of claims 1 to 3.

5. A method for screening SNP molecular markers related to the thousand-grain weight of rice, characterized in that, In addition to the method described in any one of claims 1 to 3, the method further includes the following steps: S5. Based on the obtained rice thousand-grain weight prediction model, SNP molecular markers with absolute values ​​of weight coefficients higher than a certain threshold are selected as SNP molecular markers related to rice thousand-grain weight.

6. An SNP molecular marker associated with thousand-grain weight of rice, characterized in that, The SNP molecular markers are obtained by screening using the method described in claim 5, and are selected from SNPs at the following chromosomal locations: The rice genome reference version is: Nipponbare Reference Genome Os-Nipponbare-Reference-IRGSP-1.

0.

7. The SNP molecular marker associated with thousand-grain weight of rice according to claim 6, characterized in that, The polymorphism of the SNP molecular markers is as follows:

8. A SNP molecular marker composition associated with thousand-grain weight of rice, characterized in that, It can be any of the following combinations: (I) At least two of the following SNP sites: 3_22238811, 2_25132698, 7_24300032, 1_5426089, 1_32235551, 6_23599240, 3_29315748, 5_5457754, 11_25650738, 3_30362482; (II) At least two of the following SNP sites: 3_591226, 5_27909347, 10_18384112, 11_28774264, 4_27376521, 12_1192083, 2_11134239, 11_2643669, 5_22996493, 4_31361948.

9. The SNP molecular marker composition associated with thousand-grain weight of rice according to claim 8, characterized in that, It can be any of the following combinations: (I) At least two of the following SNP sites: 3_22238811, 2_25132698, 7_24300032, 1_5426089, 1_32235551, 6_23599240, 3_29315748, 5_5457754, 11_25650738, and 3_30362482; (II) At least two of the following SNP sites: 3_591226, 5_27909347, 10_18384112, 11_28774264, 4_27376521, 12_1192083, 2_11134239, 11_2643669, 5_22996493, and 4_31361948.

10. Any of the following applications of the SNP molecular markers associated with thousand-grain weight of rice as described in claim 6 or 7, or the SNP molecular marker compositions associated with thousand-grain weight of rice as described in claim 8 or 9: (I) Predicting the thousand-grain weight of rice; (II) Identification and screening of high-yield rice varieties and strains; (III) Molecular-assisted breeding of high-yield rice; (IV) Improve rice germplasm resources.

Citation Information

Patent Citations

  • SNP marker relevant with thousand seed weight of millet and detection primer and application of SNP marker

    CN108728566A

  • SNP (Single Nucleotide Polymorphism) marker capable of improving cotton fiber strength from Zhongmiansui 70 population

    CN114317795A

  • Breeding cross-representative prediction method and system based on ensemble learning, and electronic equipment

    CN116580773A

  • Japonica rice whole genome SNP (Single Nucleotide Polymorphism) site combination, gene chip and application thereof

    CN116694806A

  • SNP (Single Nucleotide Polymorphism) molecular marker related to rice panicle number and application

    CN118703691A