Method for excavating glutinous rice wine brewing related genes

Through the integration of multi-omics data and CRISPR-Cas9 technology, a mixed linear model was constructed to reveal the complex relationship between glutinous rice gene regulation and phenotypic characteristics, solve the problem of low efficiency of glutinous rice breeding, and achieve the improvement of the quality of glutinous rice raw materials and the improvement of rice wine quality.

CN120748482APending Publication Date: 2025-10-03SHAOXING ACAD OF AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510643957.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Existing glutinous rice breeding technology has problems such as long breeding cycle, low efficiency, poor accuracy, strong uncontrollability, high cost and poor repeatability. In addition, existing genome analysis methods fail to effectively reveal the complex relationship between glutinous rice gene regulation and phenotypic characteristics, resulting in insufficient exploration of glutinous rice winemaking-related genes.

Method used

A multi-omics data integration method was used to integrate genomic, epigenomic and proteomic data by constructing a mixed linear model, and CRISPR-Cas9 technology was used to knock out or overexpress candidate genes to verify their functions, optimize the particle size and roundness of glutinous rice, and improve the quality of glutinous rice raw materials.

Benefits of technology

It significantly improves the efficiency of glutinous rice breeding, provides an efficient and accurate selection tool, improves the quality of glutinous rice raw materials, meets the market demand for high-quality rice wine, and shortens the functional research cycle.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748482A_ABST
    Figure CN120748482A_ABST
Patent Text Reader

Abstract

The invention discloses a method for excavating glutinous rice wine brewing related genes. The method comprises the following steps: firstly, collecting multiple rice germplasm resources containing indica rice, japonica rice and glutinous rice, detecting SNP / CNV variation through whole genome re-sequencing, and obtaining granularity and roundness phenotype data; secondly, performing gene-phenotype association analysis by utilizing a mixed linear model MLM and combining GEMMA software, correcting a population structure and a genetic relationship, and screening out significant associated genes; finally, the candidate genes are knocked out through a CRISPR-Cas9 technology, and phenotype verification is carried out. The method disclosed by the invention brings remarkable technical progress in the aspects of capturing a complicated incidence relation between phenotypic characteristics and genomes, enriching genome data processing and MLM model training, improving the breeding efficiency of glutinous rice crops and the like, and averagely shortens a function research period from 6-8 months of traditional transgenosis to 3-4 months; and an efficient and accurate technical means is provided for improving the quality of the glutinous rice for wine brewing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of plant crop breeding, and in particular relates to a method for mining glutinous rice wine-making related genes. Background Art

[0002] Improvement of crop shape is based on genetic variation. By creating and utilizing this genetic variation, high-quality crop varieties can be cultivated. Crop breeding techniques primarily include hybridization, mutation breeding, transgenic breeding, and genome editing. Genome editing, which precisely introduces target genes into recipient plants to achieve desired traits, offers advantages such as ease of use and the ability to precisely alter gene sequences.

[0003] Because glutinous rice has a different starch structure from indica rice and japonica rice, it has a high content of amylopectin and is easier to saccharify and ferment. It is often used in rice wine brewing. For example, Chinese Shaoxing rice wine mostly uses short round japonica glutinous rice (such as "Osmanthus Glutinous Rice") because its grains are full and round, which is suitable for traditional rice soaking and cooling processes.

[0004] The quality of glutinous rice germplasm and its preparation process are crucial attributes influencing the production of yellow rice wine. Cultivating high-quality glutinous rice germplasm plays a crucial role in improving the quality of yellow rice wine and promoting agricultural structural transformation. Currently, glutinous rice breeding techniques primarily rely on traditional hybridization, which has the disadvantages of long breeding cycles, low efficiency, poor precision, high uncontrollability, high costs, and poor reproducibility.

[0005] Research has found that the particle size and roundness of glutinous rice directly affect the physical and chemical process of rice wine brewing: raw materials with uniform and round particles can optimize starch utilization, stabilize the fermentation environment, and ultimately achieve the mellow taste, rich aroma and clear body of rice wine.

[0006] By studying the relationship between the phenotypic characteristics of glutinous rice grain size and roundness and glutinous rice genes, we can make targeted changes to the genome and improve the traits of glutinous rice crops, thereby obtaining new varieties of high-yield, high-quality glutinous rice crops for winemaking. It is imperative to accelerate the creation and cultivation of glutinous rice with excellent taste quality, which can meet the demand of consumers and the market for high-quality rice wine.

[0007] Mixed linear models (MLMs) are gene-phenotype association models that integrate fixed effects (direct association signals) and random effects (background genetic / environmental noise) to analyze the association between genetic variation and phenotypic traits. Their core advantage lies in controlling for confounding factors and improving statistical power. However, in practical applications, they are susceptible to interference from genetic heterogeneity. The same phenotype may be driven by different genes or pathways (genetic heterogeneity), resulting in weak effects from single SNPs. GWAS require testing millions of SNPs, leading to multiple comparison problems (inflated false positive rates). Furthermore, the model's assumptions are limited. MLMs assume that genetic effects are linearly additive, but nonlinear effects (such as dominance and epistasis) may exist in real biology.

[0008] How to construct a multi-level model to integrate biological information at different levels and reveal the complex relationship between glutinous rice gene regulation and phenotypic characteristics has become a difficulty in exploring glutinous rice winemaking-related genes. Summary of the Invention

[0009] In the current gene mining of glutinous rice's winemaking function, the mixed linear model (MLM) urgently needs to solve the difficulties of high-dimensional data calculation, genetic heterogeneity and model assumption limitations. This application integrates multi-omics data (such as epigenome and proteome) and reveals the complex relationship between glutinous rice gene regulation and phenotypic characteristics by constructing a multi-level model to integrate biological information at different levels.

[0010] The association analysis between genomic variation (such as SNP (single nucleotide polymorphism), CNV) and transcriptome expression patterns is the core link in analyzing the genetic basis of winemaking traits. However, existing studies have mostly focused on single-omics data (such as analyzing only the genome or transcriptome), and the genomic variation is disconnected from the transcriptome expression. Only the genomic variation (such as SNP) of the raw rice variety is detected or only the transcriptome data (such as gene expression) is analyzed. There is no correlation between the impact of these variations on the expression of winemaking-related genes or the source of the genomic variation behind them, and there is a lack of systematic integration. For example, there is a SNP variation in the starch synthase gene (Wx) of glutinous rice varieties, but it has not been verified whether the SNP leads to downregulation of gene expression and thus affects starch content.

[0011] Comparative analyses across multiple varieties are insufficient. Only genomic and transcriptomic data from a single variety are analyzed, lacking comparative analysis across multiple varieties. The differential contributions of genomic variation across varieties to winemaking traits remain unexplained. For example, while varieties A and B have different CNVs in the sugar transporter gene (SUT), their gene expression differences and their impact on winemaking efficiency have not been compared.

[0012] In view of this, the main purpose of this application is to provide a method for exploring genes related to glutinous rice winemaking, improve the quality of glutinous rice raw materials, combine traditional rice wine craftsmanship and domestic and foreign market demands, promote technological progress of the rice wine industry and enhance market competitiveness.

[0013] The overall technical process of this method is shown in the appendix of the manual. Figure 1 The specific technical solution of this application is as follows: Provide a method for mining glutinous rice wine-making related genes,

[0014] The method comprises the following steps: (1) multi-omics data collection: collecting multiple rice germplasm resources, including indica rice, japonica rice and glutinous rice varieties, and performing whole genome resequencing to detect SNP (single nucleotide polymorphism) / CNV (copy number variation) variations and phenotypic group data measurement, wherein the phenotypic group data includes particle size and roundness; (2) data fusion and weight calculation: constructing a mixed linear model (MLM) for gene-phenotype association analysis, wherein the mixed linear model formula is: y=Xα+Zβ+Wμ+e, wherein y is the phenotypic trait vector, X is the indicator matrix of fixed effects (population structure Q matrix), α is the estimated parameter vector of fixed effects, Z is the SNP (single nucleotide polymorphism) / CNV (copy number variation) genotype matrix, β is the effect of SNP (single nucleotide polymorphism) / CNV (copy number variation), vector W is the indicator matrix of random effects (kinship K matrix), μ is the predicted random individual effect vector, and e is the random residual vector; and GEMMA software is used to calculate the contribution weight of each gene to the particle size and the roundness, and to screen the candidate genes with significant association; (3) Gene function verification: The candidate genes are knocked out or overexpressed using CRISPR-Cas9 technology, and the function of the candidate genes is verified by the phenotypic changes of the edited rice germplasm.

[0015] Preferably, the collected multiple rice germplasms include indica rice varieties, japonica rice varieties and glutinous rice varieties, and the multiple rice germplasms are more than 400.

[0016] Preferably, in step (1): the whole genome resequencing is performed using the Illumina NovaSeq platform with a sequencing depth of ≥30X, SNP / CNV detection is performed by aligning to the IRGSP-1.0 reference genome via BWA-MEM, and the filtering criteria are QUAL>30 and DP>5; the particle size measurement is performed by quantifying the volume of 150 hulled polished rice grains using the water displacement method, repeated 3 times, and ±2SD outliers are excluded; the roundness measurement is performed by obtaining rice grain images using a high-resolution scanner (1200 dpi), and the aspect ratio is calculated using ImageJ software.

[0017] Preferably, in the step (2): for the SNP (single nucleotide polymorphism),

[0018] The mixed linear model formula y=Xα+Zβ+Wμ+e is more specifically:

[0019] y=X geno α+X epi γ+X prot δ+Zβ+Wμ+e(SNP (single nucleotide polymorphism));

[0020] Among them, the genotype effect (X geno α)

[0021] X geno is a genome fixed effect design matrix (dimension: n×p), representing the genotype data of the rice germplasm, composed of SNP data detected by whole genome resequencing, with each row corresponding to one rice germplasm and each column corresponding to one genotype variable (SNP);

[0022] α is the genomic fixed effect coefficient vector (dimension: p × 1), which represents the direct effect of each SNP on the phenotype, and each element corresponds to the effect value of a genotype variable;

[0023] Epigenomic effect (X epi γ):

[0024] X epi is the epigenomic fixed effect design matrix (dimension: n×q), which represents the epigenetic data of the rice accessions, with each row corresponding to one rice accession and each column corresponding to one epigenomic variable;

[0025] γ is the epigenetic group fixed effect coefficient vector (dimension: q × 1), which represents the effect of epigenetic variation on phenotype, and each element corresponds to the effect value of an epigenetic group variable;

[0026] Proteomic effects (X prot δ):

[0027] X prot Design a matrix (dimension: n × r) for the proteomic fixed effects, representing the protein expression or structural data of rice accessions, i.e., protein expression levels or post-translational modification data, with each row corresponding to a rice accession and each column corresponding to a proteomic variable;

[0028] δ is the proteomic fixed effect coefficient vector (dimension: r × 1), which represents the effect of protein expression or structural variation on the phenotype, and each element corresponds to the effect value of a proteomic variable.

[0029] Preferably, in step (2): for the CNV (copy number variation), the mixed linear model formula y=Xα+Zβ+Wμ+e is more specifically: y=X cnv α+X epiγ+X prot δ+Zβ+Wμ+e(CNV(copy number variation));

[0030] Among them, the genotype effect (X cnv α):X cnv is the CNV fixed effect matrix (dimension: n×m), encoded as copy number (0: missing, 1: normal, 2: amplified), and the data comes from resequencing CNV detection; α is the CNV effect value vector, CNV effect value (m×1), quantifying the contribution of copy number change to the phenotype, and a positive effect indicates that the copy number increase promotes the phenotype;

[0031] Epigenomic effect (X epi γ):X epi is the epigenomic fixed effect design matrix (dimension: n×q), which represents the epigenetic data of the rice accessions, with each row corresponding to one rice accession and each column corresponding to one epigenomic variable;

[0032] γ is the epigenetic group fixed effect coefficient vector (dimension: q × 1), which represents the effect of epigenetic variation on phenotype, and each element corresponds to the effect value of an epigenetic group variable;

[0033] Proteomic effects (X prot δ):X prot Design a matrix (dimension: n × r) for the proteomic fixed effects, representing the protein expression or structural data of rice accessions, i.e., protein expression levels or post-translational modification data, with each row corresponding to a rice accession and each column corresponding to a proteomic variable;

[0034] δ is the proteomic fixed effect coefficient vector (dimension: r × 1), which represents the effect of protein expression or structural variation on the phenotype, and each element corresponds to the effect value of a proteomic variable.

[0035] Preferably, the mixed linear model formula y=Xα+Zβ+Wμ+e is more specifically:

[0036] Z is the SNP (single nucleotide polymorphism) / CNV (copy number variation) genotype matrix, a random effect design matrix (dimension: n × k), which represents the kinship between individuals or the population structure. Each row corresponds to a rice accession, and each column corresponds to a random effect variable.

[0037] β is the effect of SNP (single nucleotide polymorphism) / CNV (copy number variation), a random effect coefficient vector (dimension: k×1), which obeys β~N(0,σ 2 g I) represents the influence of other unobserved factors on the phenotype, and each element corresponds to the effect value of a random effect variable.

[0038] Preferably, the mixed linear model formula y=Xα+Zβ+Wμ+e is more specifically: e is the random residual vector, the vector matrix of the residual, the residual term (dimension: n×1), captures unexplained random errors, represents the influence of other unobserved factors on the phenotype, and obeys e~N(0,σe2I).

[0039] Preferably, the mixed linear model formula y=Xα+Zβ+Wμ+e,

[0040] Z is the genetic kinship matrix (Kinship Matrix), the formula is:

[0041]

[0042] Where M is the number of SNPs (single nucleotide polymorphisms) / CNVs (copy number variations), g ik For sample i at SNP k / CNV k Genotype, p k is the minor allele frequency.

[0043] Preferably, the present application further includes constructing a dedicated database for rice grain size and roundness, integrating genomic data, phenotypic data and gene editing results of more than 400 germplasms, and supporting SNP (single nucleotide polymorphism) / CNV (copy number variation) effect value query and breeding marker screening.

[0044] Preferably, the present application further includes application in molecular design breeding of glutinous rice, wherein the gene sets of the selected excellent quality traits are combined as the final glutinous rice excellent quality gene set.

[0045] In combination with the above-mentioned technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected in this application are: the method described in this application has brought significant technological progress in capturing the complex correlation between phenotypic characteristics and genomes, enriching genome data processing and MLM model training, and improving the breeding efficiency of glutinous rice crops, providing an efficient and accurate selection tool for the breeding of glutinous rice for brewing. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0047] Figure 1 This is a specific flow chart of the glutinous rice winemaking gene mining method provided by the present invention. DETAILED DESCRIPTION

[0048] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0049] In view of the problems existing in the prior art, the present invention provides a method for mining glutinous rice winemaking related genes, which adopts the following technical solutions:

[0050] As described in the background, the particle size and roundness of glutinous rice directly impact the physical and chemical processes of rice wine brewing: uniform, rounded grains optimize starch utilization and stabilize the fermentation environment, ultimately contributing to the rich taste, rich aroma, and clear body of rice wine. The specific impacts of these two phenotypic characteristics on rice wine brewing are as follows:

[0051] ①The influence of particle size on rice wine brewing

[0052] Water absorption and cooking efficiency: Large-grain glutinous rice absorbs water slowly, requiring longer soaking time to ensure that the rice grains fully absorb water; steam penetration is difficult during cooking, which may cause excessive gelatinization of the outer layer and incomplete gelatinization of the inner layer, resulting in uneven starch release; small-grain glutinous rice absorbs water quickly and evenly, and is easy to completely gelatinize during cooking, but overly broken grains can easily cause the rice grains to stick together, hindering oxygen exchange during fermentation.

[0053] Starch release and saccharification efficiency: Uniformly granulated glutinous rice forms a uniform gelatinized starch layer after steaming, which facilitates the efficient breakdown of amylopectin by enzymes in the koji (such as α-amylase and saccharifying enzymes) into fermentable sugars. However, excessive starch exposure can lead to rapid saccharification early in the fermentation process, making fermentation temperature difficult to control and creating the risk of rancidity.

[0054] Air permeability during the fermentation process: Glutinous rice with larger particles forms a loose structure after steaming, which is conducive to the oxygen exchange and metabolic activities of microorganisms (yeast and mold) during the fermentation process; rice mash with too small particles or excessive gelatinization has high viscosity and easily forms a dense structure, which inhibits microbial activity, causing fermentation stagnation or the production of odor.

[0055] ②The influence of roundness on rice wine brewing

[0056] Physical structure and processing adaptability: Round rice grains have a smooth surface, absorb water evenly, and are intact after steaming. The pores between the rice grains are moderate, which is conducive to the uniform distribution of the koji and a stable fermentation environment. Non-round rice grains (such as long grains or broken grains): They absorb water unevenly, are easy to break or stick together after steaming, affect the uniformity of the fermentation mash, and may lead to insufficient local saccharification.

[0057] Starch gelatinization and wine body and taste: Round glutinous rice has a high content of amylopectin (usually over 95%), which forms a viscous gel-like structure after gelatinization. It can wrap the yeast and slowly release sugar, promoting the formation of a mellow taste and complex flavor. Non-round rice grains may have a slightly higher proportion of amylose, resulting in a refreshing wine but lacking the softness of traditional rice wine.

[0058] Impurities and fermentation stability: Glutinous rice with high roundness usually has fewer impurities and a lower breakage rate, which reduces the risk of bacterial contamination during the fermentation process and ensures the pure flavor of the rice wine. Rice grains with rough or broken surfaces are prone to adsorbing impurities and may introduce wild fungi, causing sourness or odor.

[0059] In summary, uniform, rounded glutinous rice grains achieve high starch conversion, increasing wine yield and alcohol content (reaching 14-20% vol), while also reducing residual sugar. Ideal grain size and roundness promote the production of flavor compounds like esters and higher alcohols, imparting layered aromas like floral, fruity, and caramel to the rice wine. Conversely, excessive fusel alcohol content can lead to a harsh taste. Rounded rice grains allow for more complete protein precipitation after fermentation, resulting in clearer wine. Uneven grains can lead to turbidity or increased sediment.

[0060] This application aims to collect as many rice germplasms as possible, including traditional rice varieties and glutinous rice varieties. Traditional rice varieties include indica rice (IR64, Zhenshan 97), japonica rice (Nipponbare, Koshihikari), etc., covering different subpopulations and geographical origins; glutinous rice varieties include Yunnuo No. 3, Tainuo No. 1, etc., whose endosperm is almost entirely amylopectin (caused by mutation of the Waxy gene). Rice germplasm with a wide geographical coverage is selected as much as possible to ensure the statistical power of subsequent association analysis.

[0061] Phenomics data refers to the collection of results from systematic, large-scale measurements of all observable traits (i.e., phenotypes) exhibited by an organism under a specific environment. This application collects the above-mentioned two phenotypic characteristics of rice germplasm: grain size and roundness.

[0062] Particle size is measured using the water displacement method, which measures the volume of 150 grains of rice. Roundness is measured using the aspect ratio. The aspect ratio is the ratio of the length to the width of a rice grain. A value closer to 1 indicates a more rounded grain and better roundness; a larger value indicates a more elongated grain and poorer roundness. A high-resolution scanner (1200 dpi) can be used to capture a flat image of the rice grains. Image recognition tools can be used to automatically identify the grain outlines and calculate the aspect ratio: maximum length (L) / maximum width (W).

[0063] Whole-genome resequencing involves sequencing the genomes of different individuals from a species with a known genome sequence, and then analyzing the differences between individuals or populations. This application performed whole-genome resequencing and variant detection on multiple rice germplasms, detecting SNPs (single nucleotide polymorphisms) and CNVs (copy number variations).

[0064] SNPs (single nucleotide polymorphisms) refer to DNA sequence polymorphisms caused by variations in a single nucleotide at the genomic level. Individual rice plants may differ in a single base at certain sites in the genome. These differences are known as SNPs. For example, the A→G mutation at position Chr6:10,256,789 can lead to changes in the expression of the Waxy gene. Detecting SNPs helps understand genetic variation.

[0065] CNVs (copy number variations) are changes in the number of copies of specific DNA segments in the genome, including deletions and duplications. These variations can involve small DNA segments or hundreds of genes. CNVs are an important component of genomic diversity and can influence gene expression, phenotype, and disease susceptibility.

[0066] This application performs data fusion and weight calculation on the measurement results of rice germplasm obtained, and adopts a mixed linear model (MLM) as a statistical model to analyze the association between traits and genes with complex genetic structure and environmental factors. Due to the complex population structure and intra-population kinship of rice populations, the mixed linear model can correct the influence of these factors on the trait-gene association analysis. By constructing a "gene-phenotype association model", that is, a mixed linear model, the contribution weight of each gene to the particle size and roundness of rice germplasm is calculated, thereby gaining a deeper understanding of the relationship between genes and the two target phenotypes.

[0067] The mixed linear model formula is y = Xα + Zβ + Wμ + e, where

[0068]

[0069]

[0070] GEMMA (Genome-wide Efficient Mixed Model Association) is a software specifically designed for genome-wide association analysis. It efficiently processes large-scale genomic and phenotypic data, taking into account factors such as population structure and kinship. GEMMA combines the phenotypic diversity of the target trait with the polymorphism of the gene. Based on the mixed linear model formula, GEMMA performs an association test for each SNP / CNV marker, calculating the strength and significance of the association with the target phenotype.

[0071] This application uses CRISPR-Cas9 technology to knock out / overexpress candidate genes, and uses the phenotypic changes of the edited species to verify the function of the candidate genes.

[0072] The CRISPR-Cas9 system consists of two main parts: the Cas9 nuclease, an enzyme that cuts double-stranded DNA; and the guide RNA (gRNA), which can specifically recognize and bind to a specific DNA sequence of the target gene. When the gRNA binds to the target DNA sequence, it guides the Cas9 nuclease to that site for cutting, creating a gap in the double-stranded DNA. The cell's own DNA repair mechanism will repair this gap, but errors may occur during the repair process, such as the insertion or deletion of some bases, which can lead to loss of gene function (knockout) or changes, such as overexpression. Although overexpression is not achieved directly through CRISPR-Cas9 cutting, it can be achieved through similar gene editing strategies, such as the introduction of strong promoters, to achieve the construction of gene overexpression effects. CRISPR-Cas9 is mainly used to accurately locate and manipulate target gene sites.

[0073] The process of knockout using CRISPR-Cas9 technology is as follows:

[0074] First, gRNA is designed: for each candidate gene, a gRNA is designed that can specifically recognize a specific sequence in the gene. The sequence is usually located in the coding region or regulatory region of the gene to ensure that the gene function can be effectively affected after cutting.

[0075] Next, construct the vector: the Cas9 gene and gRNA are constructed on the same vector, or separately (if it is a multi-gene editing system). The vector can be a plasmid, a viral vector, etc., which is used to introduce Cas9 and gRNA into rice cells.

[0076] The constructed vector is then introduced into rice cells through methods such as Agrobacterium-mediated transformation and gene gun therapy. Inside the rice cells, Cas9 and gRNA are expressed and function, cleaving the target gene.

[0077] Finally, positive plants are screened: After a period of cultivation, positive plants that have successfully introduced Cas9 and gRNA are screened through screening markers (such as antibiotic resistance genes, fluorescent protein genes, etc.).

[0078] Details of overexpression using CRISPR-Cas9 technology are as follows:

[0079] First, a strategy is selected: To achieve overexpression of candidate genes, the regulatory region (such as the promoter region) of the candidate gene can be edited using CRISPR-Cas9 technology to introduce a strong promoter sequence and enhance the transcriptional activity of the candidate gene; or an overexpression vector containing the candidate gene can be constructed and introduced into a specific location of the rice cell (such as a specific chromosome region) in combination with the CRISPR-Cas9 positioning strategy to achieve more precise regulation. Secondly, vector construction and transformation: According to the selected strategy, the corresponding vector is constructed, and the candidate gene or regulatory element is integrated with the CRISPR-Cas9 related element (if positioning is required) and then introduced into the rice cell. Positive plants are also screened out by screening markers.

[0080] Taking rice seeds as an example, the specific overexpression process is as follows:

[0081] First, target design and vector construction: Design sgRNA, avoid off-target sites (such as whole-genome alignment to ensure ≤3 mismatches) on the CRISPR-P or CHOPCHOP online platform, and design sgRNA for the target gene exon (efficiency predicted >80%); vector selection: Overexpression: pCAMBIA1300-Ubi (rice ubiquitin promoter drives the candidate gene).

[0082] Next, genetic transformation was performed using Agrobacterium-mediated transformation. Callus preparation was performed by inducing callus from mature seeds (2,4-D medium). Callus were co-cultured with Agrobacterium EHA105 carrying the vector (OD600 = 0.6, 20 minutes). Transgenic calli were screened for resistance to hygromycin (50 mg / L) for four weeks to obtain transgenic calli. The T0 generation conversion rate was approximately 30% for indica rice and up to 50% for japonica rice.

[0083] Then, mutant identification is performed: PCR / sequencing verification. Off-target detection: whole genome sequencing or PCR of predicted off-target sites.

[0084] CRISPR-Cas9 technology is highly precise and can target specific gene loci, compared to the randomness of traditional mutagenesis. It is also highly efficient, enabling mutants to be generated in the T0 generation. Through precise editing with CRISPR-Cas9, not only can the function of candidate genes be verified, but it also provides direct targets for subsequent molecular design and breeding. The method presented in this application shortens the functional research cycle from 6-8 months for traditional transgenics to 3-4 months on average.

[0085] Finally, this application needs to verify the function of the candidate gene edited by CRISPR-Cas9 technology by using phenotypic changes in the species. The specific steps of verification are as follows:

[0086] First, a control group was set up: while the gene editing operation was being performed, wild-type rice plants that had not undergone gene editing were set up as a control group.

[0087] Secondly, planting and observation: Gene-edited plants and control plants were planted under the same environmental conditions, ensuring that they were exposed to the same environmental factors such as light, temperature, water, and soil. Phenotypic changes were regularly observed and recorded at different stages of plant growth.

[0088] Then, data analysis: statistical analysis of the observed phenotypic data is performed to compare the differences between the gene-edited plants and the control plants. If the phenotype of the gene-edited plant has changed significantly, and this change is related to the knockout or overexpression of the gene, then the functional relationship between the gene and the rice grain size or other related phenotypes can be preliminarily verified. For example, if the rice grain size becomes significantly smaller after knocking out / overexpressing a gene, and the rice grain size becomes significantly larger after overexpressing the gene (achieving a similar effect through the above strategy), then it can be considered that the gene has a positive regulatory effect on the rice grain size.

[0089] Example 1: SNP (Single Nucleotide Polymorphism)

[0090] Collect multiple rice germplasms, measure phenotypic data, perform whole-genome resequencing, and detect variations in the rice germplasm.

[0091] First, 400 rice germplasms were collected.

[0092] Secondly, phenotypic data were measured for each rice germplasm, collecting two phenotypic characteristics of rice germplasm: grain size and roundness.

[0093] Then, whole genome resequencing and variation detection were performed on each rice germplasm. In this embodiment, whole genome resequencing and variation detection were performed on 400 rice germplasms to detect SNPs (single nucleotide polymorphisms).

[0094] The obtained rice germplasm measurement results were subjected to data fusion and weight calculation to construct a mixed linear model - the "gene-phenotype association model". The mixed linear model formula is y = Xα + Zβ + Wμ + e. In this embodiment, multi-omics variables are introduced into the MLM framework (y = Xα++Zβ+Wμ+e):

[0095] y=X geno α+X epi γ+X prot δ+Zβ+Wμ+e

[0096] Relying solely on genomic data (SNPs) may miss direct effects of epigenetic regulation or protein function. This example comprehensively analyzes the regulatory chain of "SNP → epigenetic modification → protein → phenotype" by incorporating genomic, epigenomic, and proteomic data. By integrating genotype, epigenomic, and proteomic data, it can more comprehensively capture factors that influence phenotypic variation and improve the model's ability to explain phenotypic variation.

[0097] in:

[0098] y: Phenotypic trait vector (y) (dimension: n × 1), representing the target trait, i.e., grain size and roundness measurements of the rice accessions. Each element corresponds to a phenotypic measurement of a rice accession.

[0099] Genotype effect (X geno α)

[0100] X geno : The genome fixed-effect design matrix (dimension: n × p), representing the genotype data of rice accessions, is composed of SNP data detected by whole-genome resequencing (e.g., genotype coding is 0 / 1 / 2), where each row corresponds to a rice accession and each column corresponds to a genotype variable (e.g., SNP).

[0101] α: Genomic fixed effect coefficient vector (dimension: p×1), representing the direct effect of each SNP on the phenotype, with each element corresponding to the effect value of a genotype variable.

[0102] Epigenomic effect (X epi γ)

[0103] X epi : Epigenomic fixed-effect design matrix (dimension: n × q), representing the epigenetic data of rice accessions (such as DNA methylation levels, chromatin accessibility data, histone modifications, etc.), with each row corresponding to a rice accession and each column corresponding to an epigenomic variable.

[0104] γ: epigenetic group fixed effect coefficient vector (dimension: q × 1), representing the impact of epigenetic variation on phenotype, each element corresponds to the effect value of an epigenetic group variable.

[0105] Proteomic effects (X prot δ)

[0106] X prot : Proteomic fixed-effect design matrix (dimension: n × r), representing the protein expression or structural data of rice germplasm, i.e., protein expression amount or post-translational modification data, each row corresponds to a rice germplasm, and each column corresponds to a proteomic variable.

[0107] δ: Proteomic fixed effect coefficient vector (dimension: r × 1), representing the effect of protein expression or structural variation on the phenotype, with each element corresponding to the effect value of a proteomic variable.

[0108] Random effects (Zβ)

[0109] Z is a random-effects design matrix (dimension n × k), representing the kinship relationship between individuals or the population structure. It is usually a kinship matrix or an environmental heterogeneity matrix, with each row corresponding to a rice accession and each column corresponding to a random-effect variable.

[0110] β: random effect coefficient vector (dimension: k×1), subject to β~N(0,σ 2 g I) represents the influence of other unobserved factors on the phenotype, which is used to control genetic background or environmental noise. Each element corresponds to the effect value of a random effect variable.

[0111] Random error term (e)

[0112] e: residual vector matrix, residual term (dimension: n×1), captures unexplained random errors, representing the influence of other unobserved factors on the phenotype. e 2 I).

[0113] The data sources and analysis of the model used in this example are as follows:

[0114] Genome fixed effect (X geno α)

[0115] Data source: Whole-genome SNP data of 400 rice accessions, detected by resequencing, and high-quality SNPs (MAF>0.05, missing rate <5%) screened.

[0116] Coding method: SNP genotypes are usually coded as 0 (homozygous reference), 1 (heterozygous), and 2 (homozygous variant).

[0117] Quantify the direct genetic effect of each SNP on particle size and roundness. For example, a SNP (α j ) was significantly positive, indicating that the allele at this locus may promote grain enlargement.

[0118] The apparent group fixed effect (X epi γ)

[0119] Data sources: Epigenomic data may include:

[0120] DNA methylation (e.g., whole-genome bisulfite sequencing to detect CpG site methylation levels); chromatin accessibility (e.g., ATAC-seq or DNase-seq data);

[0121] Histone modifications (such as H3K4me3, H3K27ac, and other marks measured by ChIP-seq).

[0122] Preprocessing: standardization (such as β value or M value conversion), batch effect correction (such as ComBat).

[0123] Epigenomic fixed effects reveal how epigenetic modifications regulate phenotypes. For example, a significant negative methylation site may suppress the expression of genes related to grain development.

[0124] Proteomic fixed effect (X prot δ):

[0125] Data source: Proteomic data were quantified using mass spectrometry techniques (e.g., LC-MS / MS) to detect the expression levels or phosphorylation status of proteins in rice grains.

[0126] Preprocessing: normalization (such as TMT calibration quantification), missing value filling (such as k-nearest neighbor method).

[0127] Proteomic fixed effects directly correlate protein function with phenotype, starch synthase (δ m ) High expression may positively affect granularity.

[0128] Random effects (Zβ)

[0129] Kinship Matrix: Calculates the genetic similarity between samples based on SNPs. The formula is:

[0130]

[0131] Where M is the number of SNPs, g ik For sample i at SNP k Genotype, p k is the minor allele frequency (MAF). Reduce false positive associations by controlling for population structure, recessive kinship, or environmental noise.

[0132] Residual term (e): Assuming it follows an independent normal distribution, the variance σ e 2 Reflects the variation not explained by the model (e.g., measurement error, unmodeled biological factors).

[0133] The data results of Example 1 are as follows:

[0134] Simulated data generation:

[0135] Genome: 400 samples*10,000 SNPs (50 causal SNPs, effect size α~N(0,0.1)).

[0136] Epigenetic group: 50 methylation sites are regulated by SNPs (γ~N(0.2α, 0.05)).

[0137] Proteome: 20 proteins whose expression levels were driven by methylation (δ~N(0.3γ, 0.1)).

[0138] Phenotype: y = X geno α+X epi γ+X prot δ+Zβ+Wμ+e,

[0139] where σ g 2 =0.4,σ e 2 =0.6.

[0140] Model output: significant variables:

[0141] SNP_Chr5: 12345 (α=0.12, p=1e-6), located in the promoter region of the grain weight gene GW2.

[0142] Methyl_Site_Chr8:67890 (γ=-0.08, p=0.003), regulates starch synthase gene expression.

[0143] Protein_SSIII (δ=0.15, p=0.001) directly promoted grain filling.

[0144] Heritability decomposition: genomic contribution 25%, epigenomic 10%, proteomic 15%, residual 50%.

[0145] The analysis identified marker loci that are closely associated with variations in two target phenotypes: grain size and roundness. These marker loci are likely located near genes associated with these two target phenotypes, providing important clues for further investigation into gene function and regulatory mechanisms.

[0146] The significant sites in this example are:

[0147] Particle size: SNP_3_128736 (p = 3.2 × 10 -8 ) is located in an unknown gene region on rice chromosome 3 and may explain 24.2% of phenotypic variation.

[0148] Roundness: SNP_5_3200000 (p = 1.8 × 10 -6) is located in an unknown gene region on rice chromosome 5 and may explain 17.8% of phenotypic variation.

[0149] In this step, multiple candidate genes related to the target phenotype are analyzed and targeted. In this embodiment, the fusion analysis targets multiple candidate genes, two of which are:

[0150] Gene ID Mutation Type Functional prediction Phenotypic effect value (β) GS5 Promoter SNP Cell size regulation +0.41mL / allele FL02 Intronic SNPs Regulation of starch synthesis +0.25mL / allele

[0151] Functional verification of the candidate genes was performed. CRISPR-Cas9 technology was used to knock out the two candidate genes, and the functions of the candidate genes were verified through phenotypic changes.

[0152] In this example, two candidate genes were knocked out, and after the candidate genes were knocked out, the gene functions were verified through phenotypic changes.

[0153] Example 2: CNV (copy number variation)

[0154] Collect multiple rice germplasms, measure phenotypic data, perform whole-genome resequencing, and detect variations in the rice germplasm.

[0155] First, 400 rice germplasms were collected.

[0156] Secondly, phenotypic data were measured for each rice germplasm, collecting two phenotypic characteristics of rice germplasm: grain size and roundness.

[0157] Then, whole-genome resequencing and variation detection were performed on each rice germplasm. In this example, whole-genome resequencing and variation detection were performed on 400 rice germplasms to detect CNV (copy number variation).

[0158] The obtained rice germplasm measurement results were subjected to data fusion and weight calculation to construct a mixed linear model - the "gene-phenotype association model". The mixed linear model formula is y = Xα + Zβ + Wμ + e. In this embodiment, multi-omics variables are introduced into the MLM framework (y = Xα + Zβ + Wμ + e):

[0159] y=X cnv α+X epi γ+X prot δ+Zβ+Wμ+e

[0160] Among them, X cnv : CNV fixed effect matrix (dimension: n×m), encoded as copy number (0: deletion, 1: normal, 2: amplification).

[0161] A: CNV effect size vector, quantifying the contribution of copy number variation to the phenotype.

[0162] Data preprocessing: Epigenetic analysis: Whole-genome methylation sequencing (WGBS) was used to detect 5mC levels and filter low-coverage (<10×) sites.

[0163] Proteomics: Quantification of key proteins in grain development (such as starch synthases and cell wall modification enzymes) based on SWATH-MS.

[0164]

[0165] The data results of Example 2 are as follows:

[0166] Simulation parameters: CNV data: 400 samples × 500 CNV regions, including 20 causal CNVs (effect size α ~ N(0,0.15)).

[0167] Epigenetic group: 30 methylation sites are regulated by CNV (γ~N(0.25α,0.1)).

[0168] Proteome: The expression of 15 proteins was driven by methylation (δ~N(0.35γ,0.08)).

[0169] Phenotype: y = X cnv α+X epi γ+X prot δ+Zβ+Wμ+e

[0170] Among them, σ g 2 =0.35,σ e 2 =0.65.

[0171] Significantly associated sites:

[0172]

[0173] Heritability decomposition: CNV contribution: 30%; epigenomic contribution: 12%; proteomic contribution: 18%; residual: 40%.

[0174] In this step, multiple candidate genes related to the target phenotype are analyzed and targeted. In this embodiment, fusion analysis targets multiple candidate genes:

[0175]

[0176] As in Example 1, the functions of the candidate genes were verified.

[0177] CRISPR-Cas9 technology was used to knock out the above-mentioned candidate genes, and the functions of the candidate genes were verified through phenotypic changes.

[0178] In the examples, gene functions of multiple candidate genes were verified using CRISPR-Cas9.

[0179] Phenotypic analysis: Grain size: The grain volume of the edited plants was measured (Micro-CT scanning).

[0180] Roundness: Calculate the length-to-width ratio (LWR) to verify the dose effect of CNV on morphology.

[0181] Molecular mechanism:

[0182] qRT-PCR: Detect the effect of CNV on the expression of adjacent genes (such as GW5).

[0183] ChIP-seq: Analyze changes in chromatin accessibility of OsSPL16-binding target genes (such as EXPANSIN).

[0184] The specific embodiments described in this application are merely illustrative of the spirit of this application. Those skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments without departing from the spirit of this application or exceeding the scope defined by the appended claims.

Claims

1. A method for mining glutinous rice winemaking related genes, characterized in that: The method comprises the following steps: (1) Multi-omics data collection: Multiple rice germplasm resources, including indica, japonica, and glutinous rice varieties, were collected and whole-genome resequencing was performed to detect SNP / CNV variations and measure phenotypic data, including grain size and roundness; (2) Data fusion and weight calculation: A mixed linear model (MLM) was constructed to perform gene-phenotype association analysis. The formula of the mixed linear model is: y=Xα+Zβ+Wμ+e Wherein, y is the phenotypic trait vector, X is the indicator matrix of fixed effects (population structure Q matrix), α is the estimated parameter vector of fixed effects, Z is the SNP / CNV genotype matrix, β is the effect of SNP / CNV, vector W is the indicator matrix of random effects (kinship K matrix), μ is the predicted random individual effect vector, and e is the random residual vector; GEMMA software was used to calculate the contribution weight of each gene to the particle size and the roundness, and to screen the candidate genes with significant association; (3) Gene function verification: The candidate gene is knocked out or overexpressed using CRISPR-Cas9 technology, and the function of the candidate gene is verified by the phenotypic changes of the edited rice germplasm.

2. The method for mining glutinous rice winemaking-related genes according to claim 1, characterized in that: The collected multiple rice germplasms include indica rice varieties, japonica rice varieties and glutinous rice varieties, and the multiple rice germplasms are more than 400.

3. The method for mining glutinous rice winemaking-related genes according to claim 1, characterized in that: In the step (1): The whole genome resequencing was performed using the Illumina NovaSeq platform with a sequencing depth of ≥30X, and SNP / CNV detection was performed by alignment to the IRGSP-1.0 reference genome using BWA-MEM, with filtering criteria of QUAL>30 and DP>5; The particle size measurement was performed by quantifying the volume of 150 hulled polished rice grains using the water displacement method, repeated three times and excluding ±2SD outliers. The roundness was measured by acquiring rice grain images using a high-resolution scanner (1200 dpi) and calculating the aspect ratio using ImageJ software.

4. The method for mining glutinous rice winemaking-related genes according to claim 1, characterized in that: In the step (2): for the SNP, The mixed linear model formula y=Xα+Zβ+Wμ+e is more specifically: y=X geno a+X epi c+X prot δ+Zβ+Wμ+e(SNP); Among them, the genotype effect (X geno α) X geno is a genome fixed effect design matrix (dimension: n×p), representing the genotype data of the rice germplasm, composed of SNP data detected by whole genome resequencing, with each row corresponding to one rice germplasm and each column corresponding to one genotype variable (SNP); α is the genomic fixed effect coefficient vector (dimension: p × 1), which represents the direct effect of each SNP on the phenotype, and each element corresponds to the effect value of a genotype variable; Epigenomic effect (X epi γ): X epi is the epigenomic fixed effect design matrix (dimension: n×q), which represents the epigenetic data of the rice accessions, with each row corresponding to one rice accession and each column corresponding to one epigenomic variable; γ is the epigenetic group fixed effect coefficient vector (dimension: q × 1), which represents the effect of epigenetic variation on phenotype, and each element corresponds to the effect value of an epigenetic group variable; Proteomic effects (X prot δ): X prot Design a matrix (dimension: n × r) for the proteomic fixed effects, representing the protein expression or structural data of rice accessions, i.e., protein expression levels or post-translational modification data, with each row corresponding to a rice accession and each column corresponding to a proteomic variable; δ is the proteomic fixed effect coefficient vector (dimension: r × 1), which represents the effect of protein expression or structural variation on the phenotype, and each element corresponds to the effect value of a proteomic variable.

5. The method for mining glutinous rice winemaking-related genes according to claim 1, characterized in that: In the step (2): for the CNV, The mixed linear model formula y=Xα+Zβ+Wμ+e is more specifically: y=X cnv a+X epi c+X prot δ+Zβ+Wμ+e(CNV); Among them, the genotype effect (X cnv α) X cnv is the CNV fixed effect matrix (dimension: n×m), encoded as copy number (0: missing, 1: normal, 2: amplified), and the data comes from resequencing CNV detection; α is the CNV effect value vector, CNV effect value (m×1), quantifying the contribution of copy number change to the phenotype, and a positive effect indicates that the copy number increase promotes the phenotype; Epigenomic effect (X epi γ): X epi is the epigenomic fixed effect design matrix (dimension: n×q), which represents the epigenetic data of the rice accessions, with each row corresponding to one rice accession and each column corresponding to one epigenomic variable; γ is the epigenetic group fixed effect coefficient vector (dimension: q × 1), which represents the effect of epigenetic variation on phenotype, and each element corresponds to the effect value of an epigenetic group variable; Proteomic effects (X prot δ): X prot Design a matrix (dimension: n × r) for the proteomic fixed effects, representing the protein expression or structural data of rice accessions, i.e., protein expression levels or post-translational modification data, with each row corresponding to a rice accession and each column corresponding to a proteomic variable; δ is the proteomic fixed effect coefficient vector (dimension: r × 1), which represents the effect of protein expression or structural variation on the phenotype, and each element corresponds to the effect value of a proteomic variable.

6. The method for mining glutinous rice winemaking-related genes according to claim 1, characterized in that: The mixed linear model formula y=Xα+Zβ+Wμ+e is more specifically: Z is the SNP / CNV genotype matrix, a random effect design matrix (dimension: n × k), which represents the relatedness between individuals or the population structure. Each row corresponds to a rice accession, and each column corresponds to a random effect variable. β is the effect of SNP / CNV, a random effect coefficient vector (dimension: k×1), which follows β~N(0,σ 2 g I) represents the influence of other unobserved factors on the phenotype, and each element corresponds to the effect value of a random effect variable.

7. The method for mining glutinous rice winemaking-related genes according to claim 1, characterized in that: The mixed linear model formula y=Xα+Zβ+Wμ+e is more specifically: e is the random residual vector, the vector matrix of residuals, the residual term (dimension: n×1), which captures unexplained random errors and represents the influence of other unobserved factors on the phenotype, and obeys e~N(0,σe2I).

8. The method for mining glutinous rice winemaking-related genes according to claim 6, characterized in that: The mixed linear model formula y=Xα+Zβ+Wμ+e, Z is the genetic kinship matrix (Kinship Matrix), the formula is: Where M is the number of SNPs / CNVs, g ik For sample i at SNP k / CNV k Genotype, p k is the minor allele frequency.

9. The method for mining glutinous rice winemaking-related genes according to claim 1, characterized in that: It further includes the construction of a dedicated database for rice grain size and roundness, integrating genomic data, phenotypic data and gene editing results of more than 400 germplasms, and supporting SNP / CNV effect value queries and breeding marker screening.

10. An application of the method according to any one of claims 1 to 8 in molecular design breeding of glutinous rice, characterized in that: The gene sets of the selected excellent quality traits were combined to form the final glutinous rice excellent quality gene set.