Method of increasing kernel weight in maize

By cloning the maize kernel shape regulating gene ZmKW7 and its homologous genes using fine mapping and gene editing technology, the problem of poor genetic stability of maize kernel weight was solved, and the effect of increasing maize kernel width and weight was achieved, thereby increasing yield.

CN116334125BActive Publication Date: 2026-05-29HUAZHONG AGRI UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAZHONG AGRI UNIV
Filing Date
2023-03-06
Publication Date
2026-05-29

Smart Images

  • Figure CN116334125B_ABST
    Figure CN116334125B_ABST
Patent Text Reader

Abstract

The present application relates to a method for increasing grain weight, and belongs to the field of molecular genetics. Two maize kernel type regulating genes are disclosed, and a method for increasing maize grain weight and yield by gene editing is disclosed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for increasing the weight of corn kernels, belonging to the field of molecular genetics. Background Technology

[0002] The genetic basis of kernel weight has attracted much attention. Although many scientists at home and abroad have conducted extensive QTL mapping studies on maize kernel weight, most of the preliminarily mapped kernel weight QTLs have poor genetic stability, are greatly affected by environmental factors, and have a low contribution rate to kernel weight, making them difficult to utilize. Statistics show that there are currently 1735 published QTLs in maize, of which less than 10% are related to kernel weight (maizeGDB, http: / / www.maizegdb.org). Kernel shape, including kernel length, width, and thickness, is an important factor determining kernel weight. Locating the genetic loci and genes controlling kernel shape is of great significance for studying the regulatory mechanisms of kernel shape, improving kernel shape traits, and increasing yield. Although some genes affecting maize kernel shape (such as...) have been identified... ZmGS3, ZmGW2 (e.g.) have been identified, but due to the complexity of the regulatory mechanism of grain shape traits, more genetic loci or genes that can affect grain shape and grain weight need to be identified and applied to maize genetic improvement.

[0003] To address the aforementioned problems, this invention provides a gene that, through fine mapping and cloning, influences maize kernel shape and weight. ZmKW7 They also identified two homologous genes. Surprisingly, these two homologous genes... ZmKW7 Their mechanisms of action are completely opposite. By editing and mutating these two genes, it is possible to increase the kernel width and / or kernel weight traits in maize, thereby increasing yield. Summary of the Invention

[0004] One of the objectives of this invention is to provide a gene that influences the kernel shape trait of maize.

[0005] The second objective of this invention is to provide a method for improving the shape of corn kernels and increasing their weight.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] This invention provides an application of a gene in regulating the kernel shape trait of maize, characterized in that: the gene is the gene with the number Zm00001d047817 or Zm00001d048574 in the MaizeGDB database (https: / / www.maizegdb.org).

[0008] The present invention also provides a method for increasing the width and / or weight of corn kernels, characterized by: inhibiting the expression and / or activity of the above-mentioned proteins in corn, and selecting plants with increased corn kernel width and / or weight.

[0009] In some implementations, the sequence of the above-mentioned protein is as follows:

[0010] MAAEINGGFLAAGGPRQHRGGLGCGRCFQNISLLHGLGIKFVLVPGTHIQIDKLLSEIGNKAKYVGQYRITDEDARKAAMDAAGRIRLTIEAKLSPGPPMLNLRRHGVIGRWHGLVDSIASGNFLGAKRRGVVNGIDYGFTEEVTKIDVSRIRERLDSDSIVVISNMGFSSSGDV or

[0011] As shown.

[0012] In some implementations, the methods for inhibiting protein expression and / or activity include any one of gene editing, RNA interference, or T-DNA insertion.

[0013] In some implementations, the gene editing described above uses the CRISPR / Cas9 method.

[0014] In some implementations, the DNA sequence of the genomic target region in maize using the CRISPR / Cas9 method described above is shown as GTATTTATTCTGGCGAGCAA or GCGAGTTGTCGTGGTCGGCG.

[0015] This invention also provides a kit for increasing corn kernel width and / or kernel weight, characterized in that it comprises any one of the following:

[0016] (1) An RNA molecule capable of the above target sequence; the RNA molecule may be an sgRNA molecule containing gRNA, crRNAs and tracrRNA structures, or a complex formed by gRNA, crRNAs and tracrRNA alone, or a complex including gRNA and crRNAs; in some embodiments, the sequence of the above RNA molecule is as shown in GUAUUUAUUCUGGCGAGCAAguuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuuuu or GCGAGUUGUCGUGGUCGGCGguuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuuuu.

[0017] (2) The DNA molecule encoding the RNA described in (1);

[0018] (3) A vector for expressing the RNA described in (1).

[0019] The present invention also provides a mutant gene, characterized in that: the mutant gene sequence is the above target sequence of the above gene replaced with GTATTTATTCTGGTGAGCAA or GCGAGTTGTCGTGGTCGAGCG respectively.

[0020] The present invention also provides the application of the above-mentioned mutant gene and its encoded protein in increasing the width and / or weight of maize kernels.

[0021] In some implementations, the above-described encoded protein amino acid sequence is as follows:

[0022] MAAEINGGFLAAGGPRQHRGGLGCGRCFQNISLLHGLGIKFVLVPGTHIQIDKLLSEIGNKAKYVGQYRITDEDARKAAMDAAGRIRLTIEAKLSPGPPMLNLRRHGVIGRWHGLVDSIASGNFLGAKRRGVVNGIDYGFTEEVTKIDVSRIRERLDSDSIVVISNMGFSSSGDV or

[0023] MAASLAAAASRLAATGGELSWSSAALPRESRGELPQRRAAPAPRFGRAVRRWGFARGRGDGGGSGGRGPLRWLVPGGVALHPGPPGQHLRRRHLRGGHRRPAPRRHSAGYIPSSWSRDQICSCSWNTCPD.

[0024] Compared with existing technologies, the beneficial effects of this invention are: the functions of the two maize kernel shape regulatory genes obtained by this invention are previously unreported, and they are related to another cloned kernel shape gene. ZmKW7 Their mechanisms of action are completely opposite. This invention also provides a method for increasing the kernel width and / or kernel weight traits of maize by gene editing and mutating these two genes. Attached Figure Description

[0025] Figure 1 Mapping results of QTLs for maize kernel type. Top: Genome-wide association signals on maize chromosome 7. Black dashed lines represent the significance threshold of genome-wide association analysis; HKW: 100-kernel weight; KL: kernel length. Bottom: Genotypic identification results of recombinant plants in the population by SNP markers, and the intervals and gene locations for fine mapping based on these markers. P12~SNP21: SNP marker locations; rec-1~rec-5: recombinant plants; SK: genotype of SK material; ZHENG58: genotype of Zheng 58 material; Heterozygous: heterozygous genotype. Detailed Implementation

[0026] The following definitions and methods are provided to better define this application and to guide those skilled in the art in its practice. Unless otherwise stated, the terms are to be understood in accordance with their conventional usage by those skilled in the art. All patent literature, academic papers, industry standards, and other publicly available publications cited herein are incorporated herein by reference in their entirety.

[0027] As used herein, “maize” means any maize plant and includes all plant varieties that can be bred with maize, including the whole plant, plant cells, plant organs, plant protoplasts, plant cell tissue cultures from which the plant can regenerate, plant callus, and complete plant cells in a plant or plant part, such as embryo, pollen, ovule, seed, leaf, flower, branch, fruit, stem, root, root tip, anther, etc. Unless otherwise indicated, nucleic acids are written from left to right in a 5' to 3' direction; amino acid sequences are written from left to right in the amino to carboxyl direction. Amino acids may be represented herein by their commonly known three-letter symbols or by the single-letter symbols recommended by the IUPAC-IUB Committee on Biochemistry Nomenclature. Similarly, nucleotides may be represented by commonly accepted single-letter codes. Numerical ranges include numbers that define the range. As used herein, “nucleic acid” includes deoxyribonucleotides or ribonucleotide polymers in single-stranded or double-stranded form, and, unless otherwise limited, includes known analogs (e.g., peptide nucleic acids) that have the basic properties of natural nucleotides and hybridize with single-stranded nucleic acids in a manner similar to that of naturally occurring nucleotides. As used herein, the term “encoding” or “encoded” in the context of a particular nucleic acid means that the nucleic acid contains the essential information to guide the translation of that nucleotide sequence into a particular protein. Codons are used to represent the information encoding the protein. As used herein, “full-length sequence” referring to a particular polynucleotide or the protein it encodes means the entire nucleic acid sequence or the entire amino acid sequence having a natural (non-synthetic) endogenous sequence. Full-length polynucleotides encode the full-length, catalytically active form of that particular protein. The terms “polypeptide,” “polypeptide,” and “protein” are used interchangeably herein to refer to polymers of amino acid residues. This term is used for amino acid polymers in which one or more amino acid residues are artificial chemical analogs of the corresponding naturally occurring amino acids. This term is also used for naturally occurring amino acid polymers. The terms “residue” or “amino acid residue” or “amino acid” are used interchangeably in this document to refer to an amino acid incorporated into a protein, polypeptide, or peptide (collectively, “protein”). Amino acids can be naturally occurring amino acids, and unless otherwise limited, may include known analogs of natural amino acids that can function in a similar manner to naturally occurring amino acids.

[0028] The term "trait" refers to the physiological, morphological, biochemical, or physical characteristics of a plant or a particular plant material or cell. In some cases, this trait is visible to the human eye, such as seed or plant size, or can be measured by biochemical techniques, such as detecting the protein, starch, or oil content of seeds or leaves, or by observing metabolic or physiological processes, such as by measuring tolerance to water deprivation or specific salt, sugar, or nitrogen concentrations, or by observing the expression levels of one or more genes, or by agronomic observations such as tolerance to osmotic stress or yield.

[0029] “Transgenic” means any cell, cell line, callus, tissue, plant part, or plant whose genome has been altered by the presence of a heterologous nucleic acid, such as a recombinant DNA construct. As used herein, the term “transgenic” includes those initial transgenic events and those resulting from those events through sexual hybridization or asexual reproduction, and does not cover genomic (chromosomal or extrachromosomal) alterations made through conventional plant breeding methods or through naturally occurring events such as random cross-fertilization, infection with a non-recombinant virus, transformation by a non-recombinant bacteria, non-recombinant transposition, or spontaneous mutation.

[0030] "Plant" includes indexes for whole plants, plant organs, plant tissues, seeds, and plant cells, as well as their offspring. Plant cells include, but are not limited to, cells from seeds, suspension cultures, plumules, meristematic regions, callus, leaves, roots, seedlings, gametophytes, sporophytes, pollen, and microspores. "Offspring" includes any subsequent generations of a plant.

[0031] In this application, the terms "comprising," "including," or variations thereof should be understood to include other elements, numbers, or steps besides those described. "Test plant" or "test plant cell" refers to a plant or plant cell in which genetic modification has taken effect, or a progeny cell of such a modified plant or cell containing the modification. "Control," "control plant," or "control plant cell" provides a reference point for measuring phenotypic changes in the test plant or plant cell.

[0032] Negative or control plants may include, for example: (a) wild-type plants or cells, i.e., plants or cells with the same genotype as the genetically modified starting material, the genetic modification producing the test plants or cells; (b) plants or plant cells with the same genotype as the starting material but transformed with an empty construct (i.e., a construct with no known effect on the target trait, such as a construct containing the target gene); (c) plants or plant cells that are non-transformed isomers of the test plants or plant cells; (d) plants or plant cells that are genetically identical to the test plants or plant cells but not exposed to conditions or stimuli that would induce the expression of the target gene; or (e) the test plants or plant cells themselves, which are under conditions where the target gene is not expressed.

[0033] Those skilled in the art will readily recognize that advances in molecular biology, such as site-specific and random mutagenesis, polymerase chain reaction methods, and protein engineering techniques, have provided a wide range of appropriate tools and procedures for modifying or engineering the amino acid sequences and potential gene sequences of proteins of interest in agriculture.

[0034] In some embodiments, the nucleotide sequence of this application may be modified to perform conserved amino acid substitutions. Principles and examples of conserved amino acid substitutions are further described below. In some embodiments, the nucleotide sequence of this application may be substituted without altering the amino acid sequence according to disclosed monocotyledonous codon preferences; for example, a codon encoding the same amino acid sequence may be substituted with a codon preferred by monocotyledons without changing the amino acid sequence encoded by the nucleotide sequence. In some embodiments, a portion of the nucleotide sequence in this application may be substituted with a different codon encoding the same amino acid sequence, thereby changing the nucleotide sequence without altering the encoded amino acid sequence. Conserved variants include those sequences that encode an amino acid sequence of one of the proteins of the embodiments due to genetic codon degeneracy. In some embodiments, a portion of the nucleotide sequence in this application may be substituted according to a codon preferred by monocotyledons. Those skilled in the art will recognize that amino acid additions and / or substitutions are generally based on the relative similarity of amino acid side-chain substituents, such as the hydrophobicity, charge, size, etc., of the substituents. Exemplary amino acid substituents having the various properties considered above are well known to those skilled in the art and include arginine and lysine; glutamic acid and aspartic acid; serine and threonine; glutamine and asparagine; and valine, leucine, and isoleucine. Guidance on appropriate amino acid substitutions that do not affect the biological activity of the target protein can be found in the model of Dayhoff et al. (1978), *Atlas of Protein Sequence and Structure* (Natl. Biomed. Res. Found., Washington, D. C.) (incorporated herein by reference). Conserved substitutions, such as replacing one amino acid with another amino acid having similar properties, can be performed. Identification of sequence consistency includes hybridization techniques. For example, a known nucleotide sequence, in whole or in part, can be used as a probe for selective hybridization with other corresponding nucleotide sequences present in cloned genomic DNA fragments or cDNA fragment groups (i.e., genomic libraries or cDNA libraries) from a selected organism. The hybridization probe may be a genomic DNA fragment, cDNA fragment, RNA fragment, or other oligonucleotide, and may be labeled with a detectable group such as 32P or other detectable markers. Thus, for example, hybridization probes can be prepared by labeling synthetic oligonucleotides based on sequences from the embodiment. Methods for preparing hybridization probes and constructing cDNA and genomic libraries are generally known in the art. Hybridization of the sequences can be performed under stringent conditions. As used herein, the terms "stringent conditions" or "stringent hybridization conditions" refer to conditions under which the probe will hybridize with its target sequence to a detectable extent (e.g., at least 2, 5, or 10 times the background) relative to hybridization with other sequences.Harshness conditions are sequence-dependent and vary across different environments. By controlling hybridization harshness and / or washing conditions, target sequences 100% complementary to the probe can be identified (homologous probe method). Alternatively, harshness conditions can be adjusted to allow for some sequence mismatches in order to detect lower similarities (heterologous probe method). Typically, probe lengths are less than about 1000 or 500 nucleotides. Typically, harshness conditions are those where the salt concentration is less than about 1.5 M Na ions at pH 7.0 to 8.3, typically about 0.01 M to 1.0 M Na ion concentration (or other salts), and the temperature conditions are: at least about 30°C for short probes (e.g., 10 to 50 nucleotides) and at least about 60°C for long probes (e.g., greater than 50 nucleotides). Harshness conditions can also be achieved by adding a destabilizing agent such as formamide. Exemplary low-tightness conditions include hybridization at 37°C using 30% to 35% formamide buffer, 1 M NaCl, and 1% SDS (sodium dodecyl sulfate), followed by washing at 50°C to 55°C in 1 to 2× SSC (20× SSC = 3.0 M NaCl / 0.3 M trisodium citrate). Exemplary medium-tightness conditions include hybridization at 37°C using 40% to 45% formamide, 1.0 M NaCl, and 1% SDS, followed by washing at 55°C to 60°C in 0.5× to 1× SSC. Exemplary high-tightness conditions include hybridization at 37°C using 50% formamide, 1 M NaCl, and 1% SDS, followed by a final wash at 60°C to 65°C in 0.1× SSC for at least about 20 minutes. Optionally, the wash buffer may contain about 0.1% to about 1% SDS. Hybridization duration is typically less than about 24 hours, typically from about 4 hours to about 12 hours. Specificity typically depends on post-hybridization washing, with key factors being the ionic strength and temperature of the final washing solution. The Tm (thermodynamic melting point) of DNA-DNA hybrids can be approximated by the formula from Meinkoth and Wahl (1984) Anal. Biochem. 138: 267-284: Tm = 81.5℃ + 16.6(logM) + 0.41(%GC) - 0.61(%formamide) - 500 / L; where M is the molar concentration of monovalent cations, %GC is the percentage of guanosine and cytosine nucleotides in the DNA, "formamide%" is the percentage of formamide in the hybridization solution, and L is the base pair length of the hybrid. Tm is the temperature at which 50% of the complementary target sequence hybridizes with a perfectly matched probe (at a given ionic strength and pH). Washing is typically performed at least until equilibration is reached and a low hybridization background level is achieved, such as for 2 hours, 1 hour, or 30 minutes. Each 1% mispairing should lower Tm by approximately 1°C; therefore, Tm, hybridization, and / or washing conditions can be adjusted to hybridize with the desired sequence of consistency.For example, if a sequence with ≥90% homology is required, the Tm can be lowered by 10°C. Typically, stringency conditions are chosen to be approximately 5°C lower than the Tm of the specific sequence and its complementary sequence at the defined ionic strength and pH. However, under very stringent conditions, hybridization and / or washing can be performed at 4°C lower than the stated Tm; under moderately stringent conditions, hybridization and / or washing can be performed at 6°C lower than the stated Tm; and under low stringency conditions, hybridization and / or washing can be performed at 11°C lower than the stated Tm.

[0035] In some embodiments, a fragment of a nucleotide sequence and the amino acid sequence it encodes is also included. As used herein, the term "fragment" refers to a portion of the nucleotide sequence of a polynucleotide of an embodiment or a portion of the amino acid sequence of a polypeptide. A fragment of the nucleotide sequence may encode a protein fragment that retains the biological activity of the native or corresponding full-length protein and thus has protein activity. Mutant proteins include biologically active fragments of native proteins containing consecutive amino acid residues that retain the biological activity of the native protein. Some embodiments also include transformed plant cells or transgenic plants containing a nucleotide sequence of at least one embodiment. In some embodiments, plants are transformed using an expression vector containing a nucleotide sequence of at least one embodiment and a promoter operatively linked thereto that drives expression in plant cells. Transformed plant cells and transgenic plants represent plant cells or plants whose genome contains a heteropolynucleotide. Generally, the heteropolynucleotide is stably integrated into the genome of the transformed plant cell or transgenic plant to pass the polynucleotide to offspring. The heteropolynucleotide may be integrated into the genome alone or as part of an expression vector. In some embodiments, the plants involved in this application include plant cells, plant protoplasts, plant cell tissue cultures capable of regenerating plants, plant callus, plant masses, and plant cells that are whole plants or parts of plants, such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruits, kernels, ears, rachis, husks, straw, roots, root tips, anthers, etc. This application also includes plant cells, protoplasts, tissues, callus, embryos, flowers, stems, fruits, leaves, and roots derived from transgenic plants of this application or their progeny, and thus at least partially containing the nucleotide sequences of this application.

[0036] In the context of nucleic acid amplification, the term "amplification" refers to any process in which additional copies of a selected nucleic acid (or its transcribed form) are produced. Common amplification methods include replication methods based on various polymerases, including polymerase chain reaction (PCR), ligase-mediated methods such as ligase chain reaction (LCR), and amplification methods based on RNA polymerases (e.g., via transcription).

[0037] An allele is “associated” with a trait when it is linked to the trait, and when the presence of an allele is an indicator that the desired trait or the form of the trait will occur in a plant containing the allele.

[0038] As used in this article, the term "quantitative trait locus" or "QTL" refers to a polymorphic locus that has at least one allele associated with differential expression of a phenotypic trait in at least one genetic context (e.g., in at least one breeding population or offspring). QTLs can function through single-gene or multi-gene mechanisms.

[0039] The term "QTL mapping" used in this article refers to the method of locating QTLs on a genetic map using methods similar to single-gene mapping, determining the distance between the QTL and the genetic marker (expressed as recombination rate). Depending on the number of markers, it can be divided into single-marker, double-marker, and multi-marker methods. Based on the statistical analysis methods, it can be divided into variance and mean analysis, regression and correlation analysis, moment estimation, and maximum likelihood estimation, etc. Based on the number of marker intervals, it can be divided into zero-interval mapping, single-interval mapping, and multi-interval mapping. In addition, there are comprehensive analysis methods that combine different methods, such as QTL composite interval mapping (CIM), multi-interval mapping (MIM), multiple QTL mapping, and multi-trait mapping (MTM).

[0040] The term "molecular marker" as used in this article refers to a specific DNA segment that reflects a certain difference in the genome of an individual or population.

[0041] The term "major gene" used in this article refers to a gene that determines a trait by a single gene. The term "minor gene" refers to several non-allelic genes that each has only a partial influence on the phenotype of the same trait; such genes are called additive genes or polygenes. In additive genes, each gene has only a small phenotypic effect, hence the name minor gene.

[0042] The term "inbred line" used in this article refers to a line obtained by self-pollinating under artificially controlled self-pollination for several generations, continuously eliminating undesirable rows of ears, and selecting individual plants with better agronomic traits for self-pollination, thereby obtaining a line with more uniform agronomic traits and a simpler genetic basis.

[0043] The term "backcross" as used in this article refers to the method of hybridizing the F1 generation with either of the two parents.

[0044] As used herein, the term "hybridization" or "hybrid" refers to the fusion of gametes (e.g., cells, seeds, or plants) that produce offspring through pollination. This term includes sexual hybridization (one plant being pollinated by another) and self-pollination (self-pollination, such as when pollen and ovules come from the same plant). The term "hybridization" refers to the gamete fusion that produces offspring through pollination.

[0045] The term "backcross" as used in this article refers to the process in which the hybrid offspring are repeatedly backcrossed with one of the parents. In a backcross scheme, the "donor" parent refers to the parental plant that possesses the desired gene or locus to be infiltrated. The "recipient" parent (used once or multiple times) or "recurrent" parent (used twice or multiple times) refers to the parental plant into which the gene or locus is infiltrated. The initial hybridization produces the F1 generation; then, the term "BC1" refers to the second use of the recurrent parent, "BC2" refers to the third use of the recurrent parent, and so on.

[0046] As used herein, the term "tightly linked" refers to a recombination frequency between two linked loci of equal to or less than about 10% (i.e., a segregation frequency of no more than 10 cM on a genome map). In other words, tightly linked loci co-segregate at least 90% of the time. Marker loci are particularly useful in this invention when they show a significant probability of co-segregation (linkage) with a desired trait (e.g., pathogen resistance). Tightly linked loci, such as marker loci and second loci, may show an intralocular recombination frequency of 10% or less, preferably about 9% or less, more preferably about 8% or less, more preferably about 7% or less, more preferably about 6% or less, more preferably about 5% or less, more preferably about 4% or less, more preferably about 3% or less, more preferably about 2% or less. In a highly preferred embodiment, the associated loci show a recombination frequency of about 1% or less, for example, about 0.75% or less, more preferably about 0.5% or less, more preferably about 0.25% or less. Two loci located on the same chromosome, and whose distance between them such that the recombination frequency between the two loci is less than 10% (e.g., approximately 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, 0.75%, 0.5%, 0.25%, or lower), are also referred to as "close to each other." In some cases, two different markers can have the same genomic map coordinates. In that case, the two markers are close enough that the recombination frequency between them is so low as to be undetectable.

[0047] Centiliter (“cM”) is a unit of measurement for recombination frequency. 1 cM is equal to the 1% probability that a marker at one locus will separate from a marker at a second locus after a single-generation cross.

[0048] A "favorable allele" is an allele at a specific locus that confers or contributes to an agronomically desired phenotype, such as improved kernel type in maize, and allows for the identification of plants with the agronomically desired phenotype. A marked "favorable" allele is a marker allele that cosegregates with the favorable phenotype.

[0049] A "gene map" is a description of gene linkages between loci on one or more chromosomes in a given species, typically presented as a graph or table. For each gene map, the distance between loci is measured by the frequency of recombination between them, and recombination between loci can be detected using various markers. A gene map is the product of the mapping population, the types of markers used, and the polymorphic potential of each marker across different populations. The order and genetic distance between loci in one gene map may differ from those in another. However, using a general frame of common markers allows for the association of information between one map and another. Those skilled in the art can use the frame of common markers to identify marker locations and loci of interest on individual gene maps.

[0050] "Genographic location" is the position on the gene map relative to the surrounding genetic markers on the same linkage group, where the specified marker can be found in a given population.

[0051] "Gene mapping" is a method for defining linkage relationships at gene loci, which is performed using genetic markers, marker segregation, and standard genetic principles of recombination frequency.

[0052] "Genetic recombination frequency" is the frequency of exchange events (recombination) between two loci. Recombination frequency can be observed after marker and / or segregation of traits following meiosis.

[0053] The term "genotype" is the genetic makeup of an individual (or group of individuals) at one or more loci, as opposed to an observable trait (phenotype). A genotype is defined by the alleles at one or more known loci that the individual has inherited from their parents. The term genotype can be used to refer to an individual's genetic makeup at a single locus, at multiple loci, or more generally, to refer to the genetic makeup of all the genes in an individual's genome.

[0054] "Germium" refers to the genetic material of an individual (e.g., a plant), a group of individuals (e.g., a plant strain, variety, or family), or a clone derived from or obtained from a strain, variety, species, or culture. Germium can be a part of an organism or cell, or can be isolated from an organism or cell. Germium typically provides genetic material and specific molecular structure that provides the physical basis for some or all of the heritable traits of an organism or cell culture. As used herein, germplasm includes cells, seeds, or tissues from which new plants can grow, or plant parts such as leaves, stems, pollen, or cells that can be cultured into a whole plant.

[0055] A “marker” is a nucleotide sequence or its encoded product (e.g., a protein) used as a reference point. For markers used to detect recombination, they need to detect differences or polymorphisms within the monitored population. For molecular markers, this means that differences at the DNA level are due to differences in multiple nucleotide sequences (e.g., SSRs, RFLPs, FLPs, and SNPs). Genomic variability can originate from any source, such as the presence and sequence of insertions, deletions, duplications, repetitive elements, point mutations, recombination events, or transposons. Molecular markers can be derived from the genome or expressed nucleic acids (e.g., ESTs) and can also refer to nucleic acids used as probes or primer pairs that can amplify sequence fragments using PCR-based methods.

[0056] Markers corresponding to genetic polymorphisms among population members can be detected using methods established in the art. These methods include, for example, DNA sequencing, PCR-based sequence-specific amplification methods, restriction fragment length polymorphism detection (RFLP), isoenzyme labeling detection, polynucleotide polymorphism detection (ASH) via allele-specific hybridization, amplified variable sequence detection of plant genomes, autonomous sequence replication detection, simple repeat sequence detection (SSR), single nucleotide polymorphism detection (SNP), or amplified fragment length polymorphism detection (AFLP). Established methods are also known for detecting expressed sequence tags (ESTs) and SSR markers derived from EST sequences, as well as randomly amplified polymorphic DNA (RAPD).

[0057] A “marker allele” or “marker locus allele” can refer to one of several polymorphic nucleotide sequences located at a marker locus in a population, which is polymorphic with respect to the marker locus.

[0058] A “labeled probe” is a nucleic acid sequence or molecule that can be used to identify the presence or absence of a marker locus by nucleic acid hybridization, such as a nucleic acid molecular probe complementary to a marker locus sequence. A labeled probe containing 30 or more adjacent nucleotides (“all or part” of the marker locus sequence) can be used for nucleic acid hybridization. Alternatively, in some respects, a molecular probe refers to any type of probe capable of distinguishing (i.e., genotyping) specific alleles present at a marker locus.

[0059] As described above, when identifying linked loci, the term "molecular marker" can be used to refer to a genetic marker, or its encoded product (e.g., a protein) used as a reference point. Markers can be derived from genomic nucleotide sequences or expressed nucleotide sequences (e.g., from spliced ​​RNA, cDNA, etc.), or from encoded polypeptides. The term also refers to nucleic acid sequences complementary to or flanked by marker sequences, such as nucleic acids used as probes or primer pairs capable of amplifying marker sequences. A "molecular marker probe" is a nucleic acid sequence or molecule that can be used to identify the presence or absence of a marker locus, such as a nucleic acid probe complementary to a marker locus sequence. Alternatively, in some respects, a molecular probe refers to any type of probe capable of distinguishing (i.e., genotype) specific alleles present at a marker locus. Nucleic acids are "complementary" when they hybridize specifically in solution, for example, according to the Watson-Crick base pairing principle. Some markers described herein are also called hybridization markers when located in insertion / deletion regions, such as the non-collinear regions described herein. This is because insertion regions are polymorphisms concerning the absence of insertions. Therefore, the marker only needs to indicate the presence or absence of the insertion / deletion region. Any suitable marker detection technique can be used to identify such hybridization markers, such as the KASP technique.

[0060] This invention located a major-effect QTL controlling maize kernel shape and weight within an 11 kb region on chromosome 7 using association and linkage analysis. This QTL contains only one functionally annotated gene, Zm00001d021956. Further gene editing, knockout, and overexpression studies confirmed that this gene controls maize kernel shape and weight traits, and it was named Zm00001d021956. ZmKW7 .

[0061] The present invention further identified two related to ZmKW7 These are genes with extremely high homology. Through gene editing experiments, it was unexpectedly discovered that these two genes regulate grain size in a similar way. ZmKW7 different.

[0062] This invention provides a method for increasing the width and / or weight of corn kernels, and provides the sequences of mutant genes and their encoded proteins that increase the width and / or weight of corn kernels.

[0063] The following examples are used to illustrate the present invention, but are not intended to limit the scope of the invention. Any modifications or substitutions made to the methods, steps, or conditions of the present invention without departing from the spirit and essence of the invention are within the scope of this application. Unless otherwise specified, the examples are conducted under conventional experimental conditions, such as those described in Sambrook et al.'s *Molecular Cloning: A Laboratory Manual* (Sambrook J & Russell DW, 2001), or according to the conditions recommended in the manufacturer's instructions. Unless otherwise specified, the chemical reagents used in the examples are all commercially available and conventional methods well known to those skilled in the art.

[0064] Example 1: Positioning of corn kernel-type QTL

[0065] This invention utilizes the maize inbred line ZHENG58 (Zheng 58, an inbred line developed in 1995, now available commercially) exhibiting large-kernel characteristics and the maize inbred line SK (bred from local varieties by Professor Li Jiansheng of China Agricultural University and Professor Yan Jianbing of Huazhong Agricultural University) exhibiting small-kernel characteristics to construct a recombinant inbred line (RIL) population. Linkage and association analyses are then used to locate QTL loci related to maize kernel type. The specific steps are as follows:

[0066] (1) Chain analysis.

[0067] The RIL population was constructed by crossing ZHENG58 with SK to obtain F1, followed by six generations of self-pollination. Phenotypic data on grain length, width, thickness, and 100-grain weight were obtained through multi-year, multi-location field trials. The method involved randomly selecting grains from the middle of 20 regularly shaped ears for measurement of grain length, width, and thickness. The 100-grain weight was measured three times for each material and averaged. The phenotypic data from different environments were integrated using proc mixed to obtain the Best Linear Unbiased Estimate (BLUP). Genotypic data for the RIL population were obtained using 56,110 SNPs from the Illumina MaizeSNP50 BeadChip (Pan et al., 2016). Data analysis was performed using Composite Interval Mapping (CIM) in WinQTLCart 2.5 with a step rate of 0.5 cM and a threshold of 2.5. The analysis results showed that there was a signal peak on chromosome 7 of maize, which may be a genetic locus affecting maize kernel shape.

[0068] (2) Association analysis.

[0069] Further association analysis and fine mapping of maize kernel-type QTLs were performed. Genotypic data for the association analysis used existing 1.25M SNPs (see: Liu H, Luo X, Niu L, et al. Distant eQTLs and Non-coding Sequences Play Critical Roles in Regulating Gene Expression and Quantitative Trait Variation in Maize[J]. Molecular Plant, 2017, 10(3):414-426.). Phenotypic data used for the association analysis were the means of two replicates per year (see Table 1). The model used was a mixed linear model in TASSEL 3.0 software, which could simultaneously control for the influence of population structure and kinship on the association results. The minimum allele frequency (MAF) of the markers used was ≥0.05. Genotypic analysis of five recombinant near-isogenic lines (R-1 to R-5) revealed that the candidate gene is located on chromosome 7 within an 11 kb region ranging from 160.60 to 163.97 Mb. This gene is found in the near-isogenic line NIL. SK and NIL ZHENG58 Resequencing analysis was performed on the 11 kb candidate region, and the maize reference genome database was queried. Only one functionally annotated gene, Zm00001d021956, was found within this region (see the schematic diagram of gene localization). Figure 1 The marker information used for localization is shown in Table 2. Based on the annotation information on the maize reference genome, Zm00001d021956 (named...) ZmKW7 The functional annotation is GNAT - transcription factor 33. GNAT is a very large family of transcription factors in organisms, and there is no publicly available data showing that this gene is associated with the kernel type trait in maize.

[0070] Example 2 Functional verification of maize kernel type gene

[0071] To further validate the function of the candidate genes, the CRISPR-Cas9 gene editing system was used to target the located genes. ZmKW7Knockout was performed, and the relationship between the gene and the trait was determined by the kernel shape phenotype of the knockout maize. Using the online software CRISPR-P 2.0 (http: / / crispr.hzau.edu.cn / CRISPR2 / ), two target editing sites, TCGTGGTGGTCATCTCCAGC and GTATTATTCTGGCGAGCAA, were selected on the gene to design gRNAs. A fusion unit sequence “U6 promoter1-gRNA1-sgRNA-U6 promoter2-gRNA2-sgRNA” was synthesized, excised from the intermediate vector PUC57 using a double enzyme digestion system, and ligated to the backbone vector CPB-ZmUbi-hspCas9 (using the same editing vector backbone as in patent CN112375130A) via homologous recombination. The recombinant vector was transformed into E. coli competent cells DH5α, and the target sequence was detected. After confirming the sequence was correct, the gene was genetically transformed into the maize inbred line KN5585. Two target sites were detected in the obtained T0 generation transformed plants, and the editing type was analyzed based on the sequencing sequence. The analysis showed that ZmKW7 has two editing types: one involves the deletion of three bases, resulting in a one-amino acid deletion (KO1); the other involves a one-base deletion, leading to premature translation termination (KO2). Analysis of the grain shape traits of the two edited materials revealed that grain length, grain width, and 100-grain weight were all reduced after gene knockout (Table 3). These results indicate that... ZmKW7 Positively regulates corn kernel size.

[0072]

[0073] Example 3: Improving maize kernel shape traits through overexpression

[0074] because ZmKW7 This gene positively regulates maize kernel size; therefore, overexpression of this gene can improve kernel shape and increase maize yield. The specific method is as follows:

[0075] First, the full-length genome sequence was amplified segment by segment. Then, the amplified sequence was recombined into the vector pCAMBIA using the ClonExpress II kit (www.vazyme.com). The constructed vector was preceded by the maize ubiquitin gene promoter, and the commonly used nos terminator was used. The recombinant vector was transformed into competent E. coli DH5α strain, and positive clones were obtained. After the target sequence was confirmed to be correct, it was transformed into the maize inbred line KN5585. The transgenic plants were then confirmed. ZmKW7 After gene expression levels were increased compared to the receptor control, the grain weight per 100 grains of the overexpression material was investigated, and it was found that... ZmKW7Gene overexpression can significantly increase the 100-kernel weight of maize. Therefore, overexpression... ZmKW7 Genes can be used to increase corn yield.

[0076] Example 4: Obtaining and editing homologous genes

[0077] This invention further analyzes ZmKW7 Among the homologous genes, two members with the highest homology were found: Zm00001d047817 and Zm00001d048574.

[0078] The two genes were edited according to the gene editing method described in Example 2. The target sequence of the Zm00001d047817 gene is GTATTTATTCTGGCGAGCAA, and the target sequence of the Zm00001d048574 gene is GCGAGTTGTCGTGGTCGGCG. Therefore, the gRNA sequences expressed by the editing vector are as follows: GUAUUUAUUCUGGCGAGCAAguuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuuuu and GCGAGUUGUCGUGGUCGGCGguuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuuuuu. After obtaining the transformed plants, the sequence of the target region and the grain shape traits of the edited plants were determined. It was found that a single base variation (GTATTTATTCTGG) occurred in the Zm00001d047817 gene. C GAGCAA becomes GTATTTATTCTGG T The edited plant (047817-KO) showed significantly reduced grain length, increased grain width, and significantly increased 100-grain weight; while the Zm00001d048574 gene underwent a single base insertion (GCGAGTTGTCGTGGTCGGCG changed to GCGAGTTGTCGTGGTCG). A GCG), leading to premature termination of translation, resulted in a significant decrease in grain length, an increase in grain width, and a significant increase in 100-grain weight in the edited plant (048574-KO) (Table 4). This indicates that Zm00001d047817 and Zm00001d048574 are related to... ZmKW7 Genes regulate corn kernel shape in completely different ways.

[0079] The above embodiments show that mutations in Zm00001d047817 and Zm00001d048574 can significantly increase the grain width and / or grain weight traits of maize. The amino acid sequence encoded by the mutated gene is as follows: MAAEINGGFLAAGGPRQHRGGLGCGRCFQNISLLHGLGIKFVLVPGTHIQIDKLLSEIGNKAKYVGQYRITDEDARKAAMDAAGRIRLTIEAKLSPGPPMLNLRRHGVIGRWHGLVDSIASGNFLGAKRRGVVNGIDYGFTEEVTKIDVSRIRERLDSDSIVVISNMGFSSSGDVLNCNTNEVATACALALEADKLICVVDGQVFDEHGRAIQYMSIEEADFLIRKRARQSDIAASYVKVVDEEGINSLHKGDNRPSLSPKAYINGYAASFRNGLGFNNGNGIYSGEQGFAIGGEERLSRSNADPFQNGLGFNNGNGIYSGEQGFAIGGEERLSRSNGYLSELVAAACVP or

[0080] As shown in the diagram, introducing the mutant gene into other small-kernel maize recipients through methods such as cross-pollination can improve the recipient's kernel shape traits, increase kernel width and / or kernel weight, and ultimately increase maize yield.

[0081] Although the present invention has been described in detail above with general descriptions and specific embodiments, modifications or improvements can be made to it, which will be obvious to those skilled in the art. Therefore, all such modifications or improvements made without departing from the spirit of the present invention fall within the scope of protection claimed by the present invention.

Claims

1. The application of a gene in regulating the kernel weight trait of maize, characterized in that: The gene in question is the gene with the number Zm00001d048574 in the MaizeGDB database.

2. A method for increasing the weight of corn kernels, characterized in that: Inhibit the expression and / or activity of the protein encoded by the gene of claim 1 in maize, and select maize plants with increased kernel weight.

3. The method according to claim 2, characterized in that: The protein sequence is as follows.

4. The method according to any one of claims 2-3, characterized in that: The methods for inhibiting protein expression and / or activity include any one of gene editing, RNA interference, and T-DNA insertion.

5. The method according to claim 4, characterized in that: The gene editing was performed using the CRISPR / Cas9 method.

6. The method according to claim 5, characterized in that: The DNA sequence of the genomic target region in maize obtained by the CRISPR / Cas9 method is shown in GCGGAGTTGTCGTGGTCGGCG.

7. A reagent kit for increasing the weight of corn kernels, characterized in that: Including any of the following: (1) An RNA molecule capable of recognizing the target sequence described in claim 6; (2) The DNA molecule encoding the RNA described in (1); (3) A vector expressing the RNA described in (1); The sequence of the RNA molecule is GGCGAGUUGUCGUGGUCGGCGguuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuuuu.