A method for increasing the size or weight of a plant seed
By using CRISPR gene editing technology to carry out non-frameshift mutations in plants such as rice, we created variant sites and mutant proteins that increase seed size or weight, solving the problem of lack of effective sites in existing technologies and achieving a significant increase in seed size and weight.
Patent Information
- Application Number
- CN202110532350.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-17
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2041-05-17
AI Technical Summary
In the existing technology, except for the D364E mutation of the rice OsPPKL1 protein, it is not clear how mutations at other sites can effectively increase seed size or weight.
Using CRISPR gene editing technology, non-frameshift mutations were performed on the genomic coding position near the D364 site of the PPKL1 homologous protein in plants such as rice, non-frameshift mutant proteins were created, and mutation sites and mutant proteins that increased seed size or weight were screened out.
It achieves a significant increase in seed size and weight without affecting the normal growth and development of the plant, and provides new mutation sites and mutant proteins to achieve the same or better effects.
Smart Images

Figure CN115368445B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for increasing the size or weight of plant seeds, and belongs to the field of plant genetic engineering. Background Art
[0002] Seed (grain) size is an important trait affecting crop yield. qGL3 / OsPPKL1 (LOC_Os03g44500) in rice encodes a protein phosphatase OsPPKL1 containing two Kelch functional domains, which plays a negative regulator role in the regulation of rice grain length. The rare allele qgl3 at D364E located in the conserved AVLDT region of the second Kelch functional domain in OsPPKL1 leads to a long-grain phenotype (Zhang X, Wang J, Huang J, et al. Rare allele of OsPPKL1 associated with grain length causes extra-large grain and asignificant yield increase in rice[J]. Proc Natl Acad Sci US A. 2012, 109(52): 21534-21539.).
[0003] Cytokinins play a crucial regulatory role in plant seed development. Mutations at position D364 in the PPKL1 protein can activate the cytokinin response, promote cell division, and thus increase seed (grain) size. However, it is unknown what other mutations at these positions can achieve the same or even better effect in increasing seed size.
[0004] To address the above problems, the present invention utilizes CRISPR gene editing vectors to mutate the genomic coding position near the D364 site of the PPKL1 homologous protein in plants, create non-frameshift mutant proteins, and identify the grain length and grain weight performance of the edited plants, thereby screening for mutation sites and mutant proteins that increase plant seed size or increase seed weight. Summary of the Invention
[0005] The present invention identifies rice mutants and uses gene editing technology to mutate the genomic coding position near the D364 site of the PPKL1 homologous protein in the plant, creates a non-frameshift mutant protein, and identifies the grain length and grain weight of the edited plants, thereby screening for mutation sites and mutant proteins that increase plant seed size or increase seed weight.
[0006] The present invention provides an amino acid variation site of a plant protein, characterized in that the site is located at any one of the following positions:
[0007] 1) positions 84-420 of the sequence shown in SEQ ID NO: 1; or
[0008] 2) positions 156-420 of the sequence shown in SEQ ID NO: 1; or
[0009] 3) positions 362-375 of the sequence shown in SEQ ID NO: 1; or
[0010] 4) the corresponding positions of an amino acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% similarity to the sequence shown in positions 84-420 of SEQ ID NO: 1; or
[0011] 5) the corresponding positions of an amino acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% similarity to the sequence shown in positions 156-420 of SEQ ID NO: 1; or
[0012] 6) the corresponding positions of an amino acid sequence having at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% similarity to the sequence shown in positions 362-375 of SEQ ID NO: 1.
[0013] The present invention also provides a protein, characterized in that the amino acid sequence of the protein comprises the above-mentioned amino acid sites, and one or more amino acid mutations occur in the amino acid sites, wherein the mutation is in the form of amino acid substitution or deletion;
[0014] In some embodiments, the variation comprises any one of the following:
[0015] 1)T365R;
[0016] 2)D364N;
[0017] 3) D364E and T365R;
[0018] 4) T365R and 366A-370W deletions;
[0019] 5) 364D~369V are missing;
[0020] 6) 362V~375S are missing;
[0021] Each amino acid mutation position corresponds to the amino acid position of the sequence shown in SEQ ID NO: 1; the sequence shown in SEQ ID NO: 1 is the amino acid sequence of rice OsPPKL protein.
[0022] The present invention also provides a nucleic acid, characterized in that the nucleic acid encodes the above-mentioned protein;
[0023] The present invention also provides a method for increasing the size or weight of plant seeds, characterized in that: the above-mentioned protein is expressed in the plant, and plants with enlarged seeds or increased seed weight are selected;
[0024] In some embodiments, the method comprises site-directed mutagenesis of an endogenous plant protein or introduction of an exogenous protein.
[0025] The present invention also provides a plant genome editing target sequence for increasing plant seed size or improving plant seed weight, characterized in that the target sequence is selected from any one or several target sequences listed in Table 4.
[0026] The present invention also provides a kit for increasing the size or weight of plant seeds, characterized in that it comprises any one of the following:
[0027] (1) sgRNA molecules, capable of targeting any one or more target sequences listed in Table 4;
[0028] (2) a DNA molecule encoding the sgRNA;
[0029] (3) A vector for expressing the sgRNA.
[0030] The above-mentioned sgRNA vector or sgRNA encoding DNA molecule or sgRNA molecule can mutate the target protein in the plant through gene editing, thereby creating the above-mentioned mutant protein, thereby achieving the effect of increasing the size or weight of plant seeds.
[0031] In some embodiments, the plant is a monocot or a dicot; the monocot includes but is not limited to rice, corn, and wheat; the dicot includes but is not limited to soybean and Arabidopsis.
[0032] The beneficial effects of the present application are that although it is known that the D364E site of the rice OsPPKL1 protein can increase grain length, there is no technical inspiration for what other sites in this protein and other homologous proteins can be mutated to achieve the same or better effect of increasing seed size. The present application uses a CRISPR gene editing vector to mutate the genomic coding position of the amino acid region upstream and downstream of the D364 site of the PPKL1 homologous protein in plants, to create non-frameshift mutant proteins, and to identify the seed grain length and weight performance of the obtained edited plants, to screen for mutant proteins with increased plant seed size or increased seed weight. By non-frameshift mutation of these mutant sites or expression of exogenous mutant proteins, the technical effect of increasing plant seed size or increasing seed weight can be achieved. BRIEF DESCRIPTION OF DRAWINGS
[0033] Figure 1 : s48 mutant phenotype. A: plant performance during the grain filling stage, scale = 20 cm; B: grain performance at the mature stage, scale = 5 mm. ZH11: control Zhonghua 11; s48(+ / -): mutant with heterozygous genotype; s48: mutant with homozygous genotype.
[0034] Figure 2 : Test of callus greening of Zhonghua 11 (ZH11) and s48 induced by combination of cytokinin and auxin. Different concentrations of t-zeatin and naphthylacetic acid (NAA) were used for analysis. The calli of ZH11 and S48 were cultured on medium with different concentrations of t-zeatin and NAA for 1 month, and then the green calli were collected for photography.
[0035] Figure 3 : Schematic diagram of T-DNA region of cytosine base editor. RB: right border of T-DNA region; U6: U6 promoter; gRNA: target recognition sequence; UBI: ubiquitin promoter; APO1: cytosine editor; XTEN: signal peptide; hSpCas9: Cas9 gene; 35S: 35S promoter; HYG: hygromycin selection marker; PolyA: terminator.
[0036] Figure 4 : CRISPR / Cas9 vector map.
[0037] Figure 5 : Amino acid sequence display of different editing types and corresponding grain performance. KY131: control air-cultivated 131.
[0038] Figure 6: Plant architecture of different edited plants. KY131: Recipient control Kongyu 131.
[0039] Figure 7 : Grain performance of overexpressing rice. NY: N364+499Y; DY: D364+499Y; NH: N364+499H; DH: D364+499H.
[0040] Figure 8 : Amino acid alignment of the editing target sites of rice and Arabidopsis PPKL homologous proteins. The area covered by the black line is the mutation site region. DETAILED DESCRIPTION
[0041] The following definitions and methods are provided to better define this application and to guide those skilled in the art in practicing this application. Unless otherwise noted, terms are to be understood according to conventional usage by those skilled in the relevant art. All patent documents, academic papers, industry standards, and other publications cited herein are hereby incorporated by reference in their entirety.
[0042] As used herein, "plant" is any plant, including whole plants, plant cells, plant organs, plant protoplasts, plant cell tissue cultures from which plant cells can be regenerated, plant calli, plant cells that are intact in plants or plant parts and include embryos, pollen, ovules, seeds, leaves, flowers, branches, fruit, roots, root tips, anthers, and the like. Nucleic acids are written left to right in 5' to 3' orientation, unless otherwise indicated; amino acid sequences are written left to right in amino to carboxyl orientation, unless otherwise indicated. Amino acids can be referred to herein by either the commonly accepted single-letter codes, or by accepted three-letter codes. Nucleotides, unless otherwise indicated, are referred to by their commonly accepted single-letter codes. Numeric ranges are inclusive of the numbers defining the range. As used herein, "nucleic acid" includes polynucleosides or polynucleotides in either single- or double-stranded form, and unless otherwise limited, includes known analogues of natural nucleotides (e.g., peptide nucleic acids) that function in a manner similar to naturally occurring nucleotides in that they hybridize with single-stranded nucleic acids in a manner similar to naturally-occurring nucleotides. As used herein, the term "encoding" or "encoded" with respect to a specified nucleic acid sequence is intended to mean that the specified nucleic acid sequence includes a nucleotide sequence, such that the nucleotide sequence encodes a specific protein. The information on the codons is used to specify the amino acid sequence of the protein. As used herein, "full-length sequence" with respect to a particular polynucleotide or its encoded protein refers to the entire nucleic acid sequence or the entire amino acid sequence having the native (non-synthetic) endogenous sequence. The full-length polynucleotide encodes the full-length, catalytically active form of the particular protein. The terms "polypeptide," "polypeptide," and "protein" are used interchangeably herein to refer to a polymer of amino acid residues. The term is used to refer to a polymer of amino acid residues, in which one or more of the amino acid residues are artificial chemical mimics of corresponding naturally occurring amino acids. The term is also used to refer to naturally occurring amino acid polymers. The terms "residue" or "amino acid residue" or "amino acid" are used interchangeably herein to refer to an amino acid that is incorporated into a protein, polypeptide, or peptide (collectively "protein"). The amino acid can be a naturally occurring amino acid and, unless otherwise limited, can include known analogs of naturally occurring amino acids that can function in a manner similar to the naturally occurring amino acid.
[0043] The term "trait" refers to a physiological, morphological, biochemical, or physical characteristic of a plant or a particular plant material or cell. In some cases, the characteristic is visible to the human eye, such as seed or plant size, or can be measured by biochemical techniques, such as detecting protein, starch, or oil content of seeds or leaves, or by observing metabolic or physiological processes, for example, by measuring tolerance to water deprivation or specific salt or sugar or nitrogen concentrations, or by observing expression levels of one or more genes, or by agronomic observations such as osmotic stress tolerance or yield.
[0044] "Transgenic" refers to any cell, cell line, callus, tissue, plant part, or plant whose genome is altered by the presence of a heterologous nucleic acid, such as a recombinant DNA construct. As used herein, the term "transgenic" includes those original transgenic events and those generated from the original transgenic events by sexual crosses or asexual propagation, and does not encompass genomic (chromosomal or extrachromosomal) alterations made by conventional plant breeding methods or by naturally occurring events, such as random cross-fertilization, non-recombinant viral infection, non-recombinant bacterial transformation, non-recombinant transposition, or spontaneous mutation.
[0045] In this application, the words "comprises," "comprising," or variations thereof are to be understood as including, in addition to the described elements, numbers, or steps, other elements, numbers, or steps. A "test plant" or "test plant cell" refers to a plant or plant cell in which a genetic modification has been effected, or a progeny of such a modified plant or cell that contains the modification. A "control" or "control plant" provides a reference point for measuring phenotypic changes in the test plant.
[0046] Negative or control plants can include, for example: (a) wild-type plants or cells, i.e., plants or cells having the same genotype as the genetically modified starting material that produced the test plant or cell; (b) plants or plant cells having the same genotype as the starting material but that have been transformed with an empty construct (i.e., with a construct that has no known effect on the trait of interest, such as a construct comprising a marker gene); (c) plants or plant cells that are non-transformed segregants of the test plant or plant cell; (d) plants or plant cells that are genetically identical to the test plant or plant cell but that have not been exposed to conditions or stimuli that would induce expression of the gene of interest; or (e) the test plant or plant cell itself, which is under conditions where the gene of interest is not expressed.
[0047] Those skilled in the art will readily recognize that advances in the field of molecular biology, such as site-specific and random mutagenesis, polymerase chain reaction methods, and protein engineering techniques, provide a wide range of appropriate tools and procedures for modifying or engineering the amino acid sequence and underlying gene sequence of proteins of agricultural interest.
[0048] In some embodiments, the nucleotide sequences of the present application can be altered to make conservative amino acid substitutions. The principles and examples of conservative amino acid substitutions are further described below. In certain embodiments, the nucleotide sequences of the present application can be substituted without changing the amino acid sequence according to the disclosed monocot codon preferences, for example, codons encoding the same amino acid sequence can be replaced with codons preferred by monocots without changing the amino acid sequence encoded by the nucleotide sequence. In some embodiments, part of the nucleotide sequence in the present application is replaced with different codons encoding the same amino acid sequence, thereby not changing the amino acid sequence encoded by the nucleotide sequence while changing the nucleotide sequence. Conservative variants include those sequences that encode the amino acid sequence of one of the proteins of the embodiments due to the degeneracy of the genetic code. In some embodiments, part of the nucleotide sequence in the present application is replaced according to the monocot codon preference. Those skilled in the art will recognize that amino acid additions and / or substitutions are generally based on the relative similarity of the amino acid side chain substituents, for example, the hydrophobicity, charge, size, etc. of the substituents. Exemplary amino acid substitution groups with various aforementioned properties are well known to those skilled in the art and include arginine and lysine; glutamic acid and aspartic acid; serine and threonine; glutamine and asparagine; and valine, leucine and isoleucine. Guidance on appropriate amino acid substitutions that do not affect the biological activity of the protein of interest can be found in the model of Dayhoff et al. (1978) Atlas of Protein Sequence and Structure (Natl. Biomed. Res. Found., Washington, DC) (incorporated herein by reference). Conservative substitutions such as replacing one amino acid with another amino acid having similar properties can be performed. Sequence identity identification includes hybridization techniques. For example, all or part of a known nucleotide sequence is used as a probe for selective hybridization with other corresponding nucleotide sequences present in cloned genomic DNA fragments or cDNA fragment groups (i.e., genomic libraries or cDNA libraries) from a selected organism. The hybridization probe can be a genomic DNA fragment, a cDNA fragment, an RNA fragment or other oligonucleotide, and can be marked with a detectable group such as 32P or other detectable markers. Thus, for example, a hybridization probe can be prepared by labeling a synthetic oligonucleotide based on the embodiment sequence. The method for preparing hybridization probes and constructing cDNA and genomic libraries is generally known in the art. The hybridization of the sequence can be carried out under stringent conditions. As used herein, the term "stringent conditions" or "stringent hybridization conditions" represents following conditions, i.e., under these conditions, relative to hybridization with other sequences, the probe will hybridize with its target sequence to a greater extent (e.g., at least 2 times, 5 times or 10 times of background) that can be detected.Stringent conditions are sequence-dependent and vary in different environments. By controlling hybridization stringency and / or controlling washing conditions, a target sequence that is 100% complementary to the probe can be identified (homologous probe method). Alternatively, stringent conditions can be adjusted to allow some sequence mismatches in order to detect lower similarities (heterologous probe method). Typically, the probe length is less than about 1000 or 500 nucleotides. Typically, stringent conditions are those in which the salt concentration is less than about 1.5 M Na ions, typically about 0.01 M to 1.0 M Na ion concentration (or other salts) at pH 7.0 to 8.3, and the temperature is: when used for short probes (e.g., 10 to 50 nucleotides), at least about 30°C; when used for long probes (e.g., greater than 50 nucleotides), at least about 60°C. Stringent conditions can also be achieved by adding destabilizing agents such as formamide. Exemplary low stringency conditions include hybridization at 37°C using 30% to 35% formamide buffer, 1M NaCl, 1% SDS (sodium dodecyl sulfate), and washing in 1× to 2× SSC (20× SSC = 3.0M NaCl / 0.3M trisodium citrate) at 50°C to 55°C. Exemplary moderate stringency conditions include hybridization at 37°C using 40% to 45% formamide, 1.0M NaCl, 1% SDS, and washing in 0.5× to 1× SSC at 55°C to 60°C. Exemplary high stringency conditions include hybridization at 37°C using 50% formamide, 1M NaCl, 1% SDS, and a final wash in 0.1× SSC at 60°C to 65°C for at least about 20 minutes. Optionally, the wash buffer may contain about 0.1% to about 1% SDS. Duration of hybridization is typically less than about 24 hours, typically about 4 hours to about 12 hours. Specificity generally depends on post-hybridization washes, with the key factors being the ionic strength and temperature of the final wash solution. The Tm (thermodynamic melting point) of a DNA-DNA hybrid can be approximated by the formula of Meinkothand Wahl (1984) Anal. Biochem. 138:267-284: Tm = 81.5°C + 16.6 (log M) + 0.41 (% GC) - 0.61 (% formamide) - 500 / L, where M is the molar concentration of monovalent cations, % GC is the percentage of guanosine and cytosine nucleotides in the DNA, "% formamide" is the percentage of formamide in the hybridization solution, and L is the base pair length of the hybrid. The Tm is the temperature (under defined ionic strength and pH) at which 50% of the complementary target sequence hybridizes to a perfectly matched probe. Washes are typically performed at least until equilibrium is reached and low background levels of hybridization are achieved, such as for 2 hours, 1 hour, or 30 minutes. Each 1% mismatch should reduce the Tm by about 1°C; thus, the Tm, hybridization, and / or wash conditions can be adjusted to hybridize to sequences of the desired identity. For example, if sequences with ≥90% identity are desired, the Tm can be reduced by 10°C.Generally, stringent conditions are selected to be about 5°C lower than the Tm for the specific sequence and its complement at a defined ionic strength and pH. However, under very stringent conditions, hybridization and / or washing can be performed at 4°C below the Tm; under moderately stringent conditions, hybridization and / or washing can be performed at 6°C below the Tm; and under low stringency conditions, hybridization and / or washing can be performed at 11°C below the Tm.
[0049] In some embodiments, fragments of nucleotide sequences and the amino acid sequences they encode are also included. As used herein, the term "fragment" refers to a portion of the nucleotide sequence of a polynucleotide of an embodiment or a portion of the amino acid sequence of a polypeptide. Fragments of nucleotide sequences can encode protein fragments that retain the biological activity of a native or corresponding full-length protein and thus have protein activity. Mutant proteins include biologically active fragments of native proteins that contain contiguous amino acid residues that retain the biological activity of the native protein. Some embodiments also include transformed plant cells or transgenic plants that contain the nucleotide sequence of at least one embodiment. In some embodiments, plants are transformed using an expression vector that contains the nucleotide sequence of at least one embodiment and a promoter that drives expression in plant cells operably linked thereto. Transformed plant cells and transgenic plants refer to plant cells or plants that contain heterologous polynucleotides in their genomes. Generally speaking, the heterologous polynucleotides are stably integrated in the genome of the transformed plant cells or transgenic plants so that the polynucleotides are passed on to future generations. The heterologous polynucleotides can be integrated into the genome individually or as part of an expression vector. In some embodiments, the plants involved in the present application include plant cells, plant protoplasts, plant cell tissue cultures that can regenerate plants, plant calli, plant masses and plant cells, which are complete plants or parts of plants, such as embryos, pollen, ovules, seeds, leaves, flowers, branches, fruits, kernels, ears, cobs, shells, stalks, roots, root tips, anthers, etc. The present application also includes plant cells, protoplasts, tissues, calli, embryos, flowers, stems, fruits, leaves and roots derived from the transgenic plants of the present application or their progeny, and thus at least partially comprising the nucleotide sequence of the present application.
[0050] The following examples are intended to illustrate the present invention but are not intended to limit the scope of the present invention. Without departing from the spirit and substance of the present invention, modifications or replacements made to the inventive method, steps or conditions are intended to fall within the scope of this application. Unless otherwise specified, the examples are based on conventional experimental conditions, such as Sambrook et al.'s Molecular Cloning Laboratory Manual (Sambrook J & Russell DW, Molecular cloning: a laboratory manual, 2001), or the conditions recommended by the manufacturer's instructions. Unless otherwise specified, the chemical reagents used in the examples are conventional commercially available reagents, and the technical means used in the examples are conventional means well known to those skilled in the art.
[0051] Example 1 Identification of large-grain rice mutants
[0052] The present invention uses sodium azide to induce mutation in rice variety Zhonghua 11 (a common rice variety, selected and bred by the Chinese Academy of Agricultural Sciences), and isolates a large-grain mutant s48 (attached Figure 1 Compared to the wild-type Zhonghua11, the s48 mutant showed a 35.3% increase in kernel length and a 52% increase in 1,000-kernel weight (Table 1). The heterozygous mutant s48(+ / -) exhibited a distinctly intermediate phenotype between Zhonghua11 and the homozygous mutant s48, indicating that the mutation is semidominant. Furthermore, the increased kernel length in s48 was due to an increase in cell number, rather than cell size.
[0053] Table 1 Grain phenotypes of rice s48 mutant
[0054]
[0055] "*" indicates significant difference compared with the wild-type control (α=0.05).
[0056] The mutant's response to cytokinin application was further analyzed. Cytokinins are widely used to induce callus greening during tissue culture regeneration. When the auxin NAA and the cytokinin t-zeatin were used in combination, the callus of s48 became greener than that of Zhonghua 11. At an appropriate concentration of auxin NAA, the callus of s48 also turned green more effectively than that of Zhonghua 11 at low concentrations of t-zeatin (see Appendix). Figure 2 Furthermore, during dark culture, the effect of cytokinin on slowing dark-induced leaf senescence in S48 was more pronounced than in Zhonghua11. Furthermore, the expression of A-type RRs, a type of cytokinin-responsive gene, was significantly increased in the panicles of the s48 mutant. These results suggest that the cytokinin response in the s48 mutant is enhanced.
[0057] The present invention further constructed a backcross population of s48 and Zhonghua 11. Resequencing of pooled DNA from the F2 segregating population revealed that the mutation in s48 occurred within a known OsPPKL1 gene. However, unlike the previously reported D364E mutation, the mutation in s48 is D364N. While N (asparagine, a polar amino acid) and E (glutamic acid, an acidic amino acid) are two distinct amino acids, mutations at D364 to either N or E can increase grain size. The amino acid sequence of OsPPKL1 is shown in SEQ ID NO. 1.
[0058] Example 2 Artificial editing of the D364 locus to alter rice grain size
[0059] The CRISPR gene editing vector was used to mutate the genomic position at the D364 site, and the grain length and weight of the edited plants were identified.
[0060] The cytosine base editor APO1 was used to perform genome site-directed editing on the site corresponding to D364. APO1 was connected to the codon-optimized hSpCas9 through the signal peptide XTEN to form the vector pCSGAPO1. The targeting sequence 5'-TAGACACAGCTGCTGGAGTC-3' was synthesized and annealed to form an oligonucleotide adapter. The oligo linker and pCSGAPO1 were used for ligation reaction, which was then digested with BsaI and ligated with T4 to produce a CRISPR / hSpCas9 plasmid vector. The T-DNA region of the vector contains a gRNA expression cassette (for targeting genome editing sites), an APO1-XTEN-hSpCas9 expression cassette (for cytosine base editing), and a HYG expression cassette (for screening markers during transformation). Schematic diagrams of each element are attached. Figure 3 .
[0061] At the same time, CRISPR / Cas9 vector was used to perform base deletion editing at the corresponding site of D364. The CRISPR / Cas9 vector used a commonly used editing vector, and the target sequence was 5'-TAGACACAGCTGCTGGAGTC-3'. The vector diagram is shown in the attached figure. Figure 4 shown.
[0062] The two editing vectors were transformed into rice Kongyu 131 (KY131, a common rice variety bred by Heilongjiang Academy of Agricultural Sciences) to obtain transformed plants. The grain phenotypes of the edited plants were investigated, and the editing sites with phenotypic changes were detected.
[0063] In plants with altered grain phenotypes, in addition to the D364E and T365R variants, several new variants were found, namely D364E-T365R, del.5, del.6, and del.14. The amino acid sequences of the target sites (360-377) corresponding to these variants, as well as the grain length and 1000-grain weight, are shown in the attached figure. Figure 5 The specific data are shown in Table 2. The grain length and 1000-grain weight of these variants were significantly increased compared with the wild type.
[0064] Table 2 Grain phenotypic traits corresponding to different OsPPKL1 mutation types
[0065]
[0066] Bold indicates amino acid substitution, "-" indicates amino acid deletion, and "*" indicates significant difference compared to the wild-type control (α=0.05).
[0067] The above results indicate that mutations at sites near D364 can cause changes in the phenotype of plant seeds while ensuring the normal growth and development of the plants (see Appendix). Figure 6 The mutation must be a non-frameshift mutation (i.e., a deletion or mutation of the amino acid near the D364 site, but without affecting the amino acids at other positions). If a frameshift mutation occurs, rice cannot grow and develop normally, which also shows the important role of the PPKL1 protein in rice growth and development.
[0068] Example 3 Overexpression of mutant protein increases rice grain length and weight
[0069] The mutant proteins P1-N (PPKL1 full-length protein 364 position is N) and P1-D (PPKL1 full-length protein 364 position is D) and P1 N -N (a truncated protein of the N-terminal 1-600 amino acids of PPKL1 with N at position 364) and P1 N -D (a truncated protein containing 1-600 amino acids at the N-terminus of PPKL1 with a D at position 364), nucleic acid molecules encoding these mutant proteins were synthesized, and overexpression vectors were constructed (the commonly used CaMV 35S was used as a promoter and OCS was used as a terminator in the expression cassette of the mutant protein nucleic acid molecule), and the rice Zhonghua 11 was transformed. The grain length and grain weight of the transgenic plants were analyzed. The results showed that the transgenic rice containing the N364 site showed longer grain length or higher 1000-grain weight than the transgenic rice containing the D364 site and the recipient Zhonghua 11 with the D364 site (data are shown in Table 3, and the grain phenotypes are shown in the Appendix). Figure 7 ).
[0070] Table 3 Grain phenotypic traits corresponding to different variant types of OsPPKL1
[0071]
[0072] ZH11 indicates ZH11 in the receptor, OE indicates overexpression, and "*" indicates a significant increase compared to the receptor control (α=0.05).
[0073] This example shows that the effect of increasing seed size and weight can also be achieved by exogenously introducing the N364 mutant protein encoding gene.
[0074] Example 4 Artificial editing of a wider range near the D364 site of homologous proteins in Arabidopsis, rice, and maize
[0075] To further identify the target region that can cause grain changes, PPKL1 homologous proteins from the model species of dicotyledonous and monocotyledonous plants, Arabidopsis thaliana, rice, and maize were selected. Using the single-base editor AOP1, targets were designed for single-base editing of the regions on both sides of the PPKL1-D364 site, specifically the region at positions 84-420 of SEQ ID NO:1, to create more non-frameshift mutants.
[0076] The selected PPKL1 homologous proteins include Arabidopsis BSL1 (At4g03080), BSL2 (At1g08420), BSL3 (At2g27210), and BSU1 (At1g03445); rice PPKL1 (Os03g44500), PPKL2 (Os05g05240), and PPKL3 (Os12g42310); and maize BSL1 (Zm00001d010887) and BSL3 (Zm00001d041307, Zm00001d031088). The LOC numbers of the genes are in parentheses. The rice gene and protein sequences are from the RGAP (http: / / rice.plantbiology.msu.edu / ) database, the Arabidopsis gene and protein sequences are from the TAIR (https: / / www.arabidopsis.org / ) database, and the maize gene and protein sequences are from the maizeGDB (http: / / www.maizegdb.org / ) database. The amino acid alignment results are shown in the Appendix. Figure 8 .
[0077] Table 4 PPKL1 homologous gene editing targets and sgRNA sequences
[0078]
[0079]
[0080] The construction process of the editing vector is as described in Example 2. The sequence U expressing the sgRNA on the vector is changed to T. The expressed sgRNA molecule also includes the sgRNA Scaffold sequence, for example:
[0081] guuuuagagcuagaaauagcaaguuaaaauaaggcuaguccguuaucaacuugaaaaaguggcaccgagucggugcuuuuuuu.
[0082] The constructed editing vector was used to transform rice Kongyu 131 (KY131), Arabidopsis Col-0, and maize KN5585 to obtain transformed plants. The grain phenotypes of the edited plants were investigated, and the editing sites with phenotypic changes were detected and the mutation types were analyzed.
[0083] Example 5: Finding more similar sites of homologous proteins
[0084] Using the amino acid sequence shown at positions 84-420 of SEQ ID NO:1, the NCBI database was searched for proteins with similar sequences, with a similarity of at least 91%. Proteins of this type were found to include proteins numbered (gene LOC number) Gm10g02760, Gm02g17040, Gm13g38850, Gm11g18090, and Gm12g31540 in soybean, proteins numbered KAF6982594.1 and AVK60242.2 in wheat, and BSL proteins from various other plants. The amino acid sequences corresponding to positions 84-420 of SEQ ID NO:1 in these proteins are also considered effective mutation sites. By creating non-frameshift mutations at these positions, increased seed size or weight can also be achieved.
[0085] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made based on the present invention. Therefore, such modifications and improvements, which do not depart from the spirit of the present invention, are intended to be within the scope of protection claimed herein. Sequence Listing <110> Weimi Biotechnology (Jiangsu) Co., Ltd., Institute of Crop Sciences, Chinese Academy of Agricultural Sciences <120> Method for increasing plant seed size or increasing plant seed weight <130> 1 <160> 1 <170> SIPOSequenceListing 1.0 <210> 1 <211> 1003 <212> PRT <213> Oryza sativa <400> 1 Met Asp Val Asp Ser Arg Met Thr Thr Glu Ser Asp Ser Asp Ser Asp 1 5 10 15 Ala Ala Ala Thr Ala Ala Ala Ser Ala Ser Val Ala Ala Gln Gly Gly 20 25 30 Leu Ala Ser Glu Thr Ser Ser Ser Ser Ser Ala Ser Ala Pro Ser Thr 35 40 45 Pro Gly Thr Pro Thr Val Ala Pro Ala Pro Ala Ala Ala Gly Ala Thr 50 55 60 Gly Pro Arg Pro Ala Pro Gly Tyr Thr Ala Val Ser Ala Val Ile Glu 65 70 75 80 Lys Lys Glu Asp Gly Pro Gly Cys Arg Cys Gly His Thr Leu Thr Ala 85 90 95 Val Pro Ala Val Gly Glu Glu Gly Thr Pro Gly Tyr Ile Gly Pro Arg 100 105 110 Leu Ile Leu Phe Gly Gly Ala Thr Ala Leu Glu Gly Asn Ser Ala Thr 115 120 125 Pro Pro Ser Ser Ala Gly Ser Ala Gly Ile Arg Leu Ala Gly Ala Thr 130 135 140 Ala Asp Val His Cys Tyr Asp Val Leu Ser Asn Lys Trp Ser Arg Leu 145 150 155 160 Thr Pro Gln Gly Glu Pro Pro Ser Pro Arg Ala Ala His Val Ala Thr 165 170 175 Ala Val Gly Thr Met Val Val Ile Gln Gly Gly Ile Gly Pro Ala Gly 180 185 190 Leu Ser Ala Glu Asp Leu His Val Leu Asp Leu Thr Gln Gln Arg Pro 195 200 205 Arg Trp His Arg Val Val Val Gln Gly Pro Gly Pro Gly Pro Arg Tyr 210 215 220 Gly His Val Met Ala Leu Val Gly Gln Arg Phe Leu Leu Thr Ile Gly 225 230 235 240 Gly Asn Asp Gly Lys Arg Pro Leu Ala Asp Val Trp Ala Leu Asp Thr 245 250 255 Ala Ala Lys Pro Tyr Glu Trp Arg Lys Leu Glu Pro Glu Gly Glu Gly 260 265 270 Pro Pro Pro Cys Met Tyr Ala Thr Ala Ser Ala Arg Ser Asp Gly Leu 275 280 285 Leu Leu Leu Cys Gly Gly Arg Asp Ala Asn Ser Val Pro Leu Ala Ser 290 295 300 Ala Tyr Gly Leu Ala Lys His Arg Asp Gly Arg Trp Glu Trp Ala Ile 305 310 315 320 Ala Pro Gly Val Ser Pro Ser Pro Arg Tyr Gln His Ala Ala Val Phe 325 330 335 Val Asn Ala Arg Leu His Val Ser Gly Gly Ala Leu Gly Gly Gly Arg 340 345 350 Met Val Glu Asp Ser Ser Ser Val Ala Val Leu Asp Thr Ala Ala Gly 355 360 365 Val Trp Cys Asp Thr Lys Ser Val Val Thr Thr Pro Arg Ile Gly Arg 370 375 380 Tyr Ser Ala Asp Ala Ala Gly Gly Asp Ala Ala Val Glu Leu Thr Arg 385 390 395 400 Arg Cys Arg His Ala Ala Ala Ala Val Gly Asp Gln Ile Phe Ile Tyr 405 410 415 Gly Gly Leu Arg Gly Gly Val Leu Leu Asp Asp Leu Leu Val Ala Glu 420 425 430 Asp Leu Ala Ala Ala Glu Thr Thr Thr Ala Ala Asn His Ala Ala Ala 435 440 445 Ser Ala Ala Ala Thr Asn Val Gln Ser Gly Arg Thr Pro Gly Arg Tyr 450 455 460 Ala Tyr Asn Asp Glu Arg Ala Arg Gln Thr Ala Pro Glu Ser Ala Gln 465 470 475 480 Asp Gly Ser Val Val Leu Gly Thr Pro Val Ala Pro Pro Val Asn Gly 485 490 495 Asp Met Tyr Thr Asp Ile Ser Pro Glu Asn Ala Val Leu Gln Gly Gln 500 505 510 Arg Arg Leu Ser Lys Gly Val Asp Tyr Leu Val Glu Ala Ser Ala Ala 515 520 525 Glu Ala Glu Ala Ile Ser Ala Thr Leu Ala Ala Val Lys Ala Arg Gln 530 535 540 Val Asn Gly Glu Met Glu Gln Leu Pro Asp Lys Glu Gln Ser Pro Asp 545 550 555 560 Ser Ala Ser Thr Ser Lys His Ser Ser Leu Ile Lys Pro Asp Ser Ile 565 570 575 Leu Ser Asn Asn Met Thr Pro Pro Pro Gly Val Arg Leu His His Arg 580 585 590 Ala Val Val Val Ala Ala Glu Thr Gly Gly Ala Leu Gly Gly Met Val 595 600 605 Arg Gln Leu Ser Ile Asp Gln Phe Glu Asn Glu Gly Arg Arg Val Ser 610 615 620 Tyr Gly Thr Pro Glu Asn Ala Thr Ala Ala Arg Lys Leu Leu Asp Arg 625 630 635 640 Gln Met Ser Ile Asn Ser Val Pro Lys Lys Val Ile Ala Ser Leu Leu 645 650 655 Lys Pro Arg Gly Trp Lys Pro Pro Val Arg Arg Gln Phe Phe Leu Asp 660 665 670 Cys Asn Glu Ile Ala Asp Leu Cys Asp Ser Ala Glu Arg Ile Phe Ser 675 680 685 Ser Glu Pro Ser Val Leu Gln Leu Lys Ala Pro Val Lys Ile Phe Gly 690 695 700 Asp Leu His Gly Gln Phe Gly Asp Leu Met Arg Leu Phe Asp Glu Tyr 705 710 715 720 Gly Ala Pro Ser Thr Ala Gly Asp Ile Ala Tyr Ile Asp Tyr Leu Phe 725 730 735 Leu Gly Asp Tyr Val Asp Arg Gly Gln His Ser Leu Glu Thr Met Thr 740 745 750 Leu Leu Leu Ala Leu Lys Val Glu Tyr Pro Gln Asn Val His Leu Ile 755 760 765 Arg Gly Asn His Glu Ala Ala Asp Ile Asn Ala Leu Phe Gly Phe Arg 770 775 780 Ile Glu Cys Ile Glu Arg Met Gly Glu Arg Asp Gly Ile Trp Thr Trp 785 790 795 800 His Arg Met Asn Arg Leu Phe Asn Trp Leu Pro Leu Ala Ala Leu Ile 805 810 815 Glu Lys Lys Ile Ile Cys Met His Gly Gly Ile Gly Arg Ser Ile Asn 820 825 830 His Val Glu Gln Ile Glu Asn Leu Gln Arg Pro Ile Thr Met Glu Ala 835 840 845 Gly Ser Val Val Leu Met Asp Leu Leu Trp Ser Asp Pro Thr Glu Asn 850 855 860 Asp Ser Val Glu Gly Leu Arg Pro Asn Ala Arg Gly Pro Gly Leu Val 865 870 875 880 Thr Phe Gly Pro Asp Arg Val Met Glu Phe Cys Asn Asn Asn Asp Leu 885 890 895 Gln Leu Ile Val Arg Ala His Glu Cys Val Met Asp Gly Phe Glu Arg 900 905 910 Phe Ala Gln Gly His Leu Ile Thr Leu Phe Ser Ala Thr Asn Tyr Cys 915 920 925 Gly Thr Ala Asn Asn Ala Gly Ala Ile Leu Val Leu Gly Arg Asp Leu 930 935 940 Val Val Val Pro Lys Leu Ile His Pro Leu Pro Pro Ala Ile Thr Ser 945 950 955 960 Pro Glu Thr Ser Pro Glu His His Ile Glu Asp Thr Trp Met Gln Glu 965 970 975 Leo Asn Ala Asn Arg Pro Pro Thr Pro Thr Arg Gly Arg Pro Gln Val 980 985 990 Ala Ala Asn Asp Arg Gly Ser Leu Ala Trp Ile 995 1000
Claims
1. A protein, characterized in that The protein is a protein obtained by subjecting the protein represented by SEQ ID NO. 1 to any of the following mutations: 1) T365R; 2) D364N; 3) D364E and T365R; 4) T365R and 366A-370W are missing; 5) 364D~369V are missing; 6) 362V~375S are missing.
2. A nucleic acid, characterized in that The nucleic acid encodes the protein according to claim 1.
3. A method for increasing the size or weight of rice seeds, characterized in that: The protein of claim 1 is expressed in the rice, and plants with enlarged seeds or increased grain weight are selected.
4. The method according to claim 3, wherein: The method comprises site-directed mutation of rice endogenous protein or introduction of exogenous protein.
5. The method according to claim 4, characterized in that: The method for site-directed mutation of rice endogenous protein is achieved by a gene editing method with 5'-TAGACACAGCTGCTGGAGTC-3' as the target.