Method for increasing plant yield
By introducing stress-responsive cis-elements, especially heat-responsive elements, into plant gene promoters and optimizing carbon assimilate allocation, the problem of crop yield reduction under high-temperature stress was solved, and high and stable yields were achieved under high-temperature conditions.
Patent Information
- Application Number
- PCT/CN2025/082642
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-30
- Filing Date
- 2025-03-14
- Publication Date
- 2025-11-20
AI Technical Summary
Existing technologies are insufficient to effectively optimize the allocation of carbon assimilates in plants under high-temperature stress, leading to reduced crop yields. Furthermore, traditional gene overexpression methods may disrupt source-sink balance, thereby reducing yields.
By introducing stress-responsive cis-elements, especially heat-responsive elements, into the promoters of yield-related genes in plants using gene editing technology, the expression of genes such as cell wall invertases can be precisely regulated, optimizing the allocation of carbon assimilates to sink organs.
Under high temperature stress, we can increase plant yield and fruit set rate, enhance the crop's ability to maintain stable yield under heat stress, and achieve high and stable yield in high temperature environments.
Smart Images

Figure PCTCN2025082642-FTAPPB-I100001 
Figure PCTCN2025082642-FTAPPB-I100002 
Figure PCTCN2025082642-FTAPPB-I100003
Abstract
Description
Methods of increasing plant yield TECHNICAL FIELD
[0001] The present invention belongs to the field of biotechnology, particularly agricultural biotechnology. More specifically, the present invention relates to methods of increasing plant yield by precisely introducing environmentally responsive elements to regulate carbon assimilate partitioning through gene editing, such as prime editing.
[0002] BACKGROUND
[0003] By 2050, global crop production needs to double to meet the expected demand due to population growth, changes in dietary structure, and increased consumption of biofuels (Long and Ort, 2010). However, current crop production rates are insufficient to ensure global food supply. This problem has become increasingly urgent due to the frequent occurrence of abiotic stress related to global climate change, leading to significant crop yield reduction (Wheeler and von Braun, 2013; Gao, 2021). The main threat to crop yield from these climate change-related stresses is heat stress (Balfagon et al., 2020). Previous studies have reported that for every 1 °C increase in temperature during the growing season, major crops in different regions will experience a 2.5% to 16% increase in yield loss (Peng et al., 2004; Battisti and Naylor, 2009). Therefore, how to quickly develop "environmentally intelligent" crops that can achieve high yield under normal conditions and stable yield under heat stress is an urgent need for global food security (Marsh et al., 2021; Naqvi et al., 2022).
[0004] Over the past few decades, much has been learned about the effects of external factors such as environmental conditions and diseases on crop yield. A large number of studies and related genes have been used for molecular breeding. However, research and attention to endogenous factors such as internal nutrient partitioning in plants have lagged far behind (Fernie et al., 2020). Most research has focused on optimizing internal nutrient partitioning through cultivation management, while ignoring the great potential of rationally designing and molecularly manipulating internal carbon partitioning mechanisms to improve crop yield. One of the main reasons is that the internal nutrient partitioning mechanism of plants is very complex, and it produces even more complex feedback mechanisms when responding to environmental changes.
[0005] The source-sink concept, first proposed by Mason and Maskell in 1928, has a history of more than 100 years and is a fundamental concept to explain how nutrients are partitioned within a plant body (Mason and Maskell, 1928). Source tissues are net producers of photosynthate (mainly carbohydrates such as sucrose), while sink tissues are net importers of photosynthate for utilization or storage. The major carbon assimilate produced by photosynthesis is sucrose, which accounts for 90% of plant biomass and is a key factor determining yield (Ruan et al., 2010; Aluko et al., 2021). Sucrose is transported from source tissues (mainly mature leaves) to various sink tissues through phloem and must be degraded to hexose or its derivatives (mainly glucose and fructose) in order to provide energy and support growth of sink organs such as roots, developing flowers, fruits, seeds, cotton fibers, and storage organs such as tubers or bulbs (Sturm and Tang, 1999; Ruan, 2014). Sucrose is degraded to glucose and fructose by invertases or to uridine diphosphate glucose and fructose by sucrose synthase (SuSy). Invertases are encoded by two gene families: acid invertases, including cell wall invertases (CWINs) and vacuolar invertases (VINs); and neutral / alkaline cytoplasmic invertases (CINs), which exist in the cytoplasm (Wang et al., 2010; Li et al., 2012). Among them, CWINs play an indispensable role in providing nutrient, energy source, and signaling molecules for growth, yield formation, and stress response in different species of plants, including tomato (Jin et al., 2009; Liu et al., 2016), rice (Wang et al., 2008), maize (Cheng et al., 1996), Vicia faba (Weber et al., 1996), cotton and Arabidopsis (Wang and Ruan, 2012), barley (Weschke et al., 2003), and cassava (Yan et al., 2019). They have co-evolved with vascular plants, and this gene family has expanded from gymnosperms to angiosperms in seed plants (Wan et al., 2018). This expansion implies an evolutionary link between CWINs and seed or fruit formation. Importantly, the coding or regulatory regions of CWIN genes have been selected during domestication of major crops such as tomato (Tieman et al., 2017; Gao et al., 2019) and rice (Wang et al., 2008). The tomato CWIN gene LIN5 was mapped to a major quantitative trait locus that determines fruit sugar content and total yield (Fridman et al., 2004).Solanum pennellii LIN5 allele (a tomato variety with single nucleotide polymorphism (SNP) near the catalytic site, which has higher sugar content in the fruit, and knockdown of the LIN5 gene leads to poor development of seeds and fruits, and a high frequency of fruit abortion (Zanor et al., 2009). In maize, loss of function of the homologous gene MINIATURE 1 (Mnl) of LIN5 leads to a typical seed development failure phenotype, resulting in a reduction of grain yield by about 70% (Cheng et al., 1996). In rice, the orthologous gene of CWIN encoding LIN5 and Mn1, GRAIN INCOMPLETE FILLING 1 (GIF1), controls the distribution of photosynthate during early grain filling and determines the final grain yield, and its regulatory region is likely to be selected during domestication (Wang et al., 2008). It has been reported that CWINs can regulate pollen development, pollen tube elongation, fertilization, ovule development, nectar production, leaf development and senescence, and fruit ripening in addition to playing an important role in seed / fruit yield in different crops (Fridman et al., 2004; Zanor et al., 2009; Vallarino et al., 2017).
[0006] The physiological basis of crop yield reduction and quality decline under high temperature is the disruption of the balance between carbon sink and source, leading to insufficient energy supply to sink organs, and thus failure of reproductive development and yield formation (Pressman et al., 2002; Rizhsky et al., 2004; Suwa et al., 2010). Heat stress rapidly inhibits carbon allocation to sink organs, leading to selective shriveling of grains or ovaries, which is a common problem and the main cause of yield loss for grain or fruit crops under heat stress (Li et al., 2012). This “strategic abandonment” is an evolutionary strategy of plants to adapt to challenging environments when nutrients are scarce in natural ecosystems, thus gaining a competitive advantage in survival competition (Rodrigues et al., 2019). Unfortunately, this sensitive environmental adaptability has been retained throughout the long history of crop domestication, and under the background of global climate change, this adaptability has become undesirable in agronomy, as it often causes irreversible yield loss (Shen et al., 2023). Agricultural ecosystems are usually different from natural ecosystems in terms of nutrient and water supply, plant spacing, shading, and pests, pathogens, and weeds. In the past few decades, with the rapid development of industrialization, especially the production and application of chemical fertilizers, mechanization, and agricultural chemicals, this difference has been further exacerbated. Therefore, the responsive inhibition of carbon allocation becomes unnecessary and seriously hinders crops from fully realizing their yield potential. Previous studies consistently showed that, compared with heat-sensitive lines, heat-tolerant tomato varieties exhibited higher CWIN activity in flowers and young fruits, and allocated more sucrose to fruits and less to vegetative organ tissues (Li et al., 2012; Liu et al., 2016).
[0007] While molecular breeding has successfully optimized crop morphology to tap into the yield potential in the agricultural ecosystem, there has been little progress in the "internal" optimization, particularly in blunting or eliminating the over-sensitive feedback inhibition mechanisms of carbon partitioning, as these mechanisms have never been directly selected for (Amthor, 2000). While attempts have been made to overexpress CWINs in various crops to improve carbon partitioning efficiency, the results often show that the disruption of source-sink balance caused by overexpression not only fails to improve efficiency, but also reduces yield (von Schaewen et al., 1990; Dickinson et al., 1991; Wang et al., 2008), which indicates the importance of fine-tuning the source-sink relationship. Notably, the reproductive development of many crops (such as tomato and cereal legumes) is more sensitive to high night temperature than to daytime temperature (Ismail and Hall, 1999). For example, when the night temperature is higher than 24°C, up to 80% of tomato flowers or fruits will abort, and the fruits produced will be smaller, seedless, and deformed (Ruan et al., 2010). This indicates that insufficient carbon assimilate supply to sink organs at night will exacerbate the impact of heat stress on yield. It is worth noting that the impact of global warming on day and night temperatures is not uniform. Compared with daytime warming, night warming is more prevalent globally (Cox et al., 2020). In order to address the threat of global warming to crop yield and the urgent need for rapid breeding, the inventors have developed a climate-responsive optimization of photosynthetic assimilate partitioning to sink organs (CROCS) strategy by rationally manipulating the expression of CWINs genes in tomato and cereal crops using prime editing technology.
[0008] SUMMARY
[0009] The present invention encompasses at least the following items:
[0010] Item 1. A method of producing a modified plant, the method comprising introducing one or more stress-responsive cis-elements into the expression regulatory sequence, such as the promoter, of one or more yield-related genes, preferably endogenous source-sink relationship-related genes, of the plant.
[0011] Item 2. The method of item 1, wherein the stress is a biotic stress or an abiotic stress, for example the abiotic stress is selected from heat stress, cold stress, osmotic stress (such as drought stress, salt stress, metal stress), water stress (flooding), light stress (insufficient or excessive light), nutrient stress (deficiency or excess of nutrient elements in the soil, such as nitrogen deficiency stress, phosphorus deficiency stress), mechanical stress (wind, hail, and other physical damage); the biotic stress is selected from fungal infection, bacterial infection, viral infection, parasitic infection, insect stress, weed competition, parasitic plant,
[0012] Preferably, the stress is heat stress.
[0013] Item 3. The method of item 1 or 2, wherein the stress-responsive cis-element comprises a cis-element selected from Table 1, preferably the stress-responsive cis-element is a heat-responsive element, more preferably the heat-responsive element comprises the sequence ATTCTAGAAT.
[0014] Item 4. The method of any one of items 1 to 3, wherein the source-sink relationship related gene is selected from the group consisting of a gene encoding a cell wall invertase (CWIN), a sucrose transporter (SUT), a hexose transporter (HT), a Sugars Will Eventually be Exported Transporter (SWEET), a sucrose-phosphate synthase (SPS), a sucrose synthase (SUS), a Glucan-water dikinase (GWD), a Source activity enhanced (SOE).
[0015] Item 5. The method of any one of items 1 to 3, wherein the source-sink relationship related gene is selected from the genes of Table 2, or the source-sink relationship related gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to a protein encoded by a gene selected from Table 2. Alternatively, the coding sequence of the source-sink relationship related gene has at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to the coding sequence of a gene selected from Table 2.
[0016] Item 6. The method of any one of items 1 to 5, wherein introduction of the stress-responsive cis-element results in the expression of the source-sink relationship related gene in the plant being up- or down-regulated, preferably up-regulated, in response to the respective stress.
[0017] Item 7. The method of any one of items 1 to 6, wherein the one or more stress-responsive cis-element is introduced into a predetermined site within the promoter, preferably the predetermined site
[0018] 1) is not located within an endogenous cis-element of the gene;
[0019] 2) is located within open chromatin; and / or
[0020] 3) close to the start codon of the gene.
[0021] Item 8. The method of item 7, wherein the selection of the pre-determined site comprises:
[0022] a) selecting a position within the promoter that does not contain possible cis-elements as a candidate site based on Plant CARE (Plant C ARE, a database of plant promoters and their cis-acting regulatory elements (ugent.be) cis-element online prediction website;
[0023] b) performing DNase I hypersensitivity prediction on the candidate sites obtained in step a) by DHS (DNase-I hypersensitive sites) online prediction website (http: / / www.epigenome.cuhk.edu.hk / ) to screen out the candidate sites with peak value less than 0.5; and
[0024] c) selecting a site closest to the start codon of the gene from the candidate sites obtained in step b) as the final pre-determined site.
[0025] Item 9. The method of item 7, wherein the pre-determined site is no more than about 2 kb, preferably no more than about 1 kb, more preferably no more than 500 bp away from the translation start codon of the gene.
[0026] Item 10. The method of any one of items 1-9, wherein a plurality (e.g., 2 to about 10 or more) of stress-responsive cis-elements are introduced into the promoter, e.g., the plurality of stress-responsive cis-elements are introduced into the promoter in tandem.
[0027] Item 11. The method of any one of items 1-10, wherein before introducing the stress-responsive cis-elements into the plant, the stress-responsive elements inserted into the pre-determined site are verified to be capable of conferring stress-responsive ability by expressing a reporter system in vitro, e.g., in a tobacco cell expressing reporter system.
[0028] Item 12. The method of any one of items 1-11, wherein the introduction of the one or more stress-responsive cis-elements can be achieved by genome editing, e.g., introducing a genome editing system into the plant.
[0029] Item 13. The method of item 12, wherein the genome editing system is a CRISPR, ZFN or TALEN based genome editing system; preferably, the genome editing system is a CRISPR based genome editing system.
[0030] Item 14. The method of item 12 or 13, wherein the method comprises
[0031] 1) introducing into the plant a genome editing system targeting the preselected site, which results in a double-strand break (DSB) in the genome at or near the preselected site;
[0032] 2) introducing a homologous recombination donor nucleic acid comprising the one or more stress-responsive cis-elements and homology arm sequences corresponding to the flanking sides of the DSB, whereby the one or more stress-responsive cis-elements are introduced into the preselected site by homologous recombination.
[0033] Item 15. The method of item 12 or 13, wherein the introduction of the one or more stress-responsive cis-elements is achieved by introducing into the plant a prime editing system.
[0034] Item 16. The method of item 15, wherein the prime editing system comprises:
[0035] i) a fusion protein comprising a CRISPR nickase and a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding the fusion protein; and
[0036] ii) a pegRNA and / or an expression construct containing a nucleotide sequence encoding the pegRNA,
[0037] wherein the pegRNA comprises, in the 5’ to 3’ direction, a guide sequence, a gRNA scaffold sequence, a reverse transcription (RT) template sequence, and a primer binding site (PBS) sequence, wherein the pegRNA is capable of forming a complex with the fusion protein and targeting the fusion protein to a first target sequence in the genome comprising or near the preselected site, resulting in a first nick within the first target sequence, wherein the reverse transcription template sequence comprises a nucleotide sequence of the one or more stress-responsive cis-elements or a complement thereof,
[0038] Preferably, the prime editing system further comprises iii) a nick gRNA and / or an expression construct containing a nucleotide sequence encoding the nick gRNA, the nick gRNA comprising a guide sequence and a gRNA scaffold sequence;
[0039] More preferably, the prime editing system further comprises iv) a Csy4 protein and / or an expression construct containing a nucleotide sequence encoding the Csy4 protein, and the 3’ end of the pegRNA and / or the 3’ and 5’ ends of the nick RNA comprise a Csy4 recognition site sequence,
[0040] More preferably, the pegRNA further comprises a tevopreQ1 motif at the 3' end.
[0041] Item 17. The method of item 16, wherein the CRISPR nickase is a Cas9 nickase, for example, the Cas9 nickase comprises the amino acid sequence set forth in SEQ ID NO: 12 or 26.
[0042] Item 18. The method of item 16 or 17, wherein the reverse transcriptase is an M-MLV reverse transcriptase or a functional variant thereof, for example, the reverse transcriptase comprises the amino acid sequence set forth in SEQ ID NO: 13.
[0043] Item 19. The method of any one of items 16-18, wherein the CRISPR nickase and the reverse transcriptase in the fusion protein are connected by a linker, for example, the linker can be the linker set forth in SEQ ID NO: 14 (33 aa linker).
[0044] Item 20. The method of any one of items 16-19, wherein the fusion protein further comprises one or more nuclear localization sequences (NLS), for example, the NLS is an SV40 NLS (the amino acid sequence is set forth in SEQ ID NO: 15); and / or
[0045] the fusion protein further comprises a LA polypeptide at the C-terminus, for example, a LA polypeptide comprising the amino acid sequence set forth in SEQ ID NO 27.
[0046] Item 21. The method of any one of items 16-20, wherein the guide sequence (also referred to as seed sequence or spacer sequence) in the pegRNA is arranged to have sufficient sequence identity (preferably 100% identity) to the first target sequence, so as to be capable of binding to the complementary strand of the first target sequence through base pairing, achieving sequence-specific targeting.
[0047] Item 22. The method of any one of items 16-21, wherein the scaffold sequence of the gRNA is set forth in SEQ ID NO: 16.
[0048] Item 23. The method of any one of items 16-22, wherein the primer binding sequence is arranged to be complementary to at least a portion of the first target sequence, preferably, the primer binding sequence is complementary to at least a portion of a 3' overhang single strand resulting from the nick in the sense strand of the first target sequence, in particular, the primer binding sequence is complementary to the nucleotide sequence at the 3' end of the 3' overhang single strand.
[0049] Item 24. The method of any one of items 16-23, wherein the RT template sequence is arranged to be complementary to at least a portion of the sequence downstream of the nick in the first target sequence, and comprises the nucleotide sequence of the one or more stress-responsive cis-elements or a complement thereof.
[0050] Item 25. The method of any one of items 16-24, the nicking gRNA does not comprise a reverse transcription (RT) template sequence and a primer binding site (PBS) sequence, and a guide sequence (also referred to as a seed sequence or spacer sequence) in the nicking gRNA is set to have sufficient sequence identity (preferably 100% identity) to a second target sequence in the genome, such that the fusion protein is capable of being targeted to the second target sequence and causing a second nick within the second target sequence, the second target sequence being on the opposite strand of the genomic DNA from the first target sequence.
[0051] Item 26. The method of item 25, the second nick is upstream or downstream of the first nick, and the first nick and the second nick are from about 1 to about 300 or more nucleotides apart.
[0052] Item 27. The method of any one of items 16-26, the Csy4 protein comprises the amino acid sequence set forth in SEQ ID NO: 17; and / or, the Csy4 recognition site comprises the nucleotide sequence set forth in SEQ ID NO: 18.
[0053] Item 28. The method of any one of items 16-27, the prime editing system comprises a first expression construct encoding a fusion protein comprising, from N-terminus to C-terminus: a Csy4 protein - a self-cleaving peptide - a NLS - a CRISPR nickase - a NLS - a linker - a reverse transcriptase - a NLS, or a Csy4 protein - a self-cleaving peptide - a NLS - a CRISPR nickase - a linker - a reverse transcriptase - a linker - a LA polypeptide - a NLS; and a second expression construct comprising: a Csy4 recognition site sequence - a pegRNA coding sequence - a Csy4 recognition site sequence - a nicking gRNA coding sequence - a Csy4 recognition site sequence, or a Csy4 recognition site sequence - a pegRNA coding sequence - a Csy4 recognition site sequence - a tevopreQl motif - a nicking gRNA coding sequence - a Csy4 recognition site sequence.
[0054] Item 29. The method of item 28, the first expression construct is driven by a 35S promoter for expression of the fusion protein, and the second expression construct is driven by a CmYLCV promoter for expression.
[0055] Item 30. The method of item 28 or 29, the first expression construct comprises the nucleotide sequence set forth in SEQ ID NO: 19 or 29.
[0056] Item 31. The method of any one of items 1-30, wherein the plant is a crop plant, for example selected from the group consisting of Solanum lycopersicum (tomato), Nicotiana benthamiana (tobacco), Capsicum annuum (pepper), Physalis pruinosa (ground cherry), Solanum melongena (eggplant), Solanum tuberosum (potato), Solanum pennellii (Pennell's tomato), Solanum chilense (Chilean tomato), Solanum habrochaites (Habrochaites tomato), Solanum pimpinellifolium (Pimpinellifolium tomato), Solanum galapagense (Galapagos tomato), Petunia hybrid (petunia), grape, strawberry, Citrullus lanatus (watermelon), Cucumis sativus (cucumber), Lactuca sativa L. (lettuce), Chinese cabbage, oilseed rape, Brassica oleracea, wheat, Oryza sativa L. (rice), maize, soybean, sunflower, sorghum, oilseed rape, alfalfa, cotton, barley, millet, sugarcane, and cassava.
[0057] Item 32. The method of any one of items 1-31, wherein the source-sink relationship associated gene is a cell wall invertase (CWIN) encoding gene, preferably the cell wall invertase (CWIN) encoding gene is a tomato LIN5 gene or a homologous gene thereof, for example a rice GIF1 gene, a maize MN1 gene, a soybean Glyma.10G074800 gene, or a wheat TraesCS2A03G0736600 gene;
[0058] or the gene is a tomato SRG1, FZY6, SOE, or a rice GRG1 gene.
[0059] Item 33. The method of item 32, wherein
[0060] i) the LIN5 gene encodes a LIN5 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 1; or the LIN5 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 2;
[0061] ii) the GIF1 gene encodes a GIF1 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 3; or the GIF1 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 4;
[0062] iii) the MN1 gene encodes a MN1 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 5; or the MN1 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 6; or
[0063] iv) the TraesCS2A03G0736600 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 7; or the TraesCS2A03G0736600 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 8; or
[0064] v) the soybean Glyma.10G074800 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 24; or the soybean Glyma.10G074800 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 25; or
[0065] vi) the SRG1 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the amino acid sequence set forth in SEQ ID NO: 42; or
[0066] vii) the FZY6 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the amino acid sequence set forth in SEQ ID NO: 43; or
[0067] viii) the SOE gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the amino acid sequence set forth in SEQ ID NO: 44; or
[0068] ix) the GRG1 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the amino acid sequence set forth in SEQ ID NO: 45.
[0069] Item 34. The method of item 32 or 33, wherein
[0070] the promoter of the tomato LIN5 gene comprises a nucleotide sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the nucleotide sequence set forth in SEQ ID NO: 9; or
[0071] the promoter of the rice GIF1 gene comprises a nucleotide sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the nucleotide sequence set forth in SEQ ID NO: 10.
[0072] Item 35. The method of any one of items 1-34, wherein
[0073] the plant is a tomato,
[0074] the stress is heat stress,
[0075] the stress-responsive cis-element is a heat-responsive element, for example the heat-responsive element comprises the sequence ATTCTAGAAT, and
[0076] The source-sink related gene is the tomato LIN5 gene.
[0077] Item 36. The method of item 35, wherein the heat responsive element is introduced into the tomato LIN5 gene promoter at 410 bp upstream of the start codon.
[0078] Item 37. The method of item 36, wherein introduction of the stress responsive cis-element results in the modified plant comprising a mutated LIN5 gene promoter set forth in SEQ ID NO: 20.
[0079] Item 38. The method of any one of items 1-34, wherein
[0080] The plant is rice,
[0081] The stress is heat stress,
[0082] The stress responsive cis-element is a heat responsive element, for example the heat responsive element comprises the sequence ATTCTAGAAT, and
[0083] The source-sink related gene is the rice GIF1 gene.
[0084] Item 39. The method of item 38, wherein the heat responsive element is introduced into the rice GIF1 gene promoter at 427 bp upstream of the start codon.
[0085] Item 40. The method of item 39, wherein introduction of the stress responsive cis-element results in the modified plant comprising a mutated GIF1 gene promoter set forth in SEQ ID NO: 21.
[0086] Item 41. A modified plant obtained according to the method of any one of items 1-40.
[0087] Item 42. The modified plant of item 41, which has increased yield, for example increased fruit / seed set, increased fruit / seed weight, in the presence and / or absence of the stress, as compared to an unmodified wild type plant.
[0088] Item 43. The modified plant of item 41, which has increased plot yield, in the presence and / or absence of the stress, as compared to an unmodified wild type plant.
[0089] Item 44. The modified plant of item 42 or 43, wherein the increased yield is about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100% or more.
[0090] BRIEF DESCRIPTION OF DRAWINGS
[0091] Figure 1. Insertion of HSE into the LIN5 promoter region confers heat- inducible expression ability.
[0092] Figure 2. Precise insertion of HSE into the endogenous LIN5 promoter region in tomato using improved PE.
[0093] Figure 3. AC cultivars modified by CROCS exhibit high yield at normal temperature and stable yield at high temperature phenotype.
[0094] Figure 4. Rice modified by CROCS exhibits high yield at normal temperature and stable yield at high temperature phenotype.
[0095] Figure 5. qPCR detection of GIF1 expression in different modern rice cultivars.
[0096] Figure 6. qPCR detection of GIF1 expression in wud-4-gif1-de with precise insertion of HSE.
[0097] Figure 7. Heat stress-induced expression of LIN5 after precise insertion of HSE into the LIN5 promoter region in M82 and YW1 backgrounds.
[0098] Figure 8. Verification of the ability of environmental response element insertion into different species LIN5 promoters to confer specific environmental response in tobacco.
[0099] Figure 9. Verification of the ability of different environmental response element insertion to confer specific environmental response to different tomato target gene promoters in tobacco.
[0100] Figure 10. Precise knock-in of different environmental response elements into different genes in tomato and rice, resulting in phenotypic changes in response to the environment.
[0101] Figure 11. Promoter cis-acting elements have a hybrid dominant effect.
[0102] DETAILED DESCRIPTION
[0103] I. DEFINITIONS
[0104] In the present application, the scientific and technical terms used herein have the meanings commonly understood by one of ordinary skill in the art, unless otherwise indicated. Also, the terms and techniques employed herein relating to protein and nucleic acid chemistry, molecular biology, cell and tissue culture, microbiology, immunology, and immunological procedures and laboratory procedures are those well known and commonly used in the corresponding art. For example, the standard recombinant DNA and molecular cloning techniques used in the present application are well known and are described more fully in Sambrook, J., Fritsch, E. F. and Maniatis, T., Molecular Cloning: A Laboratory Manual; Cold Spring Harbor Laboratory Press: Cold Spring Harbor, 1989. Also, to better understand the present application, the following definitions and explanations of the relevant terms are provided.
[0105] As used herein, the term "and / or" encompasses all combinations of the items linked by the term. For example, "A and / or B" covers "A", "B", and "A and B". For example, "A, B, and / or C" covers "A", "B", "C", "A and B", "A and C", "B and C", and "A and B and C".
[0106] The word "comprising" is used herein to mean that a protein or nucleic acid can consist of the sequence recited, or can have additional amino acids or nucleotides at either or both ends of the protein or nucleic acid, but still have the activity described in the present application. Furthermore, it is clear to one skilled in the art that the methionine encoded by the start codon at the N-terminus of a polypeptide is in some practical cases (e.g. when expressed in a particular expression system) retained, but does not materially affect the function of the polypeptide. Therefore, when describing a specific polypeptide amino acid sequence in the specification and claims of the present application, although it can not comprise the methionine encoded by the start codon at the N-terminus, it is nevertheless intended to encompass sequences comprising the methionine, and correspondingly, its encoding nucleotide sequence can comprise the start codon; vice versa.
[0107] "Polynucleotide," "nucleic acid sequence," "nucleotide sequence," or "nucleic acid fragment" are used interchangeably and are single- or double-stranded RNA or DNA polymers, optionally containing synthetic, non-natural or altered nucleotide bases. Nucleotides are referred to by their single letter designation: "A" is either adenosine or deoxyadenosine (corresponding to RNA or DNA, respectively), "C" denotes cytidine or deoxycytidine, "G" denotes guanosine or deoxyguanosine, "U" denotes uridine, "T" denotes deoxythymidine, "R" denotes purine (A or G), "Y" denotes pyrimidine (C or T), "K" denotes G or T, "H" denotes A or C or T, "I" denotes inosine, and "N" denotes any nucleotide. Although nucleotide sequences herein can be represented in DNA sequence (containing T), the corresponding RNA sequence (i.e., with U in place of T) can be readily determined by one of skill in the art when RNA is referred to.
[0108] "Polypeptide," "peptide," and "protein" are used interchangeably herein to refer to polymers of amino acid residues. The term applies to amino acid polymers in which one or more amino acid residue is an artificial chemical analogue of a corresponding naturally occurring amino acid, as well as to naturally occurring amino acid polymers. The terms "polypeptide," "peptide," "amino acid sequence," and "protein" can also include modified forms, including but not limited to glycosylation, lipid attachment, sulfation, gamma-carboxylation of glutamic acid residues, hydroxylation, and ADP-ribosylation.
[0109] "Sequence identity" has the art-recognized meaning and can be calculated as the percentage of sequence identity between two nucleic acid or polypeptide molecules or regions using published techniques. Sequence identity can be measured along the full length of a polynucleotide or polypeptide or along a region of the molecule. (See, e.g., Computational Molecular Biology, Lesk, A. M., ed., Oxford University Press, New York, 1988; Biocomputing: Informatics and Genome Projects, Smith, D. W., ed., Academic Press, New York, 1993; Computer Analysis of Sequence Data, Part I, Griffin, A. M., and Griffin, H. G., eds., Humana Press, New Jersey, 1994; Sequence Analysis in Molecular Biology, von Heinje, G., Academic Press, 1987; and Sequence Analysis Primer, Gribskov, M. and Devereux, J., eds., M Stockton Press, New York, 1991). While there are a number of methods for measuring sequence identity between two polynucleotides or polypeptides, the term "identity" is well known to one of skill in the art (Carrillo, H. & Lipman, D., SIAM J Applied Math 48: 1073 (1988)).
[0110] In peptides or proteins, suitable conservative amino acid substitutions are known to those of skill in the art and can generally be made without altering the biological activity of the resulting molecule. In general, those of skill in the art recognize that a single amino acid substitution in a non-essential region of a polypeptide will not substantially alter biological activity (see, e.g., Watson et al., Molecular Biology of the Gene, 4th Edition, 1987, The Benjamin / Cummings Pub. co., p. 224).
[0111] As used herein, "expression construct" refers to a vector, such as a recombinant vector, suitable for expression of a nucleotide sequence of interest in an organism. "Expression" refers to the production of a functional product. For example, expression of a nucleotide sequence can refer to transcription of the nucleotide sequence (e.g., to produce mRNA or functional RNA) and / or translation of the RNA into a precursor or mature protein.
[0112] An "expression construct" of the application can be a linear nucleic acid fragment, a circular plasmid, a viral vector, or, in some embodiments, can be a translatable RNA (such as an mRNA).
[0113] An "expression construct" of the application can comprise regulatory sequences of different origin and a nucleotide sequence of interest, or regulatory sequences and a nucleotide sequence of interest of the same origin but arranged in a manner different from that which normally exists in nature.
[0114] "Regulatory sequence" and "expression control sequence" are used interchangeably to refer to nucleotide sequences located upstream (5' non-coding sequences), within, or downstream (3' non-coding sequences) of a coding sequence, and which influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences can include, but are not limited to, promoters, translation leader sequences, introns, and polyadenylation recognition sequences.
[0115] A "promoter" refers to a nucleic acid segment that controls the transcription of another nucleic acid segment. In some embodiments of the application, a promoter is one that controls the transcription of a gene in a cell, whether or not it is derived from that cell. A promoter can be a constitutive promoter or a tissue-specific promoter or a developmentally-regulated promoter or an inducible promoter. A "constitutive promoter" refers to a promoter that will generally cause a gene to be expressed in a majority of cell types and a majority of the time. A "tissue-specific promoter" and "tissue-preferred promoter" are used interchangeably and refer to a promoter that is expressed, primarily but not necessarily exclusively, in one tissue or organ, and can also be expressed in a particular cell or cell type. A "developmentally-regulated promoter" refers to a promoter whose activity is determined by a developmental event. An "inducible promoter" selectively expresses an operably linked DNA sequence in response to an endogenous or exogenous stimulus (environmental, hormonal, chemical signal, etc.). Examples of promoters include, but are not limited to, a pol I, pol II, or pol III promoter. When used in plants, the promoter can be a cauliflower mosaic virus 35S promoter, a maize Ubi-1 promoter, a wheat U6 promoter, a rice U3 promoter, a maize U3 promoter, a rice actin promoter.
[0116] As used herein, the term "operably linked" refers to the linking of a regulatory sequence (such as, but not limited to, a promoter sequence, a transcription termination sequence, etc.) to a nucleic acid sequence (such as a coding sequence or open reading frame) such that the transcription of the nucleotide sequence is controlled and regulated by the transcriptional regulatory sequence. Techniques for operably linking regulatory sequence regions to nucleic acid molecules are known in the art.
[0117] "Introduction" of a nucleic acid molecule (e.g., a plasmid, a linear nucleic acid fragment, an RNA, etc.) or a protein into an organism means transformation of a cell of the organism with the nucleic acid or protein such that the nucleic acid or protein is able to function in the cell. As used herein, "transformation" includes stable transformation and transitory transformation. "Stable transformation" refers to the transfer of an exogenous nucleotide sequence into the genome of a cell resulting in genetically stable inheritance. Once stably transformed, the exogenous nucleic acid sequence is stably integrated into the genome of the organism and any successive generations thereof. "Transitory transformation" refers to the transfer of a nucleic acid molecule or protein into a cell, which performs a function without the exogenous gene being stably inherited. In transitory transformation, the exogenous nucleic acid sequence is not integrated into the genome.
[0118] As used herein, the term "plant" includes whole plants, and any descendant, cell, tissue, or part of a plant. The term "plant part" includes any part of a plant, including, for example and without limitation: seeds (including mature seeds, immature embryos without seed coats, and immature seeds); plant cuttings; plant cells; plant cell cultures; plant organs (e.g., pollen, embryos, flowers, fruits, shoots, leaves, roots, stems, and related explants). A plant tissue or plant organ can be a seed, callus, or any other group of plant cells that is organized into a structural or functional unit. A plant cell or tissue culture is capable of regenerating plants having the physiological and morphological characteristics of the plant from which it was derived and of regenerating plants having substantially the same genotype as the plant. In contrast, some plant cells are not capable of regenerating a whole plant. Regenerable cells in a plant cell or tissue culture can be embryos, protoplasts, meristematic cells, callus, pollen, leaves, anthers, roots, root tips, silk, flowers, kernels, ears, cobs, husks, or stalks.
[0119] Plant "progeny" includes any subsequent generations of a plant.
[0120] A "trait" refers to a physiological, morphological, biochemical, or physical characteristic of a cell or organism. An "agronomic trait" refers specifically to a measurable indicator of plant performance, including, but not limited to, leaf greenness, grain yield, growth rate, total biomass or rate of accumulation, fresh weight at maturity, dry weight at maturity, fruit yield, seed yield, total plant nitrogen content, fruit nitrogen content, seed nitrogen content, nitrogen content of vegetative plant tissue, total plant free amino acid content, fruit free amino acid content, seed free amino acid content, free amino acid content of vegetative plant tissue, total plant protein content, fruit protein content, seed protein content, protein content of vegetative plant tissue, herbicide resistance, drought resistance, nitrogen uptake, root lodging, harvest index, stalk lodging, plant height, ear height, ear length, disease resistance, cold tolerance, salt tolerance, and tiller number.
[0121] II. Methods of Producing Modified Plants
[0122] In one aspect, the present application provides a method of producing a modified plant, the method comprising introducing one or more stress-responsive cis-elements into the expression regulatory sequence, such as the promoter, of one or more yield-related genes, preferably endogenous source-sink relationship-related genes, of the plant. The genes are, for example, endogenous genes.
[0123] The plants described herein are preferably crop plants, including but not limited to Solanum lycopersicum (tomato), Nicotiana benthamiana (tobacco), Capsicum annuum (pepper), Physalis pruinosa (ground cherry), Solanum melongena (eggplant), Solanum tuberosum (potato), Solanum pennellii (Pennell's tomato), Solanum chilense (Chilean tomato), Solanum habrochaites (Habrochaites tomato), Solanum pimpinellifolium (Pimpinellifolium tomato), Solanum galapagense (Galapagos tomato), Petunia hybrid (petunia), grape, strawberry, Citrullus lanatus (watermelon), Cucumis sativus (cucumber), Lactuca sativa L. (lettuce), Chinese cabbage, rape, cabbage, wheat, Oryza sativa L. (rice), maize, soybean, sunflower, sorghum, rape, alfalfa, cotton, barley, millet, sugarcane, and cassava.
[0124] In some embodiments, the plant is a tomato. In some embodiments, the plant can be from different cultivars of Solanum lycopersicum (tomato). In some embodiments, the plant is a tomato cultivar Ailsa Craig, M82, Heinz 1706, Jingxian 8, VF-1, Beijing 1, TS545, TS181, TS590. Preferably, the plant is a tomato cultivar Ailsa Craig, M82, or Heinz 1706.
[0125] In some embodiments, the plant is rice. In some embodiments, the plant can be from different cultivars of Oryza sativa L. (rice). In some embodiments, the plant is a rice cultivar ZH11, Wu You rice-4, Longjing 31, or Zhongkefa 5.
[0126] In some embodiments, the introduction of the stress-responsive cis-element results in the expression of the yield-related gene, preferably the source-sink relationship-related gene, in the plant being up- or down-regulated, preferably up-regulated, in response to the corresponding stress. That is, when the plant comprising the introduced stress-responsive cis-element is subjected to the corresponding stress, the expression of the yield-related gene, preferably the source-sink relationship-related gene, is up- or down-regulated, preferably up-regulated.
[0127] As used herein, "stress" generally refers to an environmental or biological factor that adversely affects normal physiological activities of a plant. Stress conditions can result in stunted growth, abnormal development, or even death of a plant. For a crop plant, stress can particularly refer to conditions that result in a decrease in yield of the crop. However, stress can also simply refer to conditions that result in a yield of a plant, such as a crop, that does not reach a desired level, under which the plant or crop itself can not necessarily exhibit abnormal growth and development other than in yield. For example, stress can also simply refer to conditions that result in a yield of a plant, such as a crop, that does not reach a desired level, under which the plant or crop itself has been affected at the genetic level, but does not exhibit abnormal growth and development other than in yield.
[0128] As used herein, stress can be biotic or abiotic. Abiotic stress includes, but is not limited to, heat stress, cold stress, osmotic stress (e.g., drought stress, salt stress, metal stress), water stress (flooding), light stress (insufficient or excessive light, excessive or insufficient day length), nutrient stress (deficiency or excess of a nutrient element in the soil, e.g., nitrogen deficiency stress, phosphorus deficiency stress), mechanical stress (wind, hail, and other physical damage), and the like. Biotic stress includes fungal infection, bacterial infection, viral infection, parasitic infection, insect stress, weed competition, parasitic plants, and the like. In some preferred embodiments, the stress described herein is heat stress.
[0129] As used herein, "yield" refers to the amount (e.g., weight) of economically valuable parts of a harvested plant, such as a crop plant. Economically valuable parts include, but are not limited to, grains, seeds, fruits, fibers, tubers, and the like. Yield of a plant can be evaluated in different ways, such as generally by the weight of the economically valuable parts of the harvested crop per unit area of land. However, depending on the specific needs, yield of a plant can also be evaluated by other alternative or simplified methods, such as thousand kernel weight, fruit set / seeding rate, fruit / seed weight, yield per plant, plot yield, and the like.
[0130] The stress described herein will result in a decrease in yield of the plant. In some embodiments, the yield of an unmodified wild type plant is decreased in the presence of the stress, e.g., decreased by about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100%, as compared to the absence of the stress.
[0131] In some embodiments, the modified plant has increased yield, e.g., increased fruit / seed weight, increased fruit set, increased seed set, as compared to an unmodified wild type plant in the presence and / or absence of the stress. In some embodiments, the modified plant has increased yield per plant or per plot in the presence and / or absence of the stress, as compared to an unmodified wild type plant. In some embodiments, the yield is increased by about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100% or more.
[0132] In some embodiments, the modified plant has comparable fruit mass, e.g., comparable Brix content, as compared to an unmodified wild type plant in the presence and / or absence of the stress.
[0133] The stress responsive cis-elements described herein include, but are not limited to, a cis-element selected from Table 1 below:
[0134] Table 1. Stress responsive cis-elements
[0135] In some specific embodiments, the stress responsive cis-element is a heat responsive element. In some preferred embodiments, the heat responsive element comprises the sequence ATTCTAGAAT. In some preferred embodiments, the heat responsive element comprises the sequence GTTCATGAAC. In some preferred embodiments, the heat responsive element comprises the sequence AGAACGTTCT.
[0136] In some specific embodiments, the stress responsive cis-element is a light responsive element. In some preferred embodiments, the light responsive element comprises the sequence TGTGTGGTTAATATGAAGATAAGATT. In some preferred embodiments, the light responsive element comprises the sequence GTGTGTGAA.
[0137] In some specific embodiments, the stress responsive cis-element is a drought responsive element. In some preferred embodiments, the drought responsive element comprises the sequence CATGTG.
[0138] In some embodiments, the stress response cis-element is a flood response element. In some preferred embodiments, the flood response element comprises the sequence TTGACC.
[0139] The sequence of the cis-element of the application can also include the complement of the specifically shown sequence.
[0140] The yield-related genes, preferably source-sink relationship-related genes, described herein include, but are not limited to, genes encoding Cell wall invertase (CWIN), Sucrose transporter (SUT), Hexose transporter (HT), Sugars Will Eventually be Exported Transporter (SWEET), Sucrose-phosphate synthase (SPS), Sucrose synthase (SUS), Glucan-water dikinase (GWD), Source activity enhanced (SOE), and the like.
[0141] Suitable source-sink relationship-related genes in tomato include, but are not limited to, genes selected from Table 2 below:
[0142] Table 2, Source-sink relationship-related genes in tomato
[0143] The coding nucleotide sequence and protein sequence of the genes can be readily obtained by one of skill in the art from the Genomics Network (SGN) database (https: / / solgenomics.net / ), which are both incorporated herein by reference in their entirety, by the genomic accession number of the gene.
[0144] For other species, the source-sink relationship-related genes encompass homologous genes to the above-mentioned tomato genes. For example, the homologous genes encode proteins having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to the proteins encoded by the tomato genes. Alternatively, the coding sequences of the homologous genes have at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to the coding sequences of the tomato genes.
[0145] In some embodiments, the source-sink relationship related gene is a cell wall invertase (CWIN) encoding gene. In some embodiments, the cell wall invertase (CWIN) encoding gene is a tomato LIN5 gene or a homologous gene thereof, such as a rice GIF1 gene, a maize MN1 gene, a soybean Glyma.10G074800 gene, or a wheat TraesCS2A03G0736600 gene.
[0146] In some specific embodiments, the source-sink relationship related gene is a tomato LIN5 gene. In some embodiments, the LIN5 gene encodes a LIN5 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 1. In some embodiments, the LIN5 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 2.
[0147] In some specific embodiments, the source-sink relationship related gene is a rice GIF1 gene. In some embodiments, the GIF1 gene encodes a GIF1 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, even 100% to the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, the GIF1 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 4.
[0148] In some embodiments, the source-sink relationship associated gene is the maize MN1 gene. In some embodiments, the MN1 gene encodes a MN1 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 5. In some embodiments, the MN1 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 6.
[0149] In some embodiments, the source-sink relationship associated gene is the wheat TraesCS2A03G0736600 gene. In some embodiments, the TraesCS2A03G0736600 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the TraesCS2A03G0736600 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 8.
[0150] In some embodiments, the source-sink relationship associated gene is the soybean Glyma.10G074800 gene. In some embodiments, the Glyma.10G074800 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 24. In some embodiments, the Glyma.10G074800 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 25.
[0151] The "promoter of a yield-related gene, preferably of an endogenous source-sink relationship-related gene" as used herein refers to a sequence of a gene upstream of the translation initiation codon of said gene, which can regulate the expression of said gene. Typically, the promoter is a sequence of about 100 bp to about 10 kb, preferably about 2 kb, upstream of the translation initiation codon of said gene, preferably not including the initiation codon, in the genome.
[0152] In some embodiments, the promoter of the tomato LIN5 gene comprises a nucleotide sequence having at most 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 9.
[0153] In some embodiments, the promoter of the rice GIF1 gene comprises a nucleotide sequence having at most 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 10.
[0154] In some embodiments, the yield-related gene, preferably the source-sink relationship-related gene, can also be a SRG1, FZY6, SOE, GRG1 gene.
[0155] Tomato SRG1 is a plant type-related gene, and increasing the expression of SRG1 can reduce the number of lateral branches, thereby achieving the purpose of yield increase by reducing energy consumption. Tomato auxin synthesis gene FZY6 encodes a flavin monooxygenase, which can catalyze the direct conversion of indolepyruvic acid into IAA, is a rate-limiting enzyme in the tryptophan-dependent indole acetic acid biosynthesis pathway, and plays a key role in maintaining auxin content; tomato SOE belongs to the chlorophyll-binding protein gene family, and the protein encoded by this type of gene is an important component in photosynthesis, mainly involved in the assembly of photosystems and the capture and transmission of light energy, and is a gene with "source increase" potential that can promote plant photosynthesis. This gene can optimize the light energy utilization efficiency of tomato, reduce photoinhibition and photodamage, thereby supporting the normal growth of plants under continuous light conditions. GRG1 is a gravity stimulus-related gene in rice.
[0156] In some embodiments, the SRG1 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 42.
[0157] In some embodiments, the FZY6 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 43.
[0158] In some embodiments, the SOE gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 44.
[0159] In some embodiments, the GRG1 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 45.
[0160] In some embodiments, the one or more stress-responsive cis-elements are introduced into the predetermined site in the promoter. Introduction of the one or more elements at the predetermined site does not substantially affect the original gene expression regulation of the gene, and is capable of effectively exerting its stress-responsive gene expression regulation function.
[0161] In some embodiments, the predetermined site is 1) not located within an endogenous cis-element of the gene; 2) located within open chromatin; and / or 3) close to the start codon of the gene.
[0162] In some embodiments, the selection of the predetermined site comprises:
[0163] a) selecting a position in the promoter that does not contain a possible cis-element as a candidate site based on the Plant CARE (PlantCARE, a database of plant promoters and their cis-acting regulatory elements (ugent.be) cis-element online prediction website;
[0164] b) performing DNase I hypersensitivity prediction on the candidate sites obtained in step a) by the DHS (DNase-I hypersensitive sites) online prediction website (http: / / www.epigenome.cuhk.edu.hk / ), and screening out candidate sites with a peak value less than 0.5; and
[0165] c) selecting from the candidate sites obtained in step b) the site closest to the start codon of the gene as the final predetermined site.
[0166] In some preferred embodiments, the predetermined site is no more than about 2 kb, preferably no more than about 1 kb, more preferably no more than 500 bp from the translation start codon of the gene.
[0167] In some specific embodiments, the stress-responsive cis-element, such as a heat-responsive element, is introduced into the promoter of the rice GIF1 gene at 427 bp upstream of the start codon.
[0168] In some specific embodiments, the stress-responsive cis-element, such as a heat-responsive element, is introduced into the promoter of the rice GIF1 gene at 427 bp upstream of the start codon.
[0169] In some specific embodiments, the stress-responsive cis-element, such as a heat-responsive element, is introduced into the promoter of the rice GIF1 gene at 427 bp upstream of the start codon.
[0170] In some specific embodiments, the stress-responsive cis-element, such as a heat-responsive element, is introduced into the promoter of the rice GIF1 gene at 427 bp upstream of the start codon.
[0171] In some specific embodiments, the stress-responsive cis-element, such as a heat-responsive element, is introduced into the promoter of the rice GIF1 gene at 427 bp upstream of the start codon.
[0172] In some specific embodiments, the stress-responsive cis-element, such as a light-responsive element, is introduced into the promoter of the SRG1 gene at 311 bp upstream of the start codon.
[0173] In some specific embodiments, the stress-responsive cis-element, such as a heat-responsive element, is introduced into the promoter of the FZY6 gene at 491 bp upstream of the start codon.
[0174] In some specific embodiments, the stress-responsive cis-element, such as a light-responsive element, is introduced into the promoter of the SOE gene at 39 bp upstream of the start codon.
[0175] In some specific embodiments, the stress-responsive cis-element, such as a drought-responsive element, is introduced into the promoter of the GRG1 gene at 546 bp upstream of the start codon.
[0176] In some embodiments, a plurality (e.g., 2 to about 10 or more) stress response cis-elements are introduced into the promoter. In some embodiments, the plurality of stress response cis-elements are introduced into the promoter in tandem, e.g., into the predetermined site within the promoter.
[0177] In some embodiments, the method of the present application further comprises, prior to introducing the stress response cis-element into the plant, verifying that the stress response element inserted into the predetermined site is capable of conferring stress response ability by expressing a reporter system in vitro, e.g., in a tobacco cell expressing a reporter system.
[0178] In the present application, the introduction of the one or more stress response cis-elements can be achieved by genome editing. For example, in some embodiments, a genome editing system targeting the predetermined site can be introduced into the plant. The genome editing system can be a CRISPR, ZFN or TALEN based genome editing system. Preferably, the genome editing system is a CRISPR based genome editing system, preferably a prime editing system.
[0179] For example, in some embodiments, the plant can be introduced with 1) a genome editing system targeting the predetermined site, which results in a double-strand break (DSB) of the genome at or near the predetermined site; 2) a homologous recombination donor nucleic acid comprising the one or more stress response cis-elements and homology arm sequences corresponding to the two sides of the DSB, whereby the one or more stress response cis-elements are introduced into the predetermined site by homologous recombination.
[0180] In some preferred embodiments, the introduction of the one or more stress response cis-elements is achieved by introducing a prime editing system into the plant.
[0181] In some embodiments, the prime editing system comprises:
[0182] i) a fusion protein comprising a CRISPR nickase and a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding the fusion protein; and
[0183] ii) a pegRNA and / or an expression construct containing a nucleotide sequence encoding the pegRNA,
[0184] wherein the pegRNA comprises, in the 5’ to 3’ direction, a guide sequence, a gRNA scaffold sequence, a reverse transcription (RT) template sequence, and a primer binding site (PBS) sequence, wherein the pegRNA is capable of forming a complex with the fusion protein and targeting the fusion protein to a first target sequence in the genome comprising or adjacent to the predetermined site, resulting in a first nick within the first target sequence, wherein the reverse transcription template sequence comprises a nucleotide sequence of the one or more stress-responsive cis-elements or a complement thereof.
[0185] As used herein, a "target sequence" refers to a sequence of about 20 nucleotides in length in a genome characterized by a PAM (protospacer adjacent motif) sequence flanking 5' or 3'. Generally, a PAM is necessary for a complex of a CRISPR nuclease or variant thereof and a guide RNA to recognize a target sequence. For example, for Cas9 nucleases and variants thereof, the target sequence is immediately adjacent to the PAM at the 3' end, e.g., 5'-NGG-3'. Based on the presence of a PAM, one of skill in the art can readily determine target sequences in a genome that can be targeted. Moreover, depending on the location of the PAM, a target sequence can be on either strand of a genomic DNA molecule. For Cas9 or derivatives thereof, e.g., Cas9 nickases, a target sequence is preferably 20 nucleotides.
[0186] In some embodiments, the CRISPR nickase in the fusion protein is capable of forming a nick within the first target sequence in the genomic DNA. In some embodiments, the CRISPR nickase is a Cas9 nickase.
[0187] In some embodiments, the Cas9 nickase is derived from SpCas9 of S. pyogenes and comprises at least the amino acid substitution H840A relative to wild-type SpCas9. In some preferred embodiments, the Cas9 nickase is derived from SpCas9 of S. pyogenes and comprises at least the amino acid substitutions H840A, R221K, and N394K relative to wild-type SpCas9. An exemplary wild-type SpCas9 comprises the amino acid sequence set forth in SEQ ID NO: 11. In some embodiments, the Cas9 nickase nCas9(H840A) comprises the amino acid sequence set forth in SEQ ID NO: 12 or nCas9(H840A, R221K, N394K) SEQ ID NO: 26. In some embodiments, the Cas9 nickase in the fusion protein is capable of forming a nick between the -3 position nucleotide of the PAM of the first target sequence (the first nucleotide 5' of the PAM sequence is the +1 position) and the -4 position nucleotide. The nick results in the first target sequence forming a free single strand with a 3' end (3' free single strand) and a free single strand with a 5' end (5' free single strand).
[0188] In some embodiments, the reverse transcriptase in the fusion protein of the present application can be derived from different sources. In some embodiments, the reverse transcriptase is a viral-derived reverse transcriptase. For example, in some embodiments, the reverse transcriptase is an M-MLV reverse transcriptase or a functional variant thereof. In some embodiments, the reverse transcriptase is a CaMV-RT from Cauliflower mosaic virus (CaMV). In some embodiments, the reverse transcriptase is a bacterial-derived reverse transcriptase, such as a retron-RT from Escherichia coli. In some specific embodiments, the reverse transcriptase in the fusion protein of the present application comprises the amino acid sequence set forth in SEQ ID NO: 13.
[0189] In some embodiments, the fusion protein further comprises a LA polypeptide (small RNA-binding exonuclease protection factor La) at the C-terminus. The LA polypeptide is capable of binding to small RNA and protecting it from being cleaved by exonuclease. An exemplary LA polypeptide comprises the amino acid sequence set forth in SEQ ID NO: 27.
[0190] In some embodiments, the CRISPR nickase, the reverse transcriptase, and / or the LA polypeptide in the fusion protein are connected by a linker. As used herein, a "linker" can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 20-25, 25-50) or more amino acids, without secondary structure above. For example, the linker can be a flexible peptide linker. For example, the linker can be the linker set forth in SEQ ID NO: 14 (33 aa linker).
[0191] In some embodiments, the CRISPR nickase in the fusion protein is fused to the N-terminus of the reverse transcriptase, directly or through a linker. In some embodiments, the CRISPR nickase in the fusion protein is fused to the C-terminus of the reverse transcriptase, directly or through a linker.
[0192] In some embodiments of the application, the fusion protein of the application can further comprise a nuclear localization sequence (NLS). Generally, the NLS(s) in the fusion protein should be of sufficient strength to drive accumulation of the fusion protein in the nucleus of a cell in an amount that enables it to perform its base editing function. Generally, the strength of the nuclear localization activity is determined by the number, position, specific NLS(es) used, or a combination of these factors, of the NLS(s) in the fusion protein. In some embodiments, the NLS is an SV40 NLS (amino acid sequence set forth in SEQ ID NO: 15).
[0193] The guide sequence (also referred to as seed sequence or spacer sequence) in the pegRNA of the application is arranged to have sufficient sequence identity (preferably 100% identity) to the first target sequence, so as to be able to bind to the complementary strand of the first target sequence through base pairing, enabling sequence-specific targeting.
[0194] A variety of scaffold sequences suitable for gRNAs for CRISPR nuclease (e.g., Cas9)-based genome editing are known in the art, and these can be used in the application. In some particular embodiments, the scaffold sequence of the gRNA is set forth in SEQ ID NO: 16.
[0195] In some embodiments, the primer binding sequence is arranged to be complementary to at least a portion of the first target sequence, preferably the primer binding sequence is complementary to at least a portion of a 3' overhang single strand resulting from the nick in the sense strand of the first target sequence, in particular to the nucleotide sequence of the 3' end of the 3' overhang single strand. When the 3' overhang single strand of the sense strand binds to the primer binding sequence via base pairing, the 3' overhang single strand can serve as a primer to perform reverse transcription of a reverse transcription (RT) template sequence immediately adjacent to the primer binding sequence under the action of a reverse transcriptase in the fusion protein, extending a DNA sequence corresponding to the reverse transcription (RT) template sequence.
[0196] The primer binding sequence can depend on the length of the overhang single strand formed in the target sequence by the CRISPR nickase used, however, it should have a minimum length ensuring specific binding. In some embodiments, the primer binding sequence can be 4-20 nucleotides in length, for example 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides in length.
[0197] In some embodiments, the RT template sequence is arranged to be complementary to at least a portion of a sequence downstream of the nick in the first target sequence, and comprises a nucleotide sequence of the one or more stress-responsive cis-elements or a complement thereof.
[0198] In some embodiments, the RT template sequence can be about 1-300 or more nucleotides in length, for example 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300 nucleotides or more polynucleotides in length.
[0199] In some embodiments, the prime editing system further comprises iii) a nick gRNA and / or an expression construct containing a nucleotide sequence encoding the nick gRNA, the nick gRNA comprising a guide sequence and a gRNA scaffold sequence. In some preferred embodiments, the nick gRNA does not comprise a reverse transcription (RT) template sequence and a primer binding site (PBS) sequence.
[0200] The guide sequence (also referred to as the seed sequence or spacer sequence) in the nicking gRNA of the present application is configured to have sufficient sequence identity (preferably 100% identity) to a second target sequence in the genome, such that the fusion protein is capable of targeting the second target sequence and causing a second nick within the second target sequence, which is located on the opposite strand of the genomic DNA from the first target sequence. In some embodiments, the first nick and the second nick are separated by about 1 to about 300 or more nucleotides, e.g., 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300 nucleotides or more. In some embodiments, the nick formed by the nicking gRNA is upstream or downstream of the nick formed by the pegRNA (both upstream and downstream refer to the strand of DNA where the pegRNA target sequence is located). In some embodiments, the guide sequence in the nicking gRNA has sufficient sequence identity (preferably 100% identity) to the opposite strand (the modified strand) of the pegRNA target sequence after the editing event occurs, such that the nicking gRNA only targets the nick target sequence that is created only after the pegRNA-induced targeting and modification of the target sequence is complete. In some embodiments, the PAM of the nick target sequence is located within the complement of the pegRNA target sequence.
[0201] In some embodiments, the prime editing system further comprises iv) a Csy4 protein and / or an expression construct comprising a nucleotide sequence encoding the Csy4 protein.
[0202] In some embodiments, the 3’ end of the pegRNA and / or the nicking RNA comprises a Csy4 recognition site sequence on one side or on both the 3’ and 5’ ends.
[0203] In some embodiments, the Csy4 protein described herein comprises the amino acid sequence set forth in SEQ ID NO: 17. The Csy4 recognition site comprises the nucleotide sequence set forth in SEQ ID NO: 18.
[0204] In some embodiments, the pegRNA further comprises a tevopreQ1 motif at the 3’ end. The tevopreQ1 motif can prevent degradation of the pegRNA. An exemplary tevopreQ1 motif comprises the nucleotide sequence set forth in SEQ ID NO: 28.
[0205] In some embodiments, the expression constructs of items i), ii), iii), and / or iv) in the prime editing system described herein can be separate expression constructs, or can be the same expression construct in any combination. For example, i) and iii) can be the same construct. Alternatively, ii) and iv) can be the same construct.
[0206] In some embodiments, the prime editing system described herein comprises a first expression construct encoding a fusion protein comprising, from N-terminus to C-terminus: a Csy4 protein - a self-cleaving peptide - a NLS - a CRISPR nickase - a NLS - a linker - a reverse transcriptase - a NLS; and a second expression construct comprising: a Csy4 recognition site sequence - a pegRNA coding sequence - a Csy4 recognition site sequence - a nick gRNA coding sequence - a Csy4 recognition site sequence.
[0207] In some embodiments, the prime editing system described herein comprises a first expression construct encoding a fusion protein comprising, from N-terminus to C-terminus: a Csy4 protein - a self-cleaving peptide - a NLS - a CRISPR nickase - a NLS - a linker - a reverse transcriptase - a linker - a LA polypeptide - a NLS; and a second expression construct comprising: a Csy4 recognition site sequence - a pegRNA coding sequence - a Csy4 recognition site sequence - a tevopreQ1 motif - a nick gRNA coding sequence - a Csy4 recognition site sequence.
[0208] In some embodiments, the first expression construct drives expression of the fusion protein by a 35S promoter, and the second expression construct drives expression by an RNA polymerase II promoter, such as a CmYLCV promoter (SEQ ID NO: 23).
[0209] As used herein, “self-cleaving peptide” means a peptide that can achieve self-cleavage within a cell. For example, the self-cleaving peptide can comprise a protease recognition site, such that it is recognized and specifically cleaved by a protease within the cell.
[0210] Alternatively, the self-cleaving peptide can be a 2A polypeptide. 2A polypeptides are a class of short peptides from viruses that self-cleave during translation. When two different proteins of interest are expressed in the same reading frame with a 2A polypeptide, the two proteins of interest are produced in nearly a 1 : 1 ratio. Commonly used 2A polypeptides can be P2A from porcine techovirus-1, T2A from Thosea asigna virus, E2A from equine rhinitis A virus, and F2A from foot-and-mouth disease virus. Among them, P2A has the highest cleavage efficiency and is therefore preferred. A variety of functional variants of these 2A polypeptides are also known in the art and can be used in the present application.
[0211] In some embodiments, the coding sequences of the proteins / polypeptides described herein can be codon-optimized according to the species to be applied.
[0212] Codon optimization refers to a method of modifying a nucleic acid sequence in order to enhance expression in a host cell of interest by replacing at least one codon of the native sequence (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) with a codon that is more frequently or most frequently used in the genes of the host cell while maintaining the native amino acid sequence. Different species exhibit particular biases for certain codons of a particular amino acid. Codon bias (differences in codon usage between organisms) often correlates with the efficiency of translation of messenger RNA (mRNA), which is believed to be dependent on the properties of the codons being translated and the availability of particular transfer RNA (tRNA) molecules. The predominance of selected tRNAs within a cell generally reflects the frequency with which codons are used in protein synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables can be readily obtained, for example, in the Codon Usage Database available at www.kazusa.orjp / codon / , and these tables can be adapted in different ways. See, Nakamura Y. et al., “Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).
[0213] In some embodiments, the first expression construct comprises the nucleotide sequence set forth in SEQ ID NO: 19.
[0214] In some embodiments, the first expression construct comprises the nucleotide sequence set forth in SEQ ID NO: 29.
[0215] In some embodiments, the second expression construct comprises the sequence set forth below:
[0216] wherein italic represents the CmYLCV promoter sequence,
[0217] bold represents the scaffold sequence of gRNA,
[0218] underlined represents the Csy4 recognition site sequence,
[0219] N x1 represents the first guide sequence of pegRNA,
[0220] N x2 represents the primer binding sequence and reverse transcription template sequence of pegRNA,
[0221] N x3 represents the second guide sequence of nicked gRNA,
[0222] N is A, T, C or G; xi, x2 or x3 is any integer, preferably, xi or x2 is 20.
[0223] In some embodiments, the second expression construct comprises the sequence set forth below:
[0224] wherein italic represents the CmYLCV promoter sequence,
[0225] bold represents the scaffold sequence of gRNA,
[0226] underlined represents the Csy4 recognition site sequence,
[0227] N x1 represents the first guide sequence of pegRNA,
[0228] N x2 represents the primer binding sequence and reverse transcription template sequence of pegRNA,
[0229] N x3 represents the second guide sequence of nicked gRNA,
[0230] N is A, T, C, or G; x1, x2, or x3 is any integer, preferably x1 or x2 is 20. The capital letters represent the tevopreQ1 motif.
[0231] In some embodiments, the plant is a tomato. In some embodiments, the stress response cis-element is a heat response element, e.g., the heat response element comprises the sequence ATTCTAGAAT, GTTCATGAAC, or AGAACGTTCT. In some embodiments, the source-sink related gene is a tomato LIN5 gene. In some embodiments, introduction of the heat response cis-element results in the modified plant comprising a mutated LIN5 gene promoter set forth in SEQ ID NO: 20, 30, or 31. In some embodiments, the second expression construct comprises the nucleotide sequence set forth in SEQ ID NO: 22.
[0232] In some embodiments, the plant is a tomato. In some embodiments, the stress response cis-element is a drought response element, e.g., the drought response element comprises the sequence CATGTG. In some embodiments, the source-sink related gene is a tomato LIN5 gene. In some embodiments, introduction of the drought response cis-element results in the modified plant comprising a mutated LIN5 gene promoter set forth in SEQ ID NO: 32.
[0233] In some embodiments, the plant is a tomato. In some embodiments, the stress response cis-element is a flood response element, e.g., the flood response element comprises the sequence TTGACC. In some embodiments, the source-sink related gene is a tomato LIN5 gene. In some embodiments, introduction of the flood response cis-element results in the modified plant comprising a mutated LIN5 gene promoter set forth in SEQ ID NO: 33.
[0234] In some embodiments, the plant is a tomato. In some embodiments, the stress response cis-element is a light response element, e.g., the light response element comprises the sequence TGTGTGGTTAATATGAAGATAAGATT. In some embodiments, the source-sink related gene is a tomato LIN5 gene. In some embodiments, introduction of the light response cis-element results in the modified plant comprising a mutated LIN5 gene promoter set forth in SEQ ID NO: 34.
[0235] In some embodiments, the plant is Solanum lycopersicum. In some embodiments, the stress-responsive cis-element is a heat-responsive element, e.g., the heat-responsive element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is the Solanum lycopersicum FZY6 gene. In some embodiments, introduction of the heat-responsive cis-element results in the modified plant comprising a mutated FZY6 gene promoter set forth in SEQ ID NO: 36.
[0236] In some embodiments, the plant is Solanum lycopersicum. In some embodiments, the stress-responsive cis-element is a heat-responsive element, e.g., the heat-responsive element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is the Solanum lycopersicum FZY6 gene. In some embodiments, introduction of the heat-responsive cis-element results in the modified plant comprising a mutated FZY6 gene promoter set forth in SEQ ID NO: 36.
[0237] In some embodiments, the plant is Solanum lycopersicum. In some embodiments, the stress-responsive cis-element is a light-responsive element, e.g., the light-responsive element comprises the sequence AAATTGTGA. In some embodiments, the source-sink related gene is the Solanum lycopersicum SOE gene. In some embodiments, introduction of the light-responsive cis-element results in the modified plant comprising a mutated SOE gene promoter set forth in SEQ ID NO: 37.
[0238] In some embodiments, the plant is Oryza sativa. In some embodiments, the stress-responsive cis-element is a heat-responsive element, e.g., the heat-responsive element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is the Oryza sativa GIF1 gene. In some embodiments, introduction of the stress-responsive cis-element results in the modified plant comprising a mutated GIF1 gene promoter set forth in SEQ ID NO: 21.
[0239] In some embodiments, the plant is Oryza sativa. In some embodiments, the stress-responsive cis-element is a drought-responsive element, e.g., the drought-responsive element comprises the sequence TACCGACAT. In some embodiments, the source-sink related gene is the Oryza sativa GRG1 gene. In some embodiments, introduction of the drought-responsive cis-element results in the modified plant comprising a mutated GRG1 gene promoter set forth in SEQ ID NO: 38.
[0240] In some embodiments, the plant is maize. In some embodiments, the stress- responsive cis-element is a heat-responsive element, e.g., the heat-responsive element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is a maize LIN5 gene. In some embodiments, the introduction of the heat-responsive cis-element results in the modified plant comprising a mutated maize LIN5 gene promoter set forth in SEQ ID NO: 39.
[0241] In some embodiments, the plant is soybean. In some embodiments, the stress- responsive cis-element is a heat-responsive element, e.g., the heat-responsive element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is a soybean LIN5 gene. In some embodiments, the introduction of the heat-responsive cis-element results in the modified plant comprising a mutated soybean LIN5 gene promoter set forth in SEQ ID NO: 40.
[0242] In some embodiments, the plant is wheat. In some embodiments, the stress- responsive cis-element is a heat-responsive element, e.g., the heat-responsive element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is a wheat LIN5 gene. In some embodiments, the introduction of the heat-responsive cis-element results in the modified plant comprising a mutated wheat LIN5 gene promoter set forth in SEQ ID NO: 41.
[0243] In the methods of the application, the genome editing system, e.g., a prime editing system, can be introduced into a plant by various methods well known to those skilled in the art. Methods that can be used to introduce the prime editing system of the application into a plant include, but are not limited to, biolistics, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway, and ovary injection.
[0244] In some embodiments, the introduction comprises transforming the genome editing system, e.g., a prime editing system, of the application into an isolated plant cell or tissue, and then regenerating the transformed plant cell or tissue into a whole plant.
[0245] In other embodiments, the genome editing system, e.g., a prime editing system, of the application can be transformed into a specific location on a whole plant, such as a leaf, a shoot tip, a pollen tube, an ear shoot, or a hypocotyl. This is particularly suitable for transformation of plants that are difficult to regenerate via tissue culture.
[0246] In some embodiments of the application, an in vitro expressed protein and / or an in vitro transcribed RNA molecule (e.g., the expression construct is an in vitro transcribed RNA molecule) is directly transformed into the plant. The protein and / or RNA molecule is capable of effecting genome editing in the plant cell, followed by degradation by the cell, avoiding integration of exogenous nucleotide sequences into the plant genome.
[0247] III. Modified Plants
[0248] In another aspect, the present application provides a modified plant, wherein the expression of one or more endogenous yield-related genes, preferably endogenous source-sink relationship-related genes, of said plant comprises one or more stress-responsive cis-elements in the regulatory sequence, such as the promoter, of said genes. In some embodiments, the modified plant is obtained or obtainable by the methods of the present application.
[0249] The plants of the present application are preferably crop plants, including but not limited to Solanum lycopersicum (tomato), Nicotiana benthamiana (tobacco), Capsicum annuum (pepper), Physalis pruinosa (ground cherry), Solanum melongena (eggplant), Solanum tuberosum (potato), Solanum pennellii (Pennell's tomato), Solanum chilense (Chilean tomato), Solanum habrochaites (Habrochaites tomato), Solanum pimpinellifolium (Pimpinellifolium tomato), Solanum galapagense (Galapagos tomato), Petunia hybrid (petunia), grape, strawberry, Citrullus lanatus (watermelon), Cucumis sativus (cucumber), Lactuca sativa L. (lettuce), Chinese cabbage, rape, cabbage, wheat, Oryza sativa L. (rice), maize, soybean, sunflower, sorghum, rape, alfalfa, cotton, barley, millet, sugarcane, and cassava.
[0250] In some specific embodiments, the plant is a tomato. In some embodiments, the plant can be from different cultivars of Solanum lycopersicum (tomato). In some embodiments, the plant is a tomato cultivar Ailsa Craig, M82, Heinz 1706, Jingxian 8, VF-1, Beijing 1, TS545, TS181, TS590. Preferably, the plant is a tomato cultivar Ailsa Craig.
[0251] In some embodiments, the plant is rice. In some embodiments, the plant can be from different cultivars of rice. In some embodiments, the plant is rice cultivar ZH11, Wuyou rice-4, Longjing 31 or Zhongkefa 5.
[0252] In some embodiments, the introduction of the stress responsive cis-element results in the expression of the yield-related gene, preferably the source-sink relationship-related gene, in the plant being up- or down-regulated, preferably up-regulated, in response to the corresponding stress. That is, when the plant comprising the introduced stress responsive cis-element is subjected to the corresponding stress, the expression of the yield-related gene, preferably the source-sink relationship-related gene, is up- or down-regulated, preferably up-regulated.
[0253] The stress described herein can be biotic or abiotic stress. Abiotic stress includes, but is not limited to, heat stress, cold stress, osmotic stress (such as drought stress, salt stress, metal stress), water stress (flooding), light stress (insufficient or excessive light, too long or too short), nutrient stress (deficiency or excess of nutrient elements in the soil, such as nitrogen deficiency stress, phosphorus deficiency stress), mechanical stress (wind, hail, and other physical damage), and the like. Biotic stress includes fungal infection, bacterial infection, viral infection, parasitic infection, insect stress, weed competition, parasitic plant, and the like. In some preferred embodiments, the stress described herein is heat stress.
[0254] In some embodiments, the modified plant has increased yield, e.g., increased fruit / seed set, increased fruit / seed weight, compared to the unmodified wild type plant in the presence and / or absence of the stress. In some embodiments, the modified plant has increased plot yield compared to the unmodified wild type plant in the presence and / or absence of the stress. In some embodiments, the yield is increased by about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100% or more.
[0255] The stress responsive cis-element described herein includes, but is not limited to, a cis-element selected from Table 1.
[0256] In some embodiments, the stress responsive cis-element is a heat responsive element. In some preferred embodiments, the heat responsive element comprises the sequence ATTCTAGAAT. In some preferred embodiments, the heat responsive element comprises the sequence GTTCATGAAC. In some preferred embodiments, the heat responsive element comprises the sequence AGAACGTTCT.
[0257] In some embodiments, the stress response cis-element is a light response element. In some preferred embodiments, the light response element comprises the sequence TGTGTGGTTAATATGAAGATAAGATT. In some preferred embodiments, the light response element comprises the sequence GTGTGTGAA.
[0258] In some embodiments, the stress response cis-element is a drought response element. In some preferred embodiments, the drought response element comprises the sequence CATGTG.
[0259] In some embodiments, the stress response cis-element is a flood response element. In some preferred embodiments, the flood response element comprises the sequence TTGACC.
[0260] The yield-related genes, preferably source-sink relationship-related genes, described herein include, but are not limited to, genes encoding cell wall invertase (CWIN), sucrose transporter (SUT), hexose transporter (HT), SWEET (Sugars Will Eventually be Exported Transporter), sucrose-phosphate synthase (SPS), sucrose synthase (SUS), and the like.
[0261] Suitable source-sink relationship-related genes in tomato include, but are not limited to, genes selected from the genes of Table 2.
[0262] For other species, the source-sink relationship-related genes encompass homologous genes to the above-mentioned tomato genes. For example, the homologous genes encode proteins having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to the proteins encoded by the tomato genes. Alternatively, the coding sequences of the homologous genes have at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to the coding sequences of the tomato genes.
[0263] In some embodiments, the source-sink relationship related gene is a cell wall invertase (CWIN) encoding gene. In some embodiments, the cell wall invertase (CWIN) encoding gene is a tomato LIN5 gene or a homologous gene thereof, such as a rice GIF1 gene, a maize MN1 gene, a soybean Glyma.10G074800 gene, or a wheat TraesCS2A03G0736600 gene.
[0264] In some specific embodiments, the source-sink relationship related gene is a tomato LIN5 gene. In some embodiments, the LIN5 gene encodes a LIN5 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 1. In some embodiments, the LIN5 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 2.
[0265] In some specific embodiments, the source-sink relationship related gene is a rice GIF1 gene. In some embodiments, the GIF1 gene encodes a GIF1 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, even 100% to the amino acid sequence set forth in SEQ ID NO: 3. In some embodiments, the GIF1 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 4.
[0266] In some embodiments, the source-sink relationship associated gene is the maize MN1 gene. In some embodiments, the MN1 gene encodes a MN1 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 5. In some embodiments, the MN1 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 6.
[0267] In some embodiments, the source-sink relationship associated gene is the wheat TraesCS2A03G0736600 gene. In some embodiments, the TraesCS2A03G0736600 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 7. In some embodiments, the TraesCS2A03G0736600 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 8.
[0268] In some embodiments, the source-sink relationship associated gene is the soybean Glyma.10G074800 gene. In some embodiments, the Glyma.10G074800 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 24. In some embodiments, the Glyma.10G074800 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 25.
[0269] A "promoter of a yield-related gene, preferably of an endosperm sink-source relationship-related gene" as used herein refers to a sequence located upstream of the translation start codon of said gene, which can modulate the expression of said gene. Typically, the promoter is a sequence of about 100 bp to about 10 kb, preferably about 2 kb, 5' upstream of the translation start codon of said gene, preferably not including the start codon, in the genome.
[0270] In some embodiments, the promoter of the tomato LIN5 gene comprises a nucleotide sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the nucleotide sequence set forth in SEQ ID NO: 9.
[0271] In some embodiments, the promoter of the rice GIF1 gene comprises a nucleotide sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the nucleotide sequence set forth in SEQ ID NO: 10.
[0272] In some embodiments, the yield-related gene, preferably the sink-source relationship-related gene, can also be a SRG1, FZY6, SOE, or GRG1 gene.
[0273] In some embodiments, the SRG1 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the amino acid sequence set forth in SEQ ID NO: 42.
[0274] In some embodiments, the FZY6 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the amino acid sequence set forth in SEQ ID NO: 43.
[0275] In some embodiments, the SOE gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the amino acid sequence set forth in SEQ ID NO: 44.
[0276] In some embodiments, the GRG1 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 45.
[0277] In some embodiments, the one or more stress-responsive cis-elements are introduced into a predetermined site in the promoter. Introduction of the one or more elements at the predetermined site does not substantially affect the original gene expression regulation of the gene, and is capable of effectively exerting its stress-responsive gene expression regulation function.
[0278] In some embodiments, the predetermined site is 1) not located within an endogenous cis-element of the gene; 2) located within open chromatin; and / or 3) close to the start codon of the gene.
[0279] In some embodiments, the selection of the predetermined site comprises:
[0280] a) selecting a position in the promoter that does not contain a possible cis-element as a candidate site based on the Plant CARE (PlantCARE, a database of plant promoters and their cis-acting regulatory elements (ugent.be) cis-element online prediction website;
[0281] b) performing DNase I hypersensitivity prediction on the candidate sites obtained in step a) by the DHS (DNase-I hypersensitive sites) online prediction website (http: / / www.epigenome.cuhk.edu.hk / ), and screening out candidate sites with a peak value less than 0.5; and
[0282] c) selecting a site closest to the start codon of the gene from the candidate sites obtained in step b) as the final predetermined site.
[0283] In some preferred embodiments, the predetermined site is no more than about 2 kb, preferably no more than 1 kb, more preferably no more than 500 bp, away from the translation start codon of the gene.
[0284] In some specific embodiments, the stress-responsive cis-element, such as a heat-responsive element, is introduced into the promoter of the tomato LIN5 gene at 410 bp or 452 b upstream of the start codon.
[0285] In some embodiments, the stress response cis-element, e.g., heat response element, is introduced into the promoter of the rice GIF1 gene at a position 5' to the start codon at position 427 bp.
[0286] In some embodiments, a plurality (e.g., 2 to about 10 or more) stress response cis-elements are introduced into the promoter. In some embodiments, the plurality of stress response cis-elements are introduced into the promoter in tandem, e.g., at the predetermined position in the promoter.
[0287] In some embodiments, the plant is tomato. In some embodiments, the stress is heat stress, the stress response cis-element is a heat response element, e.g., the heat response element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is the tomato LIN5 gene. In some embodiments, introduction of the stress response cis-element results in the modified plant comprising a mutated LIN5 gene promoter of SEQ ID NO: 20. In some embodiments, the heat stress conditions for cultivation of tomato are, e.g., about 3-10°C higher during the day and about 2-5°C higher at night compared to normal growth conditions. In some embodiments, the heat stress conditions for greenhouse cultivation of tomato are about 32°C / about 21°C (day / night), and the corresponding normal conditions are about 28°C / about 19°C (day / night). In some embodiments, the heat stress conditions for open field cultivation of tomato are about 38°C / about 24°C (day / night), and the corresponding normal conditions are about 32°C / about 22°C (day / night).
[0288] In some embodiments, the plant is tomato. In some embodiments, the stress is heat stress, the stress response cis-element is a heat response element, e.g., the heat response element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is the tomato LIN5 gene. In some embodiments, introduction of the stress response cis-element results in the modified plant comprising a mutated LIN5 gene promoter of SEQ ID NO: 20. In some embodiments, the heat stress conditions for cultivation of tomato are, e.g., about 3-10°C higher during the day and about 2-5°C higher at night compared to normal growth conditions. In some embodiments, the heat stress conditions for greenhouse cultivation of tomato are about 32°C / about 21°C (day / night), and the corresponding normal conditions are about 28°C / about 19°C (day / night). In some embodiments, the heat stress conditions for open field cultivation of tomato are about 38°C / about 24°C (day / night), and the corresponding normal conditions are about 32°C / about 22°C (day / night).
[0289] In some embodiments, the plant is Solanum lycopersicum. In some embodiments, the stress response cis-element is a drought response element, e.g., the drought response element comprises the sequence CATGTG. In some embodiments, the source-sink related gene is a Solanum lycopersicum LIN5 gene. In some embodiments, introduction of the drought response cis-element results in the modified plant comprising a mutated LIN5 gene promoter set forth in SEQ ID NO: 32.
[0290] In some embodiments, the plant is Solanum lycopersicum. In some embodiments, the stress response cis-element is a flood response element, e.g., the flood response element comprises the sequence TTGACC. In some embodiments, the source-sink related gene is a Solanum lycopersicum LIN5 gene. In some embodiments, introduction of the flood response cis-element results in the modified plant comprising a mutated LIN5 gene promoter set forth in SEQ ID NO: 33.
[0291] In some embodiments, the plant is Solanum lycopersicum. In some embodiments, the stress response cis-element is a light response element, e.g., the light response element comprises the sequence TGTGTGGTTAATATGAAGATAAGATT. In some embodiments, the source-sink related gene is a Solanum lycopersicum LIN5 gene. In some embodiments, introduction of the light response cis-element results in the modified plant comprising a mutated LIN5 gene promoter set forth in SEQ ID NO: 34.
[0292] In some embodiments, the plant is Solanum lycopersicum. In some embodiments, the stress response cis-element is a light response element, e.g., the light response element comprises the sequence GTGTGTGAA. In some embodiments, the source-sink related gene is a Solanum lycopersicum SRG1 gene. In some embodiments, introduction of the light response cis-element results in the modified plant comprising a mutated SRG1 gene promoter set forth in SEQ ID NO: 35.
[0293] In some embodiments, the plant is Solanum lycopersicum. In some embodiments, the stress response cis-element is a heat response element, e.g., the heat response element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is a Solanum lycopersicum FZY6 gene. In some embodiments, introduction of the heat response cis-element results in the modified plant comprising a mutated FZY6 gene promoter set forth in SEQ ID NO: 36.
[0294] In some embodiments, the plant is Solanum lycopersicum. In some embodiments, the stress-responsive cis-element is a light-responsive element, e.g., the light-responsive element comprises the sequence AAATTGTGA. In some embodiments, the source-sink related gene is a Solanum lycopersicum SOE gene. In some embodiments, introduction of the light-responsive cis-element results in the modified plant comprising a mutated SOE gene promoter set forth in SEQ ID NO: 37.
[0295] In some embodiments, the plant is Oryza sativa. In some embodiments, the stress-responsive cis-element is a drought-responsive element, e.g., the drought-responsive element comprises the sequence TACCGACAT. In some embodiments, the source-sink related gene is an Oryza sativa GRG1 gene. In some embodiments, introduction of the drought-responsive cis-element results in the modified plant comprising a mutated GRG1 gene promoter set forth in SEQ ID NO: 38.
[0296] In some embodiments, the plant is Zea mays. In some embodiments, the stress-responsive cis-element is a heat-responsive element, e.g., the heat-responsive element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is a Zea mays LIN5 gene. In some embodiments, introduction of the heat-responsive cis-element results in the modified plant comprising a mutated Zea mays LIN5 gene promoter set forth in SEQ ID NO: 39.
[0297] In some embodiments, the plant is Glycine max. In some embodiments, the stress-responsive cis-element is a heat-responsive element, e.g., the heat-responsive element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is a Glycine max LIN5 gene. In some embodiments, introduction of the heat-responsive cis-element results in the modified plant comprising a mutated Glycine max LIN5 gene promoter set forth in SEQ ID NO: 40.
[0298] In some embodiments, the plant is Triticum aestivum. In some embodiments, the stress-responsive cis-element is a heat-responsive element, e.g., the heat-responsive element comprises the sequence ATTCTAGAAT. In some embodiments, the source-sink related gene is a Triticum aestivum LIN5 gene. In some embodiments, introduction of the heat-responsive cis-element results in the modified plant comprising a mutated Triticum aestivum LIN5 gene promoter set forth in SEQ ID NO: 41.
[0299] IV. Guide editing systems for modifying monocot plants
[0300] In another aspect, the present application provides a prime editing system for gene editing in a dicotyledonous plant, comprising:
[0301] i) a fusion protein comprising a CRISPR nickase and a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding said fusion protein;
[0302] ii) a pegRNA and / or an expression construct containing a nucleotide sequence encoding said pegRNA, wherein said pegRNA comprises, in the 5’ to 3’ direction, a first guide sequence, a scaffold sequence, a reverse transcription (RT) template sequence, and a primer binding site (PBS) sequence, wherein said pegRNA is capable of forming a complex with said fusion protein and targeting said fusion protein to a first target sequence in the genome, resulting in a first nick within said first target sequence; and
[0303] iii) optionally, a nick gRNA and / or an expression construct containing a nucleotide sequence encoding said nick gRNA, said nick gRNA comprising a second guide sequence and a scaffold sequence, wherein said nick gRNA is capable of forming a complex with said fusion protein and targeting said fusion protein to a second target sequence in the genome, resulting in a second nick within said second target sequence;
[0304] wherein said pegRNA and / or said nick RNA comprises a Csy4 recognition site sequence at its 3’ end or at both 3’ and 5’ ends, respectively.
[0305] In some embodiments, the prime editing system further comprises iv) a Csy4 protein and / or an expression construct containing a nucleotide sequence encoding said Csy4 protein.
[0306] Dicotyledonous plants described herein include, but are not limited to, Arabidopsis thaliana (thale cress), Solanum lycopersicum (tomato), Nicotiana benthamiana (tobacco), Capsicum annuum (pepper), Physalis pruinosa (ground cherry), Solanum melongena (eggplant), Solanum tuberosum (potato), Solanum pennellii (pennell tomato), Solanum chilense (Chilean tomato), Solanum habrochaites (habrochaites tomato), Solanum pimpinellifolium (pimpinellifolium tomato), Solanum galapagense (Galapagos tomato), Petunia hybrid (petunia), grape, strawberry, Citrullus lanatus (watermelon), Cucumis sativus (cucumber), Lactuca sativa L. (lettuce), Chinese cabbage, rape, cabbage, soybean, cotton, alfalfa, cassava.
[0307] In some preferred embodiments, the dicotyledonous plant is a tomato. In some embodiments, the plant can be from different cultivars of Solanum lycopersicum (tomato). In some embodiments, the plant is tomato cultivar Ailsa Craig, M82, Heirloom 1, Jingcai 8, VF-1, Beijing 1, TS545, TS181, TS590. Preferably, the plant is tomato cultivar Ailsa Craig.
[0308] As used herein, "target sequence" refers to a sequence of about 20 nucleotides in length in the genome characterized by a 5' or 3' flanking PAM (protospacer adjacent motif) sequence. Generally, a PAM is necessary for the complex of a CRISPR nuclease or variant thereof and a guide RNA to recognize a target sequence. For example, for Cas9 nucleases and variants thereof, the target sequence is immediately adjacent to the PAM at the 3' end, e.g., 5'-NGG-3'. Based on the presence of a PAM, one of skill in the art can readily determine target sequences in a genome that can be targeted. Moreover, depending on the location of the PAM, the target sequence can be on either strand of the genomic DNA molecule. For Cas9 or derivatives thereof, e.g., Cas9 nickases, the target sequence is preferably 20 nucleotides.
[0309] In some embodiments, the CRISPR nickase in the fusion protein is capable of forming a nick within a first target sequence in genomic DNA. In some embodiments, the CRISPR nickase is a Cas9 nickase.
[0310] In some embodiments, the Cas9 nickase is derived from SpCas9 of S. pyogenes and comprises at least the amino acid substitution H840A relative to wild-type SpCas9. In some preferred embodiments, the Cas9 nickase is derived from SpCas9 of S. pyogenes and comprises at least the amino acid substitutions H840A, R221K, and N394K relative to wild-type SpCas9. An exemplary wild-type SpCas9 comprises the amino acid sequence set forth in SEQ ID NO: 11. In some embodiments, the Cas9 nickase nCas9(H840A) comprises the amino acid sequence set forth in SEQ ID NO: 12 or nCas9(H840A, R221K, N394K) SEQ ID NO: 26. In some embodiments, the Cas9 nickase in the fusion protein is capable of forming a nick between the -3 position nucleotide of the PAM of the first target sequence (the first nucleotide 5' of the PAM sequence is the +1 position) and the -4 position nucleotide. The nick results in the first target sequence forming a free single strand with a 3' end (3' free single strand) and a free single strand with a 5' end (5' free single strand).
[0311] In some embodiments, the reverse transcriptase in the fusion protein of the present application can be derived from different sources. In some embodiments, the reverse transcriptase is a viral-derived reverse transcriptase. For example, in some embodiments, the reverse transcriptase is an M-MLV reverse transcriptase or a functional variant thereof. In some embodiments, the reverse transcriptase is a CaMV-RT from Cauliflower mosaic virus (CaMV). In some embodiments, the reverse transcriptase is a bacterial-derived reverse transcriptase, such as a retron-RT from Escherichia coli. In some specific embodiments, the reverse transcriptase in the fusion protein of the present application comprises the amino acid sequence set forth in SEQ ID NO: 13.
[0312] In some embodiments, the fusion protein further comprises a LA polypeptide (small RNA-binding exonuclease protection factor La) at the C-terminus. The LA polypeptide is capable of binding to small RNA and protecting it from being cleaved by exonucleases. An exemplary LA polypeptide comprises the amino acid sequence set forth in SEQ ID NO: 27.
[0313] In some embodiments, the CRISPR nickase, the reverse transcriptase, and / or the LA polypeptide in the fusion protein are connected by a linker. As used herein, a "linker" can be a non-functional amino acid sequence of 1-50 (e.g., 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, or 20-25, 25-50) or more amino acids that does not have secondary structure above. For example, the linker can be a flexible peptide linker. For example, the linker can be the linker set forth in SEQ ID NO: 14 (33 aa linker).
[0314] In some embodiments, the CRISPR nickase in the fusion protein is fused to the N-terminus of the reverse transcriptase, directly or through a linker. In some embodiments, the CRISPR nickase in the fusion protein is fused to the C-terminus of the reverse transcriptase, directly or through a linker.
[0315] In some embodiments of the application, the fusion protein of the application can further comprise a nuclear localization sequence (NLS). Generally, one or more NLS in the fusion protein should be of sufficient strength to drive accumulation of the fusion protein in the nucleus of a cell in an amount that enables it to perform its base editing function. Generally, the strength of nuclear localization activity is determined by the number, position, specific NLS(es) used, or a combination of these factors, of NLS in the fusion protein. In some embodiments, the NLS is an SV40 NLS (amino acid sequence set forth in SEQ ID NO: 15).
[0316] The guide sequence (also referred to as seed sequence or spacer sequence) in the pegRNA of the application is arranged to have sufficient sequence identity (preferably 100% identity) to the first target sequence, so as to be able to bind to the complementary strand of the first target sequence through base pairing, achieving sequence-specific targeting.
[0317] A variety of scaffold sequences suitable for gRNAs for CRISPR nuclease (e.g., Cas9)-based genome editing are known in the art, which can be used in the application. In some particular embodiments, the scaffold sequence of the gRNA is set forth in SEQ ID NO: 16.
[0318] In some embodiments, the primer binding sequence is arranged to be complementary to at least a portion of the first target sequence, preferably the primer binding sequence is complementary to at least a portion of a 3' overhang single strand resulting from the nick in the sense strand of the first target sequence, in particular to the nucleotide sequence of the 3' end of the 3' overhang single strand. When the 3' overhang single strand of the sense strand binds to the primer binding sequence via base pairing, the 3' overhang single strand can serve as a primer to perform reverse transcription of a reverse transcription (RT) template sequence immediately adjacent to the primer binding sequence under the action of a reverse transcriptase in the fusion protein, to extend a DNA sequence corresponding to the reverse transcription (RT) template sequence.
[0319] The primer binding sequence can depend on the length of the overhang single strand formed in the target sequence by the CRISPR nickase used, however, it should have a minimum length to ensure specific binding. In some embodiments, the primer binding sequence can have a length of 4-20 nucleotides, for example, a length of 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 nucleotides.
[0320] The RT template sequence described herein can be any sequence. By reverse transcription as described above, its sequence information can be integrated into the DNA strand where the target sequence is located (i.e. the strand containing the PAM of the target sequence), and then through the DNA repair action of the cell, a DNA double strand containing the sequence information of the RT template sequence is formed. In some embodiments, the RT template sequence contains a desired modification. For example, the desired modification includes one or more substitutions, deletions and / or additions of nucleotides. For example, the modification includes one or more substitutions selected from the group consisting of: C to T substitution, C to G substitution, C to A substitution, G to T substitution, G to C substitution, G to A substitution, A to T substitution, A to G substitution, A to C substitution, T to C substitution, T to G substitution, T to A substitution; and / or a deletion of one or more nucleotides, for example, 1 to about 100 or more, for example, 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotides; and / or an insertion of one or more nucleotides, for example, 1 to about 100 or more, for example, 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100 nucleotides. Preferably, the RT template sequence contains the nucleotide sequence of the one or more stress-responsive cis-elements or a complement thereof.
[0321] In some embodiments, the RT template sequence is configured to be complementary to at least a portion of the sequence downstream of the first target sequence nick and comprises a desired modification. The desired modification includes substitution, deletion and / or addition of one or more nucleotides.
[0322] In some embodiments, the RT template sequence can be about 1-300 or more nucleotides in length, for example 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300 nucleotides or more polynucleotides in length. Preferably, the RT template sequence is 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23 nucleotides in length.
[0323] In some preferred embodiments, the nicking gRNA does not comprise a reverse transcription (RT) template sequence and a primer binding site (PBS) sequence.
[0324] The guide sequence (also referred to as seed sequence or spacer sequence) in the nicking gRNA of the present application is configured to have sufficient sequence identity (preferably 100% identity) to a second target sequence in the genome, such that the fusion protein is capable of targeting the second target sequence and causing a second nick within the second target sequence, the second target sequence being located on the opposite strand of the genomic DNA from the first target sequence. In some embodiments, the first nick and the second nick are separated by about 1- about 300 or more nucleotides, for example 1, 2, 3, 4, 5, about 10, about 20, about 30, about 40, about 50, about 75, about 100, about 125, about 150, about 175, about 200, about 225, about 250, about 275, about 300 nucleotides or more polynucleotides. In some embodiments, the nick formed by the nicking gRNA is upstream or downstream of the nick formed by the pegRNA (both the upstream and downstream refer to the DNA strand where the pegRNA target sequence is located). In some embodiments, the guide sequence in the nicking gRNA has sufficient sequence identity (preferably 100% identity) to the opposite strand (modified) of the pegRNA target sequence after the editing event occurs, such that the nicking gRNA only targets the nicked target sequence that is generated only after the completion of the pegRNA-induced target sequence targeting and modification. In some embodiments, the PAM of the nicked target sequence is located within the complement of the pegRNA target sequence.
[0325] In some embodiments, the 3’ end of the pegRNA and / or the 3’ and 5’ ends of the nicking RNA comprise a Csy4 recognition site sequence.
[0326] In some embodiments, the Csy4 protein described herein comprises the amino acid sequence set forth in SEQ ID NO: 17. The Csy4 recognition site comprises the nucleotide sequence set forth in SEQ ID NO: 18.
[0327] In some embodiments, the pegRNA further comprises a tevopreQ1 motif at the 3’ end. The tevopreQ1 motif can prevent degradation of the pegRNA. An exemplary tevopreQ1 motif comprises the nucleotide sequence set forth in SEQ ID NO: 28.
[0328] In some embodiments, the expression constructs of items i), ii), iii), and / or iv) of the prime editing system described herein can be separate expression constructs, or can be the same expression construct in any combination. For example, i) and iii) can be the same construct. Alternatively, ii) and iv) can be the same construct.
[0329] In some embodiments, the prime editing system described herein comprises a first expression construct encoding a fusion protein comprising, from N- to C-terminus: a Csy4 protein - a self-cleaving peptide - a NLS - a CRISPR nickase - a NLS - a linker - a reverse transcriptase - a NLS; and a second expression construct comprising: a Csy4 recognition site sequence - a pegRNA coding sequence - a Csy4 recognition site sequence - a nicking gRNA coding sequence - a Csy4 recognition site sequence.
[0330] In some embodiments, the prime editing system described herein comprises a first expression construct encoding a fusion protein comprising, from N- to C-terminus: a Csy4 protein - a self-cleaving peptide - a NLS - a CRISPR nickase - a NLS - a linker - a reverse transcriptase - a linker - a LA polypeptide - a NLS; and a second expression construct comprising: a Csy4 recognition site sequence - a pegRNA coding sequence - a Csy4 recognition site sequence - a tevopreQ1 motif - a nicking gRNA coding sequence - a Csy4 recognition site sequence.
[0331] In some embodiments, the first expression construct drives expression of the fusion protein by a 35S promoter, and the second expression construct drives expression by an RNA polymerase II promoter, such as a CmYLCV promoter (SEQ ID NO: 23).
[0332] As used herein, "self-cleaving peptide" means a peptide that can achieve self- cleavage within a cell. For example, the self-cleaving peptide can comprise a protease recognition site, thereby being recognized and specifically cleaved by a protease within the cell.
[0333] Alternatively, the self-cleaving peptide can be a 2A polypeptide. 2A polypeptides are a class of short peptides from viruses, whose self-cleavage occurs during translation. When two different proteins of interest are expressed in the same reading frame with a 2A polypeptide, the two proteins of interest are generated almost in a 1 : 1 ratio. Commonly used 2A polypeptides can be P2A from porcine techovirus-1, T2A from Thosea asigna virus, E2A from equine rhinitis A virus, and F2A from foot-and-mouth disease virus. Among them, P2A has the highest cleavage efficiency and is therefore preferred. A variety of functional variants of these 2A polypeptides are also known in the art, which can also be used in the present application.
[0334] In some embodiments, the coding sequences of the proteins / polypeptides described herein can be codon-optimized according to the species to be applied.
[0335] Codon optimization refers to a method of modifying a nucleic acid sequence to enhance expression in a host cell of interest by replacing at least one codon of a native sequence (e.g., about or more than about 1, 2, 3, 4, 5, 10, 15, 20, 25, 50, or more codons) with a codon that is more frequently or most frequently used in the genes of the host cell while maintaining the native amino acid sequence. Different species exhibit particular biases for certain codons of a particular amino acid. Codon bias (differences in codon usage between organisms) often correlates with the efficiency of translation of messenger RNA (mRNA), which is believed to be dependent on the properties of the codons being translated and the availability of a particular transfer RNA (tRNA) molecule. The predominance of selected tRNAs within a cell generally reflects the frequency with which codons are used in protein synthesis. Thus, genes can be tailored for optimal gene expression in a given organism based on codon optimization. Codon usage tables can be readily obtained, for example, in the Codon Usage Database available at www.kazusa.orjp / codon / , and these tables can be adapted in different ways. See, Nakamura Y. et al., “Codon usage tabulated from the international DNA sequence databases: status for the year 2000. Nucl. Acids Res., 28:292 (2000).
[0336] In some embodiments, the first expression construct comprises the nucleotide sequence set forth in SEQ ID NO: 19.
[0337] In some embodiments, the first expression construct comprises the nucleotide sequence set forth in SEQ ID NO: 29.
[0338] In some embodiments, the second expression construct comprises the sequence set forth below:
[0339] wherein italicized represents a CmYLCV promoter sequence,
[0340] bold represents a scaffold sequence of the gRNA,
[0341] underlined represents a Csy4 recognition site sequence,
[0342] N x1 represents a first guide sequence of the pegRNA,
[0343] N x2the first guide sequence representing the pegRNA,
[0344] N x3 the second guide sequence representing the nicking gRNA,
[0345] N is A, T, C or G; x1, x2 or x3 is any integer, preferably x1 or x2 is 20.
[0346] In some embodiments, the second expression construct comprises the following sequences:
[0347] wherein italic represents the CmYLCV promoter sequence,
[0348] bold represents the scaffold sequence of the gRNA,
[0349] underlined represents the Csy4 recognition site sequence,
[0350] N x1 the first guide sequence representing the pegRNA,
[0351] N x2 the primer binding sequence and reverse transcription template sequence representing the pegRNA,
[0352] N x3 the second guide sequence representing the nicking gRNA,
[0353] N is A, T, C or G; x1, x2 or x3 is any integer, preferably x1 or x2 is 20. Capital letters represent the tevopreQ1 motif.
[0354] In another aspect, the present application provides a method of producing a genetically modified dicot plant, comprising introducing the prime editing system of the present application into at least one of said dicot plants. In some embodiments, the introduction of the prime editing system of the present application results in a genetic modification in the genome of said at least one dicot plant.
[0355] In the present application, the genetic modification can be located anywhere in the plant genome, for example within a functional gene such as a protein-coding gene, or for example can be located in a gene expression regulatory region such as a promoter region or enhancer region, thereby achieving a modification of the function of the gene or a modification of the gene expression. The modification in the sequence of the genome of the cell can be detected by T7EI, PCR / RE or sequencing methods.
[0356] In some embodiments, the method further comprises screening the at least one plant for a plant having a desired modification.
[0357] In the methods of the application, the prime editing system can be introduced into the plant in a variety of ways known to those skilled in the art. Methods that can be used to introduce the prime editing system of the application into a plant include, but are not limited to, biolistics, PEG-mediated protoplast transformation, Agrobacterium-mediated transformation, plant virus-mediated transformation, pollen tube pathway, and ovary injection.
[0358] In some embodiments, the introducing comprises transforming the prime editing system of the application into an isolated plant cell or tissue, and then regenerating the transformed plant cell or tissue into a whole plant.
[0359] In other embodiments, the prime editing system of the application can be transformed into a specific location on a whole plant, such as a leaf, shoot tip, pollen tube, young ear, or hypocotyl. This is particularly suitable for transformation of plants that are difficult to regenerate via tissue culture.
[0360] In some embodiments of the application, an in vitro expressed protein and / or an in vitro transcribed RNA molecule (e.g., the expression construct is an in vitro transcribed RNA molecule) is directly transformed into the plant. The protein and / or RNA molecule is capable of effecting genome editing in the plant cell, and is subsequently degraded by the cell, avoiding integration of exogenous nucleotide sequences into the plant genome.
[0361] In some embodiments of the application, wherein the modified genomic region is associated with a plant trait, such as an agronomic trait, the modification substitution results in the plant having an altered (preferably improved) trait, such as an agronomic trait, relative to a wild type plant.
[0362] In some embodiments, the method further comprises the step of screening the plants for the desired modification and / or the desired trait, such as an agronomic trait.
[0363] In some embodiments of the application, the method further comprises obtaining progeny of the genetically modified dicot plant. Preferably, the genetically modified dicot plant or its progeny has the desired modification and / or the desired trait, such as an agronomic trait.
[0364] In another aspect, the application also provides a genetically modified dicot plant or its progeny or a part thereof, wherein the dicot plant is obtained by the above-described methods of the application. Preferably, the genetically modified plant or its progeny has the desired genetic modification and / or the desired trait, such as an agronomic trait.
[0365] In another aspect, the present application also provides a method for breeding a plant, comprising crossing a first genetically modified dicot plant obtained by the method described above with a second dicot plant which does not contain the modification, thereby introducing the modification into the second dicot plant. Preferably, the first genetically modified dicot plant has a desired trait such as an agronomic trait. Examples
[0366] The present application can be further understood by reference to the specific examples described herein, which are intended to be purely exemplary of the application and do not limit the scope of the application. Obviously, many modifications and variations of the present application are possible in light of the above teachings, the scope of the application is not limited by the specific examples described herein. It is therefore to be understood that within the scope of the appended claims, and their equivalents, changes can be made to the
[0367] Experimental materials and methods
[0368] Plant materials and growth conditions
[0369] Tomato (Solanum lycopersicum) cultivar Ailsa Craig (AC) was used in this study. Seeds were directly sown in soil in 72-well plug trays. Tomato seedlings were transplanted into pots after they had grown 3-4 true leaves and were grown under long-day conditions (16 hours light / 8 hours dark) in a growth chamber equipped with LED lights (Philips Lighting IBRS, 10461, 5600 VB, NL) under a 25°C / 23°C day / night temperature regime, or in a greenhouse under natural light; or plants were transplanted into soil and grown in a plastic sunlight greenhouse or under natural light outdoors.
[0370] Rice variety ZH11 was used in this study. For seedling production, rice seeds were first soaked on filter paper and then sown in 96-well plastic plates. Seedlings were transplanted into the field after one month of growth. Rice plants were grown in a field in Beijing, China (39.9° N, 116.3° E) from May to October each year; and in a field in Hainan, China (18.25° N, 109.50° E) from February to May each year. The planting density, fertilization, ploughing, irrigation, sowing, and harvesting of tomato and rice were performed according to the local farm cultivation management methods.
[0371] Heat treatment method
[0372] 1. Heat treatment of AC material under sunlight greenhouse planting mode (32°C day / 21°C night, 50 days): After the first ear inflorescence end fruit of AC plants transplanted into the sunlight greenhouse set fruit, cover plastic film for warming (simulated heat treatment), leave 40 cm ventilation opening at the bottom for ventilation. Phenotypic records and data statistics were taken after the first and second ear fruits were fully matured. Temperature changes during AC growth were monitored in real time by sensors.
[0373] 2. Heat treatment of AC material under open-air planting mode (38°C day / 24°C night, 36 days): After the first ear inflorescence end fruit of AC plants transplanted into the open-air environment set fruit, cover plastic film for warming (simulated heat treatment), do not cover plastic film at the head and tail of the warming shed for ventilation. Phenotypic records and data statistics were taken after the first and second ear fruits were fully matured. Temperature changes during AC growth were monitored in real time by sensors.
[0374] 3. Heat treatment of rice in paddy field: ZH11 rice was transplanted into the paddy field and grew to the heading stage during which the outdoor daytime temperature was 31°C and the nighttime temperature was 22°C. To increase the temperature, an additional layer of transparent plastic film was covered on the plants, with the bottom open for air circulation. The average daytime temperature was 36°C and the nighttime temperature was 24°C. Fully matured ears were harvested for yield analysis. Sensors were used to monitor temperature changes during ZH11 growth in real time.
[0375] Yield determination
[0376] To ensure the reliability of yield determination and as much as possible to correspond to the five yield determination standards (Khaipho-Burch et al., 2023), yield evaluation was conducted on tomato AC material and rice ZH11 material planted in different plots or different geographical locations and different growth seasons.
[0377] The statistical standard of AC yield per unit area under sunlight greenhouse conditions is the total weight of the first three fruits of each tomato plant per square meter. Under normal and heat treatment conditions, WT and ac-lin5-de were each counted in 4 plots, i.e. 16 tomato plants were harvested for WT and ac-lin5-de under normal and heat treatment conditions; representative plants were selected from the 16 tomato plants for phenotype acquisition and yield index determination such as individual plant yield and individual fruit weight. The statistical standard of AC yield per unit area under open-air mode is the total weight of all fruits on each tomato plant per square meter. Under normal and heat treatment conditions, WT and ac-lin5-de were each counted in 5 plots, i.e. 15 tomato plants were harvested for WT and ac-lin5-de under normal and heat treatment conditions. Representative plants were selected from the 15 tomato plants for phenotype acquisition and yield index determination such as individual plant yield and individual fruit weight.
[0378] Phenotype acquisition and data statistics were performed after the rice material was fully matured. In Beijing, WT and gif1-de were each set up in 3 plots under normal and heat treatment conditions, with a protective row area of 1.6 x 1.6 meters. A total of 24 plants were randomly selected from the 3 repeated plots under normal and heat treatment conditions, i.e. 8 plants from each plot, for representative picture acquisition of WT and gif1-de under normal and heat treatment conditions, and for yield index statistical data analysis such as tiller number, panicle length, primary panicle branch number, secondary panicle branch number, grain number per panicle, seed setting rate, 1000-grain weight, individual plant yield, etc. The yield measurement area of the rice plot was about 2 square meters (1.4 m x 1.4 m) after removing the protective row, and a total of 64 rice plants were included in each plot. After mechanical threshing, all the plants were treated in a 37°C oven for 3 days, and then weighed for yield measurement.
[0379] CWIN family phylogenetic tree analysis
[0380] The protein sequences of CWIN family members of Arabidopsis, rice, maize and tomato were downloaded from the Phytozome v13 database (https: / / phytozome-next.jgi.doe.gov / ). The sequence phylogenetic tree of all CWIN proteins was constructed using MEGAX software. The phylogenetic tree was constructed by maximum likelihood method with 500 bootstrap repeats using MEGAX (Kumar et al., 2018).
[0381] sHSP gene family promoter phylogenetic tree analysis and HSE element analysis
[0382] The sequences of 1K upstream of the Arabidopsis, rice and tomato sHSP family members promoters were downloaded from Phytozome v13 database (https: / / phytozome-next.jgi.doe.). The MEGAX software was used to construct the phylogenetic tree of all sHSP family members promoters sequences. To identify heat shock elements (HSEs), the Arabidopsis, rice and tomato promoter sequences were subjected to element screening using MEME (https: / / meme-suite.org / meme / tools / meme). The parameters were set as “Find only palindromes” and “Max and min width of elements 6 and 15 bases, respectively”. Sequences with E-value less than 10 and P-value less than 0.0001 must be displayed in the output results.
[0383] Plasmid construction, CRISPR / Cas9 mutagenesis, plant transformation and genotype identification
[0384] To engineer 22C-PPE and MH-PPE vectors, M-MLV reverse transcriptase and nCas9 (H840A) were codon-optimized for both dicotyledonous and monocotyledonous plants. Subsequently, pegRNA and nicking sgRNA (BGI) were synthesized commercially according to the insertion position of HSEs in LIN5 and GIF1 promoters, amplified using primers and cloned into Csy4-PE and pCAM-PE, respectively, by homologous recombination. The final binary vectors were introduced into tomato AC and rice ZH11 by Agrobacterium-mediated transformation. The T0 transgenic plants were transplanted into soil and grown under standard greenhouse conditions. DNA was extracted from 3 different leaf samples of each plant using the CTAB method to extract genomic DNA, and the mutations generated by Csy4-PE and pCAM-PE were genotyped by PCR amplification. The mutations in the target genes of T0 and T1 transgenic positive plants were identified using a combination of CELI digestion of PCR products and Sanger sequencing. The gene sequences of tomato and rice were obtained from the Sol Genomics Network (SGN) database (https: / / solgenomics.net / ) and rice Annotation Project (RAP) database (https: / / rapdb.dna.affrc.go.jp / ), respectively.
[0385] Prediction of cis-elements and DNase I hypersensitive sites on promoters
[0386] To avoid the disruption of the expression pattern of the gene of LIN5 by the insertion of HSE, first, the cis-acting elements of the 1k sequence upstream of the LIN5 promoter were predicted using the online prediction website PlantCARE, (a database of cis-acting regulatory elements, enhancers, and repressors in plants. http: / / bioinformatics.psb.ugent.be / webtools / plantcare / html / ), and multiple promoter candidate regions without cis-acting elements in the prediction results were selected for standby. Next, the openness of the chromatin upstream of the LIN5 promoter in the 17-day post-anthesis and 47-day post-anthesis tomato fruit tissues was predicted by the online prediction website of DNase I hypersensitive sites (http: / / www.epigenome.cuhk.edu.hk / jbrowse2 / ), and the relatively inactive chromatin region was selected as a candidate insertion region. Combining the promoter candidate regions predicted by cis-acting elements and DHSs, the region with the overlap of the two and as close to the start codon ATG as possible was finally selected as the target region of the insertion element, and the target was designed for editing in this region. If the closest overlapping region to the start codon does not have a usable PAM, the second closest overlapping region to the start codon is selected, if there is still no usable PAM, the third one is selected, and so on. The determination of the insertion site of the promoter of the rice GIF1 gene is similar to that of LIN5.
[0387] Transcriptional activity assay
[0388] Dual luciferase reporter system was used for transcriptional activity analysis. In the experiment, LIN5 own promoter (ProLIN5:LUC), LIN5 promoter inserted with HSE (ProLIN5+HSE:LUC) and LIN5 promoter inserted with scrambled HSE (ProLIN5+scrambled HSE:LUC) were used as reporter genes respectively. Firefly REN gene driven by CaMV 35S promoter (35S:REN) was used as internal control. Empty vector was used as negative control. Four constructs were injected into the same tobacco leaf, each occupying one quarter of the leaf area, and the infiltrated area was marked with a marker pen for easy sampling later. After 48 hours of infiltration, the tobacco leaves were subjected to time gradient (0 hour, 0.5 hour, 1 hour, 2 hour, 4 hour, 8 hour) heat treatment in a plant incubator. The samples were quickly punched with a 5mm diameter puncher, and then the heat-treated tobacco leaves from different infiltrated areas were placed in 2ml centrifuge tubes and quickly frozen in liquid nitrogen. The leaf samples from at least three different plants were mixed as one biological repeat, and three biological repeats were set.
[0389] To determine luciferase LUC and REN activity, the leaf homogenate was extracted with 100ul lysis buffer and centrifuged at 4°C for 15min (12000rpm). The luciferase and Renilla activity assay was performed according to the manufacturer's protocol of the Luciferase Assay System (Promega) with slight modifications: 10ul of effective plant extract, 40ul of LARII and 40ul of Stop&Glo reagent were used. The measurement was performed using a GloMax 96 microplate luminometer (Promega) with a 2 second delay and a 10 second measurement. Three independent leaves from different plants were taken as one repeat, three repeats were set, and the average of the three samples was taken as the LUC / REN ratio. The transcriptional activity ratio of LUC to REN reflects its transcriptional activity. The luciferase and Renilla activity assay was performed according to the manufacturer's protocol of the Luciferase Assay System (Promega) with slight modifications: 10ul of effective plant extract, 40ul of LARII and 40ul of Stop&Glo reagent were used. The measurement was performed using a GloMax 96 microplate luminometer (Promega) with a 2 second delay and a 10 second measurement. Three independent leaves from different plants were taken as one repeat, three repeats were set, and the average of the three samples was taken as the LUC / REN ratio. The transcriptional activity ratio of LUC to REN reflects its transcriptional activity.
[0390] RNA extraction and quantitative RT-PCR (qPCR)
[0391] Total RNA was extracted with TRIzol reagent (Ambion, 391304). M-MLV reverse transcriptase (Tiangen, Y1420) was used for reverse transcription. qPCR was performed using gene-specific primers, TB Green Premix Ex Taq II (Takara, RR820) reaction system on a CFX96 Real-Time system (Bio-Rad) according to the manufacturer's instructions, with UBI as the internal control. These experiments were independently repeated three times.
[0392] Example 1: Integration of HSE into LIN5 promoter can make it obtain heat- responsive up-regulation ability
[0393] When plants are exposed to high temperature, a set of genes will be rapidly activated to cope with the stress. The promoter region of these genes usually contains heat shock cis-elements (HSEs), which can be recognized by heat shock transcription factors (HSFs) and rapidly initiate expression (Bonner et al., 1994; Waters, 1995; Richter et al., 2010). HSFs bind to DNA sequences with nGAAnnTTCn or nTTCnnGAAn HSE matrices, and there is no difference in preference for the two sequences (Xiao and Lis, 1988; Bonner et al., 1994). Therefore, the inventors propose that the targeted integration of HSEs into the LIN5 promoter can confer heat-responsive activation characteristics, thereby optimizing the activity of the library under stress response by fine-tuning the expression of LIN5 (Fig. 1A). To determine the optimal HSE, the inventors performed motif analysis on the promoters of all small heat shock protein targeted genes (sHSPs) in Arabidopsis, tomato, and rice, and found that nTTCnnGAAn or nGAAnnTTCn palindromic repeats were enriched (Fig. 1B). In the middle region of the repeat sequence, the CTAGA motif often appears in the promoter region of various plant stress response genes (Suzuki et al., 2011; Fragkostefanakis et al., 2015; Arce et al., 2018). Therefore, ATTCTAGAAT was used as the minimum HSE unit inserted into the LIN5 promoter region.
[0394] To alleviate the interference of HSE insertion on LIN5 expression pattern, the inventors established a selection criterion for targeting site by combining bioinformatics analysis and in planta transient expression test. First, the insertion site should not contain any functional cis-acting elements; second, the ideal insertion site should be located in an open chromatin region, according to the whole genome open chromatin identification database, such as DNase-I hypersensitive sites sequencing (DNase-seq), Assays for Transposase-Accessible Chromatin sequencing (ATAC-seq) or ChIP sequencing (ChIP-Seq) etc., the binding peak of which is relatively weak. Third, the insertion site is preferably located near the start codon, generally within 1000 bp. According to the above three selection criteria for insertion site, the inventors analyzed the LIN5 promoter sequence (2 kb upstream of the translation initiation codon) using the PlantCARE database to avoid the putative cis-elements. Next, the inventors also analyzed the chromatin openness of LIN5 promoter according to the published DNase-seq and ATAC-seq data, the binding peak of which is relatively weak (Lü et al., 2018). Finally, the site located 410 bp upstream of the start codon was selected as the insertion site of HSE (Fig. 1B).
[0395] Next, whether the chimeric LIN5 promoter with HSE is effective in planta was tested. Dual luciferase reporter assay was performed by transient expression in tobacco leaves (Fig. 1C). In this assay, luciferase (LUC) was driven by the native LIN5 promoter (ProLIN5) or the chimeric LIN5 promoter with HSE (ProLIN5-HSE) or scrambled HSE (AGGGTATTTT) (ProLIN5-ScramHSE), the latter of which served as a negative control (Fig. 1C and 1D). Meanwhile, 35S overexpressed Renilla served as an internal control. The results showed that the relative expression level of LUC driven by ProLIN5-HSE gradually increased after 2 hours of heat treatment (Fig. 1E). In contrast, neither ProLIN5 nor ProLIN5-ScramHSE could respond to heat stress, indicating that the insertion of HSE indeed conferred the LIN5 promoter with heat-responsive upregulation ability. In summary, the inventors established a set of scheme for designing stress-responsive chimeric promoters, which integrated large-scale screening of heat-responsive cis-elements, determination of appropriate insertion sites within the promoter region, and rapid efficiency verification in planta.
[0396] Example 2, Establishing CROCS breeding strategy by developing efficient prime editing system
[0397] Although transgenic strategy has been an important approach for crop improvement (Yu et al., 2022), the frequent occurrence of co-suppression phenomena such as post-transcriptional gene silencing and expression attenuation hinders the application of transgenic technology in crop breeding (Vaucheret et al., 1998). The heat-responsive upregulation property of LIN5 chimeric promoter prompted us to explore the possibility of knocking in HSE into the endogenous LIN5 gene promoter using genome editing (Fig. 2A). The biggest challenge to achieve this goal lies in the efficiency and accuracy of inserting DNA fragments into endogenous genes. Prime editing (PE) has been considered as a game-changing technology in gene editing since its inception (Marzec and Hensel, 2020), almost covering all editing requirements, including base substitution, transition, and even fragment insertion and deletion (Anzalone et al., 2019). Unfortunately, although PE systems have relatively high editing efficiency in monocot crops such as rice, there is still a lack of efficient PE systems in dicot plants (Lu et al., 2021; Vu et al., 2023; Yao et al., 2024).
[0398] PE contains a fusion protein composed of Cas9(H840A) nuclease and reverse transcriptase (RT), an engineered prime editing guide RNA (pegRNA) that also contains a primer binding site (PBS) sequence, a RT template (including the desired edit), and a nick-sgRNA that can target the non-editing strand. Preventing pegRNA from circularizing and maintaining the integrity of PBS and RT template is crucial for pegRNA to function in PE systems. In addition, degradation of pegRNA and nick-sgRNA 3' sequences by exonucleases also reduces PE efficiency (Liu et al., 2021; Nelson et al., 2022).
[0399] Based on the basic components contained in the prime editing system, and the processing principle of pegRNA and nick-sgRNA mediated by Csy4, the inventors restructured the prime editing system (Figure 2B). To this end, taking the third generation PE system (PE3) as a reference, the fusion protein of Csy4 protein and Cas9 (H840A) nickase and reverse transcriptase was co-expressed driven by 35s promoter, and CmYLCV promoter was used to co-express pegRNA and nick-sgRNA to form a fused transcript, and pegRNA and nick-sgRNA were flanked by Csy4 recognition sites. Csy4 nuclease can specifically recognize Csy4 sites, cut and release pegRNA and nick-sgRNA from the fused transcript. With the processing of Csy4, the 20 nt Csy4 recognition site remains at the 3' end of pegRNA to form a hairpin structure, which becomes the extended pegRNA and nick-sgRNA. In order to further improve the prime editing system, the structure of the prime editor was optimized, including the use of tomato codon-optimized RT, the addition of SV40 nuclear localization sequence at the N terminus of nCas9 and the C terminus of RT, and the addition of a 33-amino-acid linker containing SV40 nuclear localization sequence between nCas9 and RT (Figure 2B).
[0400] To verify the effectiveness of the Csy4-PE system, we also constructed and compared three expression forms based on the transcription of pegRNA and sgRNA sequences downstream of the U6 promoter and the transcription of two guide RNAs mediated by tRNAGly driven by the CmYLCV promoter (Figure 2B). According to the nucleotide sequence of HSE (ATTCTAGAAT) and the insertion site of the LIN5 promoter tested above (410 bp upstream of the start codon), a construct consisting of pegRNA and nick-sgRNA was designed (Figure 2B). The three constructs were transformed into tomato MT varieties for efficiency comparison. PCR genotyping and large-scale sequencing showed that no editing was detected for the U6-PE vector (Figure 2C), and no precise editing was detected for the tRNA-PE (Figure 2C). Remarkably, the editing efficiency of the optimized PE system was as high as 53.33%, with a precise editing efficiency as high as 13.33% (Figures 2C-E), and it had a high editing efficiency in dicotyledonous plants, reaching a practical level (Lu et al., 2021; Vu et al., 2023; Yao et al., 2024).
[0401] We also detected the short fragment insertion efficiency of Csy4-PE at different genetic sites and in different tomato varieties, and the results showed that the precise insertion efficiency was higher than 10% (Figure 2C-K), which indicated that our Csy4 PE system could have a relatively ideal editing efficiency in tomato.
[0402] In addition, on the basis of Csy4PE, a small RNA-binding exonuclease protection factor La was fused to the Cas9 nickase and M-MLV end to form Csy4PE7 (Yan et al., 2024), and further site (R221K and N394K) point mutations were made to the Cas9 nickase to form Csy4PE7max (Figure 2L). In addition, a structural RNA motif tevopreQ1 was added at the end of the Csy4 recognition site at the 3' end of the PegRNA expression cassette to form epegRNA (Figure 2M) to further prevent degradation of pegRNA and improve PE efficiency (Nelson et al., 2022).
[0403] The Csy4PE7max was used to determine the insertion of HSE elements into the LIN5 promoter, and it was found that the precise editing efficiency of Csy4PE7max with epegRNA (23.53%) was nearly 2 times higher than that of the Csy4PE system (13.3%), indicating that Csy4PE7max with epegRNA has higher precise editing efficiency (Figure 2N).
[0404] Example 3, Verification of the reliability of CROCS breeding strategy in tomato cultivars and different cultivation modes
[0405] The above guide editing construct was transformed into the classic tomato cultivar Alisa Craig (AC), and 37 independent transgenic lines were generated. Genotyping results showed that 9 lines were edited, and 5 lines had ideal editing effects. The precise editing efficiency reached 13.5% (5 / 37), confirming that the optimized PE system has high editing efficiency in different tomato varieties.
[0406] Next, a strain ac-lin5-de with precise knock-in of HSE was selected, and cas9 free T2 generation was identified for phenotypic analysis. To comprehensively evaluate the fruit yield and quality that match the actual tomato production, these plants were grown under optimal and heat stress conditions in a plastic sunlight greenhouse and open field, respectively (Figure 3). Under normal conditions in the sunlight greenhouse (day / night, 28°C / 19°C), ac-lin5-de had significantly improved fruit set (Figure 3A and 3B), and the fruit size was significantly increased, with a 9% increase in fruit weight compared to WT plants (Figure 3A and 3D). Remarkably, the average plot yield of ac-lin5-de was increased by 47% (Figure 3C), while the Brix content of the fruit was not compromised (Figure 3E).
[0407] Heat stress (day / night, 32°C / 21°C) treatment significantly reduced the fruit set (Figure 3F and 3G) and fruit weight (Figure 3I) of AC. Compared to the optimal state (Figure 3C), the average plot yield (Figure 3H) was reduced by 36.68%. The fruit weight (Figure 3I) and average plot yield (Figure 3H) of ac-lin5-de plants were 35% and 33% higher than those of WT plants, respectively. Notably, knocking in HSE in the LIN5 promoter rescued 43.61% of the yield loss (Figure 3C and 3H). In addition, the fruit Brix content of ac-lin5-de plants was also improved (Figure 3J).
[0408] To investigate the yield performance of ac-lin5-de plants in the open field, tomato plants were grown in summer and equipped with real-time temperature sensors for temperature monitoring. These tomato plants were grown naturally in the open field without any artificial intervention such as pruning or tying branches except for regular watering (Figure 3K and 3L). It was found that under normal conditions (day / night, 32°C / 22°C), the fruit set of ac-lin5-de was significantly higher than that of WT, resulting in an increase of nearly 14% in the average plot yield, while the fruit weight and Brix content were not affected (Figure 3M to P).
[0409] To create heat stress conditions in the open field cultivation environment, tomatoes were covered with transparent plastic film after they began to flower, and temperature sensors were provided to monitor temperature dynamics. Likewise, all plants under heat stress conditions were grown naturally, without human intervention (FIGS. 3Q and 3R). As expected, heat stress (day / night, 38°C / 24°C) significantly reduced the fruit set rate and yield of WT and ac-lin5-de compared to normal conditions (FIGS. 3S-T), although fruit weight and Brix content did not show significant differences (FIGS. 3U and 3V), and the average plot yield of ac-lin5-de was nearly 28% higher than that of WT under heat stress conditions (FIG. 3T). Importantly, when the average daytime temperature increased by 6°C and the average nighttime temperature increased by 2°C, AC plants reduced production by 20.28%, but ac-lin5-de plants overcame the reduction in production caused by high temperatures, indicating that precise knock-in of HSE can achieve stable yield in open fields under higher temperature conditions. These results show that the CROCS strategy can enable conventional tomato cultivars to achieve high yield under normal conditions and stable yield under heat stress in different cultivation modes.
[0410] Example 4, Rapid Creation of CROCS Strategy for High Yield in Favorable Conditions and Stable Yield in Adverse Conditions in Rice
[0411] It has been reported that a large portion of carbon assimilates produced by photosynthesis is transported to the stem, rather than being converted into grain for human consumption (Lafitte and Travis, 1984). In addition, heat stress during the grain filling stage of rice can inhibit the conversion of sucrose to glucose and fructose, thereby disrupting source-sink balance, leading to seed shriveling and yield loss (Peng et al., 2004; Wang et al., 2008). Under normal conditions, the empty grain rate of common rice varieties is about 10-20% (Shen et al., 2023). However, under stress conditions, especially when the temperature is higher than 35°C, the empty grain rate can be as high as 30-40% or even higher, thereby significantly reducing rice yield (Arshad et al., 2017; Wang et al., 2019; Ren et al., 2021). Therefore, in rice breeding and production, increasing the number and size of panicles, and reducing the empty grain rate, increasing grain filling rate and increasing thousand-grain weight, are equally important for achieving high yield in rice.
[0412] The conservation of CWIN function in different plant species, based on the fact that the inventors successfully optimized the expression of LIN5 in tomato varieties to improve yield and quality, prompted the inventors to explore whether the CROCS strategy could be applied to cereal crops. GRAIN INCOMPLETE FILLING 1 (GIF1), a CWIN homolog of tomato LIN5, controls sucrose partitioning during grain filling. Accumulated mutations in the GIF1 regulatory region were selected during domestication (Wang et al., 2008; Sun et al., 2014). However, the application of these natural variant alleles often requires multiple generations of backcrossing to introduce the trait, thus severely limiting their role in rice breeding.
[0413] To develop a fast and accurate method to optimize GIF1 function without genetic background limitations, we used the prime editing technology to insert HSE into the promoter of GIF1 to achieve heat-responsive optimization of source-sink relationships. Although prime editing has been used to replace a single or a few nucleotides (<5 bp) in monocot plants such as rice (Jiang et al., 2022), the efficiency and accuracy of prime editing to insert DNA sequences longer than 5 bp are still relatively low. Therefore, the strategy of optimizing the tomato PE system described above was adopted to design the rice PE system. Based on the pYLCRISPR / Cas9-MH vector (Ma et al., 2015), a PE system was designed by introducing a codon-optimized M-MLV reverse transcriptase and a novel RNA structure motif evopreQ1 (Nelson et al., 2022) at the 3' end of the pegRNA to protect the pegRNA from exonuclease degradation, named pCAM-PE (Fig. 4A). According to the three criteria for selecting target sites in the promoter region described above, the 427 bp upstream of the start codon of the GIF1 gene was selected as the target site for inserting HSE (ATTCTAGAAT). “ZH11” was chosen as the target variety for prime editing. ZH11 is a rice variety cultivated in northern China, with high yield, good grain quality, strong disease resistance, and resistance to abiotic stress, but its empty grain rate is usually 15-20%, which shows the potential for yield increase of this excellent variety (Nguyen et al., 2022). If the empty grain rate can be reduced and the seed setting rate increased, while increasing the thousand-grain weight, it is expected to breed a new rice variety with stress resistance, high yield, and good quality. Then, the PE construct was transformed into ZH11. A total of 37 independent transgenic lines were obtained, of which 3 transgenic lines (gif1-de) precisely inserted HSE at 427 bp upstream of the start codon of GIF1. The precise editing efficiency reached 8.1% (Fig. 4B and C), which is relatively high in the current use of single pegRNA PE systems to precisely insert DNA sequences longer than 5 bp in monocot plants.
[0414] To investigate the yield performance of gif1-de plants, yield trait evaluation was performed under normal and higher temperature conditions. The planting density of ZH11 and gif1-de plants was the density commonly used by local farmers, and a protective row was planted in each plot to exclude the influence of edge effects on yield measurement. The cultivation management of plants was closely related to that of local farms. Under normal conditions, gif1-de had no significant difference in tiller number (4F), ear length, primary branch number of ear and 1000-grain weight (Fig. 4G) compared with ZH11 (Fig. 4D, E). However, the plant height (Fig. 4H), seed setting rate (Fig. 4K), secondary branch number of ear (Fig. 4I) and grain number per ear (Fig. 4J) of gif1-de mutant were all significantly increased, which was similar to the effect on tomato plants (Fig. 3). These phenotypes indicated that the increase of sink activity might stimulate the strength of source organs. Notably, the seed setting rate of gif1-de was increased by 7% compared with ZH11 (Fig. 4K), indicating that gif-de significantly reduced the proportion of empty grains. Accordingly, the grain yield per plant and per plot of gif-de plants was increased by about 20% and 13%, respectively (Fig. 4L and Fig. 4M), indicating that the precise insertion of HSE into GIF1 promoter by PE alleviated the stubborn problem of grain abortion caused by uneven energy allocation among grains due to competition for carbon assimilates at different positions, thereby improving the grain yield.
[0415] To create a higher temperature condition for rice fields, the rice plants were covered with transparent plastic film at the booting stage, and the temperature changes were monitored by sensors. The results showed that the daytime temperature increased by about 5°C, and the nighttime temperature increased by about 3°C (Fig. 4N, O). The inventors found that after the temperature increased, the seed setting rate of WT plants decreased by 16.3% (Fig. 4K and Fig. 4U), and the yield per plant and per plot decreased by 32.2% and 37.9%, respectively (Fig. 4L and 4V; Fig. 4M and Fig. 4W), indicating that high temperature treatment indeed affected the yield. Then, the performance of WT and gif1-de under higher temperature was compared. Although the plant height (Fig. 4R) and tiller number (Fig. 4P) of WT and gif1-de plants did not change significantly, the seed setting rate, yield per plant, and grain yield per plot of gif1-de were 26%, 26%, and 25% higher than those of WT, respectively (Fig. 4U, 4V, and 4W). Remarkably, the precise insertion of HSE into the GIF1 promoter could rescue 40.9% of the yield loss caused by high temperature (Fig. 4L and 4V). Further investigation into the reasons for the increased yield of gif1-de under high temperature found that the number of secondary branches increased significantly (Fig. 4S), leading to a significant increase in the number of grains per panicle (Fig. 4T). Consistently, previous studies have shown that secondary and higher-order branches are more prone to flower and seed abortion than other parts of the rice panicle, due to inefficient carbon partitioning, independent of carbon source capacity (Seki et al., 2015; Lauxmann et al., 2016). In addition, the seed setting rate (Fig. 4U) and 1000-grain weight (Fig. 4Q) of gif1-de increased by 10.5%, 11.8%, and 4.86% compared with WT, respectively. In summary, the results showed that the application of CROCS strategy in rice not only increased grain yield under favorable conditions, but also reduced yield loss under adverse conditions.
[0416] Example 5, CROCS strategy rapidly creates different varieties of rice with high yield under favorable conditions and stable yield under adverse conditions
[0417] The inventors selected three recently approved and widely planted varieties, Wuyou 4, Longjing 31, and Zhongkefa 5, from elite rice varieties, and analyzed the expression of GIF1 under normal conditions and heat stress. By 2023, the annual planting area of the above three varieties in China was about 217,800 mu, 1,722,600 mu, and 178,400 mu, respectively.
[0418] The results are shown in Fig. 5. Heat stress significantly inhibited the expression of GIF1 in these three varieties, and the degree of heat inhibition was no different from ZH11, indicating that GIF1 has not been evolutionarily selected in high-quality modern varieties, and there is still a lot of room for improvement in the improvement of modern varieties. The CROCS strategy has broad adaptability for the targeted improvement of modern cultivated varieties with “stable yield under normal temperature and increased yield under adverse conditions”.
[0419] To investigate the broad applicability of the CROCS strategy, the inventors used the same prime editing construct targeting the GIF1 gene in ZH11 to knock in HSE into the GIF1 promoter of “Wuyou Rice-4”. Two independent TO lines (named wyd-4-gif1-de) were obtained, which have the desired precise insertion of HSE for GIF1 expression detection. The inventors found that, under normal conditions, the expression level of GIF1 in wyd-4-gif1-de was significantly higher than that in wild type (Figure 6A). Under heat stress, the HSE knock-in can alleviate the inhibition of heat stress on GIF1 expression, and the GIF1 expression is enhanced by 1.96 times compared with the wild type (Figure 6B), which is similar to the results in ZH11 (Figure 6C and Figure 6D). Preliminary yield determination results of TO generation of WT and wyd-4-gif1-de showed that the thousand-grain weight and harvest index of wyd-4-gif1-de were significantly improved compared with WT, thereby improving the yield per plant of rice (Figure 6E-H). These results and the yield data of ZH11 show that the rational design of GIF1 expression by CROCS provides an effective method to release its potential in crop improvement, which is a supplement to traditional breeding methods and creates valuable genetic variation that may be overlooked or underutilized in traditional breeding practices.
[0420] Example 6, Feasibility of CROCS breeding strategy in different tomato backgrounds
[0421] To verify the feasibility of the CROCS breeding strategy in different tomato backgrounds, m82-lin5-de mutants were generated in the M82 background. Fortunately, the inventors obtained homozygous mutants with precise insertion of HSE in TO plants. The expression of LIN5 under normal and heat conditions was evaluated. It was found that the expression of LIN5 in M82 was significantly inhibited by heat stress (Figure 7A). The expression of LIN5 in m82-lin5-de mutants was 1.5 times that of M82 wild type plants (Figure 7B). Notably, it was nearly 3 times higher than the wild type under heat stress in m82-lin5-de mutants (Figure 7C), which reproduced the effect observed in Alisa Craig.
[0422] Considering that modern commercial tomato varieties are mainly F1 hybrids, their parental materials are usually kept secret, it is necessary to extensively search for inbred lines that are still commercially planted to verify the broad applicability of the CROCS breeding strategy in modern tomato varieties. Fortunately, the inventors found a modern inbred line variety, Original Wonder 1 (YW1), which is still widely planted in the Beijing area. It is favored by consumers for its excellent flavor and is mainly grown in northern China, where it prefers warm climates. The inventors introduced a guide editing vector that generates lin5-de into YW1 to obtain precisely edited plants yw1-lin5-de. Expression analysis of LIN5 under normal and heat stress conditions showed that heat stress significantly inhibited the expression of LIN5 in YW1 (Figure 7D). However, after targeting insertion of HSE, the expression of LIN5 under normal conditions was slightly increased, and the expression of LIN5 in yw1-lin5-de gene edited plants was activated by heat induction, with nearly two-fold increase in expression (Figures 7E and 7F), which is similar to the results of M82 and Alisa Craig. The results of these three varieties, which are suitable for different climates, have different growth habits, and come from different countries, further confirm the reliability and universal applicability of the CROCS breeding strategy.
[0423] Example 7. Verification of the CROCS breeding strategy by different cis-elements and different yield-related genes
[0424] The above examples obtained plants with high yield in favorable conditions and stable yield in adverse conditions by inserting HSE (ATTCTAGAAT) into the CWIN gene (tomato LIN5, rice GIF1) promoter in different tomato and rice varieties. This example further inserts cis-elements that respond to different stresses into different yield-related genes to further confirm the reliability and universal applicability of the breeding strategy of the present application in vivo and in vitro.
[0425] 7.1. Tobacco transient verification experiment
[0426] 1) Insertion of heat-responsive elements into the promoters of LIN5 homologous genes in different plants
[0427] The promoters of LIN5 homologous genes that are highly expressed in corn, soybean, and wheat grains were analyzed, and positions 485 bp, 360 bp, and 421 bp away from the first exon were selected as insertion sites for HSE (ATTCTAGAAT) in the promoters of corn, soybean, and wheat, respectively. Rapid in vitro verification was performed by a tobacco in vitro transient expression system.
[0428] Results are shown in Figure 8. For the LIN5 homologous gene promoters of corn, soybean and wheat, the insertion of HSEs did not show significant difference compared with the wild type promoter without inserted elements under normal conditions, but could respond to heat (40°C, 4h) to drive LUC to be induced and significantly up-regulated under heat treatment (Figure 8A-F).
[0429] 2) Different elements inserted into the tomato LIN5 promoter
[0430] In addition to the above-mentioned HSE (ATTCTAGAAT, named HSE-1) inserted into the tomato, rice and different species, the applicant also selected the heat responsive element HSE-2 (GTTCATGAAC), the heat responsive element HSE-3 (AGAACGTTCT), the light responsive element LRE-1 (TGTGTGGTTAATATGAAGATAAGATT), the drought responsive element DRE (CATGTG), and the flood responsive element FRE (TTGACC) to be inserted into the position of 452bp upstream of the tomato LIN5 promoter, and transient expression in tobacco was carried out for rapid verification.
[0431] The results show that the LIN5 promoter inserted with HSE-2 can respond to heat (40°C, 4h) to drive LUC to be induced and significantly up-regulated, and the LIN5 promoter inserted with HSE-3 can drive LUC to be up-regulated 1.25 times compared with the control under normal conditions, and significantly up-regulated 1.95 times under heat treatment (Figure 9A-D).
[0432] LRE-1 can endow the LIN5 promoter with the ability to significantly respond to light, and can induce LUC gene to be up-regulated 4 times compared with the wild type promoter without inserted elements under 140μmol light intensity and 4h continuous light (Figure 9E, F).
[0433] After 4h continuous waterlogging, the LIN5 promoter inserted with FRE can respond to waterlogging to induce LUC to be significantly up-regulated (Figure 9K, L).
[0434] Using 300mM Mannitol to simulate drought conditions, and continuous treatment for 4h, the insertion of DRE enables the LIN5 promoter to respond to drought to drive LUC to be significantly up-regulated (Figure 91, J).
[0435] 3) Light responsive element inserted into the promoter of tomato SRG1 gene
[0436] In addition, the light responsive element LRE-2 (GTGTGTGAA) was also inserted into the promoter of the tomato SRG1 gene to drive LUC to be expressed in tobacco. The experimental method is similar to the above. The tomato SRG1 is a plant type related gene, and increasing the expression of SRG1 can reduce the number of lateral branches, so as to achieve the purpose of increasing yield by reducing sink energy consumption.
[0437] The results show that the insertion of LRE-2 can endow the SRG1 promoter with the ability to respond to light significantly (Fig. 9G, H).
[0438] The above three sets of experimental results verify that the insertion of different environmental response elements can endow the promoters of different target genes of different species with the new ability to respond to the environment specifically.
[0439] 7.2. In vivo verification experiments of tomato and rice
[0440] In order to verify the universality of the CROCS breeding strategy in vivo, in tomato and rice, different environmental response elements were precisely knocked into different genes by using the guide editing system, and the expression and phenotype of the improved genes were analyzed and observed in vivo. The specific elements, target genes and precise knock-in frequencies are shown in Fig. 10A.
[0441] Firstly, the tomato auxin synthesis gene FZY6 was selected, which encodes a flavin monooxygenase that can catalyze the direct conversion of indole pyruvic acid into IAA, and is a rate-limiting enzyme in the tryptophan-dependent indole acetic acid biosynthesis pathway, playing a key role in maintaining the content of auxin (Liu et al., 2016). Using the improved guide editing system of the present application, a heat response element HSE (ATTCTAGAAT) was inserted at the position 491 bp upstream of the promoter, and the precise editing plant fzy6-de was obtained, with a precise editing efficiency of 11.76% (Fig. 10A and B). FZY6 is mainly expressed in tomato flower buds and young fruits, and also expressed in roots. In order to quickly verify the expression level change of FZY6 in vivo, we took the roots for expression determination, and the qRT-PCR expression results showed that the expression level of FZY6 in the roots of fzy6-de was 1.27 times that of the wild type under normal conditions, and 1.42 times that of the wild type under heat treatment (Fig. 10C, D). Phenotype observation found that after heat treatment (35℃, 2 days), the root length of fzy6-de was significantly shorter than that of WT (Fig. 10E, F). It has been reported that auxin at a certain concentration can inhibit root elongation, which implies that the knock-in of HSE can intelligently regulate the synthesis level of auxin mediated by FZY6 in different tissues of plants, meaning that by changing the temperature, which is an environmentally friendly way, the synthesis level of endogenous auxin can be regulated to improve the fruit setting rate and yield, without the need to use external application of auxin analogues which can easily cause environmental pollution to improve the yield of horticultural crops such as tomato and grape.
[0442] SOE belongs to the chlorophyll-binding protein gene family, and the protein encoded by this gene is an important component of photosynthesis, mainly involved in the assembly of photosystem and the capture and transmission of light energy. It is a gene with "source increasing" potential that can promote plant photosynthesis (Velez-Ramirez et al., 2014). The gene can optimize the light energy utilization efficiency of tomato, reduce photoinhibition and photodamage, and thus support the normal growth of plants under continuous light conditions. The present application successfully inserted the light response element LRE (AAATTGTGA) into the domestication site of the 5'UTR of the SOE gene, and obtained the precise editing plant soe-de, and the precise editing efficiency was 6.67% (Figure 10A). The wild type tomato and soe-de were treated with continuous light (120 μmol, 2 weeks), and soe-de was more resistant to continuous light than the wild type, and the leaf chlorophyll content was higher (Figures 10G-I). Continuous light and increased chlorophyll content promoted the photosynthetic capacity of tomato, and the photosynthate could continuously generate energy supply for the development of sink organs, thereby improving crop yield.
[0443] In addition, a drought response element DRE (TACCGACAT) was knocked into a gravity stimulus-related gene GRG1 in rice (Zhao et al., 2018), and the editing efficiency of the precise editing material grg1-de was as high as 15% (Figure 10A). Using mannitol to simulate drought conditions, after 2 and 8 h of treatment on solid medium containing 200 mM mannitol, 1 cm of root tip was taken, and the expression level of GRG1 was detected. The results showed that compared with the control, the expression level of GRG1 with DRE element knocked in was increased by 1.8 times under 2 h of drought stress, and by 1.5 times under 4 h of stress. The knock-in of DRE endowed GRG1 with a new ability to respond to drought in vivo. The results of root length observation showed that the knock-in of DRE promoted the growth of GRG1 roots under drought conditions, and the elongation of root system helped to better absorb water and nutrients from deep soil, so that the wild type rice could be more drought-resistant under drought conditions, suggesting the potential to recover the loss of rice yield caused by drought, and providing materials for creating water-saving rice germplasm (Figures 10J, K).
[0444] Example 8, Promoter cis-acting elements have a hybrid dominant effect
[0445] The hybrid dominance effect is gradually overturning the traditional theoretical framework of trait inheritance in maize hybrid breeding. For a long time, the genetic basis of hybrid advantage has focused on the interaction between coding genes and dominance, while the hybridization effect of cis-regulatory elements has been ignored (Bertolini et al., 2025). In recent years, research has found that the hybridization of parent promoters in maize hybrids can drive the nonlinear increase in downstream gene expression levels (such as superdominant expression) through the functional complementation and synergy of cis-acting elements, thereby breaking through the limitations of single environmental adaptability under homozygous conditions (Xiao et al., 2021). This effect not only manifests in the enhancement of independent traits such as photosynthesis and stress resistance, but also achieves the coordinated optimization of multiple traits (such as the coupling phenotype of "high photosynthetic efficiency-drought resistance-high grain-filling efficiency") through the dynamic reconstruction of regulatory networks. Advances in molecular biology techniques (such as single-cell sequencing and three-dimensional genomics) have further revealed that hybrid promoters can remodel chromatin spatial rearrangement and epigenetic modification to form "molecular switch" effects, exhibiting more flexible transcriptional regulation capabilities in environmental signal responses. Although this field has made breakthroughs in model promoter research, the universality of hybrid promoters in complex agronomic traits, the dose sensitivity of hybridization effects, and the stability across generations still need to be systematically analyzed. In the future, with the cross-penetration of synthetic biology and artificial intelligence technology, designing "modular" hybrid promoters and accurately predicting their phenotypic outputs will become the core engine for the construction of intelligent breeding systems, opening up new paths for the efficient utilization of maize and even crop hybrid advantages.
[0446] This study found that promoter precision editing mutants in tomato and rice varieties still have the ability to significantly up-regulate gene expression in a hybrid state (Fig. 11A-D), which implies that creating excellent paternal varieties through the CROCS strategy and crossing them with excellent maternal varieties can quickly achieve hybrid seed production.
[0447] Discussion
[0448] Since the source-sink theory was proposed in 1928 (Mason and Maskell, 1928), it has played a crucial role in the basic research of plant physiology and developmental biology (Ruan et al., 2010; Lemoine et al., 2013; Rodrigues et al., 2019; Fernie et al., 2020) and agricultural production and food security (Chang and Zhu, 2017; Smith et al., 2018; Aluko et al., 2021), especially through the optimization of source-sink relationship by cultivation management, which has been widely used in agricultural production and yield improvement. With the deepening of the understanding of the regulation mechanism of source-sink relationship and the development of gene engineering technologies such as overexpression and gene silencing or knockout using CRISPR / Cas9, specific source-sink related genes have been used to improve crop yield (Wang et al., 2008; Li et al., 2013; Nam et al., 2022; Singh et al., 2022). Although these methods have successfully improved the yield of some crops, they often disrupt the balance of source-sink relationship and cause negative effects (Dickinson et al., 1991; Sun et al., 2014). In addition, many source-sink related genes need to be moderately up-regulated to function (Nam et al., 2022), and there is still a lack of effective gene editing tools to achieve higher expression levels. Therefore, the reasonable design of source-sink balance and the optimization of climate response have not been achieved. With the intensification of the impact of global climate change, it is urgent to overcome this challenge in a limited time (Lobell et al., 2008; Razzaq et al., 2021).
[0449] In this study, the inventors developed a CROCS strategy and established an efficient prime editing system that can precisely insert stress-responsive elements into the promoter regions of important source-sink related genes, enabling these genes to respond to stress by increasing their expression without disrupting their expression patterns. This precisely reverses the inhibition of carbon allocation by stress, improving the yield of different crops under different cultivation modes, thereby greatly reducing the yield loss caused by stress. The strategy of the present invention realizes the rapid and accurate optimization of source-sink relationship, achieving the long-term goal of agricultural production: high yield under favorable conditions and stable yield under adverse conditions. It provides a new breeding strategy for optimizing reasonable source-sink design and improving crop yield in the future. In the present invention, the strategy is verified in vivo and in vitro by combining multiple different stress-responsive elements, multiple different target genes, multiple different crop plants or different cultivation species, fully demonstrating the reliability and universal applicability of the breeding strategy of the present invention.
[0450] Plants have retained many developmental plasticity features during evolution as survival strategies to adapt to the competition in natural ecosystems. However, a considerable portion of these traits are too sensitive to environmental changes, which are not ideal traits for crops in agricultural ecosystems and often lead to yield loss (Shen et al., 2023). In addition to the carbon allocation inhibition caused by environmental stress, there are also some traits in crops that will quickly respond to weak light or water deficiency. For example, tomatoes grown in greenhouses often encounter weak light in winter, which leads to rapid stem growth and causes the plant to grow too tall. Short-term water deficiency can also trigger overreaction of crops, such as flower and fruit drop (Sawicki et al., 2015; Reichardt et al., 2020). Therefore, if these overly sensitive environmental adaptation mechanisms can be appropriately weakened in agricultural ecosystems where light or water can be timely supplemented, yield loss can be greatly reduced. The CROCS breeding strategy of inserting environmental response elements in the promoter region of key genes provides a practical method to solve this problem.
[0451] With the rapid development of multi-omics, a large number of structural variations in promoter regions, especially variations in cis-regulatory elements, have been discovered (Alonge et al., 2020; Wang et al., 2020). These variations are selected by domestication and help to improve crop yield traits, quality traits, and stress resistance. However, these natural structural variations often exist in specific genetic backgrounds, and breeders must invest a lot of time and effort to introduce beneficial structural variations into different breeding germplasm resources. The CROCS breeding strategy of the present invention overcomes this limitation, directly introducing cis-regulatory elements into specific genetic backgrounds, thereby enabling the rapid generation of ideal regulatory variants and cleverly responding to environmental changes. This strategy can accurately and quickly design gene expression of important agronomic traits, enabling it to have the ability to intelligently adapt to environmental changes. This strategy provides a novel and widely applicable breeding method that can meet the urgent need to cultivate climate-smart crops to cope with global climate change. It also provides an effective gene editing tool and a feasible operating procedure for basic scientific research on plant stress response. More importantly, in addition to customizing high-yield excellent varieties by optimizing source-sink relationships and stabilizing yield under stress conditions, the breeding strategy of the present invention can also quickly endow plants with environmental adaptability and developmental plasticity by inserting climate-responsive cis-regulatory elements. The latter enables plants to reshape crop structure and nutrient uptake efficiency by sensing environmental changes. This greatly reduces the demand for human labor, fertilizers, pesticides, and energy in agricultural production, and greatly reduces the large amount of resources wasted in cultivation and management. This breeding strategy has the potential to improve the profitability of farms while reducing the impact on the environment. More broadly, the introduction of climate-responsive elements will drive the direct evolution of plant gene regulatory sequences, endowing genes with new spatiotemporal expression patterns. This in turn restructures the molecular network of gene expression to cope with environmental changes, providing a new perspective for studying the molecular relationship between genotype, phenotype, and environment and their adaptation to ecosystems.
[0452] The patent documents or non-patent documents mentioned in this text are incorporated herein by reference in their entirety.
[0453] REFERENCES
[0454] Alonge, M., Wang, X., Benoit, M., Soyk, S., Pereira, L., Zhang, L., Suresh, H., Ramakrishnan, S., Maumus, F., Ciren, D., et al. Major impacts of widespread structural variation on gene expression and crop improvement in tomato. Cell 2020: 182(1): 145-161 e123.
[0455] Aluko, O.O., Li, C., Wang, Q., and Liu, H. Sucrose utilization for improved crop yields: A Review Article. Int. J. Mol. Sci. 2021: 22(9): 4704.
[0456] Amthor, J.S. The mccree-de wit-penning de vries-thornley respiration paradigms: 30 years later. Ann. Bot. 2000: 86(1): 1-20.
[0457] Anzalone, A.V., Randolph, P.B., Davis, J.R., Sousa, A.A., Koblan, L.W., Levy, J.M., Chen, P.J., Wilson, C., Newby, G.A., Raguram, A., et al. Search-and-replace genome editing without double-strand breaks or donor DNA. Nature 2019: 576(7785): 149-157.
[0458] Arce, D., Spetale, F., Krsticevic, F., Cacchiarelli, P., Las Rivas, J., Ponce, S., Pratta, G., and Tapia, E. Regulatory motifs found in the small heat shock protein (sHSP) gene family in tomato. BMC Genom. 2018: 19 (Suppl 8): 860.
[0459] Arshad, M. S., Farooq, M., Asch, F., Krishna, J. S. V., Prasad, P. V. V., and Siddique, K. H. M. Thermal stress impacts reproductive development and grain yield in rice. Plant physiology and biochemistry : PPB 2017: 115 57-72.
[0460] Balfagon, D., Zandalinas, S. I., Mittler, R., and Gomez-Cadenas, A. High temperatures modify plant responses to abiotic stress conditions. Physiol. Plant 2020: 170 (3): 335-344.
[0461] Battisti, D. S., and Naylor, R. L. Historical Warnings of Future Food Insecurity with Unprecedented Seasonal Heat. Science (New York, N.Y.) 2009: 323 (5911): 240-244.
[0462] Bertolini, E., Rice, B. R., Braud, M., Yang, J., Hake, S., Strable, J., Lipka, A. E., and Eveland, A. L. Regulatory variation controlling architectural pleiotropy in maize. Nature Communications 2025: 16(1): 2140.
[0463] Bonner, J. J., Ballou, C., and Fackenthal, D. L. Interactions between DNA-bound trimers of the yeast heat shock factor. Mol. Cell Biol. 1994: 14(1): 501-508.
[0464] Chang, T.-G., and Zhu, X.-G. Source-sink interaction: a century old concept under the light of modern molecular systems biology. J. Exp. Bot. 2017: 68(16): 4417-4431.
[0465] Cheng, W. H., Taliercio, E. W., and Chourey, P. S. The Miniature 1 seed locus of maize encodes a cell wall invertase required for normal development of endosperm and maternal cells in the pedicel. Plant Cell 1996: 8(6): 971-983.
[0466] Cox, D. T. C., Maclean, I. M. D., Gardner, A. S., and Gaston, K. J. Global variation in diurnal asymmetry in temperature, cloud cover, specific humidity and precipitation and its association with leaf area index. Glob. Chang. Biol. 2020: 26(12): 7099-7111.
[0467] Dickinson, C. D., Altabella, T., and Chrispeels, M. J. Slow-growth phenotype of transgenic tomato expressing apoplastic invertase. Plant Physiol. 1991: 95(2): 420-425.
[0468] Fernie, A. R., Bachem, C. W. B., Helariutta, Y., Neuhaus, H. E., Prat, S., Ruan, Y. L., Stitt, M., Sweetlove, L. J., Tegeder, M., Wahl, V., et al. Synchronization of developmental, molecular and metabolic aspects of source-sink interactions. Nat. Plants 2020: 6(2): 55-66.
[0469] Fragkostefanakis, S., Simm, S., Paul, P., Bublak, D., Scharf, K.-D., and Schleiff, E. Chaperone network composition in Solanum lycopersicum explored by transcriptome profiling and microarray meta-analysis. Plant Cell Environ. 2015: 38(4): 693-709.
[0470] Fridman, E., Carrari, F., Liu, Y. S., Fernie, A. R., and Zamir, D. Zooming in on a quantitative trait for tomato yield using interspecific introgressions. Science (New York, N.Y.) 2004: 305(5691): 1786-1789.
[0471] Gao, C. Genome engineering for crop improvement and future agriculture. Cell 2021: 184(6): 1621-1635.
[0472] Gao, L., Gonda, I., Sun, H., Ma, Q., Bao, K., Tieman, D. M., Burzynski-Chang, E. A., Fish, T. L., Stromberg, K. A., Sacks, G. L., et al. The tomato pan-genome uncovers new genes and a rare allele regulating fruit flavor. Nat. Genet. 2019: 51(6): 1044-1051.
[0473] Ismail, A. M., and Hall, A. E. Reproductive-stage heat tolerance, leaf membrane thermostability and plant morphology in cowpea. Crop Sci. 1999: 39(6): 1762-1768.
[0474] Jiang, Y., Chai, Y., Qiao, D., Wang, J., Xin, C., Sun, W., Cao, Z., Zhang, Y., Zhou, Y., Wang, X. C., et al. Optimized prime editing efficiently generates glyphosate-resistant rice plants carrying homozygous TAP-IVS mutation in EPSPS. Mol. Plant 2022: 15 (11): 1646-1649.
[0475] Jin, Y., Ni, D. A., and Ruan, Y. L. Posttranslational elevation of cell wall invertase activity by silencing its inhibitor in tomato delays leaf senescence and increases seed weight and fruit hexose level. Plant Cell 2009: 21 (7): 2072-2089.
[0476] Khaipho-Burch, M., Cooper, M., Crossa, J., de Leon, N., Holland, J., Lewis, R., McCouch, S., Murray, S. C., Rabbi, I., Ronald, P., et al. Genetic modification can improve crop yields-but stop overselling it. Nature 2023: 621 (7979): 470-473.
[0477] Kumar, S., Stecher, G., Li, M., Knyaz, C., and Tamura, K. MEGA X: molecular evolutionary genetics analysis across computing platforms. Molecular biology and evolution 2018: 35 (6): 1547-1549.
[0478] Lafitte, H. R., and Travis, R. L. Photosynthesis assimilate partitioning in closely related lines of rice exhibiting different sink: source relationships. Crop Sci. 1984: 24(3): cropsci1984.0011183X002400030004x.
[0479] Lauxmann, M. A., Annunziata, M. G., Brunoud, G., Wahl, V., Koczut, A., Burgos, A., Olas, J. J., Maximova, E., Abel, C., Schlereth, A., et al. Reproductive failure in Arabidopsis thaliana under transient carbohydrate limitation: flowers and very young siliques are jettisoned and the meristem is maintained to allow successful resumption of reproductive growth. Plant Cell Environ. 2016: 39(4): 745-767.
[0480] Lemoine, R., La Camera, S., Atanassova, R., Dedaldechamp, F., Allario, T., Pourtau, N., Bonnemain, J. L., Laloi, M., Coutos-Thevenot, P., Maurousset, L., et al. Source-to-sink transport of sugar and regulation by environmental factors. Front. Plant Sci. 2013: 4272.
[0481] Li, B., Liu, H., Zhang, Y., Kang, T., Zhang, L., Tong, J., Xiao, L., and Zhang, H. Constitutive expression of cell wall invertase genes increases grain yield and starch content in maize. Plant Biotechnol. J. 2013: 11 (9): 1080-1091.
[0482] Li, Z., Palmer, W. M., Martin, A. P., Wang, R., Rainsford, F., Jin, Y., Patrick, J. W., Yang, Y., and Ruan, Y. L. High invertase activity in tomato reproductive organs correlates with enhanced sucrose import into, and heat tolerance of, young fruit. J. Exp. Bot. 2012: 63 (3): 1155-1166.
[0483] Liu, Y., Yang, G., Huang, S., Li, X., Wang, X., Li, G., Chi, T., Chen, Y., Huang, X., and Wang, X. Enhancing prime editing by Csy4-mediated processing of pegRNA. Cell Res. 2021: 31 (10): 1134-1136.
[0484] Liu, Y. H., Offler, C. E., and Ruan, Y. L. Cell wall invertase promotes fruit set under heat stress by suppressing ROS-independent cell death. Plant Physiol. 2016: 172 (1): 163-180.
[0485] Lobell, D. B., Burke, M. B., Tebaldi, C, Mastrandrea, M. D., Falcon, W. P., and Naylor, R. L. Prioritizing climate change adaptation needs for food security in 2030. Science (New York, N. Y.) 2008: 319(5863): 607-610.
[0486] Long, S. P., and Ort, D. R. More than taking the heat: crops and global change. Curr. Opin. Plant Biol. 2010: 13(3): 241-248.
[0487] Lü, P., Yu, S., Zhu, N., Chen, Y.-R., Zhou, B., Pan, Y., Tzeng, D., Fabi, J. P., Argyris, J., Garcia-Mas, J., et al. Genome encode analyses reveal the basis of convergent evolution of fleshy fruit ripening. Nat. Plants 2018: 4(10): 784-791.
[0488] Lu, Y., Tian, Y., Shen, R., Yao, Q., Zhong, D., Zhang, X., and Zhu, J. K. Precise genome modification in tomato using an improved prime editing system. Plant Biotechnol. J. 2021: 19(3): 415-417.
[0489] Marsh, J. I., Hu, H., Gill, M., Batley, J., and Edwards, D. Crop breeding for a changing climate: integrating phenomics and genomics with bioinformatics. Theor. Appl. Genet. 2021: 134(6): 1677-1690.
[0490] Marzec, M., and Hensel, G. Prime editing: game changer for modifying plant genomes. Trends Plant Sci. 2020: 25(8): 722-724.
[0491] Mason, T.G., and Maskell, E.J. Studies on the transport of carbohydrates in the cotton plant: II. the factors determining the rate and the direction of movement of sugars1. Ann. Bot. 1928: os-42(3): 571-636. Nam, H., Gupta, A., Nam, H., Lee, S., Cho, H.S., Park, C., Park, S., Park, S.J., and Hwang, I. JULGI-mediated increment in phloem transport capacity relates to fruit yield in tomato. Plant Biotechnol. J. 2022: 20(8): 1533-1545.
[0492] Naqvi, R.Z., Siddiqui, H.A., Mahmood, M.A., Najeebullah, S., Ehsan, A., Azhar, M., Farooq, M., Amin, I., Asad, S., Mukhtar, Z., et al. Smart breeding approaches in post-genomics era for developing climate-resilient food crops. Front. Plant Sci. 2022: 13972164.
[0493] Nelson, J. W., Randolph, P. B., Shen, S. P., Everette, K. A., Chen, P. J., Anzalone, A. V., An, M., Newby, G. A., Chen, J. C., Hsu, A., et al. Engineered pegRNAs improve prime editing efficiency. Nat. Biotechnol. 2022: 40(3): 402-410.
[0494] Nguyen, H. T., Mantelin, S., Ha, C. V., Lorieux, M., Jones, J. T., Mai, C. D., and Bellafiore, S. Insights into the genetics of the Zhonghua 11 resistance to meloidogyne graminicola and its molecular determinism in rice. Front. Plant Sci. 2022: 13.
[0495] Peng, S., Huang, J., Sheehy, J. E., Laza, R. C., Visperas, R. M., Zhong, X., Centeno, G. S., Khush, G. S., and Cassman, K. G. Rice yields decline with higher night temperature from global warming. Proc. Natl. Acad. Sci. USA 2004: 101(27): 9971-9975.
[0496] Pressman, E., Peet, M. M., and Pharr, D. M. The effect of heat stress on tomato pollen characteristics is associated with changes in carbohydrate concentration in the developing anthers. Ann. Bot. 2002: 90(5): 631-636.
[0497] Razzaq, A., Kaur, P., Akhter, N., Wani, S. H., and Saleem, F. Next-generation breeding strategies for climate-ready crops. Front. Plant Sci. 2021 : 12620420.
[0498] Reichardt, S., Piepho, H. P., Stintzi, A., and Schaller, A. Peptide signaling for drought-induced tomato flower drop. Science (New York, N.Y.) 2020: 367(6485): 1482-1485.
[0499] Ren, Y., Huang, Z., Jiang, H., Wang, Z., Wu, F., Xiong, Y., and Yao, J. A heat stress responsive NAC transcription factor heterodimer plays key roles in rice grain filling. J. Exp. Bot. 2021 : 72(8): 2947-2964.
[0500] Richter, K., Haslbeck, M., and Buchner, J. The heat shock response: life on the verge of death. Mol. Cell 2010: 40(2): 253-266.
[0501] Rizhsky, L., Liang, H., Shuman, J., Shulaev, V., Davletova, S., and Mittler, R. When defense pathways collide. The response of Arabidopsis to a combination of drought and heat stress. Plant Physiol. 2004: 134(4): 1683-1696.
[0502] Rodrigues, J., Inze, D., Nelissen, H., and Saibo, N.J.M. Source-sink regulation in crops under water deficit. Trends Plant Sci. 2019: 24(7): 652-663.
[0503] Ruan, Y.L. Sucrose metabolism: gateway to diverse carbon use and sugar signaling. Annu. Rev. Plant Biol. 2014: 6533-67.
[0504] Ruan, Y.L., Jin, Y., Yang, Y.J., Li, G.J., and Boyer, J.S. Sugar input, metabolism, and signaling mediated by invertase: roles in development, yield potential, and response to drought and heat. Mol. Plant 2010: 3(6): 942-955.
[0505] Sawicki, M., Barka, E., Clément, C., Vaillant-Gaveau, N., and Jacquard, C. Cross-talk between environmental stresses and plant metabolism during reproductive organ abscission. J. Exp. Bot. 2015: 66(7): 1707-1719.
[0506] Seki, M., Feugier, F. G., Song, X. J., Ashikari, M., Nakamura, H., Ishiyama, K., Yamaya, T., Inari-Ikeda, M., Kitano, H., and Satake, A. A mathematical model of phloem sucrose transport as a new tool for designing rice panicle structure for high grain yield. Plant Cell Physiol. 2015: 56(4): 605-619.
[0507] Shen, S., Ma, S., Wu, L., Zhou, S.-L., and Ruan, Y.-L. Winners take all: competition for carbon resource determines grain fate. Trends Plant Sci. 2023: 28(8): 893-901.
[0508] Singh, J., Das, S., Jagadis Gupta, K., Ranjan, A., Foyer, C. H., and Thakur, J. K. Physiological implications of SWEETs in plants and their potential applications in improving source-sink relationships for enhanced yield. Plant Biotechnol. J. 2022.
[0509] Smith, M. R., Rao, I. M., and Merchant, A. Source-sink relationships in crop plants and their influence on yield development and nutritional quality. Front. Plant Sci. 2018: 91889.
[0510] Sturm, A., and Tang, G. Q. The sucrose-cleaving enzymes of plants are crucial for development, growth and carbon partitioning. Trends Plant Sci. 1999: 4(10): 401-407.
[0511] Sun, L., Yang, D. L., Kong, Y., Chen, Y., Li, X. Z., Zeng, L. J., Li, Q., Wang, E. T., and He, Z. H. Sugar homeostasis mediated by cell wall invertase GRAIN INCOMPLETE FILLING 1 (GIF1) plays a role in pre-existing and induced defence in rice. Mol. Plant Pathol. 2014: 15(2): 161-173.
[0512] Suwa, R., Hakata, H., Hara, H., El-Shemy, H. A., Adu-Gyamfi, J. J., Nguyen, N. T., Kanai, S., Lightfoot, D. A., Mohapatra, P. K., and Fujita, K. High temperature effects on photosynthate partitioning and sugar metabolism during ear expansion in maize (Zea mays L.) genotypes. Plant physiology and biochemistry: PPB 2010: 48(2-3): 124-130.
[0513] Suzuki, N., Sejima, H., Tam, R., Schlauch, K., and Mittler, R. Identification of the MBF1 heat-response regulon of Arabidopsis thaliana. Plant J. 2011: 66(5): 844-851.
[0514] Tieman, D., Zhu, G., Resende, M. F., Jr., Lin, T., Nguyen, C., Bies, D., Rambla, J. L., Beltran, K. S., Taylor, M., Zhang, B., et al. A chemical genetic roadmap to improved tomato flavor. Science (New York, N.Y.) 2017: 355(6323): 391-394.
[0515] Vallarino, J. G., Yeats, T. H., Maximova, E., Rose, J. K., Fernie, A. R., and Osorio, S. Postharvest changes in LIN5-down-regulated plants suggest a role for sugar deficiency in cuticle metabolism during ripening. Phytochemistry 2017: 142 11-20.
[0516] Vaucheret, H., Beclin, C., Elmayan, T., Feuerbach, F., Godon, C., Morel, J. B., Mourrain, P., Palauqui, J. C., and Vernhettes, S. Transgene-induced gene silencing in plants. Plant J. 1998: 16(6): 651-659.
[0517] Velez-Ramirez, A. I., van Ieperen, W., Vreugdenhil, D., van Poppel, P. M. J. A., Heuvelink, E., and Millenaar, F. F. A single locus confers tolerance to continuous light and allows substantial yield increase in tomato. Nature Communications 2014: 5(1).
[0518] von Schaewen, A., Stitt, M., Schmidt, R., Sonnewald, U., and Willmitzer, L. Expression of a yeast-derived invertase in the cell wall of tobacco and Arabidopsis plants leads to accumulation of carbohydrate and inhibition of photosynthesis and strongly influences growth and phenotype of transgenic tobacco plants. The EMBO journal 1990: 9(10): 3033-3044.
[0519] Vu, T. V., Nguyen, N. T., Kim, J., Hong, J. C., and Kim, J. Y. Prime editing: Mechanism insight and recent applications in plants. Plant Biotechnol. J. 2023.
[0520] Wan, H., Wu, L., Yang, Y., Zhou, G., and Ruan, Y. L. Evolution of sucrose metabolism: the dichotomy of invertases and beyond. Trends Plant Sci. 2018: 23(2): 163-177.
[0521] Wang, E., Wang, J., Zhu, X., Hao, W., Wang, L., Li, Q., Zhang, L., He, W., Lu, B., Lin, H., et al. Control of rice grain-filling and yield by a gene with a potential signature of domestication. Nat. Genet. 2008: 40(11): 1370-1374.
[0522] Wang, L., and Ruan, Y.-L. New insights into roles of cell wall invertase in early seed development revealed by comprehensive spatial and temporal expression patterns of GhCWIN1 in cotton. Plant Physiol. 2012: 160 777-787.
[0523] Wang, L., Li, X. R., Lian, H., Ni, D. A., He, Y. K., Chen, X. Y., and Ruan, Y. L. Evidence that high activity of vacuolar invertase is required for cotton fiber and Arabidopsis root elongation through osmotic dependent and independent pathways, respectively. Plant Physiol. 2010: 154 (2): 744-756.
[0524] Wang, X., Gao, L., Jiao, C., Stravoravdis, S., Hosmani, P. S., Saha, S., Zhang, J., Mainiero, S., Strickler, S. R., Catala, C., et al. Genome of Solanum pimpinellifolium provides insights into structural variants during tomato breeding. Nat. Commun. 2020: 11 (1): 5817.
[0525] Wang, Y., Wang, L., Zhou, J., Hu, S., Chen, H., Xiang, J., Zhang, Y., Zeng, Y., Shi, Q., Zhu, D., et al. Research progress on heat stress of rice at flowering stage. Rice Sci. 2019: 26 (1): 1-10.
[0526] Waters, E. R. The molecular evolution of the small heat-shock proteins in plants. Genetics 1995: 141(2): 785-795.
[0527] Weber, H., Borisjuk, L., and Wobus, U. J. T. P. J. Controlling seed development and seed size in Vicia faba: a role for seed coat-associated invertases and carbohydrate state. Plant J. 1996: 10(5): 823-834.
[0528] Weschke, W., Panitz, R., Gubatz, S., Wang, Q., Radchuk, R., Weber, H., and Wobus, U. The role of invertases and hexose transporters in controlling sugar ratios in maternal and filial tissues of barley caryopses during early development. Plant J. 2003: 33(2): 395-411.
[0529] Wheeler, T., and von Braun, J. Climate change impacts on global food security. Science (New York, N.Y.) 2013: 341(6145): 508-513.
[0530] Xiao, H., and Lis, J. T. Germline transformation used to define key features of heat-shock response elements. Science (New York, N.Y.) 1988: 239(4844): 1139-1142.
[0531] Xiao, Y., Jiang, S., Cheng, Q., Wang, X., Yan, J., Zhang, R., Qiao, F., Ma, C., Luo, J., Li, W., et al. The genetic mechanism of heterosis utilization in maize improvement. Genome Biology 2021 : 22(1): 148.
[0532] Yan, J., Oyler-Castrillo, P., Ravisankar, P., Ward, C.C., Levesque, S., Jing, Y., Simpson, D., Zhao, A., Li, H., Yan, W., et al. Improving prime editing with an endogenous small RNA-binding protein. Nature 2024: 628(8008): 639-647.
[0533] Yan, W., Wu, X., Li, Y., Liu, G., Cui, Z., Jiang, T., Ma, Q., Luo, L., and Zhang, P. Cell wall invertase 3 affects cassava productivity via regulating sugar allocation from source to sink. Front. Plant Sci. 2019: 10541.
[0534] Yao, Q., Shen, R., Shao, Y., Tian, Y., Han, P., Zhang, X., Zhu, J.-K., and Lu, Y. Efficient and multiplex gene upregulation in plants through CRISPR-Cas-mediated knockin of enhancers. Molecular Plant 2024: 17(9): 1472-1483.
[0535] Yu, H., Yang, Q., Fu, F., and Li, W. Three strategies of transgenic manipulation for crop improvement. Front. Plant Sci. 2022: 13948518.
[0536] Zanor, M. I., Osorio, S., Nunes-Nesi, A., Carrari, F., Lohse, M., Usadel, B., Kuhn, C., Bleiss, W., Giavalisco, P., Willmitzer, L., et al. RNA interference of LIN5 in tomato confirms its role in controlling Brix content, uncovers the influence of sugars on the levels of fruit hormones, and demonstrates the importance of sucrose cleavage for normal fruit development and fertility. Plant Physiol. 2009: 150(3): 1204-1218.
[0537] Zhao, Y., Zhang, H., Xu, J., Jiang, C., Yin, Z., Xiong, H., Xie, J., Wang, X., Zhu, X., Li, Y., et al. Loci and natural alleles underlying robust roots and adaptive domestication of upland ecotype rice in aerobic conditions. PLoS Genet. 2018: 14(8): e1007521.
[0538] Part of the sequence information involved in the present application:
Claims
1. A method of producing a modified plant, said method comprising introducing one or more stress-responsive cis-elements into the expression regulatory sequence, such as the promoter, of one or more yield-related genes, preferably endogenous source-sink relationship-related genes, of said plant.
2. The method of claim 1, wherein said stress is a biotic stress or an abiotic stress, for example said abiotic stress is selected from the group consisting of heat stress, cold stress, osmotic stress (such as drought stress, salt stress, metal stress), water stress (flooding), light stress (insufficient or excessive light), nutrient stress (deficiency or excess of a nutrient element in the soil, such as nitrogen deficiency stress, phosphorus deficiency stress), mechanical stress (wind, hail, etc. physical damage); said biotic stress is selected from the group consisting of fungal infection, bacterial infection, viral infection, parasitic infection, insect stress, weed competition, parasitic plant, Preferably, said stress is heat stress.
3. The method of claim 1 or 2, wherein said stress-responsive cis-element comprises a cis-element selected from Table 1, preferably said stress-responsive cis-element is a heat-responsive element, more preferably said heat-responsive element comprises the sequence ATTCTAGAAT.
4. The method of any one of claims 1-3, wherein the source-sink relationship-related gene is selected from the group consisting of a gene encoding a cell wall invertase (CWIN), a sucrose transporter (SUT), a hexose transporter (HT), a SWEET (Sugars Will Eventually be Exported Transporter), a sucrose-phosphate synthase (SPS), a sucrose synthase (SUS), a Glucan-water dikinase (GWD), a Source activity enhanced (SOE).
5. The method of any one of claims 1-3, wherein the source-sink relationship-related gene is selected from the genes of Table 2, or said source-sink relationship-related gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to a protein encoded by a gene selected from Table 2. Alternatively, the coding sequence of said source-sink relationship-related gene has at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity to the coding sequence of a gene selected from Table 2.
6. The method of any one of claims 1-5, wherein the introduction of said stress-responsive cis-element results in the expression of said source-sink relationship-related gene in said plant being up- or down-regulated, preferably up-regulated, in response to the corresponding stress.
7. The method of any one of claims 1-6, wherein the one or more stress-responsive cis-elements are introduced into the promoter at a predetermined site, preferably, the predetermined site is selected from the group consisting of: 1) not located within an endogenous cis-element of the gene; 2) located within open chromatin; and / or 3) proximal to the start codon of the gene.
8. The method of claim 7, wherein the selection of the predetermined site comprises: a) selecting a position within the promoter that does not contain a possible cis-element as a candidate site based on Plant CARE (PlantCARE, a database of plant promoters and their cis-acting regulatory elements (ugent.be), an online cis-element prediction website; b) performing DNase I hypersensitivity prediction on the candidate sites obtained from step a) by DHS (DNase-I hypersensitive sites) online prediction website (http: / / www.epigenome.cuhk.edu.hk / ) to screen out the candidate sites with peak value less than 0.5; and c) selecting the site closest to the start codon of the gene from the candidate sites obtained from step b) as the final predetermined site.
9. The method of claim 7, wherein the predetermined site is within about 2 kb, preferably within about 1 kb, more preferably within 500 bp, from the translation start codon of the gene.
10. The method of any one of claims 1-9, wherein a plurality (e.g., 2 to about 10 or more) of stress-responsive cis-elements are introduced into the promoter, for example, the plurality of stress-responsive cis-elements are introduced into the promoter in tandem.
11. The method of any one of claims 1-10, wherein the stress-responsive elements inserted at the predetermined site are verified to be capable of conferring stress-responsive ability by in vitro expression of a reporter system, for example, in a tobacco cell expression reporter system, prior to introducing the stress-responsive cis-elements into the plant.
12. The method of any one of claims 1-11, wherein the introduction of the one or more stress-responsive cis-elements is achieved by genome editing, for example, introducing a genome editing system into the plant.
13. The method of claim 12, wherein the genome editing system is a CRISPR, ZFN or TALEN based genome editing system; preferably, the genome editing system is a CRISPR based genome editing system.
14. The method of claim 12 or 13, wherein the method comprises 1) introducing into the plant a genome editing system targeting the predetermined site, which results in a genomic double-strand break (DSB) at or near the predetermined site; 2) introducing a homologous recombination donor nucleic acid comprising the one or more stress-responsive cis-elements and homologous arm sequences corresponding to the two sides of the DSB, whereby the one or more stress-responsive cis-elements are introduced into the predetermined site by homologous recombination. 15. The method of claim 12 or 13, wherein the introduction of the one or more stress- responsive cis-elements is achieved by introducing into the plant a prime editing system.
16. The method of claim 15, wherein the prime editing system comprises: i) a fusion protein comprising a CRISPR nickase and a reverse transcriptase and / or an expression construct containing a nucleotide sequence encoding the fusion protein; and ii) a pegRNA and / or an expression construct containing a nucleotide sequence encoding the pegRNA, wherein the pegRNA comprises, in the 5’ to 3’ direction, a guide sequence, a gRNA scaffold sequence, a reverse transcription (RT) template sequence, and a primer binding site (PBS) sequence, wherein the pegRNA is capable of forming a complex with the fusion protein and targeting the fusion protein to a first target sequence comprising or adjacent to the predetermined site in the genome, resulting in a first nick within the first target sequence, wherein the reverse transcription template sequence comprises a nucleotide sequence of the one or more stress-responsive cis-elements or a complement thereof, preferably, the prime editing system further comprises iii) a nick gRNA and / or an expression construct containing a nucleotide sequence encoding the nick gRNA, the nick gRNA comprising a guide sequence and a gRNA scaffold sequence; more preferably, the prime editing system further comprises iv) a Csy4 protein and / or an expression construct containing a nucleotide sequence encoding the Csy4 protein, and the pegRNA and / or the 3’ end of the nick RNA comprises a Csy4 recognition site sequence on one side or both sides of the 3’ and 5’ ends, more preferably, the pegRNA further comprises a tevopreQl motif at the 3’ end.
17. The method of claim 16, wherein the CRISPR nickase is a Cas9 nickase, for example, the Cas9 nickase comprises an amino acid sequence set forth in SEQ ID NO: 12 or 26.
18. The method of claim 16 or 17, wherein the reverse transcriptase is an M-MLV reverse transcriptase or a functional variant thereof, for example, the reverse transcriptase comprises an amino acid sequence set forth in SEQ ID NO:
13.
19. The method of any one of claims 16-18, wherein the CRISPR nickase and the reverse transcriptase in the fusion protein are connected by a linker, for example, the linker can be a linker set forth in SEQ ID NO: 14 (33 aa linker).
20. The method of any one of claims 16-19, wherein the fusion protein further comprises one or more nuclear localization sequences (NLS), for example, the NLS is an SV40 NLS (amino acid sequence set forth in SEQ ID NO: 15); and / or the fusion protein further comprises a LA polypeptide at the C-terminus, for example, a LA polypeptide comprising an amino acid sequence set forth in SEQ ID NO 27. 21. The method of any one of claims 16-20, wherein the guide sequence (also referred to as seed sequence or spacer sequence) in the pegRNA is configured to have sufficient sequence identity (preferably 100% identity) to the first target sequence to enable sequence-specific targeting by base pairing to the complementary strand of the first target sequence.
22. The method of any one of claims 16-21, wherein the scaffold sequence of the gRNA is set forth in SEQ ID NO:
16.
23. The method of any one of claims 16-22, wherein the primer binding sequence is configured to be complementary to at least a portion of the first target sequence, preferably the primer binding sequence is complementary to at least a portion of a 3' overhang single strand resulting from the nick in the sense strand of the first target sequence, in particular the primer binding sequence is complementary to the nucleotide sequence at the 3' end of the 3' overhang single strand.
24. The method of any one of claims 16-23, wherein the RT template sequence is configured to be complementary to at least a portion of the sequence downstream of the nick in the first target sequence, and comprises the nucleotide sequence of the one or more stress-responsive cis-elements or a complement thereof.
25. The method of any one of claims 16-24, wherein the nicking gRNA does not comprise a reverse transcription (RT) template sequence and a primer binding site (PBS) sequence, and the guide sequence (also referred to as seed sequence or spacer sequence) in the nicking gRNA is configured to have sufficient sequence identity (preferably 100% identity) to a second target sequence in the genome to enable the fusion protein to target the second target sequence and cause a second nick within the second target sequence, the second target sequence being on the opposite strand of the genomic DNA from the first target sequence.
26. The method of claim 25, wherein the second nick is upstream or downstream of the first nick, and the first nick and the second nick are separated by about 1 to about 300 or more nucleotides.
27. The method of any one of claims 16-26, wherein the Csy4 protein comprises the amino acid sequence set forth in SEQ ID NO: 17; and / or, the Csy4 recognition site comprises the nucleotide sequence set forth in SEQ ID NO:
18.
28. The method of any one of claims 16-27, wherein the prime editing system comprises a first expression construct encoding a fusion protein comprising, from N- to C-terminus: a Csy4 protein - a self-cleaving peptide - NLS - a CRISPR nickase - NLS - a linker - a reverse transcriptase - NLS, or a Csy4 protein - a self-cleaving peptide - NLS - a CRISPR nickase - NLS - a linker - a reverse transcriptase - a linker - a LA polypeptide - NLS; and a second expression construct comprising: a Csy4 recognition site sequence - a pegRNA coding sequence - a Csy4 recognition site sequence - a nicking gRNA coding sequence - a Csy4 recognition site sequence, or a Csy4 recognition site sequence - a pegRNA coding sequence - a Csy4 recognition site sequence - a tevopreQl motif - a nicking gRNA coding sequence - a Csy4 recognition site sequence.
29. The method of claim 28, wherein the first expression construct drives expression of the fusion protein by a 35S promoter, and the second expression construct drives expression by a CmYLCV promoter.
30. The method of claim 28 or 29, wherein the first expression construct comprises the nucleotide sequence set forth in SEQ ID NO: 19 or 29.
31. The method of any one of claims 1-30, wherein the plant is a crop plant, for example selected from the group consisting of Solanum lycopersicum (tomato), Nicotiana benthamiana (tobacco), Capsicum annuum (pepper), Physalis pruinosa (ground cherry), Solanum melongena (eggplant), Solanum tuberosum (potato), Solanum pennellii (Pennell's tomato), Solanum chilense (Chilean tomato), Solanum habrochaites (Habrochaites tomato), Solanum pimpinellifolium (Pimpinellifolium tomato), Solanum galapagense (Galapagos tomato), Petunia hybrid (petunia), grape, strawberry, Citrullus lanatus (watermelon), Cucumis sativus (cucumber), Lactuca sativa L. (lettuce), Chinese cabbage, oilseed rape, Brassica oleracea, wheat, Oryza sativa L. (rice), maize, soybean, sunflower, sorghum, oilseed rape, alfalfa, cotton, barley, millet, sugarcane, and cassava.
32. The method of any one of claims 1-31, wherein the source-sink relationship associated gene is a cell wall invertase (CWIN) encoding gene, preferably the cell wall invertase (CWIN) encoding gene is a tomato LIN5 gene or a homologous gene thereof, such as a rice GIF1 gene, a maize MN1 gene, a soybean Glyma.10G074800 gene, or a wheat TraesCS2A03G0736600 gene; or the gene is a tomato SRG1, FZY6, SOE, or a rice GRG1 gene.
33. The method of claim 32, wherein i) the LIN5 gene encodes a LIN5 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 1; or the LIN5 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 2; ii) the GIF1 gene encodes a GIF1 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99% sequence identity, even 100% to the amino acid sequence set forth in SEQ ID NO: 3; or the GIF1 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 4; iii) the MN1 gene encodes a MN1 protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 5; or the MN1 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 6; or iv) the TraesCS2A03G0736600 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 7; or the TraesCS2A03G0736600 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 8; or v) the Glyma.10G074800 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 24; or the Glyma.10G074800 gene comprises a coding sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 25; or vi) the SRG1 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 42; or vii) the FZY6 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 43; or viii) the SOE gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO: 44; or ix) the GRG1 gene encodes a protein having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the amino acid sequence set forth in SEQ ID NO:
45.
34. The method of claim 32 or 33, wherein the promoter of the Solanum lycopersicum LIN5 gene comprises a nucleotide sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 9; or the promoter of the Solanum lycopersicum LIN5 gene comprises a nucleotide sequence having at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity to the nucleotide sequence set forth in SEQ ID NO: 9; or The promoter of the rice GIF1 gene comprises a nucleotide sequence having at most 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or even 100% sequence identity with the nucleotide sequence set forth in SEQ ID NO:
10.
35. The method of any one of claims 1-34, wherein the plant is a tomato, the stress is heat stress, the stress-responsive cis-element is a heat-responsive element, for example the heat-responsive element comprises the sequence ATTCTAGAAT, and the source-sink related gene is a tomato LIN5 gene.
36. The method of claim 35, wherein the heat-responsive element is introduced into the tomato LIN5 gene promoter at 410 bp upstream of the start codon.
37. The method of claim 36, wherein introduction of the stress-responsive cis-element results in the modified plant comprising a mutated LIN5 gene promoter set forth in SEQ ID NO:
20.
38. The method of any one of claims 1-34, wherein the plant is a rice, the stress is heat stress, the stress-responsive cis-element is a heat-responsive element, for example the heat-responsive element comprises the sequence ATTCTAGAAT, and the source-sink related gene is a rice GIF1 gene.
39. The method of claim 38, wherein the heat-responsive element is introduced into the rice GIF1 gene promoter at 427 bp upstream of the start codon.
40. The method of claim 39, wherein introduction of the stress-responsive cis-element results in the modified plant comprising a mutated GIF1 gene promoter set forth in SEQ ID NO:
21.
41. A modified plant obtained according to the method of any one of claims 1-40.
42. The modified plant of claim 41, having increased yield, for example increased fruit / seed set, increased fruit / seed weight, in the presence and / or absence of the stress, as compared to an unmodified wild type plant.
43. The modified plant of claim 41, having increased plot yield in the presence and / or absence of the stress, as compared to an unmodified wild type plant.
44. The modified plant of claim 42 or 43, the yield increase is about 5%, about 10%, about 15%, about 20%, about 30%, about 40%, about 50%, about 60%, about 70%, about 80%, about 90%, about 100% or more.
Citation Information
Patent Citations
Method for inserting exogenous sequence in genome at fixed point
CN117126876A
Methods and Compositions for Improvement in Seed Yield
US20160138038A1
Promoter, promoter control elements, and combinations, and uses thereof
WO2012009551A1