Replicons for precise genome editing in plants
Patent Information
- Authority / Receiving Office
- EP · EP
- Patent Type
- Applications
- Current Assignee / Owner
- INARI AGRICULTURE TECHNOLOGY INC
- Filing Date
- 2024-07-12
- Publication Date
- 2026-05-20
AI Technical Summary
Existing methods for transient expression of nucleic acids in plants face challenges in achieving sufficient copy numbers, limiting the efficiency of genome editing techniques like HDR-based methods.
Development of soybean mild mottle virus (SbMMV)-based replicons that include an exogenous nucleic acid and a SbMMV long intergenic region (LIR), enhancing the production of nucleic acids in plant host cells.
The SbMMV replicons demonstrate superior amplification of nucleic acids, improving the efficiency of genome editing techniques by enabling precise integration of repair templates into target sequences.
Smart Images

Figure US2024037927_23012025_PF_FP_ABST
Abstract
Description
[0001] REPLICONS FOR PRECISE GENOME EDITING IN PLANTS
[0002] CROSS REFERENCE TO RELATED APPLICATIONS
[0003] This application claims priority to U.S. Provisional Application No. 63 / 526,880, filed July 14, 2023, which is incorporated by reference in its entirety.
[0004] FIELD
[0005] This relates to agricultural biotechnology, more specifically compositions and methods for modifying plants.
[0006] INCORPORATION OF ELECTRONIC SEQUENCE LISTING
[0007] The Sequence Listing is submitted as an XML file named “Sequence. xml,” created on July 12, 2024, 33,209 bytes, which is incorporated by reference herein.
[0008] BACKGROUND
[0009] Transient expression of a nucleic acid is a convenient tool for plant molecular biology, as it is a simplified process that is less labor intensive than generating stable transgenic lines. However, achieving a sufficient amount of exogenous nucleic acids (e.g., sufficient copy number) can be difficult in transient systems. Thus, improved compositions and methods for producing large quantities of nucleic acids of interest in plant host cells are needed.
[0010] SUMMARY
[0011] Disclosed herein are soybean mild mottle virus (SbMMV)-based replicons that have superior production of nucleic acids of interest in plant host cells. The replicons are useful, for example, for producing a repair template for use in homology dependent repair (HDR)-based genome editing techniques, or expressing a nucleic acid of interest in a plant cell (e.g., expressing a gene). The SbMMV replicons include an exogenous nucleic acid and a SbMMV long intergenic region (LIR). The exogenous nucleic acid is any nucleic acid of interest that is exogenous to SbMMV. In some examples, the exogenous nucleic acid is a repair template or an expression cassette. In some examples, the SbMMV replicon includes a sequence encoding replication initiator protein (Rep) and RepA. In some examples, the expression cassette includes a gene and / or regulatory element that is expressed in the plant host cell. In some examples, the repair template includes from 5’ to 3’ : a left homology arm (LHA), an insert sequence, and a right homology arm (RHA). The LHA and RHA are complementary to a target sequence, respectively. In some examples, the insert sequence introduces an insertion, deletion, or substitution in a target nucleic acid upon homologous recombination.
[0012] Further disclosed arc nucleic acids or vectors including the SbMMV replicon. The nucleic acids or vectors can further include one or more regulatory elements (e.g., promoter, terminator, etc.), selectable markers (e.g., resistance marker or fluorescent marker), nucleases, or other elements (e.g., transfer DNA border sequences). In some examples, the vector is a T-DNA vector or a vector suitable for biolistic transformation. Also disclosed are cells including the disclosed nucleic acids or vectors, and plants or plant parts including such cells. In some examples, the cells, plants, or plant parts are soybean (Glycine max) cells, soybean plants, or soybean plant parts.
[0013] Methods of producing a nucleic acid in a plant cell, or modifying genetic material of a plant cell, are also disclosed. Such methods include transforming a plant cell with a nucleic acid or vector disclosed herein. In some examples, a method of modifying genetic material of a plant cell includes transforming a nucleic acid or vector including a SbMMV replicon including a repair template. In some examples, the method further includes introducing a nuclease to induce a double-stranded break in a target nucleic acid. In some examples, a method of producing a nucleic acid in a plant cell includes transforming a plant cell with a nucleic acid or vector including a SbMMV replicon including an expression cassette. In some examples, the plant cell is a soybean cell (Glycine max).
[0014] The foregoing and other features of this disclosure will become more apparent from the following detailed description of several aspects which proceeds with reference to the accompanying figures.
[0015] BRIEF DESCRIPTION OF THE FIGURES
[0016] FIG. 1 shows a diagram of pIN4612, pIN4801, and pIN4802 constructs (elements are not to scale).
[0017] FIG. 2 shows that pIN4612 produces superior GFP expression in soybean.
[0018] FIG. 3 shows a diagram of pIN4614 and pIN4612 constructs (elements are not to scale).
[0019] FIG. 4 shows an exemplary diagram of a SbMMV replicon (elements are not to scale).
[0020] FIG. 5 shows an exemplary diagram of a repair template (elements are not to scale).
[0021] SEQUENCES
[0022] The nucleic acid and amino acid sequences listed herein are shown using standard letter abbreviations for nucleotide bases and amino acids, as defined in 37 C.F.R. § 1.822. Only one strand of each nucleic acid sequence is shown, but the complementary strand is understood as included by any reference to the displayed strand. Unless otherwise stated, nucleic acid sequences in the text of this specification are given, when read from left to right, in the 5’ to 3’ direction. Nucleic acid sequences may be provided as DNA or as RNA, as specified; disclosure of one necessarily defines the other, as well as necessarily defines the exact complements.
[0023] SEQ ID NO: 1 an exemplary SbMMV long intergenic region (LIR) sequence.
[0024] ATCGGTAAAAGCAAATGTACCCCCAATTGCCCCCCCCTTCTAAAACTCTATACAATTGGGGGTA ATGGGGGTGAATATATACCTACTACTATTAAATTCTCTTTAGCGAGGATTTCAAATCCGCCACG TGTACAAAAGGCCATCCCTTATAATATTACAGGGATGGCCGCGCCCGCCCCCCCTTTATTGTGG CGCCCACATCTTTGATTTCAGCCAATCAGGAAGCAGCCTCAATCCTTATTTATCTGTTTCTCCAC TATAAACTTGTTGCGCAAGTTGGTATTTGAATTAAAG SEQ ID NO: 2 is an exemplary nucleic acid encoding SbMMV Rep / RepA.
[0025] TTAATAAAGATTGAATTTTATATCATATTTAATCTCGAAATGGGATACAATGACGAACTGCTCA
[0026] AGCAATACTTTGTTGATAGCAGTCTGGACACTTAAAATAGAAATTGAACCTATACTATTCAAGA
[0027] ACGACATCATTAAAAACTTGAAAGCTCGAAAAAATCTCCAGTTCTGACTCTGTGAGTGGAGAA
[0028] AGATCTTTAAATCCAGGAAGCATTTGTGAATCTCCAACTTCCTCTTGACACAGTTGTTGAACAT
[0029] TATTCGAATGTTCATCTCGTCGTATGGGAGACCCATCACCCCATAGTCTTCCCGGAGGAGTTTG
[0030] AAACAGAGGGGATTGTTCACCTCCCAGATATACACGCCACGCTCTAGTTCCGATTGTGTGAGTA
[0031] ACTCCCCTGTGCGAGAATCCATGGTCTGTGCAGTCGAGGTGAACGTATATAGAACACCCACAA
[0032] GGACAGTCTATGCGACGCCGTCTAGTCGCCCTATGACGCTTTGCGTCCCTATGTTGAGCCTTGA
[0033] TGTTCGGTGGAGAACAGTGGCTCTTCGAGGGTGTAGAAGATCGCATTTTTTAAAGCCCACTCCT
[0034] TAAGCGCTGCATTCGACTCCTCATCCATATATTCTTTATAGGAAGATCTGGGGCCAGGATTGCA
[0035] GAGGAAAATTGTTGGTATTCCACCTTTAATTTGAGTGGGCTTGCCGTACTTCACGTTGCTTTGCC
[0036] AGTCTCTTTGGGACCCCATGAATTCTTTCATATGCTTTAGATAGTGGGGATCTACGTCATCGAT
[0037] GACGTTGTACCAGGCATCGTTTGAATAGACCTTGGCGCTGAGATCTAAATGACCACATAGGTA
[0038] ATTGTGCCTACCAAGGCTTCTGGCCCACATCGTCTTCCCCGTCCTGGATTCACCTTCCAGGACG
[0039] ATGCTTATGGGTCTCATCGGCCGCGCAGCGGAAGCTTTCACATTATCTGATGCCCATTGAGACA
[0040] GAACTTCCGGCACATTGTTAAATTGTTCAACTTTAAAAGGCGATTTGTAAATCGTCGTTGGAGG
[0041] CGTAAAAATCTTATTGAGGTTGCTAATAACATTATGATACTGAAAAACATAATCTCGTGGAAGC
[0042] TGCTCTTTGATGATCTGTAAAGCAGCTTCAGCGGAACCAGCATTTAATGCCGACGCACATGCGT
[0043] CGTTAGCATTTTTGCAACCTCCTCTGGCACTTCGACCATCAATCTGGAAAGACCCCCATTGTAT
[0044] GGTGTCTCCGTCTTTGTCGATGTATGTTTTGACGTCCGAACTTGATTTAGCTCCCTGAATGTTCG
[0045] GATGGAAATGTGTTGACCTTCTTGGGGAAGTGAGGTCGAAGTGTCTTTCGTTCGTAATTTGACA
[0046] CTTTCCTTCAAACTGGATAAGCACATGGAGATGTGGTTCCCCATTCTCGTGTAGTTCACGAGCT
[0047] ATCTTGATGAACTTCTTGTTAGATGGACATTGAATGCTCTGCAATTGTGAAAGTGCTTCTTCTTT
[0048] TGTGAGAGAACATCTGGGATATGTGAGGAAAATGTTCTTTGCCTTCACACAAAAATAACCACT
[0049] CCGGGGCAT
[0050] SEQ ID NO: 3 is an exemplary nucleic acid SbMMV replicon 5’ backbone sequence.
[0051] ATCGGTAAAAGCAAATGTACCCCCAATTGCCCCCCCCTTCTAAAACTCTATACAATTGGGGGTA
[0052] ATGGGGGTGAATATATACCTACTACTATTAAATTCTCTTTAGCGAGGATTTCAAATCCGCCACG
[0053] TGTACAAAAGGCCATCCCTTATAATATTACAGGGATGGCCGCGCCCGCCCCCCCTTTATTGTGG
[0054] CGCCCACATCTTTGATTTCAGCCAATCAGGAAGCAGCCTCAATCCTTATTTATCTGTTTCTCCAC
[0055] TATAAACTTGTTGCGCAAGTTGGTATTTGAATTAAAGATGTGGGATCTGGCGCGCC
[0056] SEQ ID NO: 4 is an exemplary nucleic acid SbMMV replicon 3’ backbone sequence.
[0057] TTAATAAAGATTGAATTTTATATCATATTTAATCTCGAAATGGGATACAATGACGAACTGCTCA
[0058] AGCAATACTTTGTTGATAGCAGTCTGGACACTTAAAATAGAAATTGAACCTATACTATTCAAGA
[0059] ACGACATCATTAAAAACTTGAAAGCTCGAAAAAATCTCCAGTTCTGACTCTGTGAGTGGAGAA
[0060] AGATCTTTAAATCCAGGAAGCATTTGTGAATCTCCAACTTCCTCTTGACACAGTTGTTGAACAT
[0061] TATTCGAATGTTCATCTCGTCGTATGGGAGACCCATCACCCCATAGTCTTCCCGGAGGAGTTTG
[0062] AAACAGAGGGGATTGTTCACCTCCCAGATATACACGCCACGCTCTAGTTCCGATTGTGTGAGTA
[0063] ACTCCCCTGTGCGAGAATCCATGGTCTGTGCAGTCGAGGTGAACGTATATAGAACACCCACAA
[0064] GGACAGTCTATGCGACGCCGTCTAGTCGCCCTATGACGCTTTGCGTCCCTATGTTGAGCCTTGA
[0065] TGTTCGGTGGAGAACAGTGGCTCTTCGAGGGTGTAGAAGATCGCATTTTTTAAAGCCCACTCCT
[0066] TAAGCGCTGCATTCGACTCCTCATCCATATATTCTTTATAGGAAGATCTGGGGCCAGGATTGCA
[0067] GAGGAAAATTGTTGGTATTCCACCTTTAATTTGAGTGGGCTTGCCGTACTTCACGTTGCTTTGCC
[0068] AGTCTCTTTGGGACCCCATGAATTCTTTCATATGCTTTAGATAGTGGGGATCTACGTCATCGAT
[0069] GACGTTGTACCAGGCATCGTTTGAATAGACCTTGGCGCTGAGATCTAAATGACCACATAGGTA
[0070] ATTGTGCCTACCAAGGCTTCTGGCCCACATCGTCTTCCCCGTCCTGGATTCACCTTCCAGGACG
[0071] ATGCTTATGGGTCTCATCGGCCGCGCAGCGGAAGCTTTCACATTATCTGATGCCCATTGAGACA
[0072] GAACTTCCGGCACATTGTTAAATTGTTCAACTTTAAAAGGCGATTTGTAAATCGTCGTTGGAGG
[0073] CGTAAAAATCTTATTGAGGTTGCTAATAACATTATGATACTGAAAAACATAATCTCGTGGAAGC
[0074] TGCTCTTTGATGATCTGTAAAGCAGCTTCAGCGGAACCAGCATTTAATGCCGACGCACATGCGT
[0075] CGTTAGCATTTTTGCAACCTCCTCTGGCACTTCGACCATCAATCTGGAAAGACCCCCATTGTAT GGTGTCTCCGTCTTTGTCGATGTATGTTTTGACGTCCGAACTTGATTTAGCTCCCTGAATGTTCG
[0076] GATGGAAATGTGTTGACCTTCTTGGGGAAGTGAGGTCGAAGTGTCTTTCGTTCGTAATTTGACA
[0077] CTTTCCTTCAAACTGGATAAGCACATGGAGATGTGGTTCCCCATTCTCGTGTAGTTCACGAGCT
[0078] ATCTTGATGAACTTCTTGTTAGATGGACATTGAATGCTCTGCAATTGTGAAAGTGCTTCTTCTTT
[0079] TGTGAGAGAACATCTGGGATATGTGAGGAAAATGTTCTTTGCCTTCACACAAAAATAACCACT
[0080] CCGGGGCATATCGGTAAAAGCAAATGTACCCCCAATTGCCCCCCCCTTCTAAAACTCTATACAA
[0081] TTGGGGGTAATGGGGGTGAATATATACCTACTACTATTAAATTCTCTTTAGCGAGGATTTCAAA
[0082] TCCGCCACGTGTACAAAAGGCCATCCCTTATAATATTACAGGGATGGCCGCGCCCGCCCCCCCT
[0083] TTATTGTGGCGCCCACATCTTTGATTTCAGCCAATCAGGAAGCAGCCTCAATCCTTATTTATCTG
[0084] TTTCTCCACTATAAACTTGTTGCGCAAGTTGGTATTTGAATTAAAG
[0085] SEQ ID NO: 5 is an exemplary Cas9 amino acid sequence.
[0086] MDKKYSIGLDIGTNSVGWAVITDDYKVPSKKFKVLGNTDRHSIKKNLIGALLFDSGEIAEATRLKRT
[0087] ARRRYTRRKNRICYLQEIFSNEMAKVDDSFFHRLEESFLVEEDKKHERHPIFGNIVDEVAYHEKYPTI
[0088] YHLRKKLVDSTDKADLRLIYLALAHMIKFRGHFLIEGDLNPDNSDVDKLFIQLVQTYNQLFEENPIN
[0089] ASGVDAKAILSARLSKSRRLENLIAQLPGEKKNGLFGNLIALSLGLTPNFKSNFDLAEDAKLQLSKD
[0090] TYDDDLDNLLAQIGDQYADLFLAAKNLSDAILLSDILRLNSEITKAPLSASMIKRYDEHHQDLTLLK
[0091] ALVRQQLPEKYKEIFFDQSKNGYAGYIDGGASQEEFYKFIKPILEKMDGTEELLAKLNREDLLRKQR
[0092] TFDNGSIPHQIHLGELHAILRRQEDFYPFLKDNREKIEKILTFRIPYYVGPLARGNSRFAWMTRKSEE
[0093] TITPWNFEEVVDKGASAQSFIERMTNFDKNLPNEKVLPKHSLLYEYFTVYNELTKVKYVTEGMRKP
[0094] AFLSGEQKKAIVDLLFKTNRKVTVKQLKEDYFKKIECFDSVEISGVEDRFNASLGTYHDLLKIIKDK
[0095] DFLDNEENEDILEDIVLTLTLFEDREMIEERLKTYAHLFDDKVMKQLKRRRYTGWGRLSRKLINGIR
[0096] DKQSGKTILDFLKSDGFANRNFMQLIHDDSLTFKEDIQKAQVSGQGDSLHEHIANLAGSPAIKKGIL
[0097] QTVKVVDELVKVMGRHKPENIVIEMARENQTTQKGQKNSRERMKRIEEGIKELGSQILKEHPVENT
[0098] QLQNEKLYLYYLQNGRDMYVDQELDINRLSDYDVDHIVPQSFIKDDSIDNKVLTRSDKNRGKSDN
[0099] VPSEEVVKKMKNYWRQLLNAKLITQRKFDNLTKAERGGLSELDKAGFIKRQLVETRQITKHVAQIL
[0100] DSRMNTKYDENDKLIREVKVITLKSKLVSDFRKDFQFYKVREINNYHHAHDAYLNAVVGTALIKK
[0101] YPKLESEFVYGDYKVYDVRKMLAKSEQEIGKATAKYFFYSNIMNFFKTEITLANGEIRKRPLIETNG
[0102] ETGEIVWDKGRDFATVRKVLSMPQVNIVKKTEVQTGGFSKESILPKRNSDKLIARKKDWDPKKYG
[0103] GFDSPTVAYSVLVVAKVEKGKSKKLKSVKELLGITIMERSSFEKNPIDFLEAKGYKEVRKDLIIKLPK
[0104] YSLFELENGRKRMLASAGELQKGNELALPSKYVNFLYLASHYEKLKGSPEDNEQKQLFVEQHKHY
[0105] LDEIIEQISEFSKRVILADANLDKVLSAYNKHRDKPIREQAENIIHLFTLTNLGAPTAFKYFDTTIDRK
[0106] RYTSTKEVLDATFIHQSITGLYETRIDLSQLGGD
[0107] SEQ ID NO: 6 is an exemplary 5’ transfer DNA (T-DNA) border sequence (LB).
[0108] TGGCAGGATATATTGTGGTGTAAAC
[0109] SEQ ID NO: 7 is an exemplary 3’ transfer DNA (T-DNA) border sequence (RB).
[0110] TGACAGGATATATTGGCGGGTAAAC
[0111] SEQ ID NO: S is an exemplary EPSPS resistance gene.
[0112] TCAAGCAGCCTTAGTGTCGGAGAGTTCGATCTTAGCTCCAAGACCAGCCATCAAATCCATGAA
[0113] CTCTGGGAAGCTAGTAGCGATCATAGTAGCATCATCAACAGTAACAGGGTTTTCAGAAACGAG
[0114] ACCCATAACGAGGAAGCTCATAGCGATACGGTGATCGAGGTGGGTAGCGACAGCTGCTCCAGA
[0115] AGCGTTACCGAGACCCTTACCGTCAGGACGACCACGCACGACGAGAGAAGTCTCACCTTCATC
[0116] GCAATCAACACCGTTGAGCTTGAGACCGTTTGCGACAGCAGAAAGACGGTCGCTTTCCTTAAC
[0117] ACGGAGTTCTTCCAAACCGTTCATAACGGTAGCACCTTCAGCGAATGCAGCTGCAACAGCGAG
[0118] AATTGGATACTCGTCGATCATAGAAGGAGCACGGTCTTCTGGAACAGTAACACCCTTCAAAGT
[0119] AGAAGAACGAACACGCAAGTCAGCCACGTCTTCTCCACCAGCAAGACGTGGGTTGATCACTTC
[0120] GATGTCGGCACCCATTTCCTGCAGAGTCAAGATGAGACCAGTACGGGTTGGGTTCATCAAAAC
[0121] GTTAAGGATGGTGACGTCGGAACCTGGAACAAGCAAGGCAGCAACCAATGGGAAAGCAGTAG
[0122] AGGATGGATCACCTGGAACATCAATCACTTGACCGGTGAGCTTACCACGACCTTCAAGACGGA
[0123] TGGTACGCACACCGTCAGCATCAGTCTCAACGGTAAGGTTAGCACCAAAACCTTGAAGCATCT TTTCAGTGTGGTCACGAGTCATGATTGGCTCGATAACAGTGGTGATACCTGGGGTGTTGAGACC
[0124] AGCAAGCAGAACAGCGGACTTCACTTGAGCGGAAGCCATAGGTACCCTGTAGGTGATTGGCGT
[0125] TGGAGTCTTTGGTCCACGCAAGGTAACTGGAAGACGATCACCGTCTTCAGACTTCACTTGCACA
[0126] CCCATTTCGCGAAGTGGGTTCAACACACGACCCATTGGACGCTTAGTGAGAGAAGCGTCACCA
[0127] ATGAAAGTGCTATCGAAATCGTAAACACCAACAAGACCCATAGTCAAACGGCAACCAGTTGCA
[0128] GCGTTACCGAAATCGAGAGGAGCCTCAGGAGCAAGGAGTCCACCGTTACCAACACCATCAATG
[0129] ATCCAAGTATCACCTTCCTTACGGATTCTGGCACCCATAGCTTGCATAGCCTTACCAGTGTTGA
[0130] TAACATCTTCACCTTCCAAAAGACCGGTGATACGAGTTTCACCGCTAGCGAGACCTCCAAACAT
[0131] GAAGGACCTGTGGGAGATAGACTTGTCACCTGGAATACGGACGGTTCCAGAAAGACCAGAGG
[0132] ACTTACGAGCAGTTGCTGGACGGCTGCTTGCACCGTGAAGCATGCACGCCGTGGAAACAGAAG
[0133] ACATGACCTTAAGAGGACGAAGCTCAGAGCCAATTAACGTCATCCCACTCTTCTTCAATCCCCA
[0134] CGACGACGAAATCGGATAAGCTCGTGGATGCTGCTGCGTCTTCAGAGAAACCGATAAGGGAGA
[0135] TTTGCGTTGACTGGATTTCGAGAGATTGGAGATAAGAGATGGGTTCTGCACACCATTGCAGATT
[0136] CTGCTAACTTGCGCCAT
[0137] SEQ ID NO: 9 is an exemplary left homology arm (pIN4614).
[0138] CGGCAGAAGCTTGACAACTTGCAAAGGGGAGGGAGGACCCACACACACAGGATTACACATGA
[0139] TAGTCTTTTTTTTTATTCTAGAAACTCTTTGTATTTCCAGGAAGACCTGTAAGTGGGAGTGAGGGAAGTAAAGAAAAAAATCATAGAGTGAAAAAAATATATTAAGTTAATCTATCTATAATAAAAT
[0140] GAGTTGCTAATAGAATAAAAAAAATTCTTTATTGGATGAGAAATATCCATGCATCCGGTGTAC
[0141] AGCCAATCAATTAAGAACCGATTCCAAGGGCCCCATTCTGTGTTTATGACCAGCGGAGAGAGA
[0142] AAGAGAGTGGAAAAGTGGAAAGGATAACTCGTCACATATCATGACTAATGTGATTTCAGCACA
[0143] AACATCCCAGACCCAAAGGCAATTGTTGAGATGAGCCGAATTAAGAGACAATATATTAGTAGT AGTGGGAGGACCTGCACGGATGCACTTGATTTATTTCTAGCCTTTTTTCATTTTTTGTTTTTCCT ATAGAACTTGTATTTTCTCTAGCTAAATCTAAATTATGTATTGCAACATATTTTAAAGGAGGGG CAAAACACACTCGATCTACCTTCAAAGAAACTTTTAAGACATTATAAATATGGAGCCCACACTC
[0144] CCTTCTTCCTCAGTAAGATCATCAACTTGTATTCCTCACTTCCTTAGCTCCTCCTCTTCCTTGTTC CTCCTCTTACAATGGCAAAAATGCCTTTAGAGCCTCTAATAG
[0145] SEQ ID NO: 10 is an exemplary right homology arm (pIN4614).
[0146] TGGGGAGAGTCATAGGAGAAGTTCTTGATTCTTTCACCACAAGCACAAAAATGACTGTGAGTT
[0147] ACAACAAGAAGCAAGTTTACAATGGCCATGAGCTCTTCCCTTCCACTGTCAACACCAAGCCCA AGGTTGAGATTGAGGGTGGTGATATGAGGTCCTTCTTTACACTGGTATATATATTTCTTTTATCT CTTCTCTTCTTATTCCTTTTCTTCTTTAAAGAACAAGATTTTTTTGGGGAAAAAAAAAGAGAAA AAACTTAAGCTGTTGCCAGTGTGTTTCTTTGTTATTTTCTGTAATCATGGTCACACGGTATGTCT
[0148] CTTTTATTGGAGTTTTCTTTACAAATACTGAATCATTAAGCAAATGTCCCCCTTTTTCGGTGCAG
[0149] ATCATGACTGACCCTGATGTTCCTGGCCCTAGTGACCCTTATCTGAGAGAGCACTTGCACTGGT ACTTAATATAAAAATTAACTAACTTAAGTAGTCATTTATTACATAAATAACACACCACACCCAC CCCACATATATTGTTATACCGGTTTCTCTTAATAACTTAACCTCTTTAAAAGTTATGTACTGGTT ATCCATCTGATATTCACGTTAGATGTCTGTGTATTTTATTATCAAGTTTGAAGTTTTGCAAGATA
[0150] TATGCATTCCTTTCATCACATCAAACTATGGACAGCCAGGAAGGTAATAACAAATACTTACATC ACATGAAGAAGGTACACACATGCCTTTGTATTCTGCAAAAGTAGCTTG
[0151] SEQ ID NO: 11 is an exemplary SbMMV replicon (pIN4614). (NNN)nmarks the location of the insert sequence, which can be any desired sequence.
[0152] ATCGGTAAAAGCAAATGTACCCCCAATTGCCCCCCCCTTCTAAAACTCTATACAATTGGGGGTA ATGGGGGTGAATATATACCTACTACTATTAAATTCTCTTTAGCGAGGATTTCAAATCCGCCACG TGTACAAAAGGCCATCCCTTATAATATTACAGGGATGGCCGCGCCCGCCCCCCCTTTATTGTGG CGCCCACATCTTTGATTTCAGCCAATCAGGAAGCAGCCTCAATCCTTATTTATCTGTTTCTCCAC TATAAACTTGTTGCGCAAGTTGGTATTTGAATTAAAGATGTGGGATCTGGCGCGCCCGGCAGA AGCTTGACAACTTGCAAAGGGGAGGGAGGACCCACACACACAGGATTACACATGATGCTTTTT TTGACCTGATCGAGCTGATGGGAGAGAAAAAACTGGGAAATTTTATATTTTTAAATATTTTTTA CAAATTTTTTCAGAACTAGGATGGGATAGAAATAAATAAAAAAAGTACACAATAAGAGTTGCT AATAGAATAAAAAAAATTCTTTATTGGATGAGAAATATCCATGCATCCGGTGTACAGCCAATC
[0153] AATTAAGAACCGATTCCAAGGGCCCCATTCTGTGTTTATGACCAGCGGAGAGAGAAAGAGAGT
[0154] GGAAAAGTGGAAAGGATAACTCGTCACATATCATGACTAATGTGATTTCAGCACAAACATCCC
[0155] AGACCCAAAGGCAATTGTTGAGATGAGCCGAATTAAGAGACAATATATTAGTAGTAGTGGGAG
[0156] GACCTGCACGGATGCACTTGATTTATTTCTAGCCTTTTTTCATTTTTTGTTTTTCCTATAGAACTT
[0157] GTATTTTCTCTAGCTAAATCTAAATTATGTATTGCAACATATTTTAAAGGAGGGGCAAAACACA
[0158] CTCGATCTACCTTCAAAGAAACTTTTAAGACATTATAAATATGGAGCCCACACTCCCTTCTTCC
[0159] TCAGTAAGATCATCAACTTGTATTCCTCACTTCCTTAGCTCCTCCTCTTCCTTGTTCCTCCTCTTA
[0160] CAATGGCAAAAATGCCTTTAGAGCCTCTAATAG(NNN)„TGGGGAGAGTCATAGGAGAAGTTCTT
[0161] GATTCTTTCACCACAAGCACAAAAATGACTGTGAGTTACAACAAGAAGCAAGTTTACAATGGC
[0162] CATGAGCTCTTCCCTTCCACTGTCAACACCAAGCCCAAGGTTGAGATTGAGGGTGGTGATATGA
[0163] GGTCCTTCTTTACACTGGTATATATATTTCTTTTATCTCTTCTCTTCTTATTCCTTTTCTTCTTTAA
[0164] AGAACAAGATTTTTTTGGGGAAAAAAAAAGAGAAAAAACTTAAGCTGTTGCCAGTGTGTTTCT
[0165] TTGTTATTTTCTGTAATCATGGTCACACGGTATGTCTCTTTTATTGGAGTTTTCTTTACAAATACT
[0166] GAATCATTAAGCAAATGTCCCCCTTTTTCGGTGCAGATCATGACTGACCCTGATGTTCCTGGCC
[0167] CTAGTGACCCTTATCTGAGAGAGCACTTGCACTGGTACTTAATATAAAAATTAACTAACTTAAG
[0168] TAGTCATTTATTACATAAATAACACACCACACCCACCCCACATATATTGTTATACCGGTTTCTCT
[0169] TAATAACTTAACCTCTTTAAAAGTTATGTACTGGTTATCCATCTGATATTCACGTTAGATGTCTG
[0170] TGTATTTTATTATCAAGTTTGAAGTTTTGCAAGATATATGCATTCCTTTCATCACATCAAACTAT
[0171] GGACAGCCAGGAAGGTAATAACAAATACTTACATCACATGAAGAAGGTACACACATGCCTTTG
[0172] TATTCTGCAAAAGTAGCTTGGAGCTCGGATCCGCATGCAAGCTTGGCGTAATCAGCAATTAATA
[0173] AAGATTGAATTTTATATCATATTTAATCTCGAAATGGGATACAATGACGAACTGCTCAAGCAAT
[0174] ACTTTGTTGATAGCAGTCTGGACACTTAAAATAGAAATTGAACCTATACTATTCAAGAACGACA
[0175] TCATTAAAAACTTGAAAGCTCGAAAAAATCTCCAGTTCTGACTCTGTGAGTGGAGAAAGATCTT
[0176] TAAATCCAGGAAGCATTTGTGAATCTCCAACTTCCTCTTGACACAGTTGTTGAACATTATTCGA
[0177] ATGTTCATCTCGTCGTATGGGAGACCCATCACCCCATAGTCTTCCCGGAGGAGTTTGAAACAGA
[0178] GGGGATTGTTCACCTCCCAGATATACACGCCACGCTCTAGTTCCGATTGTGTGAGTAACTCCCC
[0179] TGTGCGAGAATCCATGGTCTGTGCAGTCGAGGTGAACGTATATAGAACACCCACAAGGACAGT
[0180] CTATGCGACGCCGTCTAGTCGCCCTATGACGCTTTGCGTCCCTATGTTGAGCCTTGATGTTCGGT
[0181] GGAGAACAGTGGCTCTTCGAGGGTGTAGAAGATCGCATTTTTTAAAGCCCACTCCTTAAGCGCT
[0182] GCATTCGACTCCTCATCCATATATTCTTTATAGGAAGATCTGGGGCCAGGATTGCAGAGGAAAA
[0183] TTGTTGGTATTCCACCTTTAATTTGAGTGGGCTTGCCGTACTTCACGTTGCTTTGCCAGTCTCTTT
[0184] GGGACCCCATGAATTCTTTCATATGCTTTAGATAGTGGGGATCTACGTCATCGATGACGTTGTA
[0185] CCAGGCATCGTTTGAATAGACCTTGGCGCTGAGATCTAAATGACCACATAGGTAATTGTGCCTA
[0186] CCAAGGCTTCTGGCCCACATCGTCTTCCCCGTCCTGGATTCACCTTCCAGGACGATGCTTATGG
[0187] GTCTCATCGGCCGCGCAGCGGAAGCTTTCACATTATCTGATGCCCATTGAGACAGAACTTCCGG
[0188] CACATTGTTAAATTGTTCAACTTTAAAAGGCGATTTGTAAATCGTCGTTGGAGGCGTAAAAATC
[0189] TTATTGAGGTTGCTAATAACATTATGATACTGAAAAACATAATCTCGTGGAAGCTGCTCTTTGA
[0190] TGATCTGTAAAGCAGCTTCAGCGGAACCAGCATTTAATGCCGACGCACATGCGTCGTTAGCATT
[0191] TTTGCAACCTCCTCTGGCACTTCGACCATCAATCTGGAAAGACCCCCATTGTATGGTGTCTCCG
[0192] TCTTTGTCGATGTATGTTTTGACGTCCGAACTTGATTTAGCTCCCTGAATGTTCGGATGGAAATG
[0193] TGTTGACCTTCTTGGGGAAGTGAGGTCGAAGTGTCTTTCGTTCGTAATTTGACACTTTCCTTCAA
[0194] ACTGGATAAGCACATGGAGATGTGGTTCCCCATTCTCGTGTAGTTCACGAGCTATCTTGATGAA
[0195] CTTCTTGTTAGATGGACATTGAATGCTCTGCAATTGTGAAAGTGCTTCTTCTTTTGTGAGAGAAC
[0196] ATCTGGGATATGTGAGGAAAATGTTCTTTGCCTTCACACAAAAATAACCACTCCGGGGCATATC
[0197] GGTAAAAGCAAATGTACCCCCAATTGCCCCCCCCTTCTAAAACTCTATACAATTGGGGGTAATG
[0198] GGGGTGAATATATACCTACTACTATTAAATTCTCTTTAGCGAGGATTTCAAATCCGCCACGTGT
[0199] ACAAAAGGCCATCCCTTATAATATTACAGGGATGGCCGCGCCCGCCCCCCCTTTATTGTGGCGC
[0200] CCACATCTTTGATTTCAGCCAATCAGGAAGCAGCCTCAATCCTTATTTATCTGTTTCTCCACTAT
[0201] AAACTTGTTGCGCAAGTTGGTATTTGAATTAAAG
[0202] SEQ ID NO: 12 is an exemplary AtUbilO promoter.
[0203] CCTGTTAATCAGAAAAACTCAGATTAATCGACAAATTCGATCGCACAAACTAGAAACTAACAC
[0204] CAGATCTAGATAGAAATCACAAATCGAAGAGTAATTATTCGACAAAACTCAAATTATTTGAAC
[0205] AAATCGGATGATATCTATGAAACCCTAATCGAGAATTAAGATGATATCTAACGATCAAACCCA GAAAATCGTGTTCGATCTAAGATTAACAGAATCTAAACCAAAGAACATATACGAAATTGGGAT
[0206] CGAACGAAAACAAAATCGAAGATTTTGAGAGAATAAGGAACACAGAAATTTACCTTGATCACG
[0207] GTAGAGAGAATTGAGAGAAAGTTTTTAAGATTTTGAGAAATTGAAATCTGAATTGTGAAGAAG
[0208] AAGAGCTCTTTGGGTATTGTTTTATAGAAGAAGAAGAAGAAAAGACGAGGACGACTAGGTCAC
[0209] GAGAAAGCTAAGGCGGTGAAGCAATAGCTAATAATAAAATGACACGTGTATTGAGCGTTGTTT
[0210] ACACGCAAAGTTGTTTTTGGCTAATTGCCTTATTTTTAGGTTGAGGAAAAGTATTTGTGCTTTGA
[0211] GTTGATAAACACGACTCGTGTGTGCCGGCTGCAACCACTTTGACGCCGTTTATTACTGACTCGT
[0212] CGACAACCACAATTTCTAACGGTCGTCATAAGATCCAGCCGTTGAGATTTAACGATCGTTACGA
[0213] TTTATATTTTTTTAGCATTATCGTTTTATTTTTTAAATATACGGTGGAGCTGAAAATTGGCAATA
[0214] ATTGAACCGTGGGTCCCACTGCATTGAAGCGTATTTCGTATTTTCTAGAATTCTTCGTGCTTTAT
[0215] TTCTTTTCCTTTTTGTTTTTTTTTGCCATTTATCTAATGCAAGTGGGCTTATAAAATCAGTGAATT
[0216] TCTTGGAAAAGTAACTTCTTTATCGTATAACATATTGTGAAATTATCCATTTCTTTTAATTTTTT
[0217] AGTGTTATTGGATATTTTTGTATGATTATTGATTTGCATAGGATAATGACTTTTGTATCAAGTTG
[0218] GTGAACAAGTCTCGTTAAAAAAGGCAAGTGGTTTGGTGACTCGATTTATTCTTGTTATTTAATT
[0219] CATATATCAATGGATCTTATTTGGGGCCTGGTCCATATTTAACACTCGTGTTCAGTCCAATGAC
[0220] CAATAATATTTTTTCATTAATAACAATGTAACAAGAATGATACACAAAACATTCTTTGAATAAG
[0221] TTCGCTATGAAGAAGGGAACTTATCCGGTCCTAGATCATCAGTTCATACAAACCTCCATAGAGT
[0222] TCAACATCTTAAACAAGAATATCCTGATCCGTTGACCTGCAGCTCGAC
[0223] SEQ ID NO: 13 is an exemplary SlUbilO promoter,
[0224] ATCGTATCCAGTGCACCATATTTTTTGGCGATTACCACTCATATTATTGTGTTTAGTAGATATTT
[0225] TAGGTGCATAATTGATCTCTTCTTTAAAACTAGGGGCACTTATTATTATACATCCACTTGACACT
[0226] TGCTTTAGTTGGCTATTTTTTTTATTTTTTATTTTTTGTCAACTACCCCAATTTAAATTTTATTTGA
[0227] TTAAGATATTTTTATGGACCTACTTTATAATTAAAAATATTTTCTATTTGAAAAGGAAGGACAA
[0228] AAATCATACAATTTTGGTCCAACTACTCCTCTCTTTTTTTTTTTGGCTTTATAAAAAAGGAAAGT
[0229] GATTAGTAATAAATAATTAAATAATGAAAAAAGGAGGAAATAAAATTTTCGAATTAAAATGTA
[0230] AAAGAGAAAAAGGAGAGGGAGTAATCATTGTTTAACTTTATCTAAAGTACCCCAATTCGATTT
[0231] TACATGTATATCAAATTATACAAATATTTTATTAAAATATAGATATTGAATAATTTTATTATTCT
[0232] TGAACATGTAAATAAAAATTATCTATTATTTCAATTTTTATATAAACTATTATTTGAAATCTCAA
[0233] TTATGATTTTTTAATATCACTTTCTATCCATGATAATTTCAGCTTAAAAAGTTTTGTCAATAATT
[0234] ACATTAATTTTGTTGATGAGGATGACAAGATTTCGGTCATCAATTACATATACACAAATTGAAA
[0235] TAGTAAGCAACTTGATTTTTTTTCTCATAATGATAATGACAAAGACACGAAAAGACAATTCAAT
[0236] ATTCACATTGATTTATTTTTATATGATAATAATTACAATAATAATATTCTTATAAAGAAAGAGA
[0237] TCAATTTTGACTGATCCAAAAATTTATTTATTTTTACTATACCAACGTCACTAATTATATCTAAT
[0238] AATGTAAAACAATTCAATCTTACTTAAATATTAATTTGAAATAAACTATTTTTATAACGAAATT
[0239] ACTAAATTTATCCAATAACAAAAAGGTCTTAAGAAGACATAAATTCTTTTTTTGTAATGCTCAA
[0240] ATAAATTTGAGTAAAAAAGAATGAAATTGAGTGATTTTTTTTTAATCATAAGAAAATAAATAAT
[0241] TAATTTCAATATAATAAAACAGTAATATAATTTCATAAATGGAATTCAATACTTACCTCTTAGA
[0242] TATAAAAAATAAATATAAAAATAAAGTGTTTCTAATAAACCCGCAATTTAAATAAAATATTTA
[0243] ATATTTTCAATCAAATTTAAATAATTATATTAAAATATCGTAGAAAAAGAGCAATATATAATAC
[0244] AAGAAAGAAGATTTAAGTACAATTATCAACTATTATTATACTCTAATTTTGTTATATTTAATTTC
[0245] TTACGGTTAAGGTCATGTTCACGATAAACTCAAAATACGCTGTATGAGGACATATTTTAAATTT
[0246] TAACCAATAATAAAACTAAGTTATTTTTAGTATATTTTTTTGTTTAACGTGACTTAATTTTTCTTT
[0247] TCTAGAGGAGCGTGTAAGTGTCAACCTCATTCTCCTAATTTTCCCAACCACATAAAAAAAAAAT
[0248] AAAGGTAGCTTTTGCGTGTTGATTTGGTACACTACACGTCATTATTACACGTGTTTTCGTATGAT
[0249] TGGTTAATCCATGAGGCGGTTTCCTCTAGAGTCGGCCATACCATCTATAAAATAAAGCTTTCTG
[0250] CAGCTCATTTTTTCATCTTCTATCTGATTTCTATTATAATTTCTCTGAATTGCCTTCAAATTTCTC
[0251] TTTCAAGGTTAGAATTTTTCTCTATTTTTTGGTTTTTGTTTGTTTAGATTCTGAGTTTAGTTAATC
[0252] AGGTGCTGTTAAAGCCCTAAATTTTGAGTTTTTTTCGGTTGTTTTGATGGAAAATACCTAACAAT
[0253] TGAGTTTTTTCATGTTGTTTTGTCGGAGAATGCCTACAATTGGAGTTCCTTTCGTTGTTTTGATG
[0254] AGAAAGCCCCTAATTTGAGTGTTTTTCCGTCGATTTGATTTTAAAGGTTTATATTCGAGTTTTTT
[0255] TCGTCGGTTTAATGAGAAGGCCTAAAATAGGAGTTTTTCTGGTTGATTTGACTAAAAAAGCCAT
[0256] GGAATTTTGTGTTTTTGATGTCGCTTTGGTTCTCAAGGCCTAAGATCTGAGTTTCTCCGGTTGTT
[0257] TTGATGAAAAAGCCCTAAAATTGGAGTTTTTATCTTGTGTTTTAGGTTGTTTTAATCCTTATAAT
[0258] TTGAGTTTTTTCGTTGTTCTGATTGTTGTTTTTATGAATTTCCTGCA SEQ ID NO: 14 is an exemplary AtU6 promoter.
[0259] CTTCGTTGAACAACGGAAACTCGACTTGCCTTCCGCACAATACATCATTTCTTCTTAGCTTTTTT
[0260] TCTTCTTCTTCGTTCATACAGTTTTTTTTTGTTTATCAGCTTACATTTTCTTGAACCGTAGCTTTC
[0261] GTTTTCTTCTTTTTAACTTTCCATTCGGAGTTTTTGTATCTTGTTTCATAGTTTGTCCCAGGATTA
[0262] GAATGATTAGGCATCGAACCTTCAAGAATTTGATTGAATAAAACATCTTCATTCTTAAGATATG
[0263] AAGATAATCTTCAAAAGGCCCCTGGGAATCTGAAAGAAGAGAAGCAGGCCCATTTATATGGGA
[0264] AAGAACAATAGTATTTCTTATATAGGCCCATTTAAGTTGAAAACAATCTTCAAAAGTCCCACAT
[0265] CGCTTAGATAAGAAAACGAAGCTGAGTTTATATACAGCTAGAGTCGAAGTAGTGATTG
[0266] SEQ ID NO: 15 is an exemplary pea rbcs E9 terminator.
[0267] ACAACTACAAGTGTTTTACTCCTCATATTAACTTCGGTCATTAGAGGCCACGATTTGACACATT
[0268] TTTACTCAAAACAAAATGTTTGCATATCTCTTATAATTTCAAATTCAACACACAACAAATAAGA
[0269] GAAAAAACAAATAATATTAATTTGAGAATGAACAAAAGGACCATATCATTCATTAACTCTTCT
[0270] CCATCCATTTCCATTTCACAGTTCGATAGCGAAAACCGAATAAAAAACACAGTAAATTACAAG
[0271] CACAACAAATGGTACAAGAAAAACAGTTTTCCCAATGCCATAATACTCGAAC
[0272] SEQ ID NO: 16 is an exemplary AtHSPt terminator.
[0273] ATTTTATTTGTACATGTAGTGGCATTAGGTATTGGTAGATATTAATTGTATAGTGTTTGGTTGGTGCCATATATTTATCATAAAAATGACTTTTAGATAGTTGGACATTTGATAAGATGTTAGTTCGAT
[0274] CATTATAATGAATAAACAAATGTTTCTATAATCCATTGTGAATGTTTTGTTGGATCTCTTCTGCA
[0275] GCATATAACTACTGTATGTGCTATGGTATGGACTATGGAATATGATTAAAGATAAG
[0276] SEQ ID NO: 17 is an exemplary pea3A terminator.
[0277] CAGGCCTCCCAGCTTTCGTCCGTATCATCGGTTTCGACAACGTTCGTCAAGTTCAATGCATCAG
[0278] TTTCATTGCCCACACACCAGAATCCTACTAAGTTTGAGTATTATGGCATTGGAAAAGCTGTTTT
[0279] CTTCTATCATTTGTTCTGCTTGTAATTTACTGTGTTCTTTCAGTTTTTGTTTTCGGACATCAAAAT
[0280] GCAAATGGATGGATAAGAGTTAATAAATGATATGGTCCTTTTGTTCATTCTCAAATTATTATTA
[0281] TCTGTTGTTTTTACTTTAATGGGTTGAATTTAAGTAAGAAAGGAACTAACAGTGTGATATTAAG
[0282] GTGCAATGTTAGACATATAAAACAGTCTTTCACCTCTCTTTGGTTATGTCTTGAATTGGTTTGTT
[0283] TCTTCACTTATCTGTGTAATCAAGTTTACTATGAGTCTATGATCAAGTAATTATGCAATCAAGTT
[0284] AAGTACAGTATAGGCTT
[0285] DETAILED DESCRIPTION
[0286] I. Introduction
[0287] Geminiviral vectors can be used for transient expression of a nucleic acid. Here it is shown that a replicon based on a specific Geminivirus, soybean mild mottle virus (SbMMV), is particularly effective at generating high copies of a nucleic acid of interest in a plant host cell. In brief, the V2 (coat protein) and VI (movement protein) regions were replaced with a nucleic acid of interest, while the C1 / C2 region and stem loop region were retained. The Examples show that the SbMMV replicon had surprisingly superior amplification over other Geminivirus-based replicons, such as those based on bean yellow dwarf virus (BeYDV) or soybean blistering mosaic virus (SbBMV). When used to amplify a repair template in the context of CRISPR-based genome editing, the SbMMV replicon improved the precise integration of a repair template into a target, and thus increased the efficiency of CRISPR-based editing. II. Summary of Terms
[0288] Unless otherwise noted, technical terms are used according to conventional usage. Definitions of many common terms in molecular biology may be found in Krebs et al. (eds.), Lewin ’s genes XII, published by Jones & Bartlett Learning, 2017. As used herein, the singular forms “a,” “an,” and “the,” refer to both the singular as well as plural, unless the context clearly indicates otherwise. For example, the term “a replicon” includes singular or plural replicons and can be considered equivalent to the phrase “at least one replicon.” As used herein, the term “comprises” means “includes.” It is further to be understood that any and all base sizes or amino acid sizes, and all molecular weight or molecular mass values, given for nucleic acids or polypeptides are approximate, and are provided for descriptive purposes, unless otherwise indicated. Although many methods and materials similar or equivalent to those described herein can be used, particular suitable methods and materials are described herein. In case of conflict, the present specification, including explanations of terms, will control. In addition, the materials, methods, and examples are illustrative only and not intended to be limiting. To facilitate review of the various aspects, the following explanations of terms are provided:
[0289] CRISPR / Cas Systems: CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) / Cas (CRISPR-associated) systems, or CRISPR systems, use RNA-guided nucleases termed CRISPR-associated or “Cas” endonucleases (e.g., Cas9 or Cpfl ) to cleave a target nucleotide sequence. In a typical CRISPR / Cas system, a Cas endonuclease is directed to a target nucleotide sequence (e.g., a site in the genome that is to be sequence-edited) by sequence-specific, non-coding “guide RNA” (gRNA) that targets single- or double-stranded nucleotide sequences.
[0290] Three classes (I-III) of CRISPR systems have been identified. The well characterized class II CRISPR systems use a single Cas endonuclease (rather than multiple Cas proteins). One class II CRISPR system includes a type II Cas endonuclease such as Cas9, a CRISPR RNA (“crRNA”), and a trans-activating crRNA (“tracrRNA”). The crRNA contains a “guide RNA,” typically a 16 to 24 nucleotide RNA sequence that corresponds to (e.g., is identical or nearly identical to, or alternatively is complementary or nearly complementary to) a respective 16 to 24 nucleotide target DNA sequence. The crRNA also contains a region that binds to the tracrRNA to form a partially double-stranded structure which is cleaved by RNase
[0291] III, resulting in a crRNA / tracrRNA hybrid. The crRNA / tracrRNA hybrid then directs the Cas9 endonuclease to recognize and cleave the target DNA sequence. Efficient gene editing has also been achieved using a chimeric “single guide RNA” (“sgRNA”), an engineered (synthetic) single RNA molecule that contains both a tracrRNA (for binding the nuclease) and at least one crRNA (to guide the nuclease to the sequence targeted for editing) and mimics a naturally occurring crRNA-tracrRNA complex.
[0292] Another class II CRISPR system includes the type V endonuclease Cpfl (Cas 12a), which is a smaller endonuclease than is Cas9; examples include AsCpfl (from Acidaminococcus sp.) and LbCpfl (from Lachnospiraceae sp.). Cpfl -associated CRISPR arrays are processed into mature crRNAs without the requirement of a tracrRNA; in other words, a Cpfl system requires only the Cpfl nuclease and a crRNA to cleave the target DNA sequence. Cpfl endonucleases, are associated with T-rich PAM sites, e.g., 5’-TTN. Cpfl can also recognize a 5' -CT A PAM motif. Cpfl cleaves the target DNA by introducing an offset or staggered double-strand break with a 4- or 5-nucleotide 5' overhang, for example, cleaving a target DNA with a 5-nucleotide offset or staggered cut located 18 nucleotides downstream from (3’ from) from the PAM site on the coding strand and 23 nucleotides downstream from the PAM site on the complementary strand; the 5-nucleotide overhang that results from such offset cleavage allows more precise genome editing by DNA insertion by homologous recombination than by insertion at blunt-end cleaved DNA. See, e.g., Zetsche et al. (2015) Cell, 163:759 - 771.
[0293] Other CRISPR nucleases useful in methods and compositions of the disclosure include C2cl and C2c3 (see, e.g., Shmakov et al. (2015) Mol. Cell, 60:385 - 397) and CasX (Casl2e) and CasY (Casl2d) (see e.g., Burstein et al. (2016) Nature, doi:10.1038 / nature21059). Like other CRISPR nucleases, C2cl from Alicyclobacillus acidoterrestris (AacC2cl) requires a guide RNA and PAM recognition site; C2cl cleavage results in a staggered seven-nucleotide double-strand break in the target DNA (see. e.g., Yang et al. (2016) Cell, 167:1814 - 1828. el2) and is reported to have high mismatch sensitivity, thus reducing off-target effects. Other CRISPR nucleases include nucleases identified from the genomes of uncultivated microbes, such as CasX (Casl2e) and CasY (Casl2d) (see, e.g., Burstein et al. (2016) Nature, doi: 10.1038 / nature21059). Casl2i, Casl2f, Casl2j, and genome editing systems based on these nucleases and corresponding RNAs have also been developed.
[0294] For the purposes of gene editing, CRISPR arrays can be designed to contain one or multiple guide RNA sequences corresponding to a desired target DNA sequence; see, for example, Cong et al. (2013) Science, 339:819-823; Ran et al. (2013) Nature Protocols, 8:2281 - 2308. At least 16 or 17 nucleotides of gRNA sequence is typically needed for Cas9 mediated DNA cleavage to occur; and at least 16 nucleotides of gRNA sequence are typically needed to achieve for Cpfl mediated DNA cleavage. In practice, guide RNA sequences generally have a length of 17 - 24 nucleotides (frequently 19, 20, or 21 nucleotides) and exact complementarity (perfect base-pairing) to a targeted nucleic acid sequence. Guide RNAs having less than 100% sequence identity to the target sequence can also be used (e.g., a gRNA with a length of 20 nucleotides and between 1 - 4 mismatches to the target sequence), but increase the potential for off-target effects.
[0295] The target DNA sequence must generally be adjacent to a “protospacer adjacent motif’ (“PAM”) that is specific for a given Cas endonuclease; however, PAM sequences are short and relatively non-specific, appearing throughout a given genome. CRISPR endonucleases identified from various prokaryotic species have unique PAM sequence requirements; examples of PAM sequences include 5’-NGG (Streptococcus pyogenes), 5’-NNAGAA (Streptococcus thermophilus CRISPR1), 5’-NGGNG (Streptococcus thermophilus CRISPR3), and 5’-NNNGATT (Neisseria meningitidis). Some endonucleases, e.g., Cas9 endonucleases, are associated with G-rich PAM sites, e.g., 5’-NGG, and perform blunt-end cleaving of the target DNA at a location 3 nucleotides upstream from (5’ from) the PAM site.
[0296] Complementarity: The ability of a nucleic acid to form hydrogen bond(s) with another nucleic acid sequence by either traditional Watson-Crick base pairing or non-traditional pairing types. Percent complementarity indicates the percentage of residues in a nucleic acid molecule which can form hydrogen bonds with a second nucleic acid sequence (e.g., 5, 6, 7, 8, 9, and 10 out of 10 being 50%, 60%, 70%, 80%, 90%, and 100% complementarity, respectively).
[0297] Contacting: Placing an agent in direct physical association; includes both in solid and liquid form, and can take place either in vivo or in vitro. Contacting includes contacting a cell (e.g., a soybean cell) by placing an agent (e.g., a nucleic acid or vector) in direct physical association with the cell.
[0298] Control: A reference standard. A control can be a positive or negative control. In some examples, the control is a measurement (e.g., expression of a target) obtained prior to modifying a host cell (e.g., introducing an exogenous nucleic acid). In some examples, the control is a historical control or a standard reference or range (a typical measurement or range observed for a particular population, such as a typical measurement (e.g., gene expression) or range for an unmodified plant (e.g., soybean)).
[0299] Expression (of a nucleic acid): As used herein, expression of a nucleic acid includes transcription and / or translation of the nucleic acid.
[0300] Expression Cassette: A nucleic acid fragment designed for expression of a particular gene (or genes) in a host cell. Expression cassettes can be included, for example, in a vector or replicon. An expression cassette can include regulatory elements, such as promoters and / or terminators.
[0301] Exogenous: Originating from a different source. For example, a nucleic acid molecule that is exogenous to a cell is a nucleic acid that originated from a source other than the cell itself (e.g., a synthetic nucleic acid that is introduced into a cell). In another example, a nucleic acid molecule that is exogenous to a replicon is a nucleic acid that originated from a source other than the replicon itself.
[0302] Geminivirus: A single-stranded DNA virus that replicates in a host nucleus using a rolling circle mechanism. Replication initiator protein (Rep) and Rep A are encoded on the complementary-sense transcript (C transcript). Rep mRNA is generated when a short intron is spliced in the C transcript, while unspliced mRNA allows translation of Rep A. Rep mediates nicking and ligating functions during rolling circle replication. The LIR contains a bi-directional promoter and a stem-loop structure that initiates rollingcircle replication. Exemplary Gemini viruses include, but are not limited to: Soybean Mild Mottle Virus (SbMMV; see, e.g., NCBI Reference Sequence NC_014140.1), Soybean Blistering Mosaic Virus (SbBMV; see, e.g., NCBI Reference Sequence NC_038463), and Bean Yellow Dwarf Virus (BeYDV; see, e.g., NCBI Reference Sequence Y11023).
[0303] Homolog -Directed Repair (HDR): The repair of one or more double-stranded breaks in DNA using homologous recombination with a donor template (repair template). In molecular biology applications, HDR mechanisms can be used to facilitate site-specific gene / genome editing; for example, a double stranded break can be induced in a target sequence by introducing into a cell a nuclease targeting the sequence (such as a Cas nuclease), and a repair template containing a desired nucleotide sequence can be integrated or “knocked-in” via HDR at the location of the double-stranded break. In some implementations, the repair template disclosed herein is a donor template for HDR. Increase or Decrease: A positive (increase) or negative (decrease) difference relative to a reference value, such as a control. The difference can be a qualitative or quantitative. In some examples, the difference is statistically significant (e.g., P-Value less than 0.05 or 0.01). In some examples, the difference is an increase relative to a control of at least 5%, such as at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 150%, at least 200%, at least 250%, at least 300%, at least 350%, at least 400%, at least 500%, or greater than 500%. In some examples, the difference is a decrease relative to a control of at least 5%, such as at least 10%, at least 20%, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, or 100%.
[0304] Isolated or Purified: An “isolated” or “purified” biological component (such as a nucleic acid, protein, or cell) is one that has been substantially separated from other biological components in the environment in which the component occurs, e.g., separated from other chromosomal and extra- chromosomal DNA and RNA, proteins and / or cells. Nucleic acids and proteins that have been “isolated” include nucleic acids and proteins purified by standard purification methods. The term also embraces nucleic acids and proteins prepared by recombinant expression in a host cell as well as chemically synthesized nucleic acids.
[0305] Absolute purity or isolation is not required, it is intended as a relative term. Thus, for example, a purified / isolated protein, nucleic acid, or cell preparation is one in which the protein, nucleic acid, or cell is more enriched than the protein, nucleic acid, or cell is in its initial environment. In one example, a preparation is purified / isolated such that the protein, nucleic acid, or cell represents at least 50% of the total content of the preparation. A substantially purified protein or nucleic acid is at least 60%, at least 70%, at least 80%, at least 90%, at least 95%, at least 98%, or at least 99% pure. Thus, in one specific, non-limiting example, a substantially purified protein or nucleic acid is 90% free of other components.
[0306] Modified Plant: A plant that includes an artificial genetic modification. A modified plant is not, or otherwise excludes, a naturally occurring plant. In some examples, a modified plant is a genome-edited plant (e.g., by CRISPR-based editing). In other examples, a modified plant is a transgenic plant (a plant that includes a transgene).
[0307] Operably Linked: A nucleic acid sequence is “operably linked” when it is placed in a functional relationship with a second nucleic acid sequence. For instance, a promoter is operably linked to a coding sequence if the promoter affects the transcription or expression of the coding sequence. Generally, operably linked DNA sequences are contiguous and, where necessary to join two protein-coding regions, are in the same reading frame.
[0308] Producing (a nucleic acid): Generating copies of a nucleic acid. Producing includes, for example, amplifying, replicating, or multiplying a nucleic acid. In some examples, a nucleic acid is produced in a plant host cell, such as a soybean cell.
[0309] Promoter: A nucleic acid control sequence that directs transcription of a nucleic acid. A promoter includes necessary nucleic acid sequences near the start site of transcription. A promoter also optionally includes distal enhancer or repressor elements. A “constitutive promoter” is a promoter that is continuously active and is not subject to regulation by external signals or molecules. In contrast, the activity of an “inducible promoter” is regulated by an external signal or molecule (for example, a transcription factor). In some examples, the vectors provided herein include a pol III promoter (e.g., U6), a pol II promoter, ubiquitin promoter, Cauliflower Mosaic Virus (CaMV) 35S promoter, RUBISCO promoter, or combinations thereof.
[0310] Recombinant: A nucleic acid or protein that has a sequence made by an artificial combination of two otherwise separated segments of sequence (e.g., a “chimeric” sequence). This artificial combination can be accomplished by chemical synthesis or by manipulation of isolated segments of nucleic acids, for example, by standard molecular biology techniques (e.g., cloning).
[0311] Regulatory Element: A term that includes promoters, enhancers, internal ribosomal entry sites (IRES), and other expression control elements (e.g., transcription termination signals, such as polyadenylation signals and poly-U sequences). Regulatory elements include those that direct constitutive expression of a nucleotide sequence in many types of host cells and those that direct expression of the nucleotide sequence only in certain host cells (e.g., tissue-specific regulatory sequences). Regulatory elements may also direct expression in a temporal-dependent manner, such as in a cell-cycle dependent or developmental stage-dependent manner, which may or may not also be tissue or cell-type specific.
[0312] Repair Template: A nucleic acid fragment used to modify a target nucleic acid via homologydependent repair (HDR). HDR is a mechanism Eukaryotic cells use to repair double-stranded DNA breaks. Thus, the term “repair” does not refer to repairing function of a target, but rather repair of a double-stranded break induced in a target nucleic acid (for example, by a nuclease). In general, a repair template includes an insert sequence (e.g., a sequence to introduce specific mutations, insertions, or deletions into the target) flanked by sequence homologous to the target nucleic acid. The repair template is incorporated into the target through HDR resulting in modification of the target nucleic acid. A DNA repair template can be single stranded or double stranded. In some examples, the DNA repair template is single stranded.
[0313] Replicon: A nucleic acid molecule that replicates as a unit.
[0314] Reporter: A protein whose expression is linked to the expression of a gene of interest. Exemplary reporter proteins include fluorescent proteins and chemiluminescent molecules, such as infrared-fluorescent proteins (IFPs), GFP, YFP, RFP, CFP, mRFPl, mCherry, mOrange, DsRed, tdTomato, mKO, tagRFP, EGFP, mEGFP, mOrange2, mScarlet, maple, tagRFP-T, firefly luciferase, Renilla luciferase, and click beetle luciferase (see, e.g., US Pat. Pub. No. 2010 / 0122355). In some examples, the reporter protein is positioned downstream of and in frame with a gene of interest, such that the reporter protein is co-expressed with the gene of interest.
[0315] Sequence Identity: The degree of similarity between amino acid or nucleic acid sequences. Sequence identity is frequently measured in terms of percentage identity (or percent identity); the higher the percentage, the more similar the two sequences are. Homologs of a polypeptide (or nucleotide sequence) will possess a relatively high degree of sequence identity when aligned using standard methods. Methods of alignment of sequences for comparison have been described. The NCBI Basic Local Alignment Search Tool (BLAST) tool is often used and is available from several sources, including the National Center for Biotechnology Information (blast.ncbi.nlm.nih.gov / Blast.cgi). Various types of BLAST are available, for example, blastp, blastn, blastx, tblastn and tblastx. A description of how to determine sequence identity using this program is available on the NCBI website and other resources. In some examples, percent sequence identity is determined by using BLAST with default parameters, for example, parameters provided below in Table 1.
[0316] Table 1: Exemplary default parameters for BLAST programs.
[0317] Transformed: A transformed cell is a cell into which an exogenous nucleic acid molecule has been introduced by a molecular biology technique. As used herein, the term transformation encompasses all techniques by which a nucleic acid molecule might be introduced into such a cell, including chemical methods (e.g., calcium-phosphate transfection), physical methods (e.g., electroporation, microinjection, particle bombardment), fusion (e.g., liposomes), lipofection, nucleofection, receptor-mediated endocytosis (e.g., DNA-protein complexes, viral envelope / capsid-DNA complexes), agrobacterium-mediated transformation, biolistics (particle gun accelerator or gene gun), or other transduction and / or transfection methods.
[0318] Transgene: A nucleic acid that is artificially introduced into an organism in which the nucleic acid does not naturally occur.
[0319] Vector: A nucleic acid molecule that can be introduced into a host cell (for example, by transformation), thereby producing a transformed host cell. A vector can include nucleic acid sequences that permit it to replicate in a host cell, such as an origin of replication. Recombinant DNA vectors are vectors containing recombinant DNA. A vector can also include one or more selectable marker genes and other genetic elements. Often vectors are plasmids, however, they can also be viral vectors, cosmids, or artificial chromosomes. In some examples, the vector is a transfer DNA (T-DNA) vector suitable for agrobacterium- mediated transformation or a vector suitable for biolistics.
[0320] III. Nucleic Acids and Vectors
[0321] Disclosed herein are nucleic acids including a recombinant soybean mild mottle virus (SbMMV) replicon. The SbMMV replicon includes an exogenous nucleic acid (e.g., exogenous to SbMMV) and a sequence comprising a SbMMV long intergenic region (LIR). In some implementations, the SbMMV replicon includes a 5’ SbMMV LIR, a 3’ SbMMV LIR, or a 5’ and a 3’ SbMMV LIR relative to the exogenous nucleic acid. In some aspects of the disclosure, the exogenous nucleic acid includes or consists of a repair template having complementary sequence to a target nucleic acid. In other aspects of the disclosure, the exogenous nucleic acid includes or consists of an expression cassette. The repair template or expression cassette are exogenous to SbMMV, for example, the repair template or expression cassette are from a source other than SbMMV. Thus, the SbMMV replicon is a recombinant nucleic acid molecule that is not naturally occurring.
[0322] The LIR facilitates replication of the SbMMV replicon. An exemplary SbMMV LIR sequence is provided herein as SEQ ID NO: 1. In some examples, the LIR includes or consists of a sequence having at least 80% sequence identity to SEQ ID NO: 1; for example, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 1. In some examples, the LIR includes or consists of a sequence having at least 90% sequence identity to SEQ ID NO: 1. In some examples, the LIR includes or consists of a sequence having at least 95% sequence identity to SEQ ID NO: 1. In some examples, the LIR includes or consists of SEQ ID NO: 1. In some examples, the LIR includes a stem loop structure, for example, a SbMMV stem loop.
[0323] In some implementations, the SbMMV replicon includes a 5’ SbMMV replicon backbone sequence. In some examples, the 5’ SbMMV replicon backbone sequence includes or consists of a sequence having at least 80% sequence identity to SEQ ID NO: 3, for example, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 3. In some examples, the 5’ SbMMV replicon backbone sequence includes or consists of a sequence having at least 95% sequence identity to SEQ ID NO: 3. In further examples, the 5’ SbMMV replicon backbone sequence includes or consists of a sequence having at least 98% sequence identity to SEQ ID NO: 3. In some examples, the 5’ SbMMV replicon backbone sequence includes or consists of SEQ ID NO: 3.
[0324] In some implementations, the SbMMV replicon includes a 3’ SbMMV replicon backbone sequence. In some examples, the 3’ SbMMV replicon backbone sequence includes or consists of a sequence having at least 80% sequence identity to SEQ ID NO: 4, for example, at least 85%, at least 90%, at least 95%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 4. In some examples, the 3’ SbMMV replicon backbone sequence includes or consists of a sequence having at least 95% sequence identity to SEQ ID NO: 4. In further examples, the 3’ SbMMV replicon backbone sequence includes or consists of a sequence having at least 98% sequence identity to SEQ ID NO: 4. In some examples, the 3’ SbMMV replicon backbone sequence includes or consists of SEQ ID NO: 4.
[0325] In a specific, non-limiting example, the SbMMV replicon includes or consists of from 5’ to 3’: a sequence having at least 90% sequence identity to SEQ ID NO: 3, an exogenous nucleic acid (e.g., repair template or an expression cassette), and a sequence having at least 90% sequence identity to SEQ ID NO: 4. In another example, the SbMMV replicon includes or consists of from 5’ to 3’ : a sequence having at least 95% sequence identity to SEQ ID NO: 3, an exogenous nucleic acid (e.g., repair template or an expression cassette), and a sequence having at least 95% sequence identity to SEQ ID NO: 4. In a further example, the SbMMV replicon includes or consists of from 5’ to 3’ : a sequence having at least 98% sequence identity to SEQ ID NO: 3, an exogenous nucleic acid (e.g., repair template or an expression cassette), and a sequence having at least 98% sequence identity to SEQ ID NO: 4. In some examples, the SbMMV replicon includes or consists of from 5’ to 3’ : SEQ ID NO: 3 or a degenerate variant thereof, an exogenous nucleic acid (e.g., repair template or an expression cassette), and SEQ ID NO: 4 or a degenerate variant thereof. A degenerate variant refers to a nucleic acid sequence variant due to degeneracy (redundancy) of the genetic code.
[0326] In some aspects, the SbMMV replicon includes the repair template. In some examples, the repair template includes an insert sequence flanked by a left homology arm (LHA) portion and a right homology arm (RHA) portion (see, e.g., FIG. 5). The LHA and RHA portions include a sequence complementary to a target nucleic acid to facilitate homology-dependent repair of the target. In some examples, the repair template includes or consists of from 5’ to 3’: a LHA portion complementary to a target nucleic acid, an insert sequence, and a RHA portion complementary to a target nucleic acid. In some examples, the insert sequence is inserted (e.g., incorporated) into a target nucleic acid upon homologous recombination with the LHA and RHA. In some examples, the LHA and / or RHA consist of, or essentially consist of, a sequence complementary to the target nucleic acid.
[0327] In some examples, the LHA includes at least 20 base pairs complementary to a target nucleic acid, for example, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1,000, at least 1,250, at least 1,500, at least 1,750, at least 2,000, at least 2,250, at least 2,500, or more, base pairs complementary to a target nucleic acid. In some examples, the LHA includes 20 to 3,000 base pairs complementary to a target nucleic acid, for example, 20 to 2,500, 20 to 2,000, 20 to 1,750, 20 to 1,500, 20 to 1,000, 20 to 750, 20 to 500, 20 to 250, 20 to 100, 20 to 50, 30 to 2,500, 30 to 2,000, 30 to 1,750, 30 to 1,500, 30 to 1,000, 30 to 750, 30 to 500, 30 to 250, 30 to 100, 30 to 50, 50 to 2,500, 50 to 2,000, 50 to 1,750, 50 to 1,500, 50 to 1,000, 50 to 750, 50 to 500, 50 to 250, 50 to 100, 100 to 2,500, 100 to 2,000, 100 to 1,750, 100 to 1,500, 100 to 1,000, 100 to 750, 100 to 500, 100 to 250, 100 to 100, 250 to 2,500, 250 to 2,000, 250 to 1,750, 250 to 1,500, 250 to 1,000, 250 to 750, 250 to 500, 400 to 2,500, 400 to 2,000, 400 to 1,750, 400 to 1,500, 400 to 1,000, 400 to 750, 400 to 500, 500 to 2,500, 500 to 2,000, 500 to 1,750, 500 to 1,500, 500 to 1,000, 500 to 750, 1000 to 2,500, 1000 to 2,000, 1000 to 1,750, 1000 to 1,500, 1,500 to 2,500, 1,500 to 2,000, or 1,500 to 1,750, base pairs complementary to a target nucleic acid. In a non-limiting example, the LHA includes 250 to 2,000 base pairs complementary to a target nucleic acid. In another non-limiting example, the LHA includes 500 to 2,000 base pairs complementary to a target nucleic acid.
[0328] In some examples, the RHA includes at least 20 base pairs complementary to a target nucleic acid, for example, at least 25, at least 30, at least 40, at least 50, at least 75, at least 100, at least 150, at least 200, at least 250, at least 300, at least 350, at least 400, at least 450, at least 500, at least 550, at least 600, at least 650, at least 700, at least 750, at least 800, at least 850, at least 900, at least 950, at least 1,000, at least 1,250, at least 1,500, at least 1,750, at least 2,000, at least 2,250, at least 2,500, or more, base pairs complementary to a target nucleic acid. In some examples, the RHA includes 20 to 3,000 base pairs complementary to a target nucleic acid, for example, 20 to 2,500, 20 to 2,000, 20 to 1,750, 20 to 1,500, 20 to 1,000, 20 to 750, 20 to 500, 20 to 250, 20 to 100, 20 to 50, 30 to 2,500, 30 to 2,000, 30 to 1,750, 30 to 1,500, 30 to 1,000, 30 to 750, 30 to 500, 30 to 250, 30 to 100, 30 to 50, 50 to 2,500, 50 to 2,000, 50 to 1,750, 50 to 1,500, 50 to 1,000, 50 to 750, 50 to 500, 50 to 250, 50 to 100, 100 to 2,500, 100 to 2,000, 100 to 1,750, 100 to 1,500, 100 to 1,000, 100 to 750, 100 to 500, 100 to 250, 100 to 100, 250 to 2,500, 250 to 2,000, 250 to 1,750, 250 to 1,500, 250 to 1,000, 250 to 750, 250 to 500, 400 to 2,500, 400 to 2,000, 400 to 1,750, 400 to 1,500, 400 to 1,000, 400 to 750, 400 to 500, 500 to 2,500, 500 to 2,000, 500 to 1,750, 500 to 1,500, 500 to 1,000, 500 to 750, 1000 to 2,500, 1000 to 2,000, 1000 to 1,750, 1000 to 1,500, 1,500 to 2,500, 1,500 to 2,000, or 1,500 to 1,750, base pairs complementary to a target nucleic acid. In a non-limiting example, the RHA includes 250 to 2,000 base pairs complementary to a target nucleic acid. In another non-limiting example, the RHA includes 500 to 2,000 base pairs complementary to a target nucleic acid.
[0329] In some implementations, the genomic target sequences of the LHA and RHA are adjacent, for example, when the repair template introduces a substitution or insertion. In some implementations, there is a gap sequence (a segment of genomic sequence that is not complementary to the LHA or RHA) between the genomic target sequences of the LHA and RHA, for example when the repair template is used to introduce a deletion and / or insertion. In some examples, the gap sequence is less than lOkb, for example, less than 9.5kb, less than 9kb, less than 8.5kb, less than 8kb, less than 7.5kb, less than 7kb, less than 6.5kb, less than 6kb, less than 5.5kb, less than 5kb, less than 4.5kb, less than 4kb, less than 3.5kb, less than 3kb, less than 2.5kb, less than 2kb, less than lOOObp, less than 900bp, less than 800bp, less than 700bp, less than 600bp, less than 500bp, less than 400bp, less than 300bp, less than 200bp, less than lOObp, or less than 50bp. In some examples, the gap sequence is 0 (the LHA and RHA genomic target sequences are contiguous in the genome) to lOkb, for example, 0 to 9kb, 0 to 8kb, 0 to 7kb, 0 to 6kb, 0 to 5kb, 0 to 4kb, 0 to 3kb, 0 to 2kb, 0 to lOObp, 0 to 50bp, or 0 to 25bp. The gap sequence in the genome is replaced with the repair template upon HDR.
[0330] In some implementations, the repair template is designed to introduce a modification into a target, such as an insertion, deletion, or one or more substitutions (e.g., point mutations), upon homologous recombination with the target. In some examples, the repair template increases or decreases expression of a target nucleic acid upon homologous recombination. In some examples, the repair template introduces a modification that disrupts, restores, or otherwise changes the function of a protein encoded by a target nucleic acid upon homologous recombination.
[0331] In some implementations, the insert sequence is used to introduce one or more mutations (e.g., substitutions) in a target nucleic acid, for example, the insert sequence can be homologous to a target except at least one mismatched base pair relative to a target nucleic acid, for example, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 15, at least 20, at least 25, at least 50, or more, mismatched base pairs relative to a target nucleic acid. In some examples, the insert sequence includes a sequence having 0% to 99% sequence identity to a target sequence, for example, 0% to 95%, 0% to 90%, 0% to 80%, 0% to 75%, 0% to 70%, 0% to 60%, 0% to 55%, 0% to 50%, 0% to 45%, 0% to 40%, 0% to 35%, 0% to 30%, 0% to 25%, 0% to 20%, 0% to 15%, 0% to 10%, 5% to 95%, 5% to 90%, 5% to 80%, 5% to 75%, 5% to 70%, 5% to 60%, 5% to 55%, 5% to 50%, 5% to 45%, 5% to 40%, 5% to 35%, 5% to 30%, 5% to 25%, 5% to 20%, 5% to 15%, 5% to 10%, 10% to 95%, 10% to 90%, 10% to 80%, 10% to 75%, 10% to 70%, 10% to 60%, 10% to 55%, 10% to 50%, 10% to 45%, 10% to 40%, 10% to 35%, 10% to
[0332] 30%, 10% to 25%, 10% to 20%, 10% to 15%, 10% to 10%, 25% to 95%, 25% to 90%, 25% to 80%, 25% to
[0333] 75%, 25% to 70%, 25% to 60%, 25% to 55%, 25% to 50%, 25% to 45%, 25% to 40%, 25% to 35%, 25% to
[0334] 30%, 50% to 95%, 50% to 90%, 50% to 80%, 50% to 75%, 50% to 70%, 50% to 60%, 75% to 95%, 75% to
[0335] 90%, 75% to 80%, 90% to 99%, 91% to 99%, 92% to 99%, 93% to 99%, 94% to 99%, 95% to 99%, 96% to
[0336] 99%, 97% to 99%, or 98% to 99% sequence identity to a target sequence.
[0337] The insert sequence can be used to introduce an exogenous nucleic acid into a target, for example, the insert sequence is from a source other than the target nucleic acid (e.g., the insert sequence is a synthetic sequence and the target nucleic acid is a soybean genomic nucleic acid). In some examples, the insert sequence includes a transgene, regulatory element (e.g., promoter, terminator, enhancer, or other regulatory element), and / or a selectable marker (e.g., a resistance gene or fluorescent marker). In some examples, the insert sequence includes a promoter and / or transcription terminator.
[0338] Suitable promoters include constitutive, conditional, inducible, and temporally or spatially specific promoters (e.g., a tissue specific promoter, a developmentally regulated promoter, or a cell cycle regulated promoter). Exemplary promoters include, but are not limited to, RNA polymerase II promoter (Pol II), type HI RNA polymerase III promoter (Pol III) (e.g., U6 (e.g., Arabidopsis thaliana U6 (AtU6) or maize, tomato, or soybean U6, such as those disclosed in PCT / US2015 / 018104)), ubiquitin promoter (e.g., Solatium lycopersicum (SlUbilO), Arabidopsis thaliana (AtUbilO)), figwort mosaic virus (FMV) promoter, RUBISCO promoter, or pyruvate phosphate dikinase (PDK) promoter. Non-limiting examples of constitutive promoters include CaMV 35S promoter, as disclosed in US Patents 5,858,742 and 5,322,938, rice actin promoter as disclosed in US Patent 5,641,876, maize chloroplast aldolase promoter as disclosed in US Patent 7,151,204, and opaline synthase (NOS) and octapine synthase (OCS) promoter from Agrobacterium tumefaciens.
[0339] Suitable transcription terminators include terminators that are functional in plant and / or bacterial cells. Exemplary terminators include, but are not limited to, HSPt (e.g., Arabidopsis thaliana HSPt (AtHSPt)), nopaline synthase (NOS) terminator, Cauliflower Mosaic Virus (CaMV) 35S terminator, rbcs E9 terminator (e.g., pea rbcs E9), and U6 poly-T terminator.
[0340] Suitable exemplary selectable markers include, but are not limited to, resistance markers (e.g., herbicide or antibiotic resistance), fluorescent markers (e.g., GFP, YFP, RFP, etc.), bioluminescent markers (e.g., luciferase), betalain markers (e.g., RUBY, see, He et al. “A reporter for noninvasively monitoring gene expression and plant transformation” Horticulture Research, vol. 7 art. 152, 2020), and other reporters (e.g., % / a-glucuronidase (GUS)). In some implementations, the insert sequence includes a selectable marker.
[0341] In some aspects, the SbMMV replicon includes an expression cassette. The expression cassette can be used to produce or express a nucleic acid and / or protein of interest in a host cell (e.g., a plant cell, such as a soybean cell). In some implementations, the expression cassette encodes or includes an RNA (e.g., CRISPR-RNA, siRNA), gene, regulatory element (e.g., promoter, transcription terminator, or other regulatory element), and / or selectable marker (e.g., a resistance gene or fluorescent marker). Exemplary regulatory elements and selectable markers are disclosed herein. The gene can be any gene of interest. In some examples, the gene is a transgene.
[0342] In some aspects, the SbMMV replicon can further include a nucleic acid encoding replication initiator protein (Rep) and / or RepA. In some examples, the Rep and / or RepA are from SbMMV. In some examples, the SbMMV replicon includes a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 2; for example, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 2. In some examples, the SbMMV replicon includes a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 2. In some examples, the SbMMV replicon includes a nucleic acid sequence having at least 95% sequence identity to SEQ ID NO: 2. In some examples, the SbMMV replicon includes a nucleic acid sequence having at least 98% sequence identity to SEQ ID NO: 2. In some examples, the SbMMV replicon includes SEQ ID NO: 2, or a degenerate variant thereof.
[0343] In some implementations, the SbMMV replicon includes a SbMMV LIR, an exogenous nucleic acid (e.g., a repair template or an expression cassette), and a sequence encoding SbMMV Rep / RepA. In some implementations, the SbMMV replicon includes or consists of (a) a 5’ and / or 3’ SbMMV LIR (relative to b), (b) an exogenous nucleic acid (e.g., a repair template or expression cassette), and (c) a sequence encoding SbMMV Rep / RepA. In some implementations, the SbMMV replicon includes or consists of, from 5’ to 3’, (a) a 5’ SbMMV LIR, (b) an exogenous nucleic acid (e.g., a repair template or expression cassette), (c) a sequence encoding SbMMV Rep / RepA, and (d) a 3’ SbMMV LIR.
[0344] In some aspects, Rep / RepA as described above is included in a nucleic acid disclosed herein, but is not within a region encoding the SbMMV replicon. In some examples, Rep / RepA are encoded by a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 2; for example, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 2. In some examples, Rep / RepA are encoded by a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 2. In some examples, Rep / RepA are encoded by a nucleic acid sequence having at least 95% sequence identity to SEQ ID NO: 2. In some examples, a sequence encoding Rep / RepA includes or consists of SEQ ID NO: 2, or a degenerate variant thereof. A degenerate variant refers to a nucleic acid sequence variant due to degeneracy (redundancy) of the genetic code. While the nucleic acid sequence of degenerate variants vary, they encode the same amino acid sequence.
[0345] The nucleic acid can further include a nuclease (e.g., zinc finger nuclease (ZFN), Argonaute protein, transcription activator-like effector nuclease (TALEN) or a CRISPR-associated (Cas) nuclease). In some examples, the nuclease is an endonuclease. In some examples, the nuclease is codon optimized for expression in a plant cell. In some examples, the nucleic acid includes a SbMMV replicon including a repair template, and the nucleic acid further includes a nuclease. In some implementations, the nuclease induces a break in a target nucleic acid sequence, and a repair template disclosed herein facilitates repair of the break via HDR.
[0346] CRISPR technology (including, e.g., Cas9 (e.g., SpCas9), Cpfl (Casl2a), type I multisubunit systems, Casl2d, Casl2e, Casl4 / Casl2f, Casl2i, and Casl2j) have been described, for example, in US Patent Application Publications 2016 / 0138008A1 and US2015 / 0344912A1, and in US Patents 8,697,359, 8,771,945, 8,945,839, 8,999,641, 8,993,233, 8,895,308, 8,865,406, 8,889,418, 8,871,445, 8,889,356, 8,932,814, 8,795,965, 8,906,616, and Koonin et al. “Discovery of Diverse CRISPR-Cas Systems and Expansion of the Genome Engineering Toolbox.” Biochemistry, doi: 10.1021 / acs.biochem. 3c00159, Epub ahead of print, May 16, 2023. Cpfl (Casl2a) endonuclease and corresponding guide RNAs and PAM sites are disclosed in US Patent Application Publication 2016 / 0208243 Al. Other CRISPR nucleases useful for editing genomes include C2cl and C2c3 (see, Shmakov et al. (2015) Mol. Cell, 60:385 - 397) and CasX and CasY (see, Burstein et al. (2016) Nature, 542: 237-241). In some examples, the nuclease is a Cas nuclease (e.g., a type II Cas nuclease, a Cas9, a type V Cas nuclease, a Cpfl, a CasY, a CasX, a C2cl, a C2c3). In some implementations, the Cas nuclease is a plant-codon-optimized Cas9 (pco-Cas9) from Streptococcus pyogenes and S. thermophilus as disclosed in PCT / US2015 / 018104. In some implementations, the nucleic acid further includes a guide RNA (gRNA) to direct a Cas nuclease to a particular target. Plant RNA promoters for expressing CRISPR guide RNA and plant codon-optimized CRISPR Cas9 endonuclease are disclosed, for example, in International Patent Application PCT / US2015 / 018104. Methods of using CRISPR technology for genome editing in plants are discussed, for example, in US Patent Application Publications US 2015 / 0082478A1 US 2015 / 0059010A1, and Koonin et al. Biochemistry, 2023, doi: 10.1021 / acs.biochem.
[0347] Zinc finger nucleases (ZFNs) are engineered proteins comprising a zinc finger DNA-binding domain fused to a nucleic acid cleavage domain. The zinc finger binding domain provides specificity and can be engineered to specifically recognize any desired target DNA sequence. The construction and use of ZFNs in plants and other organisms has been described, for example, in Urnov et al. (2010) Nature Rev. Genet., 11:636 - 646. The zinc finger DNA binding domains are derived from the DNA-binding domain of a large class of eukaryotic transcription factors called zinc finger proteins (ZFPs). The DNA-binding domain of ZFPs typically contains a tandem array of at least three zinc “fingers” each recognizing a specific triplet of DNA. Numerous strategies can be used to design the binding specificity of the zinc finger binding domain. One approach, termed “modular assembly,” relies on the functional autonomy of individual zinc fingers with DNA. In this approach, a given sequence is targeted by identifying zinc fingers for each component triplet in the sequence and linking them into a multifinger peptide. Several alternative strategies for designing zinc finger DNA binding domains have also been developed. These methods are designed to accommodate the ability of zinc fingers to contact neighboring fingers as well as nucleotides bases outside their target triplet. Typically, the engineered zinc finger DNA binding domain has a novel binding specificity, compared to a naturally-occurring zinc finger protein. Engineering methods include, for example, rational design and various types of selection. Rational design includes, for example, the use of databases of triplet (or quadruplet) nucleotide sequences and individual zinc finger amino acid sequences, in which each triplet or quadruplet nucleotide sequence is associated with one or more amino acid sequences of zinc fingers which bind the particular triplet or quadruplet sequence. See, e. g., US Patents 6,453,242 and 6,534,261. Exemplary selection methods (e.g., phage display and yeast two-hybrid systems) are well known and described. In addition, enhancement of binding specificity for zinc finger binding domains has been described in US Patent 6,794,136. In addition, individual zinc finger domains may be linked together using any suitable linker sequences. Examples of linker sequences are publicly known, see, e.g., US Patents 6,479,626; 6,903,185; and 7,153,949. The nucleic acid cleavage domain is non-specific and is typically a restriction endonuclease, such as Fokl. The endonuclease dimerizes to cleave DNA, thus, cleavage by FokI as part of a ZFN requires two adjacent and independent binding events to permit dimer formation. The requirement for two DNA binding events enables more specific targeting of long and potentially unique recognition sites. Fokl variants with enhanced activities have been described; see, e. g., Guo et al. (2010) J. Mol. Biol., 400:96 - 107.
[0348] Transcription activator like effectors (TAEEs) are DNA binding proteins found in pathogenic Xanthomonas sp. that can modulate gene expression in host plants and facilitate the colonization and survival of the bacterium. TALEs act as transcription factors and modulate expression of resistance genes in the plants. Recent studies of TALEs have revealed the code linking the repetitive region of TALEs with their target DNA-binding sites. TALEs comprise a highly conserved and repetitive region consisting of tandem repeats of mostly 33 or 34 amino acid segments. The repeat monomers differ from each other mainly at amino acid positions 12 and 13. A strong correlation between unique pairs of amino acids at positions 12 and 13 and the corresponding nucleotide in the TALE-binding site has been found. The simple relationship between amino acid sequence and DNA recognition of the TALE binding domain allows for the design of DNA binding domains of any desired specificity. TALEs can be linked to a non-specific DNA cleavage domain to prepare genome editing proteins, referred to as TAL-effector nucleases, or TALENs. As in the case of ZFNs, a non-specific restriction endonuclease, such as Fokl, can be used. For a description of the use of TALENs in plants, see Mahfouz et al. (2011) Proc. Natl. Acad. Sci. USA, 108:2623 - 2628 and Mahfouz (2011) GM Crops, 2:99 - 103.
[0349] Argonautes are proteins that can function as sequence-specific endonucleases by binding a polynucleotide (e. g., a single-stranded DNA or single-stranded RNA) that includes sequence complementary to a target nucleotide sequence) that guides the Argonaut to the target nucleotide sequence and effects site-specific alteration of the target nucleotide sequence; see, e. g., US Patent Application Publication 2015 / 0089681.
[0350] The nucleic acid can further include a regulatory element (e.g., a promoter or transcription terminator), nuclease (e.g., a Cas nuclease), and / or selectable marker. Suitable regulatory elements, nucleases, and selectable markers are described herein. In some implementations, the nucleic acid further includes a transfer DNA border sequence. In some examples, the nucleic acid includes a 5’ T-DNA border sequence (also known as left border (LB)) and a 3’ T-DNA border sequence (also known as right border (RB)) relative to the SbMMV replicon. Exemplary T-DNA border sequences include, but at not limited to, SEQ ID Nos: 6 and 7. In some implementations, the 5’ T-DNA border sequence comprises a sequence having at least 90% sequence identity to SEQ ID NO: 6, for example, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 6. In a non-limiting example, the 5’ T-DNA border sequence comprises or consists of SEQ ID NO: 6. In some implementations, the 3’ T-DNA border sequence comprises a sequence having at least 90% sequence identity to SEQ ID NO: 7, for example, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 7. In a non-limiting example, the 3’ T-DNA border sequence comprises or consists of SEQ ID NO: 7.
[0351] The SbMMV replicon or nucleic acid can further include spacer and / or linker sequences between discrete elements disclosed herein, for example, nucleic acid sequences that maintain a desired spacing and / or reading frame. The SbMMV replicon can include spacer and / or linker sequences, for example, adjacent to the LIR, insert sequence, expression cassette, Rep / RepA, the 3’ backbone sequence, and / or 5’ backbone sequence. The nucleic acid can include a spacer and / or linker sequence, for example, adjacent to the SbMMV replicon, regulatory elements, and / or selectable markers, or other genes or elements.
[0352] In some examples, the exogenous nucleic acid, expression cassette, or SbMMV replicon includes a self-cleaving ribozyme, for example, a hammerhead (HH) or hepatitis delta virus (HDV) ribozyme. In some examples, an RNA, for example a guide RNA or crRNA, is flanked by self-cleaving ribozymes (e.g., a 5’ and 3’ ribozyme relative to the RNA).
[0353] Also provided are vectors that include a nucleic acid disclosed herein. Vectors typically comprise one or more regulatory elements, such as a promoter, operably linked to one or more polynucleotides (such as a nucleic acid disclosed herein). Expression of such polynucleotides can be controlled by promoter selection, such as constitutive, conditional, inducible, or temporally or spatially specific promoters (e.g., a tissue specific promoter, a developmentally regulated promoter, or a cell cycle regulated promoter). A vector can also include nucleic acid sequences that permit it to replicate in a host cell (e.g., a soybean cell), such as an origin of replication. An integrating vector is capable of integrating itself into a host nucleic acid. An expression vector is a vector that contains the necessary regulatory sequences to allow transcription and translation of inserted gene or genes.
[0354] One type of vector is a “plasmid,” which refers to a circular double-stranded DNA loop into which additional DNA segments can be inserted, such as by standard molecular cloning techniques. Another type of vector is a viral vector, wherein viral-derived DNA or RNA sequences are present in the vector for packaging into a virus (e.g., retroviruses, replication defective retroviruses, adenoviruses, replication defective adenoviruses, and adeno-associated viruses). Viral vectors also include polynucleotides carried by a virus for transfection into a host cell. Certain vectors are capable of autonomous replication in a host cell into which they are introduced (e.g., bacterial vectors having a bacterial origin of replication). Other vectors are integrated into the genome of a host cell upon introduction into the host cell and, thereby, are replicated along with the host genome.
[0355] In some examples, the vector is a plant expression vector. Expression vectors can include a DNA segment near the 3' end of an expression cassette that acts as a signal to terminate transcription and directs polyadenylation of a resultant mRNA. These are commonly referred to as “3'-untranslated regions” or “3'- UTRs” or “polyadenylation signals”. Useful 3' elements include: Agrobacterium tumefaciens nos 3', tml 3', tmr 3', tms 3', ocs 3', and tr7 3' elements disclosed in U.S. Pat. No. 6,090,627, and 3' elements from plant genes such as the heat shock protein 17, ubiquitin, and fructose- 1,6-biphosphatase genes from wheat (Triticum aestivum), and the glutelin, lactate dehydrogenase, and beta-tubulin genes from rice (Oryza sativa). In some implementations, the vector includes additional elements for improving delivery of a polynucleotide to a plant cell (including a plant protoplast), for example, a cell-penetrating peptide, localization signal, transit, or targeting peptide; or sequences to stabilize RNAs (such as a guide RNA).
[0356] In some examples, the vector is a transfer DNA (T-DNA) or vector suitable for biolistics (gene gun). T-DNA vectors are artificial vectors derived from the naturally occurring Ti plasmid found in bacterial species of Agrobacterium. Typically, T-DNA vectors are plasmids. Exemplary T-DNA vector series include, but are not limited to, pBIN (e.g., pBIN19), pPVP (e.g., pPZP200), pCB (e.g., pCB301), pCAMBIA (e.g., pCAMBIA-1 00), pGreen (e.g., pGreenOOOO), pLSU (e.g., pLSU-1) and pLX (e.g., pLX-B2). In binary systems, a vir helper plasmid (e.g., EHA101, EHA105, AGL-1, LBA4404, and GV2260), contains the vir genes of the Ti plasmid of Agrobacterium, and is co-transformed with a T-DNA vector to facilitate transfer and integration of a DNA of interest.
[0357] In some implementations, a nucleic acid disclosed herein is directly introduced into a plant cell by biolistic-based transformation (e.g., gene gun). Biolistic-based transformation includes methods of delivering nucleic acids into cells by high-speed particle bombardment, for example, a particle (e.g. , gold or tungsten or magnetic particle) coated in DNA is delivered into a cell by a particle gun accelerator or gene gun, or by magnetic force. The size of particles used in biolistics is generally in the “microparticle” range, for example, gold microcarriers in the 0.6, 1.0, and 1.6 micrometer size ranges (see, e.g., instruction manual for the Helios® Gene Gun System, Bio-Rad, Hercules, CA), however, successful biolistics delivery using larger (40 nanometer) nanoparticles has also been reported, see, e.g., O’Brian and Lummis (2011) BMC Biotechnol., 11:66 - 71. Biolistic methods can be used to introduce circular, linear, double-stranded, or single-stranded nucleic acids.
[0358] Further provided are cells (such as a host cell) including a nucleic acid or vector disclosed herein. In some implementations, the cell is a bacterial cell, such as an agrobacterium cell. In some implementations, the cell is a plant cell (including protoplasts), such as a soybean cell (Glycine max). Plants or plant parts (including callus) including a cell disclosed herein are included in the disclosure, for example, a plant or plant part regenerated from a cell transformed with a nucleic acid or vector disclosed herein. IV. Methods
[0359] Also disclosed are methods of producing or expressing a nucleic acid in a plant cell (including protoplasts, or plant cells located in an intact plant, plant part, or plant tissue), including transforming (introducing) a recombinant nucleic acid or vector including or consisting of a SbMMV replicon into a plant cell, thereby generating a transformed plant cell. In some examples, the SbMMV replicon includes an expression cassette. In some examples, the SbMMV replicon including an expression cassette is used to express an exogenous gene in a host plant (e.g., soybean plant). Nucleic acid production or expression in the transformed plant cell can be stable or transient. In some examples, the nucleic acid is produced or expressed transiently. In some examples, the plant cell is a soybean cell (Glycine max).
[0360] Also disclosed are methods of modifying genetic material of a plant cell (including protoplasts, or plant cells located in an intact plant, plant part, or plant tissue), including transforming the plant cell with a nucleic acid or vector including a SbMMV replicon, thereby generating a transformed plant cell. In some examples, the SbMMV replicon includes a repair template. In some examples, transforming the plant cell with the nucleic acid or vector disclosed herein produces a transformed plant cell containing an induced genetic modification (e.g., an insertion, deletion, or substitution). In some implementations, the modified genetic material is genomic DNA. In some implementations, modifying the genetic material of the plant cell includes introducing an insertion, deletion, or substitution into a target nucleic acid. In some implementations, modifying the genetic material is facilitated by a homology-dependent repair (HDR) mechanism. In some examples, the plant cell is a soybean cell (Glycine max).
[0361] In some examples, the method of modifying genetic material of a plant cell includes transforming the plant cell with a nucleic acid or vector including a SbMMV replicon including a repair template, as disclosed herein. In some examples, the method further includes introducing a nuclease, for example, a nuclease capable of inducing a double-stranded DNA break that can be repaired via HDR with the repair template. Thus, in some implementations, the repair template and nuclease target the same gene or loci. In some examples, the method further includes introducing a guide RNA (gRNA), for example, to guide a Cas nuclease to a particular target gene or loci. The nucleic acid or vector including the SbMMV replicon, the nuclease, and / or the gRNA can all be included on the same vector, or different vectors. In some examples, a host plant (e.g., a soybean plant) is engineered to express an inducible nuclease and / or gRNA, which is induced following transformation with the nucleic acid or vector including the SbMMV replicon. While the nucleic acid or vector including the SbMMV replicon, the nuclease, and / or the gRNA, need not be simultaneously introduced, in some examples they are introduced in the plant cell simultaneously or at substantially the same time. In some examples, the nucleic acid or vector including the SbMMV replicon also includes a Cas nuclease and a gRNA. In some implementations, the methods further include transforming the plant cell with a vir plasmid, or other replicon containing the vir genes from Agrobacterium. Transforming includes any suitable method of introducing a nucleic acid or vector into a host cell. In some implementations, a plant cell is transformed with a nucleic acid or vector disclosed herein by bacterial-mediated (e.g., Agrobacterium sp., Rhizobium sp., Sinorhizobium sp., Mesorhizobium sp., Bradyrhizobium sp., Azobacter sp., Phyllobacterium sp.) transformation. In some examples, the plant cell is transformed via agrobacterium-mediated transformation. In some implementations, the plant cell is transformed with a nucleic acid or vector disclosed herein by biolistic-based transformation (high speed particle bombardment, e.g., a gene gun), for example, a particle (e.g., gold or tungsten or magnetic particle) coated in DNA is delivered by a biolistic-type technique or with magnetic force. The size of the particles used in biolistics is generally in the “microparticle” range, for example, gold microcarriers in the 0.6, 1.0, and 1.6 micrometer size ranges (see, e.g., instruction manual for the Helios® Gene Gun System, Bio-Rad, Hercules, CA), however, successful biolistics delivery using larger (40 nanometer) nanoparticles has been reported, see. e.g., O’Brian and Lummis (2011) BMC Biotechnol., 11 :66 - 71. Biolistic methods can be used to introduce circular, linear, double-stranded, or single-stranded nucleic acids. In some examples, the methods include transforming a plant cell with a nucleic acid disclosed herein, wherein the nucleic acid is single stranded and introduced by biolistic transformation.
[0362] In some examples, production or expression of an exogenous nucleic acid (e.g., repair template or expression cassette) achieved with the SbMMV replicon disclosed herein is significantly higher than production or expression using other geminivirus replicons (e.g., SbBMV or BeYDV). In some examples, production or expression is at least 25% higher, for example, at least 30%, at least 40%, at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 100%, at least 125%, at least 150%, at least 175%, at least 200%, at least 225%, at least 250%, at least 275%, at least 300%, at least 400%, at least 500%, at least 600%, at least 700%, at least 800%, at least 900%, at least 1000% higher, or more, relative to other geminivirus replicons (e.g., SbBMV or BeYDV). In some examples, production or expression is at least two-fold greater than production or expression achieved using other geminivirus replicons (e.g., SbBMV or BeYDV), such as at least 3-fold greater, at least 4-fold greater, at least 5-fold greater, at least 6-fold greater, at least 7-fold greater, at least 8-fold greater, at least 9-fold greater, at least 10-fold greater, or more, than production or expression achieved using other geminivirus replicons (e.g., SbBMV or BeYDV).
[0363] In some implementations, the plant cell (including protoplasts) is an isolated plant cell (e.g., a plant cell isolated from a whole plant or plant part or tissue, or a plant cell in suspension or plate culture). In some examples, the isolated plant cell is obtained or isolated from a whole plant, or plant part or tissue, for example (and without limitation), an intact nodal bud, a shoot apex or shoot apical meristem, a root apex or root apical meristem, lateral meristem, intercalary meristem, a seedling (e.g., a germinating seed or small seedling or a larger seedling with one or more true leaves), a whole seed (e.g., an intact seed, or a seed with part or all of its seed coat removed or treated to make permeable), a halved seed or other seed fragment, an embryo (e.g., a mature dissected zygotic embryo, a developing embryo, a dry or rehydrated or freshly excised embryo), or callus. In some implementations, the plant cell is not an isolated plant cell, for example, the plant cell is located in an intact or growing plant or in a plant part or tissue. In such examples, the method is performed in situ or in planta.
[0364] Generally, transformed plant cells (including protoplasts) are capable of division and further differentiation. In some examples, a transformed plant cell is regenerated to a whole plant (e.g., a transformed soybean cell is regenerated to a soybean plant). In some examples, the methods further include one or more steps of growing or regenerating a plant from a transformed plant cell, for example, a transformed plant cell including, producing or expressing a nucleic acid (e.g., a nucleic acid including a SbMMV replicon), or including an induced genetic modification as disclosed herein, thereby generating a modified plant. In such examples, the grown or regenerated plant contains at least some cells or tissues producing or expressing the transformed nucleic acid or including the induced genetic modification. In some examples, a callus is produced from the transformed plant cell, and plantlets and plants are produced from the callus. In other examples, whole seedlings or plants are grown directly from the transformed plant cell without a callus stage. Thus, additional aspects of the disclosure are directed to whole seedlings and plants grown or regenerated from transformed plant cells produced by the methods disclosed herein, as well as the seeds of such plants. In some examples, the grown or regenerated plant exhibits a phenotype associated with expression of the nucleic acid or the induced genetic modification. Non-limiting phenotypes include herbicide resistance, improved tolerance of abiotic stress (e.g., tolerance of temperature extremes, drought, or salt) or biotic stress (e. g., resistance to bacterial or fungal pathogens), improved utilization of nutrients or water, modified lipid, carbohydrate, or protein composition, improved flavor or appearance, improved storage characteristics (e.g., resistance to bruising, browning, or softening), increased yield, altered morphology (e.g., floral architecture or color, plant height, branching, root structure), or expression of a selectable marker.
[0365] The methods can include a selection step of selecting transformed plant cells (or seedlings or plants grown or regenerated therefrom) with a desired phenotype, for example, transformed plant cells (or seedlings or plants) can be exposed to conditions permitting expression of a phenotype of interest; e.g., selection for herbicide resistance can include exposing the population of plant cells (or seedlings or plants) to an amount of herbicide or other substance that inhibits growth or is toxic, allowing identification and selection of those resistant plant cells (or seedlings or plants) that survive treatment. Plant cells (or seedlings or plants grown or regenerated therefrom) can be selected based on manifestation of a desired phenotype. Such plants can be selected, for example, for further analysis or plant breeding.
[0366] The plant cell can be haploid, diploid, or polyploid. In some examples, the plant cell is haploid or can be induced to become haploid. Examples of haploid cells include but are not limited to plant cells obtained from haploid plants and plant cells obtained from reproductive tissues, e.g., from flowers, developing flowers or flower buds, ovaries, ovules, megaspores, anthers, pollen, and microspores. In some examples where the plant cell is haploid, the method of modifying the genetic material of the plant cell can further include a step of chromosome doubling (e.g., by spontaneous chromosomal doubling by meiotic non- reduction, or by using a chromosome doubling agent such as colchicine, oryzalin, or trifluralin) to produce a doubled haploid plant cell that is homozygous for the induced genetic modification. Thus, aspects of the disclosure are related to haploid plant cells having the altered target nucleotide sequence as well as a doubled haploid plant cells or a doubled haploid plant that is homozygous for an induced genetic modification. Another aspect of the disclosure is related to a hybrid plant having at least one parent plant that is a doubled haploid plant provided by the method. Production of doubled haploid plants by these methods provides homozygosity in one generation, instead of requiring several generations of self-crossing to obtain homozygous plants; this may be particularly advantageous in slow-growing plants, such as fruit and other trees, or for producing hybrid plants that are offspring of at least one doubled-haploid plant.
[0367] The plant cell can be obtained from a dicot or a monocot plant. Non-limiting plants include row crop plants, fruit-producing plants and trees, vegetables, trees, and ornamental plants including ornamental flowers, shrubs, trees, groundcovers, and turf grasses. Specific, non-fimiting exampfes of commercially important plants include: alfalfa (Medicago saliva), almonds (Primus dulcis), apples (Malus x domestica), apricots (Primus armeniaca, P. brigantine, P. mandshurica, P. mume, P. sibirica), asparagus (Asparagus officinalis), bananas (Musa spp.), barley (Hordeum vulgare), beans (Phaseolus spp.), blueberries and cranberries (Vaccinium spp.), cacao (Theobroma cacao), canola and rapeseed or oilseed rape, (Brassica napus), carnation (Dianthus caryophyllus), carrots (Daucus carota sativus), cassava (Manihot esculentum), cherry (Primus avium), chickpea (Cider arietinum), chicory (Cichorium intybus), chili peppers and other capsicum peppers (Capsicum annum, C. frutescens, C. chinense, C. pubescens, C. baccatum), chrysanthemums (Chrysanthemum spp.), coconut (Cocos nucifera), coffee (Coffea spp. including Coffea arabica and Coffea canephora), cotton (Gossypium hirsutum L.), cowpea (Vigna unguiculata), cucumber (Cucumis sativus), currants and gooseberries (Ribes spp.), eggplant or aubergine Solanum melongena), eucalyptus (Eucalyptus spp.), flax (Linum usitatissumum L.), geraniums (Pelargonium spp.), grapefruit (Citrus x paradisi), grapes (Vitus spp.) including wine grapes (Vitus vinifera), guava (Psidium guajava), irises (Iris spp.), lemon (Citrus Union), lettuce (Lactuca saliva), limes (Citrus spp.), maize (Zea mays L.), mango (Mangifera indica), mangosteen (Garcinia mangostana), melon (Cucumis melo), millets (Setaria spp., Echinochloa spp., Eleusine spp, Panicum spp., Pennisetum spp.), oats (Avena sativa), oil palm (Ellis quineensis), olive (Olea europaea), onion (Allium cepa), orange (Citrus sinensis), papaya (Carica papaya), peaches and nectarines (Prunus persica), pear (Pyrus spp.), pea (Pisa sativum), peanut (Arachis hypogaea), peonies (Paeonia spp.), petunias (Petunia spp.), pineapple (Ananas comosus), plantains (Musa spp.), plum (Prunus domestica), poinsettia (Euphorbia pulcherrima), Polish canola (Brassica rapa), poplar (Populus spp.), potato (Solanum tuberosum), pumpkin (Cucurbita pepo), rice (Oryza sativa L.), roses (Rosa spp.), rubber (Hevea brasiliensis), rye (Secale cereale), safflower (Carthamus tinctorius L), sesame seed (Sesame indium), sorghum (Sorghum bicolor), soybean (Glycine max), squash (Cucurbita pepo), strawberries (Fragaria spp., Fragaria x ananassa), sugar beet (Beta vulgaris), sugarcanes (Saccharum spp.), sunflower (Helianthus annus), sweet potato (Ipomoea batatas), tangerine (Citrus tangerina), tea (Camellia sinensis), tobacco (Nicotiana tabacum L.), tomato (Lycopersicon esculentum), tulips (Tulipa spp.), turnip (Brassica rapa rapa), walnuts (Juglans spp. L.), watermelon (Citrulus lanatus), wheat (Triticum aestivum), and yams (Discorea spp.). In a non-limiting example, the plant cell is a soybean cell (Glycine max).
[0368] In some examples, the plant cell is obtained from a crop plant characterized as being of or derived from an “elite” germplasm or genetic background, for example, from an inbred crop plant that is an elite strain of germplasm, or from a hybrid crop plant that is the progeny of at least one elite strain of germplasm (e.g., progeny of an inbred male parent of a first elite strain and an inbred female parent of a second elite strain). As used herein, an “elite” strain or line of a crop plant is one that has resulted from usually multiple rounds of breeding and selection for superior performance, e.g., superior yield or other agronomic trait. As used herein, “line” or “strain” includes plants that share identical parentage and are generally inbred to some degree, and which are generally homozygous at most genetic loci. Plants of a given line or strain exhibit a consistent and predictable phenotype and agronomic performance. A plant is “homozygous” when it has only one type of allele at a given locus, e.g., a diploid plant with two identical copies of an allele at a given locus.
[0369] V. Transgenic Plants
[0370] Also disclosed herein are transgenic plants including an exogenous nucleic acid (e.g., exogenous to the plant). In some implementations, the exogenous nucleic acid includes or consists of SbMMV replication initiator protein (Rep) and RepA. A transgenic plant including SbMMV Rep and RepA is useful, for example, for facilitating replication of Geminivirus replicons in the transgenic plant. A transgenic plant including SbMMV Rep and RepA can be further transformed with a Geminivirus replicon (e.g., a SbMMV replicon disclosed herein) that includes or does not include Rep and / or RepA. In some implementations, the exogenous nucleic acid includes or consists of a nucleic acid disclosed herein (e.g., a nucleic acid including a SbMMV replicon disclosed herein). In some examples, the transgenic plant is a transgenic soybean plant (Glycine max).
[0371] In some examples, the SbMMV Rep and RepA is encoded by a nucleic acid sequence having at least 80% sequence identity to SEQ ID NO: 2; for example, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to SEQ ID NO: 2. In some examples, the SbMMV Rep and RepA is encoded by a nucleic acid sequence having at least 90% sequence identity to SEQ ID NO: 2. In some examples, the SbMMV Rep and RepA is encoded by a nucleic acid sequence having at least 95% sequence identity to SEQ ID NO: 2. In some examples, the SbMMV Rep and RepA is encoded by a nucleic acid sequence having at least 98% sequence identity to SEQ ID NO: 2. In some examples, the SbMMV Rep and RepA is encoded by a nucleic acid including or consisting of SEQ ID NO: 2, or a degenerate variant thereof.
[0372] In some implementations, the transgenic plant further includes an endonuclease. Exemplary endonucleases are described herein (e.g., zinc finger nuclease (ZFN), Argonaute protein, transcription activator-like effector nuclease (TALEN) or a CRISPR-associated (Cas) nuclease). In some examples, the transgenic plant includes a Cas nuclease, for example, a Cas9 (e.g., SpCas9), Cpfl (Casl2a), type I multisubunit systems, Casl2d, Casl2e, Casl4 / Casl2f, Casl2i, or Casl2j. In some implementations, the transgenic plant further includes a guide RNA.
[0373] Also encompassed by this disclosure is progeny of a transgenic plant, such as progeny of a transgenic plant transformed with a nucleic acid disclosed herein (e.g., SbMMV Rep and RepA, or a SbMMV replicon). The term “progeny” refers to any subsequent generation of a plant. Progeny is measured using the following nomenclature: Fl refers to the first generation progeny, F2 refers to the second generation progeny, F3 refers to the third generation progeny, and so on. In some examples, the progeny is Fl. A progeny plant can be obtained, for example, by cloning or breeding (e.g., self-pollination of a parent plant, or crossing two distinct parent plants) a plant disclosed herein.
[0374] Clauses
[0375] (Clause 1) A nucleic acid comprising a recombinant soybean mild mottle virus (SbMMV) replicon, wherein the SbMMV replicon comprises: (i) an exogenous nucleic acid, wherein the exogenous nucleic acid is exogenous to SbMMV; and (ii) a sequence comprising SbMMV long intergenic region (LIR).
[0376] (Clause 2) The nucleic acid of clause 1, wherein the nucleic acid further comprises SbMMV replication initiator protein (Rep) and RepA.
[0377] (Clause 3) The nucleic acid of clause 1 or clause 2, wherein the exogenous nucleic acid comprises a repair template having complementary sequence to a target nucleic acid.
[0378] (Clause 4) The nucleic acid of any one of the prior clauses, wherein the repair template comprises from 5’ to 3’ : (I) a left homology arm (LHA) portion complementary to the target nucleic acid, (II) an insert sequence, and (III) a right homology arm (RHA) portion complementary to the target nucleic acid.
[0379] (Clause 5) The nucleic acid of any one of the prior clauses, wherein the LHA and / or RHA comprise at least 30 base pairs complementary to the target nucleic acid, respectively.
[0380] (Clause 6) The nucleic acid of any one of the prior clauses, wherein the LHA and / or RHA comprise 250 to 2,000 base pairs complementary to the target nucleic acid, respectively.
[0381] (Clause 7) The nucleic acid of any one of the prior clauses, wherein the insert sequence is exogenous to the target nucleic acid.
[0382] (Clause 8) The nucleic acid of any one of the prior clauses, wherein the insert sequence comprises at least one mismatched base pair relative to the target nucleic acid.
[0383] (Clause 9) The nucleic acid of any one of the prior clauses, wherein the insert sequence comprises a transgene, regulatory element, and / or selectable marker.
[0384] (Clause 10) The nucleic acid of any one of the prior clauses, wherein the LIR comprises a stem loop sequence.
[0385] (Clause 11) The nucleic acid of any one of the prior clauses, wherein the repair template introduces an insertion, deletion, or substitution in the target nucleic acid upon homologous recombination.
[0386] (Clause 12) The nucleic acid of any one of the prior clauses, wherein the repair template increases or decreases expression of a gene upon homologous recombination. (Clause 13) The nucleic acid of any one of the prior clauses, wherein the exogenous nucleic acid comprises an expression cassette.
[0387] (Clause 14) The nucleic acid of any one of the prior clauses, wherein the expression cassette further comprises a gene and / or regulatory element.
[0388] (Clause 15) The nucleic acid of any one of the prior clauses, wherein the regulatory element is a promoter or transcription terminator.
[0389] (Clause 16) The nucleic acid of any one of the prior clauses, further comprising a nuclease and / or a guide RNA (gRNA).
[0390] (Clause 17) The nucleic acid of any one of the prior clauses, further comprising a resistance marker.
[0391] (Clause 18) The nucleic acid of any one of the prior clauses, wherein the resistance marker is an herbicide and / or antibiotic resistance gene.
[0392] (Clause 19) The nucleic acid of any one of the prior clauses, further comprising a transfer DNA (T-DNA) border sequence.
[0393] (Clause 20) The nucleic acid of any one of the prior clauses, comprising a 5’ T-DNA border sequence (LB) and a 3’ T-DNA border sequence (RB) relative to the SbMMV replicon.
[0394] (Clause 21) The nucleic acid of any one of the prior clauses, wherein the SbMMV replicon comprises a 5’ SbMMV LIR and a 3' SbMMV LIR relative to the exogenous nucleic acid.
[0395] (Clause 22) The nucleic acid of any one of the prior clauses, wherein the SbMMV replicon comprises from 5’ to 3’: a sequence having at least 90% sequence identity to SEQ ID NO: 3, the repair template or expression cassette, and a sequence having at least 90% sequence identity to SEQ ID NO: 4.
[0396] (Clause 23) The nucleic acid of any one of the prior clauses, wherein the SbMMV replicon comprises from 5’ to 3’: a sequence having at least 98% sequence identity to SEQ ID NO: 3, the repair template or expression cassette, and a sequence having at least 98% sequence identity to SEQ ID NO: 4.
[0397] (Clause 24) The nucleic acid of any one of the prior clauses, wherein the SbMMV replicon comprises from 5’ to 3' : SEQ ID NO: 3 or a degenerate variant thereof, the repair template or expression cassette, and SEQ ID NO: 4 or a degenerate variant thereof.
[0398] (Clause 25) A vector comprising the nucleic acid of any one of clauses 1 to 24.
[0399] (Clause 26) The vector of clause 25, wherein the vector is a T-DNA vector or a biolistic vector.
[0400] (Clause 27) The vector of claim 25 or 26, wherein the vector is the T-DNA vector.
[0401] (Clause 28) A cell comprising the nucleic acid of any one of clauses 1 to 24 or the vector of any one of clauses 25 to 27.
[0402] (Clause 29) The cell of clause 28, wherein the cell is a plant cell or a bacterial cell.
[0403] (Clause 30) The cell of clause 28 or 29, wherein the cell is a soybean cell.
[0404] (Clause 31) A plant or plant part comprising the cell of any one of clauses 28 to 30.
[0405] (Clause 32) The plant or plant part of clause 31, wherein the plant or plant part is a soybean plant or plant part. (Clause 33) A method of producing a nucleic acid in a plant cell, comprising transforming the nucleic acid of any one of clauses 1 to 24 or the vector of any one of clauses 25 to 27 into the plant cell.
[0406] (Clause 34) The method of clause 33, wherein the nucleic acid is produced transiently.
[0407] (Clause 35) A method of modifying genetic material of a plant cell, comprising transforming the nucleic acid of any one of clauses 1 to 24 or the vector of any one of clauses 25 to 27 into the plant cell.
[0408] (Clause 36) The method of any one of clauses 33 to 35, wherein the plant cell is a soybean cell. (Clause 37) The method of clause 35 or claim 36, wherein the genetic material is genomic DNA. (Clause 38) The method of any one of clauses 33 to 37, wherein transforming comprises agrobacterium- mediated transformation or biolistic transformation.
[0409] (Clause 39) The method of any one of clauses 33 to 38, further comprising generating a plant from the transformed plant cell.
[0410] (Clause 40) The method of any one of clauses 35 to 39, wherein modifying genetic material of the plant cell comprises introducing an insertion, deletion, or substitution in a target nucleic acid.
[0411] (Clause 41) The method of any one of clauses 35 to 40, wherein modifying genetic material of the plant cell comprises homology-dependent repair of a target nucleic acid.
[0412] (Clause 42) A transgenic plant comprising SbMMV replication initiator protein (Rep) and RepA. (Clause 43) The transgenic plant of clause 42, wherein the transgenic plant is a transgenic soybean plant. (Clause 44) The transgenic plant of clause 42 or 43, wherein the SbMMV Rep and RepA is encoded by a nucleic acid sequence having at least 90%, at least 95%, or at least 98% sequence identity to SEQ ID NO: 2. (Clause 45) The transgenic plant of any one of clauses 42 to 44, wherein the SbMMV Rep and RepA comprises or consists of SEQ ID NO: 2, or a degenerate variant thereof.
[0413] (Clause 46) The transgenic plant of any one of clauses 42 to 45, further comprising an endonuclease.
[0414] EXAMPLES
[0415] The following examples are provided to illustrate particular features of certain aspects of the disclosure, but the scope of the claims should not be limited to those features exemplified.
[0416] Example 1 Transgene Expression
[0417] Geminivirus replicons were generated from soybean mild mottle virus (SbMMV), soybean blistering mosaic virus (SbBMV), and bean yellow dwarf virus (BeYDV). A majority of the VI and V2 regions of each respective virus were replaced with a GFP expression cassette (GFP operably linked to a ubiquitin promoter with an HSPt terminator; see, FIG. 1). The C1 / C2 region, which encodes Rep / RepA, and LIR / stem loop of the respective viruses were retained. Since one application of a Geminivirus replicon is to provide large amounts of a repair template for CRISPR-based editing, the replicons were included in a T- DNA construct that also encoded a Cas nuclease and gRNA, although neither the Cas nor guide RNA were needed for the expression of GFP. The T-DNA construct also included an EPSPS selection marker (glyphosate resistance), and T-DNA left and right border sequences (LB and RB). The resulting constructs were referred to as pIN4612 (which included the SbMMV replicon), pIN4801 (which included the SbBMV replicon), and pIN4802 (which included the BeYDV replicon) (see, FIG. 1).
[0418] The constructs were cloned into T-DNA binary plasmid, and were ready for transformation. Agrobacterium rhizogenes-mediated transformation was used to induce transgenic hairy root growth from soybean cotyledons (see, e.g. Chen et al., The Crop Journal 6(2) : 162-171, 2017)). In brief, half seed soybean explants were prepared two days after germination. Explants were inoculated with Agrobacterium rhizogenes carrying the pIN4612, pIN4801, and pIN4802 constructs. Six days after transformation, samples were evaluated for GFP fluorescence (see, FIG. 2). pIN4612 was found to achieve higher GFP expression due to superior amplification of the Geminivirus replicon.
[0419] Example 2 Genome Editing
[0420] In view of the superior expression obtained from the SbMMV-based replicon (pIN4612), it was selected for further testing.
[0421] A SbMMV-based replicon carrying a repair template for an endogenous target (Glyma.l9G194300) was generated by substituting the GFP expression cassette of pIN4612 described in Example I with a repair template. The repair template contained a 42 base pair exogenous insert sequence flanked by a left homology arm (LHA) and a right homology arm (RHA), which respectively contained about 750 base pairs of homologous sequence to the endogenous target (Glyma.l9G194300). The construct carrying the repair template was referred to as pIN4614 (see, FIG. 3). pIN2198, which carries the Glyma.l9G194300 repair template but no Geminivirus elements, was included as a “low copy” repair template reference. pIN4614, pIN4612, and pIN2198 were transformed into soybean explants as described in Example 1. The genomic sequences of transformed explants were characterized to score for the presence or absence of the repair template. The results are shown in the Table 2.
[0422] Table 2. Presence or absence of Glyma.l9G 194300 repair template. pIN4612 produced no precise insertions as it does not include a repair template. The “low copy” reference, pIN2198, produced only 1 plant with an insertion (67 tested; about 1.4%). In contrast, pIN4614 produced 7 plants with an insertion (79 tested; about 8.8%).
[0423] It will be apparent that the precise details of the methods or compositions described may be varied or modified without departing from the spirit of the described aspects of the disclosure. We claim all such modifications and variations that fall within the scope and spirit of the claims below.
Claims
We claim:
1. A nucleic acid comprising a recombinant soybean mild mottle virus (SbMMV) replicon, wherein the SbMMV replicon comprises:(i) an exogenous nucleic acid, wherein the exogenous nucleic acid is exogenous to SbMMV; and(ii) a sequence comprising SbMMV long inter genic region (LIR).
2. The nucleic acid of claim 1, wherein the nucleic acid further comprises a nucleic acid encoding replication initiator protein (Rep) and RepA.
3. The nucleic acid of claim 1, wherein the exogenous nucleic acid comprises a repair template having complementary sequence to a target nucleic acid.
4. The nucleic acid of claim 3, wherein the repair template comprises from 5’ to 3’ :(I) a left homology arm (LHA) portion complementary to the target nucleic acid,(II) an insert sequence, and(III) a right homology arm (RHA) portion complementary to the target nucleic acid.
5. The nucleic acid of claim 4, wherein the LHA and / or RHA comprise at least 30 base pairs complementary to the target nucleic acid, respectively.
6. The nucleic acid of claim 5, wherein the LHA and / or RHA comprise 250 to 2,000 base pairs complementary to the target nucleic acid, respectively.
7. The nucleic acid of claim 4, wherein the insert sequence is exogenous to the target nucleic acid.
8. The nucleic acid of claim 4, wherein the insert sequence comprises at least one mismatched base pair relative to the target nucleic acid.
9. The nucleic acid of claim 4, wherein the insert sequence comprises a transgene, regulatory element, and / or selectable marker.
10. The nucleic acid of claim 4, wherein the LIR comprises a stem loop sequence.
11. The nucleic acid of claim 4, wherein the repair template introduces an insertion, deletion, or substitution in the target nucleic acid upon homologous recombination.
12. The nucleic acid of claim 4, wherein the repair template increases or decreases expression of a gene upon homologous recombination.
13. The nucleic acid of claim 1, wherein the exogenous nucleic acid comprises an expression cassette.
14. The nucleic acid of claim 13, wherein the expression cassette further comprises a gene and / or regulatory element.
15. The nucleic acid of claim 14, wherein the regulatory element is a promoter or transcription terminator.
16. The nucleic acid of claim 1, further comprising a nuclease and / or a guide RNA (gRNA).
17. The nucleic acid of claim 1, further comprising a resistance marker.
18. The nucleic acid of claim 17, wherein the resistance marker is an herbicide and / or antibiotic resistance gene.
19. The nucleic acid of claim 1, further comprising a transfer DNA (T-DNA) border sequence.
20. The nucleic acid of claim 19, comprising a 5’ T-DNA border sequence (LB) and a 3’ T-DNA border sequence (RB) relative to the SbMMV replicon.
21. The nucleic acid of claim 1, wherein the SbMMV replicon comprises a 5’ SbMMV LIR and a 3’ SbMMV LIR relative to the exogenous nucleic acid.
22. The nucleic acid of claim 1, wherein the SbMMV replicon comprises from 5’ to 3’ : a sequence having at least 90% sequence identity to SEQ ID NO: 3, the repair template or expression cassette, and a sequence having at least 90% sequence identity to SEQ ID NO: 4.
23. The nucleic acid of claim 1, wherein the SbMMV replicon comprises from 5’ to 3’ : a sequence having at least 98% sequence identity to SEQ ID NO: 3, the repair template or expression cassette, and a sequence having at least 98% sequence identity to SEQ ID NO: 4.
24. The nucleic acid of claim 1, wherein the SbMMV replicon comprises from 5’ to 3’:SEQ ID NO: 3 or a degenerate variant thereof, the repair template or expression cassette, and SEQ ID NO: 4 or a degenerate variant thereof.
25. A vector comprising the nucleic acid of claim 1.
26. The vector of claim 25, wherein the vector is a T-DNA vector or a biolistic vector.
27. The vector of claim 26, wherein the vector is the T-DNA vector.
28. A cell comprising the nucleic acid of any one of claims 1 to 24 or the vector of any one of claims 25 to 27.
29. The cell of claim 28, wherein the cell is a plant cell or a bacterial cell.
30. The cell of claim 28 or 29, wherein the cell is a soybean cell.
31. A plant or plant part comprising the cell of any one of claims 28 to 30.
32. The plant or plant part of claim 31, wherein the plant or plant part is a soybean plant or plant part.
33. A method of producing a nucleic acid in a plant cell, comprising transforming the nucleic acid of any one of claims 1 to 24 or the vector of any one of claims 25 to 27 into the plant cell.
34. The method of claim 33, wherein the nucleic acid is produced transiently.
35. A method of modifying genetic material of a plant cell, comprising transforming the nucleic acid of any one of claims 1 to 24 or the vector of any one of claims 25 to 27 into the plant cell.
36. The method of any one of claims 33 to 35, wherein the plant cell is a soybean cell.
37. The method of claim 35 or claim 36, wherein the genetic material is genomic DNA.
38. The method of any one of claims 33 to 37, wherein transforming comprises agrobacterium- mediated transformation or biolistic transformation.
39. The method of any one of claims 33 to 38, further comprising generating a plant from the transformed plant cell.
40. The method of any one of claims 35 to 39, wherein modifying genetic material of the plant cell comprises introducing an insertion, deletion, or substitution in a target nucleic acid.
41. The method of any one of claims 35 to 40, wherein modifying genetic material of the plant cell comprises homology-dependent repair of a target nucleic acid.
42. A transgenic plant comprising SbMMV replication initiator protein (Rep) and RepA.
43. The transgenic plant of claim 42, wherein the transgenic plant is a transgenic soybean plant.
44. The transgenic plant of claim 42 or claim 43, wherein the SbMMV Rep and RepA is encoded by a nucleic acid sequence having at least 90%, at least 95%, or at least 98% sequence identity to SEQ ID NO: 2.
45. The transgenic plant of any one of claims 42 to 44, wherein the SbMMV Rep and RepA comprises or consists of SEQ ID NO: 2, or a degenerate variant thereof.
46. The transgenic plant of any one of claims 42 to 45, further comprising an endonuclease.