Removal of recombinant DNA from plant cell genome
By combining the Cre recombinase system controlled by the Ntm19 promoter with Cas nuclease and gRNA, the laborious and time-consuming problem of recombinant DNA excision in plant somatic cell genomes in existing technologies was solved, achieving efficient and precise genome modification and improving the quality of regenerated plants.
Patent Information
- Application Number
- CN202380075439.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-26
- Filing Date
- 2023-10-24
- Publication Date
- 2025-09-16
AI Technical Summary
Existing methods for removing recombinant DNA after plant transformation are laborious and time-consuming, usually requiring screening of multiple progeny plants, and may result in abnormal phenotypes or sterility in regenerated plants. It is difficult to accurately and effectively remove unwanted genetic elements from the plant somatic cell genome at an early stage.
The Cre recombinase system under the control of the Ntm19 promoter, combined with Cas nuclease and gRNA, automatically removes the inserted recombinant genetic elements by expressing the site-specific recombinase in plant somatic cells, and utilizes lox recombination sites and guide RNA to achieve efficient genome modification and recombination, avoiding additional treatment and chemical induction.
It achieves efficient and precise excision of recombinant DNA before plant somatic cell regeneration, reduces the risk of off-target cutting and non-genetic mutations, and improves transformation efficiency and the quality of regenerated plants.
Smart Images

Figure CN120659533A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of plant molecular biology and relates to the excision of recombinant DNA from the genome of somatic plant cells. Background Art
[0002] To date, several methods have been developed to remove transgenes after plant transformation. Such methods include: (1) sexual crossing and segregation (Gao et al., 2016; Char et al., 2017); (2) co-transformation using two constructs, one with a selectable marker and one with the gene of interest, to allow subsequent removal of the selectable marker gene by genetic segregation of progeny of transgenic plants generated following independent insertion of each of these constructs (Komari et al., 1996); (3) homologous recombination to perform excision between direct repeats (Puchta, 2000); (4) transposable element-based systems (Yoder and Goldsbrough, 1994; Gao et al., 2015); and (5) several site-specific recombination systems, such as Cre / lox from bacteriophage P1 (Hoess et al., 1982) and Flp / Frt from Saccharomyces cerevisiae (Cox, 1983).
[0003] Recombinases have been delivered via secondary transformation (Odell et al., 1990) or by hybridization (Bayley et al.). These procedures are laborious and time-consuming, often requiring screening of multiple progeny plants to recover those that have undergone excision. Recombinases have also been delivered by transient expression, for example, using viral-based vectors (Kopertekh and Schiemann, 2005). Similarly, excision is also performed after the T0 generation and multiple plants need to be screened.
[0004] An alternative approach has been developed based on the use of germline-specific auto-excision vectors containing the Cre recombinase gene under the control of a germline-specific promoter, thereby obtaining marker-free progeny plants (Verweire et al., 2007).
[0005] Mlynarova et al. (2006) reported the use of the microspore-specific Ntm19 promoter to drive Cre gene expression. Excision of the marker gene occurred during microsporogenesis with an efficiency close to 100% in tobacco seeds.
[0006] Control of excision can be further achieved by placing the recombinase under the control of an inducible / chemical promoter, an expression system that allows for spatial and temporal control regulated by external or internal signals, resulting in the automatic excision of the recombinase and a selectable marker gene placed within the boundaries of the excision site (Chong-Pérez and Angenon, 2013).
[0007] The expression of morphogenetic genes has been shown to improve plant transformation efficiency and / or improve callus formation. However, its expression in regenerated plants usually damages the quality of plants, such as by causing abnormal phenotypes or sterility. Developed to excise morphogenetic genes after transformation but before bud regeneration to increase the production system of normal fertile plants. The morphogenetic gene excision (Vilardell et al., 1991) carried out using the drought-inducible Rab17 promoter that drives Cre recombinase expression has been described. Although this method works, the drying step reduces event recovery, and not all events have achieved excision (Lowe et al., 2016).
[0008] Developmentally regulated promoters such as oleosin (Ole) (Anand et al., 2017b), globulin 1 (Glb1) (Belanger and Kriz, 1991), early embryo response gene (End2) (Casper et al., 2005) and lipid transfer protein 2 (Ltp2) (Kalla et al., 1994) have been used to drive Cre-mediated excision of morphogenetic genes in early embryonic development. Although excision events have been generated, the frequency of conversion and high-quality events is low, presumably due to premature expression of Cre, resulting in morphogenetic genes not being excised in time. The use of heat shock inducible promoters Hsp17.7 and Hsp26 (driving Cre expression) resulted in a higher frequency of T0 conversion, gene excision and high-quality event recovery. These methods also allow for the simultaneous excision of morphogenetic genes and selectable marker genes (Wang, N. et al., 2020).
[0009] In addition to morphogenic genes, other examples of genes or genetic elements that are introduced into the genome of a plant during the transformation process or early stages of regeneration that are beneficial but unwanted in the regenerated plant are, for example, transposons or genes that induce double-strand breaks in the genome like TALEN, CRISPR / Cas, homing endonucleases, etc. After the introduction of the intended double-strand breaks and their repair via non-homologous end joining or homologous recombination, the extended expression of such genes may lead to off-target cleavage and the accumulation of unwanted mutations in the genome.
[0010] Therefore, there is a need in the art for the precise and efficient excision of transgenes from the genome of cells, particularly somatic cells, before or during regeneration of shoots and / or plants from these cells. Summary of the Invention
[0011] Surprisingly, we found that expressing one or more elements of a system for excising genetic elements from the genome under the control of the Ntm19 promoter in somatic cells (Custers et al. 1997, PMB 35, 689-699) allows for precise and efficient excision from the genome and regeneration of shoots lacking the genetic element to be excised. We showed that if a Cas enzyme is included on the excision element for introducing a genomic modification into the cell genome, a high percentage of regenerated shoots will contain the desired genomic alteration.
[0012] As an example of the use of this system based on the use of the Ntm19 promoter, we describe a rapid procedure for removing nuclease components shortly after the introduction of one or more targeted genomic modifications. This procedure is based on the automated excision of both the Cas nuclease, guide RNA (gRNA), and Cre recombinase using Cre recombinase under the control of the Ntm19 promoter prior to TO regeneration.
[0013] The method is based on the design of a construct containing a Cas nuclease, gRNA, and Cre recombinase flanked by lox recombination sites. A selectable marker gene is located outside the recombination sites.
[0014] The first step involves inserting a nuclease construct and introducing one or more targeted modifications into the plant genome. Following completion of the one or more targeted genomic modifications, the second step involves removing the inserted construct using a site-specific Cre recombinase under the control of the Ntm19 promoter. We discovered that, in addition to expression in microspores, this promoter surprisingly exhibits activity in meristematic cells, allowing activation of the site-specific recombinase and removal of sequences between excision boundaries by recombination or excision during the early tissue culture stage, before T0 plants are transferred to the greenhouse, without the need for additional manipulation, physical, or chemical induction. Genome editing and subsequent removal of the nuclease components and recombinase are achieved in a short timeframe, significantly reducing or eliminating the risk of off-target effects and the generation of new somatic, non-inherited mutations.
[0015] A first embodiment of the present invention is a method for excising or deleting one or more recombinant genetic elements from the genome of a somatic cell (e.g., a non-gametophytic cell) of a transgenic plant, the method comprising the following steps:
[0016] a. introducing one or more recombinant genetic elements into the genome of the somatic cell, each recombinant genetic element comprising
[0017] i. the first excision recognition site, and
[0018] ii. a second excision recognition site, and iii. a polynucleotide encoding an excision component capable of excising a recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the Ntm19 promoter,
[0019] b. expressing the excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site, wherein the excision component recognizes the excision site and excises the recombinant DNA located between the first excision recognition site and the second excision recognition site.
[0020] Excision or deletion of the one or more recombinant genetic elements or the recombinant DNA can be achieved by, for example, recombination performed by a site-specific recombinase, or by homology-directed repair (HDR) or non-homologous end joining (NHEJ) after introducing a double-strand break or nick in the excision recognition sites of i and ii.
[0021] The excision component comprises a resection protein, and if the resection protein is a nucleic acid-guided DNA endonuclease, the excision component further comprises two guide RNAs that guide the nucleic acid-guided DNA endonuclease protein to the first and second excision recognition sites.
[0022] The excision recognition site can be a site-specific recombination site, such as a lox or att-site; a recognition site for a homing endonuclease (such as a LAGLIDADG-type homing endonuclease), such as I-CreI or I-Sce-I; a recognition site for a rare restriction enzyme; a PAM site adjacent to a sequence complementary to a guide RNA (guiding CRISPR / Cas enzymes, such as Cas9, Cas12a, b, c, etc., CasX, CasY, etc.); other enzymes with recombinase, endonuclease or nickase activity, such as a Zn finger protein or TALEN fused to a peptide having such activity.
[0023] The excision protein can be a site-specific recombinase. The site-specific recombinase can be selected from the group consisting of FLP, Cre, SSV1, lambda Int, phi C31 Int, HK022, R, Gin, Tn1721, CinH, ParA, Tn5053, Bxb1, TP907-1, or U153.
[0024] The nucleic acid-guided DNA endonuclease can be a CRISPR / Cas enzyme, such as Cas9, Cas12a, b, c, etc., CasX, CasY, etc.
[0025] The resection protein may further be a homing endonuclease, such as a LAGLIDADG-type homing endonuclease, e.g., I-Crel or I-Sce-I homing endonuclease; or a rare restriction enzyme; or other enzyme with recombinase, endonuclease or nickase activity, such as, for example, a Zn finger protein or a TALEN fused to a peptide with such activity.
[0026] In preferred embodiments, if the recognition sites are sites for introducing double-strand breaks or nicks, these recognition sites are cut or nicked by the same activity (e.g., have the same sequence, or contain a sequence with sufficient homology to hybridize to the same guide RNA).
[0027] Preferably, a Cas system, more preferably a Cas12a system or a Cre / lox recombinase system, is used to excise the recombinant genetic element.
[0028] Recombinant genetic elements can be introduced into the genome of plant somatic cells by any means known in the art (such as particle bombardment, protoplast electroporation, viral infection, Agrobacterium-mediated transformation, magnetofection) using a repair template and CRISPR / Cas nuclease or nickase, etc. Preferably, they are introduced using Agrobacterium-mediated transformation (such as Agrobacterium rhizogenes or Agrobacterium tumefaciens-mediated transformation).
[0029] The plant somatic cell can be any plant somatic cell, preferably a dicotyledonous plant somatic cell, more preferably a soybean plant somatic cell. The plant cell can be a leaf cell, a stem cell, a root cell, a shoot cell, a cotyledon cell, an epicotyl cell, an embryonic cell, a callus cell, a protoplast, or a meristematic cell. Preferably, the cell is a meristematic cell, more preferably a meristematic cell derived from a dicotyledonous plant, most preferably a soybean plant. The meristematic cell can be a cell of the shoot apex meristem.
[0030] In the method of the present invention, in addition to the polynucleotide encoding the elements necessary for excision of the recombinant DNA located between excision sites i and ii (flanked by the excision sites at both ends), the recombinant genetic elements introduced into the plant cell genome in step a may further comprise additional recombinant elements that remain stably integrated in the plant somatic cell genome after excision. Such recombinant elements may be the gene of interest located outside the first and second excision sites but not between them, or the gene of interest flanking the first or second excision sites.
[0031] In the case where the first excision recognition site is different from the second excision recognition site and is recognized by different nucleases or different guide RNAs (guiding a different Cas nuclease or nickase), at least one of the polynucleotides encoding the nuclease is functionally linked to the Ntm19 promoter. Preferably, in this case, each of the two polynucleotides encoding the nuclease is functionally linked to the Ntm19 promoter.
[0032] As used herein, expressing the excision component refers to expressing the component in somatic cells. This can be an early tissue culture stage of expression.
[0033] Another embodiment of the present invention is a method for transiently expressing a gene of interest in a plant somatic cell (e.g., a non-gametophytic cell), the method comprising the steps of:
[0034] a. introducing one or more recombinant genetic elements into the genome of the somatic cell, each recombinant genetic element comprising
[0035] i. the first excision recognition site, and
[0036] ii. a second excision recognition site, and iii. a polynucleotide encoding an excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the Ntm19 promoter and the target gene, and
[0037] b. expressing the excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site, wherein the excision component recognizes the excision site and excises the recombinant DNA located between the first excision recognition site and the second excision recognition site.
[0038] As used herein, transient expression may also mean temporal expression, or may be the transient or transient presence of a gene of interest in the genome of a cell.
[0039] In another embodiment, the recombinant DNA located between the first excision recognition site and the second excision recognition site further comprises a sequence encoding at least one morphogenic gene, and / or a sequence encoding a genome editing component for editing the target sequence in the plant somatic cell, or a sequence encoding a selectable marker.
[0040] Another embodiment of the present invention is a method for improving the regeneration of transgenic plants, plant parts, calli, plant organs from plant somatic cells, the method comprising the steps of:
[0041] a. introducing one or more recombinant genetic elements into the genome of the somatic cell, each recombinant genetic element comprising
[0042] i. the first excision recognition site, and
[0043] ii. a second excision recognition site, and iii. a polynucleotide encoding an excision component capable of excising a recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the Ntm19 promoter and at least one morphogenic gene functionally linked to the promoter,
[0044] b. expressing the excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site.
[0045] The term "morphogenic gene" refers to a gene or portion of a gene that, when expressed in a cell, improves or enhances the regeneration of a plant organ, plant tissue, or plant part from a cell; or that encodes a regulatory factor that enhances the expression of an endogenous gene that improves or enhances the regeneration of a plant organ, plant tissue, or plant part from a cell.
[0046] Preferably, the morphogenetic genes are selected from the list comprising, more preferably from the list consisting of: ESR1 (Banno et al. (2001) Plant Cell 13(12)), WIND1 (Iwase et al. (2017) Plant Cell 29(1)), WUS and ESR1 (Xu et al. (2021) Sci Adv 7(33)), WOX5 or 11 (Liu et al. (2018) Plant Cell Phys 59(4))(Liu et al. (2014) Plant Cell 26(3)), GRF5 (Kong et al. (2020) Front Plant Sci 11), GRF-GIF (Debernardi et al. (2020) Nat Biotechnol 38(11)), miRNA156 (Zhang et al. (2015) Plant Cell 27(2), AGL15 (Zheng et al. (2013) Plant Physiol 161(4)), LEC1 / LEC2 (Guo et al. (2013) PLoS ONE 8(8)), RKD4 (Gordon-Kamm et al. (2019) Plants 8(2)), Baby Boom (Boutilier et al. (2002) Plant Cell 14(8)), and PLT (Kareem et al. (2015) Current Biology 25(8)).
[0047] In another embodiment, the method of the present invention further comprises the step of selecting a cell in which excision of the recombinant DNA between said first excision recognition site and said second excision recognition site has occurred. In another embodiment, the method of the present invention further comprises the steps of regenerating a bud or plantlet from said somatic cell and selecting a bud or plantlet in which excision of the recombinant DNA between said first excision recognition site and said second excision recognition site has occurred.
[0048] Cells, or buds or plantlets in which excision of the recombinant DNA between the first excision recognition site and the second excision recognition site has occurred can be selected as follows, for example using molecular techniques, such as PCR techniques using primers for specific amplification of the recombinant DNA between the first excision recognition site and the second excision recognition site as described in the Examples herein; or PCR techniques using primers flanking the first excision recognition site and the second excision recognition site, and determining whether the recombinant DNA has been excised based on the size of the amplified product; or using sequencing methods. Alternatively, cells, or buds or plantlets can be selected based on the presence or absence of a positive selectable marker or a negative selectable marker. It will be clear to those skilled in the art that in order to select cells, or buds or plantlets in which excision of the recombinant DNA between the first excision recognition site and the second excision recognition site has occurred, they must be distinguished from cells into which recombinant genetic elements have never been introduced. Therefore, the present invention is also suitable for detecting genetic elements that are not between the first excision recognition site and the second excision recognition site, such as genetic elements outside the first excision recognition site and the second excision recognition site, or footprints left by the excision recognition site after excision. Such genetic elements can, for example, be detected using molecular techniques such as PCR techniques or using sequencing methods, or using selection or screening for expression in the recombinant DNA outside the first and second excision recognition sites or the gene of interest, such as selection for expression of a selectable marker such as a marker that confers herbicide tolerance or screening for expression of a screenable marker such as a gene that causes a color change or other visible change.
[0049] Shoots may be regenerated from somatic cells as described in the art, for example by cultivation on shoot induction medium as known in the art.
[0050] Another embodiment of the present invention is a method of producing a plant or shoot comprising an edit in a target sequence, the method comprising:
[0051] a. Introducing one or more recombinant genetic elements into the genome of a somatic cell, each recombinant genetic element comprising
[0052] i. First excision recognition site,
[0053] ii. a second excision recognition site, and iii. a polynucleotide encoding an excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the ntm19 promoter, and a sequence encoding a genome editing component for editing the target sequence, and
[0054] b. regenerating buds from said somatic cells,
[0055] c. selecting buds that comprise the edit in the target sequence and in which excision of the recombinant DNA between the first excision recognition site and the second excision recognition site has occurred, and optionally
[0056] d. growing a plant from the shoot.
[0057] Further provided are seeds produced by plants produced using the methods of the invention, such as seeds comprising an edit in the target sequence.
[0058] Another embodiment provides a method for removing a genome editing component shortly after introduction of a targeted genome modification, the method comprising:
[0059] a. Introducing one or more recombinant genetic elements into the genome of a somatic cell, each recombinant genetic element comprising
[0060] i. First excision recognition site,
[0061] ii. a second excision recognition site, and iii. a polynucleotide encoding an excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the ntm19 promoter, and a sequence encoding a genome editing component for editing the target sequence, and
[0062] b. regenerating buds from said somatic cells,
[0063] c. Selecting buds that contain the edit in the target sequence and in which excision of the recombinant DNA between the first excision recognition site and the second excision recognition site has occurred.
[0064] In yet another embodiment, the sequence encoding the genome editing components for editing the target sequence encodes a site-directed nuclease, such as a nucleic acid-guided DNA endonuclease, or such as a Cas nuclease and a guide RNA.
[0065] In another embodiment, the recombinant genetic element further comprises a gene of interest outside the first excision recognition site and the second excision recognition site.
[0066] The target gene outside the first excision recognition site and the second excision recognition site can be any target gene, such as a selectable marker or a screenable marker, or a gene that confers herbicide tolerance, or a gene that confers pest resistance, or a gene that confers stress tolerance, or a gene for increasing yield, or a gene that improves the quality of plants or plant products.
[0067] In yet another embodiment, the first excision recognition site and the second excision recognition site are lox sites, and wherein the excision component is a Cre recombinase protein; and in another embodiment, the excision component is a nucleic acid-guided DNA endonuclease and one or two guide RNAs that guide the nucleic acid-guided DNA endonuclease protein to the first excision recognition site and the second excision recognition site.
[0068] One embodiment of the present invention is a recombinant construct comprising
[0069] i. the first excision recognition site, and
[0070] ii. a second excision recognition site, and a second excision recognition site between the first excision recognition site and the second excision recognition site.
[0071] iii. A polynucleotide encoding an excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the Ntm19 promoter.
[0072] The recombinant construct may further comprise a polynucleotide encoding a morphogenic gene or a polynucleotide encoding a nucleic acid-guided DNA endonuclease protein between the first excision recognition site and the second excision recognition site, wherein the nucleic acid-guided DNA endonuclease protein excises the recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the Ntm19 promoter.
[0073] Further embodiments of the present invention are methods and constructs as defined above, wherein said Ntm19 promoter comprises a sequence selected from the group consisting of:
[0074] a) a nucleic acid molecule having the sequence of SEQ ID NO: 1, and
[0075] b) a nucleic acid molecule having a sequence that is at least 80% (e.g., at least 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%), more preferably 90% (e.g., at least 91%, 92%, 93%, 94%, 95%, 96%, 97%), even more preferably 98%, and most preferably 99% identical to the sequence having SEQ ID NO: 1 over at least 250, 300, 400, 500, 600, preferably 700, more preferably 800, even more preferably 900 consecutive nucleic acid base pairs, and most preferably over the entire length, and
[0076] c) a fragment of at least 100 consecutive bases, preferably at least 200, 300, 400 or 500 consecutive bases, more preferably at least 600, 700 or 800 consecutive bases, most preferably at least 900 or 950 consecutive bases of the nucleic acid molecule of a) or b), which fragment has the same activity as the corresponding nucleic acid molecule having the sequence of SEQ ID NO: 1; in the most preferred embodiment, the fragment is a fragment comprising the 3' end of the sequence of SEQ ID NO: 1, and
[0077] d) a nucleic acid molecule that is the complement or reverse complement of any of the previously mentioned nucleic acid molecules in a) to c), and
[0078] e) a nucleic acid molecule that hybridizes to a nucleic acid molecule comprising at least 500, 600, 700, 800, 900, 950 or all consecutive nucleotides of SEQ ID NO: 1, or its complement under conditions equivalent to hybridization in 7% sodium dodecyl sulfate (SDS), 0.5 M NaPO4, 1 mM EDTA at 50°C, preferably 55°C, more preferably 60°C, even more preferably 65°C, most preferably 68°C, and washing in 2X SSC, 0.1% SDS at 50°C, preferably 55°C, more preferably 60°C, even more preferably 65°C, most preferably 68°C,
[0079] And wherein the sequence in b) to e) induces approximately the same promoter activity as SEQ ID NO: 1. "Approximately the same promoter activity" means the same tissue specificity and an expression that deviates from the expression directed by SEQ ID NO: 1 by no more than 75%, preferably no more than 50%, more preferably no more than 25%, and most preferably no more than 10%.
[0080] Vectors comprising the recombinant constructs of the present invention are further embodiments of the present invention.
[0081] The present invention also encompasses the following cells, preferably plant somatic cells, more preferably dicot somatic cells, and most preferably soybean somatic cells comprising the recombinant construct or vector of the present invention.
[0082] definition
[0083] Abbreviations: GFP-green fluorescent protein, GUS-β-glucuronidase, BAP-6-benzylaminopurine; 2,4-D-2,4-dichlorophenoxyacetic acid; MS-Mullahig and Skoog medium; NAA-1-naphthaleneacetic acid; MES-2-(N-morpholino-ethanesulfonic acid), IAA-indoleacetic acid; Kan: kanamycin sulfate; GA3-gibberellic acid; Timentin TM : Ticarcillin disodium / clavulanate potassium, microl: microliter.
[0084] It should be understood that the present invention is not limited to a specific method or scheme. It should also be understood that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the scope of the present invention, which will be limited only by the appended claims. It must be noted that, unless the context clearly indicates otherwise, as used herein and in the appended claims, the singular forms "a / an" and "the" include plural referents. Thus, for example, "a vector" refers to one or more vectors and includes equivalents thereof known to those skilled in the art. The term "about" is used herein to mean approximately, roughly, roughly, or around... When the term "about" is used in conjunction with a numerical range, it modifies the range by extending the boundaries above and below the listed values. Generally, the term "about" is used herein to modify the numerical value above and below the stated value by a difference of 20% (preferably 10%). As used herein, the word "or" means any one member of a particular list and also includes any combination of members of that list. When used in this specification and the following claims, the words "comprise," "comprising," "include," "including," and "includes" are intended to specify the presence of one or more stated features, integers, components, or steps, but they do not preclude the presence or addition of one or more other features, integers, components, steps, or groups thereof. For clarity, certain terms used in this specification are defined and used as follows:
[0085] Antiparallel: "Antiparallel" herein refers to two nucleotide sequences that are paired through hydrogen bonding between complementary base residues, wherein the phosphodiester bonds run in the 5'-3' direction in one nucleotide sequence and in the 3'-5' direction in the other nucleotide sequence.
[0086] Antisense: The term "antisense" refers to a nucleotide sequence that is in the reverse direction relative to the normal orientation of transcription or function and thereby expresses an RNA transcript that is complementary to a target gene mRNA molecule expressed in a host cell (e.g., it can hybridize to a target gene mRNA molecule or single-stranded genomic DNA by Watson-Crick base pairing) or an RNA transcript that is complementary to a target DNA molecule (such as, for example, genomic DNA present in a host cell).
[0087] Coding region: As used herein, the term "coding region" when used with respect to a structural gene refers to the nucleotide sequence encoding the amino acids found in a nascent polypeptide due to the translation of an mRNA molecule. In eukaryotes, the coding region is defined on the 5'-side by the nucleotide triplet "ATG" encoding the initiator methionine, and on the 3'-side by one of the three triplets specifying the stop codon (i.e., TAA, TAG, TGA). In addition to containing introns, the genomic form of a gene may also include sequences located at the 5' and 3' ends of the sequence present on the RNA transcript. These sequences are referred to as "flanking" sequences or regions (these flanking sequences are located 5' or 3' of the non-translated sequences present on the mRNA transcript). The 5' flanking region may contain regulatory sequences, such as promoters and enhancers that control or influence the transcription of the gene. The 3' flanking region may contain sequences that guide transcription termination, post-transcriptional cleavage, and polyadenylation.
[0088] Complementarity: "Complementarity" or "complementarity" refers to two nucleotide sequences that comprise antiparallel nucleotide sequences that are able to pair with each other (by the base pairing rules) when hydrogen bonds are formed between the complementary base residues of the antiparallel nucleotide sequences. For example, the sequence 5'-AGT-3' is complementary to the sequence 5'-ACT-3'. Complementarity can be "partial" or "overall" complementarity. "Partial" complementarity is a situation where one or more nucleic acid bases do not match according to the base pairing rules. "Overall" or "complete" complementarity between nucleic acid molecules is a situation where every nucleic acid base matches every other base under the base pairing rules. The degree of complementarity between nucleic acid molecule chains has a significant effect on the efficiency and strength of hybridization between nucleic acid molecule chains. As used herein, the "complementary sequence" of a nucleic acid sequence refers to a nucleotide sequence whose nucleic acid molecule exhibits overall complementarity with the nucleic acid molecule of that nucleic acid sequence.
[0089] Donor DNA molecule: As used herein, the terms "donor DNA molecule", "repair DNA molecule" or "template DNA molecule" are all used interchangeably herein to refer to a DNA molecule having a sequence to be introduced into the genome of a cell. It may be flanked at the 5' and / or 3' ends by sequences homologous or identical to sequences in the target region of the genome of the cell. It may contain sequences that are not naturally present in the corresponding cell, such as ORFs, non-coding RNAs, or regulatory elements that should be introduced into the target region, or it may contain sequences that are homologous to the target region except for at least one mutation, gene editing: the sequence of the donor DNA molecule may be added to the genome, or it may replace a sequence in the genome having the length of the donor DNA sequence.
[0090] Double-stranded RNA: A "double-stranded RNA" or "dsRNA" molecule comprises a sense RNA segment of nucleotide sequence and an antisense RNA segment of nucleotide sequence, both of which comprise nucleotide sequences that are complementary to each other, allowing the sense and antisense RNA segments to pair and form a double-stranded RNA molecule.
[0091] Endogenous: An "endogenous" nucleotide sequence refers to a nucleotide sequence that is present in the genome of an untransformed plant cell.
[0092] Enhanced expression: "enhancing" or "increasing" the expression of a nucleic acid molecule in a plant cell are used equivalently herein and mean that the expression level of the nucleic acid molecule in a plant, a part of a plant or a plant cell after applying the method of the present invention is higher than the expression of the nucleic acid molecule in the plant, a part of a plant or a plant cell before applying the method, or that the expression level is higher compared to a reference plant lacking the recombinant nucleic acid molecule of the present invention. For example, the reference plant comprises the same construct only lacking the corresponding NEENA. As used herein, the terms "enhanced" or "increased" are synonymous and mean higher, preferably significantly higher, expression of the nucleic acid molecule to be expressed. As used herein, "enhancing" or "increasing" the level of an agent (such as a protein, mRNA or RNA) means that the level is increased relative to essentially the same plant, part of a plant or a plant cell lacking the recombinant nucleic acid molecule of the present invention (e.g., lacking a NEENA molecule, a recombinant construct of the present invention or a recombinant vector) grown under essentially the same conditions. As used herein, "enhancement" or "increase" of the level of a substance (e.g., preRNA, mRNA, rRNA, tRNA, snoRNA, snRNA, and / or protein products encoded thereby) expressed by a target gene means an increase of 50% or more, e.g., 100%-fold or more, preferably 200%-fold or more, more preferably 5-fold or more, even more preferably 10-fold or more, most preferably 20-fold or more, e.g., 50-fold, relative to a cell or organism lacking the recombinant nucleic acid molecule of the present invention. The enhancement or increase can be determined by methods familiar to the skilled person. Thus, the enhancement or increase in the amount of a nucleic acid or protein can be determined, for example, by immunological detection of the protein. In addition, techniques such as protein assays, fluorescence, Northern hybridization, nuclease protection assays, reverse transcription (quantitative RT-PCR), ELISA (enzyme-linked immunosorbent assay), Western blots, radioimmunoassays (RIA) or other immunoassays, and fluorescence-activated cell analysis (FACS) can be used to measure specific proteins or RNA in plants or plant cells. Depending on the type of protein product induced, its activity or effect on the phenotype of the organism or cell can also be determined. Methods for determining the amount of protein are known to the skilled person. Examples which may be mentioned are: the micro-Biuret method (Goa J (1953) Scand J Clin Lab Invest 5: 218-222), the Folin-Ciocalteau method (Lowry OH et al. (1951) J Biol Chem 193: 265-275) or the measurement of the absorption of CBB G-250 (Bradford MM (1976) Analyt Biochem 72: 248-254).As an example of quantifying protein activity, detection of luciferase activity is described in the Examples below.
[0093] Expression: "Expression" refers to the biosynthesis of a gene product, preferably the transcription and / or translation of a nucleotide sequence (e.g., an endogenous gene or a heterologous gene) in a cell. For example, in the case of a structural gene, expression involves the transcription of the structural gene into mRNA and, optionally, the subsequent translation of the mRNA into one or more polypeptides. In other cases, expression may refer solely to the transcription of the DNA harboring the RNA molecule.
[0094] Expression construct: As used herein, "expression construct" means a DNA sequence capable of directing the expression of a specific nucleotide sequence in an appropriate plant part or plant cell, comprising a promoter functional in the plant part or plant cell into which it is to be introduced, operably linked to a nucleotide sequence of interest, optionally operably linked to a termination signal. If translation is desired, it typically also contains sequences required for proper translation of the nucleotide sequence. The coding region may encode a protein of interest, but may also encode a functional RNA of interest, such as RNAa, siRNA, snoRNA, snRNA, microRNA, ta-siRNA, or any other non-coding regulatory RNA, in either the sense or antisense orientation. An expression construct comprising a nucleotide sequence of interest may be chimeric, meaning that one or more of its components is heterologous to one or more of the other components. An expression construct may also be naturally occurring but has been obtained in a recombinant format for heterologous expression. However, typically, an expression construct is heterologous to the host, meaning that the specific DNA sequence of the expression construct is not naturally present in the host cell and must have been introduced into the host cell or an ancestor of the host cell through a transformation event. The expression of the nucleotide sequence in the expression construct can be under the control of a constitutive promoter or an inducible promoter, which initiates transcription only when the host cell is exposed to some specific external stimulus. In the case of plants, the promoter can also be specific to a particular tissue, organ, or developmental stage.
[0095] Foreign: The term "foreign" refers to any nucleic acid molecule (e.g., a gene sequence) that is introduced into the genome of a cell by experimental manipulation, and can include sequences found in that cell, so long as the introduced sequence contains some modification (e.g., point mutations, the presence of a selectable marker gene, etc.) and is therefore different from the naturally occurring sequence.
[0096] Functional connection: The term "functional connection" or "functionally connected" should be understood to mean, for example, that a regulatory element (e.g., a promoter) and a nucleic acid sequence to be expressed and, if appropriate, additional regulatory elements (e.g., terminators or NEENAs) are arranged in sequence in such a way that each of these regulatory elements can perform its intended function to allow, modify, promote or otherwise influence the expression of the nucleic acid sequence. As synonyms, the terms "operably linked" or "operably connected" can be used. The result of expression depends on the arrangement of the nucleic acid sequence relative to the sense or antisense RNA. For this purpose, a direct connection in the chemical sense is not necessarily required. Genetic control sequences (e.g., enhancer sequences) can also exert their function on the target sequence from a more distant position or actually from a position away from other DNA molecules. A preferred arrangement is one in which the nucleic acid sequence to be recombinantly expressed is located after the sequence acting as a promoter, so that the two sequences are covalently linked to each other. The distance between the promoter sequence and the nucleic acid sequence to be recombinantly expressed is preferably less than 200 base pairs, particularly preferably less than 100 base pairs, and very particularly preferably less than 50 base pairs. In a preferred embodiment, the nucleic acid sequence to be transcribed is located behind a promoter such that the start of transcription is identical to the desired start of the chimeric RNA of the invention. Functionally linked and expression constructs can be generated by conventional recombination and cloning techniques as described (e.g., Maniatis T, Fritsch EF and Sambrook J (1989) Molecular Cloning: A Laboratory Manual, 2nd ed., Cold Spring Harbor Laboratory, Cold Spring Harbor, NY; Silhavy et al. (1984) Experiments with Gene Fusions, Cold Spring Harbor Laboratory, Cold Spring Harbor, NY; Ausubel et al. (1987) Current Protocols in Molecular Biology, Greene Publishing Assoc. and Wiley Interscience; Gelvin et al. (eds.) (1990) Plant Molecular Biology Manual; Kluwer Academic Publishers, Dordrecht, The Netherlands). However, further sequences, for example as linkers with specific cleavage sites for restriction enzymes or as signal peptides, can also be located between the two sequences. The insertion of sequences can also lead to the expression of fusion proteins.Preferably, the expression construct consisting of the connection of regulatory regions (eg promoter and nucleic acid sequence to be expressed) can be present in vector-integrated form and inserted into the plant genome, for example by transformation.
[0097] Gene: The term "gene" refers to a region operably linked to appropriate regulatory sequences that can regulate the expression of a gene product (e.g., a polypeptide or functional RNA) in some way. A gene includes untranslated regulatory regions (e.g., promoters, enhancers, repressors, etc.) of DNA before (upstream) and after (downstream) the coding region (open reading frame, ORF), and, where applicable, intervening sequences (i.e., introns) between separate coding regions (i.e., exons). As used herein, the term "structural gene" is intended to mean a DNA sequence that is transcribed into mRNA that is then translated into an amino acid sequence unique to a specific polypeptide.
[0098] Genome and genomic DNA: The term "genome" or "genomic DNA" refers to the heritable genetic information of a host organism. The genomic DNA includes the DNA of the cell nucleus (also called chromosomal DNA), but also includes the DNA of plastids (e.g., chloroplasts) and other organelles (e.g., mitochondria). Preferably, the term genome or genomic DNA refers to the chromosomal DNA of the cell nucleus.
[0099] As used herein, genome editing (also referred to as gene editing, genome engineering) refers to the targeted modification of genomic DNA, wherein DNA can be inserted, deleted, modified or replaced in the genome. Genome editing can use sequence-specific enzymes (e.g., endonucleases, nickases, base transfer enzymes) and / or donor nucleic acids (e.g., dsDNA, oligonucleotides) to introduce desired changes in DNA. Sequence-specific nucleases that can be programmed to recognize specific DNA sequences include meganucleases (MGNs), zinc finger nucleases (ZFNs), TAL effector nucleases (TALENs), and RNA-guided or DNA-guided nucleases or nickases, such as Cas9, Cpf1, CasX, CasY, C2c1, C2c3, certain Argonaut-based systems (see, e.g., Osakabe and Osakabe, Plant Cell Physiol. 2015 Mar;56(3):389-400; Ma et al., Mol Plant. 2016 Jul 6;9(7):961-74; Bortesie et al., Plant Biotech J, 2016, 14; Murovec et al., Plant Biotechnol. J. Plant Biotechnol. 15:917-926, 2017; Nakade et al., Bioengineered 8, No. 3:265-273, 2017; Burstein et al., Nature 542, 37-241; Komor et al., Nature 533, 420-424, 2016; all incorporated herein by reference). The donor nucleic acid can be used as a template for repairing DNA breaks induced by sequence-specific nucleases.
[0100] Editing is the modification of a target sequence by genome editing. The target sequence can be any target sequence in the genome of a cell.
[0101] Heterologous: The term "heterologous" with respect to a nucleic acid molecule or DNA refers to a nucleic acid molecule that is operably linked or manipulated to become operably linked to a second nucleic acid molecule (e.g., a promoter) that is not operably linked to the second nucleic acid molecule in nature (e.g., in the genome of a WT plant) or is operably linked to the second nucleic acid molecule in nature (e.g., in the genome of a WT plant).
[0102] Preferably, the term "heterologous" with respect to a nucleic acid molecule or DNA (e.g., an enhancer) refers to a nucleic acid molecule that is operably linked, or manipulated to become operably linked, to a second nucleic acid molecule (e.g., a promoter) with which it is not operably linked in nature.
[0103] A heterologous expression construct comprising a nucleic acid molecule and one or more regulatory nucleic acid molecules (e.g., a promoter or transcription termination signal) linked thereto is, for example, a construct produced by experimental manipulation, wherein a) the nucleic acid molecule, b) the regulatory nucleic acid molecule, or c) both (i.e., (a) and (b)) are not located in their natural (original) genetic environment or have been modified by experimental manipulation, examples of modifications being substitutions, additions, deletions, inversions, or insertions of one or more nucleotide residues. The natural genetic environment refers to the natural chromosomal locus in the organism of origin or to the presence in a genomic library. In the case of a genomic library, the natural genetic environment of the nucleic acid molecule sequence is preferably retained, at least partially retained. This environment is flanked on at least one side by the nucleic acid sequence and has a sequence length of at least 50 bp, preferably at least 500 bp, particularly preferably at least 1,000 bp, and very particularly preferably at least 5,000 bp. When a naturally occurring expression construct (e.g., a naturally occurring combination of a promoter and a corresponding gene) is modified by non-natural, synthetic, "artificial" methods (e.g., mutagenesis), the naturally occurring expression construct becomes a transgenic expression construct. Such methods have been described (US 5,565,350; WO 00 / 15815). For example, it is believed that the following protein-encoding nucleic acid molecule is heterologous relative to a promoter that is operably linked to a promoter that is not the natural promoter of the molecule. Preferably, the heterologous DNA is not endogenous or naturally associated with the cell into which it is introduced, but is obtained from another cell or is synthesized. Heterologous DNA also includes a DNA sequence that contains some modified endogenous DNA sequences, multiple copies of non-naturally occurring endogenous DNA sequences, or another DNA sequence that is physically connected to it that is not naturally associated. Typically, although not necessarily, heterologous DNA encodes an RNA or protein that is not normally produced by the cell expressing the DNA.
[0104] Hybridization: as defined herein, term "hybridization" is a process in which substantially complementary nucleotide sequences anneal to each other. The hybridization process can occur completely in solution, that is, two complementary nucleic acids are all in solution. The hybridization process can also occur when one of the complementary nucleic acids is fixed to a substrate (such as magnetic beads, agarose beads or any other resin). In addition, the hybridization process can occur when one of the complementary nucleic acids is fixed on a solid support such as nitrocellulose or a nylon membrane or is fixed on, for example, a siliceous glass support (the latter being referred to as a nucleic acid array or microarray or being referred to as a nucleic acid chip) by, for example, photolithography. In order to allow hybridization to occur, nucleic acid molecules are usually denatured or chemically denatured to melt a double strand into two single strands and / or remove hairpins or other secondary structures from single-stranded nucleic acids.
[0105] The term "stringency" refers to the conditions under which hybridization occurs. The stringency of hybridization is affected by conditions (such as temperature, salt concentration, ionic strength and hybridization buffer composition). Generally, low stringency conditions are selected to be approximately 30°C lower than the thermal melting point (Tm) of the specific sequence under the ionic strength and pH defined. Medium stringency conditions are when the temperature is lower than 20°C of Tm, and high stringency conditions are when the temperature is lower than 10°C of Tm. High stringency hybridization conditions are typically used to separate hybridization sequences with high sequence similarity to the target nucleic acid sequence. However, due to the degeneracy of the genetic code, nucleic acid can deviate from and still encode substantially identical polypeptides in sequence. Therefore, sometimes medium stringency hybridization conditions may be needed to identify this type of nucleic acid molecules.
[0106] Tm is the temperature (under defined ionic strength and pH) at which 50% of the target sequence hybridizes to a perfectly matched probe. Tm depends on the solution conditions as well as the base composition and length of the probe. For example, longer sequences hybridize specifically at higher temperatures. The maximum rate of hybridization is obtained at about 16°C below Tm up to 32°C. The presence of monovalent cations in the hybridization solution reduces the electrostatic repulsion between the two nucleic acid chains, thereby promoting hybrid formation; this effect is visible for sodium concentrations up to 0.4M (for higher concentrations, this effect is negligible). For each percentage of formamide, formamide reduces the melting temperature of DNA-DNA and DNA-RNA duplexes by 0.6°C to 0.7°C, and the addition of 50% formamide allows hybridization at 30°C to 45°C, although the hybridization rate is reduced. Base pair mismatches reduce the hybridization rate and thermal stability of the duplex. On average and for large probes, Tm is reduced by about 1°C / % base mismatch. Depending on the type of hybrid, Tm can be calculated using the following formula:
[0107] DNA-DNA hybrids (Meinkoth and Wahl, Anal. Biochem., 138: 267-284, 1984):
[0108] Tm=81.5℃+16.6xlog[Na+]a+0.41x%[G / Cb]-500x[Lc]-1-0.61x%formamide
[0109] DNA-RNA or RNA-RNA hybrids:
[0110] Tm=79.8+18.5(log10[Na+]a)+0.58(%G / Cb)+11.8(%G / Cb)2-820 / Lc
[0111] oligo-DNA or oligo-RNA hybrids:
[0112] For <20 nucleotides: Tm = 2(ln)
[0113] For 20-35 nucleotides: Tm = 22 + 1.46 (ln)
[0114] a or for other monovalent cations, but is only accurate in the range 0.01–0.4 M.
[0115] bFor %GC, accurate only in the range of 30% to 75%.
[0116] c L = length of the duplex in base pairs.
[0117] d Oligo, oligonucleotide; ln, effective length of primer = 2×(number of G / C) + (number of A / T).
[0118] Non-specific binding can be controlled using any of a number of known techniques (e.g., blocking the membrane with a protein-containing solution, adding heterologous RNA, DNA, and SDS to the hybridization buffer, and treating with RNase). For non-related probes, a series of hybridizations can be performed by varying one of the following: (i) gradually lowering the annealing temperature (e.g., from 68° C. to 42° C.) or (ii) gradually lowering the formamide concentration (e.g., from 50% to 0%). The skilled artisan is aware of the various parameters that can be varied and maintained or altered during hybridization to maintain stringency.
[0119] In addition to hybridization conditions, the specificity of hybridization typically also depends on the function of washing after hybridization. In order to remove the background produced by non-specific hybridization, the sample is washed with a dilute salt solution. The key factors of this type of washing include the ionic strength and temperature of the final wash solution: the lower the salt concentration and the higher the wash temperature, the higher the stringency of the washing. Washing conditions are typically carried out under hybridization stringency or lower than hybridization stringency. Positive hybridization produces a signal that is at least twice the background signal. Usually, suitable stringent conditions for nucleic acid hybridization determination or gene amplification detection procedures are as described above. More stringent or less stringent conditions can also be selected. Those skilled in the art know that the various parameters that can change and maintain or change stringency conditions during washing are known.
[0120] For example, typical high stringency hybridization conditions for DNA hybrids longer than 50 nucleotides include hybridization in 1xSSC at 65°C or hybridization in 1xSSC and 50% formamide at 42°C, followed by washing in 0.3xSSC at 65°C. Examples of medium stringency hybridization conditions for DNA hybrids longer than 50 nucleotides include hybridization in 4xSSC at 50°C or hybridization in 6xSSC and 50% formamide at 40°C, followed by washing in 2xSSC at 50°C. The length of the hybrid is the expected length of the nucleic acid hybridization. When nucleic acids of known sequence are hybridized, the hybrid length can be determined by aligning the sequences and identifying the conserved regions described herein. 1xSSC is 0.15M NaCl and 15mM sodium citrate; the hybridization solution and washing solution may additionally contain 5xDenhardt's reagent, 0.5%-1.0% SDS, 100 μg / ml denatured fragmented salmon sperm DNA, and 0.5% sodium pyrophosphate. Another example of high stringency conditions is hybridization in 0.1x SSC containing 0.1 SDS and optionally 5x Denhardt's reagent, 100 μg / ml denatured fragmented salmon sperm DNA, 0.5% sodium pyrophosphate at 65°C, followed by a wash in 0.3x SSC at 65°C.
[0121] For the purpose of defining the level of stringency, reference may be made to Sambrook et al. (2001) Molecular Cloning: a laboratory manual, 3rd ed., Cold Spring Harbor Laboratory Press, CSH, New York or Current Protocols in Molecular Biology, John Wiley & Sons, New York (1989 and annual updates).
[0122] "Identity": When used in reference to comparing two or more nucleic acid or amino acid molecules, "identity" means that the sequences of the molecules share a degree of sequence similarity such that the sequences are partially identical.
[0123] Enzyme variants can be defined by their sequence identity when compared with the parent enzyme.Sequence identity is usually provided in the form of "sequence identity %" or "identity %". In order to determine the identity percentage between two amino acid sequences in the first step, a paired sequence alignment is generated between the two sequences, wherein the two sequences are aligned (that is, paired global alignment) over their complete length. With the program implementing Needleman and Wunsch algorithm (J.Mol.Biol. [J.Molecular Biology] (1979) 48, 443-453 pages), preferably by using program "NEEDLE" (European Molecular Biology Open Software Suite, EMBOSS) with program default parameters (gap open=10.0, gap extension=0.5 and matrix=EBLOSUM62) to generate an alignment. Preferred alignment for the purpose of the present invention is an alignment from which the highest sequence identity can be determined.
[0124] The following example is intended to illustrate two nucleotide sequences, but the same calculations apply to protein sequences:
[0125] Sequence A: AAGATACTG, length: 9 bases
[0126] Sequence B: GATCTGA, length: 7 bases
[0127] Therefore, the shorter sequence is sequence B.
[0128] Produces a pairwise global alignment of two sequences showing their full length, resulting in
[0129] Sequence A: AAGATACTG-
[0130] ||||||
[0131] Sequence B: --GAT-CTGA
[0132] The "I" symbol in the alignment indicates an identical residue (this means a base for DNA or an amino acid for protein). The number of identical residues is 6.
[0133] The "-" symbol in the alignment indicates a gap. The number of gaps introduced by the alignment within sequence B is 1. The number of gaps introduced by the alignment at the boundaries of sequence B is 2, while the number of gaps introduced at the boundaries of sequence A is 1.
[0134] The alignment length for the full length of aligned sequences is 10.
[0135] Thus, according to the present invention, pairwise alignments of shorter sequences showing the full length are generated, resulting in:
[0136] Sequence A: GATACTG-
[0137] ||||||
[0138] Sequence B: GAT-CTGA
[0139] Thus, according to the present invention, a pairwise alignment of sequence A showing the entire length is generated, resulting in:
[0140] Sequence A: AAGATACTG
[0141] ||||||
[0142] Sequence B: --GAT-CTG
[0143] Thus, according to the present invention, a pairwise alignment showing the entire length of sequence B is generated, resulting in:
[0144] Sequence A: GATACTG-
[0145] ||||||
[0146] Sequence B: GAT-CTGA
[0147] The alignment length of the shorter sequence shown is 8 (there is a gap which is included in the alignment length of the shorter sequence).
[0148] Therefore, the alignment length showing the full length of sequence A is 9 (meaning that sequence A is a sequence of the present invention).
[0149] Therefore, the alignment length showing the full length of sequence B is 8 (meaning that sequence B is a sequence of the present invention).
[0150] After comparing two sequences, in the second step, the identity value is determined according to the comparison produced. For the purpose of this description, the identity percentage is calculated as follows: identity %=(same residue / the length of the comparison area of the corresponding sequence of the present invention showing the complete length) * 100. Therefore, the sequence identity related to the comparison of two amino acid sequences according to the present embodiment is calculated by dividing the number of identical residues by the length of the comparison area of the corresponding sequence of the present invention showing the complete length. This value is multiplied by 100 and obtains " identity % ". According to the example provided above, identity %: for sequence A as a sequence of the present invention, it is (6 / 9) * 100 = 66.7%; for sequence B as a sequence of the present invention, it is (6 / 8) * 100 = 75%.
[0151] The terms "introducing," "introduction," and the like, with respect to the introduction of a donor DNA molecule at a target site of a target DNA, refer to any introduction of the sequence of the donor DNA molecule into the target region, such as by physically integrating the donor DNA molecule or a portion thereof into the target region or using the donor DNA as a template for a polymerase to introduce the sequence of the donor DNA molecule or a portion thereof into the target region.
[0152] Isogenic: Genetically identical organisms (eg, plants) except that they may differ by the presence or absence of heterologous DNA sequences.
[0153] Separated: As used herein, the term "isolated" means a material that has been artificially removed and exists apart from its original natural environment and is therefore not a natural product. An isolated material or molecule (such as a DNA molecule or enzyme) can exist in a purified form or can exist in a non-natural environment (such as a transgenic host cell). For example, a naturally occurring polynucleotide or polypeptide present in a living plant is not isolated, while the same polynucleotide or polypeptide separated from some or all of the coexisting materials in the natural system is isolated. Such polynucleotides can be part of a vector and / or such polynucleotides or polypeptides can be part of a composition and will be isolated because such a vector or composition is not part of its original environment. Preferably, the term "isolated" when used with respect to a nucleic acid molecule (such as in "isolated nucleic acid sequence") refers to a nucleic acid sequence that is identified and separated from at least one contaminating nucleic acid molecule that is typically associated with it in its natural source. An isolated nucleic acid molecule is a nucleic acid molecule that exists in a form or environment different from that in which it is found in nature. In contrast, non-isolated nucleic acid molecules are nucleic acid molecules that are found in the state in which they exist in nature, such as DNA and RNA. For example, a given DNA sequence (e.g., a gene) is found on a host cell chromosome close to adjacent genes; an RNA sequence (e.g., a specific mRNA sequence encoding a specific protein) is found in a cell as a mixture with a variety of other mRNAs encoding a variety of proteins. However, an isolated nucleic acid sequence comprising, for example, SEQ ID NO: 1 includes, for example, a nucleic acid sequence that normally contains SEQ ID NO: 1 in a cell, wherein the nucleic acid sequence is at a chromosomal position or extrachromosomal position different from that in the natural cell, or is otherwise flanked by nucleic acid sequences that are different from the nucleic acid sequence found in nature. An isolated nucleic acid sequence can exist in single-stranded or double-stranded form. When an isolated nucleic acid sequence is used to express a protein, the nucleic acid sequence will contain at least a portion of the sense strand or coding strand (i.e., the nucleic acid sequence can be single-stranded). Alternatively, it can contain both the sense strand and the antisense strand (i.e., the nucleic acid sequence can be double-stranded).
[0154] Minimal promoter: A promoter element, particularly a TATA element, that is inactive or has greatly reduced promoter activity in the absence of upstream activators. In the presence of appropriate transcription factors, a minimal promoter functions to allow transcription.
[0155] Non-coding: The term "non-coding" refers to sequences of a nucleic acid molecule that do not encode part or all of the expressed protein. Non-coding sequences include, but are not limited to, introns, enhancers, promoter regions, 3' untranslated regions, and 5' untranslated regions.
[0156] Nucleic acid and nucleotide: The terms "nucleic acid" and "nucleotide" refer to naturally occurring or synthetic or artificial nucleic acids or nucleotides. The terms "nucleic acid" and "nucleotide" include deoxyribonucleotides or ribonucleotides or any nucleotide analogs and polymers, or hybrids thereof in single-stranded or double-stranded, sense or antisense form. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), and complementary sequences as well as explicitly indicated sequences. The term "nucleic acid" is used interchangeably herein with "gene," "cDNA," "mRNA," "oligonucleotide," and "polynucleotide." Nucleotide analogs include nucleotides having modifications in the chemical structure of the base, sugar, and / or phosphate, including but not limited to 5-position pyrimidine modifications, 8-position purine modifications, modifications on the cytosine exocyclic amine, substitutions of 5-bromo-uracil, and the like; and 2'-position sugar modifications, including but not limited to sugar-modified ribonucleotides in which the 2'-OH is replaced by a group selected from H, OR, R, halogen, SH, SR, NH2, NHR, NR2, or CN. Short hairpin RNAs (shRNAs) can also contain unnatural elements, such as unnatural bases (eg, ionosin and xanthine), unnatural sugars (eg, 2'-methoxyribose), or unnatural phosphodiester linkages (eg, methylphosphonate, phosphorothioate, and peptides).
[0157] Nucleic acid sequence: The phrase "nucleic acid sequence" refers to a single-stranded or double-stranded polymer of deoxyribonucleotide or ribonucleotide bases read from the 5' end to the 3' end. It includes chromosomal DNA, self-replicating plasmids, infectious polymers of DNA or RNA, and DNA or RNA that plays a primary structural role. "Nucleic acid sequence" also refers to a continuous list of abbreviations, letters, characters or words representing nucleotides. In one embodiment, a nucleic acid can be a "probe", which is a relatively short nucleic acid, generally less than 100 nucleotides in length. Typically, nucleic acid probes are about 50 nucleotides to about 10 nucleotides in length. The "target region" of a nucleic acid is the portion of the nucleic acid that is identified as the object of interest. The "coding region" of a nucleic acid is the portion of the nucleic acid that, when placed under the control of appropriate regulatory sequences, is transcribed and translated in a sequence-specific manner to produce a specific polypeptide or protein. The coding region is considered to encode such a polypeptide or protein.
[0158] Oligonucleotide: The term "oligonucleotide" refers to an oligomer or polymer of ribonucleic acid (RNA) or deoxyribonucleic acid (DNA) or mimetics thereof, as well as functionally similar oligonucleotides having non-naturally occurring portions. Such modified or substituted oligonucleotides are generally preferred over the native forms because they possess desirable properties such as, for example, enhanced cellular uptake, enhanced affinity for nucleic acid targets, and increased stability in the presence of nucleases. An oligonucleotide preferably comprises two or more nucleomonomers covalently coupled to each other by a linkage (e.g., phosphodiester) or alternative linkage.
[0159] Overhang: An "overhang" is a relatively short single-stranded nucleotide sequence at the 5'- or 3'-hydroxyl end of a double-stranded oligonucleotide molecule (also called an "extension," "bulging end," or "sticky end").
[0160] Plant: Generally understood to mean any eukaryotic unicellular or multicellular organism capable of photosynthesis, or its cells, tissues, organs, parts or propagation materials (such as seeds or fruits). For the purposes of the present invention, all genera and species of higher and lower plants of the plant kingdom are included. Annual, perennial, monocotyledonous and dicotyledonous plants are preferred. The term includes mature plants, seeds, sprouts and seedlings, as well as derived parts, propagation materials (such as seeds or microspores), plant organs, tissues, protoplasts, callus and other cultures (e.g., cell cultures), and any other type of plant cell classified as producing functional or structural units. Mature plants refer to plants at any desired developmental stage beyond seedlings. Seedlings refer to young, immature plants in early developmental stages. Annual, perennial, monocotyledonous and dicotyledonous plants are preferred host organisms for generating transgenic plants. In addition, it is also advantageous to express genes in all ornamental plants, useful or ornamental trees, flowers, cut flowers, shrubs or lawns. Plants that may be mentioned by way of example and not limitation are angiosperms; bryophytes, such as, for example, Hepaticae (livers) and Musci (mosses); ferns, such as ferns, horsetails and club mosses; gymnosperms, such as conifers, cycads, ginkgo and Gnetatae; algae, such as Chlorophyceae, Phaeophyceae, Rhodophyceae, Myxophyceae, Xanthophyceae, Bacillariophyceae (diatoms) and Euglenophyceae.Preference is given to plants used for food or feed purposes, such as Leguminosae, such as peas, alfalfa and soybeans; Gramineae, such as rice, maize, wheat, barley, sorghum, millet, rye, triticale or oats; Umbelliferae, especially the genus Daucus, very especially the species Daucus carota (carrot), and the genus Apium, very especially the species Apium Graveolens dulce (celery) and many other plants; Solanaceae, especially the genus Lycopersicon, very especially the species Lycopersicon esculentum (tomato), and the genus Solanum, very especially the species Solanum tuberosum (potato) and Solanum melongena (eggplant) and many other plants (such as tobacco); and Capsicum, very especially the species Capsicum. annuum / pepper) and many other plants; the family Leguminosae, in particular the genus Glycine, very especially the species Glycine max / soybean, alfalfa, pea, lucerne, beans or peanuts and many other plants; and the family Cruciferae / Brassicacae, in particular the genus Brassica, very especially the species Brassica napus (rape), Brassica campestris (beet), Brassica oleracea cv Tastie (cabbage), Brassica oleracea cv Snowball Y (cauliflower) and Brassica oleracea cv Emperor (broccoli); and the genus Arabidopsis, very especially the species Arabidopsis thaliana thaliana) and many other plants; the Compositae family, especially the genus Lactuca, very especially the species Lactuca sativa (lettuce) and many other plants; the Asteraceae family, such as sunflower, marigold, lettuce or calendula and many other plants; the Cucurbitaceae family, such as melon, pumpkin (squash), or zucchini and linseed.Further preferred are cotton, sugar cane, hemp, flax, chillies, and various tree, nut and vine species.
[0161] Polypeptide: The terms "polypeptide," "peptide," "oligopeptide," "polypeptide," "gene product," "expression product," and "protein" are used interchangeably herein to refer to a polymer or oligomer of consecutive amino acid residues.
[0162] Preprotein: A protein that is typically targeted to an organelle, such as a chloroplast, and still contains its transit peptide.
[0163] "Precisely" with respect to the introduction of a donor DNA molecule in a target region means that the sequence of the donor DNA molecule is introduced into the target region without any insertions, deletions, duplications or other mutations as compared to the unaltered DNA sequence of the target region not contained in the sequence of the donor DNA molecule.
[0164] Primary transcript: As used herein, the term "primary transcript" refers to an immature RNA transcript of a gene. A "primary transcript" for example still contains introns, and / or does not yet contain a poly A tail or cap structure, and / or lacks other modifications (such as, for example, trimming or editing) necessary for proper function as a transcript.
[0165] Promoter: The terms "promoter" and "promoter sequence" are equivalent and, as used herein, refer to a DNA sequence that, when linked to a nucleotide sequence of interest, is capable of controlling the transcription of the nucleotide sequence of interest into RNA. Such promoters can be found, for example, in the following public databases: http: / / www.grassius.org / grasspromdb.html; http: / / mendel.cs.rhul.ac.uk / mendel.php?topic=plantprom; http: / / ppdb.gene.nagoya-u.ac.jp / cgi-bin / index.cgi. The promoters listed therein can be used in the methods of the present invention and are incorporated herein by reference. The promoter is located 5' (i.e., upstream) of the nucleotide sequence of interest that is transcribed into mRNA under its control, adjacent to the transcription start site, and provides a site for specific binding by RNA polymerase and other transcription factors to initiate transcription. The promoter comprises, for example, at least 10 kb, such as 5 kb or 2 kb, adjacent to the transcription start site. It can also comprise at least 1500bp, preferably at least 1000bp, more preferably at least 500bp, even more preferably at least 400bp, at least 300bp, at least 200bp or at least 100bp adjacent to the transcription start site. In another preferred embodiment, the promoter comprises at least 50bp, for example, at least 25bp, adjacent to the transcription start site. The promoter does not comprise exons and / or introns or 5' untranslated regions. The promoter can be, for example, heterologous or homologous to the corresponding plant. If the polynucleotide sequence is derived from an exogenous species, or if derived from the same species but modified relative to its original form, the polynucleotide sequence is "heterologous" to the organism or the second polynucleotide sequence. For example, a promoter operably linked to a heterologous coding sequence refers to a coding sequence from a species different from the promoter source species, or if from the same species, the coding sequence (for example, a genetically engineered coding sequence or an allele thereof from a different ecotype or variant) is not naturally associated with the promoter. Suitable promoters can be derived from the genes of the host cell in which expression should occur, or from the pathogen of this host cell (for example, a plant or a plant pathogen such as a plant virus). A plant-specific promoter is a promoter suitable for regulating expression in a plant. It can be derived from a plant, it can be derived from a plant pathogen, or it can be a synthetic promoter designed by a person. If the promoter is an inducible promoter, the transcription rate increases in response to an inducing agent. Similarly, the promoter can be regulated in a tissue-specific or tissue-preferred manner so that it is only active in or primarily in transcribing the associated coding region in one or more specific tissue types (such as leaves, roots or meristems).The term "tissue specificity" refers to a promoter that can guide the selective expression of a target nucleotide sequence for a specific tissue type (for example, petal), and the same target nucleotide sequence is relatively lacking in different types of tissues (for example, root). The tissue specificity of the promoter can be evaluated as follows, for example, a reporter gene and a promoter sequence can be operably connected to generate a reporter construct, the reporter construct is introduced into the genome of the plant so that the reporter construct is integrated into every kind of tissue of the transgenic plant gained, and the expression of the reporter gene in the different tissues of the transgenic plant (for example, detecting the activity of mRNA, protein or protein encoded by the reporter gene) is detected. The expression level relative to the reporter gene in other tissues is detected, and the expression level of the reporter gene in one or more tissues is higher. The promoter is specific to the tissue that detects a higher expression level. The term "cell type specificity" refers to a promoter that can guide the selective expression of a target nucleotide sequence in a specific type of cell when applied to a promoter, and the same target nucleotide sequence is relatively lacking in different types of cells in the same tissue. The term "cell type specificity" also means that a promoter can promote the selective expression of a target nucleotide sequence in the region within a single tissue when applied to a promoter. The cell type specificity of the promoter can be assessed using methods well known in the art (e.g., GUS activity staining, GFP protein or immunohistochemical staining). The term "constitutive" when used with respect to a promoter or expression derived from a promoter means that the promoter is capable of directing an operably linked nucleic acid molecule to be transcribed in most plant tissues and cells throughout the entire life cycle of a plant or plant part in the absence of stimulation (e.g., heat shock, chemicals, light, etc.). Typically, a constitutive promoter is capable of directing transgene expression in substantially any cell and any tissue.
[0166] Promoter specificity: The term "specificity" when referring to a promoter refers to the expression pattern conferred by the respective promoter. Specificity describes the tissue and / or developmental state of a plant or part thereof in which a promoter confers expression of a nucleic acid molecule under its control. Promoter specificity can also include environmental conditions under which a promoter can be activated or downregulated, such as induction or repression by biological or environmental stresses, such as cold, drought, wounding, or infection.
[0167] Purified: As used herein, the term "purified" refers to a molecule (nucleic acid or amino acid sequence) that is removed, isolated, or separated from its natural environment. A "substantially purified" molecule is at least 60% free, preferably at least 75% free, and more preferably at least 90% free from other components with which it is naturally associated. A purified nucleic acid sequence can be an isolated nucleic acid sequence.
[0168] Recombinant: The term "recombinant" with respect to nucleic acid molecules, DNA or genetic elements refers to nucleic acid molecules produced by recombinant DNA technology. Recombinant nucleic acid molecules can also include molecules that are not themselves naturally occurring, but have been modified, altered, mutated or otherwise manipulated by humans. Preferably, a "recombinant nucleic acid molecule" is a non-naturally occurring nucleic acid molecule that differs from the sequence of a naturally occurring nucleic acid molecule by at least one nucleic acid. A "recombinant nucleic acid molecule" can also include a "recombinant construct" that comprises a non-naturally occurring nucleic acid molecule sequence, preferably operably linked, in that order. Preferred methods for producing such recombinant nucleic acid molecules can include cloning techniques, directed or non-directed mutagenesis, synthesis or recombination techniques.
[0169] Sense: The term "sense" is understood to mean a nucleic acid molecule having a sequence that is complementary or identical to a target sequence, such as a sequence that binds to a protein transcription factor and is involved in the expression of a given gene. According to a preferred embodiment, the nucleic acid molecule comprises a gene of interest and an element that allows the expression of the gene of interest.
[0170] Significant increase or decrease: for example, an increase or decrease in enzyme activity or gene expression that is greater than the error range inherent in the measurement technique, preferably an increase or decrease of about 2-fold or more, more preferably an increase or decrease of about 5-fold or more, and most preferably an increase or decrease of about 10-fold or more, compared to the control enzyme activity or expression in a control cell.
[0171] Small nucleic acid molecules: "Small nucleic acid molecules" are understood to be molecules consisting of nucleic acids or derivatives thereof (such as RNA or DNA). They can be double-stranded or single-stranded and are between about 15 and about 30 bp (e.g., between 15 and 30 bp), more preferably between about 19 and about 26 bp (e.g., between 19 and 26 bp), even more preferably between about 20 and about 25 bp (e.g., between 20 and 25 bp). In particularly preferred embodiments, the oligonucleotides are between about 21 and about 24 bp, for example, between 21 and 24 bp. In a most preferred embodiment, the small nucleic acid molecules are between about 21 bp and about 24 bp, for example, 21 bp and 24 bp.
[0172] Substantially complementary: In its broadest sense, the term "substantially complementary" when used herein in relation to a nucleotide sequence related to a reference or target nucleotide sequence means a nucleotide sequence having a percentage identity of at least 60%, more ideally at least 70%, even more ideally at least 80% or 85%, preferably at least 90%, more preferably at least 93%, still more preferably at least 95% or 96%, yet more preferably at least 97% or 98%, yet more preferably at least 99% or most preferably 100% (the latter being equivalent to the term "identical" in this context) between the substantially complementary nucleotide sequence and the exact complement of the reference or target nucleotide sequence. Preferably, identity to the reference sequence is assessed over at least 19 nucleotides, preferably at least 50 nucleotides, more preferably over the entire length of the nucleic acid sequence (if not otherwise specified below). Sequence comparisons are performed based on the algorithm of Needleman and Wunsch (Needleman and Wunsch (1970) J Mol. Biol. 48:443-453; as defined above) using the default GAP analysis of the SEQWEB application of the University of Wisconsin GCG, GAP. A nucleotide sequence that is "substantially complementary" to a reference nucleotide sequence is hybridized to the reference nucleotide sequence under low stringency conditions, preferably medium stringency conditions, and most preferably high stringency conditions (as defined above).
[0173] As used herein, "target site" means the following location in the genome at which a double-strand break or one or a pair of single-strand breaks (nicks) are induced using recombinant technology (such as Zn-fingers, TALENs, restriction enzymes, homing endonucleases, RNA-guided nucleases, RNA-guided nickases (such as CRISPR / Cas nucleases or nickases), etc.).
[0174] Transgene: As used herein, the term "transgene" refers to any nucleic acid sequence that is introduced into the genome of a cell through experimental manipulation. A transgene can be an "endogenous DNA sequence" or a "heterologous DNA sequence" (i.e., "foreign DNA"). The term "endogenous DNA sequence" refers to a nucleotide sequence that is naturally present in the cell into which it is introduced, provided that it does not contain some modification relative to the naturally occurring sequence (e.g., point mutations, the presence of a selectable marker gene, etc.).
[0175] Transgenic: The term transgenic when referring to an organism means one that has been transformed, preferably stably transformed, with a recombinant DNA molecule that preferably comprises a suitable promoter operably linked to a DNA sequence of interest.
[0176] Vector: As used herein, the term "vector" refers to a nucleic acid molecule capable of transporting another nucleic acid molecule to which it is linked. One type of vector is a genomic integrating vector or "integration vector," which can be integrated into the chromosomal DNA of a host cell. Another type of vector is an episomal vector, a nucleic acid molecule capable of extrachromosomal replication. A vector capable of directing the expression of a gene to which it is operably linked is referred to herein as an "expression vector." Unless the context makes it clear otherwise, in this specification, "plasmid" and "vector" are used interchangeably. Expression vectors designed to produce RNA as described herein in vitro or in vivo can contain sequences recognized by any RNA polymerase, including mitochondrial RNA polymerase, RNA pol I, RNA pol II, and RNA pol III. According to the present invention, these vectors can be used to transcribe the desired RNA molecule in a cell. A plant transformation vector is understood to be a vector suitable for a plant transformation process.
[0177] Wild-type: The term "wild-type," "native," or "naturally derived" with respect to an organism, polypeptide, or nucleic acid sequence means that the organism is naturally occurring or obtainable in at least one naturally occurring organism that has not been altered, mutated, or otherwise manipulated by man. BRIEF DESCRIPTION OF THE DRAWINGS
[0178] Figure 1 and Figure 1B
[0179] For Agrobacterium strain GA00804 [(EHA105(pTiEHA105)(pBas05397)], the InDel% distribution in TO plants is shown in FIG1 , and Figure 1B The distribution of InDel% grouped by 2m-EPSPS copy number is shown in FIG.
[0180] Figure 2 and Figure 2B
[0181] For Agrobacterium strain GA00817 [(SHA017(pRi1599)(pBas05397)], the InDel% distribution in TO plants is shown in FIG2 , and Figure 2B The distribution of InDel% grouped by 2m-EPSPS copy number is shown in FIG.
[0182] Figure 3
[0183] Design of a construct for excising the Cas nuclease component. The excision construct contains the DNA fragment to be excised, which consists of the Cas12a nuclease gene, sgRNA, and Cre recombinase under the control of the pNtm19 promoter. The DNA fragment to be excised is flanked by two loxP recombination sites in a forward orientation. After excision, only the 2mepsps gene remains inserted into the genome. Dashed line: Genomic sequence flanking the transgene insert.
[0184] Examples
[0185] Chemicals and common methods
[0186] Unless otherwise indicated, the cloning procedure carried out for the purposes of the present invention is carried out as described in (Sambrook et al., 1989), including restrictive digestion, agarose gel electrophoresis, nucleic acid purification, nucleic acid connection, transformation of bacterial cells, selection and cultivation. The sequence analysis of recombinant DNA is carried out using Sanger technology (Sanger et al., 1977) with laser fluorescence DNA sequencer (Applied Biosystems, Foster City, California, USA). Unless otherwise described, chemicals and reagents are obtained from Sigma Aldrich (Sigma Aldrich, St. Louis, USA), Promega (Madison, Wisconsin, USA), Duchefa (Haarlem, Netherlands) or Invitrogen (Invitrogen) (Carlsbad, California, USA). Restriction endonucleases are from New England Biolabs (New England Biolabs) (Ipswich, Massachusetts, USA). Oligonucleotides were synthesized by Integrated DNA Technologies (Coralville, IA, USA).
[0187] Example 1: Soybean transformation
[0188] Mature seeds of Thorne were used for stable transformation. Seeds were surface sterilized using chlorine gas in a desiccator for approximately 16 hours as described by Di et al., 1996. Agrobacterium transformation using half-seed explants was performed essentially as described by Paz et al. (2006) and Luth et al. (2015). Co-cultivation of half-seed explants was performed using the disarmed strain EHA105 of Agrobacterium tumefaciens (Hood et al., 1993) and the disarmed strain SHA017 [K599(pRi2659)] of Agrobacterium rhizogenes (Mankin et al., 2007; WO 2006 / 024509), both of which harbor the T-DNA vector pBas05397. After 5 to 6 days of co-cultivation, glyphosate-resistant shoots were selected on selection medium containing 0.075 mM glyphosate.
[0189] Example 2: Plant transformation vector
[0190] The plant transformation vector pBas05397 contains a Cas12a nuclease expression cassette, a gRNA cassette, and a Cre recombinase cassette between the lox sites, and a 2m-EPSPS gene expression cassette outside the lox sites ( Figure 3). The Cas12a expression cassette comprises a soybean codon-optimized V-type CRISPR-associated protein Cas12a gene from Lachnospiraceae bacteria ND2006, which is driven by a promoter (Grefen et al., 2010) and a 3'35S terminator from the ubiquitin 10 (UBQ10) gene of Arabidopsis thaliana. The Cre recombinase is under the control of the tobacco (Nicotiana tabacum) Ntm19 promoter (Oldenhof et al., 1996) and a 3'35S terminator. The gRNA cassette comprises a gRNA (Sun et al., 2015) under the control of a promoter of a soybean (Glycine max) small nuclear RNA (U6-10) gene and is designed to direct the Lb Cas12a nuclease to the FAD2 target sequence (TS) TTTA-GTCCCTTATTTCTCATGGAAAAT (SEQ ID NO: 2). In the gRNA cassette, the sequence encoding the spacer RNA is between the direct repeats (DR) (Zetsche et al., 2016) and is preceded by a tRNA truncated at the 5' end (RNA from the glycine transfer RNA gene of wheat (Triticum aestivum) (Marcu KB. et al., 1977)). The T-DNA vector pBas05397 further contains a 2m-EPSPS gene expression cassette with the 2m-EPSPS gene under the control of the Arabidopsis thaliana Ph4a748_ABC histone promoter and a 3'His terminator to allow selection on glyphosate. The plant transformation vector pBas05019 is identical to pBas05397, but does not contain Cre recombinase and does not contain lox sites.
[0191] Example 3: Genotyping of transformants
[0192] To genotype individual soybean transformants, one top leaf of an "in vitro" shoot approximately 10 cm in size was harvested in a 1 mL tube on a 96-well plate for genomic DNA extraction.
[0193] The copy number of the 2m-EPSPS gene was determined by real-time PCR (Ingham et al., 2001). InDel efficiency was measured using ddPCR drop-off assay. ddPCR assays were designed using Primer3Plus software, using a modified setup compatible with the applied master mix. In order to avoid losing the binding site, primers and reference probes were designed to be away from the cleavage site. PCR primers were designed according to the following guidelines: primer length was 17-24 bases, primer melting temperature was 55°C to 60°C (wherein the ideal temperature was 58°C, and the melting temperature difference between the two primers was no more than 2°C), primer GC content was 35%-65%, and amplicon size was 100-250 bases. When one or more base substitutions, insertions or deletions were introduced at the FAD2 target site, the drop-off probe was designed to lose its binding site. The sequences of probes and primers are shown in Table 1.
[0194] Table 1. Primers and probes used in ddPCR dropout assays to determine editing frequency at FAD2 target sites.
[0195] Primer / probe type sequence SEQ ID NO Forward primer TGATTGCTCACGAGTGT 3 Reverse primer CAGAGACATTGAAGGCTAAA 4 Reference probe (FAM) CCGTGATGAAGTGTTTGTCCCA 5 Shedding probe (HEX) CTCATGGAAAATAAGCCATCG 6
[0196] To screen for "in vitro" buds in which the Cas12a nuclease was removed by Cre / lox recombinase under the control of the Ntm19 promoter, the buds were further analyzed by PCR using a primer set (HT-19-001 forward / HT-22-007 reverse) for specific amplification of the Cas12a nuclease (Table 2).
[0197] Table 2. Primers used for PCR to screen for the presence or absence of Cas12a in buds.
[0198] Primers sequence SEQ ID NO HT-19-001 AAGAAGAAGAGGAAGGTGGG 7 HT-22-007 GAGAGGTCAGCATCAGCATAC 8
[0199] To identify different types of mutations and their frequencies at the FAD2 target site, NGS was performed. The region surrounding the target site was PCR amplified using the primer pair HT-21-080 forward / HT-21-081 reverse with Q5 High-Fidelity Polymerase (M0492L) to amplify a 354 bp region (Table 3).
[0200] Table 3. Primer pairs used to amplify a 354 bp target region for NGS.
[0201] Primers sequence SEQ ID NO HT-21-080 CCTCATTGCATGGCCAATC 9 HT-21-081 CCAGAGACATTGAAGGCTAAATAC 10
[0202] Example 4: Soybean genome editing by removing nuclease components during tissue culture
[0203] Half seeds were co-cultivated with the Agrobacterium tumefaciens loss-of-function strain GA00804 (EHA105 (pTiEHA105) (pBas05397)) or the Agrobacterium rhizogenes loss-of-function strain GA00817 (SHA017 (pRi2659) (pBas05397)). For each strain, 80 independent glyphosate-resistant shoots were analyzed, and the percentage of InDel at the FAD2 target site was determined by ddPCR dropout and the copy number of the 2m-EPSPS gene was determined by real-time PCR ( Figure 1A 、 1B Single copy 2m-EPSPS events were obtained at frequencies of 60% and 67% using Agrobacterium strains GA00804 and GA00817, respectively. Figure 1B and 2B ).
[0204] In order to screen " in vitro " buds in which Cas12a nuclease is removed by Cre / lox recombinase under the control of Ntm19 promoter, buds are further analyzed by PCR using primer sets (HT-19-001 forward / HT-22-007 reverse) to carry out specific amplification of Cas12a nuclease. For GA00804 or GA00817 strains, PCR analysis of the presence or absence of Cas12a nuclease carried out by transformants using primer sets (HT-19-001 forward / HT-22-007 reverse) showed that 33% (26 of 80 transformants obtained by GA00804) and 40% (32 of 80 transformants obtained by GA00817) did not produce PCR products, indicating that the nuclease component has been removed. When the transformation vector pBas05019 was identical to the transformation vector pBAS05397 but lacked the Cre recombinase and lox sites, 99% (158 / 160) of the transformants still contained the Cas12a nuclease.
[0205] NGS analysis was performed on a subset of 1 copy 2m-EPSPS GA00804 and GA00817 transformants, from which Cas12a nuclease components were removed by Cre / lox recombination. For each specific mutation, the mutation frequency at the FAD2 target locus was assessed by calculating the percentage of sequence reads with a specific mutation (deletion) as the ratio of the total number of reads. These data are summarized in Table 4. Mutation is mainly deletion. For example, event TMGM0139-058-01$001 has 4 types of deletions with 8, 12, 14 and 22 nucleotides, occurring at a frequency of 24.7%, 21.5%, 26.4% and 26.8% respectively. The control event TMGM0139-Ctrl006-01$001 derived from half a seed explant not co-cultivated with Agrobacterium shows that the frequency of WT readings is about 100%, as expected. Table 4 further shows that the editing frequencies obtained by ddPCR dropout correlate very well with those obtained by NGS, demonstrating that our assay is reliable.
[0206] These data show that the nuclease component can be removed in the early tissue culture stage by Cre / lox recombination using Cre recombinase under the control of the Ntm19 promoter before transferring the TO plants to the greenhouse, thereby avoiding the generation of new somatic non-inherited mutations at later developmental stages.
[0207] Table 4.1 Percentage (%) of reads for each specific mutation at the FAD2 target site of copy 2m-EPSPS events in which the Cas12a nuclease component was removed by Cre / lox recombination performed by Cre recombinase under the control of the Ntm19 promoter.
[0208]
[0209]
[0210] The edited plants lacking Cas12a nuclease were transferred to the greenhouse, seeds were produced by selfing, and the inheritance of the edits was evaluated in the T1 progeny.
[0211] Example 5: Genome editing of rapeseed by removing nuclease components during tissue culture
[0212] In another experiment, the removal of the nuclease component in the tissue culture stage was tested in rapeseed (Brassica napus) protoplasts. Rapeseed protoplasts were separated from the leaves of 4 to 7 week-old sterile growth strains. Healthy leaves were cut into thin strips with a sharp razor blade. The thin strips were infiltrated with a cell wall lytic enzyme solution (1.5% cellulase R10 and 0.75% macerate R10 in 10mM KCl and 0.6M mannitol, pH 7.5), and incubated overnight with gentle shaking (40rpm) at 24°C in the dark. After enzymatic digestion, the protoplasts released were collected by filtering the mixture through a 40-μm nylon mesh, and were resuspended in the W5 solution. The resuspended protoplasts were kept on ice and allowed to settle by gravity, and the cell pellets were then resuspended in the MMG. For transformation, 200 μl of cells (2.5 x 10 5 ) were mixed with 20 μg of plasmid DNA and 220 μl of freshly prepared polyethylene glycol (PEG) solution. The mixture was incubated in the dark for 15-20 min. After removing the PEG solution, protoplasts were embedded in an alginate layer (as described by Kielkowska and Adamus (2012)) and incubated at approximately 25°C. Protoplast-derived colonies were picked and transferred to a shoot induction medium (Murashige & Skoog, 1962) with a high cytokinin / auxin ratio, i.e., 1 mg / L BAP + 0.1 mg / L NAA or BAP or 3 mg / L BAP + 0.1 mg / L NAA + 0.1 mg / L GA3, and selected on glyphosate. Further plant regeneration was performed as described by De Block et al. (1989).
[0213] The plasmid DNA for conversion contains Cas12a nuclease expression cassette, gRNA box and Cre recombinase cassette between lox sites, and contains 2m-EPSPS gene expression cassette outside of lox sites.Cas12a expression cassette includes V-type CRISPR associated protein Cas12a gene, which is driven by promoter (Grefen et al., 2010) and 3'35S terminator from ubiquitin 10 (UBQ10) gene of Arabidopsis thaliana.Cre recombinase is under the control of tobacco (Nicotiana tabacum) Ntm19 promoter (Oldenhof et al., 1996) and 3'35S terminator.GRNA box includes gRNA under the control of polymerase III type promoter in Arabidopsis thaliana U6 snRNA gene, and is designed for Cas12a nuclease to be directed to BnFAD2 target sequence. In the gRNA cassette, the sequence encoding the spacer RNA is between the direct repeats (DR) (Zetsche et al., 2016) and is preceded by a truncated tRNA at the 5' end (RNA from the wheat glycine transfer RNA gene (Marcu KB. et al., 1977)). The plasmid further contains a 2m-EPSPS gene expression cassette with the 2m-EPSPS gene under the control of the Arabidopsis thaliana Ph4a748_ABC histone promoter and a 3'His terminator to allow selection on glyphosate. As a control, a plasmid identical to this plasmid but without Cre recombinase and lox sites was used.
[0214] For the transfected plasmids and the control plasmids, independent glyphosate-resistant shoots were analyzed to determine the % InDel at the FAD2 target site and to determine the copy number of the 2m-EPSPS gene.
[0215] In order to screen " in vitro " buds in which Cas12a nuclease is removed by Cre / lox recombinase under the control of Ntm19 promoter, buds are further analyzed by PCR using a primer set for specific amplification of Cas12a nuclease. PCR analysis using a primer set to the presence or absence of Cas12a nuclease shows that transformants that do not produce PCR products can be obtained, indicating that the nuclease component is removed. In the case of using a control plasmid, all transformants still contain Cas12a nuclease.
Claims
1. A method for excising one or more recombinant genetic elements from the genome of a somatic cell of a transgenic plant, the method comprising the steps of: a. introducing one or more recombinant genetic elements into the genome of the somatic cell, each recombinant genetic element comprising i. First excision recognition site, ii. a second excision recognition site, and a second excision recognition site between the first excision recognition site and the second excision recognition site. iii. a polynucleotide encoding an excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the ntm19 promoter, and b. expressing the excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site, wherein the excision component recognizes the excision site and excises the recombinant DNA located between the first excision recognition site and the second excision recognition site.
2. A method for transiently expressing a target gene in plant somatic cells, the method comprising the following steps: a. introducing one or more recombinant genetic elements into the genome of the somatic cell, each recombinant genetic element comprising i. First excision recognition site, ii. a second excision recognition site, and a second excision recognition site between the first excision recognition site and the second excision recognition site. iii. a polynucleotide encoding a resection component capable of resecting the recombinant DNA located between the first resection recognition site and the second resection recognition site under the control of the ntm19 promoter and the target gene, and b. expressing the excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site, wherein the excision component recognizes the excision site and excises the recombinant DNA located between the first excision recognition site and the second excision recognition site.
3. The method of claim 1 or 2, wherein the recombinant DNA located between the first excision recognition site and the second excision recognition site further comprises a sequence encoding at least one morphogenic gene, and / or a sequence encoding a genome editing component for editing the target sequence in the plant somatic cell, and / or a sequence encoding a selectable marker.
4. A method for improving the regeneration of transgenic plants from plant somatic cells, the method comprising the steps of: a. introducing one or more recombinant genetic elements into the genome of the somatic cell, each recombinant genetic element comprising i. the first excision recognition site, and ii. a second excision recognition site, and a second excision recognition site between the first excision recognition site and the second excision recognition site. iii. a polynucleotide encoding an excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the ntm19 promoter and at least one morphogenic gene functionally linked to the promoter, b. expressing the excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site, wherein the excision component recognizes and excises the recombinant DNA located between the first excision recognition site and the second excision recognition site.
5. The method of any one of claims 1 to 4, further comprising the step of selecting cells in which excision of the recombinant DNA between the first excision recognition site and the second excision recognition site has occurred.
6. The method of any one of claims 1 to 4, further comprising the steps of regenerating shoots or plantlets from the somatic cells and selecting shoots or plantlets in which excision of the recombinant DNA between the first excision recognition site and the second excision recognition site has occurred.
7. A method of producing a plant or shoot comprising an edit in a target sequence, the method comprising: a. Introducing one or more recombinant genetic elements into the genome of a somatic cell, each recombinant genetic element comprising i. First excision recognition site, ii. a second excision recognition site, and a second excision recognition site between the first excision recognition site and the second excision recognition site. iii. a polynucleotide encoding an excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the ntm19 promoter and a sequence encoding a genome editing component for editing the target sequence, and b. regenerating buds from said somatic cells, c. selecting buds comprising the edit in the target sequence and in which excision of the recombinant DNA between the first excision recognition site and the second excision recognition site has occurred, and optionally d. growing a plant from the shoot.
8. A method for removing a genome editing component shortly after introduction of a targeted genome modification, the method comprising: a. Introducing one or more recombinant genetic elements into the genome of a somatic cell, each recombinant genetic element comprising i. First excision recognition site, ii. a second excision recognition site, and a second excision recognition site between the first excision recognition site and the second excision recognition site. iii. a polynucleotide encoding an excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the ntm19 promoter and a sequence encoding a genome editing component for editing the target sequence, and b. regenerating buds from said somatic cells, c. Selecting buds that comprise the edit in the target sequence and in which excision of the recombinant DNA between the first excision recognition site and the second excision recognition site has occurred.
9. The method of claim 7 or 8, wherein the sequence encoding the genome editing components for editing the target sequence encodes a site-directed nuclease, such as a nucleic acid-guided DNA endonuclease, or such as a Cas nuclease and a guide RNA.
10. The method according to any one of claims 1 to 9, wherein the recombinant genetic element further comprises a gene of interest outside the first excision recognition site and the second excision recognition site.
11. The method of any one of claims 1 to 10, wherein the first excision recognition site and the second excision recognition site are lox sites, and wherein the excision component is a Cre recombinase protein.
12. The method of any one of claims 1 to 10, wherein the excision component is a nucleic acid-guided DNA endonuclease and one or two guide RNAs that guide the nucleic acid-guided DNA endonuclease protein to the first and second excision recognition sites.
13. A recombinant construct comprising i. the first excision recognition site, and ii. a second excision recognition site, and a second excision recognition site between the first excision recognition site and the second excision recognition site. iii. A polynucleotide encoding an excision component capable of excising the recombinant DNA located between the first excision recognition site and the second excision recognition site under the control of the ntm19 promoter.
14. The method of any one of claims 1 to 12 or the construct of claim 13, wherein the ntm19 promoter comprises a sequence selected from the group consisting of: a) a nucleic acid molecule having the sequence of SEQ ID NO: 1, and b) a nucleic acid molecule having a sequence that is at least 80% identical to SEQ ID NO: 1, and c) a fragment of at least 100 consecutive bases of the nucleic acid molecule of I) or II), which has the same activity as the corresponding nucleic acid molecule having the sequence of SEQ ID NO: 1, and d) a nucleic acid molecule that is the complement or reverse complement of any of the previously mentioned nucleic acid molecules in I) to III), and e) a nucleic acid molecule that hybridizes to a nucleic acid molecule comprising at least 50 consecutive nucleotides of SEQ ID NO: 1, or its complement, under conditions equivalent to hybridization in 7% sodium dodecyl sulfate (SDS), 0.5 M NaPO4, 1 mM EDTA at 50°C and washing in 2X SSC, 0.1% SDS at 50°C.
15. A vector comprising the recombinant construct according to claim 13 or 14.
16. A cell comprising the recombinant construct of claim 13 or 14 or the vector of claim 15.
Citation Information
Patent Citations
Compounds and methods for site directed mutations in eukaryotic cells
US5565350A
RAC-like genes from maize and methods of use
WO2000015815A1
Disarmed agrobacterium strains, ri-plasmids, and methods of transformation based thereon
WO2006024509A2