Increasing haploid editing efficiency
Patent Information
- Authority / Receiving Office
- IL · IL
- Patent Type
- Applications
- Current Assignee / Owner
- SYNGENTA CROP PROTECITON AG
- Filing Date
- 2024-12-20
- Publication Date
- 2026-08-01
AI Technical Summary
Current methods for introducing transgenic traits into elite inbred maize lines are inefficient, requiring multiple generations of backcrossing to eliminate the transformed parent's genome while retaining the transgenic trait, a process that can take several years.
The use of haploid induction followed by a heat treatment during the haploid editing process to increase the efficiency of gene editing in plant genomic DNA, allowing for the direct introduction of edited haploid progeny that retain the desired transgenic trait.
This method significantly accelerates the introduction of transgenic traits into maize lines, reducing the time required to produce homozygous inbred lines from decades to just two generations, and enhances the haploid induction and editing rates.
Smart Images

Figure 00000095_0000
Abstract
Description
INCREASING HAPLOID EDITING EFFICIENCYRELATED APPLICATION
[0001] This application claims priority to Patent Cooperative Treaty (PCT) Application No. PCT / CN2023 / 140616, filed on December 21, 2023. The entire content of said application is herein incorporated by reference for all purposes.FIELD
[0002] This invention is related to the field of plant biotechnology, specifically agriculture biotechnology and gene editing, as well as plant breeding. The presently disclosed subject matter relates to using a haploid inducing line (whether existing or created) and transforming the haploid inducing line so that it contains DNA coding for cellular machinery capable of editing genes.SEQUENCE LISTING
[0003] This application is accompanied by an XML formatted sequence listing entitled 83020-WO-REG-ORG-P-l.xml and created on December 2, 2024. The XML file is approximately 209 kilobytes in size and is filed concurrently with the specification. The sequence listings contained in the XML file are part of the specification and are incorporated herein by reference in their entirety.BACKGROUND
[0004] Plant transformation, that is, the stable integration of foreign DNA (“transgenes”) into a plant genome, has been used for decades to add new and useful traits to crops. While some plant lines, e.g., maize, are relatively easy to transform (i.e., accepting of transgenic DNA), most lines are not. For example, most elite inbred lines, which are produced by self-pollination over several generations to obtain a pure or nearly pure homozygous genome and which are used as parent lines to create commercially valuable hybrids, often cannot be transformed with foreign DNA. Thus, to move a transgenic trait into an inbred line, the transgenic trait must first be transformed into a transformable maize line. In maize, for example, that transformed maize line is rarely suitable for use as a parent line in breeding platforms. Therefore, the transformed maize line is crossed into an inbred line to create a progeny plant which will comprise, in a heterozygous manner, the genomes of both the inbred parent and the transformed parent. Then, that progeny plant comprising the transgene must be backcrossed into the inbred line forapproximately six or seven generations to eliminate, as much as possible, the genome contributed by the transformed parent while retaining the transgenic trait. This trait introgression process generally takes between three to seven years.
[0005] An important tool in plant breeding is haploid induction (HI), which is a class of plant phenomena characterized by loss of one parent's set of chromosomes (i.e., the chromosomes from the haploid inducer parent) from the embryo at some time during or after fertilization. Loss of one set of chromosomes often occurs during early embryo development. Haploid induction is also known as gynogenesis if the inducer line is used as the male in the cross or androgenesis if the inducer line is used as the female in the cross. HI has been observed in numerous plant species, such as sorghum, barley, wheat, maize, Arabidopsis, and many other species. Haploids are valuable when they are doubled (referred to as double haploid (DH) plants) and used to produce homozygous breeding lines. In homozygous lines, all genes on each pair of chromosomes in every cell of the plant are identical. These homozygous lines are 100-percent inbred lines, which otherwise would have to be produced by repeated forced self- pollinations. The haploid method lets breeders produce inbred lines within just two generations, while traditional breeding takes 10 generations. The most efficient way to produce doubled haploids in com is through haploid induction.
[0006] In maize, haploid seed or embryos can be produced by making crosses between a haploid inducer male (i.e., “haploid inducer pollen”) and virtually any ear that one chooses. In the case of maternal HI systems, e.g., matrilineal -based systems, haploids are produced when the haploid inducer pollen DNA (i.e., haploid inducer paternal DNA) is not fully transmitted and / or maintained through the first cell divisions of the embryos. The resulting kernels have haploid embryos that contain only the maternal DNA plus normal (fertilized) triploid endosperm. In the case of paternal HI systems, e.g., CENH3-based or igl-based systems, haploids are produced after the egg is fertilized by the sperm cell and the maternal chromosomes are lost upon cell division. The resulting kernels have haploid embryos that contain only the paternal DNA plus normal (fertilized) triploid endosperm. Regardless of the HI system used, the resulting phenotype is not fully penetrant, with some ovules containing haploid embryos and others containing diploid embryos, aneuploid embryos, chimeric embryos, or aborted embryos. After haploid induction, haploid embryos or seeds are typically segregated from diploid and aneuploid siblings using a phenotypic or genetic marker screen and grown or cultured into haploid plants. These plants are then converted either naturally or via chemical manipulation (e.g., using an anti -microtubule agent such as colchicine,pronamide, dithipyr, or trifluralin) into doubled haploid (DH) plants which then produce inbred seed.SUMMARY
[0007] The Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. This Summary is not intended to identify key or essential features of the claimed subject matter nor is it intended to be used as an aid in limiting the scope of the claimed subject matter.
[0008] In one aspect, provided is a method of editing plant genomic DNA, comprising (a) providing an egg-donor plant comprising plant genomic DNA that is to be edited; (b) pollinating the egg-donor plant with a pollen-donor plant, wherein the pollen-donor plant expresses a DNA modification enzyme and, optionally, a guide nucleic acid; (c) applying a heat treatment to the pollinated egg-donor plant of step b; and (d) producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the egg-donor plant and does not comprise the genome of the pollen-donor haploid inducer plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and, optionally, the guide nucleic acid delivered by the pollen-donor haploid inducer plant.
[0009] In another aspect, provided is a method of editing plant genomic DNA, comprising: (a) providing a pollen-donor plant that expresses a DNA modification enzyme and, optionally, a guide nucleic acid; (b) applying a heat treatment to pollen of the pollen-donor plant; (c) pollinating an egg-donor plant with heat-treated pollen of the pollen-donor plant, wherein the egg-donor plant comprises the plant genomic DNA that is to be edited; and (d) producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the egg-donor plant and does not comprise the genome of the pollen-donor plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and, optionally, the guide nucleic acid delivered by the pollen-donor plant.
[0010] In some embodiments of these methods, the pollen-donor plant is a haploid inducer plant. In some embodiments, the pollen-donor plant is a paternal haploid inducer plant. In some embodiments, the paternal haploid inducer plant comprises a knock-out mutation in MATL gene.
[0011] In another aspect, provided is a method of editing plant genomic DNA, comprising: (a) providing a pollen-donor plant comprising plant genomic DNA that is to be edited; (b)pollinating an egg-donor plant with pollen from the pollen-donor plant, wherein the egg-donor plant expresses a DNA modification enzyme and, optionally, a guide nucleic acid; (c) applying a heat treatment to the pollinated egg-donor plant of step b.; and (d) producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the pollendonor plant and does not comprise the genome of the egg-donor plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and, optionally, the guide nucleic acid delivered by the egg-donor plant.
[0012] A method of editing plant genomic DNA, comprising: (a) providing a pollen-donor plant comprising plant genomic DNA that is to be edited; (b) applying a heat treatment to pollen of the pollen-donor plant; (c) pollinating an egg-donor plant with heat-treated pollen of the pollen-donor plant, wherein the egg-donor plant expresses a DNA modification enzyme and, optionally, a guide nucleic acid; and (d) producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the pollen-donor plant and does not comprise the genome of the egg-donor plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and, optionally, the guide nucleic acid delivered by the egg-donor plant.
[0013] In some embodiments of these methods, the egg-donor plant is a haploid inducer plant. In some embodiments, the egg-donor plant is a maternal haploid inducer plant. In some embodiments, the maternal haploid inducer plant comprises a mutation in a CENH3 gene. In some embodiments, the maternal haploid inducer plant is heterozygous for the mutation in the CENH3 gene.
[0014] In embodiments of the provided methods, at least one of the egg-donor plant or the pollen-donor plant is a maize plant. In some embodiments, the pollen-donor plant is a maize plant. In some embodiments, the egg-donor plant is a maize plant. In some embodiments, the egg-donor plant is a wheat plant.
[0015] In some embodiments, the pollen-donor plant is a maize plant. In some embodiments, the maize plant selected and / or derived from the lines Stock 6, RWK, RWS, UH400, NP2222RS, or NP2222.
[0016] In some embodiments, the DNA modification enzyme is a site-directed nuclease selected from the group consisting of meganucleases (MNs), zinc-finger nucleases (ZFNs), transcription-activator like effector nucleases (TALENs), and Cas nucleases. In some embodiments, the Cas nuclease is a Type II Cas nuclease, a Type IV Cas nuclease, or a TypeV Cas nuclease. In some embodiments, the Type II Cas nuclease is a Cas9 nuclease, a Cas9 nickase, a nuclease-inactive Cas9, or a Cas9 fused to a heterologous domain. In some embodiments, the Type V Cas nuclease is a Casl2a nuclease, a Casl2a nickase, a nucleaseinactive Cas 12a, or a Cas 12a fused to a heterologous domain. In some embodiments, the guide nucleic acid is a guide RNA.
[0017] In some embodiments, the haploid progeny is treated with a chromosome doubling agent, thereby creating an edited doubled haploid progeny. In some embodiments, the chromosome doubling agent is colchicine, pronamide, dithipyr, trifluralin, or another known anti-microtubule agent.
[0018] In some embodiments, the pollen-donor plant expresses a marker gene. In some embodiments, the marker gene is selected from the group consisting of Rl, R1-SCM2, Rl-nj, GUS, PMI, PAT, GFP, RFP, CFP, Bl, CI, anthocyanin pigments, and any other marker gene.
[0019] In another aspect, provided is an edited haploid plant produced by any method of the present disclosure.
[0020] In another aspect, provided is a progeny plant of an edited haploid plant that is produced by any method of the present disclosure.BRIEF DESCRIPTION OF THE FIGURES
[0021] Figures 1 A-B show confocal microscopy images of maize pollen grains stained with DAPI DNA stain. Figure 1 A shows control pollen grain with typical sperm pair morphology, i.e., long, thin, and wispy. Figure IB shows a pair of sperm after a one-hour heat treatment at 45°C. The sperm exhibited diffused staining, a broadened shape, and overall larger size.BRIEF DESCRIPTION OF THE SEQUENCES IN THE SEQUENCE LISTING
[0022] SEQ ID NO: 1 is the nucleotide sequence encoding construct 27145.
[0023] SEQ ID NO: 2 is the nucleotide sequence encoding construct 27146.
[0024] SEQ ID NO: 3 is the nucleotide sequence encoding construct 27680.
[0025] SEQ ID NOs: 4-6 represent TaqMan® quantitative PCR assay 3895, using primersTCCTTGTTCCGTCTTTTGCAG (SEQ ID NO: 4) and AAGGCAAAAGGAGGGAACTGAT (SEQ ID NO: 5) and probe TACCTCGGCGACGCC (SEQ ID NO: 6).
[0026] SEQ ID NO: 7 is an exemplary MAIL variant cDNA nucleotide sequence.
[0027] SEQ ID NO: 8 is an exemplary CENH3 variant DNA nucleotide sequence.
[0028] SEQ ID NO: 9 is the nucleotide sequence encoding construct 28255.
[0029] SEQ ID NO: 10 is the nucleotide sequence encoding construct 28291.
[0030] SEQ ID NO: 11 is the nucleotide sequence encoding construct 28292.
[0031] SEQ ID NO: 12 is the nucleotide sequence encoding construct 28293.
[0032] SEQ ID NO: 13 is the nucleotide sequence encoding construct 28294.
[0033] SEQ ID NO: 14 is the nucleotide sequence encoding construct 25072.
[0034] SEQ ID NO: 15 is the nucleotide sequence encoding construct 28825.
[0035] SEQ ID NO: 16 is the nucleotide sequence encoding construct 28834.DETAILED DESCRIPTIONI. Terminology
[0036] All technical and scientific terms used herein, unless otherwise defined below, are intended to have the same meaning as commonly understood by one of ordinary skill in the art. References to techniques employed herein are intended to refer to the techniques as commonly understood in the art, including variations on those techniques and / or substitutions of equivalent techniques that would be apparent to one of skill in the art. While the following terms are believed to be well understood by one of ordinary skill in the art, the following definitions are set forth to facilitate explanation of the presently disclosed subject.
[0037] As used in herein, the singular forms “a,” “an,” and “the” include plural referents unless the content clearly dictates otherwise. Thus, for example, reference to “an antibody” optionally includes a combination of two or more such molecules and the like.
[0038] The term “about” as used herein refers to the usual error range for the respective value readily known to the skilled person in this technical field, for example ± 20%, ± 10%, or ± 5% are within the intended meaning of the recited value.
[0039] As used herein, the term “comprising” or “comprise” is open-ended. When used in connection with a subject nucleic acid (or amino acid sequence), it refers to a nucleic acid sequence (or an amino acid sequence) that includes the subject sequence as a part or as its entire sequence.
[0040] The term “plurality” refers to more than one entity. Thus, a “plurality of individuals” refers to at least two individuals. In some embodiments, the term plurality refers to more than half of the whole. For example, in some embodiments a “plurality of a population” refers to more than half the members of that population.
[0041] As used herein, the term “and / or” when used in the context of listing of entities, refers to the entities being present singly or in combination. Thus, for example, the phrase “A, B, C, and / or D” includes A, B, C, and D individually, but also includes any and all combinations and subcombinations of A, B, C, and D (e.g., AB, AC, AD, BC, BD, CD, ABC, ABD, and BCD). In some embodiments, one of more of the elements to which the “and / or” refers can also individually be present in single or multiple occurrences in the combinations(s) and / or subcombination(s).
[0042] The term “HI-Edit efficiency” refers to the measurement of progeny plants produced from a HI-Edit cross that are both edited and are (or were, if doubled) haploid. “HI-Edit efficiency,” “haploid editing rate,” “HI-Edit rate,” and “HER” are used interchangeably throughout. HI-Edit efficiency is usually expressed as a percentage.
[0043] “HI-Edit window” means a period of time during which the editing may occur prior to genome elimination, typically a few hours after pollination to three days after pollination. This period of time is not precise, and ambient environmental factors, as well as biological factors, such as pollen vigor, can affect the duration and / or initiation of the HI-Edit window.
[0044] “Heat treatment” refers to any treatment provided to the plant, room, or maize ear to increase the temperature thereof. In one example (the “heated-room treatment”), a heat treatment maintains the glasshouse temperature at 35°C during the day and 25°C during the night. In a second example (the “heat pack treatment”), a heat pack may be used as the heat treatment.
[0045] “Heat pack” as used herein, refers to a device which generates heat, whether by exothermal chemical reaction, such as a therapeutic heating patch, (see, e.g., THERMACARE® Muscle Pain Therapy HeatWraps (thermacare.com / heat-wraps / muscle- pain-therapy (last visited Aug. 15, 2023)) or by electrical resistance, such as an electric blanket (see, e.g., CONAIRCOMFORT® Standard Heating Pad (conair.com / en / standard-heating- pad / HP40.html (last visited Aug. 15, 2023)). The heat pack is wrapped around the plant or the portion of the plant where haploid induction occurs. For example, for corn plants, the heat pack is typically wrapped around the ear with an optional insulation barrier to keep the eartemperature at approximately 30-40°C. Other heat sources may be used as well. An insulation barrier may comprise wool, fiberglass, and the like.
[0046] The term “haploid inducer” or “haploid inducer line” refers to a plant line that triggers development of a pollinated egg cell into an embryo that contains only a haploid genome (i.e., induce the formation of haploid progeny). Progeny that results from crosses with a haploid inducer line lack the haploid inducer genome. The haploid inducer may be a pollen-donor haploid inducer or an egg-donor haploid inducer. The haploid inducer may be a paternal haploid inducer or a maternal haploid inducer. A “haploid inducer plant” is a plant of a haploid inducer line. One example of a haploid inducer plant is a maize plant comprising a mutation in ZmMATL. Another example is a maize plant comprising a mutation in a CENH3 gene. Yet another example is a maize plant — wildtype with respect to MAIL and CENH3 — used in a wide cross to pollinate a wheat plant. Haploid inducer plants are not limited to maize, which is recited here only for illustration.
[0047] The term “wide cross” as used herein refers to the process of undertaking a cross where one parent is from outside the immediate gene pool of the other. A wide cross may refer to the crossing between two different species or genera and may be used to move genes and to create new crop species.
[0048] The term “plant” as used herein can refer to a whole plant or any part or component of a plant at any stage of development and includes reference to a cell or tissue culture derived from a plant. Thus, for example, “plant” can refer to components or organs, e.g., leaves, stems, roots, plant tissues, seeds, and / or plant cells.
[0049] The term “plant cell” as used herein refers to a structural and physiological unit of a plant, comprising a protoplast and a cell wall. The plant cell may be in form of an isolated single cell or a cultured cell or as a part of higher organized unit such as, for example, plant tissue, a plant organ, or a whole plant. The plant cell may be derived from or part of an angiosperm or gymnosperm. The plant cell may be a monocotyledonous plant cell (e.g., a maize cell, a rice cell, a sorghum cell, a sugarcane cell, a barley cell, a wheat cell, an oat cell, a turf grass cell, or an ornamental grass cell) or a dicotyledonous plant cell (e.g., a tobacco cell, a pepper cell, an eggplant cell, a sunflower cell, a crucifer cell, a flax cell, a potato cell, a cotton cell, a soybean cell, a sugar bee cell, or an oilseed rape cell).
[0050] The term “plant cell culture” as used herein refers to cultures of plant units such as, for example, protoplasts, cell culture cells, cells in plant tissues, pollen, pollen tubes, ovules, embryo sacs, zygotes, and embryos at various stages of development.
[0051] A “plant organ” is a distinct and visibly structured and differentiated part of a plant such as a root, stem, leaf, flower bud, or embryo.
[0052] The term “plant tissue” as used herein refers to a group of plant cells organized into a structural and functional unit. Any tissue of a plant in planta or in culture is included. This term includes, but is not limited to, whole plants, plant organs, plant seeds, tissue culture, and any group of plant cells organized into structural and / or functional units. The use of this term in conjunction with, or in the absence of, any specific type of plant tissue as listed above or otherwise embraced by this definition is not intended to be exclusive of any other type of plant tissue.
[0053] The term “plant part” as used herein refers to a part of a plant, including single cells and cell tissues, such as plant cells that are intact in plants, cell clumps, and tissue cultures from which plants can be regenerated. Examples of plant parts include, but are not limited to, single cells and tissues from pollen, ovules, zygotes, leaves, embryos, roots, root tips, anthers, flowers, flower parts, fruits, stems, shoots, cuttings, and seeds; as well as pollen, ovules, egg cells, zygotes, leaves, embryos, roots, root tips, anthers, flowers, flower parts, fruits, stems, shoots, cuttings, scions, rootstocks, seeds, protoplasts, calli, and the like.
[0054] The terms “variety” or “cultivar” mean a group of similar plants that by structural or genetic features and / or performance can be distinguished from other varieties within the same species.
[0055] As used herein, the terms “progeny,” “progeny plant,” and “offspring” refer to a plant generated from vegetative or sexual reproduction from one or more parent plants. The term “progeny” can refer to any descent of a particular cross or parent plant. Typically, progeny plants result from the breeding of two individuals, although some species (particularly some plants and hermaphroditic animals) can be selfed or cloned (i.e., the same plant acts as the donor of both male and female gametes). The descendant(s) can be, for example, of the Fl, the F2, or any subsequent generation. In some embodiments, “progeny” plants result from a HI- Edit method or a Hot-Edit method.
[0056] As used herein the term “event” refers to a recombinant plant produced by genetic modification of a plant cell or tissue and regeneration of said plant cell or tissue. Examples of events include a gene editing event, e.g., gene editing by a site-directed nuclease (SDN); and a transformation event with an expression cassette that includes a gene of interest. The term “event” also refers to the original modified plant (by gene editing, transformation, or other methods) and / or progeny of the original modified plant comprising the genetic modification. The term “event” also refers to progeny produced by a sexual outcross between the modified plant and another plant line, wherein the progeny comprises the genetic modification. Even after repeated backcrossing to a recurrent parent, the modified DNA from the transformed parent, or the inserted DNA with the flanking DNA from the transformed parent, is present in the progeny of the cross at the same chromosomal location. In the case of a transgenic event, the term “event” also refers to DNA from the original transformant comprising (1) the inserted DNA and (2) the flanking genomic sequence immediately adjacent to the inserted DNA that would be expected to be transferred to a progeny that receives inserted DNA. The inserted DNA includes the transgene of interest as the result of a sexual cross of one parental line that includes the inserted DNA (e.g., the original transformant and progeny resulting from selfing) and a parental line that does not contain the inserted DNA. Normally, transformation of plant tissue produces multiple events, each of which represents insertion of a DNA construct into a different location in the genome of a plant cell. Based on the expression of the transgene, absence of deleterious effects, or other desirable characteristics, a particular event is selected.
[0057] The term “trait introgression” or “introgression” refers to the incorporation of a desired trait into an existing plant line (also referred to as an “elite” or “inbred” line; these lines have stable genetics and are practically homozygous across their chromosomes). Introgression involves the transfer of genetic material from a donor line to an existing, i.e., recipient, elite line, such that the benefits of that trait are incorporated into existing elite germplasm. Plant progeny are crossed back into their inbred parent line for many generations and selected for desirable traits.
[0058] A plant referred to herein as “haploid” has a reduced number of chromosomes (n) in the haploid plant, and its chromosome set is equal to that of the gamete. In a haploid organism, only half of the normal number of chromosomes are present. Thus, haploids of diploid (2n) organisms (e.g., maize) exhibit monoploidy (In); haploids of tetrapioid (4n) organisms (e.g., ryegrasses) exhibit diploidy (2n); haploids of hexapioid (6n) organisms (e.g., wheat) exhibit triploidy (3n); etc.
[0059] A plant referred to here as a “doubled haploid” is generated by doubling the haploid set of chromosomes. A plant or seed that is obtained from a doubled haploid plant that is selfed to any number of generations may still be identified as a doubled haploid plant. A doubled haploid plant is considered a homozygous plant. A plant is considered to be doubled haploid if it is fertile, even if the entire vegetative part of the plant does not consist of the cells with the doubled set of chromosomes; that is, a plant will be considered doubled haploid if it contains viable gametes, even if it is chimeric in vegetative tissues.
[0060] The term “quantitative trait locus” or “QTL” refers to a region of DNA that is associated with a particular phenotypic trait, i.e., a phenotype that can be measured numerically and varies in degree and which can be attributed to polygenic effects, i.e., the product of two or more genes and their environment. Typically, QTLs underlie continuous traits (those traits which vary continuously, e.g., haploid induction rate) as opposed to qualitative (i.e., discrete) traits.
[0061] The term “allele(s)” means any of one or more alternative forms of a gene, all of which alleles relate to at least one trait or characteristic. In a diploid cell, the two alleles of a given gene occupy corresponding loci on a pair of homologous chromosomes. In some instances (e.g., for QTLs) it is more accurate to refer to “haplotype” (i.e., an allele of a chromosomal segment) instead of “allele.” However, in those instances, the term “allele” should be understood to comprise the term “haplotype.” If two individuals (e.g., two plants) possess the same allele at a particular locus and the alleles were inherited from one common ancestor (i.e., the alleles are copies of the same parental allele), the alleles are termed “identical by descent.” The alternative is that the alleles are “identical by state,” i.e., the alleles appear to be the same but are derived from two different copies of the allele. Identity by descent information is useful for linkage studies; both identity by descent and identity by state information can be used in association studies, although identity by descent information can be particularly useful.
[0062] The term “haplotype” can refer to the set of alleles an individual inherited from one parent. A diploid individual thus has two haplotypes. The term “haplotype” can be used in a more limited sense to refer to physically linked and / or unlinked genetic markers (e.g., sequence polymorphisms) associated with a phenotypic trait. The phrase “haplotype block” (sometimes also referred to in the literature simply as a haplotype) refers to a group of two or more genetic markers that are physically linked on a single chromosome (or a portion thereof). Typically,each block has a few common haplotypes, and a subset of the genetic markers (i.e., a “haplotype tag”) can be chosen that uniquely identifies each of these haplotypes.
[0063] The term “genotype” and variants thereof refers to the genetic composition of an organism, including, for example, whether a diploid organism is heterozygous (i.e., has two different alleles for a given gene or QTL) or homozygous (i.e., has the same allele for a given gene or QTL) for one or more genes or loci (e.g., a SNP, a haplotype, a gene mutation, an insertion, or a deletion).
[0064] “Phenotype” is understood within the scope of the present disclosure to refer to a distinguishable characteristic(s) of a genetically controlled trait. The phrase “phenotypic trait” refers to the appearance or other detectable characteristic of an individual, resulting from the interaction of its genome with the environment.
[0065] As used herein, the term “marker” can be used to refer to a genetic marker, as defined above, or an encoded product thereof (e.g., a protein) used as a point of reference when identifying the presence / absence of a locus. A marker can be derived from genomic nucleotide sequences or from expressed nucleotide sequences (e.g., from an RNA, a cDNA, etc.). The term also refers to nucleotide sequences complementary to or flanking the marker sequences, such as nucleotide sequences used as probes and / or primers capable of amplifying the marker sequence. Nucleotide sequences are “complementary” when they specifically hybridize in solution (e.g., according to Watson-Crick base pairing rules). The term “marker” also refers to the genetic markers that indicate a trait by the absence of the nucleotide sequences complementary to or flanking the marker sequences, such as nucleotide sequences used as probes and / or primers capable of amplifying the marker sequence.
[0066] The term “marker-based selection” is understood within the scope of the present disclosure to refer to the use of genetic markers to detect one or more nucleic acids from the plant, where the nucleic acid is associated with a desired trait to identify plants that carry genes for desirable (or undesirable) traits, so that those plants can be used (or avoided) for any purpose, e.g., in a transformation program or in a selective breeding program.
[0067] The term “tester” or “tester plant” is understood within the scope of the present disclosure to refer to a plant used to characterize genetically a trait in a plant to be tested. Typically, the plant to be tested is crossed with a tester plant and the segregation ratio of the trait in the progeny of the cross is scored. The term “tester” can also refer to a line or individual with a standard genotype, known characteristics, and established performance. A “testerparent” is an individual from a tester line that is used as a parent in a sexual cross. Typically, the tester parent is unrelated to and genetically different from the individual to which it is crossed. A tester is typically used to generate Fl progeny when crossed to individuals or inbred lines for phenotypic evaluation.
[0068] The term “seed set” refers to a measure of the portion of a maize ear that produces embryos (i.e., kernels or seeds). Seed set may be expressed qualitatively (e.g., low, good, or high) or quantitatively. In a quantitative measurement, the measurement may be given as either a percentage or as a number of seeds per ear. The term generally refers to the percentage or number of normal kernels (i.e., non-aborted, endosperm-viable kernels). For normal maize lines (i.e., not haploid inducer lines), a seed set above 80% (or above 300 kernels per ear) is considered a good seed set. For haploid inducer lines, seed set tends to be lower; a seed set above 50% (e.g., above 60%, above 70%, or above 80%) or above 180 kernels per ear (e.g., above 200, above 220, above 260, or above 280) is generally considered a high seed set.
[0069] A “gene” is a defined region that is located within a genome that can include, in addition to coding nucleic acid sequence, other sequences such as regulatory sequences responsible for the control of the expression (i.e., transcription) and translation of the coding portion (wherein a protein is produced by the gene). Genes can include both coding and noncoding regions (e.g., introns, regulatory elements, promoters, enhancers, termination sequences, and 5' and 3' untranslated regions). A gene typically expresses mRNA, functional RNA, or a specific protein, including regulatory sequences. Genes may or may not be capable of being used to produce a functional protein. In some embodiments, a gene refers to only the coding region. The term “native gene” refers to a gene as found in nature.
[0070] A gene may be “isolated” by which is meant a nucleic acid molecule that is substantially or essentially free from components normally found in association with the nucleic acid molecule in its natural state. Such components include other cellular material, culture medium from recombinant production, and / or various chemicals used in chemically synthesizing the nucleic acid molecule. “Isolated” does not necessarily mean that the preparation is technically pure (homogeneous), but it is sufficiently pure to provide the nucleic acid in a form in which it can be used for the intended purpose.
[0071] Thus, an “isolated” nucleic acid molecule is a nucleic acid molecule or nucleotide sequence that is not immediately contiguous with nucleotide sequences with which it is immediately contiguous (one on the 5' end and one on the 3' end) in the naturally occurringgenome of the organism from which it is derived. Accordingly, in one embodiment, an isolated nucleic acid includes some or all of the 5’ non-coding (e.g., promoter) sequences that are immediately contiguous to a coding sequence. The term therefore includes, for example, a recombinant nucleic acid that is incorporated into a vector, into an autonomously replicating plasmid or virus, or into the genomic DNA of a prokaryote or eukaryote, or which exists as a separate molecule (e.g., a cDNA or a genomic DNA fragment produced by PCR or restriction endonuclease treatment), independent of other sequences. It also includes a recombinant nucleic acid that is part of a hybrid nucleic acid molecule encoding an additional RNA or polypeptide sequence. An “isolated” nucleic acid molecule can also include a polynucleotide derived from and inserted into the same natural, original cell type but which is present in a nonnatural state, e.g., present in a different copy number and / or under the control of different regulatory sequences than that found in the native state of the nucleic acid molecule.
[0072] The terms “nucleic acid” and “polynucleotide” are used interchangeably and as used herein refer to deoxyribonucleic acids (DNA) or ribonucleic acids (RNA) and polymers thereof in either single-stranded or double-stranded form, as well as to both sense and anti-sense strands of RNA, cDNA, genomic DNA, mitochondrial DNA, and synthetic forms and mixed polymers of the above. In higher plants, DNA is the genetic material while RNA is involved in the transfer of information contained within DNA into proteins. A “genome” is the entire body of genetic material contained in each cell of an organism. It is understood that when an RNA is described, its corresponding cDNA is also described, wherein uridine is represented as thymidine. In particular embodiments, a nucleotide refers to a ribonucleotide, deoxynucleotide, or a modified form of either type of nucleotide, or combinations thereof. In addition, a polynucleotide disclosed herein may include either or both naturally occurring and modified nucleotides linked together by naturally occurring and / or non-naturally occurring nucleotide linkages. The nucleic acid molecules may be modified chemically or biochemically or may contain non-natural or derivatized nucleotide bases, as will be readily appreciated by those of skill in the art. Such modifications include, for example, labels, methylation, substitution of one or more of the naturally occurring nucleotides with an analogue, internucleotide modifications such as uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoramidates, carbamates, and the like), charged linkages (e.g., phosphorothioates, phosphorodithioates, and the like), pendent moieties (e.g., polypeptides), intercalators (e.g., acridine, psoralen, and the like), chelators, alkylators, and modified linkages (e.g., alpha anomeric nucleic acids and the like). The above term is also intended to include any topologicalconformation, including single-stranded, double-stranded, partially duplexed, triplex, hair- pinned, circular, and padlocked conformations. A reference to a nucleic acid sequence encompasses its complement unless otherwise specified. Thus, a reference to a nucleic acid molecule having a particular sequence should be understood to encompass its complementary strand, with its complementary sequence. Nucleotide sequences are “complementary” when they specifically hybridize in solution (e.g., according to Watson-Crick base pairing rules). The term also includes codon-optimized nucleic acids that encode the same polypeptide sequence. It is also understood that nucleic acids can be unpurified, purified, or attached, for example, to a synthetic material such as a bead or column matrix.
[0073] The term “promoter” as used herein refers to a nucleotide sequence, usually upstream (5’) to its coding sequence, which controls the expression of the coding sequence by providing the recognition for RNA polymerase and other factors required for proper transcription. “Promoter regulatory sequences” consist of proximal and more distal upstream elements. Promoter regulatory sequences influence the transcription, RNA processing or stability, or translation of the associated coding sequence. Regulatory sequences include enhancers, promoters, untranslated leader sequences, introns, and polyadenylation signal sequences. They include natural and synthetic sequences as well as sequences that may be a combination of synthetic and natural sequences. An “enhancer” is a DNA sequence that can stimulate promoter activity and may be an innate element of the promoter or a heterologous element inserted to enhance the level or tissue specificity of a promoter. It is capable of operating in both orientations (normal or flipped) and is capable of functioning even when moved either upstream or downstream from the promoter. The meaning of the term “promoter” includes “promoter regulatory sequences.”
[0074] As used herein, the term “primer” refers to an oligonucleotide that is capable of annealing to a nucleic acid target (in some embodiments, annealing specifically to a nucleic acid target) allowing a DNA polymerase and / or reverse transcriptase to attach thereto, thereby serving as a point of initiation of DNA synthesis when placed under conditions in which synthesis of a primer extension product is induced (e.g., in the presence of nucleotides and an agent for polymerization such as DNA polymerase and at a suitable temperature and pH). In some embodiments, one or more pluralities of primers are employed to amplify plant nucleic acids (e.g., using polymerase chain reaction or PCR).
[0075] As used herein, the term “reference sequence” in the context of a nucleic acid sequence refers to a defined nucleotide sequence used as a basis for nucleotide sequence comparison.
[0076] The term “corresponding to” in the context of nucleic acid sequences as used in the present disclosure refers to certain positions or certain regions of a nucleotide sequence of interest that align with these positions or regions of a reference sequence when the two sequences are optimally aligned but that are not necessarily in these exact numerical positions of the two sequences. While optimal alignment and scoring can be accomplished manually, the process is facilitated by a computer-implemented alignment algorithm. Readily available sequence comparison and multiple sequence alignment algorithms are, respectively, the Basic Local Alignment Search Tool (BLAST) and ClustalW / ClustalW2 / Clustal Omega programs available on the Internet (e.g., the website of the EMBL-EBI). Other suitable programs include, but are not limited to, GAP, BestFit, Plot Similarity, and FASTA, which are part of the Accelrys GCG Package available from Accelrys, Inc. of San Diego, Calif., United States of America. See also Smith & Waterman, 1981; Needleman & Wunsch, 1970; Pearson & Lipman, 1988; Ausubel et al., 1988; and Sambrook & Russell, 2001.
[0077] One example of an algorithm that is suitable for determining percent sequence identity and sequence similarity is the BLAST algorithm, which is described in Altschul et al., 1990. In some embodiments, a percentage of sequence identity refers to sequence identity over the full length of a nucleic acid or polypeptide sequence.
[0078] Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions in sequences that encode proteins), alleles, SNPs, and complementary sequences as well as the sequence explicitly indicated.
[0079] The terms “polypeptide,” “peptide,” and “protein” are used interchangeably herein to refer to a polymer of amino acid residues. As used herein, the terms encompass amino acid chains of any length, including full-length proteins, wherein the amino acid residues are linked by covalent peptide bonds.
[0080] The terms “identity” or “substantial identity,” as used in the context of a polynucleotide or polypeptide sequence described herein, refers to a sequence that has at least 60% sequence identity to a reference sequence. Alternatively, percent identity can be any integer from 60% to 100%. Exemplary embodiments include at least: 60%, 65%, 70%, 75%,80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, or 99%, as compared to a reference sequence using the programs described herein; preferably BLAST using standard parameters, as described below. One of skill will recognize that these values can be appropriately adjusted to determine corresponding identity of proteins encoded by two nucleotide sequences by taking into account codon degeneracy, amino acid similarity, reading frame positioning, and the like.
[0081] For sequence comparison, typically one sequence acts as a reference sequence to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters can be used, or alternative parameters can be designated. The sequence comparison algorithm then calculates the percent sequence identities for the test sequences relative to the reference sequence, based on the program parameters.
[0082] A “comparison window,” as used herein, includes reference to a segment of any one of the number of contiguous positions selected from the group consisting of from 20 to 600, usually about 50 to about 200, more usually about 100 to about 150 in which a sequence may be compared to a reference sequence of the same number of contiguous positions after the two sequences are optimally aligned. Methods of alignment of sequences for comparison are well- known in the art. Optimal alignment of sequences for comparison may be conducted by the local homology algorithm of Smith and Waterman Add. APL. Math. 2:482 (1981), by the homology alignment algorithm of Needleman and Wunsch J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson and Lipman Proc. Natl. Acad. Sci. (U.S.A.) 85: 2444 (1988), by computerized implementations of these algorithms (e.g., BLAST), or by manual alignment and visual inspection.
[0083] Unless otherwise stated, identity and similarity will be calculated by the Needleman- Wunsch global alignment and scoring algorithms (Needleman and Wunsch (1970) J. Mol. Biol. 48(3):443-453) as implemented by the “needle” program, distributed as part of the EMBOSS software package (Rice, P., Longden, I., and Bleasby, A., EMBOSS: The European Molecular Biology Open Software Suite, 2000, Trends in Genetics 16, (6) pp276-277, versions 6.3.1 available from EMBnet at embnet.org / resource / emboss and emboss.sourceforge.net, among other sources) using default gap penalties and scoring matrices (EBLOSUM62 for protein and EDNAFULL for DNA). Equivalent programs may also be used. The term “equivalentprogram” refers to any sequence comparison program that, for any two sequences in question, generates an alignment having identical nucleotide residue matches and an identical percent sequence identity when compared to the corresponding alignment generated by needle from EMBOSS version 6.3.1.
[0084] Additional mathematical algorithms are known in the art and can be utilized for the comparison of two sequences. See, for example, the algorithm of Karlin and Altschul (1990) Proc. Natl. Acad. Sci. USA 87:2264, modified as in Karlin and Altschul (1993) Proc. Natl. Acad. Sci. USA 90:5873-5877. Such an algorithm is incorporated into the BLAST programs of Altschul et al. (1990) J. Mol. Biol. 215:403. BLAST nucleotide searches can be performed with the BLASTN program (nucleotide query searched against nucleotide sequences) to obtain nucleotide sequences homologous to nucleic acid molecules of the invention or with the BLASTX program (translated nucleotide query searched against protein sequences) to obtain protein sequences homologous to nucleic acid molecules of the invention. BLAST protein searches can be performed with the BLASTP program (protein query searched against protein sequences) to obtain amino acid sequences homologous to protein molecules of the invention or with the TBLASTN program (protein query searched against translated nucleotide sequences) to obtain nucleotide sequences homologous to protein molecules of the invention. To obtain gapped alignments for comparison purposes, Gapped BLAST (in BLAST 2.0) can be utilized as described in Altschul et al. (1997) Nucleic Acids Res. 25:3389. Alternatively, PSI-Blast can be used to perform an iterated search that detects distant relationships between molecules. See Altschul et al. (1997) supra. When utilizing BLAST, Gapped BLAST, and PSI- Blast programs, the default parameters of the respective programs (e.g., BLASTX and BLASTN) can be used. Alignment may also be performed manually by inspection.II. Introduction
[0085] In the present disclosure, “haploid-induction editing” of a genome (also referred to as “HI-Edi ’), employs the haploid-induction phenomenon to deliver gene editing machinery to the genome of a recipient plant cell. In such a procedure, a first plant is crossed with a second plant to obtain haploid progeny in which the chromosomes of the haploid inducer line are eliminated, and the haploid chromosomes of the recipient plant have the desired edit. In the methods provided in this disclosure, the inventors determined that applying a heat treatment during a HI-Edit method results in an increased haploid induction rate and / or an increased haploid editing rate in the haploid progeny. The heat treatment can be applied to an egg-donorplant pollinated with pollen from a pollen-donor plant or to pollen from a pollen-donor plant that is then used to pollinate an egg-donor plant. One of the egg-donor plant or the pollen-donor plant expresses gene editing machinery (i.e., a DNA modification enzyme and, optionally, a guide nucleic acid, expressed via transgene expression), while the other comprises plant genomic DNA that is to be edited. For example, in some instances, the egg-donor plant expresses the gene editing machinery, and the pollen-donor plant comprises plant genomic DNA that is to be edited. In other instances, the pollen-donor plant expresses the gene editing machinery, and the egg-donor plant comprises plant genomic DNA that is to be edited. The haploid progeny have the genome of the donor plant that comprised the plant genomic DNA that was to be edited, and the genome of the haploid progeny has been modified by the gene editing machinery delivered by the other donor plant. An increase in the haploid induction rate (HIR) and / or the haploid editing rate (HER) translates into a commercially meaningful increase in developing new inbred lines.
[0086] There is a relatively short window for HI-Edit to occur before the gene editing machinery transgene is lost. In some embodiments, the gene editing machinery transgene is lost some time between fertilization and genome elimination. It is generally understood that in maize it takes about 10 to 20 hours after pollination for the pollen tube to grow down through the style (silk) and transmit the sperm cells to the ovule for double fertilization to take place. After double fertilization, the embryo begins to develop. In haploid inducers, some of the embryos will exhibit genome elimination of the inducer genome sometime after fertilization, and this proceeds either immediately following fertilization or within the first couple of cell divisions of the embryo, which may take anywhere from several hours to a few days. After genome elimination, the non-inducer genome remains, and the embryo is haploid (or, if a partial inducer genome is present, the embryo may be aneuploid - aneuploid embryos, when they are found, are typically discarded).
[0087] Studies show heat stress can increase haploid induction frequency. See, e.g., Jin et al., (2023) and Ahmadli et al. (2023). In some embodiments, heat may increase expression of certain genes, including, the gene editing machinery transgene. A heat treatment may increase the expression of the gene editing machinery (i.e. a DNA modification enzyme and, optionally, a guide nucleic acid) from the transfer DNA (T-DNA) in the HI-Edit pollen or zygote, by opening the chromatin around it, based on information of heat stress responses in plants. See, e.g., Huang, et al. (2023); Liang (2021); Das & Mathur (2023); Perrella, et al. (2022). This effect may be particularly useful for HI-Edit because sperm cell chromatin is remarkablycompact during late pollen development, during pollen tube growth during style transit, and during fertilization. This compaction may block efficient expression of the gene editing machinery, or it may block efficient editing of the target site. Therefore, without wishing to be bound by theory, increased chromatin opening or relaxation may lead to either increased expression of the gene editing machinery transgene(s) or increased target site accessibility so that the gene editing machinery can make the edits to the genome more effectively. Either of these phenomena would lead to higher haploid editing rates. Methods that combine heat application with EH-Edit are termed “Hot-Edit.”III. Haploid Induction Editing With Heat Treatment (Hot-Edit)
[0088] Provided herein are methods of editing plant genomic DNA. In general, the methods comprise providing a first parental plant as an egg-donor plant and a second parental plant as a pollen-donor plant such that at least one progeny is produced from pollination of the two parental plants comprises a genome that has been edited when compared to one of the two parental plants in the cross. In some embodiments, a heat treatment step is applied before, during, or after the crossing event between the first and second plants. One or more progeny plants may be produced with the methods disclosed herein. A progeny plant that is produced using the methods disclosed herein is an edited haploid.
[0089] In the methods disclosed herein, one parental plant, i.e., the target plant, comprises the plant genomic DNA that is to be edited and the other comprises components capable of performing gene editing. In some embodiments, the components capable of performing gene editing comprise a DNA modification enzyme, and optionally, a guide nucleic acid. In some embodiments, the first and second parental plants are different species. In some embodiments, both the first and second parental plants are the same species.
[0090] In some embodiments, the first parental plant is an egg-donor plant, and the second parental plant is a pollen-donor plant. In some embodiments, the egg-donor plant comprises the plant genomic DNA that is to be edited. In some embodiments, the pollen-donor plant expresses the DNA modification enzyme and the optional guide nucleic acid. In some embodiments, the method comprises producing at least one edited haploid progeny, wherein the haploid progeny comprises the genome of the egg-donor plant and does not comprise the genome of the pollen-donor plant. Further, in some embodiments, the genome of the haploid progeny has been modified by the DNA modification enzyme and the optional guide nucleic acid delivered by the pollen-donor plant.
[0091] In some embodiments, the first parental plant is a pollen-donor plant, and the second parental plant is an egg-donor plant. In some embodiments, the pollen-donor plant comprises the plant genomic DNA that is to be edited. In some embodiments, the egg-donor plant expresses the DNA modification enzyme and the optional guide nucleic acid.
[0092] In some embodiments, the method comprises producing at least one edited haploid progeny, wherein the haploid progeny comprises the genome of the pollen-donor plant and does not comprise the genome of the egg-donor plant. Further, in some embodiments, the genome of the haploid progeny has been modified by the DNA modification enzyme and the optional guide nucleic acid delivered by the egg-donor plant.
[0093] In some embodiments, the egg-donor plant is pollinated with the pollen-donor plant. In some embodiments, after the pollination step, a heat treatment is applied to the pollinated egg-donor plant.
[0094] In some embodiments, a heat treatment is applied to the pollen-donor plant. In some embodiments, a heat treatment is applied to pollen from the pollen-donor plant. In some embodiments, the heat treatment is applied before the pollination step. In some embodiments, the egg-donor plant is pollinated with the heat-treated pollen of the pollen-donor plant.
[0095] In some embodiments, the first parental plant is a pollen-donor plant, and the second parental plant is an egg-donor plant. In some embodiments, the egg-donor plant comprises the plant genomic DNA that is to be edited. In some embodiments, the pollen-donor plant expresses the DNA modification enzyme and the optional guide nucleic acid.A. Haploid Induction via HI-Edit
[0096] Haploid induction (HI) can be used to introduce genome edits in nascent seeds of diverse monocot and dicot species by a method termed “HI-Edit” for “haploid-induction editing.” HI-Edit employs a haploid-inducer plant modified to express gene editing machinery to deliver the editing machinery to the genome to be edited of a recipient plant. In such a procedure, the first plant is crossed with the second plant to obtain haploid progeny in which the chromosomes of the haploid inducer line are eliminated, and the haploid chromosomes of the recipient plant have the desired edit. HI-Edit methods are detailed in PCT publication WO 2018 / 102816 and are also described in Kelliher, T et al. (2019) One-step genome editing of elite crop germplasm during haploid induction. Nature Biotech 37: 287-292.
[0097] Commonly, during haploid induction, both parent lines used in the induction cross are both diploids, so their gametes (egg cells and sperm cells) are haploids. Haploid induction is frequently a medium to low penetrance trait of the inducer line, so the resulting progeny, depending on the species or situation, may be either diploid (if no genome loss takes place) or haploids (if genome loss does indeed take place). If the parent line that is crossed to the haploid inducer is not diploid, but rather a tetrapioid, hexapioid, or other plant of higher ploidy, the “haploid” progeny produced will have a gametic chromosome number, e.g., diploids if the parent is tetrapioid, or triploids if the parent is hexapioid. Therefore, as used herein, “haploids” possess half the number of chromosomes of either parent.
[0098] Haploid induction can occur during self-pollination or intercrossing of two lines within the same species, or it can occur during wide crosses, where it can be viewed as a hybridization barrier, preventing the formation of interspecific hybrids. In maize, a commonly employed method of inducing haploids is through the use of HI alleles at several genomic loci that can promote efficient haploid induction. In wheat, rice, barley, brassica, and other crops, a common method of inducing haploids is by wide cross to maize pollen. For example, one could use com pollen on wheat, millet pollen on wheat, barley pollen on other barley species, or any other wide crossing method. In those cases of gynogenetic haploid induction, it would be preferable for the paternal line to contain the editing machinery, because it is the paternal (pollen-derived) DNA that is eliminated in the haploid induction process. In wheat, wide cross to maize pollen results in haploid induction regardless of parent genotype or lineage; it works with almost any wheat crossed by almost any maize pollen.
[0099] In maize, the most commonly employed method of inducing haploids is through the use of an intraspecific haploid inducer male line, which is primarily triggered by rearrangements of, mutations in, and / or recombinations, insertion, or deletions within a quantitative trait locus ("QTL") of chromosome 1, specifically the MATRILINEAL (MATL) gene, also known as NOT LIKE DADI (NLD1) and PHOSPHOLIPASE Al (PLA1) (with the notable exception of the igl type haploid induction, which is a result of a mutation in the INDETERMINATE GAMETOPHYTE1 gene on chromosome 3). HI maize lines contain the QTL on Chromosome 1, which is responsible for at least 66% of the variation in haploid induction. The QTL causes haploid induction at different rates when it is introgressed into various backgrounds. All maize haploid inducer lines used in the seed industry are derivatives of the founding HI line, known as Stock6, and all have the haploid inducer chromosome 1 QTL mutation. In wheat, the most common method of inducting haploids is by wide cross to maizepollen. Regardless of parent genotype or lineage, this works with almost any wheat crossed by almost any maize pollen.
[0100] In some embodiments, the maize plants used in the methods provided in this disclosure comprise a HI allele at the MATRILINEAL (MATL) gene, which is the patatin-like phospholipase A2a gene (PLPA2a, maize B73 gene ID GRMZM2G471240 on chromosome 1; also known as Zm00001d029412 [B73_v5]; also known as NOT LIKE DAD (NLD) and PHOSPHOLIPASE Al (PLA1; ZmPLAl)). In some embodiments, the HI allele is a loss-of- function mutation in MATL (generally referred to as matl). In some embodiments, the variant allele comprises a four base pair insertion frameshift mutation in the MATL coding sequence. In some embodiments, the four base pair insertion corresponds to the four nucleotides at positions 1146-1149 of SEQ ID NO:7. In some embodiments, the variant allele comprises a different mutation (i.e., other than the four base pair insertion mutation) or different mutations resulting in a loss-of-function in the protein product encoded by MATL. Any assay that is able to identify a loss-of-function mutation in MATL may be used to identify the plants described herein. Exemplary methods of identifying plants having a loss-of-function mutation in MATL are described in PCT / US2022 / 022271, filed May 22, 2022, incorporated herein by reference in its entirety.
[0101] In some embodiments, the maize plants used in the methods provided in this disclosure comprise a HI allele at at least one quantitative trait locus (QTL) allele associated with increased haploid induction (HI-QTL). In some embodiments, the maize plants are at least heterozygous (e.g., heterozygous or homozygous) for a HI allele at at least one HI-QTL. In some embodiments, the maize plants are homozygous for a HI allele at at least one HI- QTL. In some embodiments, maize plants that are homozygous for a HI allele at a HI-QTL display more efficient haploid induction relative to maize plants that are heterozygous for the HI allele at the HI-QTL. In some embodiments, the maize plants comprise a HI allele at the qhir8 HI-QTL on chromosome 9 as described, for example, in PCT / US2022 / 022271.
[0102] In cases of androgenic haploid induction, the editing machinery would be optimally present in the maternal parent because the maternal chromosomes are eliminated in the haploid induction process. In some embodiments, haploid plants can be generated through seeds by manipulating a single centromere protein, i.e., the centromere-specific histone CENH3. See Maruthachalam & Chan 2010; Wang et al. 2021. In some embodiments, the plant comprises a mutation in the CENH3 gene, a representative cDNA sequence of which is set forth in SEQ IDNO: 8. When a haploid inducer line with cenh3 null mutants expressing altered CENH3 proteins are crossed to wild type, chromosomes from the haploid inducer line are eliminated. Genome elimination may be caused by centromere failure due to CENH3 dilution during post-meiotic cell divisions that precede gamete formation. The cenh3 approach can be used to create paternal haploids, and it can be used to create maternal haploids as well.
[0103] In some embodiments, the pollen-donor plant of the provided methods is a haploid inducer plant. For example, the pollen-donor plant can be a paternal haploid inducer plant. In some embodiments, the paternal haploid inducer plant comprises a knock-out mutation in a MAIL gene.
[0104] In some embodiments, the egg-donor plant of the provided methods is a maternal haploid inducer plant. In some embodiments, the maternal haploid inducer plant comprises a mutation in a CENH3 gene. In some embodiments, the maternal haploid inducer plant is heterozygous for the mutation in the CENH3 gene.
[0105] After HI, the haploid embryos or seeds are typically segregated from diploid and aneuploid siblings using a phenotypic or genetic marker screen and grown or cultured into haploid plants. These plants are then converted either naturally or via chemical manipulation (e.g., using an anti -microtubule agent such as colchicine, pronamide, dithipyr, or trifluralin) into double haploid (DH) plants, which then produce inbred seed.
[0106] The production of DH plants enables plant breeders to obtain inbred lines without multi -generational inbreeding, thus decreasing the time required to produce homozygous plants. DH plants provide an invaluable tool to plant breeders, particularly for generating inbred lines, quantitative trait locus (QTL) mapping, cytoplasmic conversion, trait introgression, and F2 screening for high throughput trait improvement. A great deal of time is spared as homozygous lines are essentially generated in one generation, negating the need for multi- generational single-seed descent (conventional inbreeding). In particular, because DH plants are entirely homozygous, they are very amenable to quantitative genetics studies.
[0107] In some embodiments, heat treatment increases HIR of a HI-Edit method. Because heat treatment may cause chromatin to open or relax, heat treatment may cause increases in HIR and HER in the context of different expression cassettes and their components, events, and species. In some embodiments, heat treatment relaxes and / or unpacks chromatin. HI efficiency can be represented as haploid induction rate (HIR), which is the percentage of total progeny embryos that are haploid from a cross between a haploid inducer line and another lineand comprise the edit directed by the gene editing machinery expressed by one of the donor plant parents. As discussed above, haploid induction is frequently a medium to low penetrance trait of the inducer line - as such, not all progeny from a cross are haploids.
[0108] In maize, for example, HIR refers to the number of haploid kernels divided by the total number of kernels produced following pollination. For example, the HIR can be determined in maize by harvesting test-crossed ears after pollination (e.g., at about 15 to 20 days after pollination). Embryos from the kernels can be isolated and incubated in appropriate media (referred to as embryo rescue media) suitable for maintaining embryo viability. In some embodiments, the rescue media used for HIR determination comprise 4.43 grams of Murashige and Skoog basal media with vitamins, 30 grams of sucrose, and 70 mg of salicylic acid. The embryos in the rescue media can be placed under a condition to allow the expression of a marker gene such as a color indicator gene (e.g., Rl, R1-SCM2, Rl-nj, GUS, PMI, PAT, GFP, RFP, CFP, Bl, CI, or anthocyanin pigments). In exemplary embodiments, where the marker gene is R1-SCM2 gene, the embryos are placed under 100-400 micromol light for 16-24 hours at 22-31°C until some of the embryos turn purple due to the expression of the R1-SCM2 gene. See an exemplary protocol, e.g., as described in WO 2015 / 104358. The purple (diploid) and cream-colored (haploid) embryos can be counted from each ear. The HIR (frequency of haploids) can be determined based on the number of haploids (with the R1-SCM2 color marker change) over the total embryos. The number of true haploids may be confirmed by genotyping using, for example, a TaqMan® target site assay (Applied Biosystems), next generation sequencing (NGS). NGS data can be used to confirm the haploid editing and identify the sequences of the edits and can be used to calculate the final HER.B. Heat Treatment
[0109] Heat treatment can be applied before, during, and / or after a pollination event between two plants. In some embodiments, heat treatment is applied to an egg-donor plant before pollination. In some embodiments, heat treatment is applied to a pollen-donor plant before pollination. In some embodiments, heat treatment is applied to the egg-donor plant after the pollination step, i.e., after the egg-donor plant has been pollinated. In some embodiments, heat treatment is applied to pollen from the pollen-donor plant before the pollination step.
[0110] In some embodiments, heat treatment is applied during the “HI-Edit window,” which covers the period of time during which the editing may occur prior to genome elimination. The timing of treatments ranged from a few hours to three days after pollination. In someembodiments, heat treatment is applied for no less than 10 hours, no less than 15 hours, no less than 20 hours, no less than 25 hours, or no less than 30 hours. In some embodiments, heat treatment is applied for about 0-10 hours, about 10-20 hours, or about 20-30 hours. In some embodiments, heat treatment is applied in a cycle of about 15 hours at 25-30°C and about 9 hours at 16-20°C, wherein the cycle may be repeated 0, 1, 2, or more times.[OHl] As can be appreciated by one of ordinary skill in the art, there are many ways to apply heat treatment to plants such that the heat is maintained at about a constant temperature and applied evenly to the plant part. For example, incubator temperatures or room temperatures (e.g., a greenhouse) may be set at target temperatures. In another example, a heat source is placed in close proximity to or is attached to plant parts or plant organs where haploid induction occurs. Insulation materials or barriers may be used to aid in the application of heat. In some embodiments, the heat source is placed in close proximity or is attached to the plant egg and / or plant pollen. In some embodiments, the plant is maize; thus, in some embodiments, the heat source is placed in close proximity or is attached to the maize husk, ear, tassel, pollen, kernel, and / or egg.
[0112] In some embodiments, heat packs are used. Commercial heat packs are readily available to one of skill in the art, e.g., Thermacare® Muscle Pain Therapy HeatWraps, thermacare.com / heat-wraps / muscle-pain-therapy (last visited Aug. 15, 2023). In the case of heat packs, heat packs may be exchanged from time to time (e.g., every 4 hours, every 8 hours, every 12 hours, every 16 hours, and so on) to maintain about constant heat application. In some embodiments, heat packs are attached around the plant or around a portion of the plant where pollination is occurring. For example, in some embodiments, for maize, heat packs can be attached around the maize husks. For example, in some embodiments, for maize, heat packs can be attached to maize ears.
[0113] In some embodiments, heat is applied at about 16-20°C, about 20-25°C, about 25- 30°C, and about 30-36°C. In some embodiments, temperature of about 23°C, about 24°C, about 25°C, about 26°C, about 27°C, 28°C, about 29°C, about 30°C, about 31°C, about 32°C, about 33°C, about 34°C, about 35°C, about 36°C, about 37°C, about 38°C, about 39°C, or about 40°C is maintained. In some embodiments, a daytime temperature of about 23°C, about 24°C, about 25°C, about 26°C, about 27°C, 28°C, about 29°C, about 30°C, about 31°C, about 32°C, about 33°C, about 34°C, about 35°C, about 36°C, about 37°C, about 38°C, about 39°C, or about 40°C is maintained. In some embodiments, a nighttime temperature of about 23°C,about 24°C, about 25°C, about 26°C, about 27°C, 28°C, about 29°C, about 30°C, about 31°C, about 32°C, about 33°C, about 34°C, about 35°C, about 36°C, about 37°C, about 38°C, about 39°C, or about 40°C is maintained. In some embodiments, heat is applied within about 5°C of the target temperature.
[0114] In some embodiments, the internal temperature of the plant or the portion of the plant to which the heat treatment is being applied is maintained within about 5°C of the target temperature. In some embodiments, the internal temperature of the plant or the portion of the plant to which the heat treatment is being applied is about 23°C, about 24°C, about 25°C, about 26°C, about 27°C, about 28°C, about 29°C, about 30°C, about 31°C, about 32°C, about 33°C, about 34°C, about 35°C, about 36°C, about 37°C, about 38°C, about 39°C, or about 40°C. For example, where the plant is a maize plant, the internal temperatures referenced here may be the internal temperature of the maize ears.C. Donor Plants
[0115] The methods and compositions disclosed herein are useful for editing genomic DNA of a variety of plants. In some embodiments, the egg-donor plant is a monocot or a dicot. In some embodiments, the pollen-donor plant is a monocot or a dicot. Exemplary monocots include maize, wheat, rice, barley, oats, triticale, sorghum, pearl millet, teosinte, bamboo, sugar cane, asparagus, onion, and garlic.
[0116] In some embodiments, the pollen-donor plant is maize, wheat, rice, soybean, sunflower, tomato, Arabidopsis, cucumber, barley, oat, triticale, sorghum, pearl millet, asparagus, onion, garlic, teosinte, bamboo, or sugar cane. In some embodiments, the egg-donor plant is maize, wheat, rice, soybean, sunflower, tomato, Arabidopsis, cucumber, barley, oat, triticale, sorghum, pearl millet, asparagus, onion, or garlic, teosinte, bamboo, or sugar cane. In some embodiments, the progeny produced by the provided methods is maize, wheat, rice, soybean, sunflower, tomato, Arabidopsis, cucumber, barley, oat, triticale, sorghum, pearl millet, asparagus, onion, or garlic, teosinte, bamboo, or sugar cane.
[0117] In some embodiments, the plant is maize. The maize plant may be derived from any known heterotic group. A heterotic group is a set of genetically related genotypes that show similar hybrid performance when crossed with individuals from another genetically distinct germplasm group (Melchinger, A.E. and Gumber, R.K. (1998). Overview of Heterosis and Heterotic Groups in Agronomic Crops. In Concepts and Breeding of Heterosis in Crop Plants (eds K.R. Lamkey and J.E. Staub); doi: 10.2135 / cssaspecpub25.c3). Aside from traitintrogression, a goal of plant breeding is to make genetic improvements in varietal lines and also parental lines of hybrids. An effective hybrid breeding program makes genetic improvements to parent lines in both the hybrid’s maternal parent heterotic group and the hybrid’s paternal parent heterotic group. Therefore, it is advantageous to make genetic improvements in all heterotic groups used in a breeding program. Table 1 below shows the common heterotic groups to which various germplasms belong. A maize plant from one heterotic group can be used to cross with a maize plant from any of the other heterotic groups to edit its genome and improve its traits.
[0118] In some embodiments, the paternal parent and / or the maternal parent of the methods described above belong to any of the heterotic groups in Table 1 above. In some embodiments, the paternal parent belongs to a different heterotic group than the maternal parent. In some embodiments, the maize plant comprises a Stiff Stalk germplasm, a Non-Stiff Stalk germplasm, a Non-Stiff Stalk lodent germplasm, a tropical germplasm, or a subtropical germplasm. In other embodiments, the maize plant comprises a germplasm classified into any other heterotic group known to one of skill in the art (see, e.g., L. Reid, et al., 2011, “Genetic diversity analysis of 119 Canadian maize inbred lines based on pedigree and simple sequence repeat markers,” Can. J. Plant Sci. 91 : 651-661 and M. Mikel and J. Dudley, 2006, “Evolution of North American Dent Com from Public to Proprietary Germplasm,” Crop Sci. 46: 1193-1205, each of which is incorporated herein by reference in its entirety). The maize plants of the present disclosure may also be derived from any publicly known or proprietary line. In some embodiments, the maize plant is derived from any of lines Stock 6, RWK, RWS, UH400, NP2222RS, and / or NP2222. In other embodiments, the maize plant is derived from any other line of interest.
[0119] In some embodiments, the maize plants described herein comprise at least one selectable marker to facilitate screening and selection of offspring of interest (e.g., offspring kernels that have become haploid). As used herein, the term selectable marker encompasses screening or reporter markers (e.g., color indicators that can be used to visually screen foroffspring of interest) and selection markers (e.g., antibiotic resistance genes that can be used for antibiotic-mediated enrichment of offspring of interest). In some embodiments, the plants comprise a selectable marker gene. The selectable marker gene may be, for example, a mutation of an endogenous gene or a transgene. In some embodiments, the selectable marker gene encodes a detectable protein product. In some embodiments, the plants are heterozygous for a selectable marker. In some embodiments, the plants are homozygous for a selectable marker. In some embodiments, the selectable marker gene encodes a pigment or other detectable product that will only be present in diploid embryos, facilitating selection of haploid embryos, as detailed below and in the Examples. In some embodiments, the selectable marker may include any one of GUS, PMI, PAT, GFP, RFP, CFP, Bl, CI, NPTII, HPT, ACC3, AADA, high oil content (see, e.g., Melchinger et al. 2013. Sci. Reports 3:2129 and Chaikam et al. 2019. Theor. andAppl. Genet. 132:3227-3243), R-navajo (R-nj), Rl- scutellum (R1-SCM2), and / or an anthocyanin pigment. Other selectable marker genes are known to a person skilled in the art (see, e.g., Ziemienowicz. 2001. Acta Physiologiae Plantarum 23:363-374). In some embodiments, the selectable marker comprises an antibiotic resistance gene.D. Site-directed nucleases
[0120] Gene editing in the Hot-Edit methods provided in this disclosure may be performed using various site-directed nucleases (SDNs). Generally, the desired outcomes in SDN- mediated genome editing are 1) to target SDNs to cleave DNA at a specific genomic site in a host (e.g., a plant cell), and 2) to use the host’s natural repair mechanisms to introduce specific genomic changes at the cleavage site. The changes can include small deletions, substitutions, or the addition of a number of nucleotides. Such targeted edits can result in a new and desired characteristic (e.g., enhanced nutrient uptake or decreased allergen production) and / or a reduction in an undesirable characteristic (e.g., herbicide susceptibility).
[0121] SDN applications have generally been divided into three categories: SDN-1, SDN-2, and SDN-3. SDN-1 produces a double-stranded break in a genome without the addition of foreign DNA. When such a break is repaired by the host (e.g., via non-homologous end joining; NHEJ), mutations or deletions can be introduced. If these mutations or deletions are in a gene, the gene can be silenced or knocked out. SDN-2 uses template DNA to introduce a predicted modification at the target cleavage site (e.g., via HDR), but does not result in insertion of recombinant DNA. SDN-3 also uses template DNA to introduce recombinant or exogenous DNA templates (e.g., a transgene) at the target cleavage site.
[0122] Suitable SDNs include, but are not limited to, CRISPR-associated (Cas) proteins or Cas nucleases, meganucleases (MNs), zinc finger nucleases (ZFNs), transcription activatorlike effector nucleases (TALENs), RNA-binding proteins (RBPs), CRISPR-associated RNA binding proteins, recombinases, flippases, transposases, Argonaute (Ago) proteins (e.g., prokaryotic Argonaute (pAgo), archaeal Argonaute (aAgo), eukaryotic Argonaute (eAgo), and Natronobacterium gregoryi Argonaute (NgAgo)), adenosine deaminases acting on RNA (ADAR), CRISPR-Cas-inspired RNA targeting (CIRT) system, Pumilio / fem-3 binding factor (PUF), homing endonuclease, or any functional fragment thereof, any derivative thereof, any variant thereof, and any fragment thereof. Exemplary SDNs suitable for use are described further below.
[0123] In some embodiments, the SDN is a naturally-occurring SDN. Exemplary naturally- occurring SDNs are known in the art (see for example, Makarova et al., 2017, Cell 168: 328- 328. el, and Shmakov et al., 2017, Nat Rev Microbiol 15(3): 169-182). In some embodiments, an SDN binds a DNA-targeting polynucleotide (e.g., a guide RNA) and is thereby directed to a specific sequence within a target DNA and cleaves the target DNA.
[0124] The gene-editing machinery (e.g., a DNA modifying enzyme such as an SDN, and an optional guide nucleic acid) introduced into the plants can be controlled by any promoter that can drives recombinant gene expression in that plant. In some embodiments, the promoter is a constitutive promoter. In some embodiments, the promoter is a tissue-specific promoter, e.g., a pollen-specific promoter or a sperm cell specific promoter, a zygote specific promoter, or a promoter that is highly expressed in sperm, eggs, and zygotes (e.g., prOsActinl). Suitable promoters are disclosed in U.S. Pat. No. 10,519,456 and International Appl. No. PCT / CN2023 / 110941. Exemplary promoters are shown in Table 2 below. Promoters can be used in a transgene to drive high sperm cell expression of editing machinery to boost the efficiency of simultaneous editing and doubled-haploid induction (SEDHI).1. CRISPR-Cas Systems
[0125] In some embodiments, the haploid-inducer plant expresses a CRISPR-Cas system. A CRISPR-Cas system can comprise at least one guide nucleic acid, such as a guide RNA (gRNA), complexed with a DNA modification enzyme, such as a Cas protein, for targeted regulation of gene expression and / or activity or nucleic acid editing. An RNA-guided Cas protein (e.g., a Cas nuclease such as a Cas9 nuclease or a Casl2 nuclease) can specifically bind a target polynucleotide (e.g., DNA) in a sequence-dependent manner. The Cas protein, if possessing nuclease activity, can cleave the DNA (Gasiunas, G., et al, “Cas9-crRNA ribonucleoprotein complex mediates specific DNA cleavage for adaptive immunity in bacteria,” Proc Natl Acad Sci USA (2012) 109:E2579-E2 86; Jinek, M., et al, “A programmable dual-RNA-guided DNA endonuclease in adaptive bacterial immunity,” Science (2012) 337:816-821; Sternberg, S. H., et al, “DNA interrogation by the CRISPR RNA-guided endonuclease Cas9,” Nature (2014) 507:62; Deltcheva, E., et al, “CRISPR RNA maturation by trans-encoded small RNA and host factor RNase III,” Nature (2011) 471 :602-607). DNA cleavage (e.g., double-strand breaks) can result in DNA break repair which allows for the introduction of gene modification(s) (e.g., nucleic acid editing). DNA break repair can occur via non-homologous end joining (NHEJ), microhomology-mediated end joining (MMEJ), or homology-directed repair (HDR). CRISPR-Cas systems have been widely used for programmable genome editing in a variety of organisms and model systems (Cong, L., et al, “Multiplex genome engineering using CRISPR Cas systems,” Science (2013) 339:819-823; Jiang, W., et al, “RNA-guided editing of bacterial genomes using CRISPR-Cas systems,” Nat.Biotechnol. (2013) 31 : 233-239; Sander, J. D. & Joung, J. K, “CRISPR-Cas systems for editing, regulating and targeting genomes,” Nature Biotechnol. (2014) 32:347-355).
[0126] In some embodiments, the Cas protein forms a complex with a guide nucleic acid. In some embodiments, the Cas protein comprises an RNA-binding protein (RBP) optionally complexed with a guide nucleic acid, such as a guide RNA (gRNA), which is able to form a complex with a Cas protein. In some instances, RNA-guided Cas proteins recognize DNA targets that are complementary to a portion of the gRNA known as a CRISPR RNA (crRNA) sequence. The target sequence is often referred to as a protospacer, and the part of the crRNA sequence that is complementary to the protospacer is often referred to as a spacer. In order to function (e.g., to cleave DNA), many Cas proteins also require a specific protospacer adjacent motif (PAM), an approximately 2 to 6 base pair DNA sequence immediately following the protospacer sequence.
[0127] Any suitable CRISPR-Cas system can be used. A CRISPR-Cas system can be referred to using a variety of naming systems. Exemplary naming systems are provided in Makarova, K.S. et al., “An updated evolutionary classification of CRISPR-Cas systems,” Nat Rev Microbiol (2015) 13 :722-736 and Shmakov, S. et al, “Discovery and Functional Characterization of Diverse Class 2 CRISPR-Cas Systems,” Mol Cell (2015) 60: 1-13. A CRISPR-Cas system can be a type I, a type II, a type III, a type IV, a type V, a type VI system, or any other suitable CRISPR-Cas system. A CRISPR-Cas system as used herein can be a Class 1, Class 2, or any other suitably classified CRISPR-Cas system. Class 1 or Class 2 determination can be based upon the genes encoding the effector module. Class 1 systems generally have a multi-subunit crRNA-effector complex, whereas Class 2 systems generally have a single protein, such as Cas9, Cpfl, C2cl, C2c2, C2c3 or a crRNA-effector complex. A Class 1 CRISPR-Cas system can use a complex of multiple Cas proteins to effect regulation. A Class 1 CRISPR-Cas system can comprise, for example, type I (e.g., I, IA, IB, IC, ID, IE, IF, IU), type III (e g., Ill, IIIA, IIIB, IIIC, IIID), and type IV (e.g, IV, IVA, IVB) CRISPR-Cas type. A Class 2 CRISPR-Cas system can use a single large Cas protein to effect regulation. A Class 2 CRISPR-Cas systems can comprise, for example, type II (e.g., II, IIA, IIB) and type V CRISPR-Cas type. CRISPR systems can be complementary to each other, and / or can lend functional units in trans to facilitate CRISPR locus targeting.(a) Cas Proteins
[0128] A Cas protein can be from any suitable organism. Non-limiting examples of suitable organisms include Streptococcus pyogenes, Streptococcus thermophilus, Streptococcus sp., Staphylococcus aureus, Nocardiopsis dassonvillei, Streptomyces pristinae spiralis, Streptomyces viridochromo genes, Streptomyces viridochromogenes, Streptosporangium roseum, Streptosporangium roseum, AlicyclobacHlus acidocaldarius, Bacillus pseudomycoides, Bacillus selenitireducens, Exiguobacterium sibiricum, Lactobacillus delbrueckii, Lactobacillus salivarius, Microscilla marina, Burkholderiales bacterium, Polar omonas naphthalenivorans, Polar omonas sp., Crocosphaera watsonii, Cyanothece sp., Microcystis aeruginosa, Pseudomonas aeruginosa, Synechococcus sp., Acetohalobium arabaticum, Ammonifex degensii, Caldicelulosiruptor becscii, Candidatus Desulforudis, Clostridium botulinum, Clostridium difficile, Finegoldia magna, Natranaerobius thermophilus, Pelotomaculum thermopropionicum, Acidithiobacillus caldus, Acidithiobacillus ferrooxidans, Allochromatium vino sum, Marinobacter sp., Nitrosococcus halophilus, Nitrosococcus watsoni, Pseudoalter omonas haloplanktis, Ktedonobacter racemifer, Methanohalobium evestigatum, Anabaena variabilis, Nodularia spumigena, Nostoc sp., Arthrospira maxima, Arthrospira platensis, Arthrospira sp., Lyngbya sp., Microcoleus chthonoplastes, Oscillatoria sp., Petrotoga mobilis, Thermosipho africanus, Acaryochloris marina, Leptotrichia shahii, and Francisella novicida. In some embodiments, the organism is Streptococcus pyogenes (S. pyogenes). In some embodiments, the organism is Staphylococcus aureus (S. aureus). In some embodiments, the organism is Streptococcus thermophilus (S. thermophilus).
[0129] A Cas protein can be derived from a variety of bacterial species including, but not limited to, Lachnospiraceae bacterium, Veillonella atypical, Fusobacterium nucleatum, Filifactor alocis, Solobacterium moorei, Coprococcus catus, Treponema denticola, Peptoniphilus duerdenii, Catenibacterium mitsuokai, Streptococcus mutans, Listeria innocua, Staphylococcus pseudintermedius, Acidaminococcus intestine, Olsenella uli, Oenococcus kitaharae, Bifidobacterium bifidum, Lactobacillus rhamnosus, Lactobacillus gasseri, Finegoldia magna, Mycoplasma mobile, Mycoplasma gallisepticum, Mycoplasma ovipneumoniae, Mycoplasma canis, Mycoplasma synoviae, Eubacterium rectale, Streptococcus thermophilus, Eubacterium dolichum, Lactobacillus coryniformis subsp. Torquens, Ilyobacter polytropus, Ruminococcus albus, Akkermansia muciniphila, Acidothermus cellulolyticus, Bifidobacterium longum, Bifidobacterium dentium,Corynebacterium diphtheria, Elusimicrobium minutum, Nitratifractorsalsuginis, Sphaerochaeta globus, Fibrobacter succinogenes subsp. Succinogenes, Bacteroides fragilis, Capnocytophaga ochracea, Rhodopseudomonas palustris, Prevotella micans, Prevotella ruminicola, Flavobacterium columnare, Aminomonas paucivorans, Rhodospirillum rubrum, Candidatus Puniceispirillum marinum, Verminephrobacter eiseniae, Ralstonia syzygii, Dinoroseobacter shibae, Azospirillum, Nitrobacter hamburgensis, Bradyrhizobium, Wolinellasuccinogenes, Campylobacter jejuni subsp. Jejuni, Helicobacter mustelae, Bacillus cereus, Acidovorax ebreus, Clostridium perfringens, Parvibaculum lavamentivorans, Roseburia intestinalis, Neisseria meningitidis, Pasteurella multocida subsp. Multocida, Sutterella wadsworthensis, proteobacterium, Legionella pneumophila, Parasutterella excrementihominis, Wolinella succinogenes, and Francisella novicida. In some embodiments, the organism is Lachnospiraceae bacterium.
[0130] Non-limiting examples of Cas proteins include c2cl, C2c2, c2c3, Casl, CaslB, Cas2, Cas3, Cas4, Cas5, Cas5e (CasD), Cas6, Cas6e, Cas6f, Cas7, Cas8, Cas8a, Cas8al , Cas8a2, Cas8b, Cas8c, Cas9 (Csnl or Csxl2), CaslO, CaslOd, Casl2a, Casl2b, Casl2i, Casl2j, Casl2L, Casl2e, Casl2c, Casl2d, Casl2g, Casl2h, TnpB, Casl3a, Casl3b, Casl4, CasF, CasG, CasH, Cpfl, Csyl, Csy2, Csy3, Csel (CasA), Cse2 (CasB), Cse3 (CasE), Cse4 (CasC), Cscl, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmrl , Cmr3, Cmr4, Cmr5, Cmr6, Csbl, Csb2, Csb3, Csxl7, Csxl4, CsxlO, Csxl6, CsaX, Csx3, Csxl, Csxl5, Csfl, Csf2, Csf3, Csf4, Cul966, and homologs or modified versions thereof. In some embodiments, the Cas protein is a Cas9 protein. In some embodiments, the Cas protein is a Casl2a protein. In some embodiments, the Cas protein is a nickase that generates single-strand nicks in DNA. In some embodiments, the Cas protein is a Cas9 nickase or a Casl2a nickase.
[0131] A Cas protein can comprise one or more domains. Non-limiting examples of domains include guide nucleic acid recognition and / or binding domains, nuclease domains (e.g., DNase or RNase domains, RuvC, and HNH), DNA binding domains, RNA binding domains, helicase domains, protein-protein interaction domains, and dimerization domains. A guide nucleic acid recognition and / or binding domain can interact with a guide nucleic acid. A nuclease domain can comprise catalytic activity for nucleic acid cleavage. A nuclease domain can lack catalytic activity to prevent nucleic acid cleavage. A Cas protein can be a chimeric Cas protein that is fused to other proteins or polypeptides. A Cas protein can be a chimera of various Cas proteins, for example, comprising domains from different Cas proteins. In some embodiments, the Cas protein is a Cas9 fused to a heterologous domain or a Casl2a fused to a heterologous domain.
[0132] A Cas protein used herein can be an active variant, inactive variant, or fragment of a wild-type or modified Cas protein. A Cas protein can comprise an amino acid change such as a deletion, truncation, insertion, substitution, variant, mutation, fusion, chimera, or any combination thereof relative to a wild-type version of the Cas protein. A Cas protein can be a polypeptide with at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild-type exemplary Cas protein. A Cas protein can be a polypeptide with at most about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, or 100% sequence identity and / or sequence similarity to a wild-type exemplary Cas protein. Variants or fragments can comprise at least about 5%, 10%, 20%, 30%, 40%, 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity or sequence similarity to a wild-type or modified Cas protein or a portion thereof. Variants or fragments can be targeted to a nucleic acid locus in complex with a guide nucleic acid while lacking nucleic acid cleavage activity.
[0133] In some embodiments, a modified Cas protein has decreased function relative to the unmodified form. In some embodiments, a modified Cas protein is deficient in a function of the unmodified form. For example, a nuclease deficient Cas protein retains the ability to bind DNA but lacks or has reduced nucleic acid cleavage activity. A Cas nuclease (e.g., retaining wild-type nuclease activity, having reduced nuclease activity, and / or lacking nuclease activity) can function in a CRISPR / Cas system to regulate the level and / or activity of a target gene or protein (e.g., decrease, increase, or elimination). The Cas protein can bind to a target polynucleotide and prevent transcription by physical obstruction or edit a nucleic acid sequence to yield non-functional gene products. In some embodiments, the modified Cas protein has no more than 90%, no more than 80%, no more than 70%, no more than 60%, no more than 50%, no more than 40%, no more than 30%, no more than 20%, no more than 10%, no more than 5%, or no more than 1% of the function (e.g., nuclease activity) of the wild-type Cas protein (e.g., Cas9a or Casl2a). In some embodiments, the modified Cas protein has no substantial function of the wild-type Cas protein. When a Cas protein is a modified form that has no substantial nucleic acid-cleaving activity, it can be referred to as enzymatically inactive, nuclease-inactive, and / or “dead” (abbreviated by “d”). A dead Cas protein (e.g., dCas, dCasl2a) can bind to a target polynucleotide but may not cleave the target polynucleotide. In some embodiments, a Cas9 protein provided herein is a nuclease-inactive Cas9 protein. In some embodiments, a Casl2a protein provided herein is a nuclease-inactive Casl2a protein.
[0134] In some embodiments, a modified Cas protein can be a modified Cas “base editor.” Base editing enables direct, irreversible conversion of one target DNA base into another in a programmable manner, without requiring DNA cleavage or a donor DNA molecule. For example, Komor et al. (2016, Nature, 533:420-424), teach a Cas9-cytidine deaminase fusion, where the Cas9 has also been engineered to be inactivated and not induce double-stranded DNA breaks. Additionally, Gaudelli et al. (2017, Nature, doi: 10.1038 / nature24644) teach a catalytically impaired Cas9 fused to a tRNA adenosine deaminase, which can mediate conversion of an A / T to G / C in a target DNA sequence. In some embodiments, a Casl2a protein provided herein is a modified Casl2a base editor.
[0135] A Cas protein can be modified to optimize regulation of gene expression. A Cas protein can be modified to increase or decrease nucleic acid binding affinity, nucleic acid binding specificity, and / or enzymatic activity. Cas proteins can also be modified to change any other activity or property of the protein, such as stability. For example, one or more nuclease domains of the Cas protein can be modified, deleted, or inactivated, or a Cas protein can be truncated to remove domains that are not essential for the function of the protein or to optimize (e.g., enhance or reduce) the activity of the Cas protein for regulating gene expression. In some embodiments, the Cas protein is a modified Cas protein that comprises one or more human- induced mutations. Exemplary modified Casl2a proteins are described, for example, in International Appl. No. PCT / CN2023 / 073486.
[0136] One or a plurality of the nuclease domains (e.g., RuvC, HNH) of a Cas protein can be deleted or mutated so that they are no longer functional or comprise reduced nuclease activity. For example, in a Cas protein comprising at least two nuclease domains (e.g., Cas9, Casl2a), if one of the nuclease domains is deleted or mutated, the resulting Cas protein, known as a nickase, can generate a single-strand break at a CRISPR RNA (crRNA) recognition sequence within a double- stranded DNA but not a double-strand break. Such a nickase can cleave the complementary strand or the non-complementary strand but may not cleave both. In some embodiments, double strand break targeting specificity is improved by targeting a nickase to opposite strands at two nearby loci. If a nickase cleaves the single strand at both loci, a double strand break is formed and can be repaired as described herein. If all of the nuclease domains of a Cas protein (e.g., RuvC nuclease domains in a Cas 12a protein) are deleted or mutated, the resulting Cas protein can have a reduced or no ability to cleave both strands of a double-stranded DNA. In some embodiments, a Cas9 protein provided herein is a Cas9 nickase protein. In some embodiments, a Casl2a protein provided herein is a Casl2a nickase protein.
[0137] Also provided herein are fusion proteins comprising any of the Cas proteins described above and a heterologous domain. As used throughout, a “fusion protein” is a protein comprising two different polypeptide sequences, e.g., a Cas9 or Casl2a protein sequence as described above and a heterologous polypeptide sequence, that are joined or linked to form a single polypeptide. In some embodiments, the two amino acid sequences are encoded by separate nucleic acid sequences that have been joined so that they are transcribed and translated to produce a single polypeptide. The Cas9 or Cas 12a protein and the heterologous domain can be linked in any order and orientation relative to each other. For example, the C’ terminal end of the Cas9 or Casl2a protein may be linked to the N’ terminal end or the C’ terminal end of the heterologous domain. The Cas9 or Cas 12a protein and the heterologous domain may also be separated by one or more additional fusion protein domains, as described below.
[0138] Exemplary heterologous domains include deaminase domains, transcription factor domains, nuclease domains, reverse-transcriptase domains, transposase domains, integrase domains, uracil DNA glycosylase inhibitor domains, recombinase domains, nickase domains, methyltransferase domains, methylase domains, acetylase domains, acetyltransferase domains, transcriptional activator domains, and transcriptional repressor domains. See, e.g., WO 2021 / 061507. In some embodiments, the heterologous domain is a Trex domain or a Cro domain. Exemplary Cas fusion proteins with these heterologous domains are described in International Appl. Nos. PCT / US2023 / 068974 and PCT / US2023 / 068977.
[0139] In some embodiments, the fusion proteins provided herein comprise one or more linkers. Linkers, also referred to as spacers, as used herein are flexible molecules or a flexible stretch of molecules that join or connect two portions (e.g., domains) of a fusion protein or a variant Cas9 or Casl2a protein as provided herein. In some embodiments, the linker is a polypeptide. Proteins with domains joined by polypeptide linkers are referred to as fusion proteins. In some embodiments, the linker is a non-peptide linker. Proteins with domains joined by polypeptide linkers are referred to as modified proteins. It will be understood that, where fusion proteins are discussed throughout the present disclosure, modified proteins are generally also contemplated, where feasible. Linkers may be short or long, flexible or rigid. See, e.g., WO 2021 / 061507, WO 2020 / 168102, and US 2021 / 0017506. Exemplary linkers are described, for example, in International Appl. Nos. PCT / US2023 / 068974 and PCT / US2023 / 068977.(b) Guide Nucleic Acids
[0140] In some embodiments, the Cas protein can be complexed with the at least one guide nucleic acid polynucleotide. In some embodiments, the polynucleotide can be deoxyribonucleic acid (DNA). In some cases, the DNA sequence can be single- stranded or doubled-stranded. In some embodiments, the polynucleotide is a ribonucleic acid, i.e., a guide RNA (gRNA). In some embodiments, the gRNA is expressed from a gRNA cassette.
[0141] In some embodiments, the Cas protein can be complexed with the at least one guide RNA polynucleotide. The at least one guide RNA polynucleotide can comprise a nucleic acidtargeting region that comprises a complementary sequence to a nucleic acid sequence on the targeted polynucleotide such as the targeted genomic loci or genes to confer sequence specificity of Cas protein targeting. In some embodiments, the at least one guide RNA polynucleotide can comprise two separate nucleic acid molecules, which can be referred to as a double guide nucleic acid, or a single nucleic acid molecule, which can be referred to as a single guide nucleic acid (e.g., single guide RNA or sgRNA).
[0142] In some embodiments, the guide nucleic acid is a single guide nucleic acid comprising a fused CRISPR RNA (crRNA) and a transactivating crRNA (tracrRNA). A crRNA can comprise the nucleic acid-targeting segment (e.g., spacer region) of the guide nucleic acid and a stretch of nucleotides that can form one half of a double-stranded duplex of the Cas proteinbinding segment of the guide nucleic acid. A tracrRNA can comprise a stretch of nucleotides that forms the other half of the double-stranded duplex of the Cas protein-binding segment of the gRNA. A stretch of nucleotides of a crRNA can be complementary to and hybridize with a stretch of nucleotides of a tracrRNA to form the double-stranded duplex of the Cas proteinbinding domain of the guide nucleic acid.
[0143] In some embodiments, the guide nucleic acid is a single guide nucleic acid comprising a crRNA but lacking a tracrRNA. In some embodiments, the guide nucleic acid is a double guide nucleic acid comprising non-fused crRNA and tracrRNA. An exemplary double guide nucleic acid can comprise a crRNA-like molecule and a tracrRNA-like molecule. An exemplary single guide nucleic acid can comprise a crRNA-like molecule. An exemplary single guide nucleic acid can comprise a fused crRNA-like molecule and a tracrRNA-like molecule.
[0144] Whether a Cas protein requires a crRNA molecule only or whether it requires both a crRNA molecule and a tracrRNA molecule (whether covalently linked or not) depends on the CRISPR-associated Cas protein used.
[0145] In some embodiments, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) can be between 18 to 72 nucleotides in length. The nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) can have a length of from about 12 nucleotides to about 100 nucleotides. For example, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) can have a length of from about 12 nucleotides (nt) to about 80 nt, from about 12 nt to about 50 nt, from about 12 nt to about 40 nt, from about 12 nt to about 30 nt, from about 12 nt to about 25 nt, from about 12 nt to about 20 nt, from about 12 nt to about 19 nt, from about 12 nt to about 18 nt, from about 12 nt to about 17 nt, from about 12 nt to about 16 nt, or from about 12 nt to about 15 nt. Alternatively, the DNA-targeting segment can have a length of from about 18 nt to about 20 nt, from about 18 nt to about 25 nt, from about 18 nt to about 30 nt, from about 18 nt to about 35 nt, from about 18 nt to about 40 nt, from about 18 nt to about 45 nt, from about 18 nt to about 50 nt, from about 18 nt to about 60 nt, from about 18 nt to about 70 nt, from about 18 nt to about 80 nt, from about 18 nt to about 90 nt, from about 18 nt to about 100 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about 20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, from about 20 nt to about 60 nt, from about 20 nt to about 70 nt, from about 20 nt to about 80 nt, from about 20 nt to about 90 nt, or from about 20 nt to about 100 nt. The length of the nucleic acid-targeting region (e.g., spacer region) can be at least 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30 or more nucleotides. The length of the nucleic acid-targeting region (e.g., spacer sequence) can be at most 5, 10, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 30, 35, 40, 45, or more nucleotides.
[0146] In some embodiments, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) is 20 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) is 19 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) is 18 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) is 17 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) is 16 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacerregion) is 21 nucleotides in length. In some embodiments, the nucleic acid-targeting region of a guide nucleic acid (e.g., spacer region) is 22 nucleotides in length.
[0147] The nucleotide sequence of the guide nucleic acid that is complementary to a nucleotide sequence (target sequence) of the target nucleic acid can have a length of, for example, at least about 12 nucleotides (nt), at least about 15 nt, at least about 18 nt, at least about 19 nt, at least about 20 nt, at least about 25 nt, at least about 30 nt, at least about 35 nt, or at least about 40 nt. The nucleotide sequence of the guide nucleic acid that is complementary to a nucleotide sequence (target sequence) of the target nucleic acid can have a length of from about 12 nt to about 80 nt, from about 12 nt to about 50 nt, from about 12 nt to about 45 nt, from about 12 nt to about 40 nt, from about 12 nt to about 35 nt, from about 12 nt to about 30 nt, from about 12 nt to about 25 nt, from about 12 nt to about 20 nt, from about 12 nt to about19 nt, from about 19 nt to about 20 nt, from about 19 nt to about 25 nt, from about 19 nt to about 30 nt, from about 19 nt to about 35 nt, from about 19 nt to about 40 nt, from about 19 nt to about 45 nt, from about 19 nt to about 50 nt, from about 19 nt to about 60 nt, from about 20 nt to about 25 nt, from about 20 nt to about 30 nt, from about 20 nt to about 35 nt, from about20 nt to about 40 nt, from about 20 nt to about 45 nt, from about 20 nt to about 50 nt, or from about 20 nt to about 60 nt.
[0148] A protospacer sequence (i.e., target sequence) of a targeted polynucleotide (for example, on plant genomic DNA) can be identified by identifying a protospacer-adjacent motif (PAM) within a region of interest and selecting a region of a desired size upstream or downstream of the PAM as the protospacer. A corresponding spacer sequence can be designed by determining the complementary sequence of the protospacer region.
[0149] A spacer sequence (i.e., nucleic acid-targeting region) can be identified using a computer program (e.g., machine readable code). The computer program can use variables such as predicted melting temperature, secondary structure formation, predicted annealing temperature, sequence identity, genomic context, chromatin accessibility, % GC, frequency of genomic occurrence, methylation status, presence of SNPs, and the like.
[0150] The percent complementarity between the nucleic acid-targeting sequence (e.g., a spacer sequence of the at least one guide polynucleotide as disclosed herein) and the target nucleic acid (e.g., a protospacer sequence of the one or more target loci as disclosed herein) can be at least 50%, at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100%. The percentcomplementarity between the nucleic acid-targeting sequence and the target nucleic acid can be at least 60%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 97%, at least 98%, at least 99%, or 100% over about 20 contiguous nucleotides.
[0151] The Cas protein binding segment of a guide nucleic acid can have a length of from about 10 nucleotides to about 100 nucleotides, e.g., from about 10 nucleotides (nt) to about 20 nt, from about 20 nt to about 30 nt, from about 30 nt to about 40 nt, from about 40 nt to about 50 nt, from about 50 nt to about 60 nt, from about 60 nt to about 70 nt, from about 70 nt to about 80 nt, from about 80 nt to about 90 nt, or from about 90 nt to about 100 nt. For example, the Cas protein-binding segment of a guide nucleic acid can have a length of from about 15 nt to about 80 nt, from about 15 nt to about 50 nt, from about 15 nt to about 40 nt, from about 15 nt to about 30 nt or from about 15 nt to about 25 nt.
[0152] The dsRNA duplex of the Cas protein-binding segment of the guide nucleic acid can have a length from about 6 base pairs (bp) to about 50 bp. For example, the dsRNA duplex of the protein-binding segment can have a length from about 6 bp to about 40 bp, from about 6 bp to about 30 bp, from about 6 bp to about 25 bp, from about 6 bp to about 20 bp, from about 6 bp to about 15 bp, from about 8 bp to about 40 bp, from about 8 bp to about 30 bp, from about 8 bp to about 25 bp, from about 8 bp to about 20 bp or from about 8 bp to about 15 bp. For example, the dsRNA duplex of the Cas protein-binding segment can have a length from about from about 8 bp to about 10 bp, from about 10 bp to about 15 bp, from about 15 bp to about 18 bp, from about 18 bp to about 20 bp, from about 20 bp to about 25 bp, from about 25 bp to about 30 bp, from about 30 bp to about 35 bp, from about 35 bp to about 40 bp, or from about 40 bp to about 50 bp.
[0153] The percent complementarity between the nucleotide sequences that hybridize to form the dsRNA duplex of the protein-binding segment can be at least about 60%. For example, the percent complementarity between the nucleotide sequences that hybridize to form the dsRNA duplex of the protein-binding segment can be at least about 65%, at least about 70%, at least about 75%, at least about 80%, at least about 85%, at least about 90%, at least about 95%, at least about 98%, or at least about 99%. In some cases, the percent complementarity between the nucleotide sequences that hybridize to form the dsRNA duplex of the proteinbinding segment is 100%.
[0154] Guide nucleic acids can include modifications or sequences that provide for additional desirable features (e.g., modified or regulated stability, subcellular targeting;tracking with a fluorescent label; a binding site for a protein or protein complex; and the like). Examples of such modifications include, for example, a 5’ cap (a 7-methylguanylate cap (m7G)); a 3’ polyadenylated tail (a 3’ poly(A) tail); a riboswitch sequence (e.g., to allow for regulated stability and / or regulated accessibility by proteins and / or protein complexes); a stability control sequence; a sequence that forms a dsRNA duplex (a hairpin)); a modification or sequence that targets the RNA to a subcellular location (e.g., nucleus, mitochondria, chloroplasts, and the like); a modification or sequence that provides for tracking (e.g., direct conjugation to a fluorescent molecule, conjugation to a moiety that facilitates fluorescent detection, a sequence that allows for fluorescent detection, and so forth); a modification or sequence that provides a binding site for proteins (e.g., proteins that act on DNA, including transcriptional activators, transcriptional repressors, DNA methyl transferases, DNA demethylases, histone acetyltransferases, histone deacetylases, and combinations thereof.
[0155] A guide nucleic acid can comprise one or more modifications (e.g., a base modification, a backbone modification), to provide the nucleic acid with a new or enhanced feature (e.g., improved stability). A guide nucleic acid can comprise a nucleic acid affinity tag. A nucleoside can be a base-sugar combination. The base portion of the nucleotide can be a heterocyclic base. The two most common classes of such heterocyclic bases are the purines and the pyrimidines. Nucleotides can be nucleosides that further include a phosphate group covalently linked to the sugar portion of the nucleoside. For those nucleosides that include a pentofuranosyl sugar, the phosphate group can be linked to the 2', the 3', or the 5' hydroxyl moiety of the sugar. In forming guide nucleic acids, the phosphate groups can covalently link adjacent nucleosides to one another to form a linear polymeric compound. In turn, the respective ends of this linear polymeric compound can be further joined to form a circular compound; however, linear compounds can be suitable. In addition, linear compounds can have internal nucleotide base complementarity and can therefore fold in a manner as to produce a fully or partially double-stranded compound. Further, within guide nucleic acids, the phosphate groups can commonly be referred to as forming the intemucleoside backbone of the guide nucleic acid. The linkage orbackbone of the guide nucleic acid can be a 3' to 5' phosphodiester linkage.
[0156] Additional guide nucleic acid modifications are described, for example in International Application Nos. PCT / US2023 / 068974 and PCT / US2023 / 068977.
[0157] In some embodiments, the at least one gRNA polynucleotide disclosed herein can bind to at least a portion of a genome (e.g., a plant genome) or a gene (e.g., a plant gene). In some cases, the at least one gRNA polynucleotide is capable of forming a complex with a Cas protein to direct the Cas protein to target the portion of a target nucleic acid (e.g., a site in a genome or a gene).2. Meganucleases (MNs)
[0158] In some embodiments, the SDN is a meganuclease (MN). Meganucleases generally refer to rare-cutting endonucleases or homing endonucleases that can be highly specific. MNs can recognize DNA target sites ranging from at least 12 base pairs in length, e.g., from 12 to 40 base pairs, 12 to 50 base pairs, or 12 to 60 base pairs in length. MNs can be modular DNA- binding nucleases such as any fusion protein comprising at least one catalytic domain of an endonuclease and at least one DNA binding domain or protein specifying a nucleic acid target sequence. The DNA-binding domain can contain at least one motif that recognizes singlestranded DNA or double-stranded DNA. An MN can generate a double-strand break. A doublestrand break in DNA can result in DNA break repair which allows for the introduction of gene modification(s) (e.g., nucleic acid editing). DNA break repair can occur via NHEJ or HDR. In HDR, a donor DNA repair template or template polynucleotide that contains homology arms flanking sites of the target DNA can be provided. The MN can be monomeric or dimeric. In some embodiments, the MN is naturally-occurring (found in nature) or wild-type, and in other instances, the MN is non-natural, artificial, engineered, synthetic, rationally designed, or manmade. In some embodiments, the MN of the present disclosure includes an I-Crel MN, I-Ceul MN, I-Msol MN, I-Scel MN, variants thereof, derivatives thereof, and fragments thereof. Detailed descriptions of useful MNs and their application in gene editing are found, e.g., in Silva et al., Curr Gene Ther, 2011, 11(1): 11-27; Zaslavoskiy et al., BMC Bioinformatics, 2014, 15: 191; Takeuchi et al., Proc Natl Acad Sci USA, 2014, 111(11):4061-4066, and U.S. Patent Nos. 7,842,489; 7,897,372; 8,021,867; 8,163,514; 8,133,697; 8,021,867; 8,119,361; 8,119,381; 8,124,36; and 8,129,134.3. Zinc Finger Nucleases (ZFNs)
[0159] In some embodiments, the SDN is a zinc finger nuclease (ZFN). ZFNs refer to a fusion between a cleavage domain, such as a cleavage domain of Fokl, and at least one zinc finger motif (e.g., at least 2, 3, 4, or 5 zinc finger motifs) which can bind polynucleotides such as DNA and RNA. The heterodimerization at certain positions in a polynucleotide of two individual ZFNs in certain orientation and spacing can lead to cleavage of the polynucleotide.For example, a ZFN binding to DNA can induce a double-strand break in the DNA. In order to allow two cleavage domains to dimerize and cleave DNA, two individual ZFNs can bind opposite strands of DNA with their C-termini at a certain distance apart. In some cases, linker sequences between the zinc finger domain and the cleavage domain can require the 5' edge of each binding site to be separated by about 5-7 base pairs. In some cases, a cleavage domain is fused to the C-terminus of each zinc finger domain. Exemplary ZFNs include, but are not limited to, those described in Urnov et al., Nature Reviews Genetics, 2010, 11 : 636-646; Gaj et al., Nat Methods, 2012, 9(8):805-7; U.S. Patent Nos. 6,534,261; 6,607,882; 6,746,838; 6,794,136; 6,824,978; 6,866,997; 6,933,113; 6,979,539; 7,013,219; 7,030,215; 7,220,719; 7,241,573; 7,241,574; 7,585,849; 7,595,376; 6,903,185; 6,479,626; and U.S. Publication Nos. 2003 / 0232410 and 2009 / 0203140.
[0160] In some embodiments, an SDN comprising a ZFN can generate a double-strand break in a target polynucleotide, such as DNA. A double-strand break in DNA can result in DNA break repair which allows for the introduction of gene modification(s) (e.g., nucleic acid editing). DNA break repair can occur via NHEJ or HDR. In HDR, a donor DNA repair template or template polynucleotide that contains homology arms flanking sites of the target DNA can be provided. In some embodiments, a ZFN is a zinc finger nickase which induces site-specific single-strand DNA breaks or nicks, thus resulting in HR. Descriptions of zinc finger nickases are found, e.g., in Ramirez et al., Nucl Acids Res, 2012, 40(12):5560-8; Kim et al., Genome Res, 2012, 22(7): 1327-33.4. Transcription Activator-Like Effector Nucleases (TALENs)
[0161] In some embodiments, the SDN is a transcription activator-like effector nuclease (TALEN; TAL-effector nuclease). TALENs refer to engineered transcription activator-like effector nucleases that generally contain a central domain of DNA-binding tandem repeats and a cleavage domain. TALENs can be produced by fusing a TAL effector DNA binding domain to a DNA cleavage domain. In some cases, a DNA-binding tandem repeat comprises 33-35 amino acids in length and contains two hypervariable amino acid residues at positions 12 and 13 that can recognize at least one specific DNA base pair. A transcription activator-like effector (TALE) protein can be fused to a nuclease such as a wild-type or mutated Fokl endonuclease or the catalytic domain of Fokl. Several mutations to Fokl have been made for its use in TALENs, which, for example, improve cleavage specificity or activity. Such TALENs can be engineered to bind any desired DNA sequence. TALENs can be used to generate gene modifications (e.g., nucleic acid sequence editing) by creating a double-strand break in a targetDNA sequence, which in turn, undergoes NHEJ or HR. A double-strand break in DNA can result in DNA break repair which allows for the introduction of gene modification(s) (e.g., nucleic acid editing). DNA break repair can occur via NHEJ or HDR. In HDR, a donor DNA repair template or template polynucleotide that contains homology arms flanking sites of the target DNA can be provided. In some cases, a single-stranded donor DNA repair template is provided to promote HR. Detailed descriptions of TALENs and their uses for gene editing are found, e.g., in U.S. Patent Nos. 8,440,431; 8,440,432; 8,450,471; 8,586,363; and 8,697,853; Scharenberg et al., Curr Gene Ther, 2013, 13(4):291-303; Gaj et al., Nat Methods, 2012, 9(8):805-7; Beurdeley et al., Nat Commun, 2013, 4: 1762; and Joung and Sander, Nat Rev Mol Cell Biol, 2013, 14(I):49-55.IV. Recombinant Nucleic Acids
[0162] Also provided herein are recombinant nucleic acids, i.e., DNA constructs, that may be used with methods, e.g., Hot-Edit methods, of the present disclosure. Each DNA construct is a vector that is contemplated to have the necessary functional elements that direct and regulate transcription of the inserted nucleic acid. These functional elements include, but are not limited to, a promoter, regions upstream or downstream of the promoter, such as enhancers and terminators, that may regulate the transcriptional activity of the promoter, an origin of replication, appropriate restriction sites to facilitate cloning of inserts adjacent to the promoter, antibiotic resistance genes or other markers which can serve to select for cells containing the vector or the vector containing the insert, RNA splice junctions, a transcription termination region, or any other region which may serve to facilitate the expression of the inserted gene or hybrid gene. See generally, Sambrook et al. Molecular Cloning: A Laboratory Manual, 4thed., Cold Spring Harbor Laboratory Press, Cold Spring Harbor, 2012. The vector, for example, can be a plasmid. In some embodiments of the DNA constructs and vectors provided herein, the constructs and vectors comprise one or more elements that are disclosed in Tables 10-18 and in any combination. In some embodiments, the vectors are constructs 27145 (SEQ ID NO: 1), 27146 (SEQ ID NO: 2), 27680 (SEQ ID NO: 3), 28255 (SEQ ID NO: 9), 28291 (SEQ ID NO: 10). 28292 (SEQ ID NO: 11). 28293 (SEQ ID NO: 12), 28294 (SEQ ID NO: 13), 25072 (SEQ ID NO: 14), 28825 (SEQ ID NO: 15), or 28834 (SEQ ID NO: 16).V. Transformation Methods
[0163] The recombinant nucleic acids disclosed herein may also be used in transformation of a transgenic cell, plant cell, plant and / or plant part. Transformation of a cell may be stableor transient. Transformation can refer to the transfer of a nucleic acid molecule into the genome of a host cell, resulting in genetically stable inheritance. In some embodiments, the introduction into a plant, plant part and / or plant cell is via bacterial-mediated transformation, particle bombardment transformation, calcium-phosphate-mediated transformation, cyclodextrin- mediated transformation, electroporation, liposome-mediated transformation, nanoparticle- mediated transformation, polymer-mediated transformation, virus-mediated nucleic acid delivery, whisker-mediated nucleic acid delivery, microinjection, sonication, infiltration, polyethylene glycol-mediated transformation, protoplast transformation, or any other electrical, chemical, physical and / or biological mechanism that results in the introduction of nucleic acid into the plant, plant part and / or cell thereof, or any combination thereof.
[0164] Procedures for transforming plants are well known and routine in the art and are described throughout the literature. Non-limiting examples of methods for transformation of plants include transformation via bacterial-mediated nucleic acid delivery (e.g. via bacteria from the genus Agrobacterium), viral -mediated nucleic acid delivery, silicon carbide or nucleic acid whisker-mediated nucleic acid delivery, liposome mediated nucleic acid delivery, microinjection, microparticle bombardment, calcium-phosphate-mediated transformation, cyclodextrin-mediated transformation, electroporation, nanoparticle-mediated transformation,, sonication, infiltration, PEG-mediated nucleic acid uptake, as well as any other electrical, chemical, physical (mechanical) and / or biological mechanism that results in the introduction of nucleic acid into the plant cell, including any combination thereof. General guides to various plant transformation methods known in the art include Miki et al. (“Procedures for Introducing Foreign DNA into Plants” in Methods in Plant Molecular Biology and Biotechnology, Glick, B. R. and Thompson, J. E., Eds. (CRC Press, Inc., Boca Raton, 1993), pages 67-88), and Rakowoczy-Trojanowska (Cell Mol Biol Lett 7:849-858 (2002)).
[0165] Agrobaclerium-m <) &i transformation is a commonly used method for transforming plants because of its high efficiency of transformation and because of its broad utility with many different species, ^grotocterzwm-mediated transformation typically involves transfer of the binary vector carrying the foreign DNA of interest to an appropriate Agrobacterium strain that may depend on the complement of vir genes carried by the host Agrobacterium strain either on a co-resident Ti plasmid or chromosomally (Uknes et al. 1993, Plant Cell 5: 159-169). The transfer of the recombinant binary vector to Agrobacterium can be accomplished by a tri-parental mating procedure using Escherichia coli carrying the recombinant binary vector, a helper A. coli strain that carries a plasmid that is able to mobilizethe recombinant binary vector to the target Agrobacterium strain. Alternatively, the recombinant binary vector can be transferred to Agrobacterium by nucleic acid transformation (Hbfgen and Willmitzer 1988, Nucleic Acids Res 16:9877).
[0166] Transformation of a plant by recombinant Agrobacterium usually involves cocultivation of the Agrobacterium with explants from the plant and follows methods well known in the art. Transformed tissue is typically regenerated on selection medium carrying an antibiotic or herbicide resistance marker between the binary plasmid T-DNA borders.
[0167] Another method for transforming plants, plant parts and plant cells involves propelling inert or biologically active particles at plant tissues and cells. See, e.g., U.S. Patent Nos. 4,945,050; 5,036,006 and 5,100,792. Generally, this method involves propelling inert or biologically active particles at the plant cells under conditions effective to penetrate the outer surface of the cell and afford incorporation within the interior thereof. When inert particles are utilized, the vector can be introduced into the cell by coating the particles with the vector containing the nucleic acid of interest. Alternatively, a cell or cells can be surrounded by the vector so that the vector is carried into the cell by the wake of the particle. Biologically active particles (e.g., dried yeast cells, dried bacteria, or a bacteriophage, each containing one or more nucleic acids sought to be introduced) also can be propelled into plant tissue. As used herein, the phrase “biolistic transformation” refers to a method of introducing RNA or DNA into cells (e.g., plant cells) directly, in which RNA or DNA is mixed with heavy metal particles (e.g., tungsten or gold) and released into the cell (e.g., the plant cell) using high speed pressure to allow the RNA or DNA to penetrate the cell (e.g., to penetrate the plant cell wall).VI. Methods of Genotyping
[0168] A variety of means can be used to genotype an individual (e.g., a plant) to determine whether a HI event has occurred. A variety of means can be used to determine the genotype at a polymorphic site of interest such as a gene (e.g., MATL, CENH3), a QTL, or a mitochondrial genome locus. In some embodiments, a genotyping assay is used to determine whether a sample (e.g., a nucleic acid sample) contains a specific variant allele (e.g., genomic edit, mutation, or QTL marker) or haplotype. For example, enzymatic amplification of nucleic acid from an individual can be conveniently used to obtain nucleic acid for subsequent analysis. The presence or absence of a specific variant allele (e.g., mutation or QTL marker) or haplotype in one or more loci of interest can also be determined directly from the individual’s nucleic acid without enzymatic amplification. In certain embodiments, an individual is genotyped at one,two, three, four, five, or more polymorphic sites such as a single nucleotide polymorphism (SNP) in one or more loci of interest. In some embodiments, an individual is genotyped at one, two, three, four, five, or more polymorphic sites in one or more loci of interest in the mitochondrial genome (e.g., to distinguish NA cytotype individuals from individuals of other cytotypes).
[0169] Genotyping of nucleic acid from an individual, whether amplified or not, can be performed using any of various techniques. Useful techniques include, without limitation, assays such as polymerase chain reaction (PCR) based analysis assays, sequence analysis assays, electrophoretic analysis assays, restriction length polymorphism analysis assays, hybridization analysis assays, allele-specific hybridization, oligonucleotide ligation allelespecific elongation / ligation, allele-specific amplification, single-base extension, molecular inversion probe, invasive cleavage, selective termination, restriction length polymorphism, sequencing, single strand conformation polymorphism (SSCP), single strand chain polymorphism, mismatch-cleaving, and denaturing gradient gel electrophoresis, all of which can be used alone or in combination.
[0170] Material containing nucleic acid is routinely obtained from individual plants. Such material is any biological matter from which nucleic acid can be prepared. As a non-limiting example, material can be plant parts (e.g., leaves, stems, roots, flowers or flower parts, fruits, pollen, egg cells, zygotes, seeds, cuttings, cell or tissue cultures, or any other part or product of a plant) or any plant tissue or other plant part that comprises nucleic acid. In one embodiment, a method of the present disclosure (e.g., in Example 4 below) is practiced with a leaf punch from a seedling, which can be obtained readily by non-invasive means and used to prepare genomic and / or mitochondrial DNA. In another embodiment, genotyping involves amplification of an individual’s nucleic acid using the polymerase chain reaction (PCR).
[0171] Any of a variety of different primers can be used to amplify an individual’s nucleic acid by PCR in order to determine the presence or absence of a variant allele (e.g., mutation or QTL marker) in a plant or method of the present disclosure. As understood by one skilled in the art, primers for PCR analysis can be designed based on the sequence flanking the polymorphic site(s) of interest in the gene of interest. As a non-limiting example, a sequence primer can contain from about 15 to about 30 nucleotides of a sequence upstream or downstream of the polymorphic site of interest in the gene or locus of interest. Such primers generally are designed to have sufficient guanine and cytosine content to attain a high meltingtemperature which allows for a stable annealing step in the amplification reaction. Several computer programs, such as Primer Select, are available to aid in the design of PCR primers.
[0172] An allelic discrimination assay (e.g., a TaqMan® assay available from Applied Biosystems) can be useful for genotyping an individual at a polymorphic site to thereby determine the presence or absence of a particular variant allele (e.g., mutation or QTL marker) or haplotype in the gene or locus of interest. In a TaqMan® allelic discrimination assay, a specific fluorescent dye-labeled probe for each allele is constructed. The probes contain different fluorescent reporter dyes such as FAM and TET to differentiate amplification of each allele. In addition, each probe has a quencher dye at one end which quenches fluorescence by fluorescence resonance energy transfer. During PCR, each probe anneals specifically to complementary sequences in the nucleic acid from the individual. The 5’ nuclease activity of Taq polymerase is used to cleave only probe that hybridizes to the allele. Cleavage separates the reporter dye from the quencher dye, resulting in increased fluorescence by the reporter dye. Thus, the fluorescence signal generated by PCR amplification indicates which alleles are present in the sample. Mismatches between a probe and allele reduce the efficiency of both probe hybridization and cleavage by Taq polymerase, resulting in little to no fluorescent signal. Those skilled in the art understand that improved specificity in allelic discrimination assays can be achieved by conjugating a DNA minor groove binder (MGB) group to a DNA probe as described, e.g., in Kutyavin et al., Nuc. Acids Research 28:655-661 (2000). Minor groove binders include, but are not limited to, compounds such as dihydrocyclopyrroloindole tripeptide (DPI3).
[0173] Sequence analysis can also be useful for genotyping an individual according to the methods described herein to determine the presence or absence of a particular variant allele (e.g., mutation or QTL marker) or haplotype in the gene or locus of interest. As is known by those skilled in the art, a variant allele of interest can be detected by sequence analysis using the appropriate primers, which are designed based on the sequence flanking the polymorphic site of interest in the gene or locus of interest. For example, a variant allele in a gene or locus of interest can be detected by sequence analysis using primers designed by one of skill in the art. Additional or alternative sequence primers can contain from about 15 to about 30 nucleotides of a sequence that corresponds to a sequence about 40 to about 400 base pairs upstream or downstream of the polymorphic site of interest in the gene or locus of interest. Such primers are generally designed to have sufficient guanine and cytosine content to attain a high melting temperature which allows for a stable annealing step in the sequencing reaction.A sequence analysis may be performed by any manual or automated process by which the order of nucleotides in a nucleic acid is determined, and encompasses, without limitation, chemical and enzymatic methods.
[0174] Electrophoretic analysis also can be useful in genotyping an individual according to the methods of the present disclosure to determine the presence or absence of a particular variant allele (e.g., mutation or QTL marker) or haplotype in the gene or locus of interest. Electrophoretic analysis includes a process whereby charged molecules, e.g., one or more amplified fragments of nucleic acids, are moved through a stationary medium under the influence of an electric field. Methods of electrophoretic analysis, and variations thereof, are well known in the art, as described, for example, in Ausubel et al., Current Protocols in Molecular Biology Chapter 2 (Supplement 45) John Wiley & Sons, Inc. New York (1999).
[0175] Restriction fragment length polymorphism (RFLP) analysis can also be useful for genotyping an individual according to the methods of the present disclosure to determine the presence or absence of a particular variant allele (e.g., mutation or QTL marker) or haplotype in the gene or locus of interest (see, Jarcho et al. in Dracopoli et al., Current Protocols in Human Genetics pages 2.7.1-2.7.5, John Wiley & Sons, New York; Innis et al., (Ed.), PCR Protocols, San Diego: Academic Press, Inc. (1990)). RFLP analysis may be performed on PCR amplification products.
[0176] In addition, allele-specific oligonucleotide hybridization can be useful for genotyping an individual in the plants or methods described herein to determine the presence or absence of a particular variant allele (e.g., mutation or QTL marker) or haplotype in the gene or locus of interest. Allele-specific oligonucleotide hybridization is based on the use of a labeled oligonucleotide probe having a sequence perfectly complementary, for example, to the sequence encompassing the variant allele. Under appropriate conditions, the variant allelespecific probe hybridizes to a nucleic acid containing the variant allele but does not hybridize to the one or more other alleles, which have one or more nucleotide mismatches as compared to the probe. If desired, a second allele-specific oligonucleotide probe that matches an alternate (e.g., wild-type) allele can also be used. Similarly, the technique of allele-specific oligonucleotide amplification can be used to selectively amplify, for example, a variant allele by using an allele-specific oligonucleotide primer that is perfectly complementary to the nucleotide sequence of the variant allele, but which has one or more mismatches as compared to other alleles (Mullis et al., supra). One skilled in the art understands that the one or morenucleotide mismatches that distinguish between the variant allele and other alleles are often located in the center of an allele-specific oligonucleotide primer to be used in the allele-specific oligonucleotide hybridization. In contrast, an allele-specific oligonucleotide primer to be used in PCR amplification generally contains the one or more nucleotide mismatches that distinguish between the variant and other alleles at the 3' end of the primer.
[0177] A heteroduplex mobility assay (HMA) is another well-known assay that can be used for genotyping in the plants or methods of the present disclosure to determine the presence or absence of a particular variant allele (e.g., mutation or QTL marker) or haplotype in the gene or locus of interest. HMA is useful for detecting the presence of a variant allele since a DNA duplex carrying a mismatch has reduced mobility in a polyacrylamide gel compared to the mobility of a perfectly base-paired duplex (see, Delwart et al., Science, 262: 1257-1261 (1993); White et al., Genomics, 12:301-306 (1992)).
[0178] The technique of single strand conformational polymorphism (SSCP) can also be useful for genotyping in the plants or methods described herein to determine the presence or absence of a particular variant allele (e.g., mutation or QTL marker) or haplotype in the gene or locus of interest (see, Hayashi, Methods Applic., 1 :34-38 (1991)). This technique is used to detect variant alleles based on differences in the secondary structure of single-stranded DNA that produce an altered electrophoretic mobility upon non-denaturing gel electrophoresis. Variant alleles are detected by comparison of the electrophoretic pattern of the test fragment to corresponding standard fragments containing known alleles.
[0179] Denaturing gradient gel electrophoresis (DGGE) can also be useful in the plants or methods of the present disclosure to determine the presence or absence of a particular variant allele (e.g., mutation or QTL marker) or haplotype in the gene or locus of interest. In DGGE, double-stranded DNA is electrophoresed in a gel containing an increasing concentration of denaturant; double-stranded fragments made up of mismatched alleles have segments that melt more rapidly, causing such fragments to migrate differently as compared to perfectly complementary sequences (see, Sheffield et al., “Identifying DNA Polymorphisms by Denaturing Gradient Gel Electrophoresis” in Innis et al., supra, 1990).
[0180] Other molecular methods useful for genotyping an individual are known in the art and useful in the plants or methods of the present disclosure. Such well-known genotyping approaches include, without limitation, automated sequencing and RNase mismatch techniques (see, Winter et al., Proc. Natl. Acad. Sci., 82:7575-7579 (1985)). Furthermore, one skilled inthe art understands that, where the presence or absence of multiple variant alleles is to be determined, individual variant alleles can be detected by any combination of molecular methods. See, in general, Birren et al. (Eds.) Genome Analysis: A Laboratory Manual Volume 1 (Analyzing DNA) New York, Cold Spring Harbor Laboratory Press (1997). In addition, one skilled in the art understands that multiple variant alleles can be detected in individual reactions or in a single reaction (a “multiplex” assay).VII. Plants
[0181] In another aspect, provided are the plants, plant parts, or seeds produced from the Hot-Edit methods provided in this disclosure. Plants produced as described above can be propagated to produce progeny plants, and the progeny plants that have stably incorporated into their genome the gene edit introduced by the gene editing machinery of one of the donor parent plants, which can be selected for and can be further propagated if desired. In some embodiments, a plant cell, seed, or plant part or harvest product can be obtained from the plant produced as above, and the plant cell, seed, or plant part can be screened for evidence of the gene edit.
[0182] In some embodiments, plant products can be harvested from the plant disclosed above and processed to produce processed products, such as flour, meal, oil, starch, and the like. These processed products are also within the scope of this invention provided that they comprise a gene edit introduced by the methods disclosed herein. Other plant products include but are not limited to protein concentrate, protein isolate, seed hulls, meal, flower, and oil.VIII. Exemplary Embodiments
[0183] As used below, any reference to a series of Embodiments is to be understood as a reference to each of those Embodiments disjunctively (e.g., "Embodiments 1-4" is to be understood as "Embodiments 1, 2, 3, or 4").
[0184] Embodiment 1 is a method for increasing the editing efficiency in a plant cell, comprising applying a heat treatment to the cell, wherein the cell comprises plant genomic DNA, a site-directed nuclease, and a guide nucleic acid; wherein the editing efficiency in the heat-treated plant cell is increased compared to a control plant cell; and wherein the plant cell is a haploid plant cell.
[0185] Embodiment 2 is the method of Embodiment 1, wherein the site-directed nuclease is selected from the group consisting of a meganuclease (MN), a zinc-finger nuclease (ZFN), a transcription-activator like effector nuclease (TALEN), and a Cas nuclease.
[0186] Embodiment 3 is the method of Embodiment 2, wherein the Cas nuclease is a Type II Cas nuclease, a Type IV Cas nuclease, or a Type V Cas nuclease.
[0187] Embodiment 4 is the method of Embodiment 3, wherein the Type II Cas nuclease is a Cas9 nuclease, a Cas9 nickase, a nuclease-inactive Cas9, or a Cas9 fused to a heterologous domain.
[0188] Embodiment 5 is the method of Embodiment 3, wherein the Type V Cas nuclease is a Cas 12a nuclease, a Cas 12a nickase, a nuclease-inactive Cas 12a, or a Cas 12a fused to a heterologous domain.
[0189] Embodiment 6 is the method of any one of Embodiments 1-5, wherein the guide nucleic acid is a guide RNA.
[0190] Embodiment 7 is the method of Embodiment 1, wherein the haploid plant cell is obtained by crossing a donor plant with a haploid inducer plant.
[0191] Embodiment 8 is the method of Embodiment 7, wherein the haploid inducer plant is a maternal haploid inducer plant.
[0192] Embodiment 9 is the method of Embodiment 8, wherein the maternal haploid inducer plant comprises a knock-out mutation in a MATL gene.
[0193] Embodiment 10 is the method of Embodiment 7, wherein the haploid inducer plant is a paternal haploid inducer plant.
[0194] Embodiment 11 is the method of Embodiment 10, wherein the paternal haploid inducer plant comprises a heterozygous mutation in the CENH3 gene.
[0195] Embodiment 12 is the method of Embodiment 1, wherein the haploid cell is treated with a chromosome doubling agent, thereby creating a doubled haploid cell.
[0196] Embodiment 13 is the method of Embodiment 12, wherein the chromosome doubling agent is colchicine, pronamide, dithipyr, trifluralin, or another known anti -microtubule agent.
[0197] Embodiment 14 is the method of Embodiment 1, wherein the heat treatment comprises a temperature between 30°C and 40°C, inclusive.
[0198] Embodiment 15 is the method of Embodiment 14, wherein the heat treatment comprises a temperature between 34°C and 39°C.
[0199] Embodiment 16 is the method of Embodiment 15, wherein the heat treatment comprises a temperature of 35°C.
[0200] Embodiment 17 is the method of Embodiment 1, wherein the heat treatment comprises a duration between 12 hours and 72 hours.
[0201] Embodiment 18 is the method of Embodiment 17, wherein the heat treatment comprises a duration of 24 hours.
[0202] Embodiment 19 is the method of Embodiment 1, wherein the heat treatment comprises placing the cell in a chamber with an increased ambient temperature.
[0203] Embodiment 20 is the method of Embodiment 1, wherein the heat treatment comprises applying a heat pack.
[0204] Embodiment 21 is a method of editing plant genomic DNA, comprising: providing an egg-donor plant comprising plant genomic DNA that is to be edited; pollinating the eggdonor plant with a pollen-donor plant, wherein the pollen-donor plant expresses a DNA modification enzyme and, optionally, a guide nucleic acid; applying a heat treatment to the pollinated egg-donor plant of step b; and producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the egg-donor plant and does not comprise the genome of the pollen-donor haploid inducer plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and, optionally, the guide nucleic acid delivered by the pollen-donor haploid inducer plant.
[0205] Embodiment 22 is a method of editing plant genomic DNA, comprising: providing a pollen-donor plant that expresses a DNA modification enzyme and, optionally, a guide nucleic acid; applying a heat treatment to pollen of the pollen-donor plant; pollinating an egg-donor plant with heat-treated pollen of the pollen-donor plant, wherein the egg-donor plant comprises the plant genomic DNA that is to be edited; and producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the egg-donor plant and does not comprise the genome of the pollen-donor plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and, optionally, the guide nucleic acid delivered by the pollen-donor plant.
[0206] Embodiment 23 is the method of Embodiments 21 or 22, wherein the pollen-donor plant is a haploid inducer plant.
[0207] Embodiment 24 is the method of Embodiments 21 or 22, wherein the pollen-donor plant is a maternal haploid inducer plant.
[0208] Embodiment 25 is the method of Embodiment 24, wherein the maternal haploid inducer plant comprises a knock-out mutation in a MATL gene.
[0209] Embodiment 26 is a method of editing plant genomic DNA, comprising: providing a pollen-donor plant comprising plant genomic DNA that is to be edited; pollinating an eggdonor plant with pollen from the pollen-donor plant, wherein the egg-donor plant expresses a DNA modification enzyme and, optionally, a guide nucleic acid; applying a heat treatment to the pollinated egg-donor plant of step b; and producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the pollen-donor plant and does not comprise the genome of the egg-donor plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and, optionally, the guide nucleic acid delivered by the egg-donor plant.
[0210] Embodiment 27 is a method of editing plant genomic DNA, comprising: providing a pollen-donor plant comprising plant genomic DNA that is to be edited; applying a heat treatment to pollen of the pollen-donor plant; pollinating an egg-donor plant with heat-treated pollen of the pollen-donor plant, wherein the egg-donor plant expresses a DNA modification enzyme and, optionally, a guide nucleic acid; and producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the pollen-donor plant and does not comprise the genome of the egg-donor plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and the optional guide nucleic acid delivered by the egg-donor plant.
[0211] Embodiment 28 is the method of Embodiments 26 or 27, wherein the egg-donor plant is a haploid inducer plant.
[0212] Embodiment 29 is the method of Embodiments 26 or 27, wherein the egg-donor plant is a paternal haploid inducer plant.
[0213] Embodiment 30 is the method of Embodiments 29, wherein the paternal haploid inducer plant comprises a mutation in a CENH3 gene.
[0214] Embodiment 31 is the method of Embodiments 30, wherein the paternal haploid inducer plant is heterozygous for the mutation in the CENH3 gene.
[0215] Embodiment 32 is the method of any one of Embodiments 21-31, wherein at least one of the egg-donor plant or the pollen-donor plant is a maize plant.
[0216] Embodiment 33 is the method of Embodiment 32, wherein the maize plant is selected and / or derived from the lines Stock 6, RWK, RWS, UH400, NP2222RS, or NP2222.
[0217] Embodiment 34 is the method of any one of Embodiments 21-33, wherein the pollendonor plant is a maize plant.
[0218] Embodiment 35 is the method of any one of Embodiments 21-33, wherein the eggdonor plant is a maize plant.
[0219] Embodiment 36 is the method of any one of Embodiments 21-35, wherein the DNA modification enzyme is a site-directed nuclease selected from the group consisting of a meganuclease (MN), a zinc-finger nuclease (ZFN), a transcription-activator like effector nuclease (TALEN), and a Cas nuclease.
[0220] Embodiment 37 is the method of Embodiment 36, wherein the Cas nuclease is a Type II Cas nuclease, a Type IV Cas nuclease, or a Type V Cas nuclease.
[0221] Embodiment 38 is the method of Embodiment 37, wherein the Type II Cas nuclease is a Cas9 nuclease, a Cas9 nickase, a nuclease-inactive Cas9, or a Cas9 fused to a heterologous domain.
[0222] Embodiment 39 is the method of Embodiment 38, wherein the Type V Cas nuclease is a Casl2a nuclease, a Casl2a nickase, a nuclease-inactive Casl2a, or a Casl2a fused to a heterologous domain.
[0223] Embodiment 40 is the method of any one of Embodiments 21-39, wherein the guide nucleic acid is a guide RNA.
[0224] Embodiment 41 is the method of any one of Embodiments 21-40, wherein the haploid progeny is treated with a chromosome doubling agent, thereby creating an edited doubled haploid progeny.
[0225] Embodiment 42 is the method of Embodiment 41, wherein the chromosome doubling agent is colchicine, pronamide, dithipyr, trifluralin, or another known anti -microtubule agent.
[0226] Embodiment 43 is the method of any one of Embodiments 21-42, wherein the pollendonor plant expresses a marker gene.
[0227] Embodiment 44 is the method of Embodiment 43, wherein the marker gene is selected from the group consisting of Rl, R1-SCM2, Rl-nj, GUS, PMI, PAT, GFP, RFP, CFP, Bl, CI, anthocyanin pigments, or any other marker gene.
[0228] Embodiment 45 is an edited haploid plant produced by the method of any one of Embodiments 21-44.
[0229] Embodiment 46 is a progeny plant of the edited haploid plant of Embodiment 45.EXAMPLES
[0230] The following Examples relate to the testing of post-pollination environmental conditions to increase overall haploid induction rate (EUR) and haploid editing rate (HER). Application of heat with HI-Edit methods was of particular interest.Example 1. Chromatin is Less Compact at Elevated Temperatures.
[0231] To test whether heat relaxed or otherwise allowed chromatin to become uncompact, maize pollen grains were imaged with 4',6-diamidino-2-phenylindole (DAPI) to examine the sperm nuclear size and morphology under control (Figure 1 A) and heat treatment (Figure IB). Fresh maize pollen was mixed 1 : 10 v / v in mineral oil and incubated at room temperature (RT, control) or at 45°C for one hour, and then fixed in 3: 1 ethanol: acetic acid, taken through a dehydration series, and stained with DAPI, according to the protocol provided in S. Heuer, et al., The MADS box gene ZmMADS2 is specifically expressed in maize pollen and during pollen tube growth, SEXUAL PLANT REPRODUCTION 13: 21-27 (2000). The heat treatment caused an enlargement of the sperm nuclei (Figure IB), consistent with the hypothesis that a heat treatment during HI-Edit may relax the compact chromatin characteristic of sperm nuclei.Example 2. Heat Increases the Haploid Editing Rate (HER).
[0232] To test whether heat treatment increased the haploid induction rate (“HIR”) or haploid editing rate (“HER”), several CRISPR editing events and constructs for HI-Edit efficiency were compared under control or heat treatments. The timing of treatments ranged from a few hours to three days after pollination. This range, termed the “HI-Edit window,” covers the period of time during which the editing may occur prior to genome elimination. Normal growthconditions were utilized for control experiments, i.e., 15-hour days at 25-30°C and 9-hour nights at 16-20°C.
[0233] Two heat treatment conditions were tested. In the first treatment, termed the “heated pack treatment,” heat packs were attached to maize ears, wrapped around the husks with an insulation barrier between the heat pack and the ear, and changed every 12 hours to maintain the heat treatment. The heat packs used are consumer products, which can reach temperatures of approximately 50°C for approximately 12 hours. See, e.g., Thermacare® Muscle Pain Therapy HeatWraps, thermacare.com / heat-wraps / muscle-pain-therapy / (last visited Aug. 15, 2023). In this method the internal temperature of the ears reached about 35°C. Results are shown in Tables 3-4 below.
[0234] In the second heat treatment, termed the “heated room treatment,” the daytime temperature of the glasshouse room was raised to 35°C, and the nighttime temperature was 25°C. Results are shown in Table 5 below.
[0235] Three different constructs were tested: 27145, 27146, and 27680 (Tables 10-12). These three constructs all contained a Casl2a cassette, a phosphomannose isomerase (PMI) selectable marker cassette, and a gRNA cassette expressing a gRNA capable of complexing with the Cast 2a protein to edit the Waxyl gene in maize.
[0236] The inducer line used was NP3003RS, a male inducer comprising the mall. dmp haploid induction genes and the R1-SCM2 color marker gene. The testers were Tester 8 and Tester 2, which are stiff stalk and non-stiff stalk, respectively. At least 16, and usually more than 20 ears, in the control treatments were pollinated, harvested, and embryo rescued to identify edited haploids. In the heat treatments, between six and ten ears were tested per treatment. The HIR was calculated based on the R1-SCM2 color marker change. The haploids were then sampled with TaqMan® target site assay (Applied Biosystems), and then the putative edited haploids were sequenced by next generation sequencing (NGS). The TaqMan® data served as confirmation of the true haploids using genotyping calls (>99% of color-change identified haploids were called as true haploids using genotyping). The NGS sequencing data was used to confirm the haploid editing and identify the sequences of the edits, and was used to calculate the final HER.
[0237] The results of the trials for editing at the Waxyl gRNA site and GL2 gRNA control site using the NP3003RS male haploid inducer line are shown in Tables 3-5 below. Tables 3-5 show Maternal haploid induction rate, haploids per ear, and haploid editing rate (HER) fromheat treated ears / plants and from the controls. HERs were determined as described in WO 2015 / 104358.
[0238] There was a major increase in Waxyl HER (HI-Edit efficiency) in both heat heated pack and heated room treatments versus control (GL2 HER) - about a five-fold improvement in the heated room treatment (Table 5), and greater than an eight-fold improvement in the heated pack treatment (Tables 3-4). This was a surprisingly large increase in editing efficiency.
[0239] In the heated pack treatment, the HER jumped in all three events (Table 3), averaging an increase from less than 1% in the control to greater than 8% in the Tester 2 background. In the Tester 8 background, the HER was at 5% while in its control, HER was at 0.2% (Table 2). Thus, each event was affected positively and significantly by the heated pack treatment. There was a loss of seed set in both heated pack treatments, but HIR was not affected. The equivalent HIR is surprising as a negative result, given the publication of a heat treatment increasing the HIR of a CENH3 mutant in Arabidopsis (See Jin, et al., supra). The lowered seed set in the heat treatment might be mitigated by a more focused time period for the heat treatment (for instance, from 10 to 34 hours after pollination (1 day) rather than from 3 to 75 hours after (3 days), as was approximately done in this example).
[0240] The degree of increase of HER was particularly surprising, as it far exceeded the editing rate increase seen in heat treatments during transformation. The fact that this treatment worked across different cassettes and events indicates it may boost for HI-Edit efficiency of the cassette design, promoters used, or specific constructs tested. Without wishing to be bound by theory, this data suggests that the reason for the increased HER may be a combination of higher enzymatic activity and higher expression of the CRISPR transgene.Example 3. Additional HER testing.
[0241] The additional constructs, comprising Casl2a-encoding sequences linked to different promoter and terminator combinations, as listed in Table 6, were tested with control and heat treatments using the same trial design as in Example 2 using Tester 2. Each construct carries the same gRNAs (Waxy and GL2). At least 170 haploids per event were tested for the control, and the haploid number for the heat treatment ranged from 38 to 166 (Table 6). The data againshows a dramatic increase in haploid editing rate, a key driver of HI-Edit efficiency. Haploid selection and editing event detection (TaqMan® and NGS) were performed as described in Example 2 above.
[0242] Certain events from the first trial (focused on the prZmRZDP, prZmVSP, and prSoUbi4 promoters) were selected for receiving the control and heat pack treatments again, assuming that they had enough seeds to perform the experiment after the expected homozygous Casl2a-positive events were selected. The testers were Tester 2 and Tester 8 again. At least 129 haploids per tester x event combination were tested for control, and between 1 and 226 haploids were tested for the heat treatment. Haploid selection and editing event detection (TaqMan® and NGS) were performed as described in Example 2 above. The results are shown in Table 7 (tester 8) and Table 8 (tester 2) below.Table 6. Constructs for additional HER testing (heated pack treatment for Tester 2).Table 7. Results for additional events (heated pack treatment for Tester 8).Table 8. Results for additional events (heated pack treatment for Tester 2).Example 4. Paternal Haploid Induction System
[0243] Transformable CENH3 paternal haploid inducer lines were transformed with the vector 27241, described previously in International Application Nos. PCT / US2017 / 064512, PCT / CN2023 / 110941, and PCT / US2021 / 036605. Vector 27241 carries Casl2a, PMI, and two gRNA cassettes designed to edit the WAXY1 gene, among others. The CENH3 paternal haploid inducer lines were heterozygous (cenh3 + / -) for a 19 bp mutation in the coding sequence of cenh3 (produced by Cast 2a editing), which led to a frameshift and premature stop codon. While the CENH3 paternal haploid inducer lines segregate for this CENH3 paternal haploid inducer allele, the lines lack the matrilineal inducer allele. However, they carried the R1-SCM2 color marker in a homozygous condition that allowed for color-based haploid selection. Due to the CENH3 + / - status, the lines had a 2 to 12% HIR when outcrossed as females.
[0244] The donor plants used to produce immature embryos for transformation of 27241 were genotyped for the CENH3 mutation a gRNA cut site using the assay known as TaqMan® quantitative PCR assay 3895. The primers used in the assay were TCCTTGTTCCGTCTTTTGCAG (SEQ ID NO: 4) andAAGGCAAAAGGAGGGAACTGAT (SEQ ID NO: 5); the TaqMan® probe sequence was TACCTCGGCGACGCC (SEQ ID NO: 6). Transformation donor plants from the lines PlantHIe75 and PlantHIe78 were selected to not have the CENH3 mutation, whereas those plants from line PlantHIe77 were selected to have the CENH3 mutation. The rationale for this was to understand if the presence of the CENH3 knockout allele in a heterozygous condition impacted the transformation rate. The transformation rates of the resulting experiments were as follows (Table 9).
[0245] The TO events deriving from PlantHIe79 and PlantHIe80 were all wild type for the CENH3 mutation, but events from PlantHIe81 were a mixture of wild type (12) and heterozygous (8), proving that CENH3+ / - embryos are somewhat transformable, albeit less efficiently. However, the CENH3 mutation in a heterozygous condition had a significantly lower transformation rate (2.7%) and lower seed set (explants / ear) than the experiments usingCENH3 + / + donors. The advantage to producing a TO event is that one does not have to reintroduce the knockout allele and can simply self-pollinate a TO event, thus recovering T1 plants that are CENH3 + / - and homozygous for the CRISPR transgene. In contrast, a TO event that is CENH3 + / + needs to have the mutation reintroduced. Accordingly, many events from all four experiments which did not have the CENH3 mutation were maintained. To reintroduce the CENH3 mutation to the CENH3 + / + TOs from all three experiments, CENH3 (+ / -) plants from lines PlantHIe75 and PlantHIe78 were used to pollinate the TO event ears. This had the effect of reintroducing the CENH3 mutation to those CRISPR+ lines, but the resulting Fl offspring could only be hemizygous for the CRISPR transgene.
[0246] While those Fl offspring could theoretically be used for HI-Edit, the expected 1 to 1 segregation of the CRISPR machinery will be such that 50% of the haploid progeny would not have had a chance to be edited. Therefore, the events shown in Table 10 were advanced to the next generation, by selfing Casl2a+ (27241), CENH3 + / - plants (type A) and using those plants as pollen donors onto the ears of Casl2a+ (27241), CENH3 + / + (type B) plants. Through this method, seed from segregating for Cast 2a and CENH3 was produced. Then, HI-Edit donors (Casl2a-homozygous, CENH3 + / -) were selected in the next generation.
[0247] While producing the seed for the HI-Edit trial, Fl zygote assays were performed for some of the events using the type B plants (from Table 5) as females and using Tester 2 and Tester 8 as males. This assay was a reliable proxy for HI-Edit haploid editing rate, and thereforewas used as a means for selecting events with a higher HI-Edit efficiency. Crossing the type B plants from Table 8 as females with males being Tester 2 or Tester 8 led to the production of Fl seeds. These were germinated and leaf punches were taken from V2 stage seedlings for TaqMan® assays in order to identify the individuals that carried the Casl2a transgene and were edited at the gRNA target sites. Plants carrying Cast 2a that showed new editing at the Waxyl site (i.e., there was no or very little amplification of the “wild type” assay for Waxyl) were then subjected to PCR and NGS analysis of the Waxyl and other target sites. The parental edit(s) at the target site were determined and routinely found in the Fl offspring with -50% read abundance. The criterium for zygote editing was recovery of a new edit having a >30% read abundance in addition to the parental edit. With this criterium, the zygote editing rate averaged 30% at the Waxyl target.
[0248] The eight events in Table 11 were prioritized for Fl zygote analysis. That prioritization was based on consistent editing at the Waxy and other target sites seen in the TaqMan® data from the prior generation (represented in Table 8, supra). Note that unlike in Fl zygote tests where the matrilineal inducer was used as a male, in this Fl zygote test the CRISPR transgenic plants were females, because in CENH3 HI-Edit, the CRISPR machinery is coming from the female side of the cross (i.e., the egg cell). This is important because for the Fl zygote data to be a relevant proxy of HI-Edit rate, the direction of the cross needed to be the same as what will occur during the HI-Edit trial. Based on the information in Table 9 and the seed availability data shown in Table 5, event PlantHIe94, PlantHIe85, and PlantHIe84 were select for HI-Edit trials.
[0249] In the HI-Edit efficiency test, four testers representing diverse maize germplasm (one stiff stalk (Tester 8), two non-stiff stalk (Tester 2 and Tester 9), and one tropical line (Tester 10)) were crossed as males onto the T2 Casl2a-homozygous, CENH3 + / - knockout ears from the three selected events. Haploids were selected using the color marker assay and then were germinated and sampled for sequencing. The haploid editing rate, haploid induction rate, and seed set were thus measured. A heat treatment was provided to a subset of ears, using either the room temperature control or the heat pack method (wrapping heat packs around the ears).
[0250] For each of the 3 cenh3 + / - HI-Edit events in the crossing matrix, we aimed to produce 100 plants that were Casl2a-homozygous, cenh3 + / -. From the 100 plants for each event, we aimed to cross 25 by each of the four tester lines, and with each plant expected to make 1 or 2 ears, that means we aimed to harvest 30-40 pollinated ears per event, per tester. Each ear was expected to yield 20 to 70 haploids per ear and the pollinated ears would then be used for the three different heat treatments (hot room, heat pad, or control) which would be initiated immediately after pollination. These numbers would allow us to reach our goal to evaluate the haploid editing rate in more than 300 haploids from three events by four testers by three treatments.
[0251] Due to segregation distortion, the segregation of cenh3 + / - progeny plants from cenh3 + / - parents is about 10 to 40%. To obtain at least 100 plants that were Casl2a- homozygous, cenh3 + / -, at least 400 seeds from ears whose parents were Casl2a-HOM, cenh3 + / - and at least 1600 seeds from ears whose parents were Casl2a-HET, cenh3 + / - were needed.
[0252] For the event PlantHIe85, sufficient T2 seeds (more than 1000) were produced that were homozygous positive for Cast 2a (“Casl2a+-HOM”) and segregating for CENH3, resulting in at least 100 Casl2a-HOM, cenh3 + / - HI-Edit plants. For the event PlantHIe89, more than 1600 F2 seeds were planted to have sufficient plants for selection of 100 HI-Edit donors. These plants were strong and healthy.
[0253] In the HI-Edit efficiency test, the four tester lines were crossed as males onto ears from the three selected events (Casl2a-homozygous, cenh3 + / - HI-Edit donor ears). Induced ears were provided heat treatments (controls did not receive heat treatment) starting about 6 to 10 hours after pollination.
[0254] In the hot-room treatment, the plants were moved to a glasshouse room where the temperature was set to 39°C (compared to 26.6°C for the control room) for a duration of 24 hours. Plants were moved back to the control room after treatment. Ear temperatures were monitored with a VersaLog WF-TH 8-channel thermistor datalogger. The VersaLog datalogger was equipped with eight 10KOHM 3969K NTC thermistors (#MA100GG103BN, Amphenol Thermometries). One thermistor was used to monitor air temperature of the room. Additional thermistors were used to monitor temperatures of individual ears. A representative subsample of primary or secondary ears were monitored. These additional thermistors were installed to rest between the embryos and inner husk of the ear. To install an ear thermistor, a 3mm hole was pierced through the ear husk at the ear base and the thermistor was fed upward to the mid-section of ear. The ears in the hot room reached an average ear temperature of 29.98°C during the night and 32.08°C during the day.
[0255] In the heat-pad treatment, the ears were directly treated with a heat pad wrapped around the ear that was set to 34°C. Heating was regulated by a temperature controller (Drok XY-T01) and a hysteresis setting of 0.5°C. A 10K OHM 3969K NTC thermistor (#MA100GG103BN) was installed on the temperature controller. The temperature controller thermistor was inserted into the ear, resting between the embryos and inner husk mid-section of the ear. The heat pad consisted of two 14 centimeters by 5 centimeters electric heatingpads (ADAFRUIT item no. 4308), wired in parallel and covered in aluminum foiled tape. 5V DC power was supplied to the heating pad elements from the temperature controller. The heat pads were removed after 24 hours of heat treatment. The average internal temperature of the ears during this treatment measured 32.16°C during the night and 33.12°C during the day.The internal temperature of the ears was monitored independently with a VersaLog WF-TH 8-channel thermistor datalogger. A representative subsample of primary or secondary ears were monitored.
[0256] The induced ears were harvested approximately 15 days after pollination (DAP) and sent to the embryo rescue lab for processing. Total kernel counts were obtained to calculate the haploid induction rate (HIR) and the Haploid Editing Rate (HER). After extraction, the embryos were sown onto Murashige and Skoog (MS) media and placed in a Percival growth chamber with continuous light (150 pmol m2s ') at 28 °C for 24 hrs. Total embryos (seed set) per ear was tallied. The colorless putative haploid embryos were advanced, and the haploid induction rate was determined. Averaging across events and testers, the heat pad reduced the seed set per ear compared to control, resulting in fewer haploids per ear. The hot room modestly reduced seed set but also modestly increased the haploid induction rate, often leading to a higher average haploids per ear, depending on the specific event and cross.
[0257] In contrast to the mixed results for haploid yield, the haploid editing rate (HER) dramatically increased in both the hot room and heat pad treatments compared to control. On average, there was about a 3x increase in haploid editing rate in the hot room compared to the control (no heat), and a 4x increase in the heat pad compared to the control. This is observed in all event and tester combinations for both hot room and heat pad (Table 13 and Table 14, based on the embryos sent for genotyping).
[0258] This finding that the heat treatments dramatically improved the haploid editing rates was true for both the Waxyl and the UPL3 gRNAs. This was true not only in the groups of embryos sent directly to the 96-well blocks, but was also true in the plants that were germinated and sent to the greenhouse for leaf sampling (Table 15).* “HR” indicates Hot Room treatmentExample 5. Cas-UBA Fusion.
[0259] Most plant proteins are degraded by the ubiquitin / 26S proteasome pathway. Proteins destined for degradation are first covalently tagged with a ubiquitin (“UBI”). UBI receptors deliver ubiquitinated substrates to the 26S proteasome, which are degraded by the proteasome. See Dikic, I., Wakatsuki, S., & Walters, K. J. (2009). Ubiquitin-binding domains - from structures to functions. Nature Reviews Molecular Cell Biology, 10(10), 659-671. doi.org / 10.1038 / nrm2767. UBI receptors (UBL-UBA) are very stable and have long half-lives, which protect these receptors from proteasomal degradation by inhibiting multi-ubiquitin chain assembly or preventing the generation of initiation sites for degradation. A few UBA chimeric protein studies identified that the UBA domain can be used to enhance the stability of the target protein, prolong its half-life, and successfully improve the activity of the target protein. See, e.g., In-Cheol Jang, et al., The Plant Journal, 2012 r / Xuelian Zheng, et al., Frontiers in Plant Science, 2020.
[0260] We tested whether Casl2a stability could be improved by fusing the UBA2 domain of AtRAD23 and assessing whether it could increase HER (haploid editing rate) during HI- Edit. Two constructs, 27680 (encoding for LbCasl2a and gRNAs targeting ZmWxl and ZmG12) as the control and 28825 (encoding for LbCasl2a-UBA2 fusion and the same gRNAs) as the fusion were used to generate stable events. Editing rates were tested in the male and female inducer lines. In the male haploid inducer line NP3003RS, E0 editing rates were measured to prove functionality of the Casl2a-UBA2 fusion, and the E0-F1 zygote editing rate (ZER%) was measured to see if UBA2 fusion confers to higher ZER. El ZER and HER data was measured by crossing the Casl2a-homozygous plants as male to Tester 2 and Tester 8. From the female side, the El Casl2a-homozygous plants were used to cross with Tester 2 and Tester 8, and the ZER% of ZmWxl site was measured. No significant difference was seen atthe EO stage (88.2% for Casl2a versus 88.9% for Cash 12a-UBA2 fusion), showing the Casl2a- UBA2 fusion was functional. However, the E0-F1 ZER showed significant difference between Cast 2a alone and the Casl2a-UBA2 fusion.
[0261] Results show that when used as male, the Casl2a-UBA2 fusion significantly increased HER% and ZER% compared to Casl2a alone (27680). Haploid editing quality was not affected. Applying the heat treatment significantly increased ZER%, HER%, and haploidediting quality. See Table 18, Table 19. When used as female, the Casl2a-UBA2 fusion conferred about a 20% increase in ZER% across all the treatments. See Table 20.Table 18. Average haploid editing rate (HER %) as determined by NGS.REFERENCES1. Blomme, J., et al., “The heat is on: a simple method to increase genome editing efficiency in plants,” BMC Plant Biol. 2022 Mar 24;22(1): 142.2. Kurokawa, S., et al., “A Simple Heat Treatment Increases SpCas9-Mediated Mutation Efficiency in Arabidopsis,” Plant Cell Physiol (2021).3. Malzahn, A.A., et al., “Application of CRISPR-Casl2a temperature sensitivity for improved genome editing in rice, maize, and Arabidopsis,” BMC Biol. 17:9 (2019).4. Li, B., et al., “The application of temperature sensitivity CRISPR / LbCpfl (LbCasl2a) mediated genome editing in allotetraploid cotton (G. hirsutum) and creation of nontrans- genic, gossypol-free cotton,” Plant Biotechnol J. 19:221-3 (2021).5. An, Y., “Efficient Genome Editing in Populus Using CRISPR / Casl2a,” Front Plant Sci., Vol. 11, Article 593938 (2020); doi: 10.3389 / fpls.2020.593938.6. Milner, M.J., et al.. “Turning Up the Temperature on CRISPR: Increased Temperature Can Improve the Editing Efficiency of Wheat Using CRISPR / Cas9,” Front Plant Sci. 2020; Vol. 11, Article 583374 (2020); doi: 10.3389 / fpls.2020.583374.7. Huang, Y., et al., “HSFAla modulates plant heat stress responses and alters the 3D chromatin organization of enhancer-promoter interactions,” NATURE COMMUNICATIONS 14: 469 (2023); doi: 10.1038 / s41467-023-36227-3.8. Liang, Z., “Reorganization of the 3D chromatin architecture of rice genomes during heat stress,” BMC BIOL Vol. 19, Article 53 (2021); doi: 10.1186 / sl2915-021-00996- 4.9. Das, J.R. & S. Mathur, S., “HSFAla: the quarterback of heat stress response and 3D- chromatin organization,” TRENDS PLANT SCI. Vol 28, Issue 11, pp. 1198-1200 (2023); doi: 10.1016 / j.tplants.2023.07.008.10. Perrella, G., et al., “Epigenetic regulation of thermomorphogenesis and heat stress tolerance,” NEW PHYTOLOGIST Vol. 234, Issue 4, pp. 1144-1160 (2022); doi: 10.1111 / nph. l 7970.11. Wang, Z., et al., “A simple and highly efficient strategy to induce both paternal and maternal haploids through temperature manipulation,” NAT. PLANTS Vol. 9, Issue 5, pp. 699-705 (2023); doi.org / 10.1038 / s41477-023-01389-x.12. Maruthachalam, M. & Simon W L Chan, S.W.. “Haploid plants produced by centromere-mediated genome elimination.” Nature, Vol. 464, No. 7288, pp. 615-618 (2010); doi: 10.1038 / nature08842.13. Wang, N. et al. “Haploid induction by a maize cenh3 null mutant.” Science Advances, Vol. 7, No. 4, Article eabe2299 (2021); doi: 10.1126 / sciadv.abe2299.14. Jin, C., et al., "Heat stress promotes haploid formation during CENH3 -mediated genome elimination in Arabidopsis,” PLANT REPROD 36, 147-155 (2023); doi: 10.1007 / s00497-023-00457-8.15. Ahmadli, U., et al., “High temperature increases centromere-mediated genome elimination frequency and enhances haploid induction in Arabidopsis,” PLANT COMMUNICATIONS Vol. 4, Issue 3, Article 100507 (2023), doi:10.1016 / j.xplc.2022.100507.CONSTRUCTS & COMPONENTSConstruct 28291Construct 28294
[0262] All patents, patent publications, patent applications, journal articles, books, technical references, and the like discussed in the instant disclosure are incorporated herein by reference in their entirety for all purposes.
[0263] It is to be understood that the figures and descriptions of the disclosure have been simplified to illustrate elements that are relevant for a clear understanding of the disclosure. It should be appreciated that the figures are presented for illustrative purposes and not as construction drawings. Omitted details and modifications or alternative embodiments are within the purview of persons of ordinary skill in the art.
[0264] It can be appreciated that, in certain aspects of the disclosure, a single component may be replaced by multiple components, and multiple components may be replaced by a single component, to provide an element or structure or to perform a given function or functions. Except where such substitution would not be operative to practice certain embodiments of the disclosure, such substitution is considered within the scope of the disclosure.
[0265] The examples presented herein are intended to illustrate potential and specific implementations of the disclosure. It can be appreciated that the examples are intended primarily for purposes of illustration of the disclosure for those skilled in the art. There may be variations to these diagrams or the operations described herein without departing from thespirit of the disclosure. For instance, in certain cases, method steps or operations may be performed or executed in differing order, or operations may be added, deleted or modified.
[0266] Where a range of values is provided, it is understood that each intervening value, to the smallest fraction of the unit of the lower limit, unless the context clearly dictates otherwise, between the upper and lower limits of that range is also specifically disclosed. Any narrower range between any stated values or unstated intervening values in a stated range and any other stated or intervening value in that stated range is encompassed. The upper and lower limits of those smaller ranges may independently be included or excluded in the range, and each range where either, neither, or both limits are included in the smaller ranges is also encompassed within the technology, subject to any specifically excluded limit in the stated range. Where the stated range includes one or both of the limits, ranges excluding either or both of those included limits are also included.
[0267] In the foregoing description, numerous specific details are set forth to provide a more thorough understanding of the present invention. However, it will be apparent to one of skill in the art that the invention described in this disclosure may be practiced without one or more of these specific details. In other instances, well-known features and procedures well known to those skilled in the art have not been described in order to avoid obscuring the invention. Embodiments of the disclosure have been described for illustrative and not restrictive purposes. Although the present invention is described primarily with reference to specific embodiments, it is also envisioned that other embodiments will become apparent to those skilled in the art upon reading the present disclosure, and it is intended that such embodiments be contained within the present inventive methods. Accordingly, the present disclosure is not limited to the embodiments described above or depicted in the drawings, and various embodiments and modifications can be made without departing from the scope of the claims below.
Claims
WHAT IS CLAIMED IS:
1. A method of editing plant genomic DNA, comprising: a. providing an egg-donor plant comprising plant genomic DNA that is to be edited; b. pollinating the egg-donor plant with a pollen-donor plant, wherein the pollendonor plant expresses a DNA modification enzyme and, optionally, a guide nucleic acid; c. applying a heat treatment to the pollinated egg-donor plant of step b; and d. producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the egg-donor plant and does not comprise the genome of the pollen-donor haploid inducer plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and, optionally, the guide nucleic acid delivered by the pollen-donor haploid inducer plant.
2. A method of editing plant genomic DNA, comprising: a. providing a pollen-donor plant that expresses a DNA modification enzyme and, optionally, a guide nucleic acid; b. applying a heat treatment to pollen of the pollen-donor plant; c. pollinating an egg-donor plant with heat-treated pollen of the pollen-donor plant, wherein the egg-donor plant comprises the plant genomic DNA that is to be edited; and d. producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the egg-donor plant and does not comprise the genome of the pollen-donor plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and, optionally, the guide nucleic acid delivered by the pollen-donor plant.
3. The method of claim 1 or 2, wherein the pollen-donor plant is a haploid inducer plant.
4. The method of claim lor 2, wherein the pollen-donor plant is a maternal haploid inducer plant.
5. The method of claim 4, wherein the maternal haploid inducer plant comprises a knock-out mutation in a. MA IL gene.
6. A method of editing plant genomic DNA, comprising: a. providing a pollen-donor plant comprising plant genomic DNA that is to be edited; b. pollinating an egg-donor plant with pollen from the pollen-donor plant, wherein the egg-donor plant expresses a DNA modification enzyme and, optionally, a guide nucleic acid; c. applying a heat treatment to the pollinated egg-donor plant of step b; and d. producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the pollen-donor plant and does not comprise the genome of the egg-donor plant, and (ii) the genome of the haploid progeny has been modified by the DNA modification enzyme and, optionally, the guide nucleic acid delivered by the egg-donor plant.
7. A method of editing plant genomic DNA, comprising: a. providing a pollen-donor plant comprising plant genomic DNA that is to be edited; b. applying a heat treatment to pollen of the pollen-donor plant; c. pollinating an egg-donor plant with heat-treated pollen of the pollen-donor plant, wherein the egg-donor plant expresses a DNA modification enzyme and, optionally, a guide nucleic acid; and d. producing at least one edited haploid progeny, wherein (i) the haploid progeny comprises the genome of the pollen-donor plant and does not comprise the genome of the egg-donor plant, and (ii) the genome of the haploid progeny hasbeen modified by the DNA modification enzyme and the optional guide nucleic acid delivered by the egg-donor plant.
8. The method of claim 6 or 7, wherein the egg-donor plant is a haploid inducer plant.
9. The method of claim 6 or 7, wherein the egg-donor plant is a paternal haploid inducer plant.
10. The method of claim 9, wherein the paternal haploid inducer plant comprises a mutation in a CENH3 gene.
11. The method of claim 10, wherein the paternal haploid inducer plant is heterozygous for the mutation in the CENH3 gene.
12. The method of any one of claims 1-11, wherein at least one of the egg-donor plant or the pollen-donor plant is a maize plant.
13. The method of claim 12, wherein the maize plant is selected and / or derived from the lines Stock 6, RWK, RWS, UH400, NP2222RS, or NP2222.
14. The method of any one of claims 1-13, wherein the pollen-donor plant is a maize plant.
15. The method of any one of claims 1-13, wherein the egg-donor plant is a maize plant.
16. The method of any one of claims 1-15, wherein the DNA modification enzyme is a site-directed nuclease selected from the group consisting of a meganuclease (MN), a zinc-finger nuclease (ZFN), a transcription-activator like effector nuclease (TALEN), and a Cas nuclease.
17. The method of claim 16, wherein the Cas nuclease is a Type II Cas nuclease, a Type IV Cas nuclease, or a Type V Cas nuclease.
18. The method of claim 17, wherein the Type II Cas nuclease is a Cas9 nuclease, a Cas9 nickase, a nuclease-inactive Cas9, or a Cas9 fused to a heterologous domain.
19. The method of claim 18, wherein the Type V Cas nuclease is a Casl2a nuclease, a Cas 12a nickase, a nuclease-inactive Cas 12a, or a Cas 12a fused to a heterologous domain.
20. The method of any one of claims 1-19, wherein the guide nucleic acid is a guide RNA.
21. The method of any one of claims 1-20, wherein the haploid progeny is treated with a chromosome doubling agent, thereby creating an edited doubled haploid progeny.
22. The method of claim 21, wherein the chromosome doubling agent is colchicine, pronamide, dithipyr, trifluralin, or another known anti -microtubule agent.
23. The method of any one of claims 1-22, wherein the pollen-donor plant expresses a marker gene.
24. The method of claim 23, wherein the marker gene is selected from the group consisting of Rl, R1-SCM2, Rl-nj, GUS, PMI, PAT, GFP, RFP, CFP, Bl, CI, anthocyanin pigments, and any other marker gene.
25. An edited haploid plant produced by the method of any one of claims 1-24.
26. A progeny plant of the edited haploid plant of claim 25.