Recombinant maize b chromosome sequences and uses thereof
By integrating specific genome-modifying enzymes and target DNA onto the maize B chromosome, the problems of trait complexation and linkage redundancy risks in transgenic plants were solved, achieving efficient integration of transgenic traits onto the B chromosome and improving trait complexation efficiency and stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MONSANTO TECHNOLOGY LLC
- Filing Date
- 2016-09-30
- Publication Date
- 2026-05-08
AI Technical Summary
In existing technologies for transgenic plant breeding, as the number of traits increases, the risk of linkage redundancy of traits on chromosome A increases, and the resource and technical challenges increase, making it difficult to effectively utilize chromosome B to resolve trait complication and reduce the risk of linkage redundancy.
By integrating specific genome-modifying enzymes into the maize B chromosome and inserting target DNA, transgenic traits can be developed on the B chromosome using recombinant nucleic acid technology. This includes using site-specific genome-modifying enzymes such as endonucleases, recombinases, and transposases to integrate gene expression cassettes for insecticide resistance and herbicide tolerance.
This method enables efficient integration of target DNA onto chromosome B, reduces the risk of linkage redundancy on chromosome A, and improves the trait synergy efficiency and stability of transgenic plants.
Smart Images

Figure CN108513583B_ABST
Abstract
Description
[0001] Cross-references to related applications and incorporation of sequence listings
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 236,709, filed October 2, 2015; U.S. Provisional Patent Application No. 62 / 237,048, filed October 5, 2015; and U.S. Provisional Patent Application No. 62 / 240,770, filed October 13, 2015, all of which are incorporated herein by reference in their entirety. The sequence lists contained in the files “P34350US00_SEQ.text” (3,202,544 bytes (measured in MS Windows operating system), created on October 2, 2015, and filed on October 2, 2015 with U.S. Provisional Patent Application No. 62 / 236,709), “P34350US01_SEQ.txt” (3,202,584 bytes (measured in MS Windows operating system), created on October 5, 2015, and filed on October 5, 2015 with U.S. Provisional Patent Application No. 62 / 237,048), and “P34350US02_SEQ.txt” (3,655,665 bytes (measured in MS Windows operating system), created on October 13, 2015, and filed on October 13, 2015 with U.S. Provisional Patent Application No. 62 / 240,770) are incorporated herein by reference in their entirety. The sequence list in computer-readable form is submitted electronically with this application and is incorporated herein by reference in its entirety. The sequence list is contained in a file named 61653_ANNIV_ST25.txt, which is 3,640,410 bytes in size (measured in MS Windows operating system) and was created on September 30, 2016. Background Technology
[0003] B chromosomes are supernumerary chromosomes found in many organisms, and they differ from standard nuclear chromosomes (A chromosomes) in that they rarely carry active genes. These chromosomes are not essential for life.
[0004] The first transgenic plants were produced in the 1990s, resulting from the random insertion of transgenic DNA into the nucleus A chromosome (e.g., Roundup). (Soybeans). With the increasing number of transgenic events conferring individual traits (e.g., insect resistance, drought tolerance, herbicide tolerance, quality traits), plant breeders and farmers expect crops to possess a variety of traits. As the number of traits per plant increases, the resources and technological challenges required to produce breeding complexes of multiple traits present on chromosome A in superior germplasm increase accordingly. Furthermore, with increased traits, there is a risk of linkage carryover. In maize, trait complexation and the risk of chromosome A linkage carryover can be addressed by developing next-generation transgenic traits on chromosome B. Brief description of the attached diagram
[0006] Figure 1 Image of maize root tip cells stained with DAPI fluorescent dye. The arrows indicate the 20 B chromosomes in the nucleus of a cell from a single plant. The image is magnified 40 times.
[0007] Figure 2 Localization of maize B chromosome / A chromosome translocations using B chromosome primers. Control PCR amplicon regions are indicated: BCHR004 for the proximal region; BCHR006 for the proximal euchromatin region; and BCHR007 for the distal region. Regions of PCR amplification primers identified from the unique B chromosome sequence are indicated as primers SEQ ID NO:889+890; SEQ ID NO:891+892; and SEQ ID NO:893+894. The abbreviations on the B chromosome diagram are as follows: CK represents the centromere region of the B chromosome; PH represents the proximal heterochromatin region; PE1 / PE2 represents the proximal euchromatin region; DH1-4 represent the distal heterochromatin region; and DE represents the distal euchromatin region. Invention Overview
[0009] Several embodiments relate to a recombinant nucleic acid comprising: a nucleic acid sequence having at least 0.25 Kb, 0.5 Kb, 0.75 Kb, 1 Kb, 1.25 Kb, 1.5 Kb, 1.75 Kb, 2 Kb, 2.25 Kb, 2.5 Kb, 2.75 Kb, 3 Kb, 3.25 Kb, 3.5 Kb, 3.75 Kb, 4 Kb, 4.25 Kb, 4.5 Kb, 4.75 Kb, or 5 Kb having at least 90%, at least 91%, at least 92%, at least 93%, 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with a maize B chromosome sequence selected from the group consisting of SEQ ID NO: 1-126 and 128-888; and at least one target DNA, wherein at least one of the target DNAs is integrated into the maize B chromosome sequence to produce the recombinant nucleic acid. In some embodiments, the target DNA is integrated at double-strand breaks generated by one or more site-specific genome-modifying enzymes. In some embodiments, the target DNA encodes a site-specific genome-modifying enzyme. In some embodiments, the target DNA encodes one or more endonucleases, recombinases, transposases, helicases, or any combination thereof. In some embodiments, the target DNA comprises one or more gene expression cassettes selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA binding gene expression cassettes, selectable marker gene expression cassettes, RNAi construct expression cassettes, site-specific genome-modifying enzyme gene expression cassettes, or expression cassettes encoding one or more of CRIPR-related proteins, tracr RNA, and guide RNA.
[0010] Several embodiments relate to a maize plant containing recombinant nucleic acids, said recombinant nucleic acids comprising: selected from SEQ ID NO. A nucleic acid sequence comprising at least 0.25 Kb, 0.5 Kb, 0.75 Kb, 1 Kb, 1.25 Kb, 1.5 Kb, 1.75 Kb, 2 Kb, 2.25 Kb, 2.5 Kb, 2.75 Kb, 3 Kb, 3.25 Kb, 3.5 Kb, 3.75 Kb, 4 Kb, 4.25 Kb, 4.5 Kb, 4.75 Kb, or 5 Kb of the maize B chromosome sequence having at least 90%, at least 91%, at least 92%, at least 93%, 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity; and at least one target DNA, wherein at least one of the target DNAs is integrated into the maize B chromosome sequence to produce the recombinant nucleic acid. In some embodiments, the target DNA is integrated at double-strand breaks generated by one or more site-specific genome-modifying enzymes. In some embodiments, the target DNA encodes a site-specific genome-modifying enzyme. In some embodiments, the target DNA encodes one or more endonucleases, recombinases, transposases, helicases, or any combination thereof. In some embodiments, the target DNA comprises one or more gene expression cassettes selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA binding gene expression cassettes, selectable marker gene expression cassettes, RNAi construct expression cassettes, site-specific genome-modifying enzyme gene expression cassettes, or expression cassettes encoding one or more of CRIPR-related proteins, tracr RNA, and guide RNA.
[0011] Several embodiments relate to a recombinant B chromosome, said recombinant B chromosome comprising: selected from SEQ A nucleic acid sequence comprising at least 0.25 Kb, 0.5 Kb, 0.75 Kb, 1 Kb, 1.25 Kb, 1.5 Kb, 1.75 Kb, 2 Kb, 2.25 Kb, 2.5 Kb, 2.75 Kb, 3 Kb, 3.25 Kb, 3.5 Kb, 3.75 Kb, 4 Kb, 4.25 Kb, 4.5 Kb, 4.75 Kb, or 5 Kb of maize B chromosome sequence having at least 90%, at least 91%, at least 92%, at least 93%, 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity; and at least one target DNA, wherein at least one of the target DNAs is integrated into the maize B chromosome sequence to produce the recombinant nucleic acid. In some embodiments, the target DNA is integrated at double-strand breaks generated by one or more site-specific genome-modifying enzymes. In some embodiments, the target DNA encodes a site-specific genome-modifying enzyme. In some embodiments, the target DNA encodes one or more endonucleases, recombinases, transposases, helicases, or any combination thereof. In some embodiments, the target DNA comprises one or more gene expression cassettes selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA binding gene expression cassettes, selectable marker gene expression cassettes, RNAi construct expression cassettes, site-specific genome-modifying enzyme gene expression cassettes, or expression cassettes encoding one or more of CRIPR-related proteins, tracr RNA, and guide RNA.
[0012] Several embodiments relate to a method for manufacturing maize cells containing at least one target DNA integrated into the B chromosome, the method comprising selecting at least 0.25 Kb, 0.5 Kb, 0.75 Kb, 1 Kb, 1.25 Kb, 1.5 Kb, 1.75 Kb, 2 Kb, 2.25 Kb, 2.5 Kb, 2.75 Kb, 3 Kb, 3.25 Kb, 3.5 Kb, 3.75 Kb, 4 Kb, 4.25 Kb, 4.5 Kb, 4.75 Kb, or 5 Kb of a maize B chromosome sequence selected from the group consisting of SEQ ID NO: 1-126 and 128-888, having a content of at least 90%, at least 91%, at least 92%, or at least 90% of the target DNA. The target DNA exhibits 3%, 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity; selection of a site-specific genomic modification enzyme that specifically cleaves the target site; introduction of the site-specific genomic modification enzyme into the maize cell; introduction of the target DNA into the maize cell; integration of the target DNA into the target site; and selection of transgenic maize cells containing at least one of the target DNAs integrated into the B chromosome. In some embodiments, the target DNA is integrated at double-strand breaks generated by one or more site-specific genomic modification enzymes. In some embodiments, the target DNA encodes a site-specific genomic modification enzyme. In some embodiments, the target DNA encodes one or more endonucleases, recombinases, transposases, helicases, or any combination thereof. In some embodiments, the target DNA comprises one or more gene expression cassettes, wherein the gene expression cassettes are selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA binding gene expression cassettes, selectable marker gene expression cassettes, RNAi construct expression cassettes, site-specific genome modifying enzyme gene expression cassettes, or expression cassettes encoding one or more of CRIPR-related proteins, tracr RNA, and guide RNA.
[0013] Several embodiments relate to a method for providing maize cells with a site-specific genome-modifying enzyme, the method comprising integrating at least one site-specific genome-modifying enzyme into a component selected from SEQ ID NO. The maize B chromosome sequence comprising NO:1-126 and 128-888 consists of sequences having at least 0.25 Kb, 0.5 Kb, 0.75 Kb, 1 Kb, 1.25 Kb, 1.5 Kb, 1.75 Kb, 2 Kb, 2.25 Kb, 2.5 Kb, 2.75 Kb, 3 Kb, 3.25 Kb, 3.5 Kb, 3.75 Kb, 4 Kb, 4.25 Kb, 4.5 Kb, 4.75 Kb, or 5 Kb having at least 90%, at least 91%, at least 92%, at least 93%, 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity; and maize cells containing site-specific genomic modifying enzymes integrated into said B chromosome. In some embodiments, the site-specific genomic modifying enzyme specifically cleaves one or more target sites in chromosome A that integrates the target DNA. In some embodiments, the target DNA comprises one or more gene expression cassettes selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA binding gene expression cassettes, selectable marker gene expression cassettes, and RNAi construct expression cassettes. In some embodiments, offspring containing the target DNA but not chromosome B containing the site-specific genomic modifying enzyme are selected. In some embodiments, the chromosome B contains a negative selection marker.
[0014] Several embodiments relate to a recombinant nucleic acid comprising a 1 kb nucleic acid sequence having at least 90% sequence identity with a maize B chromosome sequence selected from the group consisting of SEQ ID NO: 1-126 and 128-888; and target DNA integrated into the maize B chromosome sequence to produce the recombinant nucleic acid. In some embodiments, the recombinant nucleic acid comprises target DNA integrated near a target site of a site-specific genomic modifying enzyme. In some embodiments, the site-specific genomic modifying enzyme is an endonuclease. In some embodiments, the site-specific genomic modifying enzyme is a recombinase. In some embodiments, the site-specific genomic modifying enzyme is a transposase. In some embodiments, the site-specific genomic modifying enzyme is a helicase. In some embodiments, the site-specific genomic modifying enzyme is any combination of endonuclease, recombinase, transposase, and helicase. In some embodiments, the target site is specific to the maize B chromosome sequence. In another embodiment, the target site is specific to a maize B chromosome sequence selected from one or more maize B chromosome sequences comprising SEQ ID NO: 1-126 and 128-888. In some embodiments, the site-specific genome-modifying enzyme is an endonuclease selected from megabase nucleases, zinc finger nucleases, transcription activator-like effector nucleases (TALENs), Argonaute nucleases, DNA-directed recombinases, DNA-directed endonucleases, RNA-directed recombinases, RNA-directed endonucleases, type I CRISPR-Cas systems, type II CRISPR-Cas systems, and type III CRISPR-Cas systems. In some embodiments, the endonuclease is selected from the group consisting of: Cpf1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, and Csf4 nucleases. In some embodiments, the site-specific genome-modifying enzyme is an RNA-directed recombinase. In some embodiments, the site-specific genome-modifying enzyme is a fusion protein comprising a recombinase and a CRISPR-related protein.In some embodiments, the recombinase is a tyrosine recombinase or a serine recombinase attached to a DNA recognition motif. In some embodiments, the recombinase is Cre recombinase, Flp recombinase, Tnp1 recombinase, PhiC31 integrase, R4 integrase, or TP-901 integrase. In some embodiments, the transposase is a DNA transposase attached to a DNA-binding domain. In some embodiments, the target DNA does not encode a peptide. In some embodiments, the target DNA does encode a peptide. In some embodiments, the target DNA comprises one or more gene expression cassettes. In some embodiments, the gene expression cassette is selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA-binding gene expression cassettes, selectable marker gene expression cassettes, RNAi construct expression cassettes, site-specific genome-modifying enzyme gene expression cassettes, expression cassettes encoding recombinant guide RNA encoding RNA-guided endonucleases, or expression cassettes encoding recombinant DNA guides. In some embodiments, the target DNA is integrated into the target site of the maize B chromosome via homology-directed repair integration. In some embodiments, the target DNA is integrated into the target site of the maize B chromosome via non-homologous end conjugation integration. In some embodiments, the target DNA and / or the maize B chromosome target site sequence are modified during the integration of the target DNA into the target site of the maize B chromosome sequence. In some embodiments, the recombinant nucleic acid is present in a maize plant, a maize plant part, a maize seed, or a maize plant cell.
[0015] Several embodiments relate to a recombinant nucleic acid comprising a 1 kb nucleic acid sequence having at least 90% sequence identity with a maize B chromosome sequence selected from the group consisting of SEQ ID NO: 1-126 and 128-888; and target DNA integrated into the maize B chromosome sequence to produce the recombinant nucleic acid. In some embodiments, the recombinant nucleic acid comprises target DNA integrated between a pair of target sites of a site-specific genomic modifying enzyme. In some embodiments, the site-specific genomic modifying enzyme is an endonuclease. In some embodiments, the site-specific genomic modifying enzyme is a recombinase. In some embodiments, the site-specific genomic modifying enzyme is a transposase. In some embodiments, the site-specific genomic modifying enzyme is a helicase. In some embodiments, the site-specific genomic modifying enzyme is any combination of an endonuclease, a recombinase, a transposase, and a helicase. In some embodiments, the target site is specific to the maize B chromosome sequence. In another embodiment, the target site is specific to a maize B chromosome sequence selected from one or more maize B chromosome sequences comprising SEQ ID NO: 1-126 and 128-888. In some embodiments, the site-specific genome-modifying enzyme is an endonuclease selected from megabase nucleases, zinc finger nucleases, transcription activator-like effector nucleases (TALENs), arginases, DNA-directed recombinases, DNA-directed endonucleases, RNA-directed recombinases, RNA-directed endonucleases, type I CRISPR-Cas systems, type II CRISPR-Cas systems, and type III CRISPR-Cas systems. In some embodiments, the endonuclease is selected from the group consisting of: Cpf1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, and Csf4 nucleases. In some embodiments, the site-specific genome-modifying enzyme is an RNA-directed recombinase. In some embodiments, the site-specific genome-modifying enzyme is a fusion protein comprising a recombinase and a CRISPR-related protein.In some embodiments, the recombinase is a tyrosine recombinase or a serine recombinase attached to a DNA recognition motif. In some embodiments, the recombinase is Cre recombinase, Flp recombinase, Tnp1 recombinase, PhiC31 integrase, R4 integrase, or TP-901 integrase. In some embodiments, the transposase is a DNA transposase attached to a DNA-binding domain. In some embodiments, the target DNA does not encode a peptide. In some embodiments, the target DNA does encode a peptide. In some embodiments, the target DNA comprises one or more gene expression cassettes. In some embodiments, the gene expression cassette is selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA-binding gene expression cassettes, selectable marker gene expression cassettes, RNAi construct expression cassettes, site-specific genome-modifying enzyme gene expression cassettes, expression cassettes encoding recombinant guide RNA encoding RNA-guided endonucleases, or expression cassettes encoding recombinant DNA guides. In some embodiments, the target DNA is integrated into the target site of the maize B chromosome via homology-directed repair integration. In some embodiments, the target DNA is integrated into the target site of the maize B chromosome via non-homologous end conjugation integration. In some embodiments, the target DNA and / or the maize B chromosome target site sequence are modified during the integration of the target DNA into the target site of the maize B chromosome sequence. In some embodiments, the recombinant nucleic acid is present in a maize plant, a maize plant part, a maize seed, or a maize plant cell.
[0016] Several embodiments relate to a recombinant nucleic acid comprising a 1 kb nucleic acid sequence having at least 90% sequence identity with a maize B chromosome sequence selected from the group consisting of SEQ ID NO: 1-126 and 128-888; and target DNA integrated into the maize B chromosome sequence to produce the recombinant nucleic acid. In some embodiments, the recombinant nucleic acid comprises one or more of the target DNA inserted near two or more target sites of a site-specific genomic modifying enzyme. In some embodiments, the site-specific genomic modifying enzyme is an endonuclease. In some embodiments, the site-specific genomic modifying enzyme is a recombinase. In some embodiments, the site-specific genomic modifying enzyme is a transposase. In some embodiments, the site-specific genomic modifying enzyme is a helicase. In some embodiments, the site-specific genomic modifying enzyme is any combination of endonuclease, recombinase, transposase, and helicase. In some embodiments, the target site is specific to the maize B chromosome sequence. In another embodiment, the target site is specific to maize B chromosome sequences selected from one or more of the group consisting of SEQ ID NO: 1-126 and 128-888. In some embodiments, the site-specific genome-modifying enzyme is an endonuclease selected from megabase nucleases, zinc finger nucleases, transcription activator-like effector nucleases (TALENs), lag nucleases, DNA-directed recombinases, DNA-directed endonucleases, RNA-directed recombinases, RNA-directed endonucleases, type I CRISPR-Cas systems, type II CRISPR-Cas systems, and type III CRISPR-Cas systems. In some embodiments, the endonuclease is selected from the group consisting of: Cpf1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, and Csf4 nucleases. In some embodiments, the site-specific genome-modifying enzyme is an RNA-directed recombinase. In some embodiments, the site-specific genome-modifying enzyme is a fusion protein comprising a recombinase and a CRISPR-related protein.In some embodiments, the recombinase is a tyrosine recombinase or a serine recombinase attached to a DNA recognition motif. In some embodiments, the recombinase is Cre recombinase, Flp recombinase, Tnp1 recombinase, PhiC31 integrase, R4 integrase, or TP-901 integrase. In some embodiments, the transposase is a DNA transposase attached to a DNA-binding domain. In some embodiments, two or more of the target DNA molecules are identical. In some embodiments, two or more of the target DNA molecules are dissimilar. In some embodiments, the target DNA does not encode a peptide. In some embodiments, the target DNA does encode a peptide. In some embodiments, the target DNA includes one or more gene expression cassettes. In some embodiments, the gene expression cassette is selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA binding gene expression cassettes, selectable marker gene expression cassettes, RNAi construct expression cassettes, site-specific genome modifying enzyme gene expression cassettes, expression cassettes encoding recombinant guide RNAs encoding RNA-guided endonucleases, or expression cassettes encoding recombinant DNA guides. In some embodiments, the target DNA is integrated into the target site on the maize B chromosome via homology-directed repair integration. In some embodiments, the target DNA is integrated into the target site on the maize B chromosome via non-homologous end conjugation integration. In some embodiments, two or more of the maize B chromosome target site sequences each contain integrated target DNA to generate two or more recombinant sequences. In some embodiments, the two or more recombinant sequences are located on the same B chromosome. In some embodiments, the two or more recombinant sequences are located on different B chromosomes. In some embodiments, the two or more recombinant sequences are located on different B chromosomes, which recombine during cell division to generate new megabase loci on the new B chromosome. In some embodiments, the target DNA and / or the target site sequence of the maize B chromosome are modified during the integration of the target DNA into the target site of the maize B chromosome sequence. In some embodiments, the recombinant nucleic acid is present in a maize plant, a maize plant part, a maize seed, or a maize plant cell.
[0017] Several embodiments relate to a recombinant nucleic acid comprising a 1 kb nucleic acid sequence having at least 90% sequence identity with a maize B chromosome sequence selected from the group consisting of SEQ ID NO: 1-126 and 128-888; and target DNA integrated into the maize B chromosome sequence to produce the recombinant nucleic acid. In some embodiments, the recombinant nucleic acid comprises one or more of the target DNA inserted near two or more target sites of a site-specific genomic modifying enzyme. In some embodiments, the site-specific genomic modifying enzyme is an endonuclease. In some embodiments, the site-specific genomic modifying enzyme is a recombinase. In some embodiments, the site-specific genomic modifying enzyme is a transposase. In some embodiments, the site-specific genomic modifying enzyme is a helicase. In some embodiments, the site-specific genomic modifying enzyme is any combination of endonuclease, recombinase, transposase, and helicase. In some embodiments, the target site is specific to the maize B chromosome sequence. In another embodiment, the target site is specific to maize B chromosome sequences selected from one or more of the group consisting of SEQ ID NO: 1-126 and 128-888. In some embodiments, the site-specific genome-modifying enzyme is an endonuclease selected from megabase nucleases, zinc finger nucleases, transcription activator-like effector nucleases (TALENs), lag nucleases, DNA-directed recombinases, DNA-directed endonucleases, RNA-directed recombinases, RNA-directed endonucleases, type I CRISPR-Cas systems, type II CRISPR-Cas systems, and type III CRISPR-Cas systems. In some embodiments, the endonuclease is selected from the group consisting of: Cpf1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, and Csf4 nucleases. In some embodiments, the site-specific genome-modifying enzyme is an RNA-directed recombinase. In some embodiments, the site-specific genome-modifying enzyme is a fusion protein comprising a recombinase and a CRISPR-related protein.In some embodiments, the recombinase is a tyrosine recombinase or a serine recombinase attached to a DNA recognition motif. In some embodiments, the recombinase is Cre recombinase, Flp recombinase, Tnp1 recombinase, PhiC31 integrase, R4 integrase, or TP-901 integrase. In some embodiments, the transposase is a DNA transposase attached to a DNA-binding domain. In some embodiments, two or more of the target DNA molecules are identical. In some embodiments, two or more of the target DNA molecules are dissimilar. In some embodiments, the target DNA does not encode a peptide. In some embodiments, the target DNA does encode a peptide. In some embodiments, the target DNA includes one or more gene expression cassettes. In some embodiments, the gene expression cassette is selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA binding gene expression cassettes, selectable marker gene expression cassettes, RNAi construct expression cassettes, site-specific genome modifying enzyme gene expression cassettes, expression cassettes encoding recombinant guide RNAs encoding RNA-guided endonucleases, or expression cassettes encoding recombinant DNA guides. In some embodiments, the target DNA is integrated into the target site on the maize B chromosome via homology-directed repair integration. In some embodiments, the target DNA is integrated into the target site on the maize B chromosome via non-homologous end conjugation integration. In some embodiments, two or more of the maize B chromosome target site sequences each contain integrated target DNA to generate two or more recombinant sequences. In some embodiments, the two or more recombinant sequences are located on the same B chromosome. In some embodiments, the two or more recombinant sequences are located on different B chromosomes. In some embodiments, the two or more recombinant sequences are located on different B chromosomes, which recombine during cell division to generate new megabase loci on the new B chromosome. In some embodiments, the target DNA and / or the target site sequence of the maize B chromosome are modified during the integration of the target DNA into the target site of the maize B chromosome sequence. In some embodiments, the recombinant nucleic acid is present in a maize plant, a maize plant part, a maize seed, or a maize plant cell.
[0018] Several embodiments relate to a method for manufacturing maize plant cells containing recombinant nucleic acids having target DNA integrated into the B chromosome, the method comprising: (a) selecting a maize B chromosome genomic locus that is compatible with the DNA of SEQ ID NO. (a) Selecting a site-specific genome-modifying enzyme that specifically binds to and cleaves the target site in the maize B chromosome genomic locus; (c) Introducing the site-specific genome-modifying enzyme into maize plant cells; (d) Optionally, when the site-specific genome-modifying enzyme in step (c) is an RNA-guided endonuclease, introducing guide RNA into maize plant cells; or (e) Optionally, when the site-specific genome-modifying enzyme in step (c) is a DNA-guided endonuclease, introducing guide DNA into maize plant cells; (f) Introducing the target DNA into the maize plant cells; (g) Integrating the target DNA proximal to the target site in the maize B chromosome genomic locus; and (g) Selecting transgenic plant cells containing the target DNA integrated into the target site in the maize B chromosome genomic locus. Several embodiments relate to a method for manufacturing maize plant cells containing recombinant nucleic acids having target DNA integrated into the B chromosome, the method comprising: (a) selecting a maize B chromosome genomic locus that is compatible with the DNA of SEQ ID NO. (a) Selecting one or more site-specific genomic modification enzymes that specifically bind to and cleave the pair of target sites in the maize B chromosome genomic locus; (c) Introducing the site-specific genomic modification enzyme into maize plant cells; (d) Optionally, introducing guide RNA into maize plant cells when one of the site-specific genomic modification enzymes in step (c) is an RNA-guided endonuclease; or (e) Optionally, introducing guide DNA into maize plant cells when one of the site-specific genomic modification enzymes in step (c) is a DNA-guided endonuclease; or (e) Introducing the target DNA into the maize plant cells; (f) Integrating the target DNA between the pair of target sites of the site-specific genomic modification enzyme in the maize B chromosome genomic locus; and (g) Selecting transgenic plant cells containing at least one target DNA integrated into at least one target site of the maize B chromosome genomic locus.Several embodiments relate to a method for manufacturing maize plant cells containing recombinant nucleic acids having at least one target DNA integrated into the B chromosome, the method comprising: (a) selecting a maize B chromosome genomic locus that is compatible with the DNA of SEQ ID NO. (a) Selecting one or more site-specific genome-modifying enzymes that specifically bind to and cleave two or more target sites in the maize B chromosome genomic locus; (c) Introducing the site-specific genome-modifying enzyme into maize plant cells; (d) Optionally, introducing guide RNA into maize plant cells when one of the site-specific genome-modifying enzymes in step (c) is an RNA-guided endonuclease; or (e) Optionally, introducing guide DNA into maize plant cells when one of the site-specific genome-modifying enzymes in step (c) is a DNA-guided endonuclease; or (e) Introducing the at least one target DNA into the maize plant cells; (f) Integrating the at least one target DNA into the vicinity of two or more target sites of the site-specific genome-modifying enzyme in the maize B chromosome genomic locus; and (g) Selecting transgenic plant cells containing at least one target DNA integrated into at least one target site of the maize B chromosome genomic locus. Several embodiments relate to a method for manufacturing maize plant cells containing recombinant nucleic acids having two or more target DNAs integrated into the B chromosome, the method comprising: (a) selecting a maize B chromosome genomic locus that is compatible with the DNA of SEQ ID NO. (a) Selecting one or more site-specific genome-modifying enzymes that specifically bind to and cleave two or more target sites in the maize B chromosome genomic locus; (c) Introducing the site-specific genome-modifying enzyme into maize plant cells; (d) Optionally, introducing guide RNA into maize plant cells when one of the site-specific genome-modifying enzymes in step (c) is an RNA-guided endonuclease; or (e) Optionally, introducing guide DNA into maize plant cells when one of the site-specific genome-modifying enzymes in step (c) is a DNA-guided endonuclease; or (e) Introducing the two or more target DNAs into the maize plant cells; (f) Integrating the two or more target DNAs into the vicinity of two or more target sites of the site-specific genome-modifying enzymes in the maize B chromosome genomic locus; and (g) Selecting transgenic plant cells containing at least one target DNA integrated into at least one target site of the maize B chromosome genomic locus. In some implementations, the site-specific genome-modifying enzyme is an endonuclease.In some embodiments, the site-specific genome-modifying enzyme is a recombinase. In some embodiments, the site-specific genome-modifying enzyme is a transposase. In some embodiments, the site-specific genome-modifying enzyme is a helicase. In some embodiments, the site-specific genome-modifying enzyme is any combination of endonuclease, recombinase, transposase, and helicase. In some embodiments, the target site is specific to a maize B chromosome sequence. In another embodiment, the target site is specific to a maize B chromosome sequence selected from one or more maize B chromosome sequences comprising SEQ ID NO: 1-126 and 128-888. In some embodiments, the site-specific genome-modifying enzyme is an endonuclease selected from megabase nucleases, zinc finger nucleases, transcription activator-like effector nucleases (TALENs), lag nucleases, DNA-directed recombinases, DNA-directed endonucleases, RNA-directed recombinases, RNA-directed endonucleases, type I CRISPR-Cas systems, type II CRISPR-Cas systems, and type III CRISPR-Cas systems. In some embodiments, the endonuclease is selected from the group consisting of: Cpf1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, and Csf4 nucleases. In some embodiments, the site-specific genome-modifying enzyme is an RNA-directed recombinase. In some embodiments, the site-specific genome-modifying enzyme is a fusion protein comprising a recombinase and a CRISPR-related protein. In some embodiments, the recombinase is a tyrosine recombinase attached to a DNA recognition motif or a serine recombinase attached to a DNA recognition motif. In some embodiments, the recombinase is Cre recombinase, Flp recombinase and Tnp1 recombinase, PhiC31 integrase, R4 integrase, or TP-901 integrase. In some embodiments, the transposase is a DNA transposase attached to a DNA-binding domain. In some embodiments, two or more of the target DNA molecules are identical. In some embodiments, two or more of the target DNA molecules are dissimilar. In some embodiments, the target DNA does not encode a peptide. In some embodiments, the target DNA does encode a peptide.In some embodiments, the target DNA comprises one or more gene expression cassettes. In some embodiments, the gene expression cassette is selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA binding gene expression cassettes, selectable marker gene expression cassettes, RNAi construct expression cassettes, site-specific genome-modifying enzyme gene expression cassettes, expression cassettes encoding recombinant guide RNA of RNA-guided endonucleases, or expression cassettes encoding recombinant DNA guides. In some embodiments, the site-specific genome-modifying enzyme is stably transformed into the maize plant cells. In some embodiments, the site-specific genome-modifying enzyme is transiently transformed into the maize plant cells. In some embodiments, the site-specific genome-modifying enzyme is constitutively expressed in the maize plant cells. In some embodiments, the site-specific genome-modifying enzyme is expressed in the maize plant cells under the control of a regulatory promoter. In some embodiments, the regulatory promoter is a heat shock promoter, a tissue-specific promoter, or a chemically induced promoter. In some embodiments, the target DNA is integrated into the target site of the maize B chromosome via homology-directed repair integration. In some embodiments, the target DNA is integrated into the target site of the maize B chromosome via non-homologous end conjugation integration. In some embodiments, two or more maize B chromosome target site sequences each contain integrated target DNA to generate two or more recombinant sequences. In some embodiments, the two or more recombinant sequences are located on the same B chromosome. In some embodiments, the two or more recombinant sequences are located on different B chromosomes. In some embodiments, the two or more recombinant sequences are located on different B chromosomes, which recombine during cell division to generate new megabase loci on the new B chromosome. In some embodiments, the target DNA and / or the maize B chromosome target site sequence are modified during the integration of the target DNA into the target site of the maize B chromosome sequence. In some embodiments, the recombinant nucleic acid is present in a maize plant, a maize plant part, a maize seed, or a maize plant cell.
[0019] In some embodiments, probes specific to maize B chromosome sequences selected from the group consisting of SEQ ID NO: 1-126 and 128-888 have been developed. In some embodiments, the probes are generated by PCR amplification. In some embodiments, the PCR primers are selected from the group consisting of SEQ ID NO: 889-894. In some embodiments, the PCR primers are used to locate maize B chromosome / A chromosome translocations. In some embodiments, the maize B chromosome sequence is used to develop maize B chromosome markers. In some embodiments, the maize B chromosome sequence is used to identify side-joint sequences of transgenes integrated into the B chromosome.
[0020] Several embodiments relate to a method for identifying sequences specific to the maize B chromosome, the method comprising one or more of the following steps: (1) identifying maize plants with a high copy number of B chromosome; (2) preparing DNA from tissue collected from the plant in step 1; (3) preparing a Forse plasmid library from the DNA in step 2 and aliquoting it into the library; (4) sequencing the collected Forse plasmid library; and (5) using bioinformatics analysis to extract sequences specific to the maize B chromosome. In some embodiments, the maize plants with a high copy number of B chromosome are selected from plants having 1, 2, 3, 4, 5, 6, 7, 8, 8, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more B chromosomes. In some embodiments, the bioinformatics analysis includes: (1) extracting high-quality, usable sequences from the Fors plasmid pool; (2) combining the sequence reads from step 1 into individual library contigs; (3) analyzing and masking the individual library contigs from step 2 for: (a) known non-repetitive maize A chromosome sequences from the reference maize genome assembly; (b) known maize mitochondrial DNA sequences in the reference maize genome assembly; (c) known maize chloroplast DNA sequences in the reference maize genome assembly; and (d) E. coli and / or Fors plasmid vector sequences; (4) analyzing the sequences pooled in step (3) relative to each other to identify B chromosome repetitive sequences or filter out maize A chromosome sequences not in the maize reference genome assembly; and (5) comparing the resulting sequences to identify the longest representative sequence of each contig, and assembling the longest contig from the overlapping contigs. In other embodiments, single-molecule real-time sequencing technology is used to generate long read sequences from maize B chromosome DNA to produce high-quality sequential assemblies of genome sequences. In other implementations, maize B chromosome contig assemblies analyzed from the Fors plasmid library are combined with long read sequences to produce assemblies with unique B chromosome sequences.
[0021] Detailed description of the invention
[0022] Despite decades of study on the maize B chromosome, a comprehensive assembly of unique maize B chromosome sequences has yet to be obtained. This disclosure provides unique maize B chromosome sequences that can be used to develop probes for B chromosome mapping, develop B chromosome genomic markers, and identify loci for the specific integration of target DNA sites into the B chromosome.
[0023] Several implementation schemes described herein relate to methods for identifying unique maize B chromosome sequences. Such unique maize B chromosome sequences can be used to develop methods for site-specific integration of target DNA at a single locus, integration of two or more target DNAs at independent loci, integration of two or more target DNAs at a single locus, or integration of two or more target DNAs at two or more loci. Integration of two or more target DNAs at independent but linked loci is also referred to as a “complex.”
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. Where a term is provided in the singular, the inventors also contemplate the plural form of that term describing aspects of this disclosure. Where there is any discrepancy between the terms and definitions used and those used in references incorporated herein by reference, the terms used herein shall have the definitions given herein. Other technical terms used have their common meaning in their field of application, as exemplified in various technical dictionaries, such as "The American..." The *Science Dictionary* (Editors of the American Heritage Dictionaries, 2011, Houghton Mifflin Harcourt, Boston and New York), the *McGraw-Hill Dictionary of Scientific and Technical Terms* (6th edition, 2002, McGraw-Hill, New York), or the *Oxford Dictionary of Biology* (6th edition, 2008, Oxford University Press, Oxford and New York) are referenced. The inventors do not wish to be bound by any mechanism or mode of action. References to these are provided solely for illustrative purposes.
[0025] Unless otherwise indicated, the practice of this disclosure employs conventional techniques of biochemistry, chemistry, molecular biology, microbiology, cell biology, genomics, and biotechnology, which are within the capabilities of those skilled in the art. See Green and Sambrook, *Molecula Cloning: A Laboratory Manual*, 4th ed. (2012); *Current Protocols in Molecular Biology* (edited by F. Mausubel et al., (1987)); *Methods in Enzymology* (Academic Press, Inc.): *PCR 2: A Practical Approach* (edited by M.J. MacPherson, B.D. Hames, and G.G. Taylor, (1995)); *Antibodies, A Laboratory Manual* (edited by Harlow and Lane, (1988); *Animal Cell Culture* (edited by R.R. Freshney, (1987)); *Recombinant Protein Purification: Princesses and Methods*, 18-1142-75, GE Healthcare Life. Sciences; CN Stewart, A. Touraev, V. Citovsky, T. Tzfira, eds. (2011) PLANT TRANSFORMATION TECHNOLOGIES (Wiley-Blackwell); and RH Smith (2013) PLANT TISSUECULTURE.TECHNIQUES AND EXPERIMENTS (Academic Press, Inc.).
[0026] All references cited in this article are included in their entirety as citations.
[0027] Unless the context clearly indicates otherwise, as used herein, the singular forms “a,” “an,” and “the” include plural referents. Thus, by way of example, references to “plant,” “the plant,” or “a single plant” also include multiple plants; and, depending on the context, the use of the term “plant” may also include genetically similar or identical offspring of that plant; the use of the term “nucleic acid” implicitly includes, in effect, numerous copies of that nucleic acid molecule; similarly, the term “probe” optionally (and typically) covers numerous similar or identical probe molecules.
[0028] As used herein, “plant” means a whole plant or a cell or tissue culture derived from a plant, including any of the following: a whole plant, a plant part or organ (e.g., leaf, stem, root, etc.), plant tissue, seed, plant cell, and / or its offspring. Offspring plants can be derived from any generation, such as F1, F2, F3, F4, F5, F6, F7, etc. A plant cell is the biological cell of a plant, taken from a plant or derived from a cell taken from a plant through culture. Plant parts include harvestable parts and parts that can be used to propagate offspring plants. Plant parts that can be used for propagation include, for example, but not limited to: seeds; fruits; cuttings; seedlings; tubers; and rhizomes. Harvestable parts of a plant can be any useful part of the plant, including, for example, but not limited to: flowers; pollen; seedlings; tubers; leaves; stems; fruits; seeds; and roots. As used herein, plant cells include protoplasts and protoplasts with cell walls. Plant cells can be protoplasts, gamete-producing cells, or cells or cell collections that can regenerate into a whole plant.
[0029] As used herein, the term “about” indicates that the value includes the inherent variation of the error of the method used to determine the value or the variation present between experiments.
[0030] As used herein, “corn” and “maize” are used interchangeably to refer to maize (Zea mays L.) and include all plant varieties that can be bred from maize, including wild maize species.
[0031] As used in this article, rye refers to rye (Secale cereale).
[0032] In one respect, the maize plant or seeds provided in this disclosure are maize. In another respect, the maize plant or seeds provided in this disclosure are cultivated maize (Zea mays ssp. mays). In yet another respect, the maize plant or seeds provided herein are domesticated strains or varieties.
[0033] As used herein, the term "B chromosome" refers to a supernumerary chromosome. B chromosomes, along with their normal diploid complement, have been found in cells. In maize, there are 20 normal diploid complement chromosomes. The normal chromosome can be referred to as the "A chromosome." B chromosomes are optional, not essential, for normal development. When two types of B chromosomes are present in a single plant, they pair up during prophase I of meiosis, and recombination may occur. B chromosomes do not pair with A chromosomes.
[0034] In one aspect, the method of this disclosure incorporates the target DNA into a superchromosome. In some embodiments, the method of this disclosure incorporates the target DNA into a maize B chromosome sequence selected from the group consisting of SEQ ID NO: 1-126 and 128-888.
[0035] According to certain aspects of this disclosure, one or more B chromosomes can be transferred to progeny plants (e.g., via haploid-induced hybridization with B chromosomes retained), allowing complete transformation into new varieties in single crosses. In other aspects, B chromosomes can be transferred to other species, thereby allowing the examination of target DNA or multiple target DNAs in other crops. For example, B chromosome transfer to oats has been demonstrated (Koo et al., Genome Research 21(6):908-914, 2011) and maize chromosome transfer to wheat (Comeau et al., Plant Science 81(1):117-125, 1992).
[0036] In some cases, such as in maize and rye, B chromosomes possess an "accumulation mechanism," allowing them to be passed on at frequencies exceeding Mendelian frequencies. For example, in maize, the sister chromatids of the B chromosome fail to separate during the second pollen (first generation) division. As a result, both sister chromatids are passed to one pollen cell, while the other pollen cell receives neither. This process is called nondisjunction, meaning that a plant with only a single B chromosome can pass on two B chromosomes to the next generation when used as a male. Such processes may be necessary during phenotypic introgression because they allow individuals that are homozygous (as opposed to hemizygous) in terms of the megabase loci carried on the B chromosome to revert to their original state upon backcrossing, provided that the B chromosome is passed on from the pollen.
[0037] Nondisjunction requires the presence of a specific portion of the B chromosome. A trans-acting element at the tip of the long arm and a cis-acting element near the centromere are needed. Minimal deletions at the tip of the long arm of the B chromosome are recoverable, and the resulting B chromosome does not exhibit nondisjunction. In some embodiments of this disclosure, for example, such deletion variants of the B chromosome may be needed for the purpose of delivering megabase loci to obtain commercial traits. In other embodiments, nondisjunction may be required, for example, to transiently deliver one or more site-specific genomic modifying enzymes that modify one or more target sequences in the A genome to produce one or more gene edits or transgenic insertions, and the B chromosome containing the one or more site-specific genomic modifying enzymes is lost in offspring.
[0038] Several embodiments involve the rapid transfer of haploid induction to new lines using the B chromosome. Because the B chromosome may be retained at a low percentage, haploid induction induced by the B chromosome may allow this effect to be transferred to other lines via single crosses. This simplifies the generation of new haploid-induced lines with desired agronomic or genetic properties. Genetic components disclosed in U.S. Patent Application 62 / 375,618 (titled Compositions and Methods for Plant Haploid Induction, filed August 16, 2016) can be incorporated into the unique B chromosome sequence described herein to produce plants containing a haploid-inducible B chromosome (HI-B chromosome). The disclosure of U.S. Patent Application 62 / 375,618 is incorporated herein by reference in its entirety. Other haploid-inducing genes, such as CENH3-based transgenic genes (Kelliher, T et al., "Maternal Haploids Are Preferentially Induced by CENH3-tailswap Transgenic Complementation in Maize", Frontiers in Plant Science 7:414 (2016)), can be incorporated into the unique B chromosome sequence described herein to produce plants containing the HI-B chromosome. To transfer haploid induction to new lines, lines containing the HI-B chromosome are crossed with the desired lines, and the progeny are screened for those that are haploid and retain the HI-B chromosome.
[0039] In one respect, the B chromosome provided herein is the maize B chromosome. In another respect, the maize B chromosome sequence or maize B chromosome genomic loci provided herein are selected from the group consisting of SEQ ID NO: 1-126 and 128-888. In one respect, the B chromosome provided herein undergoes nondisjunction. In another respect, the B chromosome provided herein undergoes nondisjunction in pollen cells. In another respect, the B chromosome provided herein accumulates in a non-Mendelian manner. In another respect, the B chromosome provided herein is truncated. In another respect, the B chromosome provided herein may contain one or more target DNAs. In another respect, the B chromosome provided herein can be used in any of the methods provided herein. In one respect, the B chromosome provided herein is heterologous. In another respect, the B chromosome provided herein is monovalent during meiosis. In another respect, the B chromosome provided herein pairs with a second B chromosome during meiosis. In another respect, the B chromosome provided herein pairs with a second B chromosome during meiosis, and the first B chromosome recombines with the second B chromosome to produce a new B chromosome. In one aspect, the B chromosome provided herein contains at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten target sites for site-specific genomic modifying enzymes. In another aspect, the B chromosome provided herein includes a translocation between the B chromosome and the A chromosome, wherein the B chromosome includes the centromere of the B chromosome. In another aspect, the first B chromosome and the second B chromosome provided herein have at least 50%, at least 60%, at least 70%, at least 80%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity.
[0040] On the one hand, the plant cells provided herein may contain one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more B chromosomes.
[0041] The methods described herein can be applied to extract unique B chromosome sequences of sufficient length to provide gene targeting at multiple sites and enable the development of B chromosome markers. For example, unique sequences can be used for site-specific genome modification applications, including site-specific integration of target DNA. The methods described herein can be applied to any B chromosome-containing genome using pooled bacterial artificial chromosomes (BACs), Forse plasmid sequencing strategies, or single-molecule real-time sequencing strategies.
[0042] The sequences disclosed herein can also be used to create gene-editing tools that impart altered transmissible properties to truncated B chromosomes, which could be used in single-cross trait integration strategies. The sequences disclosed herein can also be used to identify polymorphisms between different B chromosome origins, and these polymorphisms can be used to construct gene recombination maps of the B chromosome. The sequences disclosed herein can also be used as FISH probes or for other physical mapping strategies (BA translocation mapping) to determine the location of sequence contigs relative to each other and other landmarks on the B chromosome (centromeres, telomeres, disclosed B repeat sequences, etc.).
[0043] As used in this article, “recombinant sequence” and “recombinant nucleic acid sequence” are used interchangeably.
[0044] As used in this article, "insert" and "integrate" are used interchangeably.
[0045] In some embodiments, the target DNA may be a transgene. In some embodiments, the target DNA may be a DNA molecule in the form of a "homologous transgene," which refers to a DNA sequence derived from the crop itself or from a hermaphroditic donor plant. In some embodiments, the target DNA may contain one or more "transgenes," wherein the transgene is a DNA sequence that is not naturally present on the maize B chromosome.
[0046] In one aspect, this disclosure provides one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more recombinant nucleic acid sequences, said recombinant nucleic acid sequences comprising at least 25 base pairs, at least 50 base pairs, at least 100 base pairs, at least 250 base pairs, at least 500 base pairs, at least 1000 base pairs (1Kb), at least 1500 base pairs (1.5Kb), at least 2000 base pairs (2Kb), at least 2500 base pairs (2.5Kb), at least 3000 base pairs (... 3Kb), at least 3500 base pairs (3.5Kb), at least 4000 base pairs (4Kb), at least 4500 base pairs (4.5Kb), at least 5000 base pairs (5Kb), at least 5500 base pairs (5.5Kb), at least 6000 base pairs (6Kb), at least 6500 base pairs (6.5Kb), at least 7000 base pairs (7Kb), at least 7500 base pairs (7.5Kb), at least 8000 base pairs (8Kb), at least 8500 base pairs (8.5Kb), at least 9000 base pairs (9Kb), at least 9500 base pairs (9.5Kb), or at least 10,000 base pairs (10Kb) and selected from SEQ The sequences comprising ID NOs 1-126 and 128-888 have nucleic acid sequences with at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99.5%, or 100% sequence identity; and one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more target DNAs that can be integrated into the maize B chromosome nucleic acid sequence to produce the recombinant nucleic acid sequence. Alternatively, the recombinant nucleic acid sequences provided herein are integrated into target sites of site-specific genomic modifying enzymes, wherein the target sites are specific to the maize B chromosome DNA sequence. On the other hand, the recombinant nucleic acid sequence provided herein is integrated between two target sites for site-specific genomic modifying enzymes, wherein the target sites are specific to the maize B chromosome DNA sequence.In another aspect, one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, or twenty or more target DNAs are integrated into one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, or twenty or more target sites of site-specific genomic modifying enzymes, wherein the target sites are specific to the maize B chromosome DNA sequence. In one aspect, the recombinant nucleic acid provided herein comprises target DNA encoding a peptide. In another aspect, the recombinant nucleic acid provided herein comprises target DNA that does not encode a peptide. In some embodiments, the recombinant nucleic acid provided herein comprises target DNA encoding one or more site-specific DNA modifying enzymes. In some embodiments, the recombinant nucleic acid provided herein comprises target DNA encoding one or more guide RNAs and / or tracr RNAs. In some embodiments, the recombinant nucleic acid provided herein comprises target DNA encoding siRNA. In some embodiments, the recombinant nucleic acid provided herein comprises target DNA encoding miRNA.
[0047] In one aspect, this disclosure provides one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more recombinant nucleic acid sequences, said recombinant nucleic acid sequences comprising one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more target DNAs integrated into a B chromosome. On the other hand, this disclosure provides one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more recombinant nucleic acid sequences comprising one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more target DNAs integrated into two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more B chromosomes.
[0048] On the one hand, the target DNA provided herein is modified during integration into the B chromosome sequence. On the other hand, the target DNA provided herein is not modified during integration into the B chromosome. On the one hand, the B chromosome sequence provided herein is modified when the target DNA provided herein is integrated into it. On the other hand, the B chromosome sequence provided herein is not modified when the target DNA provided herein is integrated into it. In some embodiments, the target DNA comprises at least 5 base pairs, at least 10 base pairs, at least 15 base pairs, at least 20 base pairs, at least 25 base pairs, at least 50 base pairs, at least 100 base pairs, at least 125 base pairs, at least 150 base pairs, at least 200 base pairs, at least 250 base pairs, at least 500 base pairs, at least 1000 base pairs (1Kb), at least 1500 base pairs (1.5Kb), at least 2000 base pairs (2Kb), at least 2500 base pairs (2.5Kb), at least 3000 base pairs (3Kb), at least 3500 base pairs (3.5Kb), and selected from SEQ ID. The sequences comprising NO:1-126 and 128-888 contain portions of nucleic acid sequences having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, at least 99.5%, or 100% sequence identity to facilitate homologous recombination between the B chromosome and the target DNA.
[0049] In some implementations, the target DNA may comprise a modified nucleic acid sequence. As used herein, a "modified" nucleic acid sequence includes one or more nucleotide insertions, substitutions, deletions, duplications, or inversions.
[0050] As used herein, “target nucleic acid,” “target DNA,” or “donor” is defined as a nucleic acid / DNA sequence selected for site-specific targeted insertion into the maize genome. The target nucleic acid can have any length, such as between 2 and 50,000 nucleotides (or any integer value between or above that length) or between about 1,000 and 5,000 nucleotides (or any integer value between that length). The target DNA may include one or more gene expression cassettes encoding gene sequences that are actively transcribed and / or translated. In some embodiments, the target DNA may contain a multinucleotide sequence or an entire gene that does not contain a functional gene expression cassette (e.g., may contain only regulatory sequences such as promoters), or may not contain any gene expression elements that can be identified or any gene sequence that is actively transcribed. The target DNA may optionally contain analytical domains. After the target DNA is integrated into the maize genome, the integrated sequence is referred to as “integrated target DNA” or “inserted target DNA.” Furthermore, the target DNA may be linear or circular, and may be single-stranded or double-stranded. It can be delivered to cells as naked nucleic acid, as a complex with one or more delivery agents (such as liposomes, poloxamer, protein-encapsulated T-chains, etc.), or contained in bacterial or viral delivery agents (such as Agrobacterium tumefaciens or geminiviruses, respectively).
[0051] On the one hand, the target DNA provided herein may contain at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten genes. On the other hand, the target DNA provided herein may not contain genes. Without limitation, the genes provided herein may include insecticide resistance genes, herbicide tolerance genes, nitrogen use efficiency genes, water use efficiency genes, nutrient quality genes, DNA binding genes, selectable marker genes, RNAi constructs, site-specific genome modifying enzyme genes, recombinant guide RNAs of the CRISPR / Cas9 system, geminivirus-based expression cassettes, or plant virus expression vector systems. On the one hand, the genes provided herein contain promoter nucleic acid sequences. On the other hand, the structural genes provided herein do not contain promoter nucleic acid sequences.
[0052] Examples of suitable genes of agronomic interest envisioned in this invention will include, but are not limited to, genes relating to disease, insect, or pest tolerance; herbicide tolerance; genes relating to quality improvements such as yield, fortification, environmental or stress tolerance; or any desired variation in plant physiology, growth, development, morphology, or plant product (including starch yield) (US Patent Nos. 6,538,181, 6,538,179, 6,538,178, 5,750,876, 6,476,295); improved oil yield (US Patent Nos. 6,444,876, 6,426,447, 6,380,462); high oil yield (US Patent Nos. 6,495,739, 5,608,149, 6,483,008, 6,476,295); and improved fatty acid content (US Patent Nos. 6,495,739, 5,608,149, 6,483,008, 6,476,295). Patent numbers 6,828,475, 6,822,141, 6,770,465, 6,706,950, 6,660,849, 6,596,538, 6,589,767, 6,537,750, 6,489,461, 6,459,018); high protein yield (US Patent No. 6,380,466); fruit ripening (US Patent No. 5, 512,466); enhanced animal and human nutrition (US Patent Nos. 6,723,837, 6,653,530, 6,5412,59, 5,985,605, 6,171,640); or biopolymers (US Patent Nos. RE37,543, 6,228,623, 5,958,745 and US Patent Publication No. US20030028917). Other benefits include environmental stress resistance (US Patent No. 6,072,103); pharmaceutical peptides and secretory peptides (US Patent Nos. 6,812,379, 6,774,283, 6,140,075, 6,080,560); improved processing properties (US Patent No. 6,476,295); improved digestibility (US Patent No. 6,531,648); low raffinose content (US Patent No. 6,166,292); and industrial enzyme production (US Patent No. 6,166,292). U.S. Patent No. 5,543,576; improved flavor (U.S. Patent No. 6,011,199); nitrogen fixation (U.S. Patent No. 5,229,114); hybrid seed production (U.S. Patent No. 5,689,041); fiber production (U.S. Patent Nos. 6,576,818, 6,271,443, 5,981,834, 5,869,720); and biofuel production (U.S. Patent No. 5,998,700). As those skilled in the art will understand from this disclosure, any of these or other genetic elements, methods, and transgenes may be used in this disclosure.
[0053] The target DNA provided herein may also include sequences encoding other sequences such as messenger RNA (mRNA). The mRNA generated from the nucleic acid molecules of this disclosure may contain a 5' untranslated (5'-UTR) leader sequence. This sequence may be derived from a promoter selected to express the gene and may be specifically modified to increase or decrease mRNA translation. The 5'-UTR may also be derived from viral RNA, suitable eukaryotic genes, or synthetic gene sequences. Such "enhancer" sequences may be needed to increase or alter the translation efficiency of the resulting mRNA. This invention is not limited to constructs where the untranslated region is derived from a 5'-UTR accompanying a promoter sequence. Instead, the 5'-UTR sequence may be derived from an unrelated promoter or gene (see, for example, U.S. Patent No. 5,362,865). Examples of untranslated leader sequences include the heat shock protein leader sequence of maize and petunia (US Patent No. 5,362,865), the leader sequence of plant virus coat protein, the leader sequence of plant ribulose diphosphate carboxylase, GmHsp (US Patent No. 5,659,122), PhDnaK (US Patent No. 5,362,865), AtAnt1, TEV (Carrington and Freed, Journal of Virology, (1990) 64:1590-1597), and AGRtu.nos (GenBank accession number V00087; Bevan et al., Nucleic Acids Research (1983) 11:369-385). Other genetic components that could enhance gene expression or influence its transcription or translation have also been envisioned as genetic components.
[0054] The target DNA provided herein may also include the 3' untranslated region (3'-UTR) of a gene. The provided 3'-UTR may contain a transcription terminator or an element with equivalent function, and a polyadenylation signal that functions in plants to induce the addition of polyadenylated nucleotides to the 3' end of an RNA molecule. The DNA sequence is referred to herein as the transcription termination region. Effective polyadenylation of mRNA requires the region. RNA polymerase transcribes the encoded DNA sequence through the site of polyadenylation. Examples of suitable 3' regions are (1) 3' transcription untranslated regions containing polyadenylation signals of Agrobacterium Ti plasmid genes, such as the alpha-lipoic acid synthase (NOS; Fraley et al., Proceedings of the National Academy of Sciences, USA (1983) 80:4803-4807) gene; and (2) small subunits of plant genes, such as the soybean storage protein gene and the ribulose-1,5-bisphosphate carboxylase (ssRUBISCO) gene. The preferred example of the 3' region comes from the 3' region of the pea ssRUBISCO E9 gene (European Patent Application 0385962).
[0055] In one aspect, the target DNA provided herein may encode one or more site-specific genome-modifying enzymes. In some embodiments, the one or more site-specific genome-modifying enzymes are selected from endonucleases, recombinases, transposases, helicases, or any combination thereof. In another aspect, the target DNA provided herein may encode one or more of the following: megabase nucleases, zinc finger nucleases, transcription activator-like effector nucleases (TALENs), lag nucleases, DNA-directed recombinases, DNA-directed endonucleases, RNA-directed recombinases, RNA-directed endonucleases, type I CRISPR-Cas systems, type II CRISPR-Cas systems, and type III CRISPR-Cas systems. On one hand, the target DNA provided in this article can encode one or more of the following: Cpf1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, and Csf4 nucleases. In one respect, the target DNA provided herein can encode a dCas9-recombinase fusion protein. In another respect, the target DNA provided herein can encode one or more of the following: tyrosine recombinase, serine recombinase, Cre recombinase, Flp recombinase, Tnp1 recombinase, PhiC31 integrase, R4 integrase, or TP-901 integrase. In yet another respect, the target DNA provided herein can encode one or more of tracr RNA and / or guide RNA.
[0056] On one hand, the target DNA provided herein can encode selectable, screenable, or scoreable marker genes. These genetic components are also referred to herein as functional genetic components because they produce products that function in the identification of transformed plants or have agronomical utility. DNA acting as a selection or screening device can function in regenerative plant tissues to produce compounds that confer resistance to otherwise toxic compounds. Many screenable or selectable marker genes are known in the art and are used in this disclosure. Target genes used as selectable, screenable, or scoreable markers will include, in particular, β-glucuronidase (GUS); green fluorescent protein (GFP); luciferase (LUC); and markers that confer resistance to antibiotics such as kanamycin (Dekeyser et al., Plant). Physiology (1989) 90:217-223) or genes encoding tolerance to spectinomycin (e.g., spectinomycin aminoglycoside adenosyltransferase (aadA); US Patent No. 5,217,902); genes encoding enzymes conferring tolerance to herbicides such as glyphosate (e.g., 5-enolpyruvylshikimate-3-phosphate synthase (EPSPS): Della-Cioppa et al., Bio / Technology (1987) 5:579-584) ); US Patent No. 5,627,061; US Patent No. 5,633,435; US Patent No. 6,040,497; US Patent No. 5,094,945; WO04074443 and WO04009761; glyphosate oxidoreductase (GOX; US Patent No. 5,463,175); glyphosate decarboxylase (WO05003362 and US Patent Application 20040177399; or glyphosate N-acetyltransferase (GAT): Cast Le et al., Science (2004) 304:1151-1154; US Patent Application 20030083480), [unclear text - likely related to patents and patents], ... 193A), sulfonyl herbicides (e.g., acetylhydroxyl synthase or acetyllactone synthase that is resistant to acetyllactone synthase inhibitors such as sulfonylureas, imidazolinones, triazolopyrimidines, pyrimidinyloxybenzoates and isopyridines; (US Patent Nos. 6,225,105, 5,767,366, 4,761,373, 5,633,437, 6,613,963, 5,013,659, 5,141,870, 5,378,824, 5,605,011);Encoding ALS, GST-II), bispyribac- or glufosinate ... 275,957), atrazine (encoding GST-III), dicamba (dicamba monooxygenase; US Patent Application Publications 20030115626, 20030135879) or dicamba (modified acetyl-CoA carboxylase) (US Patent No. 6,414,222) conferring tolerance to cyclohexanedione (dimethomorph) and aryloxyphenoxypropionate (flupyridine). Other selection procedures can also be implemented, including positive selection mechanisms (e.g., using the E. coli manA gene, allowing growth in the presence of mannose) and dual selection (e.g., simultaneous use of spectinomycin and glufosinate or spectinomycin and dicamba), and will still fall within the scope of this invention.
[0057] On the one hand, the selectable markers provided in this paper are positive selection markers. Positive selection markers confer an advantage on transformed cells. On the other hand, the target DNA provided in this paper includes selectable marker genes that confer antibiotic resistance or herbicide resistance. On the other hand, the selectable markers provided in this paper are negative selection markers. On the other hand, the selectable markers provided in this paper are both positive and negative selection markers. The negative selectable markers provided in this paper can be lethal or non-lethal negative selectable markers. For example, non-lethal negative selectable marker genes can be any of those listed in U.S. Publication No. 2004-0237142, such as GGPP synthase, GA 2-oxidase gene sequences, isopentenyltransferase (IPT), CKI1 (cytokinin-independent 1), ESR-2, ESR1-A, auxin-producing genes such as indole-3-acetic acid (IAA), iaaM, iaah, roLABC, genes causing overexpression of ethylene biosynthetic enzymes, VP1 gene, AB13 gene, LEC1 gene, and Bas1 gene. Non-lethal negative selectable marker genes can be included on any nucleic acid molecule provided herein. The non-lethal negative selectable marker genes provided herein are genes that cause overexpression of a class of enzymes that use substrates in the gibberellic acid (GA) biosynthesis pathway but do not produce biologically active GA. On the other hand, the nucleic acid molecules provided herein include non-lethal negative selectable marker genes, such as the phytoene synthase gene (crtB) from Erwinia spp.
[0058] On the one hand, the target DNA provided in this paper includes promoters. Promoters contain the nucleotide base sequence that instructs RNA polymerase to associate with DNA and initiate transcription into mRNA, using one strand of the DNA as a template to produce the corresponding complementary strand of RNA. On the one hand, the promoters provided in this paper are constitutive promoters. On the other hand, the promoters provided in this paper are regulatory promoters. Furthermore, the regulatory promoters provided in this paper are thermal shock promoters, tissue-specific promoters, or chemically inducible promoters.
[0059] Many promoters active in plant cells have been described in the literature. Such promoters include, but are not limited to, the lipoic acid synthase (NOS) and octopine synthase (OCS) promoters carried on the Ti plasmid of Agrobacterium tumefaciens, cauliflower mosaic virus promoters such as the cauliflower mosaic virus (CaMV) 19S and 35S promoters and the Scrophularia mosaic virus (FMV) 35S promoter, as well as the enhanced CaMV 35S promoter (e35S). Many other plant gene promoters regulated by environmental, hormonal, chemical, and / or developmental signals can also be used to express heterologous genes in plant cells, including promoters regulated by, for example, the following: (1) heat (Callis et al., Plant Physiology, (1988) 88: 965-968); (2) light (e.g., pea RbcS-3A promoter, Kuhlemeier et al., Plant Cell, (1989) 1: 471-478; maize RbcS promoter, Schaffner et al., Plant Cell (1991) 3: 997-1012); (3) hormones, such as abscisic acid (Marcotte et al., Plant Cell, (1989) 1: 969-976); (4) trauma (e.g., Siebertz et al., Plant Cell, (1989) 961-968); or other signals or chemicals. Tissue-specific promoters are also known.
[0060] In one respect, the target DNA provided herein includes at least one, at least two, at least three, at least four, at least five, at least six, at least seven, at least eight, at least nine, or at least ten promoters. In another respect, the target DNA provided herein does not include promoters. In yet another respect, the promoters provided herein may be part of a gene.
[0061] As described below, it is preferred that the selected specific promoter be capable of eliciting expression sufficient to produce an effective amount of the target gene product. Examples of such promoters include, but are not limited to, U.S. Patent No. 6,437,217 (maize RS81 promoter), U.S. Patent No. 5,641,876 (rice actin promoter), U.S. Patent No. 6,426,446 (maize RS324 promoter), U.S. Patent No. 6,429,362 (maize PR-1 promoter), U.S. Patent No. 6,232,526 (maize A3 promoter), U.S. Patent No. 6,177,611 (constitutive maize promoter), U.S. Patent Nos. 5,322,938, 5,352,605, 5,359,142 and 5,530,196 (35S promoter), and U.S. Patent No. 6,433,252. (Maize L3 oil body protein promoter), US Patent No. 6,429,357 (rice actin 2 promoter and rice actin 2 intron), US Patent No. 5,837,848 (root-specific promoter), US Patent No. 6,294,714 (photoinducible promoter), US Patent No. 6,140,078 (salt-inducible promoter), US Patent No. 6,252,138 (pathogen-inducible promoter), US Patent No. 6,175,060 (phosphorus deficiency-inducible promoter), US Patent No. 6,635,806 (γ-coixol promoter), and US Patent Application Serial No. 09 / 757,089 (maize chloroplast aldolase promoter).Other usable promoters include the lipoic acid synthase (NOS) promoter (Ebert et al., 1987), the octopus alkaloid synthase (OCS) promoter (which is carried on the tumor-inducing plasmid of Agrobacterium tumefaciens), cauliflower mosaic virus promoters such as the cauliflower mosaic virus (CaMV) 19S promoter (Lawton et al., Plant Molecular Biology (1987) 9:315-324), the CaMV 35S promoter (Odell et al., Nature (1985) 313:810-812), the Scrophularia mosaic virus 35S promoter (US Patent Nos. 6,051,753, 5,378,619), the sucrose synthase promoter (Yang and Russell, Proceedings of the National Academy of Sciences, USA (1990) 87:4144-4148), and the R gene complex promoter (Chandler ...7) 9:315-324), the CaMV 35S promoter (Odell et al., Nature (1987) Cell (1989) 1:1175-1183) and the promoters of chlorophyll a / b binding protein genes, PC1SV (US Patent No. 5,850,019) and AGRtu.nos (GenBank Register V00087; Depicker et al., Journal of Molecular and Applied Genetics (1982) 1:561-573; Bevan et al., 1983).
[0062] Promoter heterozygotes can be constructed to enhance transcriptional activity (US Patent No. 5,106,739) or to combine desired transcriptional activity, inducibility, and tissue- or developmental specificity. Promoters that function in plants include, but are not limited to, inducible promoters, viral promoters, synthetic promoters, constitutive promoters, time-regulated promoters, space-regulated promoters, and space-time-regulated promoters. Other tissue-enhancing, tissue-specific, or developmentally regulated promoters are also known in the art and are envisioned to be useful in practicing this invention.
[0063] The promoters used in the nucleic acid molecules and transformation vectors provided in this invention can be modified as needed to affect their control characteristics. Promoters can be obtained by means of linking to an operon region, random or controlled mutagenesis, etc. Furthermore, promoters can be modified to include multiple "enhancer sequences" to assist in improving gene expression.
[0064] The nucleic acid molecules of the present invention can be incorporated into any suitable plant transformation plasmid or vector containing selectable or screenable markers and related regulatory elements as described, along with one or more nucleic acids encoded by structural genes.
[0065] Site-specific genome-modifying enzymes
[0066] As used herein, the term "double-strand break inducer" refers to any agent that can induce double-strand breaks (DSBs) on DNA molecules. In some embodiments, the double-strand break inducer is a site-specific genome-modifying enzyme.
[0067] As used herein, the term "site-specific genome-modifying enzyme" refers to any enzyme that can modify nucleotide sequences in a site-specific manner. In this disclosure, site-specific genome-modifying enzymes include endonucleases, recombinases, transposases, helicases, and any combination thereof.
[0068] Several embodiments involve facilitating recombination by providing a site-specific genome-modifying enzyme. As used herein, the term "site-specific enzyme" refers to any enzyme that can modify a nucleotide sequence in a sequence-specific manner. In some embodiments, recombination is facilitated by providing a single-strand break inducer. In some embodiments, recombination is facilitated by providing a double-strand break inducer. In some embodiments, recombination is facilitated by providing a strand separation inducer. In one aspect, the site-specific genome-modifying enzyme is selected from endonucleases, recombinases, transposases, helicases, or any combination thereof. In some embodiments, recombination occurs between chromosome B. In some embodiments, recombination occurs between chromosome B and chromosome A. In some embodiments, recombination occurs between a target DNA as described herein and a unique chromosome B sequence.
[0069] Several embodiments involve facilitating the integration of one or more target DNAs by providing a site-specific genome-modifying enzyme. As used herein, the term "site-specific enzyme" refers to any enzyme that can modify a nucleotide sequence in a sequence-specific manner. In some embodiments, the integration of one or more target DNAs is facilitated by providing a single-strand break inducer. In some embodiments, the integration of one or more target DNAs is facilitated by providing a double-strand break inducer. In some embodiments, the integration of one or more target DNAs is facilitated by providing a strand separation inducing agent. In one aspect, the site-specific genome-modifying enzyme is selected from endonucleases, recombinases, transposases, helicases, or any combination thereof.
[0070] On one hand, the endonuclease is selected from megabase nucleases, zinc finger nucleases (ZFN), transcription activator-like effector nucleases (TALEN), agnucleases (non-limiting examples of agnuclease proteins include TtAgo from *Thermus thermophilus*, PfAgo from *Pyrococcus furiosus*, NgAgo from *Natronobacterium gregoryi*), and RNA-directed nucleases such as CRISPR-associated nucleases (non-limiting examples of CRISPR-associated nucleases include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc...). 1. Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, their homologues or their modified forms).
[0071] In some embodiments, the site-specific genome-modifying enzyme is a dCas9-Fok1 fusion protein. In another embodiment, the site-specific genome-modifying enzyme is a dCas9-recombinase fusion protein. As used herein, "dCas9" refers to a Cas9 endonuclease protein having one or more amino acid mutations that render the Cas9 protein inactive but retain RNA-directed site-specific DNA binding. As used herein, "dCas9-recombinase fusion protein" is a dCas9 protein fused to dCas9 in a manner that allows the recombinase to catalyze DNA.
[0072] In some embodiments, the site-specific genome-modifying enzyme is a recombinase. Non-limiting examples of recombinases include tyrosine recombinases attached to DNA recognition motifs provided herein, selected from the group consisting of Cre recombinase, Gin recombinase, Flp recombinase, and Tnp1 recombinase. In one aspect, the Cre or Gin recombinases provided herein are tethered to a zinc finger DNA-binding domain, a TALE DNA-binding domain, or a Cas9 nuclease. In another aspect, serine recombinases attached to DNA recognition motifs provided herein are selected from the group consisting of PhiC31 integrase, R4 integrase, and TP-901 integrase. In yet another aspect, DNA transposases attached to DNA-binding domains provided herein are selected from the group consisting of TALE-piggyBac and TALE-Mutator.
[0073] Site-specific genome-modifying enzymes such as megabase nucleases, ZFN, TALEN, agnuclease proteins (non-restrictive examples of agnuclease proteins include *Thermophilus thermophilus* agnuclease (TtAgo), *PfAgo* agnuclease, *NgAgo* agnuclease, their homologues or modified forms), RNA-directed nucleases (non-restrictive examples of RNA-directed nucleases include CRISPR-related nucleases such as Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, C Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, their homologues or modified forms) and engineered RNA-directed nucleases (RGN) induce genomic modifications, such as double-stranded DNA breaks (DSBs) or single-stranded DNA breaks, at target sites on the genomic sequence, which are then repaired by the natural processes of homologous recombination (HR) or non-homologous end joining (NHEJ). Then, sequence modification occurs at the cleaved site, which may include deletions or insertions that cause gene disruption in the case of NHEJ, or integration of exogenous sequences through homologous recombination.
[0074] In one aspect of this disclosure, site-specific genomic modifying enzymes are selected to induce genomic modifications in one, several, or many individual target sequences of the maize B chromosome sequence provided herein. Following exposure to the site-specific genomic modifying enzyme, the resulting recombinant nucleic acid can be identified using various methods, including sequencing, PCR amplification, analytical analysis, or other molecular methods for detecting recombinant nucleic acid sequences. Site-specific genomic modifying enzymes can be expressed in plants, resulting in one or more genomic modifications within a genomic locus and screening for molecular changes in the resulting progeny.
[0075] Any of the target DNAs provided herein can be integrated into a target site on the B chromosome sequence by introducing the target DNA and the provided site-specific genome-modifying enzyme. Any method provided herein can utilize any of the site-specific genome-modifying enzymes provided herein.
[0076] ZFN
[0077] Zinc finger nucleases (ZFNs) are synthetic proteins characterized by engineered zinc finger DNA-binding domains fused to the cleavage domain of a FokI restriction endonuclease. ZFNs can be designed to cleave segments of almost any length of double-stranded DNA by modifying the zinc finger DNA-binding domain. ZFNs form dimers from monomers fused to the non-specific DNA-cleaving domain of a FokI endonuclease, engineered to bind to target DNA sequences.
[0078] The DNA-binding domain of a ZFN typically consists of an array of 3 to 4 zinc fingers. Amino acids at positions -1, +2, +3, and +6 relative to the origin of the zinc finger ∞-helix, which contribute to site-specific binding to the target DNA, can be altered and customized to suit a specific target sequence. Other amino acids form a common backbone to produce ZFNs with different sequence specificities. Rules for selecting the target sequence of a ZFN are known in the art.
[0079] The FokI nuclease domain requires dimerization to cleave DNA, and therefore requires two ZFNs with their C-terminal regions to bind to the opposing DNA strands (5-7 bp apart) at the cleavage site. If the two ZF binding sites are palindromic, the ZFN monomer can encircle the target site. As used herein, the term ZFN is broad and includes monomeric ZFNs that can cleave double-stranded DNA without the assistance of another ZFN. The term ZFN is also used to refer to one or both members of a pair of ZFNs engineered to cooperate in cleaving DNA at the same site.
[0080] Because the DNA-binding specificity of zinc finger domains can theoretically be reengineered using one of several methods, it is possible to construct custom ZFNs to target virtually any gene sequence. Publicly available methods for engineering zinc finger domains include context-dependent assembly (CoDA), oligomerized pool engineering (OPEN), and modular assembly.
[0081] TALEN
[0082] Transcription activator-like effector factors (TALEs) can be engineered to bind virtually any DNA sequence. TALE proteins are DNA-binding domains derived from various plant bacterial pathogens of the genus *Xanthomonas*. During infection, pathogen X secretes TALEs into host plant cells. TALEs migrate to the nucleus, where they recognize and bind to specific DNA sequences in the promoter regions of specific genes in the host genome. TALEs have a central DNA-binding domain consisting of 13–28 repeating sequences of 33–34 amino acids each. The amino acids in each monomer are highly conserved, except for highly variable amino acid residues at positions 12 and 13. These two variable amino acids are called repeat-variable diresidues (RVDs). RVDs preferentially recognize adenine, thymine, cytosine, and guanine / adenine for NI, NG, HD, and NN, respectively, and regulation of RVDs can lead to the recognition of consecutive DNA bases. This simple relationship between amino acid sequences and DNA recognition has allowed for the engineering of specific DNA-binding domains by selecting combinations of repeating segments containing appropriate RVDs. Transcription activator-like effector (TALE) DNA-binding domains can be fused with functional domains, such as recombinases, nucleases, transposases, or helicases, thereby conferring sequence specificity to those functional domains.
[0083] Transcription activator-like effector nucleases (TALENs) are artificial restriction enzymes derived by fusing a transcription activator-like effector (TALE) DNA-binding domain with a nuclease domain. As used herein, the term TALEN is broad and includes monomeric TALENs that can cleave double-stranded DNA without the assistance of another TALEN. The term TALEN is also used to refer to one or both members of a pair of TALENs that cooperate in cleaving DNA at the same site. In some embodiments, the nuclease is selected from the group consisting of PvuII, MutH, TevI, FokI, AlwI, MlyI, SbfI, SdaI, StsI, CleDORF, Clo051, and Pept071. When FokI binds to each member of the TALEN pair and fuses with the TALE domain of a DNA site flanked by a target site, the FokI monomer dimerizes and induces DSB at the target site.
[0084] In addition to the wild-type FokI cleavage domain, mutant FokI cleavage domain variants have been designed to improve cleavage specificity and activity. The FokI domain functions as a dimer, thus requiring constructs with two unique DNA-binding domains of appropriate orientation and spacing targeting sites in the target genome. The number of amino acid residues between the TALEN DNA-binding domain and the FokI cleavage domain, and the number of bases between the two individual TALEN binding sites, are parameters for achieving high levels of activity. PvuII, MutH, and TevI cleavage domains are available alternatives to FokI and its variants for use with TALE. PvuII acts as a highly specific cleavage domain when coupled to TALE (see Yank et al. 2013. PLoS One. 8:e82539). MutH is capable of introducing chain-specific notches into DNA (see Gabsalilow et al. 2013. Nucleic Acids Research. 41:e83). TevI introduces double-strand breaks at target sites in DNA (see Beurdeley et al., 2013. Nature Communications. 4: 1762).
[0085] The relationship between the amino acid sequence of the TALE binding domain and DNA recognition allows for protein design. Numerous software programs, such as DNA Works, can be used to design TALE constructs. Other methods for designing TALE constructs are known to those skilled in the art. Doyle et al. (2012) TAL Effector-Nucleotide Targeter (TALE-NT) 2.0: tools for TALeffector design and target prediction. Nucleic Acids Res. 40(W1): W117-W122; Cermak (2011). Efficient design and assembly of custom TALEN and other TALeffector-based constructs for DNA targeting. Nucleic Acids Res. 39(12): e82.
[0086] Megabase nuclease
[0087] Megabase nucleases, typically identified in microorganisms, are unique enzymes with high activity and long recognition sequences (>14 bp), thereby inducing site-specific digestion of target DNA. Engineered forms of naturally occurring megabase nucleases typically have extended DNA recognition sequences (e.g., 14–40 bp).
[0088] The engineering of megabase nucleases is more challenging than that of ZFN and TALEN because the DNA recognition and cleavage functions of megabase nucleases are entangled in a single structural domain. Novel megabase nuclease variants that recognize unique sequences and possess improved nuclease activity have been generated using specialized mutagenesis and high-throughput screening methods.
[0089] Agnuclease
[0090] The argonase family of proteins consists of DNA-guided endonucleases. Argonautes isolated from *Natronobacterium gregoryi* have been reported to be suitable for DNA-guided genome editing in human cells (Gao et al., DNA-guided genome editing using the *Natronobacterium gregoryi* Argonaute. *Nature Biotechnology* 34:768-773 (2016). Argonautes from other species have been identified (non-limiting examples of argonase proteins include *Thermophyton floccosum* argonase (TtAgo), *PfAgo* argonase, *NgAgo* argonase, homologues thereof, or modified forms thereof). Each of these unique argonases has been associated with a sequence encoding a DNA guide.
[0091] CRISPR
[0092] The CRISPR (Clustered Regularly Interspaced Short Palindromic Repeats) / Cas (CRISPR-associated) system is an alternative to synthetic proteins with DNA-binding domains that enable them to modify genomic DNA at specific sequences (e.g., ZFN and TALEN). The specificity of the CRISPR / Cas system is based on an RNA guide that recognizes the target DNA sequence using complementary base pairing. In some embodiments, the site-specific genomic modifying enzyme is a CRISPR / Cas system. In one aspect, the site-specific genomic modifying enzymes provided herein can comprise any RNA-guided Cas nuclease (non-limiting examples of RNA-guided nucleases include Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Ca...). s6, Cas7, Cas8, Cas9 (also known as Csn1 and Csx12), Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, Csf4, Cpf1, their homologues or modified forms thereof); and optionally, guide RNAs necessary to target the corresponding nucleases.
[0093] The CRISPR / Cas system is part of the adaptive immune system of bacteria and archaea, preventing foreign DNA from being affected by invading nucleic acids such as viruses by cleaving it in a sequence-dependent manner. Immunity is acquired by integrating short fragments of invading DNA, called spacers, between two adjacent repetitive sequences proximal to the CRISPR locus. The CRISPR array, including the spacers, is transcribed during subsequent collisions with the invading DNA and processed into small interfering CRISPR RNAs (crRNAs) of approximately 40 nt in length. These crRNAs combine with trans-activating CRISPR RNAs (tracrRNAs) to activate and guide the Cas9 nuclease. This cleaves the homologous double-stranded DNA sequence in the invading DNA, called the protospacer. A prerequisite for cleavage is the presence of a conserved protospacer adjacent motif (PAM) downstream of the target DNA, which typically has the sequence 5'-NGG-3' but less frequently NAG. Specificity is provided by a so-called "seed sequence" approximately 12 bases upstream of the PAM, which must match between the RNA and the target DNA. Cpf1 functions similarly to Cas9, but Cpf1 does not require tracrRNA.
[0094] As used herein, the term "target site" broadly refers to a genomic sequence selected for integration of target DNA. In some aspects, the target site is located in a genomic region. In other aspects, the target site is located in an intergenetic region. In yet another aspect, the target site may include both genomic and intergenetic regions. In one aspect, the target site provided herein is recognized and cleaved by a double-strand break inducer, such as a site-specific genome-modifying enzyme, a site-specific recombinase, or a site-specific transposase. In some embodiments, the target site comprises a nucleic acid sequence having at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity to a portion of a maize B chromosome sequence selected from the group consisting of SEQ ID NO: 1-126 and 128-888.
[0095] As used in this article, the term “recombination” refers to the exchange of nucleotide sequences between two DNA molecules.
[0096] As used herein, the term "homologous recombination" refers to the exchange of nucleotide sequences at conserved regions shared by two genomic loci. Homologous recombination includes symmetrical and asymmetrical homologous recombination. Asymmetrical homologous recombination can also be called unequal recombination.
[0097] Methods for detecting recombinants include, but are not limited to, 1) phenotypic screening; and 2) molecular marker technology, such as... Alternatively, single nucleotide polymorphism (SNP) analysis can be performed using Illumina / Infinium PCR; 3) Southern blotting; and 4) sequencing (e.g., Sanger, Illumina, 454, Pac-Bio, Ion Torrent). An example of a method used to identify the integration of target DNA into a chromosome is inverse PCR (iPCR).
[0098] On the one hand, the integration of the target DNA presented in this paper occurs via homologous recombination (HR). On the other hand, the integration of the target DNA presented in this paper occurs via non-homologous end joining (NHEJ).
[0099] As used herein, “transgenic” means a plant or seed whose genome has been altered through the stable integration of recombinant DNA. Transgenic lines include plants regenerated from the initially transformed plant cells and transgenic plants from later generations or hybrids of the transformed plant. As used herein, “exogenous” means a gene that is not normally present in the cells being transformed or does not exist in the form, structure, etc., as found in the transformed DNA segment or gene. Therefore, the term “exogenous” gene or DNA means any gene or DNA segment introduced into the recipient cell, regardless of whether a similar gene may already be present in such cells. The types of DNA included in exogenous DNA can include DNA already present in plant cells, DNA from another plant, DNA from a different organism, or externally generated DNA, such as DNA sequences containing antisense messages of a gene or synthetic or modified forms of DNA sequences encoding a gene.
[0100] Methods for transforming plant cells are well known to those skilled in the art. For example, specific descriptions of transforming plant cells by microparticle bombardment with particles coated with recombinant DNA can be found in U.S. Patents 5,015,580 (soybean), 5,550,318 (maize), 5,538,880 (maize), 5,914,451 (soybean), 6,160,208 (maize), 6,399,861 (maize), and 6,153,812 (wheat), 6,002,070 (rice), 7,122,722 (cotton), and 6,051,756. (Brassica napus), 6,297,056 (Brassica napus), and U.S. Patent Publication 20040123342 (Sugarcane), while Agrobacterium-mediated transformation is described in U.S. Patents 5,159,135 (Cotton), 5,824,877 (Soybean), 5,591,616 (Maize), 6,384,301 (Soybean), 5,750,871 (Brassica napus), 5,463,174 (Brassica napus), and 5,188,958 (Brassica napus). All patents and patent publications are incorporated herein by reference. Methods for transforming other plants can be found, for example, in the Compendium of Transgenic Crop Plants (2009), Blackwell Publishing. Any suitable method known to those skilled in the art can be used to transform plant cells with any of the provided nucleic acid molecules.
[0101] As used herein, “stable transformation” is defined as the transfer of DNA into the genomic DNA of a target cell, thereby allowing the target cell to regenerate a complete organism and pass the transferred DNA to the next generation of the transformed organism. Stable transformation requires the integration of the transferred DNA into the reproductive cells of the transformed organism. As used herein, “transient transformation” is defined as the transfer of DNA into a cell but not into the next generation of the transformed organism. Transient transformation is typically the transformation of leaf or root (e.g., non-reproductive) tissues in plants. The transformed DNA typically does not integrate into the genomic DNA of the transformed cell during transient transformation. On one hand, the methods provided herein stably transform plant cells. On the other hand, the methods provided herein transiently transform plant cells.
[0102] In one aspect, this disclosure provides a method for transforming plant cells with a nucleic acid sequence of a coding site-specific genome-modifying enzyme. In another aspect, the method provided herein includes stable transformation with a nucleic acid sequence of a coding site-specific genome-modifying enzyme. In another aspect, the method provided herein includes transient transformation with a nucleic acid sequence of a coding site-specific genome-modifying enzyme. In one aspect, the method provided herein includes constitutively expressed nucleic acid of a coding site-specific genome-modifying enzyme. In another aspect, the method provided herein includes a nucleic acid sequence of a coding site-specific genome-modifying enzyme controlled by a regulatory promoter.
[0103] In one aspect of this disclosure, the recombinant nucleic acid sequences provided herein can be fully integrated into chromosomes.
[0104] In one aspect, plant cell transformation is performed via Agrobacterium-mediated transformation, and the target nucleic acid molecule is present on one or more integrated DNA sequences (US Patent Nos. 6,265,638, 5,731,179; US Patent Application Publication US2005 / 0183170; 2003110532) or other nucleic acid sequences (e.g., vector backbones) transferred into the plant cells. The sequence that can be transferred into the plant cells may be present on a transformation vector within the bacterial strain used for transformation. In another aspect, the sequence may be present on a separate transformation vector within the bacterial strain. In yet another aspect, the sequence may be visible in separate bacterial cells or bacterial strains used collectively for transformation.
[0105] As used herein, a "transformation vector" is plasmid DNA capable of transforming plant cells. On the one hand, the transformation vectors provided herein may contain any target DNA provided herein. On the other hand, the transformation vectors provided herein may contain any nucleic acid molecules provided herein.
[0106] The DNA constructs used for transformation in the methods of this disclosure generally also contain plasmid backbone DNA segments that provide replication function and antibiotic selection in bacterial cells, such as *E. coli* origin of replication like ori322, *Agrobacterium* origin of replication like oriV or oriRi, and coding regions of selectable marker genes for Tn7 aminoglycoside adenosyltransferase (aadA) conferring resistance to spectinomycin or strepmycin, such as spec / strep or gentamicin (Gm, Gent) selectable marker genes. For plant transformation, the host bacterial strain is typically *Agrobacterium tumefaciens* ABI, C58, LBA4404, AGLO, AGL1, EHA101, or EHA105 carrying plasmids with transfer function of expression units. Other strains known to those skilled in the art of plant transformation may function in this disclosure.
[0107] To confirm the presence of foreign DNA or "transgenic" cells, various assays can be performed. These assays include, for example, "molecular biology" assays, such as Southern and Northern Blotting and PCR; "biochemical" assays, such as detecting the presence of protein products, for example by immunological means (ELISA and Western Blotting) or by enzymatic functions (e.g., GUS assay); pollen histochemistry; plant part assays, such as leaf or root assays; and analysis of the phenotype of the whole regenerated plant.
[0108] This disclosure provides a method for manufacturing transgenic plant cells containing target DNA, wherein the method includes selecting a target B chromosome locus having at least 70%, 75%, 80%, 85%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, 99%, or 100% sequence identity with any sequence selected from SEQ ID NO: 1-126 and 128-888; selecting a site-specific genomic modification enzyme that specifically binds to and cleaves the target B chromosome locus; introducing the site-specific genomic modification enzyme into a plant cell; introducing the target DNA into the plant cell; and inserting the target DNA into the target B chromosome locus; and selecting plant cells containing the target DNA integrated into the target site of the B chromosome locus. In one aspect, the method provided herein uses homology-directed repair integration to integrate the target DNA into the B chromosome locus. On the other hand, the method presented in this paper uses non-homologous end joining integration to integrate the target DNA into the B chromosome locus.
[0109] In one respect, the method provided herein integrates one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more target DNAs into one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more B chromosome loci located on a B chromosome. On the other hand, the method provided herein integrates one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more target DNAs into one or more, two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, or ten or more B chromosome loci.
[0110] On one hand, the method for manufacturing transgenic plant cells provided herein includes a target DNA modified during integration into a B chromosome sequence. On the other hand, the method for manufacturing transgenic plant cells provided herein includes a B chromosome sequence modified during the integration of the target DNA provided herein.
[0111] In one aspect, this disclosure provides a method for manufacturing transgenic plant cells, wherein two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more target DNAs are integrated into two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more target B chromosome genes on an independent B chromosome. In the genomic loci; and wherein the B chromosome recombination causes the two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more target B chromosome genomic loci to produce a megabase locus containing the two or more, three or more, four or more, five or more, six or more, seven or more, eight or more, nine or more, ten or more, fifteen or more, twenty or more, or twenty-five or more target DNA.
[0112] As used herein, a “megabase locus” refers to a segment of a transgenic trait that is genetically linked and normally inherited as a single unit. Megabase loci according to this disclosure can provide plants with one or more desired traits, including but not limited to enhanced growth, drought tolerance, salt tolerance, herbicide tolerance, insect resistance, pest resistance, disease resistance, etc. In specific embodiments, a megabase locus comprises at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, or 15 transgenic loci (events) that are physically segregated but genetically linked so that they can be inherited as a single unit. Each transgenic locus in a megabase locus may be spaced apart from each other by 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1, 1.5, 2, 2.5, 3, 5, 10, 15, or 20 cM. In some implementations, the megabase locus contains at least about 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 13, or 15 transgenic loci (events) that are not gene-linked but are located on the same B chromosome.
[0113] As used herein, centimole (“cM”) is a unit of measurement for recombination frequency and genetic distance between two loci. One cM is equal to the 1% probability that a marker at one locus will segregate from a marker at a second locus after a single-generation hybridization.
[0114] As used herein, “tight linkage” means that a marker or locus is within approximately 20 cM, 15 cM, 10 cM, 5 cM, 4 cM, 3 cM, 2 cM, 1 cM, 0.5 cM, or less than 0.5 cM of another marker or locus. For example, 20 cM means that recombination occurs between the marker and the locus at a frequency equal to or less than approximately 20%.
[0115] As used in this article, "proximal" refers to the relative positioning of two nucleic acid sequences. On the one hand, if the first nucleic acid sequence is less than 50,000 base pairs, less than 25,000 base pairs, less than 15,000 base pairs, less than 10,000 base pairs, less than 7,500 base pairs, less than 5,000 base pairs, less than 4,000 base pairs, less than 3,000 base pairs, less than 2,500 base pairs, less than 2,000 base pairs, or less than 1,500 base pairs between the second nucleic acid sequence... If the number of base pairs is less than 1000, less than 750, less than 500, less than 250, less than 100, less than 75, less than 50, less than 40, less than 30, less than 20, less than 10, less than 5, less than 3, or less than 1 base pair, then the first nucleic acid sequence is proximal to the second nucleic acid sequence.
[0116] In one aspect, this disclosure provides plant cells that are not propagation material and do not mediate the natural reproduction of plants. In another aspect, this disclosure also provides plant cells that are propagation material and mediate the natural reproduction of plants. In yet another aspect, this disclosure provides plant cells that cannot sustain themselves through photosynthesis. In yet another aspect, this disclosure provides plant somatic cells. Unlike reproductive cells, somatic cells do not mediate plant reproduction.
[0117] The plant cells or plant parts provided may be derived from seeds, fruits, leaves, cotyledons, hypocotyls, meristems, plumules, endosperm, roots, branch buds, petioles, pods, flowers, inflorescences, stems, pedicels, styles, stigmas, receptacles, petals, calyxes, pollen, anthers, filaments, ovaries, ovules, pericarps, phloem, buds, or vascular tissue. In another aspect, this disclosure provides a plant chloroplast. In another aspect, the invention provides epidermal cells, stomatal cells, trichomes, root hair cells, storage root cells, or tuber cells. In another aspect, this disclosure provides a protoplast cell. In another aspect, this disclosure provides a plant callus cell. In one aspect, any plant, plant part, or plant cell provided herein may contain any recombinant sequences provided herein.
[0118] The nucleic acid molecules provided herein include deoxynucleic acid (DNA) and ribonucleic acid (RNA) and their functional analogs, such as complementary DNA (cDNA). The nucleic acid molecules provided herein can be single-stranded or double-stranded. Nucleic acid molecules contain the nucleotide bases adenine (A), guanine (G), thymine (T), and cytosine (C). Uracil (U) replaces thymine in RNA molecules. The symbol “N” can be used to denote any nucleotide base (e.g., A, G, C, T, or U). As used herein, “complementary” in relation to a nucleic acid molecule or nucleotide base means that A is complementary to T (or U) and G is complementary to C. Two complementary nucleic acid molecules are capable of hybridizing with each other. For example, the two strands of double-stranded DNA are complementary to each other. In one aspect of this disclosure, two nucleic acid sequences are homologous or substantially homologous if they have at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% sequence identity with each other. The polynucleotide molecules of this invention comprise at least 2, at least 5, at least 10, at least 20, at least 30, at least 40, at least 50, at least 100, at least 250, at least 500, at least 1000, at least 1500, at least 2000, at least 2500, or at least 3000 nucleotide bases. As used herein, “encoding” refers to a polynucleotide that encodes an amino acid of a polypeptide. A sequence of three nucleotide bases encodes one amino acid. As used herein, “expression” refers to the transcription of RNA from a DNA molecule. In one aspect, the expression of the target DNA provided herein requires a promoter. In another aspect, the expression of the target DNA provided herein does not require a promoter. In one aspect, the expression transcripts provided herein are spliced. In another aspect, the expression transcripts provided herein are not spliced.
[0119] As used herein, the terms “polypeptide,” “peptide,” and “protein” are used interchangeably to refer to polymers containing amino acid residues. The term also applies to amino acid polymers in which one or more amino acids are chemical analogs or modified derivatives of the corresponding naturally occurring amino acids.
[0120] In one aspect of this disclosure, a nucleic acid sequence may be physically linked to another nucleic acid sequence, operatively linked to another nucleic acid sequence, or both physically and operatively linked to another nucleic acid sequence. As used herein, physical linkage means that the physically linked nucleic acid sequences are located on the same nucleic acid molecule. Physical linkage can be adjacent or proximal. The nucleic acid sequence provided herein may be adjacent to another nucleic acid sequence. As used herein, operatively linked means that the operatively linked nucleic acid sequence exhibits its desired function. For example, in one aspect of this disclosure, the provided DNA promoter sequence can induce transcription of an operatively linked DNA sequence into RNA. The nucleic acid sequence provided herein may be upstream or downstream of a physically or operatively linked nucleic acid sequence. As used herein, upstream means that the nucleic acid sequence is located before the 5' end of the linked nucleic acid sequence. As used herein, downstream means that the nucleic acid sequence is located after the 3' end of the linked nucleic acid sequence. As used herein, 5' means the start point encoding the DNA sequence or the start point of the RNA molecule. As used herein, 3' means the end encoding the DNA sequence or the end of the RNA molecule. As used herein, "opposite side" refers to the 5' or 3' side of a nucleic acid molecule. For example, the first nucleic acid sequence on the 5' side of a second nucleic acid molecule is opposite to the third nucleic acid sequence on the 3' end of the second nucleic acid molecule.
[0121] In one respect, the nucleic acid molecules provided herein are transformation vectors. In another respect, the transformation vector is a bacterial plasmid. As used herein, a "plasmid" is a DNA molecule that is physically separate from chromosomal DNA and capable of replication. In another respect, the bacterial plasmid is an Agrobacterium tumor-inducing (Ti) plasmid. In yet another respect, the transformation vector is a synthetic plasmid. As used herein, a "synthetic plasmid" is an artificial plasmid capable of performing the same function (e.g., replication) as a natural plasmid (e.g., Ti plasmid). Without limitation, those skilled in the art can create synthetic plasmids de novo by synthesizing plasmids from individual nucleotides or by splicing nucleic acid molecules from different existing plasmids together. Example
[0122] Example 1: Obtaining a unique maize B chromosome sequence
[0123] To identify sequences specific to the maize B chromosome, the following method was employed. A line (designated B73+) backcrossed multiple times with B73 maize containing multiple B chromosomes was obtained from the University of Missouri and grown from seed using standard potting and growing conditions in a greenhouse. Due to the non-Mendelian segregation of the B chromosome, root tip tissue samples were collected from individual plants grown from B73+ seeds. Chromosome smears were generated from the root tip tissue samples for karyotype analysis, essentially as described by Kato et al. (2011), Chromosome Painting for Plant Biotechnology, Plant Chromosome Engineering in Methods in Molecular Biology, JA Birchler ed., 701, 67-96. These chromosome smears were stained with the chromosome counterstain 4',6-diamidinyl-2-phenylindole (DAPI) and visually scored to determine the number of B chromosomes in each individual cell nucleus. See, for example... Figure 1 Based on this data, at least one plant was identified as containing 20 separate B chromosomes in each cell nucleus.
[0124] Leaf tissues were also collected from each plant, and genomic DNA was extracted. The genomic DNA was then used for… PCR copy number analysis was performed to confirm the number of B chromosomes in each individual plant. For this analysis, B chromosome-specific markers and A chromosome-specific genes (alcohol dehydrogenase (ADH)) were amplified in the same PCR reaction. Controls included B73 plants lacking a B chromosome and four plants from non-B73 germplasm with 1, 2, 3, or 4 copies of the B chromosome. When plotted, the signal ratio of the B chromosome-specific probe to the A chromosome-specific ADH probe yielded nearly linear results. Karyotype analysis identified the plants as having 20 copies of the B chromosome, consistent with Taqman PCR copy number analysis (20 copies of the B chromosome). See Table 1.
[0125] By utilizing the 20-fold representation of the B chromosome relative to the A chromosome in DNA extracted from identified plants with 20 copies of the B chromosome, it was recognized that the Fors plasmid sequencing strategy can be used to identify unique B chromosome sequences by sequencing fragments of maize chromosome DNA and extracting high-confidence, non-repetitive, unique B sequences from the data.
[0126] Table 1: Correlation between B chromosome copy number determined by karyotype analysis and relative B marker values determined by Taqman PCR analysis
[0127]
[0128]
[0129]
[0130] Leaf tissue was collected from single plants with 20 copies of the B chromosome, and genomic DNA was extracted using a modified CTAB method. 0.8 g of plant tissue was mechanically ground, and the powder was transferred to a 50 ml tube containing 20 ml of extraction buffer (1.5% hexadecyltrimethylammonium bromide (CTAB) (Sigma-Aldrich, Saint Louis, MO); 75 mM Tris-HCl (pH 8.0); 15 mM EDTA (pH 8.0); and 1 M NaCl). The mixture was incubated at 56°C for 20 minutes in a water bath shaker. After incubation, the tube was centrifuged, and the DNA-containing layer at the top was transferred to a new tube. Extraction was achieved by adding 1 / 10 volume of 10% CTAB (10% hexadecyltrimethylammonium bromide (CTAB) and 0.7 M NaCl) and an equal volume of chloroform / isoamyl alcohol (24:1), mixing by rotating and inverting for 20 minutes, followed by centrifugation. Transfer the supernatant to a new tube and add an equal volume of isopropanol. Centrifuge to form DNA spheres and discard the supernatant. Resuspend the DNA spheres in 5 ml of 1M NaCl and 5 μL of 10 mg / ml RNase, then reprecipitate the DNA by adding twice the volume of ethanol. Wash the DNA spheres twice with 5 ml of 70% ethanol each time, and resuspend the final DNA spheres in 1 mM Tris-HCl + 0.1 mM EDTA (pH 8.0).
[0131] use The v2.0 Fors plasmid cloning kit (Lucigen Corporation, Middleton, WI) is used to prepare Fors plasmid libraries from genomic DNA isolated from plants with 20 copies of the B chromosome. Each Fors plasmid construct contains an insert of up to approximately 40 kb of maize genome sequence. The Fors plasmid constructs are pooled into 192 libraries, each containing approximately 1000 Fors plasmid clones. Fors plasmid DNA is extracted from each library, following the manufacturer's protocol. Sequencing. Based on this sequencing, 162 out of 192 libraries contained usable sequences or 'readings'. The readings of each of the 162 libraries were assembled using the assembly software tool PCAP ("Application of a superword array in genome assembly", Xiaoqiu Huang et al., Nucleic Acids Research, 2006, Vol. 34, No. 1, 201-205) and the CLCbio assembly unit (www.clcbio.com / products / clc-assembly-cell; Qiagen, Waltham, MA). Sequence repetitions in each well were masked by using A chromosome repetitive sequences from known maize genome repetitive sequence libraries (maize repetitive sequence libraries are available from Smit, AFA, Hubley, R&Green, P RepeatMasker Open-4.0.2013-2015, www.repeatmasker.org). Next, using the assembled sequences from each library, non-repetitive sequences specific to chromosome A were filtered out as follows: The alignment software Cross_Match (with parameters: minimum match 17 - minimum score 200 - label - mask level 0 - penalty - 3) (Phil Green, www.phrap.org / phredphrapconsed.html) was used to compare the Fors plasmid library sequence assemblies with the maize reference B73 genome (Schnable et al., (2009) The B73 Maize Genome: Complexity, Diversity, and Dynamics. Science 326, 1112-5) (representing the maize chromosome A sequence), and Monsanto's internal scripts were used to analyze the alignment output and adjust and filter out matching chromosome A sequences. During this process, B73 mitochondrial genomic DNA (NCBI reference sequence: NC_007982.1) and chloroplast genomic DNA (NCBI reference sequence: NC_001666.2) were also filtered out. The next step was to align each library with the sequences of 161 other libraries, ensuring that all libraries were compared against each other. This analysis was performed to filter out “chromosome B repetitive sequences” identified by sequences appearing in 130 or more wells, or to filter out “chromosome A sequences not in the reference dataset” identified by sequences appearing only once to three times in all libraries. Finally, the remaining sequences were compared to identify the longest representative sequence for each contiguous group, where multiple matches existed among the numerous contiguous groups. Contiguous sequences larger than 2 kb were identified as unique sequences of maize chromosome B (SEQ ID NO: 1-126, 128-791).
[0132] Example 2: Maize B chromosome reference assembly
[0133] To generate a B chromosome reference assembly, genomic DNA was prepared from B73 plants containing 20 copies of the B chromosome, following the manufacturer's protocol. Sequencing. PacBio sequencing generated long read sequences to produce high-quality, sequential assemblies of the genome. The long read B chromosome sequence data were combined with unique B chromosome sequences generated from sequencing and analysis of the Fors plasmid library as described in Example 1. The combined dataset was used to generate B reference chromosome assembly sequences 792-889. This B reference chromosome can be used to map small RNA and DNA methylation patterns to identify novel maize genes and their corresponding promoters, introns, terminators, and insulators; develop new B chromosome markers; physically locate new B chromosome markers to enable experiments measuring possible recombination between B chromosomes; and design gene-editing tools such as TALEN, CRISPR, Cpf1, zinc fingers, or megabase nucleases.
[0134] Example 3: Site-directed genomic modification of maize chromosome B
[0135] The maize B chromosome sequence identified herein was used for site-directed genome modification. A megabase nuclease was engineered to target one of the unique maize B chromosome sites selected from any of SEQ ID NO: 1-126 and 128-888. Maize cells were (1) contacted with the megabase nuclease under conditions allowing double-strand breaks at the selected site, and (2) contacted with a donor DNA molecule containing the target sequence to be integrated into the selected site. The donor DNA molecule could be integrated via non-homologous end joining (NHEJ) or via homologous recombination.
[0136] Alternatively, the TALEN is engineered to target one of the unique maize B chromosome loci selected from any of SEQ ID NO: 1-126 and 128-888. Maize cells are (1) contacted with the TALEN under conditions allowing double-strand breaks at the selected site, and (2) contacted with a donor DNA molecule containing the target sequence to be integrated into the selected site. The donor DNA molecule can be integrated via non-homologous end joining (NHEJ) or via homologous recombination. The TALEN can be a single molecule or can function as a pair of molecules.
[0137] Alternatively, RNA-guided DNA nucleases such as CRISPR-Cas9 or Cpf1 endonuclease are used to target one of the unique maize B chromosome loci selected from any of SEQ ID NO: 1-126 and 128-888. Maize cells are (1) contacted with CRISPR and Cas9 endonucleases or Cpf1 endonucleases under conditions allowing double-strand breaks at the selected site, and (2) contacted with a donor DNA molecule containing the target sequence to be integrated into the selected site. The donor DNA molecule may be integrated via non-homologous end joining (NHEJ) or via homologous recombination.
[0138] The donor DNA may include one or more gene expression cassettes having one or more expression elements (e.g., promoters, introns, targeting peptides, 3'-untranslated regions / termination signals) to allow protein and / or target RNA to be expressed in plant cells. The target protein may be encoded by an insecticide resistance gene, herbicide tolerance gene, nitrogen use efficiency gene, water use efficiency gene, nutrient quality gene, DNA binding gene, selectable marker gene, transcription factor gene, site-specific endonuclease gene, megabase nuclease gene, DNA-targeted recombinase gene, or a gene encoding an RNA-guided endonuclease such as Cas9 or Cpf1. In some cases, the donor DNA may encode one or more of the following: an RNAi construct, a guide sequence capable of hybridizing to the target sequence, a tracr pairing sequence, and a tracr sequence.
[0139] The maize B chromosome sequence identified herein was used for genome editing. One or more sites were selected from sequences in any of SEQ ID NO: 1-126 and 128-888. The one or more sites were targeted using an engineered site-specific megabase nuclease, endonuclease, or DNA-targeting recombinase. Maize cells (1) were then contacted with the engineered site-specific megabase nuclease, endonuclease, or DNA-targeting recombinase under conditions that allowed double-strand breaks at the selected genomic sites, and DNA was deleted or inserted (insertion / deletion) or altered.
[0140] Example 4: Development of a maize B chromosome probe
[0141] The maize B chromosome sequences identified in this paper were used to develop a set of maize B chromosome-specific PCR primers. For example, Table 2 details three sets of PCR primers that amplify regions of B chromosome-specific sequences from SEQ ID NO:175, 177, and 183. Specifically, PCR primer pairs SEQ ID NO:889 and SEQ ID NO:890 amplify the region at coordinates 2626-2825 of SEQ ID NO:175; PCR primer pairs SEQ ID NO:891 and SEQ ID NO:892 amplify the region at coordinates 1170-1368 of SEQ ID NO:177; and PCR primer pairs SEQ ID NO:893 and SEQ ID NO:894 amplify the region at coordinates 1118-1319 of SEQ ID NO:183. The control group for PCR primers SEQ ID NO:895 and SEQ ID NO:896, identified from the sequence DU978594.1pCL7T-2CL-repetitive sequence reported by Cheng et al., Chromosome Research (2010) 18:605-619, is labeled BCHR0004. The control group for PCR primers SEQ ID NO:897 and SEQ ID NO:898, identified from the sequence DU820494pCL38T-2CL-repetitive sequence DH1, identified from the sequence DU820494pCL38T-2CL-repetitive sequence reported by Cheng et al., Chromosome Research (2010) 18:605-619, is labeled BCHR0006. Furthermore, the control group of PCR primers SEQ ID NO:899 and SEQ ID NO:900, which were identified from the StarkB repeat sequence reported by Lamb et al., Chromosome Research (2007) 15:383-398, targeting the 3-4B chromosome region of the distal heterochromatin region, was labeled BCHR0007.
[0142] Table 2. Primers for Chromosome-Specific PCR
[0143]
[0144] The PCR amplicon was localized to chromosome B using the previously characterized BA translocation line. Maize with BA translocation events was obtained from the Maize Gene Bank Reserve (www.maizegdb.org / stock_catalog). Samples used to test primers and control B-specific probes were trisomic lines containing the BA translocation (and two normal A chromosomes) but not the reciprocal AB translocation. Due to transmission issues, the Maize Gene Bank Reserve preserves these BA translocations as heterozygotes. Therefore, delivered seed packages may not necessarily contain trisomic material. If so, the translocation heterozygotes (verified by FISH screening of packages received from the library) were crossed with testcross lines to produce testcross germplasm, and color-coded markers were used to identify possible trisomy in the F1 generation, which were then planted and sampled.
[0145] DNA was extracted from tissues collected from tertiary trisomy plants containing one of the BA translocation chromosomes (specifically, TB-10L7, TB-10L37, TB-10L3, TB-10L30, TB-10L19, TB-10L26, and TB-10L36), or from samples containing the full-length B chromosome (positive controls for all probes), or from samples without the B chromosome (negative controls), or from samples from lines with spontaneously truncated B chromosomes. The DNA was used for PCR amplification using the following PCR amplicon regions: BCHR004 for the proximal centromere region; BCHR006 for the euchromatin region; and BCHR007 for the distal heterochromatin region; and PCR amplification primer pairs identified from the unique B chromosome sequence (SEQ ID NO: 889+890; SEQ ID NO: 891+892; and SEQ ID NO: 893+894). Results with or without PCR amplicons are provided in Table 3. These results are shown in... Figure 2In the diagram, regions of chromosome B located by the probe are indicated in parentheses. For example, probe BCHR0004 is located to the proximal / centromere region of chromosome B; probe BCHR0006 is located to the proximal euchromatin 2 (PE2) and DH1 regions; probe BCHR0007 is located to the distal (DH3 and DH4) regions; PCR primer pair amplicon SEQ ID NO:889+890 is located to the proximal euchromatin 2 (PE2) region; and PCR primer pairs SEQ ID NO:891+892 and SEQ ID NO:894+894 are both located to the proximal euchromatin 2 (PE2) and DH1 regions. These results demonstrate that the unique chromosome B sequences present as SEQ ID NO:1-126 and SEQ ID NO:128-888 can be used to develop PCR primer amplicones for chromosome B mapping and the development of unique chromosome B probes. Chromosome B unique PCR amplicones and probes can be used in southern and northern analyses, as well as other molecular biology and sequence-related protocols.
[0146] Table 3. Results of BA translocation plot. (+) indicates the presence of PCR amplicons, (-) indicates the absence of PCR amplicons.
[0147]
[0148]
[0149] Example 5: Insertion of transgenes on chromosome B
[0150] Methods for evaluating the detection of transgenic insertion in maize B chromosome were assessed by transforming transgenic genes into maize germplasm containing the B chromosome. Two transgenic transformation vectors (transformation vector A and transformation vector B) conferring glyphosate tolerance were transformed into immature maize embryos using Agrobacterium tumefaciens to produce transgenic seedlings, according to methods known to those skilled in the art. The F1 immature embryos used with transformation vector A were derived from a cross between parental germplasm B73+B chromosome (ranging from 6 to 15 B chromosomes) (female) × LH244 (male). The F1 immature embryos used with transformation vector B were derived from a cross between parental germplasm LH244 (female) × B73+B chromosome (ranging from 4 to 16 B chromosomes) (male). After transformation, glyphosate was used to select offspring containing the transgenic genes. Root tip samples were obtained for FISH analysis during the seedling transfer to soil step. FISH probes were designed to target transgenes randomly integrated into the maize genome during transformation, and FISH was performed essentially as described by Kato et al. (2011), including DAPI contrast staining. Based on this analysis, no chromosome B insertion was identified in the case of transformation vector A. FISH analysis of root tip tissue from seedlings produced with transformation vector B identified approximately 10 events of random transgene insertion into chromosome B. Leaf samples were obtained from plants identified by FISH analysis as having confirmed chromosome B transgene insertion, and DNA was prepared. This DNA was used for inverse PCR (Hui et al., (1998) Cell Mol Life Sci. 54:1403-1422; Tonooka and Fujishima (2009) Appl Microbiol Biotechnol 85:37-43) to identify flanking sequences of the transgene insertion. In 10 events, 5 had a single copy of the transgene inserted into chromosome B, and left and right flanking sequences were obtained (Table 4). The FISH analysis was validated by using the flanking sequences to locate the transgene insertion site to the B chromosome sequence.
[0151] Unique B chromosome sequences, such as those presented in SEQ ID NO:1-126 and SEQ ID NO:128-888, are used to select specific target sites for genome modification, select specific target sites for transgene integration, and facilitate the identification of B chromosome transgene flanking. Furthermore, including unique B chromosome sequences, such as those presented in SEQ ID NO:1-126 and SEQ ID NO:128-888, in genome sequence analysis can improve the specificity of target site selection for A and B chromosome genome modifications (and thus reduce off-target effects).
[0152] Table 4. Transgenic Flanking Zones of Chromosome B SEQ ID NO
[0153]
[0154] Example 6: Development of B chromosome markers
[0155] The B chromosome sequences identified herein were used for the development of B chromosome markers, essentially as described in US20060141495, which is incorporated herein by reference in its entirety. SNP and Indel polymorphisms were identified by sequence alignment comparisons of contigs and individual pieces from at least two independent maize lines. Genomic libraries from multiple maize lines containing the B chromosome were generated by isolating genomic DNA from different maize lines using standard methods known in the art. For the genomic libraries, the genomic DNA was digested with a restriction endonuclease (e.g., PstI), the digested DNA was size-graded on a 1% agarose gel, and the recovered DNA fragments were ligated into plasmid vectors for sequencing using standard molecular biology techniques as described in Green and Sambrook (2012). All sequences were assembled to identify non-redundant sequences as described in Examples 1 and 2. Sequence differences among multiple clones on the assembled contigs were identified as single or multiple nucleotide polymorphisms. B chromosome sequences from multiple maize lines were assembled into loci possessing one or more polymorphisms, namely SNPs and / or Indels. Candidate polymorphisms were limited by the following parameters:
[0156] (a) The minimum length of an overlap group or single-piece alignment is 200 bases.
[0157] (b) The percentage of base identity observed in regions with 15 bases on each side of a candidate SNP is at least 75%.
[0158] (c) Specify that the minimum sequence reading in the contiguous group is 4.
[0159] Once the polymorphic regions are identified, a design is then created. Probes are used to detect specific genotypes in maize DNA samples. Such probes can be designed and the proprietary Taqman assay (registered trademark) is available from Applied Biosystems (Applied Biosystems, Foster City, Calif.). To verify that the assay produces accurate results, each new assay is performed on numerous replicates of samples with known genotype identifiers representing each of the three possible genotypes, such as two homozygous alleles and one heterozygous sample. For an assay to be valid and useful, well-defined clusters of data points must be generated so that at least 90% of the data points can be assigned to one of the three genotypes, and it is observed that said assignment is correct for at least 98% of the data points. Following this validation step, the assay is applied to the hybrid progeny between two highly inbred single plants to obtain clustering data, which is then used to calculate the genomic map location of polymorphic loci. SNP analysis is also used to assist in the selection of B chromosome target sites for site-specific genomic modification, as detailed in Example 3.
Claims
1. A recombinant maize B chromosome comprising: target DNA inserted into a target locus within a maize B chromosome sequence, the target locus having the sequence SEQ ID NO:881, wherein the right wing of the target DNA has the sequence SEQ ID NO:904 and the left wing of the target DNA has the sequence SEQ ID NO:903, and wherein the target DNA encodes a peptide or RNAi construct.
2. The recombinant maize B chromosome of claim 1, wherein the target DNA is integrated at a double-strand break generated by one or more site-specific genomic modifying enzymes.
3. The recombinant maize B chromosome as described in claim 2, wherein the one or more site-specific genome-modifying enzymes are selected from endonucleases, recombinases, transposases, or any combination thereof.
4. The recombinant maize B chromosome as described in claim 2, wherein one or more site-specific genome-modifying enzymes are endonucleases, and the endonucleases are selected from the group consisting of: meganucleases, zinc finger nucleases, transcription activator-like effector nucleases, Agnucleases, DNA-directed recombinases, DNA-directed endonucleases, RNA-directed recombinases, and RNA-directed endonucleases.
5. The recombinant maize B chromosome according to any one of claims 1 to 4, wherein the target DNA comprises one or more gene expression cassettes, wherein the gene expression cassettes are selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA binding gene expression cassettes, optional marker gene expression cassettes, RNAi construct expression cassettes, or site-specific genome modifying enzyme gene expression cassettes.
6. A method for manufacturing transgenic maize cells, said transgenic maize cells containing target DNA integrated into the B chromosome: (a) Select at least one target site within a maize B chromosome sequence having the sequence of SEQ ID NO:881; (b) Select site-specific genome-modifying enzymes that specifically cleave the target site; (c) Introducing the site-specific genome-modifying enzyme into the maize cells; (d) Introducing the target DNA into the corn cells; (e) Integrating the target DNA into the target site; and (f) Select transgenic maize cells containing the target DNA integrated into the B chromosome, wherein the right wing of the target DNA has the sequence SEQ ID NO: 904 and the left wing of the target DNA has the sequence SEQ ID NO:
903.
7. The method of claim 6, wherein the target DNA is integrated into the B chromosome via homology-directed repair.
8. The method of claim 6, wherein the target DNA is integrated into the B chromosome by non-homologous end joining.
9. The method of any one of claims 6 to 8, wherein the site-specific genome-modifying enzyme is selected from endonucleases, recombinases, transposases, or any combination thereof.
10. The method of claim 9, wherein the site-specific genome-modifying enzyme is an endonuclease selected from mega-nucleases, zinc finger nucleases, transcription activator-like effector nucleases, ragnucleases, DNA-directed recombinases, DNA-directed endonucleases, RNA-directed recombinases, and RNA-directed endonucleases.
11. The method of claim 6, wherein the target DNA comprises one or more gene expression cassettes, wherein the gene expression cassettes are selected from the group consisting of: insecticide resistance gene expression cassettes, herbicide tolerance gene expression cassettes, nitrogen use efficiency gene expression cassettes, water use efficiency gene expression cassettes, nutrient quality gene expression cassettes, DNA binding gene expression cassettes, selectable marker gene expression cassettes, RNAi construct expression cassettes, or site-specific genome modifying enzyme gene expression cassettes.
12. The recombinant maize B chromosome of claim 2, wherein one or more site-specific genome-modifying enzymes are endonucleases selected from the group consisting of: type I CRISPR-Cas system, type II CRISPR-Cas system and type III CRISPR-Cas system.
13. The recombinant maize B chromosome as described in claim 2, wherein one or more site-specific genome-modifying enzymes are endonucleases selected from the group consisting of: Cpf1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, and Csf4 nucleases.
14. The recombinant maize B chromosome as described in claim 2, wherein one or more site-specific genome-modifying enzymes are dCas9-recombinase fusion proteins.
15. The recombinant maize B chromosome as described in claim 2, wherein one or more site-specific genome-modifying enzymes are tyrosine recombinase, serine recombinase, Cre recombinase, Flp recombinase, Tnp1 recombinase, PhiC31 integrase, R4 integrase, or TP-901 integrase.
16. The method of claim 9, wherein the site-specific genome-modifying enzyme is an endonuclease selected from the group consisting of: type I CRISPR-Cas system, type II CRISPR-Cas system and type III CRISPR-Cas system.
17. The method of claim 9, wherein the site-specific genome-modifying enzyme is an endonuclease selected from the group consisting of: Cpf1, Cas1, Cas1B, Cas2, Cas3, Cas4, Cas5, Cas6, Cas7, Cas8, Cas9, Cas10, Csy1, Csy2, Csy3, Cse1, Cse2, Csc1, Csc2, Csa5, Csn2, Csm2, Csm3, Csm4, Csm5, Csm6, Cmr1, Cmr3, Cmr4, Cmr5, Cmr6, Csb1, Csb2, Csb3, Csx17, Csx14, Csx10, Csx16, CsaX, Csx3, Csx1, Csx15, Csf1, Csf2, Csf3, and Csf4 nucleases.
18. The method of claim 9, wherein the site-specific genome-modifying enzyme is a dCas9-recombinase fusion protein.
19. The method of claim 9, wherein the site-specific genome-modifying enzyme is a tyrosine recombinase, a serine recombinase, a Cre recombinase, an Flp recombinase, a Tnp1 recombinase, a PhiC31 integrase, an R4 integrase, or a TP-901 integrase.
Citation Information
Patent Citations
Synthetic plant genes and method for preparation
EP0385962A1
CORESET and QCL association in beam recovery procedure
US11751183B2
Apparatus for the condensation of volatile metals such as zinc and the like
US1530154A
Methods of optimizing substrate pools and biosynthesis of poly-beta-hydroxybutyrate-co-poly-beta-hydroxyvalerate in bacteria and plants
US20030028917A1
Novel glyphosate N-acetyl transferase (GAT) genes
US20030083480A1