Centromere nucleating sequence
Compact centromere nucleating sequences, like ZeppL-LINE1 retrotransposons, facilitate stable chromosomal segregation in eukaryotic cells, addressing inefficiencies in existing centromere engineering methods and enhancing genome engineering and synthetic biology applications.
Patent Information
- Application Number
- PCT/US2025/034279
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-06-09
- Filing Date
- 2025-06-18
- Publication Date
- 2025-12-26
AI Technical Summary
Current approaches for engineering functional centromeres in artificial chromosomes are cumbersome and inefficient, relying on large repetitive arrays and fusion protein-mediated tethering, and there is a lack of well-defined sequences capable of nucleating centromeres for accurate mitotic and meiotic segregation.
The use of compact, well-defined centromere nucleating sequences, such as single retrotransposon elements like ZeppL-LINE1, to support centromere formation and stable chromosomal segregation in eukaryotic cells, independent of large repetitive arrays and epigenetic mechanisms.
Enables stable mitotic and meiotic segregation of artificial chromosomes by simplifying artificial chromosome construction and ensuring accurate transmission of nucleic acid constructs, expanding applications in genome engineering and synthetic biology.
Smart Images

Figure US2025034279_26122025_PF_FP_ABST
Abstract
Description
PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CENTROMERE NUCLEATING SEQUENCE GOVERNMENTAL RIGHTS
[0001] This invention was made with government support under PGRP 2151105 and PGRP 2151106 awarded by the National Science Foundation. The government has certain rights in the invention. CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority from Provisional Application number 63 / 661,079, filed June 18, 2024 and Provisional Application number 63 / 820,066, filed June 9, 2025. The contents of each of the aforementioned applications are hereby incorporated by reference in their entirety. INCORPORATION OF SEQUENCE LISTING
[0003] This present application contains a Sequence Listing which has been submitted in .XML format via PatentCenter and is hereby incorporated by reference in its entirety. Said WIPO Sequence Listing was created on August 25, 2022 is named DDPSC0166_PCT_SEQUENCE_LISTING and is 170 kilobytes in size. FIELD OF THE INVENTION
[0004] The present disclosure relates to centromere nucleating sequences and nucleic acid vectors and methods for stable mitotic and meiotic segregation in cells. BACKGROUND OF THE INVENTION
[0005] Artificial chromosomes offer significant potential for genetic engineering, enabling the introduction of complex DNA constructs and multi-gene pathways into eukaryotic cells. However, a major challenge in their development lies in engineering functional centromeres, which are essential for accurate segregation during mitosis and meiosis. Centromeres arePROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web typically characterized by large repetitive arrays and epigenetic features, making them difficult to manipulate and integrate into synthetic constructs.
[0006] The size and complexity of centromere-nucleating sequences further complicate artificial chromosome design. Current approaches often rely on extensive arrays of native centromeric repeats or fusion protein-mediated tethering, both of which are cumbersome and inefficient. Additionally, identifying specific, precisely defined sequences capable of nucleating centromeres remains a significant obstacle, as most centromeres are specified epigenetically rather than by distinct DNA sequences.
[0007] Accordingly, there is a clear need for compact, well-defined sequences that can reliably nucleate centromeres and support the formation of functional centromeric chromatin. Such sequences would simplify artificial chromosome construction, reduce reliance on large repetitive arrays, and enable broader applications in genome engineering and synthetic biology. SUMMARY OF THE INVENTION
[0008] One aspect of the instant disclosure encompasses a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell. The DNA construct comprises a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof. The centromere nucleating sequence is capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell. The centromere nucleating sequence can be as described herein above. In some aspects, the centromere nucleating sequence comprises a single Long Interspersed Nuclear Element (LINE)-like retrotransposon or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence comprises a single Zepp, LINE-1, Tad, CRE, Deceiver, Inkcap-like, Ty5, TRAS1, SART1, GilM, GilT, or Zorro element or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence comprises a single ZeppL-LINE1 element or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence is derived from the ZeppL element of chromosome 2b of Chlamydomonas reinhardtii strain UL1690. In some aspects, 2PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web the ZeppL element comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity of SEQ ID NO: 5. In some aspects, the ZeppL element comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity of SEQ ID NO: 5.
[0009] The DNA construct can further comprise chromatin components assembled into chromatin. In some aspects, the DNA construct can further comprise chromatin components assembled into centromeric chromatin. In some aspects, the DNA construct can further comprise centromeric chromatin comprising a centromere-specific histone H3 (CENH3) protein. The CENH3 protein can be a CENH3.1 paralog, a CENH3.2 paralog of Chlamydomonas reinhardtii CENH3, or both.
[0010] In some aspects, the CENH3.1 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2. In some aspects, the CENH3.1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2. In some aspects, the CENH3.1 protein is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 1. In some aspects, the CENH3.1 protein is encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 1.
[0011] In some aspects, the CENH3.2 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 4. In some aspects, the CENH3.2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a 3PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web nucleic acid sequence of SEQ ID NO: 4. In some aspects, the CENH3.2 protein is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 3. In some aspects, the CENH3.2 protein is encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 3.
[0012] The DNA construct can be an artificial chromosome, a mini-chromosome, or a plasmid. In some aspects, the DNA construct is an artificial linear chromosome, an artificial circular chromosome, or a mini-chromosome. In some aspects, the DNA construct further comprises one or more episomal element components selected from telomeres, replication origin, autonomous replicating sequence (ARS), selectable markers, cloning sites, polynucleotides of interest, or any combination thereof. The DNA construct can further comprise a polynucleotide of interest. In some aspects, the polynucleotides of interest is an expression construct, a selectable marker, a cloning site, or any combination thereof.
[0013] The eukaryotic cell can be a chlorophyte green algae. In some aspect, the chlorophyte green algae is a Chlamydomonas sp. The Chlamydomonas sp. can be Chlamydomonas reinhardtii. In some aspects, the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690.
[0014] An additional aspect of the instant disclosure encompasses an artificial chromosome comprising a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof and one or more chromatin components assembled into chromatin on the centromere nucleating sequence as a nucleosome. The centromere nucleating sequence is capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell. The centromere nucleating sequence can be as described herein above. In some aspects, the one or more components of a nucleosome comprises a CENH3 protein. 4PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web
[0015] One aspect of the instant disclosure encompasses a eukaryotic cell comprising a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell. The eukaryotic cell can be a chlorophyte green algae. The DNA construct can be as described herein above.
[0016] In some aspect, the chlorophyte green algae is a Chlamydomonas sp. The Chlamydomonas sp. can be Chlamydomonas reinhardtii. In some aspects, the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690.
[0017] An additional aspect of the instant disclosure encompasses a cell culture comprising a population of eukaryotic cells comprising the DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell. The DNA construct can be as described herein above.
[0018] In some aspects, the eukaryotic cell is a chlorophyte green algae. In some aspect, the chlorophyte green algae is a Chlamydomonas sp. The Chlamydomonas sp. can be Chlamydomonas reinhardtii. In some aspects, the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690.
[0019] Another aspect of the instant disclosure encompasses a method for generating an artificial chromosome. The method comprises providing or having provided a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell. The DNA construct comprises a centromere nucleating sequence and one or more episomal element components. The centromere nucleating sequence comprises a single retrotransposon element or derivatives or truncated variants thereof, wherein the centromere nucleating sequence is capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell. The one or more episomal element components are selected from telomeres, replication origin, autonomous replicating sequence (ARS), selectable markers, cloning sites, polynucleotides of interest, or any combination thereof. The method further comprises introducing the DNA construct into a eukaryotic cell and allowing sufficient time and providing conditions sufficient for centromeric chromatin to assemble at the centromere nucleating sequence. The centromere nucleating sequence, the DNA constructs, and the Eukaryotic cell can be as described herein above.
[0020] Yet another aspect of the instant disclosure encompasses a method of introducing a polynucleotide of interest into a eukaryotic cell for faithful transmission of the polynucleotide of 5PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web interest during mitotic and meiotic division of a eukaryotic cell. The method comprises providing or having provided a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell and introducing the DNA construct into a eukaryotic cell under conditions sufficient to allow centromeric chromatin to assemble at the centromere nucleating sequence, thereby enabling faithful transmission of the DNA construct during mitotic and meiotic division. The DNA construct comprises a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof and a polynucleotide of interest. The centromere nucleating sequence is capable of promoting centromere formation and supporting stable chromosomal segregation. In some aspects, the DNA construct is an artificial chromosome. The centromere nucleating sequence, the DNA constructs, and the Eukaryotic cell can be as described herein above.
[0021] One aspect of the instant disclosure encompasses a kit for introducing a polynucleotide of interest into a eukaryotic cell for faithful transmission of the polynucleotide of interest during mitotic and meiotic division of a eukaryotic cell. The kit comprises one or more DNA constructs capable of stable mitotic and meiotic segregation in a eukaryotic cell, the DNA construct comprising a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof; a eukaryotic optionally further comprising one or more of the DNA constructs; or both. The centromere nucleating sequence, the DNA constructs, and the Eukaryotic cell can be as described herein above.
[0022] In some aspects, the eukaryotic cell is a chlorophyte green algae. In some aspect, the chlorophyte green algae is a Chlamydomonas sp. The Chlamydomonas sp. can be Chlamydomonas reinhardtii. In some aspects, the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690. BRIEF DESCRIPTION OF THE FIGURES
[0023] The following drawings form part of the present specification and are included to further demonstrate certain embodiments of the present disclosure. Certain embodiments can be better understood by reference to one or more of these drawings in combination with the detailed description of specific embodiments presented herein. The patent or application file 6PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web contains at least one drawing executed in color. Copies of this patent or patent application publication with color drawing(s) will be provided by the Office upon request and payment of the necessary fee.
[0024] FIG.1 Expression patterns of CENH3.1 (Cre16.g661450) and CENH3.2 (Cre02.g104800) with representative replication-dependent histones H3 (Cre17.g713950), H4 (Cre17.g714000), H2A (Cre17.g714100), and H2B (Cre17.g714050) for comparison. Data from previously published expression estimates of synchronous diurnal wild-type cultures (Strenkert et al., 2019) were used to generate the figures. Light phase and dark phase are indicated by white and gray backgrounds, respectively. DNA replication and cell division (S / M phase) occur at the beginning of the dark period concurrent with expression of replication dependent histones. ZT, zeitgeber time. (a) The maximum expression for each gene was normalized to 1.
[0025] FIG.2 FPKM (Fragments Per Kilobase of transcript per Million mapped reads) plots without internal normalization show relative expression magnitudes of each gene.
[0026] FIG.3 Identification of two Chlamydomonas CENH3 paralogs. (a) Multiple sequence alignment of CENH3 and conventional histone H3 proteins from Chlamydomonas (Cr), Arabidopsis thaliana (At) and Zea mays (Zm) (Draizen et al., 2016; Talbert et al., 2002; Zhong et al., 2002). The alignment is colored to show conserved residues. Lines above the alignments show peptide sequences from CrCENH3.1 or CrCENH3.2 to which antibodies were raised and purified.
[0027] FIG.4 Immunoblots of SDS-PAGE separated protein lysates of wild-type using a- CENH3.1-specific, a-CENH3.2-specific, a-CENH3-conserved, or a-histone H3 (internal loading control). The black arrowhead shows the position of histone H3. The red arrowheads show the position of CENH3 proteins. Asterisks mark crossreacting non-specific antigens.
[0028] FIG.5 Full immunoblots of SDS-PAGE separated whole cell protein lysates (W) and nuclear protein enriched protein lysates (N) of wild type using α-CENH3.1-specific, α- CENH3.2-specific, α-CENH3-conserved, or α-histone H3 (internal loading control). The black arrowhead shows the position of histone H3. The red arrowheads show the position of CENH3 proteins. Asterisks mark cross-reacting non-specific antigens. Coomassie Blue and Silver 7PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web staining of parallel gels serve as loading controls for total input and protein profiles. L denotes the ladder.
[0029] FIG.6 Immunofluorescence microscopy images of wild-type UL-1690 immunostained with CENH3.1-specific antibodies (pseudo colored green) and stained with DAPI (pseudo colored blue). Background autofluorescence and residual chlorophyll is pseudo colored red. Merge contains an overlay of the three channels. Scale bar = 5 lm.
[0030] FIG.7 Immunofluorescence microscopy images of strain UL-1690. Cells were immunostained with CENH3.1-specific, CENH3.2-specific, CENH3-conserved, or IgG (negative control) antibodies and imaged as in FIG.6. Scale bar = 20 μm.
[0031] FIG.8 CrCENH3.1 and CrCENH3.2 are functionally redundant. Schematic of CENH3.1 and CENH3.2 loci and Phytozome v6.1 IDs with location of targeted AphVIII marker insertion guided by CRISPR-Cas9 cutting at the indicated gRNA sites. Black rectangles, exons; dark green rectangles, untranslated regions; black lines, introns, and intergenic regions. Red arrows show target sites of gRNA used to generate cenh3.1–2 and cenh3.2–2, respectively. Inserted AphVIII selectable marker conferring paromomycin resistance is in blue. Note that the CENH3.1 (Cre16.g661450) gene model was missing UTRs in Phytozome V6.1, so the UTR coordinates were taken from V5.6 gene predictions. The CENH3.2 (Cre02.g104800) gene model was correctly predicted in V6.1, but not in V5.6.
[0032] FIG.9 Genotyping AphVIII insertions in cenh3.1 and cenh3.2 mutants. Schematic of CENH3.1 and CENH3.2 loci with location of inserted paromomycin resistance markers (AphVIII in blue) and locations of genotyping PCR primer binding sites indicated by black arrows. cenh3.1–2 was created by insertion of AphVIII in the 2nd exon of CENH3.1 and both borders could be amplified (primer sets CENH3.1 F1 / Aph8 R and CENH3.1 R1 / Aph8 F). cenh3.2–2 was created by an insertion of AphVIII in the 2nd exon. The left junction of cenh3.2–2 was successfully amplified (primer set CENH3.2 F1 / Aph8 R), but not the right junction. The 5th exon and 3′UTR region of cenh3.2–2 could be amplified indicating that there is not a deletion extending outside of the locus (primer set CENH3.2 F2 / CENH3.2 R2). 8PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web
[0033] FIG.10 Immunoblots of SDS-PAGE separated protein lysates of wild type or independent cenh3 mutants (g1, g2, g3 and g5, g6, g7) detected using a-CENH3.1- specific, a- CENH3.2-specific, a-CENH3-conserved, or a-histone H3 (internal loading control) antibodies. The black arrowhead shows the position of histone H3. The red arrowheads show the position of CENH3 proteins. Asterisks mark cross-reacting non- specific antigens. Note that the a-CENH3-conserved immunoblot signal is reduced and somewhat weak in the mutant strains.
[0034] FIG.11 Meiotic segregation ratios of progeny from cenh3.1–2 crossed to wild-type strain (CC-1691 / 6145c), cenh3.2–2 crossed to wild-type strain (CC-1691 / 6145c).
[0035] FIG.12 Immunoblots of transgenic rescue strains for cenh3.1–2.
[0036] FIG.13 Cenh3.1–2 outcrossed with cenh3.2–2 mutant. Chi squared tests of each cross are shown below the segregation data.
[0037] FIG.14 Maps of constructs used to rescue cenh3.1–2 and cenh3.2–2 mutants with native genomic CENH3.1 or CENH3.2 genes respectively.
[0038] FIG.15 Immunoblots of transgenic rescue strains for cenh3.2–2.
[0039] FIG.16 Genotyping for indicated strains to detect presence / absence of mutant and wild type CENH3.1 and CENH3.2 alleles in rescued cenh3.1–2 cenh3.2–2 double mutant strains (the first two lanes).
[0040] FIG.17A Comparison of new Chlamydomonas genome assemblies for UL-1690 and CC-400 against gap-free reference assembly CC-5816. Plot of the 17 largest scaffolds (corresponding to 17 chromosomes) from UL-1690 HiFi aligned against the CC-5816 assembly using dotPlotly. Inset table shows the mean percent identity score for each chromosome.
[0041] FIG.17B Comparison of new Chlamydomonas genome assemblies for UL-1690 and CC-400 against gap-free reference assembly CC-5816. Plot of the 17 largest scaffolds (corresponding to 17 chromosomes) from CC-400 HiFi aligned against the CC-5816 assembly using dotPlotly (https: / / github.com / tpoorten / dotPlotly / blob / master / pafCoordsDotPlotly.R). Inset table shows the mean percent identity score for each chromosome. 9PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web
[0042] FIG.18 Comparison of five Chlamydomonas genome assemblies with regions common to all assemblies shown in dark blue and locations of long tandem repeat arrays and gaps shown in red and light blue, respectively. The x axis indicates chromosome lengths in bp.
[0043] FIG.19 CENH3 localization in the UL-1690 genome assembly. For each chromosome (represented as a black bar), six tracks from bottom to top indicate protein-coding gene density, ZeppL repeats, and normalized CUT&Tag enrichment (multiple plus unique mapping) for CENH3.1-specific, CENH3.2-specific, and CENH3-conserved antibodies. Gray regions of the assembly represent large repeat blocks missing from the CC- 4532V6.1 reference assembly hosted on Phytozome. The y axes for fold enrichment of CENH3.1-specific, CENH3.2- specific, and CENH3-conserved peaks range from 0 to 75. The y axes for ZeppL-LINE1 enrichment and gene density plot ranges from 0 to 15 and 0 to 6. The window size for all tracks is 5 kb.
[0044] FIG.20A Genome browser views of selected centromere regions showing details of CENH3.1 CUT&Tag read distribution. Blue or red bars above CENH3 peaks show upper (blue) and lower (red) size estimates for the centromere based on combined multimapping plus unique mapping reads or only unique mapping reads, respectively (see Materials and Methods section for details). Example of an averaged-size centromere from chr9. (see main text for details). FIG.20B Genome browser views of selected centromere regions showing details of CENH3.1 CUT&Tag read distribution. Blue or red bars above CENH3 peaks show upper (blue) and lower (red) size estimates for the centromere based on combined multimapping plus unique mapping reads or only unique mapping reads, respectively (see Materials and Methods section for details). Left panel, genetically defined centromere region for chr2. Right panel, possible neocentromere region on chr2 (see main text for details).
[0045] FIG.21A Relationships between ZeppL-enriched region size, centromere size and chromosome size for Chlamydomonas strain UL-1690. Upper centromere size estimates plotted against the size of ZeppL enriched regions on each chromosome with linear correlation and R2 value. See Materials and Methods for details on size estimations of centromeres and ZeppL repeat regions. 10PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web
[0046] FIG.21B Relationships between ZeppL-enriched region size, centromere size and chromosome size for Chlamydomonas strain UL-1690. Lower centromere size estimate plotted against ZeppL enriched region sizes shows no correlation. See Materials and Methods for details on size estimations of centromeres and ZeppL repeat regions.
[0047] FIG.21C Relationships between ZeppL-enriched region size, centromere size and chromosome size for Chlamydomonas strain UL-1690. ZeppL repeat region size plotted against chromosome size shows a weak positive correlation. See Materials and Methods for details on size estimations of centromeres and ZeppL repeat regions.
[0048] FIG.22A Genome browser view of CENH3 and IgG CUT&Tag read distribution around the chr2 neocentromere region. Upper browser panel shows CC-4532 genome at neocentromere region without the ZeppL insertion. Lower panel shows the UL-1690 genome at the neocentromere site with the ZeppL insertion and mapping data for CUT&Tag reads. Annotation is the same as FIG.20A and FIG.20B.
[0049] FIG.22B Genome browser view of CENH3 and IgG CUT&Tag read distribution around the chr2 neocentromere region. Genotyping the junction region on chr2 flanking the ZeppL insertion on chr2. Upper panel shows a schematic of insertion between genes Cre02.g800275 and Cre02.g141600 with genotyping primer locations indicated. ZeppL fragment (8334 bp) schematic is not scaled to actual size. Upper gel shows PCR products using the genotyping primers depicted in the above diagram. Lower panel shows genotyping of the same samples with primers for the CENH3.1 locus (CENH3.1 F1 / CENH3.1 R1 in Table 7) as a template quality control.
[0050] FIG.23 Phylogenetic analysis of near full-length ZeppL sequences (>8 kb) in four genome assemblies. Unrooted Neighbor Joining tree of 49 full-length or near full-length ZeppL elements in four strains. Bootstrap values are shown at each node. Each shape (square, circle, triangle, and diamond) represents a specific strain. Each color represents a specific ZeppL element. Filled shapes indicate an intact ZeppL element and hollow shapes represent a partially decayed copy (between 5 and 8 kb) corresponding to a full-length ZeppL element at the same location in other strains. Dashes indicate absence of a ZeppL element at that location in one or more strains. nd indicates no data due to an assembly gap in that strain. The 11PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web lower portion of the tree contains distinct, centromere-specific ZeppL elements found in the same position within the centromere in all strains where there was a gapless assembly at that centromere. The upper part of the tree contains a group of highly similar, recently active ZeppL elements (near-zero branch lengths) shared across several chromosomes, including the newly inserted element in chr2 (boxed; see main text for details). Full-length ZeppL sequences and coordinates can be found in Data S2 of Liu et al., 2025, The Plant Journal (2025) 122, e70153, the disclosure of all of which is incorporated herein in its entirety. Partly decayed ZeppL sequences and coordinates can be found in Data S3 of Liu et al., 2025, The Plant Journal (2025) 122, e70153.
[0051] FIG.24 Dot plot comparisons of centromere regions of chromosomes containing one or more full length (>8 kb, red bars) or partially decayed (>5 kb, blue bars) ZeppL elements analyzed in FIG.23. Black blocks and grey blocks along each axis represent ZeppL or non- ZeppL sequences, respectively. Centromere coordinates are listed in Table 6.
[0052] FIG.25 Dot plot comparisons of centromere regions of chromosomes containing one or more full length (>8 kb, red bars) or partially decayed (>5 kb, blue bars) ZeppL elements analyzed in FIG.23. Black blocks and grey blocks along each axis represent ZeppL or non- ZeppL sequences, respectively. Centromere coordinates are listed in Table 6.
[0053] FIG.26 Dot plot comparisons of centromere regions of chromosomes containing one or more full length (>8 kb, red bars) or partially decayed (>5 kb, blue bars) ZeppL elements analyzed in FIG.23. Black blocks and grey blocks along each axis represent ZeppL or non- ZeppL sequences, respectively. Centromere coordinates are listed in Table 6.
[0054] FIG.27 Dot plot comparisons of all CC-5816 centromeres (x axes) against corresponding centromere regions on y axes from UL-1690. These plots provide a comprehensive comparison of all centromere regions.
[0055] FIG.28 Dot plot comparisons of all CC-5816 centromeres (x axes) against corresponding centromere regions on y axes from CC-1690. These plots provide a comprehensive comparison of all centromere regions. 12PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web
[0056] FIG.29 Dot plot comparisons of all CC-5816 centromeres (x axes) against corresponding centromere regions on y axes from CC-400. These plots provide a comprehensive comparison of all centromere regions.
[0057] FIG.30 Side by side dot plot comparisons of centromere regions in chr10 and chr15 to illustrate possible structural differences in the CC-400 assembly compared to other assemblies. Annotation is the same as in Figures 27–29, where both chr10 and chr15 of CC- 400 have noticeably smaller ZeppL enriched regions. Alignment regions were extended to 200 kb on chromosome 10 and 430 kb on chromosome 15 to capture structural variations in CC-400. This extended alignment identified a ~400 kb inversion near the centromeric region of chromosome 15 in CC-400 which was confirmed manually by inspecting PacBio HiFi reads spanning the inversion junctions.
[0058] FIG.31 Alignment-free midpoint-rooted distance tree of all full-length centromere sequences from strains UL-1690, CC-1690, CC-400, and CC-5816 generated using Mash (Ondov et al., 2016) (see Materials and Methods section for details). Labels show strain name followed by chromosome number. Relative distance scale is at the bottom.
[0059] FIG.32A Functional neocentromere formed by single Zepp insertion. Upper panel shows a diagram of strain UL1690 chromosome 2 split into two sub-chromosomes after fortuitous Zepp transposition. Original chr2 assembly as a single chromosome with two regions of CENH3 occupancy at native and neo-centromeres. Lower panel Zoomed view of neo- centromere region with CENH3 occupancy in Zepp flanking regions.
[0060] FIG.32B Revised assembly of UL1690 chr2 sequences after splitting into two sub- chromosomes (2a, 2b), each with its own centromere and new telomeres at the break points. This revision is unequivocally supported by long-read sequencing. DETAILED DESCRIPTION
[0061] The present disclosure relates to the surprising and unexpected discovery of short, well defined centromere nucleating sequences capable of supporting centromere formation and faithful transmission of nucleic acid constructs during mitotic and meiotic division in eukaryotic 13PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web cells. More specifically, the inventors discovered that a single ZeppL-LINE1 (ZeppL) retrotransposon insertion can create a new functional centromere. The centromere nucleating sequences of the instant disclosure are compact, well-defined, and capable of nucleating centromeres to support stable chromosomal segregation. Unlike traditional approaches that rely on large repetitive arrays or complex epigenetic mechanisms, this invention utilizes short, defined sequences to simplify artificial chromosome construction and ensure accurate segregation. The disclosure further describes DNA constructs and episomal elements incorporating these centromere nucleating sequences, along with methods for their creation and use. By providing a streamlined approach to centromere formation, the instant disclosure addresses the challenges of centromere engineering and expands possibilities for synthetic biology, genome engineering, and biotechnology applications across diverse eukaryotic systems. I. Compositions
[0062] One aspect of the instant disclosure encompasses centromere nucleating sequences and DNA constructs comprising the centromere nucleating sequences capable of stable mitotic and meiotic segregation in a eukaryotic cell. The centromere nucleating sequence comprises a retrotransposon element that was discovered by the inventors to exhibit neocentromeric properties. Despite their small size and simple construction, the centromere nucleating sequences of the instant disclosure, when included in nucleic acid constructs, are capable of nucleating centromeres to support stable chromosomal segregation of the nucleic acid constructs during mitotic and meiotic division. The centromere nucleating sequence, retrotransposons, and nucleic acid constructs comprising the centromere nucleating sequences are described herein below. (a) Centromere nucleating sequence
[0063] One aspect of the instant disclosure encompasses centromere nucleating sequences. As used herein, the term "centromere nucleating sequence" refers to a compact, well-defined nucleic acid sequence capable of supporting the formation of centromeric chromatin that 14PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web comprise centromere-specific proteins, such as the histone H3 variant CENH3. This nucleation process involves the establishment of a specialized chromatin environment at the site of the sequence, which serves as the foundation for kinetochore assembly and spindle attachment during mitotic and meiotic division. Unlike traditional centromeres, which rely on large repetitive arrays and epigenetic mechanisms for centromere specification, the centromere nucleating sequences of the instant disclosure are sufficient to nucleate centromeres independently, enabling stable chromosomal segregation and faithful transmission of nucleic acid constructs.
[0064] Accordingly, in some aspects, centromere nucleating sequences of the instant disclosure further comprise chromatin components assembled into chromatin. In some aspects, a centromere nucleating sequence of the instant disclosure further comprises chromatin components assembled into centromeric chromatin. Chromatin is the complex of DNA and associated proteins, primarily histones, that packages and organizes genetic material within the nucleus of eukaryotic cells. The structure and composition of chromatin can vary to regulate access to the underlying DNA and to support specific cellular functions. In centromeric regions, chromatin is specialized to form "centromeric chromatin," which is distinguished by the presence of a centromere-specific histone H3 variant, such as CENH3 (also known as CENP- A). This specialized chromatin environment is essential for the assembly of the kinetochore and for proper chromosome segregation during cell division. In some aspects, a centromere nucleating sequence of the instant disclosure further comprises centromeric chromatin comprising a centromere-specific histone H3 (CENH3) protein. In some aspects, the CENH3 protein is a CENH3.1 paralog, a CENH3.2 paralog of Chlamydomonas CENH3, or both.
[0065] In some aspects, the CENH3.1 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2. In some aspects, the CENH3.1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2. In some aspects, the CENH3.1 protein is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 15PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 1. In some aspects, the CENH3.1 protein is encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 1.
[0066] In some aspects, the CENH3.2 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 4. In some aspects, the CENH3.2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 4. In some aspects, the CENH3.2 protein is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 3. In some aspects, the CENH3.2 protein is encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 3.
[0067] Importantly, the inventors discovered that a single retrotransposon, absent repetitive DNA, can function as a centromere nucleating sequence. Accordingly, in some aspects, a centromere nucleating sequence of the instant disclosure comprises a single retrotransposon element or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence does not include other repetitive sequences or tandemly arranged fragments of transposons typically found in centromeres. Retrotransposons can be as described in Section I(b) herein below.
[0068] The centromere nucleating sequences of the instant disclosure are shorter and more defined than centromeres normally present in eukaryotic cells. For instance, normal centromeres in Chlamydomonas sp. are composed of many retrotransposon repeats interspersed with satellite DNA and other repetitive elements and range from about 63.5 kb to about 175 kb in size. The interplay between retrotransposons and satellite DNA in these regions contributes to the structural complexity and epigenetic specification of centromeres, 16PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web making them challenging to manipulate and integrate into synthetic constructs. Conversely, centromere nucleating sequences of the instant disclosure can comprise a single retrotransposon element and comprise a size ranging from about 7.5 kb or less to about 9.5 kb. In some aspects, a centromere nucleating sequence of the instant disclosure ranges from about 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, or about 9.5 kb or less. In some aspects, a centromere nucleating sequence of the instant disclosure ranges from less than about 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, or about 9.5 kb. In some aspects, a centromere nucleating sequence of the instant disclosure is no more than about 5, 5.5, 6, 6.5, 7, 7.5, 8, 8.5, 9, or about 9.5 kb or less. In some aspects, a centromere nucleating sequence of the instant disclosure ranges from about 7.5, to about 9.5 kb or less. (b) Retrotransposons
[0069] The centromere nucleating sequences of the instant disclosure comprise a retrotransposon element. Transposons are mobile genetic elements that are widespread across eukaryotic genomes and are often found in centromeric regions, where they contribute to chromatin structure and genome organization. DNA transposons, which move via a "cut- and-paste" mechanism, have been identified in the centromeric regions of certain organisms. In Drosophila melanogaster, for instance, DNA transposons such as the P-element and hobo elements are enriched in heterochromatic regions adjacent to centromeres, though they are not typically found in the functional core centromere. In plants like Arabidopsis thaliana, DNA transposons including CACTA elements are found interspersed with retrotransposons in heterochromatic centromeric domains, contributing to the complex repetitive landscape of these regions.
[0070] Retrotransposons, which replicate through an RNA intermediate using a "copy-and- paste" mechanism, are the predominant transposable elements within functional centromeric domains, particularly in plants. Their ability to form dense, tandem arrays enables them to provide structural scaffolding that supports centromeric chromatin. In Zea mays (maize), centromeric retrotransposons known as CRM elements (type of Ty3-gypsy LTR retrotransposons) are specifically enriched in centromeric regions and physically associate 17PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web with the centromere-specific histone CENH3 to contribute to the formation of active centromeres. In Chlamydomonas reinhardtii, Zepp elements - non-LTR retrotransposons - are similarly enriched in centromeres and are implicated in centromere identity and structure. In contrast, in humans (Homo sapiens), the primary DNA component of centromeres is alpha- satellite DNA. LINE-1 elements (non-LTR retrotransposons) are occasionally found in centromeric regions but do not play a clearly defined structural or functional role in centromere formation.
[0071] LTR retrotransposons are prominent among the retrotransposons found in centromeres, especially in plants. These elements, characterized by their long terminal repeats (LTRs), are capable of forming tandem arrays that help organize centromeric chromatin. For example, CRM elements in maize are well-studied LTR retrotransposons that bind to CENH3 and contribute directly to the formation of functional centromeres. Other Ty3-gypsy LTR retrotransposons have been identified in the centromeres of various plant species, where they serve as essential components of the centromeric architecture. Their repetitive nature and capacity to be epigenetically marked make them key contributors to centromere specification in many plant systems.
[0072] Non-LTR retrotransposons, including LINE (Long Interspersed Nuclear Element) families, are also widespread in eukaryotic genomes. While they are less commonly associated with functional centromeres than LTR elements, they are often enriched in pericentromeric regions and other heterochromatic domains. For example, LINE-1 elements are occasionally found adjacent to human centromeres, though their functional significance in these regions remains unclear. In Chlamydomonas, Zepp elements, a class of LINE-like retrotransposons, are consistently found in centromeric regions, suggesting a potential structural or epigenetic role. Non-limiting examples of non-LTR retrotransposons (LINE-like elements) other than Zepp include Tad, CRE, Deceiver and Inkcap-like, Ty5, TRAS1 and START1, GilM and GilT, and Zorro, among others. LINE-1 (L1) elements, found in mammals, are autonomous and capable of active retrotransposition, with human-specific L1Hs elements representing one of the few lineages still retrotransposition-competent today. The Tad element, present in both insects and fungi, belongs to an ancient LINE lineage. In fungi, other LINE-like 18PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web elements such as CRE, Deceiver, and Inkcap-like have been identified. In Saccharomyces cerevisiae, Ty5 is a non-LTR element that preferentially inserts into telomeric or silent chromatin regions, a behavior reminiscent of the centromere-associated Zepp elements in Chlamydomonas. In Bombyx mori, TRAS1 and SART1 are LINE-like elements that specifically target telomeric repeats. GilM and GilT, identified in Giardia, display structural similarities to the L1 family. Finally, Zorro3, a member of the L1 clade from Candida albicans, is capable of active retrotransposition and shares mechanistic features with mammalian LINE-1 elements.
[0073] Transposons often interact with satellite DNA to form the complex repetitive structure of centromeres. Satellite DNA consists of tandemly repeated sequences that are a hallmark of centromeric regions. Transposons and fragments thereof interspersed within satellite DNA add to the structural complexity and provide additional sequences that can recruit centromeric proteins, such as CENH3. This interplay between transposons and satellite DNA contributes to the epigenetic specification of centromeres, ensuring their proper function during chromosome segregation. Together, LTR retrotransposons, non-LTR retrotransposons, and their interactions with satellite DNA highlight the diverse and essential roles of transposons in centromere biology.
[0074] Important, before the instant invention was made, it was not known that a single retrotransposon element, absent any other repetitive DNA could function as a centromere nucleating sequence. Accordingly, a retrotransposon of the instant disclosure can be a single retro transposable element or derivatives or truncated variants thereof. In some aspects, a retrotransposon of the instant disclosure comprises a single LTR or non-LTR retrotransposon or derivatives or truncated variants thereof. In some aspects, a retrotransposon of the instant disclosure comprises a single CRM, Ty3-gypsy, Zepp, LINE-1, Tad, CRE, Deceiver, Inkcap- like, Ty5, TRAS1, SART1, GilM, GilT, or Zorro element or derivatives or truncated variants thereof. In some aspects, a retrotransposon of the instant disclosure comprises a single LTR retrotransposon or derivatives or truncated variants thereof. In some aspects, a retrotransposon of the instant disclosure comprises a single non-LTR retrotransposon or derivatives or truncated variants thereof. In some aspects, the retrotransposon element is a single Long Interspersed Nuclear Element (LINE)-like retrotransposon. In some aspects, a 19PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web retrotransposon is a single Zepp, LINE-1, Tad, CRE, Deceiver, Inkcap-like, Ty5, TRAS1, SART1, GilM, GilT, or Zorro element or derivatives or truncated variants thereof.
[0075] In some aspects, the retrotransposon element is a single Long Interspersed Nuclear Element (LINE)-like retrotransposon.
[0076] In some aspects, the retrotransposon element is a single Zepp-like (ZeppL) retrotransposon element. In some aspects, the retrotransposon element is a single ZeppL retrotransposon element of a Chlamydomonas sp. In some aspects, the retrotransposon element is a single ZeppL retrotransposon element of Chlamydomonas reinhardtii. In some aspects, the retrotransposon element is a single ZeppL retrotransposon element of Chlamydomonas reinhardtii strain UL1690.
[0077] In some aspects, a retrotransposon element of the instant disclosure comprises a single ZeppL retrotransposon element comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 5. In some aspects, a retrotransposon element of the instant disclosure comprises a single ZeppL retrotransposon element comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 5. (c) DNA constructs
[0078] Another aspect of the instant disclosure encompasses DNA constructs comprising centromere nucleating sequences to enable stable mitotic and meiotic segregation of the DNA constructs in eukaryotic cells. Put another way, DNA constructs of the instant disclosure can function as episomal elements capable of replication and faithful transmission of the DNA constructs during meiosis and mitosis. These constructs leverage the ability of the centromere nucleating sequences of the instant disclosure to promote chromatin formation at the centromere, including centromeric chromatin, ensuring faithful transmission of the nucleic acid constructs during cell division. By utilizing compact and well-defined centromere nucleating sequences, these constructs overcome the challenges associated with traditional centromere 20PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web engineering approaches, providing a streamlined and efficient platform for synthetic biology and genome engineering applications. Centromere nucleating sequences can be as described in Section I(a) herein above.
[0079] As it will be recognized by individuals skilled in the art, the DNA constructs described herein can be designed in various forms, including circular or linear configurations, depending on the intended application. The DNA constructs can be plasmids, which are useful for manipulation and propagation of the constructs in bacteria or can serve as vectors for further engineering. For example, such constructs can be used to build mini-chromosomes or other specialized genetic elements for research or biotechnology purposes. Alternatively, the DNA constructs of the instant disclosure can be designed to function as chromosomes. As used herein, the term "chromosome" (and any term derived therefrom, such as artificial chromosome or mini-chromosome) refers to any DNA construct that comprises a functional centromere sufficient to enable faithful transmission and segregation of the construct during cell division, including both mitosis and meiosis. Such constructs can be linear or circular and can exist as natural, artificial, or engineered entities within a eukaryotic cell.
[0080] In some aspects, the DNA constructs are circular. In some aspects, the nucleic acid constructs of the instant disclosure are linear. In some aspects, the DNA constructs of the instant disclosure are a mini-chromosomes or an artificial chromosome.
[0081] In addition to the centromere nucleating sequence, DNA constructs of the instant disclosure can comprise a number of components that facilitate the function of the DNA constructs depending on the intended use. Non-limiting examples of components that may be included in a DNA construct in addition to the centromere nucleating sequence include telomeres, replication origin, autonomous replicating sequence (ARS), selectable markers, cloning sites, polynucleotides of interest, or any combination thereof.
[0082] In some aspects, a nucleic acid construct if the instant disclosure is an artificial linear chromosome. In some aspects, a nucleic acid construct if the instant disclosure is an artificial circular chromosome. In some aspects, a nucleic acid construct if the instant disclosure is a mini-chromosome. 21PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web
[0083] In some aspects, DNA constructs of the instant disclosure further comprise chromatin components assembled into chromatin at the centromere nucleating sequence. In some aspects, a DNA construct of the instant disclosure further comprises chromatin components assembled into centromeric chromatin. In some aspects, a DNA construct of the instant disclosure further comprises centromeric chromatin comprising a centromere-specific histone H3 (CENH3) protein. In some aspects, the CENH3 protein is a CENH3.1 paralog, a CENH3.2 paralog of Chlamydomonas CENH3, or both. CENH3.1 and CENH3.2 paralogs of Chlamydomonas CENH3 can be as described herein above in Section I(a).
[0084] In some aspects, DNA constructs of the instant disclosure can further comprise one or more polynucleotide of interest. The polynucleotide of interest can be any DNA sequence that a user wishes to introduce into a cell, such as a polypeptide encoding a protein or RNA, a biosynthetic pathway, a regulatory element, or any other functional or reporter sequence. By incorporating the polynucleotide of interest into an artificial chromosome or mini-chromosome, the DNA construct can exist as an episome that is faithfully inherited during cell division, ensuring stable maintenance and expression without integration into the host genome.
[0085] Artificial chromosomes and mini-chromosomes are powerful tools for introducing large or complex genetic payloads into eukaryotic cells. Unlike plasmids, which are often limited in size and may be lost during cell division, artificial and mini-chromosomes can carry extensive genetic material and are designed to segregate accurately with the host chromosomes. This enables long-term, stable inheritance of the introduced DNA. Additionally, artificial chromosomes avoid the risks associated with random integration into the host genome, such as insertional mutagenesis or disruption of endogenous genes, providing a safer and more predictable platform for genetic engineering, therapeutic applications, and synthetic biology. (d) Eukaryotic cell
[0086] The DNA constructs of the instant disclosure can function in any eukaryotic cell. In some aspects, the DNA constructs of the instant disclosure can function in any eukaryotic cell where centromeres can comprise retrotransposons. Eukaryotic cells include a broad range of organisms across multiple kingdoms and major clades, such as animals, fungi, protists, and 22PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web viridiplantae. For example, in the animal kingdom, suitable cells include those from mammals (such as humans and mice), birds, fish, and insects. In the fungal kingdom, examples include yeasts like Saccharomyces cerevisiae and filamentous fungi such as Neurospora crassa. Among protists, suitable examples include protozoa such as Tetrahymena and Paramecium, as well as various microalgae. In the viridiplantae clade, examples include chlorophyte green algae (such as Chlamydomonas reinhardtii and Chlorella species), charophyte green algae, and land plants (embryophytes) such as Arabidopsis thaliana, maize, and rice. This broad applicability enables the use of the disclosed DNA constructs in a wide variety of research, industrial, and therapeutic contexts.
[0087] In some aspects, the eukaryotic cell is a viridiplantae eukaryotic cell. In some aspects, the viridiplantae eukaryotic cell is a charophyte green algae. In some aspects, the viridiplantae eukaryotic cell is a land plant. In some aspects, the viridiplantae eukaryotic cell is a chlorophyte green algae. In some aspects, the chlorophyte green algae is a Chlorella sp. In some aspects, the chlorophyte green algae is a Chlamydomonas sp. In some aspects, the Chlamydomonas sp. is Chlamydomonas reinhardtii. In some aspects, the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL-1690. (e) Aspects
[0088] Another aspect of the instant disclosure encompasses a centromere nucleating sequence capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell. The centromere nucleating sequence comprises a single retrotransposon element or derivatives or truncated variants thereof. The centromere nucleating sequence can be as described in Section I(a) herein above. In some aspects, the centromere nucleating sequence comprises a single Long Interspersed Nuclear Element (LINE)-like retrotransposon or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence comprises a single Zepp, LINE-1, Tad, CRE, Deceiver, Inkcap-like, Ty5, TRAS1, SART1, GilM, GilT, or Zorro element or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence comprises a single ZeppL-LINE1 element or derivatives or truncated variants 23PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web thereof. In some aspects, the centromere nucleating sequence is derived from the ZeppL element of chromosome 2b of Chlamydomonas reinhardtii strain UL1690. In some aspects, the ZeppL element comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity of SEQ ID NO: 5. In some aspects, the ZeppL element comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity of SEQ ID NO: 5.
[0089] The centromere nucleating sequence can further comprise chromatin components assembled into chromatin. In some aspects, the centromere nucleating sequence can further comprise chromatin components assembled into centromeric chromatin. In some aspects, the centromere nucleating sequence can further comprise centromeric chromatin comprising a centromere-specific histone H3 (CENH3) protein. The CENH3 protein can be a CENH3.1 paralog, a CENH3.2 paralog of Chlamydomonas reinhardtii CENH3, or both.
[0090] In some aspects, the CENH3.1 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2. In some aspects, the CENH3.1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2. In some aspects, the CENH3.1 protein is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 1. In some aspects, the CENH3.1 protein is encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 1.
[0091] In some aspects, the CENH3.2 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ 24PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web ID NO: 4. In some aspects, the CENH3.2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 4. In some aspects, the CENH3.2 protein is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 3. In some aspects, the CENH3.2 protein is encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 3.
[0092] Yet another aspect of the instant disclosure encompasses a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell. The DNA construct comprises a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof. The centromere nucleating sequence is capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell. The centromere nucleating sequence can be as described in Section I(a) herein above. In some aspects, the centromere nucleating sequence comprises a single Long Interspersed Nuclear Element (LINE)-like retrotransposon or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence comprises a single Zepp, LINE-1, Tad, CRE, Deceiver, Inkcap-like, Ty5, TRAS1, SART1, GilM, GilT, or Zorro element or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence comprises a single ZeppL-LINE1 element or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence is derived from the ZeppL element of chromosome 2b of Chlamydomonas reinhardtii strain UL1690. In some aspects, the ZeppL element comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity of SEQ ID NO: 5. In some aspects, the ZeppL element comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity of SEQ ID NO: 5. 25PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web
[0093] The DNA construct can further comprise chromatin components assembled into chromatin. In some aspects, the DNA construct can further comprise chromatin components assembled into centromeric chromatin. In some aspects, the DNA construct can further comprise centromeric chromatin comprising a centromere-specific histone H3 (CENH3) protein. The CENH3 protein can be a CENH3.1 paralog, a CENH3.2 paralog of Chlamydomonas reinhardtii CENH3, or both.
[0094] In some aspects, the CENH3.1 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2. In some aspects, the CENH3.1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2. In some aspects, the CENH3.1 protein is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 1. In some aspects, the CENH3.1 protein is encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 1.
[0095] In some aspects, the CENH3.2 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 4. In some aspects, the CENH3.2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 4. In some aspects, the CENH3.2 protein is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 3. In some aspects, the CENH3.2 protein is encoded by a nucleic acid sequence comprising at least 26PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 3.
[0096] The DNA construct can be an artificial chromosome, a mini-chromosome, or a plasmid. In some aspects, the DNA construct is an artificial linear chromosome, an artificial circular chromosome, or a mini-chromosome. In some aspects, the DNA construct further comprises one or more episomal element components selected from telomeres, replication origin, autonomous replicating sequence (ARS), selectable markers, cloning sites, polynucleotides of interest, or any combination thereof. The DNA construct can further comprise a polynucleotide of interest. In some aspects, the polynucleotides of interest is an expression construct, a selectable marker, a cloning site, or any combination thereof.
[0097] The eukaryotic cell can be a chlorophyte green algae. In some aspect, the chlorophyte green algae is a Chlamydomonas sp. The Chlamydomonas sp. can be Chlamydomonas reinhardtii. In some aspects, the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690.
[0098] An additional aspect of the instant disclosure encompasses an artificial chromosome comprising a. a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof and one or more chromatin components assembled into chromatin on the centromere nucleating sequence as a nucleosome. The centromere nucleating sequence is capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell. In some aspects, the centromere nucleating sequence comprises a single Long Interspersed Nuclear Element (LINE)-like retrotransposon or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence comprises a single Zepp, LINE-1, Tad, CRE, Deceiver, Inkcap-like, Ty5, TRAS1, SART1, GilM, GilT, or Zorro element or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence comprises a single ZeppL-LINE1 element or derivatives or truncated variants thereof. In some aspects, the centromere nucleating sequence is derived from the ZeppL element of chromosome 2b of Chlamydomonas reinhardtii strain UL1690. In some aspects, the ZeppL element comprises a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 27PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity of SEQ ID NO: 5. In some aspects, the ZeppL element comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity of SEQ ID NO: 5.
[0099] The centromere nucleating sequence can further comprise chromatin components assembled into chromatin. In some aspects, the centromere nucleating sequence can further comprise chromatin components assembled into centromeric chromatin. In some aspects, the centromere nucleating sequence can further comprise centromeric chromatin comprising a centromere-specific histone H3 (CENH3) protein. The CENH3 protein can be a CENH3.1 paralog, a CENH3.2 paralog of Chlamydomonas reinhardtii CENH3, or both.
[0100] In some aspects, the CENH3.1 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2. In some aspects, the CENH3.1 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2. In some aspects, the CENH3.1 protein is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 1. In some aspects, the CENH3.1 protein is encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 1.
[0101] In some aspects, the CENH3.2 protein comprises at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 4. In some aspects, the CENH3.2 protein comprises at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 4. In some aspects, the CENH3.2 protein is encoded by a nucleic acid sequence comprising at least about 75%, 76%, 77%, 78%, 79%, 80%, 81%, 28PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 82%, 83%, 84%, 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 3. In some aspects, the CENH3.2 protein is encoded by a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 3.
[0102] One aspect of the instant disclosure encompasses a eukaryotic cell comprising a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell. The eukaryotic cell can be a chlorophyte green algae. The DNA construct can be as described herein above.
[0103] In some aspect, the chlorophyte green algae is a Chlamydomonas sp. The Chlamydomonas sp. can be Chlamydomonas reinhardtii. In some aspects, the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690.
[0104] An additional aspect of the instant disclosure encompasses a cell culture comprising a population of eukaryotic cells comprising the DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell. The DNA construct can be as described herein above.
[0105] In some aspect, the chlorophyte green algae is a Chlamydomonas sp. The Chlamydomonas sp. can be Chlamydomonas reinhardtii. In some aspects, the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690. II. Methods
[0106] Another aspect of the instant disclosure encompasses a method for generating an artificial chromosome. The method comprises providing or having provided a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell. The DNA construct comprises a centromere nucleating sequence and one or more episomal element components. The centromere nucleating sequence comprises a single retrotransposon element or derivatives or truncated variants thereof, wherein the centromere nucleating sequence is capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell. The one or more episomal 29PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web element components are selected from telomeres, replication origin, autonomous replicating sequence (ARS), selectable markers, cloning sites, polynucleotides of interest, or any combination thereof. The method further comprises introducing the DNA construct into a eukaryotic cell and allowing sufficient time and providing conditions sufficient for centromeric chromatin to assemble at the centromere nucleating sequence. The centromere nucleating sequence can be as described in Section I(a). The DNA constructs can be as described in Section I(c). The Eukaryotic cell can be as described in Section I(d).
[0107] Yet another aspect of the instant disclosure encompasses a method of introducing a polynucleotide of interest into a eukaryotic cell for faithful transmission of the polynucleotide of interest during mitotic and meiotic division of a eukaryotic cell. The method comprises providing or having provided a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell and introducing the DNA construct into a eukaryotic cell under conditions sufficient to allow centromeric chromatin to assemble at the centromere nucleating sequence, thereby enabling faithful transmission of the DNA construct during mitotic and meiotic division. The DNA construct comprises a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof and a polynucleotide of interest. The centromere nucleating sequence is capable of promoting centromere formation and supporting stable chromosomal segregation. In some aspects, the DNA construct is an artificial chromosome. The centromere nucleating sequence can be as described in Section I(a). The DNA constructs can be as described in Section I(c). The Eukaryotic cell can be as described in Section I(d).
[0108] DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell such as artificial chromosomes are versatile tools in biotechnology, offering a platform for the stable introduction and maintenance of large or complex genetic payloads in eukaryotic cells. They can be engineered to carry multiple genes, regulatory elements, or entire biosynthetic pathways, making them ideal for applications that require the coordinated expression of several genetic components. Because these DNA constructs are designed to replicate and segregate independently of the host genome, they enable long-term, faithful 30PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web inheritance of introduced genetic material without the risks associated with random integration, such as insertional mutagenesis or disruption of endogenous genes.
[0109] In biotechnology, the DNA constructs of the instant disclosure can be used for a wide range of purposes, including the production of recombinant proteins, metabolic engineering, and the development of genetically modified organisms with enhanced traits. They provide a robust platform for gene therapy, allowing for the delivery of therapeutic genes in a manner that supports stable, regulated expression over time. The DNA constructs of the instant disclosure can also be valuable in functional genomics, where they can be used to study gene function, regulatory networks, and chromosomal behavior in a controlled context. Their capacity to accommodate large DNA fragments and complex genetic circuits makes them indispensable for synthetic biology, enabling the construction of novel biological systems and the reprogramming of cellular functions for industrial, agricultural, and medical applications. III. Kits
[0110] Another aspect of the instant disclosure encompasses a kit for introducing a polynucleotide of interest into a eukaryotic cell for faithful transmission of the polynucleotide of interest during mitotic and meiotic division of a eukaryotic cell. The kit comprises one or more DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell, the DNA construct comprising a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof. The centromere nucleating sequence is capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell. The centromere nucleating sequence can be as described in Section I(a). The DNA constructs can be as described in Section I(c).
[0111] In some aspects, the kit further comprises one or more eukaryotic cells. The eukaryotic cells can further comprise the DNA constructs. The Eukaryotic cell can be as described in Section I(d).
[0112] The kits can further comprise transfection reagents, cell growth media, selection media, in-vitro transcription reagents, nucleic acid purification reagents, protein purification 31PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web reagents, buffers, and the like. The kits provided herein generally include instructions for carrying out the methods detailed below. Instructions included in the kits can be affixed to packaging material or can be included as a package insert. While the instructions are typically written or printed materials, they are not limited to such. Any medium capable of storing such instructions and communicating them to an end user is contemplated by this disclosure. Such media include, but are not limited to, electronic storage media (e.g., magnetic discs, tapes, cartridges, chips), optical media (e.g., CD ROM), an internet address that provides the instructions, and the like. As used herein, the term “instructions” can include the address of an internet site that provides the instructions. DEFINITIONS
[0113] Unless defined otherwise, all technical and scientific terms used herein have the meaning commonly understood by a person skilled in the art to which this invention belongs. The following references provide one of skill with a general definition of many of the terms used in this invention: Singleton et al., Dictionary of Microbiology and Molecular Biology (2nd ed.1994); The Cambridge Dictionary of Science and Technology (Walker ed., 1988); The Glossary of Genetics, 5th Ed., R. Rieger et al. (eds.), Springer Verlag (1991); and Hale & Marham, The Harper Collins Dictionary of Biology (1991). As used herein, the following terms have the meanings ascribed to them unless specified otherwise.
[0114] When introducing elements of the present disclosure or the preferred aspects(s) thereof, the articles "a", "an", "the" and "said" are intended to mean that there are one or more of the elements. The terms "comprising", "including" and "having" are intended to be inclusive and mean that there may be additional elements other than the listed elements.
[0115] As various changes could be made in the above-described cells and methods without departing from the scope of the invention, it is intended that all matter contained in the above description and in the examples given below, shall be interpreted as illustrative and not in a limiting sense. 32PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web EXAMPLES
[0116] All patents and publications mentioned in the specification are indicative of the levels of those skilled in the art to which the present disclosure pertains. All patents and publications are herein incorporated by reference to the same extent as if each individual publication was specifically and individually indicated to be incorporated by reference.
[0117] The publications discussed throughout are provided solely for their disclosure before the filing date of the present application. Nothing herein is to be construed as an admission that the invention is not entitled to antedate such disclosure by virtue of prior invention.
[0118] The following examples are included to demonstrate the disclosure. It should be appreciated by those of skill in the art that the techniques disclosed in the following examples represent techniques discovered by the inventors to function well in the practice of the disclosure. Those of skill in the art should, however, in light of the present disclosure, appreciate that many changes could be made in the disclosure and still obtain a like or similar result without departing from the spirit and scope of the disclosure, therefore all matter set forth is to be interpreted as illustrative and not in a limiting sense. Example 1. Introduction
[0119] Centromeres are defining features of chromosomes in most eukaryotes and are essential for ensuring precise mitotic and meiotic genome segregation. In many species, including plants and animals, the centromeric region of each chromosome is specified epigenetically and is marked by a specialized chromatin environment. These chromatin specializations allow assembly and maintenance of kinetochores that mediate mitotic or meiotic spindle microtubule attachment and chromosome segregation. The DNA in centromeric regions is often characterized by large arrays of transposon or satellite DNA repeats, but such repeats are insufficient to specify a new centromere. Instead, specific centromeric proteins are present at centromere regions and can recruit new centromeric proteins to maintain centromere identity. A key protein of centromeres is the histone H3 variant CENP-A / CENH3 which incorporates into centromeric nucleosomes in place of 33PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web canonical histone H3. CENH3 not only marks the location of the centromere but serves as the base of the larger kinetochore complex. Overexpression or ectopic localization of CENH3 can be sufficient to drive the formation of a new centromere. Not surprisingly, CENH3 has proven to be an essential gene in all species where it has been carefully analyze d H3.
[0120] CENH3 homologs have been identified in many species and have some unusual properties compared with canonical histones. These properties include a highly variable N- terminal tail region and substitutions at a few positions within the highly conserved histone fold domain. Unlike canonical histone H3 which shows little variability across large phylogenetic distances, there is limited sequence conservation of CENH3-specific sequences between species outside the conserved histone fold domain, with evidence of positive selection on CENH3 within some clades. The rapid evolution of CENH3 proteins is mirrored by high rates of change between species in centromeric repeat sequences.
[0121] At ~113 Mb, the Chlamydomonas reinhardtii (Chlamydomonas) genome is one of the smallest among photosynthetic eukaryotes. The Chlamydomonas nuclear genome has 17 chromosomes with the smallest being only ~3.8 Mb, closer in size to Chromosome IV of yeast (~1.5 Mb) than the smallest Arabidopsis chromosome (~26 Mb). Like yeast, Chlamydomonas can be propagated like a microbe, subjected to rapid genetic manipulation, and grown at an industrial scale. Chlamydomonas is an important model species for research in photosynthesis, cell cycle control, metabolism, and biotechnology applications. The utility of Chlamydomonas in both basic and applied research combined with its relatively small genome size and compact individual chromosomes makes it a particularly suitable model for chromosome-scale synthetic biology applications.
[0122] Any effort to carry out whole chromosome design in Chlamydomonas will require a more complete understanding of the centromeres. Centromere locations in Chlamydomonas have been estimated genetically and subsequently sequenced, revealing the presence of nested Zepp-like transposons, a class of L1 LINE elements found in Chlamydomonas relatives and other green algae. Based on the lengths of ZeppL-LINE1 (ZeppL) clusters, the centromere sizes in a complete genome of the CC-5816 strain were estimated to range from 252 to 480 kb. However, centromeres in most species are defined not by DNA sequence, but by the 34PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web presence of CENH3, which may occupy only a subset of the repeat region at each centromere. In Arabidopsis, the major centromere repeat AthCEN178 occupies regions extending over 1.5– 6.5 Mb, but the CENH3-interacting regions are only ~1–2 Mb. These and similar data from other plant species suggest that the functional centromere regions of Chlamydomonas are likely to occupy only a fraction of the sequence within the ZeppL clusters.
[0123] In the following examples, the two CENH3 paralogs in Chlamydomonas, CENH3.1 and CENH3.2, are described and it was demonstrate that predicted null mutants for either gene alone are viable but that a cenh3.1 cenh3.2 double mutant is inviable. The inventors generated chromosome-scale assemblies of Chlamydomonas strains UL-1690 and CC-400 and compared homologous centromere regions among four strains with largely contiguous assemblies across their centromeres, and uncovered evidence of recent ZeppL element mobility and degeneration. The inventors next identified CENH3.1 and CENH3.2 bound genomic regions in strain UL-1690 using CUT&Tag which validated the predicted centromere localization of the two proteins and allowed us to annotate functional centromere domains. The data show that centromeres in Chlamydomonas are relatively small, occupying on average ~ 140 kb regions embedded within centromeric ZeppL clusters. Unexpectedly, a small non-centromeric locus on chromosome 2 of strain UL-1690 contained peaks of both CENH3.1 and CENH3.2 binding centered around a single strain-specific ZeppL insertion, hinting at a potential mechanism for new centromere formation. Example 2. Validation of Chlamydomonas CENH3 paralogs
[0124] Chlamydomonas is predicted to encode two CENH3 paralogs, CENH3.1 (Cre16.g661450) and CENH3.2 (Cre02.g104800), respectively, with the CENH3.2 gene model correctly predicted in the Phytozome V6.1 assembly and annotations, but not in V5.6 which mispredicted its structure. These paralogs were identified based on their divergence from canonical histone H3 sequences in specific N- and C-terminal regions, and other features that are common to CENH3 proteins. Previous transcriptome analyses in synchronous wild-type cultures showed that both genes were expressed periodically with peaks during S / M phase along with other replication-dependent histone genes (FIG.1). However, the transcript 35PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web abundance of both CENH3 paralogs was much lower than those for canonical histone genes, and the peak transcript abundance of CENH3.1 was around threefold higher than that of CENH3.2 (FIG.2).
[0125] Three Chlamydomonas-specific CENH3-specific antibodies were raised against peptides that contained unique paralog-specific sequences for CENH3.1 and CENH3.2, and a peptide covering a region that is conserved between CENH3.1 and CENH3.2, but not found in other histone H3 sub-types (FIG.3, Materials and Methods section). All three custom CENH3 antibodies recognized proteins of the expected size (~17 kD) on immunoblots, as well as some non-specific cross-reacting proteins (FIG.4; FIG.5). All three antibodies had similar punctate nuclear immunofluorescence (IF) staining patterns as expected for centromere localized antigens (FIG.6; FIG.7). While cross reactive antigens cannot be completely ruled out as the source of the IF staining patterns, the similar nuclear staining patterns for all three antibodies and additional evidence showing CENH3.1 and CENH3.2 at centromeric regions (see below CUT&Tag results) suggest the puncta represent centromeres. Example 3. CENH3.1 and CENH3.2 are essential but functionally redundant
[0126] The inventors used CRISPR-Cas9-mediated insertional mutagenesis to generate predicted null alleles for each of the two Chlamydomonas CENH3 paralogs in a wild-type strain background, UL-1690 (FIG.8) (Materials and Methods section). Multiple insertional alleles were generated and confirmed for each paralog indicating that neither of the two genes individually have essential functions (Materials and Methods section, Tables 1 and 8). Immunoblotting with each of the paralog-specific antibodies was used to validate loss of the predicted CENH3 protein in mutant strains (FIG.10). Two null strains containing marker insertions in their second exons were used for all further experiments and referred to as cenh3.1–2 and cenh3.2–2, respectively (FIG.8; FIG.9; Table 1). Immunoblots using the antibody raised against the common epitope of CENH3.1 or CENH3.2 had reduced signal in single mutant strains suggesting a lack of compensatory upregulation of the intact paralog when the other paralog was missing (FIG.10). Neither of the cenh3 mutants had obvious growth or viability phenotypes, and both showed normal Mendelian segregation in progeny 36PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web from backcrosses to a wild-type strain (FIG.11). However, when a cenh3.1–2 mutant was crossed with a cenh3.2–2 mutant, cenh3.1–2 cen3.2–2 double mutants were not recovered among the progeny, indicating that the two paralogs redundantly carry out an essential function (FIG.13). Constructs with native genomic CENH3.1 or CENH3.2 genes (FIG.14) were introduced into corresponding cenh3.1 or cenh3.2 mutant strains, respectively, and screened for restoration of expression (Materials and Methods section). Candidate rescued strains were first screened by genotyping, then followed by direct screening for restoration of CENH3 signal on immunoblots (Materials and Methods section, FIG.12,15). Confirmation of transgene function was obtained by crossing each rescued strain to a mutant strain with a cenh3 allele in the non-rescued paralog. Progeny from the cross were screened for each of the two mutant cenh3 alleles and for the presence of the rescuing transgene. Unlike the intercross of cenh3 mutants without a rescue construct, cenh3.1–2 cenh3.2–2 double mutant progeny could be recovered when a rescuing CENH3.1 or CENH3.2 transgene was also present (FIG. 16), indicating that the transgenes expressed functional CENH3. Table 1: CRISPR-Cas9-mediated mutagenesis target site sequences and mutant junction sequences gRNA in Crispr Cas9 mutagenesis target gene label guide RNA target PAM site target location Cre02.g104 g1 CAGGCTAAGCCCGTAAGTCAAGG (SEQ ID NO: 1st exon 800 7) g2 GGCAGCCCAGCGGTCTACGCGGG (SEQ ID 2nd exon NO: 8) g3 CCGCACCGCTATCGGCCGGGCAC (SEQ ID 3rd exon NO: 9) Cre16.g661 g5 AGTCGAAGCCAGTGAGTGTGAGG (SEQ ID 1st exon 450 NO: 10) g6 TCGGCCAGGGCGAAAGGCGCAGG (SEQ ID 2nd exon NO: 11) g7 CCGGCCTGGAACCGTGGCACTGC (SEQ ID 3rd exon NO: 12) 37PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web mutant insertion location sequencing result allele cenh3.1-2 chromosome_16:26 exon2 of CENH3.1 beta2-tubulin-promoter of 52496 AphVIII TCGGCCAGGGCGAAAGGTGTAAAACGACGGCCAGTGAATT GTAATACGACT… (SEQ ID NO: 13) cenh3.2-2 chromosome_02:42 exon2 of CENH3.2 pKS-AphVIII-Lox vector 83910 backbone GGCAGCCCAGCGGTCTAGGAAACAGCTATGACCATGATTA CGCCAAGCGCG…. (SEQ ID NO: 14) Example 4. High-quality genome and centromere assemblies of UL-1690 and CC-400
[0127] Laboratory strains UL-1690 (a derivative of strain CC- 1690 / 21gr) and CC-400 (cw15) were selected for HiFi PacBio sequencing to obtain contiguous assemblies of centromere regions (Materials and Methods section). CC-1690 is a standard wild-type laboratory strain, but sequence comparisons (see below) suggested that the isolate was likely the progeny of an intercross between CC-1690 and another strain. To avoid confusion, it was designated CC-1690 isolate as UL-1690 (Umen Laboratory 1690). CC-400 is a commonly used wall deficient mutant strain background that has been separated from a common ancestor with CC-1690 for around 80 years (Davies & Plaskitt, 1971; Proschold et al., 2005). The original CC-1690 genome was assembled previously, but with a high base-error rate due to limitations of Nanopore sequencing. The genomes of the strains UL-1690 and CC-400 were assembled using Pac- Bio HiFi long-read sequencing data with hifiasm as the assembler and RagTag for scaffolding (Materials and Methods section). The accuracy of the assemblies was assessed by comparing them to a recent gapless assembly of the strain CC-5816 (FIG.17). In the UL-1690 assembly, six chromosomes are gapless, with 25 gaps across the remaining 11 chromosomes. The CC-400 assembly also has six gapless chromosomes and 16 gaps across the other 11 chromosomes (details see Table 2). The CC-4532 strain serves as the standard C. reinhardtii reference genome with ongoing updates. In the current CC-4532 assembly (Phytozome V6.1), 63 gaps remain, and only six of 17 centromeres (chr4, 5, 7, 9, 14, 16) are fully assembled. In contrast, the new HiFi assemblies include 14 fully assembled centromeres in UL-1690 (chr1, 2, 4–7, 9–11, 13–17), and all 17 centromeres in CC- 400 (Table 2). 38PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web Table 2: Gap positions in four Chlamydomonas genome assemblies. Chromosome Gap start Gap end Gap In size centromere? UL-1690 (this study) chr_01 5997407 5997506 100 chr_01 6932438 6942947 10510 chr_03 6959917 6960016 100 Yes chr_03 9126875 9126974 100 chr_05 1195540 1199200 3661 chr_07 2548269 2548368 100 chr_07 4071850 4071949 100 chr_07 5802594 5802693 100 chr_08 398025 398124 100 chr_08 609988 610480 493 chr_08 2648587 2648686 100 Yes chr_08 2688208 2688307 100 Yes chr_09 2514854 2516442 1589 chr_11 3636006 3636105 100 chr_11 3713638 3718316 4679 chr_11 3753877 3764543 10667 chr_11 3790749 3814519 23771 chr_12 6459444 6474825 15382 chr_12 6844587 6844686 100 Yes chr_13 521729 521828 100 chr_13 927211 928189 979 chr_13 1754200 1754299 100 chr_13 2394235 2396260 2026 chr_14 3885535 3885634 100 chr_16 6288862 6288961 100 CC-400 (this study) chr_01 6907777 6907876 100 chr_02 3139815 3141661 1847 chr_02 3692594 3701625 9032 chr_02 6635793 6635906 114 chr_02 7544203 7544302 100 chr_03 3351881 3351980 100 chr_05 1197110 1198976 1867 chr_07 2721390 2727265 5876 chr_07 5801715 5801814 100 chr_09 3065517 3067438 1922 chr_11 3701594 3701693 100 chr_12 6426869 6437417 10549 chr_12 7514868 7514967 100 chr_13 926703 927950 1248 39PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web chr_14 3942771 3943417 647 chr_15 1195775 1195874 100 CC-4532 (phytozome V6.1) chr_01 974545 975148 604 chr_01 1132034 1151828 19795 Yes chr_01 1233533 1243603 10071 Yes chr_01 4763312 4764524 1213 chr_01 7245087 7245186 100 chr_02 2649502 2657162 7661 Yes chr_02 4594738 4594837 100 chr_02 8379022 8390796 11775 chr_03 855690 858204 2515 chr_03 5905380 5905479 100 chr_03 6193905 6201319 7415 chr_03 7075734 7083434 7701 Yes chr_03 7107376 7111733 4358 Yes chr_03 8396039 8396138 100 chr_03 9264142 9264241 100 chr_04 297308 297407 100 chr_04 812377 834151 21775 chr_04 1819408 1819507 100 chr_04 4118423 4118522 100 chr_05 849554 860230 10677 chr_05 3221803 3248934 27132 chr_06 3255156 3255255 100 chr_06 4440051 4440150 100 Yes chr_06 5759478 5772015 12538 chr_06 6037485 6040106 2622 chr_06 6224757 6236152 11396 chr_07 5770812 5781860 11049 chr_08 2652617 2666480 13864 Yes chr_09 1656887 1694678 37792 chr_09 2286595 2288631 2037 chr_10 1537607 1537706 100 chr_10 3675839 3677465 1627 Yes chr_10 3717551 3724400 6850 Yes chr_11 1217363 1218314 952 chr_11 2987397 2988500 1104 Yes chr_11 3025131 3036253 11123 chr_12 841559 841658 100 chr_12 4806608 4807530 923 chr_12 5529804 5530311 508 chr_12 5948403 5957787 9385 chr_12 6408478 6408577 100 chr_12 6830984 6836138 5155 Yes 40PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web chr_12 8002064 8051816 49753 chr_13 4362678 4368242 5565 Yes chr_14 739740 739839 100 chr_15 65850 126876 61027 chr_15 2839746 2839845 100 chr_15 3527514 3544252 16739 Yes chr_15 3866598 4066597 200000 chr_15 4339899 4366671 26773 chr_15 4860712 5095299 234588 chr_15 5129658 5129757 100 chr_15 5173233 5173332 100 chr_15 5818877 5818976 100 chr_15 5859679 5859778 100 chr_16 2553835 2553934 100 chr_16 2930428 2939298 8871 chr_16 5318971 5319070 100 chr_16 6699191 6699931 741 chr_16 6812299 6821993 9695 chr_17 19412 19511 100 chr_17 698662 705681 7020 chr_17 6090637 6126961 36325 Yes CC-1690 chr_04 803196 803245 50 chr_12 6390452 6390501 50 chr_13 4384339 4384388 50 Yes chr_15 3797460 3797559 100
[0128] The CC-4532 and CC-1690 genome assemblies differ from UL-1690, CC-400 and CC-5816 by the presence of several large tandem repeat blocks, most noticeable on chr11 and chr15 in the latter three assemblies but not the former two (FIG.18). The chr11 block is a single ~670 kb (UL-1690) or ~860 kb (CC-5816, CC-400) insertion comprising over three thousand copies of a 184 bp repeat, while the chr15 repeat insertions are in multiple locations and involve several repeat sequences with repeat units ranging from 107 to 438 bp and occupying a total of around 1.2 Mb on chr15 in all three strains (FIG.18; Table 3). Like centromere regions, these repeats are nearly devoid of predicted protein-coding genes, highly enriched for 5-methylcytosine and may form heterochromatin-like domains, but this remains to be determined. 41PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web Table 3: Tandem repeats regions in chr11 and chr15 in genome assemblies CC-5816, CC- 400, and UL-1690. CC-5816 chr# start end period copies repeat sequence SEQ IDPROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GCAATTTGGCAGGAAGGGC GGAAGGCAAGGGGCGGGC AT T TA A T A T43PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web TATGGGGGAGTAGGGAAGA GGCCCCACGCAAACACACA A T T T TA APROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CTGTGCAAGGCATGGGGAG CCTGCGGGGCTTGTGCGG AATTT A AA45PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GTGCTACCAGCGGGTTCCG TGGCAAGTGCGGTGAATGT TT T A T A APROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web ATGTCTACCAGCCTGCACT GCACTCGTGTGCGTCTGTG AA AT47PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web Example 5. Functional centromere size estimates based on CUT&Tag analysis of CENH3
[0129] Centromeres in most eukaryotes are defined by the presence of CENH3 which is maintained at centromere repeats through an epigenetic recruitment mechanism. Centromere location and size are best estimated by determining the sub-regions of centromere repeats bound to CENH3. To identify these regions in Chlamydomonas Cleavage Under Targets and Tagmentation (CUT&Tag) was carried out with the sequenced strain UL-1690 using each of the three CENH3 antibodies described above, along with nonspecific rabbit IgG as a control for background signal. The CENH3 CUT&Tag data from each of the three antibodies showed nearly identical peak locations, with localization confined almost exclusively to ZeppL elements at centromere regions for all 17 Chlamydomonas chromosomes (FIG.19, Materials and Methods section).
[0130] Prior data from multiple species showed that the distribution of CENH3 reads is usually denser in the center of centromeres and decreases toward chromosome arms, and that the pattern of CENH3 enrichment shifts between individuals. Thus, the data represent a population average. Additionally, the highly repetitive nature of centromeric DNA often makes it impossible to know the precise origin of a read. Alignment parameters that allow reads to map to one of multiple possible locations (as well as unique locations, MAPQ value of 20 or larger) tend to overestimate the centromere DNA occupied by CENH3, likely beyond the minimal regions required for centromere function. In contrast, using only unique mapping reads provides more accurate estimates of the locations of functional centromeres but may underestimate their total size. Here, both multiple plus unique mapping and unique mapping were utilized to estimate the upper and lower centromere size limits for each chromosome (FIG.20; Table 5). The results showed that the upper estimates of centromere sizes are 65– 310 kb, which is close to the size of ZeppL-enriched regions (Table 5; FIG.21; Table 4), and the lower estimates of centromere sizes are 50–105 kb with ~85% of unique mapping regions containing ZeppL sequences (Table 5; FIG. 21; Table 4). Chromosome sizes and ZeppL- enriched region sizes exhibited a weak positive correlation (R2 = 0.14, FIG.21), but when uniquely mapping centromere regions from the lower centromere size estimates were 48PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web compared to chromosome sizes, the correlation largely disappeared, suggesting that functional centromere size may be independent of chromosome size in Chlamydomonas. Table 4: Coordinates of centromere locations and ZeppL-LINE1 regions in UL-1690 HiFi genome. C Upper centromere size Lower centromere size ZeppL-LINE1 hr estimate estimate enriched region L or 9 9 00 4 3 9 6 5 7 4 5 3 6 3 3 6 7PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 1 5860 6170 3100 52.40 5860 5955 9500 76.92 5864 6169 3057 51.13 7 000 000 00 % 000 000 0 % 006 743 37 %
[0131] Chlamydomonas has monocentric chromosomes, where the functional centromeres are limited to a single domain containing CENH3. However, on chromosome 2 a second strong peak of CENH3 CUT&Tag read enrichment was observed far from its mapped centromere and close to the end of its long arm (FIG.19 and 20), suggesting a possible secondary centromere. The total estimated size of the secondary CENH3 binding region is 30– 35 kb, which is about half the size of the genetically defined centromere on chromosome 2 (Table 5). Remarkably, a single intact ZeppL element (described above) lies at the center of the CENH3 enriched region, in between two protein-coding genes Cre02.g800275_4532.1 and Cre02.g141600_4532.1 (FIG.22). Importantly, the CENH3 CUT&Tag signal is strongly enriched not only in the ZeppL sequences, but in the unique gene-rich sequences flanking the ZeppL insertion, indicating the unambiguous presence of CENH3 at this locus (FIG.20; FIG. 21). The ZeppL insertion near the end of chromosome 2 was unique for strain UL-1690 and not found in assemblies of CC-4532, CC-5816, CC-400, or the previously assembled Nanopore CC-1690 genome (FIG. 22; Table 6). These results suggest that the ZeppL element near the end of chromosome 2 in UL-1690 could be the result of a recent transposition event.
[0132] Table 5 Chromosome length, ZeppL-LINE1 region length, and centromere size estimates for UL-1690 HiFi genome assembly 50PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web Example 7. Evidence of active ZeppL transposition and turnover in Chlamydomonas
[0133] The sizes of the predicted centromere regions for a given chromosome (measured as distance between the ZeppL repeats flanking the centromere region) are similar between strains, suggesting that the overall activity of ZeppL is not high enough to cause major changes to centromere size over the decades separating CC-400, CC-1690, and CC- 5816 (Table 6; Data S1 of Liu et al., 2025, The Plant Journal (2025) 122, e70153, the 51PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web disclosure of all of which is incorporated herein in its entirety). However, the finding of the likely ZeppL transposition activity associated with insertion at the neocentromere-like region on chromosome 2 prompted us to search for the origin of the insertion element and for additional signs of ZeppL mobility. The inventors first identified 49 full-length or near full-length ZeppL sequences (>8 kb) as potential active candidates among four genome assemblies (UL-1690, CC-400, CC-1690, CC-5816) and eleven different chromosomes (1, 2, 3, 6, 9, 10, 13, 14, 15, 16, 17) including the insertion on Chr2 (Data S2 of Liu et al., 2025, The Plant Journal (2025) 122, e70153, the disclosure of all of which is incorporated herein in its entirety). The inventors also found four partial ZeppL sequences in CC-400 and UL-1690 centromeres which were >5 kb and >99.9% identical to the corresponding full-length ZeppL element in strain CC-5816. These four elements were likely beginning a process of fragmentation and decay (Table 7; Data S3 of Liu et al., 2025, The Plant Journal (2025) 122, e70153, the disclosure of all of which is incorporated herein in its entirety). Table 6. Coordinates of ZeppL-LINE1 enriched region in UL-1690, CC-400, CC-5816, and CC-4532 assemblies. C UL-1690 ZeppL- CC-400 ZeppL-LINE1 CC-5816 ZeppL- CC-4532 ZeppL- hr LINE1 enriched enriched region LINE1 enriched LINE1 enriched region region region Start End Size Start End Size Start End Size Start End Size 1 10813 12914 2100 10815 12741 1926 11028 13049 2021 11058 13028 1970 92 00 08 01 56 55 00 62 62 05 14 09 2 25439 26164 7250 25783 26513 7299 25368 26113 7446 25921 26760 8389 86 87 1 25 18 3 72 33 1 99 89 0 3 68345 70679 2334 68699 70933 2234 68563 70796 2232 69476 71861 2385 03 93 90 68 78 10 69 48 79 60 66 06 4 93025 11585 2283 93576 11631 2274 94398 11773 2333 83419 10588 2246 3 85 32 9 86 17 8 87 99 2 71 79 5 16123 17537 1414 15991 17359 1368 15877 17233 1356 15965 17353 1388 34 44 10 95 99 04 04 56 52 26 43 17 6 44543 46996 2452 44756 47226 2470 44283 46723 2439 43564 46210 2645 50 18 68 68 73 05 93 19 26 88 47 59 7 30465 31972 1506 30199 31726 1527 30196 31721 1525 30318 31976 1657 97 08 11 20 49 29 29 62 33 73 01 28 8 26841 28204 1363 26623 27743 1120 26056 27174 1117 26341 27465 1124 42 83 41 30 36 06 32 14 82 22 92 70 9 40536 42310 1774 42745 44473 1728 41947 43686 1738 41945 43705 1760 55 66 11 01 65 64 82 49 67 01 36 35 52PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 10 35751 37721 1970 35696 37788 2092 35770 37675 1905 35826 37863 2037 64 68 04 09 80 71 21 64 43 00 70 70 11 28284 29928 1644 28999 30630 1631 28367 30012 1644 28683 30251 1567 19 50 31 08 17 09 70 16 46 39 30 91 12 67513 69614 2100 67288 69361 2073 67450 69523 2072 67136 69209 2073 45 12 67 73 76 03 32 31 99 64 65 01 13 42640 43657 1016 43025 44044 1018 42893 43910 1016 43190 43626 4357 67 22 55 29 19 90 78 34 56 99 77 8 14 21175 22239 1064 21693 22876 1183 21175 22239 1064 21826 22906 1079 12 25 13 43 68 25 25 38 13 24 15 91 15 31965 33919 1954 33431 35011 1580 32421 34375 1954 34148 36231 2082 46 99 53 84 91 07 13 76 63 58 53 95 16 32453 34966 2513 32743 35258 2515 32684 35197 2513 33283 35813 2529 74 98 24 48 81 33 00 16 16 75 32 57 17 58640 61697 3057 59851 62916 3064 59685 62631 2945 60330 63451 3120 06 43 37 26 00 74 50 24 74 81 38 57
[0134] Table 7: Coordinates and percentage identity of partially decayed ZeppL-LINE1 elements to the corresponding full length ZeppL elements in CC-5816.that were categorized as relatively stable ZeppL elements and recently active ZeppL elements (FIG. 23; FIG.24–26). The relatively stable ZeppL elements formed distinct and well- supported chromosome-specific clades in the centromeres of chr1 (two elements), 6, 9, 10, 13, 14, and 15, 16. These elements were present in all four strains though there was evidence of degeneration / fragmentation of the originally full-length ZeppL element as described above (FIG.23; Data S3 of Liu et al., 2025, The Plant Journal (2025) 122, e70153, the disclosure of all of which is incorporated herein in its entirety; Table 7). The recently active sub-group of ZeppL elements were present on chromosomes 1, 2, 3 (two elements), 10, 15, 16, 17 and included the neocentromere candidate ZeppL sequence on chromosome 2 of UL-1690. This group was characterized by their high relatedness (near-zero branch lengths) suggesting more 53PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web recent origins via active transposition between chromosomes. The branching patterns of the recently active ZeppL group to which the new insertion on chromosome 2 belongs suggest some of these elements were present in the recent common ancestor of all four laboratory strains for which data was obtained. However, there also appears to have been active movement evidenced by
[0136] presence / absence between strains for elements on chromosomes 1, 10, and 17 (FIG.23). In summary, a group of highly similar ZeppL elements shared across multiple chromosomes, including the neocentromere of chr2, suggests active ZeppL transposition and some element decay before and since the four laboratory strains separated. Example 8. Centromere variability between and within Chlamydomonas strains
[0137] The identification of ZeppL element mobility led us to compare centromere lengths and sequences more broadly among different strains and assemblies. Overall individual centromere repeat region lengths and sequences were similar between different laboratory strains (FIG. 27– 29), similar to what has been observed in other species. However, there were also some notable differences caused by structural rearrangements (FIG.30). While it is difficult to meaningfully compare different centromeres with conventional sequence alignments due to their heterogeneity in size and sequence, alignment-free methods based on k-mer counting an be used to compare centromere relatedness across all centromeres and strains.
[0138] The inventors used Mash to estimate Jaccard distances between each of 64 gap- free centromere sequences available from four genome assemblies (UL- 1690, CC-1690, CC- 400, and CC-5816) (Materials and Methods section) and visualized their divergence estimates with a distance tree (FIG.31). As expected, homologous centromeres formed distinct clades and showed varying degrees of relatedness with each other. The chr4 and chr10 centromeres were most closely related to each other while the chr2 native centromere was on average most distal to all other centromeres. For each specific centromere, the amount of inter-strain divergence also varied by chromosome. Most individual centromeres for a given chromosome were highly similar in all four strains indicated by short branch lengths connecting them to their 54PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web nearest common node; but in a few cases there was marked divergence in one or more of the strains, for example, CC-1690 chr1; CC-400 chr10; CC-400, CC-1690 chr14; CC-400 chr15. For chr10 and chr15, one centromere border differed between CC-400 and the other three strains, including a ~200 kb inversion in CC-400 that substantially changed the border of its chr15 centromere (FIG.30). For chr1 and chr14, there were internal strain-specific indels in CC-1690 and / or CC-400 that were likely responsible for higher divergence scores of their centromeres from the other strains (Figures 24-26). In summary, comparative analyses indicate that both large-scale genomic changes—such as structural rearrangements and ongoing transposon activity—and smaller-scale polymorphisms have driven the rapid and continuous evolution of centromeres, even within the relatively short time span of Chlamydomonas domestication. Example 9. DISCUSSION
[0139] CENH3 proteins and their association with centromeres have been studied in animals, fungi, and land plants, but little work has been done on centromere proteins of green algae like Chlamydomonas. Previous work defined centromeric regions in Chlamydomonas by genetic mapping and by the presence of ZeppL transposon repeats in assembled genomes. Here, two functionally overlapping CENH3 paralogs in C. reinhardtii were characterized and their expression and genomic localization were validated at predicted centromere regions using CUT&Tag with custom CENH3 antibodies. The inventors also introduce two newly assembled Chlamydomonas genomes and provide data suggesting that the transposition of ZeppL elements outside of the normal centromere domains can potentially result in neocentromere formation.
[0140] Most species encode only a single CENH3 / CENP-A gene but in C. reinhardtii the inventors found two. The two CENH3s exhibit significant divergence in their N-terminal regions (FIG.19), suggesting that they did not arise from a very recent duplication event. Interestingly, a search for CENH3 orthologs in Volvocine algal relatives of C. reinhardtii identified only single genes even for its closest sequenced relative, C. incerta. Although this suggests a potentially recent duplication within the C. reinhardtii genome, the rapid evolution of CENH3 proteins 55PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web makes this explanation difficult to distinguish from selective losses. However, the lack of CENH3 duplicates in the nearest relatives of C. reinhardtii and most other species suggests that duplications of this gene are generally not long-lived.
[0141] In some species with two CENH3 genes, one of the two paralogs appears to have a larger role than the other. An extreme example is cowpea, where the localization patterns of the two different CENH3 paralogs generally overlap but only one of the two paralogs is essential for plant growth and reproduction. Unlike the case in cowpea, the two CENH3 paralogs in Chlamydomonas could both support apparently normal mitotic growth and meiotic segregation when the other paralog was absent making them seemingly redundant. However, subtle phenotypes in the single knockout strains that would require more comprehensive analyses to detect could not be ruled out.
[0142] One of the goals of this study was to prepare genome assemblies of sufficient quality to span the centromeres in two different C. reinhardtii strains, UL-1690 and CC-400. Most, but not all, the centromeres in these two strains were completely assembled. An entirely gapless assembly of strain CC-5816 was released. Each of these genomes is superior to the current assembly of the reference strain CC-4532 and includes large repeat blocks absent from prior assemblies. Centromeres in all strains were marked by clusters of ZeppL elements that are 65–310 kb in length (Table 6). CENH3 CUT&Tag analysis confirmed that the functional centromeres lie within these repetitive regions (FIG.19). By measuring the distribution of uniquely mapped reads, it was estimated that the functional domains are on the order of 50–105 kb (Table 5), considerably smaller than the estimated size of the centromeres in Arabidopsis, where the centromeric repeat arrays range from 10 to 22 Mb and the CENH3- enriched areas are roughly to ~1–2 Mb.
[0143] The inventors observed evidence of active ZeppL transposition in laboratory strains since domestication (Figures 24–26; Data S2 of Liu et al., 2025, The Plant Journal (2025) 122, e70153, the disclosure of all of which is incorporated herein in its entirety) as well as a very recent ZeppL transposition in strain UL-1690. The newly transposed ZeppL element on the end of chr2 is associated with what appears to be a new, secondary centromere (FIG. 20; FIG. 22). Though CENH3 enrichment within the ZeppL element itself could not be 56PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web assessed (due to the presence of three other ZeppL elements with nearly identical sequence [Data S2 of Liu et al., 2025, The Plant Journal (2025) 122, e70153, the disclosure of all of which is incorporated herein in its entirety]), a total of ~30 kb of unique sequence on either side of the ZeppL element is associated with CENH3 (FIG.22), creating what appears to be a dicentric chromosome. While there are many known examples of newly formed centromeres (neocentromeres), dicentric chromosomes should be unstable because the two centromeres will frequently migrate to opposite spindle poles and initiate cycles of breakage that results in either the inactivation of one centromere or the formation of two separate chromosomes. The latter hypothesis is supported by a recent observation in maize, where activation of synthetic centromeres on maize chromosome 4 induced chromosome breakage, forming neo- chromosomes without disrupting gene expression, or normal plant development, and enabling stable inheritance. Further work will be needed to determine if the neocentromere on chr2 is sufficiently large to independently interact with the spindle, or if chr2 is more unstable in UL- 1690 than in other strains.
[0144] While it is theoretically possible that a neocentromere formed spontaneously on chromosome 2 after which ZeppL transposition occurred, the reverse seems more likely—that ZeppL transposed first and CENH3 was recruited over and around it. Most complex eukaryotes display a loose, epigenetic relationship between centromere sequence and CENH3 location. In contrast, ZeppL may have evolved sequence features that actively recruit CENH3. The best-known example of a genetic determination process affecting CENH3 / CENP-A localization in multicellular eukaryotes comes from humans, where the major centromere repeat (the alpha satellite) contains a binding site for a protein known as CENP-B that facilitates CENP-A deposition. Similarly, in holocentric plants from the Rhynchospora genus, CENH3 is strictly correlated with the major centromere tandem repeat Tyba. Gain or loss of Tyba (over evolutionary time scales) results in a corresponding gain or loss of CENH3 binding. The data suggest that a single ~8 kb ZeppL element may be sufficient to recruit CENH3. Due to the small size of its centromeres and a potentially simple route to new centromere formation driven by ZeppL transposon nucleation, Chlamydomonas may be an excellent model for engineering fully artificial plant chromosomes. 57PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web Example 10. MATERIALS AND METHODS Chlamydomonas strains and growth conditions
[0145] Chlamydomonas strains CC-400 (cw15), CC-124, CC-125, CC-620, CC-621 were obtained from the Chlamydomonas Resource Center (https: / / www.chlamycollection.org / ). UL-1690 is a mating type plus (mt+) strain, most likely derived from a cross between strains CC- 1690 (21gr, mt+) and CC-1691 (6145c, mt-) that was mistakenly labeled as CC-1690 and has since been renamed UL-1690 (Umen Laboratory 1690). All strains were maintained on Tris-acetatephosphate (TAP) + 1.5% BD Difco agar (Thermo Fisher, Cat# DF0812-07-1) plates. For asynchronous growth, cells were cultured at 25°C in TAP liquid with illumination from red (625 nm, 100 lE) and blue (465 nm, 100 lE) LEDs and aeration from bubbling with air. For synchronous growth, cells were cultured at 25°C in Sueoka’s high salt medium (HSM) under a 12 h dark:12 h light diurnal cycle with illumination from red (625 nm, 150 lE) and blue (465 nm, 150 lE) LEDs and aeration from bubbling with 1% CO2. Gamete generation, mating, and zygote germination were performed following standard protocols. Segregation analysis was done with random meiotic progeny sub-cloned after bulk germination. Antibody generation
[0146] Rabbit polyclonal antibodies were prepared and affinity purified by Pacific Immunology, Ramona, CA. The peptides were designed to uniquely identify CENH3 proteins but not canonical histone H3.To generate antibodies unique to CrCENH3.1, a peptide matching amino acids 15–39 (KAQAEAATPTKSKRPSGAAATPTRG-Cys SEQ ID NO: 69) was used; to generate antibodies unique to CrCENH3.2, a peptide matching amino acids 15–40 (EAAAQRSTRAGEAVGSPARGGTPARS- Cys SEQ ID NO: 70) was used, and to generate antibodies that would identify both CrCENH3.1 and CrCENH3.2, a peptide matching amino acids 8–16 (PARPGRKA-Cys SEQ ID NO: 71) was used, though note that the lysine at position 15 is only found on CrCENH3.1). Nuclear protein enrichment and extraction 58PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web
[0147] The nuclear protein enrichment protocol was adapted and simplified from Rommelfanger et al. (2021). Chlamydomonas cultures were grown synchronously (described above) and maintained at a cell concentration between 59105 and 19106 cells mL-1.59 108 daughter cells (light phase ZT = 1) per cell culture were harvested by centrifugation at 3000g for 5 min with 0.005% Tween 20. Pellets were flash frozen in liquid nitrogen, then thawed on ice in 1 mL nuclear isolation buffer (NIB, 25 mM HEPES pH7.0, 20 mM KCl, 20 mM MgCl2, 0.6 M sucrose, 10% glycerol, 1 mM PMSF, 5 mM DTT, 10 mM sodium butyrate, 1X Roche cOmplete Protease Inhibitor Cocktail, Roche 04693116001) + 5% Triton X-100. The thawed pellets were flash frozen in liquid nitrogen again, then subjected to tissue grinding twice (frequency 30 Hz, duration 90 sec, baskets chilled with liquid nitrogen) with three 3 mm glass beads per sample in 1.5 mL Eppendorf tubes using a Qiagen TissueLyser II. The macerated mixture was then thawed and pelleted by centrifugation at 1000g for 60 min at 4°C. Pellets were washed twice using NIB and recollected by centrifugation at 1000g for 30 min at 4°C. Nuclear-enriched pellets were resuspended in lysis buffer (19 PBS pH 7.4, 19 Roche protease Inhibitor, 500 mM PMSF) to a final concentration equivalent of ~59109 cells mL-1. The suspensions were solubilized using a Covaris ultrasonicator (peak power 150, duty factor 150, cycle 200, treatment 120 sec, 4°C) to generate protein lysates. Protein lysates were mixed 5:1 with 6X SDS protein loading buffer and boiled for 10 min, then used either directly in immunoblotting or saved for later use at -20°C. Immunoblotting
[0148] Protein lysates in 19 loading buffer were further cleared by centrifugation at 12 000g for 10 min. Supernatants were loaded onto 15% tricine sodium dodecyl sulfate- polyacrylamide (SDS-PAGE) gels for separation, then wet-transferred to polyvinylidene fluoride membranes for 1 h at 50 V. Membranes were blocked in PBS containing 5% non-fat dry milk, incubated with primary antibody (anti-CENH3.1, 1:1000; anti-CENH3.2, 1:500; anti- CENH3-conserved, 1:500) added to PBS containing 3% non-fat dry milk overnight at 4°C. Membranes were then washed in PBS containing 0.1% Tween 20 for 3910 min. Secondary antibodies coupled to horseradish peroxidase (goat anti-rabbit 1:10000, Thermo Scientific) 59PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web were added to PBS containing 3% non-fat dry milk. Membranes were incubated with a secondary antibody for 1 h at room temperature, then washed in PBS containing 0.1% Tween 20 for 3910 min. Chemiluminescent detection was done using a Bio-Rad Chemidoc XRS+ Imaging System. Immunofluorescence microscopy
[0149] Wild-type UL-1690 cells (~107 cells) were grown asynchronously in TAP (described above) and maintained at a cell concentration between 59105 and 19106 cells mL-1.59 108 cells per cell culture were harvested by centrifugation at 3000g for 5 min with 0.005% Tween 20 and resuspended in 240 lL fixation buffer (2% paraformaldehyde in PBSP [19 PBS pH7.4, 1 mM DTT, 19 Roche protease inhibitor]) and left for 30 min on ice. Fixed cells were extracted in cold methanol 3910 min at -20°C and rehydrated in 1 mL PBSP for 30 min on ice. Fixed cells were blocked for 30 min in 1 mL blocking solution I (5% BSA and 1% cold water fish gelatin in PBSP) and 30 min in 1 mL blocking solution II (10% goat serum, 90% blocking solution I) at room temperature. Cells were then incubated overnight with one of three primary antibodies (anti- CENH3.1, 1:200; anti-CENH3.2, 1:50; anti-CENH3-conserved, 1:50) in 20% blocking solution I total volume 50 lL at 4°C, then washed 3910 min in 1% blocking solution I at room temperature. Cells were incubated with 1:1000 Alexa Fluor 488 conjugated goat anti-mouse IgG (Thermo Fisher) in 20% blocking solution I total volume 100 lL for 1 h at 4°C and then incubated with 40,6-diamidino- 2-phenylindole, dihydrochloride (DAPI) at a final concentration of 5 lg mL-1 at room temperature for 5 min. Cells were washed in 19 PBS for 3 910 min, then mounted for microscopy in 9:1 Mowiol: 0.1% 1,4-phenylenediamine (PPD). Cells were imaged with a Zeiss Elyra7 Super-Resolution microscope using a 63X oil lens with 405 nm (50 mW, 4.5%), 488 nm (500 mW, 4.0%), and 561 nm (500 mW, 4.0%) diode laser lines, with filter settings 561 TV1:BP 570–620 + LP 655, 488 TV2: BP 420–480 + BP 495–550, and 405 TV2: BP 420–480 + BP 495–550.3D LEAP MODE was used to capture volumes of ~3 lm thickness in ~100 nm increments. The image stacks were further processed with Zeiss Zen Black image analysis software using the dual iterative SIM (SIM2) algorithm, 3D leap 60PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web processing mode (default parameters, sharpness 3.0) for all channels, and exported as a single maximum projection images. Chlamydomonas transformation
[0150] Chlamydomonas cultures were grown to 106 cells mL-1 under the asynchronous condition (described above) and harvested by centrifugation at 3000g for 5 min with 0.005% Tween 20. Cell pellets were resuspended and washed in TAP + 40 mM sucrose to a final concentration of 29108 cells mL-1.125 lL cell suspension (~2.59107 cells) was mixed with ~ 500 ng plasmid and transferred into a 2-mm cuvette. Electroporation was performed using a NepaGene Electroporator with poring pulse (250 V initial, 2 pulses, 8 ms each, 40% decay, 50 ms interval) and transfer pulse (20 V initial, 5 pulses with alternating polarity, 50 ms each, 40% decay, 50 ms interval). Cells were then transferred into 5 mL TAP + 40 mM sucrose and recovered under dim light at room temperature overnight before plating on antibiotic selection plates as specified below. CRISPR-Cas9 mediated mutagenesis of CENH3.1 and CENH3.2
[0151] gRNAs were designed using the CRISPR RGEN tool (http: / / www.rgenome.net / cas-designer / ). The chosen gRNA target sites are listed in Table 1. In vitro transcription of sgRNA was performed following (Yu et al., 2017).5 lg (~30 pmol) Cas9 Nuclease (Alt-RTM S.p. Cas9 Nuclease V3, IDT Cat# 1081059) was incubated with 5000 ng (~120 pmol) purified sgRNA in 3 lL IDT RNA duplex buffer at 37°C for 30 min to assemble the RNP complex for one transformation.
[0152] CRISPR-Cas9 targeted insertional mutagenesis was performed following (Greiner et al., 2017; Picariello et al., 2020) with minor modifications: Cultures of strain UL- 1690 were grown to 29106 cells mL-1 asynchronously in TAP (described above) and harvested by centrifugation at 2000g for 5 min. Pellets were resuspended in 6 mL autolysin and incubated at 33°C for 30 min, then heat shocked at 40°C for another 30 min. After heat shock, cells were collected by centrifugation at 2000g for 2 min, then resuspended and washed in TAP + 40 mM sucrose to a final cell concentration 29108 cells mL-1.125 lL cell 61PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web suspension (~2.59107 cells) was mixed with RNP complex and 500 ng Chlamydomonas paromomycin marker (linear PCR product using M13 primer sets from pKS-aphVIII-lox). The RNP and donor DNA were delivered into the cells by electroporation as described above.
[0153] Individual transformants were selected on TAP +1.5% BD agar with 20 lg mL-1 paromomycin on a light shelf (50 lE white light) at room temperature for 6 days. Transformants were then picked and re-grown individually in 96-well plates with 200 lL TAP medium for 3 days. Crude DNA preparation for genotyping followed. Transformants were screened by genotyping for insertion of the paromomycin marker at the target site. Genotyping oligos and amplification conditions are listed in Table 8. Genotyping PCR reactions were performed in a 10 lL volume containing 200 lM of each dNTP, 0.4 lM of each primer, 3% DMSO, 1 lL of homemade recombinant Taq DNA polymerase, and 1 lL of crude DNA preparation. Homemade Taq buffer final concentrations: 10 mM Tris–HCl pH8.4, 50 mM KCl, 1.5 mM MgCl2, 0.08% NP40, and 0.4 lg lL-1 BSA. Loss of CENH3 expression was further confirmed by immunoblotting (see above). Transgenic rescue of cenh3 mutants
[0154] A ~4.7 kb fragment containing the full-length genomic region of CENH3.1 and its upstream region (~1.8 kb upstream of the transcription start site) was amplified with 29 KOD Hot Start Master Mix (Sigma-Aldrich Cat# 71842) from genomic UL-1690 DNA using primers CENH3.1 comp F1 / CENH3.1 comp R1 (Table 8). The vector backbone was amplified with 29 KOD Hot Start Master Mix from pHsp70A / RbcS2-cgLuc using primers CENH3.1 IO OH R1 / CENH3.1 IO OH F1 (Table 8). The amplified vector fragment was ligated with the amplified CENH3.1 genomic region using a Takara In-Fusion kit (In-Fusion Snap Assembly Master Mix, Cat # 638947) to generate CENH3.1 rescue construct pCENH3.1.
[0155] A ~4.6 kb fragment containing the full-length genomic region of CENH3.2 with ~1.5 kb upstream of 50 UTR was amplified with 29 KOD Hot Start Master Mix (Sigma-Aldrich Cat# 71842) from genomic CC-1690 DNA using primers CENH3.2 comp F2 / CENH3.2 comp R2 (Table 8). Backbone vector pHsp70A / RbcS2-cgLuc was amplified with 29 KOD Hot Start Master Mix using primers CENH3.2 g IO OH R2 / CENH3.2g IO OH F2 (Table 8). The amplified 62PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web vector fragment was ligated to the genomic fragment of CENH3.2 using a Takara In-Fusion kit to generate CENH3.2 rescue construct pCENH3.2 (constructs maps in FIG.14).
[0156] Either pCENH3.1 or pCENH3.2 (500 ng / transformation) was co-transformed with pKS-aphVII-lox (50 ng / transformation) into Chlamydomonas cenh3.1 or cenh3.2 mutants (Methods see above) by electroporation. Transformants were selected on TAP +1.5% BD agar with 30 lg mL-1 hygromycin. Individual transformants were picked into 96-well plates and screened for the presence of CENH3 transgenes by PCR genotyping (Table 8). Crude DNA preparation for genotyping followed. All genotyping oligos and amplification conditions are listed in Table 8. Genotyping PCR reactions were performed in a 10 lL volume containing 200 lM of each dNTP, 0.4 lM of each primer, 3% DMSO, 1 lL of homemade recombinant Taq DNA polymerase, and 1 lL of crude DNA preparation. Homemade Taq buffer final concentrations: 10 mM Tris–HCl pH8.4, 50 mM KCl, 1.5 mM MgCl2, 0.08% NP40, and 0.4 lg lL-1 BSA. Table 8: Oligonucleotides used in this study. assembly of CENH3.1 and CENH3.2 constructs for mutant rescue template name sequence PCR amplic condition on s w / KOD size Hot Start (bp) Master Mix pHsp70A / Rb CENH3. CCGATCCTGGTTTCTGCAGTCGCTATCAGCCTCG 98C 30s, 2222 cS2-cgLuc 2g IO ACTGTG (SEQ ID NO: 15) 35 cycles OH R2 of (98C 15s, 70C 10s, 68C 60s), 72C 5min CENH3. TTGTCGTCACCACCAAACCACATGGGCGGGCTGGGCGTAT (SEQ ID 2g IO NO: 16) OH F2 pHsp70A / Rb CENH3. TAATGGCCTGGCTCAACAGGCGCTATCAGCCTCGACTGTG ( 2222 cS2-cgLuc 1g IO SEQ ID NO: 17) OH R1 63PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CENH3. AGGTAGGAGTGGAGAGCAGGCATGGGCGGGCTGGGCGTAT (SEQ ID 1g IO NO: 18) OH F1 genomic CENH3. ACTGCAGAAACCAGGATCGG (SEQ ID NO: 19) 98C 30s, 4689 CC1690 2 comp 35 cycles DNA F2 of (98C 15s, 68C 10s, 68C 60s), 72C 5min CENH3. TGGTTTGGTGGTGACGACAA (SEQ ID NO: 20) 2 comp R2 genomic CENH3. CCTGTTGAGCCAGGCCATTA (SEQ ID NO: 21) 4897 CC1690 1 comp DNA F1 CENH3. CCTGCTCTCCACTCCTACCT (SEQ ID NO: 22) 1 comp R1 amplification of wild-type CENH3.1 and CENH3.2 for genotyping target name sequence PCR amplic condition on s w / size recombin (bp) ant Taq and homema de Taq buffer CENH3.2 CENH3. CCCGGCTGTCACCTATTGTT (SEQ ID NO: 23) 98C 30s, 569 2 F1 39 cycles of (98C 10s, 65C 30s, 72C 30s), 72C 5min CENH3. CAGCTCCGTGGATTTCTGGT (SEQ ID NO: 24) 2 R1 CENH3.1 CENH3. ACAGTCGAAGCCAGTGAGTG (SEQ ID NO: 25) 577 1 F1 CENH3. CTACCCTCACCAATCGCGAG (SEQ ID NO: 26) 1 R1 64PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CENH3.25thCENH3. GCCAAGCGCATCACCATAAG (SEQ ID NO: 27) 547 exon-3‘UTR 2 F2 CENH3. GCTGCAAAGTACACGCATCA (SEQ ID NO: 28) 2 R2 nin f r nh3 m t nt in Cri rC 9 m di t d in rti n lmuageness target name sequence PCR conditions w / recombinant Taq and homemade Taq buffer cenh3.2 CENH3. CCCGGCTGTCACCTATTGTT (SEQ ID NO: 23) 98C 30s, 39 2 F1 cycles of (98C 10s, 60C 30s, 72C 30s), 72C 5min. Aph8 R GCATCATCGGTGTTGCATGCAG (SEQ ID NO: 29) Aph8 F GTAACGTCAGTTGATGGTACCAG (SEQ ID NO: 30) cenh3.2 CENH3. CAGCTCCGTGGATTTCTGGT (SEQ ID NO: 24) 2 R1 Aph8 R GCATCATCGGTGTTGCATGCAG (SEQ ID NO: 29) Aph8 F GTAACGTCAGTTGATGGTACCAG (SEQ ID NO: 30) cenh3.1 CENH3. ACAGTCGAAGCCAGTGAGTG (SEQ ID NO: 25) 1 F1 Aph8 R GCATCATCGGTGTTGCATGCAG (SEQ ID NO: 29) Aph8 F GTAACGTCAGTTGATGGTACCAG (SEQ ID NO: 30) cenh3.1 CENH3. CTACCCTCACCAATCGCGAG (SEQ ID NO: 26) 1 R1 Aph8 R GCATCATCGGTGTTGCATGCAG (SEQ ID NO: 29) Aph8 F GTAACGTCAGTTGATGGTACCAG (SEQ ID NO: 30) amplification of cenh3.1-2 and cenh3.2-2 for genotyping target name sequence PCR amplic condition on s w / size recombin (bp) ant Taq and hPROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web de Taq buffer cenh3.1-2 CENH3. CTACCCTCACCAATCGCGAG (SEQ ID NO: 26) 98C 30s, 530 1 R1 39 cycles of (98C 10s, 65C 30s, 72C 30s), 72C 5min Aph8 F GTAACGTCAGTTGATGGTACCAG (SEQ ID NO: 30) cenh3.2-2 CENH3. CCCGGCTGTCACCTATTGTT (SEQ ID NO: 23) 451 2 F1 Aph8 R GCATCATCGGTGTTGCATGCAG (SEQ ID NO: 29) detection of Zepp insertion at neo-centromere region of chr2 target name sequence PCR amplic condition on s w / size recombin (bp) ant Taq with and no homema Zepp de Taq inserti buffer on insertion neoChr TGTACCCCCACAGTTGCATC (SEQ ID NO: 31) 98C 30s, 973 locus 2 F2 39 cycles of (98C 10s, 60C 30s, 72C 60s), 72C 5min neoChr CGGGATACGTGTGTATGGCA (SEQ ID NO: 32) 2 R2 PacBio HiFi sequencing and genome assemblies
[0157] Chlamydomonas cultures were grown under asynchronously in TAP (described above) and harvested by centrifugation at 3000g for 5 min with 0.005% Tween 20.4 g of wet weight pellets were washed twice in TAP, flash frozen in liquid nitrogen, and sent to Arizona 66PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web Genomics Institute (AGI) for high molecular weight DNA preparation and HiFi PacBio sequencing. High molecular weight DNA was extracted from cell pellets using the protocol described in Puppo et al. (2017). DNA was sheared to 10–30 kb using Megaruptor 3 (Diagenode) and purified with PB Beads (Pacbio, Menlo Park, CA). Sequencing libraries were constructed using SMRTbell Express Template Prep kit 3.0 (Pacbio, Menlo Park, CA) and size selected to 10–25 kb on a Blue Pippin (Sage Science). The libraries were prepared for sequencing with PacBio Sequel II Sequencing kit 2.0 for HiFi library, loaded to 8M SMRT cells, and sequenced in CCS mode in the Sequel II instrument for 30 h. Approximately 829 (UL- 1690) and 899 (CC-400) HiFi data were assembled to primary contigs using Hifiasm (version 0.19.6) (parameter: -l 0). The N50 values after assembly were 15818 kb for UL-1690 and 14 254 kb for CC-400. To generate chromosome-level sequence scaffolds, RagTag (version 2.1.0) ‘correct’ and ‘scaffold’ with default parameters were used to correct potential misassembled contigs and orient contigs relative to the published CC-5816 genome.
[0158] Centromere variations observed between the sequenced strains and the reference strain are genuine and not due to assembly errors. First, assembly errors often occur near gap regions, particularly in areas with limited long-read coverage, repetitive sequences, or structural variations. These errors can be more pronounced at contig edges before pseudochromosome construction due to incomplete sequence information and misassemblies. However, in the sequenced strains, centromere regions contain fewer gaps—three gaps in UL- 1690 (chromosomes 3, 8, and 12) and no gaps in CC-400—suggesting that gaps have minimal
[0159] impact on centromere contiguity. Second, long reads spanning the edges of centromere regions (ZeppL-enriched sequences of UL-1690, CC-1690, and CC-400) and the reference assembly (CC-5816) were identified in the genome assemblies, further supporting the accuracy of the genome assemblies among these centromere variations. Gene annotation
[0160] UL-1690 gene models were annotated based on the gff3 file of the reference strain CC-4532. Liftoff (version 1.6.3) was used to map annotations from CC-4532 to the UL- 1690 assembly (parameter: - exclude_partial -polish). The longest isoforms for each gene 67PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web were kept using AGAT for gene model visualization in IGV and gene density calculation. Gene density was calculated by counting the gene number per 5 kb window size. CENH3 CUT&Tag
[0161] Active Motif CUT&Tag-IT Assay Kit (Rabbit antibody specific, Cat# 53160) was used according to the manufacturer’s instructions with specific adaptations were developed for Chlamydomonas cells. UL- 1690 was synchronized as described above and ~59107 daughter cells (light phase ZT = 1) per reaction were harvested by centrifugation at 3000g for 5 min with 0.005% Tween 20. Harvested cells were subject to autolysin treatment (described above) to remove their walls prior to incubation with concanavalin-A-coated magnetic beads supplied in the kit. The remaining steps followed the manufacturer’s instructions but using 1:705% digitonin instead of the suggested 1:505% digitonin.1 lg primary antibody (anti- CENH3.1, anti-CENH3.2, anti-CENH3-conserved), or rabbit IgG Control (Invitrogen, Rabbit IgG Isotype, Cat# 10500C) in a 50 lL reaction were used for each procedure.
[0162] All primary libraries were amplified with 14 cycles of PCR. After combining libraries together, they were purified using a double-sided size selection and bead clean up with Mag-Bind TotalPure NGS (Omega Bio-Tek, M1378-00). In the first size selection, large DNA molecules were removed using a 0.5:1 volume ratio of beads to DNA; then in the second size selection, small DNA molecules were removed using a 1:1 volume ratio of beads to DNA. Illumina sequencing was carried out with paired-end, 151- nt read lengths. CUT&Tag data analysis
[0163] Centromere regions were determined by identifying CENH3 CUT& Tag enriched regions. The adapter sequences of all Illumina reads were trimmed by TrimGlore (version 0.6.7) (parameter: --fastqc–paired --gzip). All trimmed reads were aligned to the genome sequence using Bowtie2 (version 2.4.5) (parameter: --end-to-end --very-sensitive --nomixed-- no-discordant --phred33 -I 10 -X 700). Two versions of the alignment BAM file (multiple plus unique mapping and unique mapping only) were obtained for each of the four datasets from four antibodies using Samtools (version 1.16.1). Multiple plus unique mapping was used to 68PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web estimate the maximum CENH3 CUT&Tag enriched regions. Unique mapping, which was the output after removing all reads with MAPQ scores lower than 20, was used to estimate the minimum CENH3 CUT&Tag enriched regions. All alignment BAM files, including IgG controls, were converted to bedgraph files using Bedtools ‘genomecov’ function (version 2.30.0) and then uploaded to an Integrated Genome Viewer browser for data visualization. All BAM files were converted to bed files using Bedtools ‘bamtobed’ function. The read depth for each 5 kb bin of the UL-1690 HiFi genome was calculated with Bedtools ‘coverage’ function (version 2.30.0). To eliminate noisy or false CENH3 read enrichment peaks, genomic loci with the counts of control IgG reads lower than 10 were corrected to 20 manually. The normalized CENH3 and IgG CUT&Tag read depth was calculated by dividing CUT& Tag read counts by total read counts in the same 5 kb bin. The fold enrichment of CENH3 reads was calculated by dividing the normalized CENH3 CUT&Tag read depth by the normalized IgG CUT&Tag read depth. CENH3 CUT&Tag enriched regions were plotted using the karyoploteR package in R. Centromere size estimation
[0164] Multiple plus unique mapping CENH3.1 CUT&Tag reads were used to estimate the upper centromere sizes, and unique mapping reads alone were used to estimate lower centromere sizes for each chromosome in the UL-1690 genome. For the upper centromere size estimation, genomic loci with fold enrichment >5.5 were regarded as stringent CENH3 peaks and then were merged (intervals between adjacent bins were <30 kb) using Bedtools (version 2.30.0) ‘merge’ function. The same method was used to calculate the lower centromere sizes, except that genomic positions with fold enrichment <10 were considered as stringent CENH3 peaks.18 merged CENH3 CUT&Tag enriched regions larger than 10 kb that overlap with ZeppL-enriched regions were identified as centromere positions (17 genetic centromeres +1 potential neocentromere on chr2). Tandem repeats and ZeppL-enriched region identification
[0165] The CC-5816 and CC-1690 C. reinhardtii genome assemblies were downloaded from GenBank accession number GCA_026108075.1 and JABWPN000000000. The CC-4532 69PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web PacBio assembly was downloaded from Phytozome. The coordinates of ZeppL were identified by a BLAST search using the published ZeppL sequence in UL-1690, CC-400, CC- 4532, CC- 1690, and CC-5816 genomes. For tandem repeat regions genomes were annotated using tandem repeats finder with parameter: 2778010502000 -d -h. The unique tandem repeats with copies lower than 400 and the period size lower than 125 were filtered. For ZeppL- enriched regions, Bedtools ‘merge’ was used to merge nearby ZeppL sequences if the intervals of their closest coordinates were <40 kb. ZeppL-enriched regions overlapping with CENH3 peaks are shown in Tables 4 and 6. Dot plot comparisons of centromere regions
[0166] Centromere regions defined based on ZeppL enrichment (see above) were extracted from UL-1690, CC-400, CC-1690, and CC- 5816 genomes individually. CC-4532 centromere regions were excluded due to too many assembly gaps in its centromere regions. All centromere regions from UL-1690, CC-400, CC-1690, and CC-5816 (UL-1690 contains 18 regions = 17 centromeres +1 neocentromere) were divided into bin windows of 500 bp and compared to the complete centromere regions from CC-5816 using BLAST. The output of the BLAST search was visualized using dot plots of the bit score from BLAST. Centromere comparisons between chromosomes and strains
[0167] From the 68 genetic centromere regions described above, 64 were gapless and used for alignment-free comparison using Mash software. A k-mer size of 9 was chosen empirically after sampling k = 7 to k = 15 and finding very similar results for k = 8 and above. A sketch size of 106 was used for k-mer sampling but smaller sketch sizes down to 103 yielded similar results. The output distance matrix was converted to a distance tree using the T-REX web server, and the tree was visualized and formatted using FigTree v1.4.4. Example 11. Phylogeny of full-length or near full-length ZeppL elements
[0168] All ZeppL elements larger than 8 kb were extracted from the UL-1690, CC-400, CC-1690, and CC-5816 genomes producing a set of 49 full-length or near full-length ZeppL 70PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web elements among these strains (Data S2 of Liu et al., 2025, The Plant Journal (2025) 122, e70153, the disclosure of all of which is incorporated herein in its entirety). The sequences were aligned with MAFFT and used for construction of a neighbor joining tree in MEGA11 set with the Tamura 3-parameter model + gamma (0.05) and 500 bootstrap replicates. Example 12. Functional neocentromere formed by single Zepp insertion.
[0169] Upon further analysis, it was discovered that chromosome 2 in strain UL1690 split into two sub-chromosomes after fortuitous Zepp transposition (FIG.32A). A revised assembly of UL1690 chr2 sequences after splitting into two sub-chromosomes (2a, 2b), each with its own centromere and new telomeres at the break points are shown in FIG.32B. 71PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web SEQUENCES SEQ ID NO 1 CCTGTTGAGCCAGGCCATTAATCAGGCGCAAAGAAAGGGTCAATC Genomic cPROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GACCAGCAGTCGCTATGGCGAGGACGAAACAGTCGAAGCCAGTGA GTGTGAGGATATGACTCCATGTCGGTCAACGCCGGGGTTCCAACC T T A TT A T A A T A A T73PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web AACTGAGTACCACAATCGCCAAGTTCTGTCTTGGCTGTGCTGTTG CTTGACAGCGGGACGTTGTTGTGTGTGGTTACAAGGTGGTGAGCA A T ATT T T TT TA AAA A cPROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web AGAAGCGGCGCGGGGAGCCCCAGCTCCACCACGGACGCCAGTCGC CCGTGAGCAGCGGCCGACGTGATTGCCCCCGCGCTGCCGGTCCAC A A A A A T A A A ATATT AT75PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GCAGTGCCCAAGGACATGCAGCTGGCGCGCCGCATCCGTGGGCCC ATCTACGGCATTTCCAGTAACTGAACGGCCGCTGATGTTTTCCGG TA A A A AATT T T T ATTTTAA A A AAA TPROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GCTATGACACATGGGGCGCCAGACAAGCTGATGACGATAATACGC CACGCCAGCGCCTGGAAAGAAACGCCGTCCAAAAGGAAACCTAGG T A A T AT AA T A AA T AA TT77PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GGGTGTACAAGTCAGCTGCAGCCGCATCCGAGTCCCACGACAGGG GCACACCTAGGTGCGTCACGGCGCCAGTGGTAAACGGCACCCCCG T T T AA T A T AAA T AA A78PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GCAGGCGCAGGTGACCCACGTCCCAATCCCAACATACCACACGGC CCGCCGCATCCGTTGACGGCGGCTGCAGTACGCAGCCTGGAAGGG A A A T A ATA A A A T79PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CTGGAGATTGAGGTCCGCCAGCGCGGCCAAGTACCGCCGCATCCC AGCGGTCTGCTCCTGATCCTGCCCGGCGGCTGCCCGCTGATCCAA T T A A TAT T T T APROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GTTGACGATAAAACGCCACGCCAGCGCCTGGAAAGAAACGCCGCC CAAAAGGAAACCTAGGGGGCCTGCACAGGTGATGCAAGCGTCAGC AA TT AA TTA T A A T AA T T A TT A81PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GCGCCGCCAGTGACAGACGCAGGCTTCATGCTGGGGGACCGCATG GGTATGTGGGCCTCAGGCCCGTGGGGGGCGGGCGCGCTGCTGTGG A A TTT A TT TT TA T T T T TA82PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web ATTCGCTCGAGGTCATGGGTTCGACTCCCATACACCCCTCCGATT TCGATCGGAATGAATAGACTACAGACATGCAGTGGGCTGGACGAG TTTA TTA TT ATTT T ATT A T AA TTT T83PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web TTAGGAGCAGAACCGGAGGTAAATGGAGCTCTGACGAGAGCAAGC CTACTACTAGAGTAGGTCGTTGGCGTTGCAGGTGGCACGTCCCTG A T A A A AT T T ATT T A AA T T84PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web TTTCGTTAAGGCTTCGAGAGATCCTGGGTTCGAATCCCGGTCACC CCACCAGGAAGGGCCCTGGTTTACCAAAACCCATCCACACCTTGG T T T TA A T T T A T T T T TA A85PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GCTCTGGGCCTGTCAATGGGAGAGTTCTAGATGATCTGTGTGGTC AAGGCAGTGGCTGTAGTTGTCACCTAGCTGTTGAGAACTCCAGTA T T T T T A T T TT A TTT A TA T T T86PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CCAGGCTGCCTTGGCAGTAGGGTGGTAGTGCGTCTACCACCGCAA CCTTGTAAGGATGACTTTAAAAAAAAAAAGACGTCGGAGCATGTG TT A T T A A T A T AT A T T87PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GAAGTGGTAGGCCAGCTTCGCCGCCAGCACCTGCTTCGCAATGTG CACACGGCCAACGAGAGTCAGAGACAGGGCAGCCCACAGACGCGC A AAA AT T A T TA AA T A T88PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web TCCCCACCTGGCCACGGACAGGGTGCTTGTTGTCGGTGGGGACTG GAACTGTGTCACCGATGCCAGTCAGGAGGCGGCCCCTAGCCCGTC A T T A T AA TT A T T A TT89PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web AACGGGATGCTCTCTGACGCCTTCCCAGTGCTGAACGGCTTGCCG CAGGGCAGCACCGCCTCACCACCCCTGTGGGTCATCCAGATGCAG A T A T TTT T A T A AA A T90PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GTGGGCCTCAGGCCCGCGGGGGGCGGGCGCGCTGCTGTGGAGCAC TTTGCGGGCCACCTTCTTTTTTTTTTTAAAAGTCATCCTTACAAG TT T TA A A TA A TA T AA A T91PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CGCTTCATGGGGCTCAGCCCATTTTCCAGCTGTACAAAGCTGACA CCCCTTGTTTTTTTAAAGTCATCCTTACAAGGTTGCGGTGGTAGA A T A TA T AA A T AA ATAT A A A92PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web AAAACACCAGCGCTACGCACTCGCCCCGCCCAAACCCACCCCACC CCCACAAGTCGGTTGCTTAAGCAACACCCCAGGAGAACCCAGGAG AA A A A A A AA A AAAA TAAA A AT T93PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web AAAAGCAGCGGCGGCGGTGGAGGGCGACCCTCGCATGCCGCCCCC GGATGCGTGCGTCCCTTCTGCGGGGTACGTGGCGGAGGCTGTGGG A A A A T AAA T T T TA T94PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CGTGGCGGATGGCGGGGGCCTGGGCATAAGGGAGGCGGGCGGGCT CTTTAGGAATAGGAGTCTCTGCGGTAAAGCAGAGGATAAGGAAGT ATTAAT ATA T A A A T TA AAT95PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web ACAGTACCAAAACAAGTCTAAGCATAAATCGTGAATGGATGGCGT TATGGATCCCCACTTCGTTAGCGCTGCGTGCGCTACTTGGCTATC A T TTTT T T T T T A T T T A TT A A T96PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CGCCCCCAGCGCCGCGAGCCGCGGCCACAGCTCGGCCATCGGAGC CCTCCGCCTGCAGCAGCGATTTAACCGGAAGCGGCGCCGTCACAA A T T A A A TTT A TT T T97PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GGCAGATGGACGGCACGGTGGTGACTAGGCGGGTGCGGACGGGCA GGGGCAAGGGCGCGGGGGAGACGGAGCAGAAGGCCCTGCTGCAGT T T A T TA A TA TTT T AA T98PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GATGACGCCGCGCAGCAGGTGCAGACGGCCTACATGGACTACGTG CTCCTAGACGCCCGCACCCTGCGGTCCGCTAAGCGTGGGCCGGCG A A A T A AAT A A99PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web TGCAGACCGCGTTCATCACCGGCCGCTTTTTTTTTTTTTTAGGAA TCATCCTTACAAGGTTGAACAGGGGCAGACGCGCTGCCCCCCTGC T T AAA A T T A A A AA T A A100PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web ACCAACCACCCCCAACCATCCCACCCCACCCCACCCCGACCCAGC AGACAGAGCAGGAGGCTGGCCCCCCCGCCCCCCGGGAGGGGTTGC A AAAAA A A A TAA AA A AA TA A A A101PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CGCTACGCACTCGCCCCGCCCAAACCCACCCCACCCCCACAAGTC GGTTGCTTAAGCAACACCCCAGGAGAACCGGCAGAGGAGACACAA A AAAA TAAA A AT T T A TAA T102PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GTGGTTCGGGAGGTGGTCAGCGAGCTGCGCAGGGTGATGCGGCTG CGTTTCACTGCCGCCACGCTAACCCCTGAAACCCTGTCGGCTCTG A T A TT T A A A T AA TAA T A103PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web ACCCTGACAGGGCAAGAAGAGAAGAAAAGAGAAGAAAACAGAAAG CAAACACCTCGAAAACCGCCACCAAACACCCTCATCCATCCCACC A A A A A A A A A A A T104PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web ACCCCCCCGCCTGCTGAGCTGGCGCCGGGCACGGCATGGCTGCGG GGCGGGCCGGAGCCGGAGGTGCGGCCGGGTTAGCCGCGGTGATGG A T A T TA T AA T A A105PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CCGGCTGCACACGGAACAGGCCGGTGGGGCTGGTGCTGCTGTACA TGGCTGCGGCGGCTGCGGAGACAGTGTTGCGCGTGCCCGGCCCCG T A A A TT A TT AA T T A106PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web TGGGCGGGGCGTGTGTGCGTAGCGCTGGTGTTTCCTTTTCTGGTG CGCGGTGGTGGTGGTTCTTGCGGTTAGGCTGTGGTGTAGGTTGGT TT TTT T T T TTTTT T AA T107PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web TGCCCCCCTACTGCCGGAAAGCAGCCTGGTCCCCGAGACCAGAAC CCTGACAGGGCAAGAAGAGAAGAAAAGAGAAGAAAACAGAAAGCA AA A T A A A AA A AT AT AA A108PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GCCCCCTAGGTTTCCTTTTTGGCGGCGTTTCTTTCCAGGCGCTGG CGTGGCGTTTATCGTCATCAGCTTATCTGGCGCCCATGTGTTTTA T T T T A T T T ATTT T109PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web AGCGCGCCCGCCCCCCGCGGGCCTGAGGCCCACATACCCATGCGG TCCCCCAGCATGAAGCCCGCGTCCGTCACTGGCGGCGCCGCTGCC T AA A A A T T A A T T110PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web GCCCGCCGCATCCGTTGACGGCGGCTGCAGTACGCAGCCGAGCCC CCAGTAGCATTACCGGCCATAAGCACCACTGGCACGGGCGCCGCC A T T T A TTTTTTTTTTTT A AAT111PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CATGGGCGAGCTGGGGATCAAAGCAGGCGCGCGCAGCAGCAGGCG CGAGTAGCAGCCAGAAAATGCGAGAGCACTATCGTTTGAGAGAGA T A A A T A A A A A AAA112PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web AACTCACCGACCTGGTGGACCACTTTGCTGCGCGCTCCATGCACG CTGAGGACGCCAGCCTGGTGTCGCACGGGAACCCGCTCCTGCTGC AAA AAA T T T TA AA AT T T A113PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CCTGTGCAGGCCCCCTAGGTTTCCTTTTGGGCGGCGGTTCGTTCC AGGCGCTGGCGTGGCGTTTATCGTCATCAGCTTATCTGGCGCCCA T T TTTTA T T T T A T T TgRNA sequences cenh3 mutants generated w / gRNA in g2 and g6 were used in the Crispr Cas9 additional mutagenesis characterizations guide RNA target with PAM target target gene label site location CAGGCTAAGCCCGTAAGTCAAGG Cre02.g104800 g1 (SEQ ID NO: 55) 1st exon GGCAGCCCAGCGGTCTACGCGGG g2 (SEQ ID NO: 56) 2nd exon CCGCACCGCTATCGGCCGGGCAC(SEQ g3 ID NO: 57) 3rd exon AGTCGAAGCCAGTGAGTGTGAGG 1st exon GGg6 (SEQ ID NO: 59) 2nd exon CCGGCCTGGAACCGTGGCACTGC g7 (SEQ ID NO: 60) 3rd exon gRNA expression constructs 114PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web cenh3 mutants generated w / g2 and g6 RNA i d i h G T G T G T G T G T G T115PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web (SEQ ID NO: 67)CENH3.1 rescue plasmid. Cre16.g661450 Chr16:2571785..2574609. Comprises expression Construct expressing native genomic CENH3.1 gene LOCUS Cre16.g661450 7203 bp DNA circular SYN 04-JUN-2024 DEFINITION synthetic circular DNA ACCESSION . VERSION . KEYWORDS . SOURCE synthetic DNA construct ORGANISM synthetic DNA construct REFERENCE 1 (bases 1 to 7203) AUTHORS Benson Hill TITLE Direct Submission JOURNAL Exported Jun 4, 2024 from SnapGene Viewer 7.2.1 https: / / www.snapgene.com COMMENT Sequence Label: CENH3.1 rescue Cre16.g661450 Chr16:2571785..25... FEATURES Location / Qualifiers source 1..7203 / mol_type="other DNA" / organism="synthetic DNA construct" misc_feature 1..4897 / label=Chr16:2571785..2574609 primer_bind 1..20 / label=gCre16 F1 misc_feature 1 / label=Cre16:2571785 misc_feature 1653..1904 / label=5'UTR misc_feature 1905..1907 / label=start codon misc_feature 1908..1931 / label=exon1 misc_feature 1932..2034 / label=intron1 misc_feature 2035..2161 / label=exon2 misc_feature 2162..2384 / label=intron2 misc_feature 2385..2485 / label=exon3 misc_feature 2486..2676 / label=intron3 misc_feature 2677..2751 / label=exon4 misc_feature 2752..2915 / label=intron4 misc_feature 2916..2991 / label=exon5 misc_feature 2992..3207 / label=intron5 misc_feature 3208..3272 / label=exon6 misc_feature 3270..3272 / label=stop codon 116PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web misc_feature 3273..4870 / label=3' UTR primer_bind complement(4596..4617) 8 RP LS HQ IA 8 RP LS HQ IAPROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web LATRDIAEELGGEWADRFLVLYGIAAPDSQRIAFYRLLDEFF" terminator complement(6862..6909) / label=T7 terminator / note="transcription terminator for bacteriophage T7 RNA polymerase" terminator complement(6862..6909) / label=T7 terminator / note="transcription terminator for bacteriophage T7 RNA polymerase" polyA_signal 6981..7188 / label=bGH poly(A) signal / note="bovine growth hormone polyadenylation signal" polyA_signal 6981..7188 / label=bGH poly(A) signal / note="bovine growth hormone polyadenylation signal" primer_bind join(7184..7203,1..20) / label=gCre16 F1 OH primer_bind complement(join(7184..7203,1..20)) / label=gCre06 IO OH R1 ORIGIN 1 cctgttgagc caggccatta atcaggcgca aagaaagggt caatcacagc gacaacaaca 61 tcagagcaag tgtcgcggca catccatggt ccatgcaaat gaagcgccgc ggcctagctc 121 cagccccctc gctcacctcg ccggcgtact gctgcaggat ggcctccttg taaagctcca 181 gccgctgcgc cgtgtggttg gcccattcgc cgccgccagc aggcgccgcc gccgccgcca 241 ctgccagctt ctgcccctgc catagggcat caggctctcg tgcctgtccg tctgcctcca 301 cctccatttg ctgctcctgc tgctgttgct gctgctgctg ctgctgctgc tgctgcaggg 361 cgcccaactg gtcggcgagg ttgcgggtgg cgaggaagcg cagctcgcgc agcgtgccgc 421 gggtggtgcc tgcagcggca gtcgtgcagc ccaatggcgg acacgcgcag gtcagcaaca 481 tggcggctgg tggacaccgc aaggtctcat tccagagtga tagccacgcg tgcgaccagc 541 aacaccactc actcatatgg ccatccgggc cgatggccaa aagagggtgg tccagcacac 601 tctgtggaaa ggggtgcagc aaggcagcct ggtgagacat gcgatgggaa gctgcaagcc 661 cttgcgtttg ccaggtggca tcctaaccgt gcacttgctt gaagccgcgt cgcacagacc 721 atgcatggca atcacacagg aagagagcaa tacccgctgt atcttacggc ccccgccact 781 gttgttgatg tcgcacctgc agctgctgca ccgcccgcca tgcgtcgccc tcggcagcac 841 tccgcagcgc cgactggtac agcgacacca ggcctgcctc ctgcgcctca ggcgtttcct 901 ccggggcatc tccaggctgt gcagccacat ctgtcctcgt gacgccagca ccgtcgtcct 961 gcgttgtgca cgcaggcgtg ggtgtcaaat gtgcccaatc ctggatcgta gatacagaaa 1021 cgcggcccta acacaggtct ggtatgggat ttgctcctga cctgcagccc gtcttccccg 1081 gctagcacaa ttggtccacc gagcttgctt gcaattcatg agcaccgtgc gctgtcgtcg 1141 cgtgccccgg cacttcaagt aatcctggcc ggccccgcag cggttcagga ggtagccacc 1201 cgcgtttgcc cctcaacagg cagcggctcc cccgtatcct accactttgg ttactggtgc 1261 atttttcgcc gcgattccac cgaggcggac cattatgggc taagcctccc gaactcaggg 1321 cgggccccgc cgttccgtgg tggccacgga tggccaggcc ccgcccaaca ccccctcatg 1381 cgcagacaat tctacttgat ttgcttgggc acatgaatcg actcatgctg ttcgtgccac 1441 gctgtacgcg accggtaaac tgtaagcagc gcaagattac tacgactttt cgtgcgcaaa 1501 gcgtcgtgct gtgaaatact agctatttat atgaattgca gggtacgagt tgttggcttg 1561 tgcacgtaga gcggcacctg ctattgagtt gggattgggc gggagtccgt ttttcagtgt 1621 cccatcaatg tgtgccttct gcaaacggtt cacatggaca tatactaaag catcgacgca 1681 acatagtttc tagataacaa ctcaaaggat caagcagaca gaaatccaga cgggctgtca 1741 aaatagtgcc ctggttgact cgcgcctttt gtatgaagcg ttacgcaagc aacagcctgc 1801 aagactagca gtagattacc gggattagtg caagcgttgc gcgctgtggc acgacgacct 1861 tcccaattac tgacttggac gttgacttgt gaccagcagt cgctatggcg aggacgaaac 1921 agtcgaagcc agtgagtgtg aggatatgac tccatgtcgg tcaacgccgg ggttccaacc 118PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 1981 tgcgtgggca cgttcggcag cgtggaccag cggctgacca gcctcctcgc gcaggctcgg 2041 ccagggcgaa aggcgcaggc cgaggcagca acacctacga aatcgaagcg accaagcgga 2101 gctgcagcga cgccgacgcg cgggggccga agcccgggtg gaggcacacc aacgggaaag 2161 agtgagctcc gcgcatgccc ctgcgcatgc cggagtagca gtgctttgca acagggattc 2221 caaacgcaca tgcccaatga ctctcctgcg cccagactct tgttggctat tgatgcaacg 2281 atgggagtcc tgcacgctag atggcagatg tggtgatagg ggtcgcgcca tggcatgccg 2341 tcacacgtgc gacaccaaag ggcatgcctt tgtcccccac gtagagcccc accgctaccg 2401 gcctggaacc gtggcactgc gagaaatccg caagtaccag aagtccacag acctgctcat 2461 ccgcaagctg cccttctcgc gattggtgag ggtagcacgg aaatggggtg ctctcgcccg 2521 caccgtcaac gtgtctgcgg gtcaagcaca atcctgactg cattgtgcgc tgaagaatcc 2581 ctcccctgtg cattgaacag cagtgtttgc ttctgatagc aaccccaatc catacacatt 2641 atccaacata ccaaaactgc tggcatgtgc atgcaggtcc gtgaaatcag caaccagatg 2701 ctgcgcgagc ccttccgttg gactgcggag gcactactgg cgctgcagga ggtacgcagc 2761 aagagggcgt acaaggacgg acatgcatgc tgctgttgct cctgtatcct tttcgctggg 2821 tgaggggtgg tggggtcagc ggaatcgtgt ctcagttggt cacggccaga ctgctcagac 2881 cgttgtcact gccctcccct catctgcaca cgcaggccag cgaggacatg ctggtccacc 2941 tgctggagga ctgcaacctc tgcgcaatcc acgcaaagcg cgtcaccatc agtgagtgca 3001 cgggaagcaa aggggttgca gtagcacgcc agcaggcagg gaatgggagc tctgctgggt 3061 ggcgcgcact tgcgacatgt ggcggtgcat gcatccttgt ggtcaacgct gcgcggagtg 3121 gggttgatgg gggcattggg ctgtacaggg ctagcaatga caagcattaa cttcattaca 3181 tggtatccct gcactttctg gctgcagtgc ccaaggatat gcaactggcg cggcgcattc 3241 gtggaccaat ctacggcatt tccagcaact gatgtggccg agtgtttaat gcggggcatg 3301 gcggacgtga cttgcggctg ggaggcatga agtgaaagct ggttgtggag gccgttgtgg 3361 acatgctgaa attctcgtga cgttgcataa ttgactgctg ctggtggcat cacgttgctg 3421 agcatgttga aagtggtcct gcgctgcagg gctgcgtaag attaagttat aacatggtcg 3481 acagagctgc agggtagtgg tggctgccta gtagagggaa aagctgtcag gagctaggtt 3541 ccccatcatg ctttctagta aatactgttg gaggcaaagg gatggcttcc atgacaaacg 3601 acgatgccga aaggatgtat ggggattgag ttcgcggtgt accgtgttat atgactgatt 3661 tacgatgcgg ggctgcttgc acctgctgcg ctggtaggcc aagcggccat ggcccatgag 3721 accccaggat ggctgggggc ttggctgggg gcgctcaacc ggtactggta gctgtcatgc 3781 actgccggac gggctgaaga tcgactatac tgcccctgga gtgacgctgt gctcctgcag 3841 cgaagtgcct cgtgttgagc tggcactgtg tcatgtgatg gaagctcggg gaattcggac 3901 actatgagga tgccacatga agttctgcgg tagccagtcc agtccagagc tggcccgatt 3961 gccattgcat tccttgagag caaaacaatg cctgggcttt tacaatagaa tttagcacag 4021 aacgcatagt aagcaaaata ccaccacgca aactgagtac cacaatcgcc aagttctgtc 4081 ttggctgtgc tgttgcttga cagcgggacg ttgttgtgtg tggttacaag gtggtgagca 4141 cgacgtgcat tcctgtcctt gtagaaaccg aggcccgggc cggggggccc ttataacgcc 4201 tgcaacactt atatatgcag ctagtttggt ccccgaatgg gggtgcttat gctcctgtgc 4261 tcccaaatgg aatctagtaa cgcggctacg ggacgtggct ggacagtgct ggacacatta 4321 cggcgggacg tgcctgtacg aaccgacccc ggctatgtcc gggagatgag gtagacgctt 4381 gcggcctcag actagacacg tggctcagcc cacttttccc atgcgaggga agatatcttg 4441 tcgcacagaa ccgaagcctg gctttcgatg cgctagtgga cgaggcacgg gagacatggg 4501 cgaacccggc ggcagccgga tggtgcgcgc ggctgtggct gcacagggtg tcgacgtgtg 4561 agtacgtgtg gcataggccg ggtacggatg agacgaggta ggagtggaga gcaggcacgc 4621 tcttacccga cggaactatc ttatccatct tcagagccca tgtgatgagg gcgtgcgacg 4681 gtgggaacat gctggttgcg gaccagcatg cattcccctc agatggcaag caacccggca 4741 tacacatgtg caaatataag cattggatgg atctaagggt gagcagtggc ttgcagttca 4801 gtaaaagcat gcgcgaagtg atgctgaaat ctaagcctgt caggatagca gcattactca 4861 caaaataact cccctgcccc cttgttctcg ctttcctccc gccgcttaaa ttcgctgccc 4921 agcgctgcct ctcaattgtt ccacatgtgc gactccgtct cgctgccggg ctttgctgct 4981 ggcggcgtcc tgcagcgtgc tggaacccac ttttggaccc ccatgggcgg gctgggcgta 5041 tttgaagcgg gtaccccata acttcgtata atgtatgcta tacgaagtta tggtaccgcg 5101 gccgcgtaga ggatctgttg atcagcagtt caacctgttg atagtacgta ctaagctctc 119PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 5161 atgtttcacg tactaagctc tcatgtttaa cgtactaagc tctcatgttt aacgaactaa 5221 accctcatgg ctaacgtact aagctctcat ggctaacgta ctaagctctc atgtttcacg 5281 tactaagctc tcatgtttga acaataaaat taatataaat cagcaactta aatagcctct 5341 aaggttttaa gttttataag aaaaaaaaga atatataagg cttttaaagc ttttaaggtt 5401 taacggttgt ggacaacaag ccagggatgt aacgcactga gaagccctta gagcctctca 5461 aagcaatttt gagtgacaca ggaacactta acggctgaca tgggaattag cttcacgctg 5521 ccgcaagcac tcagggcgca agggctgcta aaggaagcgg aacacgtaga aagccagtcc 5581 gcagaaacgg tgctgacccc ggatgaatgt cagctactgg gctatctgga caagggaaaa 5641 cgcaagcgca aagagaaagc aggtagcttg cagtgggctt acatggcgat agctagactg 5701 ggcggtttta tggacagcaa gcgaaccgga attgccagct ggggcgccct ctggtaaggt 5761 tgggaagccc tgcaaagtaa actggatggc tttcttgccg ccaaggatct gatggcgcag 5821 gggatcaaga tctgatcaag agacaggatg aggatcgttt cgcatgattg aacaagatgg 5881 attgcacgca ggttctccgg ccgcttgggt ggagaggcta ttcggctatg actgggcaca 5941 acagacaatc ggctgctctg atgccgccgt gttccggctg tcagcgcagg ggcgcccggt 6001 tctttttgtc aagaccgacc tgtccggtgc cctgaatgaa ctgcaggacg aggcagcgcg 6061 gctatcgtgg ctggccacga cgggcgttcc ttgcgcagct gtgctcgacg ttgtcactga 6121 agcgggaagg gactggctgc tattgggcga agtgccgggg caggatctcc tgtcatctca 6181 ccttgctcct gccgagaaag tatccatcat ggctgatgca atgcggcggc tgcatacgct 6241 tgatccggct acctgcccat tcgaccacca agcgaaacat cgcatcgagc gagcacgtac 6301 tcggatggaa gccggtcttg tcgatcagga tgatctggac gaagagcatc aggggctcgc 6361 gccagccgaa ctgttcgcca ggctcaaggc gcgcatgccc gacggcgagg atctcgtcgt 6421 gacacatggc gatgcctgct tgccgaatat catggtggaa aatggccgct tttctggatt 6481 catcgactgt ggccggctgg gtgtggcgga ccgctatcag gacatagcgt tggctacccg 6541 tgatattgct gaagagcttg gcggcgaatg ggctgaccgc ttcctcgtgc tttacggtat 6601 cgccgctccc gattcgcagc gcatcgcctt ctatcgcctt cttgacgagt tcttctgagc 6661 gggactctgg ggttcgaaat gaccgaccaa gcgacgccca acctgccatc acgagatttc 6721 gattccaccg ccgccttcta tgaaaggttg ggcttcggaa tcgttttccg ggacgccggc 6781 tggatgatcc tccagcgcgg ggatctcatg ctggagttct tcgcccaccc cgggatatcc 6841 ggatatagtt cctcctttca gcaaaaaacc cctcaagacc cgtttagagg ccccaagggg 6901 ttatgctagt tattgctcag cggtggcagc agccaactca gcttcctttc gggctttgtt 6961 agcagccgga tcttctagaa tccccagcat gcctgctatt gtcttcccaa tcctccccct 7021 tgctgtcctg ccccacccca ccccccagaa tagaatgaca cctactcaga caatgcgatg 7081 caatttcctc attttattag gaaaggacag tgggagtggc accttccagg gtcaaggaag 7141 gcacggggga ggggcaaaca acagatggct ggcaactaga aggcacagtc gaggctgata 7201 gcg / / 120PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web (SEQ ID NO: 68)Cre02.g104800 plasmid. CENH3.2 rescue Cre02.g104800 Chr2:4281439..4286742 Comprises expression Construct expressing native genomic CENH3.2 gene LOCUS Cre02.g104800 6880 bp DNA circular SYN 04-JUN-2024 DEFINITION synthetic circular DNA ACCESSION . VERSION . KEYWORDS . SOURCE synthetic DNA construct ORGANISM synthetic DNA construct REFERENCE 1 (bases 1 to 6880) AUTHORS Benson Hill TITLE Direct Submission JOURNAL Exported Jun 4, 2024 from SnapGene Viewer 7.2.1 https: / / www.snapgene.com COMMENT Sequence Label: CENH3.2 rescue Cre02.g104800 Chr2:4281439..428... FEATURES Location / Qualifiers source 1..6880 / mol_type="other DNA" / organism="synthetic DNA construct" misc_feature 1..4698 / label=Chr2:4281439..4286742 primer_bind 1..20 / label=gCre02F2 misc_feature 1 / label=Chr2:4281439 misc_feature 1821..2105 / label=5'UTR misc_feature 2106..2108 / label=start codon misc_feature 2109..2132 / label=exon1 misc_feature 2133..2212 / label=intron1 misc_feature 2213..2345 / label=exon2 misc_feature 2346..2438 / label=intron2 misc_feature 2439..2539 / label=CENH3.2 exon3 misc_feature 2540..2761 / label=intron3 misc_feature 2762..2836 / label=exon4 misc_feature 2837..3042 / label=intron4 misc_feature 3043..3118 / label=exon5 misc_feature 3119..3289 / label=intron5 misc_feature 3290..3354 / label=exon6 misc_feature 3352..3354 121PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web / label=stop codon misc_feature 3355..3843 / label=3' UTR primer_bind complement(4679..4718) / label=gCre02 R2 OH primer_bind 4679..4718 / label=gCre02 IO OH F primer_bind complement(4679..4698) / label=gCre02R2 misc_feature 4698 / label=Chr2:4286742 protein_bind 4735..4768 / label=loxP / bound_moiety="Cre recombinase" / note="Cre-mediated recombination occurs in the 8-bp core sequence (ATGTATGC) (Shaw et al., 2021)." rep_origin 4821..5177 / label=R6K gamma ori / =" li i i i f E li l i R6K; 8 RP LSHQ GLAPAELFARLKARMPDGEDLVVTHGDACLPNIMVENGRFSGFIDCGRLGVADRYQDIA LATRDIAEELGGEWADRFLVLYGIAAPDSQRIAFYRLLDEFF" terminator complement(6539..6586) / label=T7 terminator / note="transcription terminator for bacteriophage T7 RNA polymerase" polyA_signal 6658..6865 / label=bGH poly(A) signal / note="bovine growth hormone polyadenylation signal" primer_bind join(6861..6880,1..20) / label=gCre02 F2 OH primer_bind complement(join(6861..6880,1..20)) / label=gCre02 IO OH R ORIGIN 1 actgcagaaa ccaggatcgg cggctgctgc ggttgcggcg ggcgcagccg cggcagcagg 61 gccgcccgca gcgtccgcac cgtcagattc aacggcagca caatgtcggg cagccttgct 121 ggctcctcca ctcgccggac tggtgctatt cttcttctgg gcacaggagc tccggcccct 181 tgaggctgct gcggcggcgc ctgcgcccgt gcgcctacag gcgccgcctc cacagcagcg 241 gccgatgccg gcactggtac tgctgttgac gtggacggct gcggtggcgg cagcggtagg 301 cgggtgatgg ttggtaccgc tgctgcaggc accatggcag cggccgggcc tgctgctgcc 361 gctgccggcg gcggtggcgg gggtggcggc aggagcgttg ccacttgcag gcgtgtgaga 421 gccgtgaggg ccgtgagctc ctgcccggcc accagacatg cggggtagag gcacagctct 481 cgcaagtggc tcaggtgcgt gaacgccgcc aggctgtaga tgggcaggag catgccacgc 122PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 541 aagccgccac cgcagcggtt gttcagaacc aggtgggtca ggttggtcag cgccggcagc 601 accgccacac cgcccccatc tgcgacccca attccgccgc cgccgccccc gccatcgcca 661 ccctgctgtt tcggctgctt tgcgattggg tagcgcgagt caacccgcag ctcacgcagc 721 cacgtgaggg tggtgttggc aaaggaggca tcagcagcac ggccggcccg gtccaggtcc 781 tcatagctga gcatcagcgc agcaggctgc agctgcatca gcgcctgcct catgccaggc 841 gacagggggc cgggcaggtg cttcagccgc agcgaggcca ggcggcgtgg gcagctgctg 901 ctgctgctgc cgccgccgct gttgctactg ccgttgctac ccgccggtgc agggtcagcg 961 gcgccagcgc ccagtccacc agcagcattg acgtcaacac gccgaaagag ccgcagcagc 1021 agcgcggcgg cgttgccgcc gccgaggcgg ttgagatggg tcagctcgaa ccggctgaag 1081 gctggcagcg ctgtcgcgcc cgccgcgtca gccagctgcg gcgtcagaag cggcgcgggg 1141 agccccagct ccaccacgga cgccagtcgc ccgtgagcag cggccgacgt gattgccccc 1201 gcgctgccgg tccacagggc aggggcaggg cagagtcagg gcaggggaga tattgcatcc 1261 atattgcatc gcgctggtac aaaacgcccc gacatgcttt catgccgtgc attggacata 1321 acttgcacaa tcaacgcgct tacatttcgc ccgacaccga agcgctgaag tccgagtcct 1381 cggggtcgcc gatgtagttg aagcgcacca cgagagggcg gtagccgcgc gcgagcatgc 1441 gcacgaatga ctgcagcacg tttgttgatg tatgctgggc gctcacggcg cgcgcagcgg 1501 tcgggagctg gatgcggtgc tcgtcgctga gcaggtgcgc gtcggttgcc aatagcagcg 1561 gcctcgccac aagcctcagc cggctgagga agccacaccc taagcttgag cagatcagct 1621 tcatcacctc gccgtcaccg tagccggcca ggtccaaaaa agggtgccgt gcctgtgcct 1681 gtgcgtcttg cattcttcaa gtaccctttg agccgaaagc tgaaacaaat ttgaactata 1741 tgtataatgt tttacccagc gcttacgagg tcgggcgcca gaacgatgca tttcgcggtc 1801 accgcatggg cgttgttgcg cgatacacaa ctgctctacg caaatgcccc agcaaaatgg 1861 gtgctctgcg tagcgcatcc ctgtattcat tttacatgtg atgctatggc aacggcgtag 1921 ctaagctgtc agaaccctgc cccggctgtc acctattgtt tgctcgcccg cagcccgtgg 1981 cctgcgcaaa taacgttctg tcatcgttga ggcggtcacc tcgaagctgt cggaggtcgt 2041 gtacgttcgg actcctcaat tgccgcttcc tccgtttgcg aacagcccgc cgctgtgttt 2101 acaacatggc tcggactaaa caggctaagc ccgtaagtca agggctgagg cagtgcaaac 2161 aggggcgtcg cctaagactg ctcggtcttg gtgacctgct ttgtgaccgc aggcccggcc 2221 tggacgagag gcggcagccc agcggtctac gcgggccggc gaagccgtgg gctcaccggc 2281 acgcggcggc acgcccgctc gcagtggacg gagcccagcg ggtaaagcag caaccggggc 2341 caagagtgag tcttgtggct gttttcgcca gcaggcttca agcaacgggg aaggggctgg 2401 ttgctgtgct ttcctgacca cgcccctgcg atacacagag ccgcaccgct atcggccggg 2461 caccgtggcg ctgcgcgaga tccgcaagta ccagaaatcc acggagctgc tgatccgcaa 2521 gctgcccttc gcgcggctgg tgggtgcatg ggtgggggcg gcggcatagg ctcaaactct 2581 ttcgaagccc aaacgaaact gcagcaaggc gcctgtacgc tgagcaccgt gtgcttgtgc 2641 gggtttatgc ccgagaacta cagcaagcag gcaccatcca acgcagacac caccagtcac 2701 gtgagcgagt gacgtgggcg agcgtgttgc tgctcctccg ctggcgactg ctggtacgca 2761 ggtccgcgag atcagcaacc agatgctgcg cgagccgttc cgctggacgg gggaggcgct 2821 gctggcgctg caacatgtga gccgccgccc gctgtgtgcg gggggggggg ggaggcgcat 2881 gcagtggcta gcatgtagag caatgagcaa gcatatgcat gcgtgttgcc cacacaccca 2941 tctcccaacc cgccatctcc cacctgccta ggcaatgcct ctgatggcct cttccacacg 3001 cgtgagcctg ctgcccgctc gctgcccgct gcccgcatgc aggtggccga ggacgtgctg 3061 gtccacctat tcgaggagtg caacctgtgc gccatacacg ccaagcgcat caccataagt 3121 aagtaactag cttggagggt gggcaaaaaa ccgctttctt tcaatagggg tgcgtgtgct 3181 gtcgcgcctc gcagtcacca tttgttgaca gtgttgcttt tttgctcttg tgggttacac 3241 cacgcatagc gttcgctagc gaagtgtttc tacacacaat ttcctgcagt gcccaaggac 3301 atgcagctgg cgcgccgcat ccgtgggccc atctacggca tttccagtaa ctgaacggcc 3361 gctgatgttt tccggtaggc aggcagacaa ttgtgtgctg cattttaacg agcacaaagt 3421 tagacggtgt gagcaaggag gggcataggc agccaccgag gtcatgtcaa gcttgccttt 3481 cgtctcttgc cttgctactt tgcaggggat ggtatggcgg tggccgcgct caccccggca 3541 ccctatgacg ccgccgccac cttcacggca acgtggcagc aacaggcaga aagccgcaac 3601 tccggttgct gccacagcag cagtggtgat gcgtgtactt tgcagcgtgc ggcgcggcaa 3661 tcctcgcaac taccaacact ttaactgaca cacggggacc acaaacgggc tttcgcgcaa 123PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 3721 ctgccaacga acaggggcag cgcgtggcgt ggtggttggc gcggggcgcg agtgggagct 3781 gacagttgcg cgggggcgcg tgtgggaagc tgctgtggag tacggcggag ctctgcgggc 3841 tgtggcggtt ggcgcgcagc gcgagtggga gtgcgtagtt ggacagcgag gcgtgtggga 3901 agctgcaaca gagtacggcg gggcttggcg gaggctggtg gttggcgcgg cgcgcgagtg 3961 ggagtgcata gtcgtgcgcg cgccgtttgg gaagctgcgt gtgtgagctg cgtgtaaact 4021 cagagagacg catggggctg acccggtttc cataagctgc ctctgagcgt acacggtctc 4081 acacggttga cagttgggct ccagcgcgtg cgggactcgc tggagtgcat gatagtcgct 4141 ggtgggcacg cgcgggcggg gcgcgggggg aacttgcgcg aaagcccgtt tgtgatcccc 4201 gtgtgtcagt taaagtgttg gtagttgcga ggattgccgc gctgcacgcc gcacgtgcaa 4261 agtacacgca tcaccactgc tgcttgccac ttcagtacct gtcacatttt gacggtcagc 4321 cggcaggttc aagatgaagg ccgagggctt gattgaagtg cagacggcgt gcaccgacac 4381 cagaagacag actttgtccc aaccgccccc gaatcccgcg cctggagctg gctcacgcac 4441 ccgaattcgg gctaggcgca cccacaggcg cccatgggcg acatggtcgc ggctgcaggg 4501 caaggcggct gcgggagtgc ggggtcatgg cggcggcgcg ggtagccttg aagtgtttgt 4561 ctgaagacgg cacagagcga ggctgggatg agttggccac ctgtgacatg aatcgattcg 4621 cctgttccgc ctgatgctgt tgctgatggt gagcgagagc atccaagctt ttcccggttt 4681 gtcgtcacca ccaaaccaca tgggcgggct gggcgtattt gaagcgggta ccccataact 4741 tcgtataatg tatgctatac gaagttatgg taccgcggcc gcgtagagga tctgttgatc 4801 agcagttcaa cctgttgata gtacgtacta agctctcatg tttcacgtac taagctctca 4861 tgtttaacgt actaagctct catgtttaac gaactaaacc ctcatggcta acgtactaag 4921 ctctcatggc taacgtacta agctctcatg tttcacgtac taagctctca tgtttgaaca 4981 ataaaattaa tataaatcag caacttaaat agcctctaag gttttaagtt ttataagaaa 5041 aaaaagaata tataaggctt ttaaagcttt taaggtttaa cggttgtgga caacaagcca 5101 gggatgtaac gcactgagaa gcccttagag cctctcaaag caattttgag tgacacagga 5161 acacttaacg gctgacatgg gaattagctt cacgctgccg caagcactca gggcgcaagg 5221 gctgctaaag gaagcggaac acgtagaaag ccagtccgca gaaacggtgc tgaccccgga 5281 tgaatgtcag ctactgggct atctggacaa gggaaaacgc aagcgcaaag agaaagcagg 5341 tagcttgcag tgggcttaca tggcgatagc tagactgggc ggttttatgg acagcaagcg 5401 aaccggaatt gccagctggg gcgccctctg gtaaggttgg gaagccctgc aaagtaaact 5461 ggatggcttt cttgccgcca aggatctgat ggcgcagggg atcaagatct gatcaagaga 5521 caggatgagg atcgtttcgc atgattgaac aagatggatt gcacgcaggt tctccggccg 5581 cttgggtgga gaggctattc ggctatgact gggcacaaca gacaatcggc tgctctgatg 5641 ccgccgtgtt ccggctgtca gcgcaggggc gcccggttct ttttgtcaag accgacctgt 5701 ccggtgccct gaatgaactg caggacgagg cagcgcggct atcgtggctg gccacgacgg 5761 gcgttccttg cgcagctgtg ctcgacgttg tcactgaagc gggaagggac tggctgctat 5821 tgggcgaagt gccggggcag gatctcctgt catctcacct tgctcctgcc gagaaagtat 5881 ccatcatggc tgatgcaatg cggcggctgc atacgcttga tccggctacc tgcccattcg 5941 accaccaagc gaaacatcgc atcgagcgag cacgtactcg gatggaagcc ggtcttgtcg 6001 atcaggatga tctggacgaa gagcatcagg ggctcgcgcc agccgaactg ttcgccaggc 6061 tcaaggcgcg catgcccgac ggcgaggatc tcgtcgtgac acatggcgat gcctgcttgc 6121 cgaatatcat ggtggaaaat ggccgctttt ctggattcat cgactgtggc cggctgggtg 6181 tggcggaccg ctatcaggac atagcgttgg ctacccgtga tattgctgaa gagcttggcg 6241 gcgaatgggc tgaccgcttc ctcgtgcttt acggtatcgc cgctcccgat tcgcagcgca 6301 tcgccttcta tcgccttctt gacgagttct tctgagcggg actctggggt tcgaaatgac 6361 cgaccaagcg acgcccaacc tgccatcacg agatttcgat tccaccgccg ccttctatga 6421 aaggttgggc ttcggaatcg ttttccggga cgccggctgg atgatcctcc agcgcgggga 6481 tctcatgctg gagttcttcg cccaccccgg gatatccgga tatagttcct cctttcagca 6541 aaaaacccct caagacccgt ttagaggccc caaggggtta tgctagttat tgctcagcgg 6601 tggcagcagc caactcagct tcctttcggg ctttgttagc agccggatct tctagaatcc 6661 ccagcatgcc tgctattgtc ttcccaatcc tcccccttgc tgtcctgccc caccccaccc 6721 cccagaatag aatgacacct actcagacaa tgcgatgcaa tttcctcatt ttattaggaa 6781 aggacagtgg gagtggcacc ttccagggtc aaggaaggca cgggggaggg gcaaacaaca 6841 gatggctggc aactagaagg cacagtcgag gctgatagcg 124
Claims
PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web CLAIMS What is claimed is:
1. A DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell, the DNA construct comprising a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof, wherein the centromere nucleating sequence is capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell.
2. The DNA construct of claim 1, wherein the centromere nucleating sequence comprises a single Long Interspersed Nuclear Element (LINE)-like retrotransposon or derivatives or truncated variants thereof.
3. The DNA construct of claim 1 or claim 2, wherein the centromere nucleating sequence comprises a single Zepp, LINE-1, Tad, CRE, Deceiver, Inkcap-like, Ty5, TRAS1, SART1, GilM, GilT, or Zorro element or derivatives or truncated variants thereof.
4. The DNA construct of any one of the preceding claims, wherein the centromere nucleating sequence comprises a single ZeppL-LINE1 element or derivatives or truncated variants thereof.
5. The DNA construct of any one of the preceding claims, wherein the centromere nucleating sequence is derived from the ZeppL element of chromosome 2b of Chlamydomonas reinhardtii strain UL1690.
6. The DNA construct of claim 4 or claim 5, wherein the ZeppL element comprises a nucleic acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity of SEQ ID NO:
5.
7. The DNA construct of any one of the preceding claims, further comprising chromatin components assembled into chromatin.
8. The DNA construct of any one of the preceding claims, further comprising chromatin components assembled into centromeric chromatin. 125PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 9. The DNA construct of any one of the preceding claims, further comprising centromeric chromatin comprising a centromere-specific histone H3 (CENH3) protein.
10. The DNA construct of claim 9, wherein the CENH3 protein a CENH3.1 paralog, a CENH3.2 paralog of Chlamydomonas reinhardtii CENH3, or both.
11. The DNA construct of claim 10, wherein the CENH3 protein comprises an amino acid sequence comprising at least about 75% or more, at least about 85% or more, at least about 95% or more, or 100% sequence identity with a nucleic acid sequence of SEQ ID NO: 2 or SEQ ID NO:
4.
12. The DNA construct of any one of the preceding claims, wherein the DNA construct is an artificial chromosome, a mini-chromosome, or a plasmid.
13. The DNA construct of any one of the preceding claims, wherein the DNA construct is an artificial linear chromosome, an artificial circular chromosome, or a mini-chromosome.
14. The DNA construct of any one of the preceding claims, wherein the DNA construct further comprises one or more episomal element components selected from telomeres, replication origin, autonomous replicating sequence (ARS), selectable markers, cloning sites, polynucleotides of interest, or any combination thereof.
15. The DNA construct of any one of the preceding claims, further comprising a polynucleotide of interest.
16. The DNA construct of claim 15, wherein the polynucleotides of interest is an expression construct, a selectable marker, a cloning site, or any combination thereof.
17. The DNA construct of any one of the preceding claims, wherein the eukaryotic cell is a chlorophyte green algae.
18. The DNA construct of claim 17, wherein the chlorophyte green algae is a Chlamydomonas sp.
19. The DNA construct of claim 18, wherein the Chlamydomonas sp. is Chlamydomonas reinhardtii. 126PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 20. The DNA construct of claim 19, wherein the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690.
21. An artificial chromosome, the artificial chromosome comprising a. a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof; and b. one or more chromatin components assembled into chromatin on the centromere nucleating sequence as a nucleosome; wherein the centromere nucleating sequence is capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell.
22. The artificial chromosome of claim 21, wherein the one or more components of a nucleosome comprises a CENH3 protein.
23. A eukaryotic cell comprising the DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell of any one of claims 1-20.
24. The eukaryotic cell of claim 23, wherein the eukaryotic cell is a chlorophyte green algae.
25. The eukaryotic cell of claim 24, wherein the chlorophyte green algae is a Chlamydomonas sp.
26. The eukaryotic cell of claim 25, wherein the Chlamydomonas sp. is Chlamydomonas reinhardtii.
27. The eukaryotic cell of claim 26, wherein the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690. 127PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 28. A cell culture comprising a population of eukaryotic cells according to any one of claims 23 to 27, wherein each eukaryotic cell comprises a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell of any one of claims 1-20.
29. The cell culture of claim 28, wherein the eukaryotic cell is a chlorophyte green algae.
30. The cell culture of claim 29, wherein the chlorophyte green algae is a Chlamydomonas sp.
31. The cell culture of claim 30, wherein the Chlamydomonas sp. is Chlamydomonas reinhardtii.
32. The cell culture of claim 31, wherein the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690.
33. A method for generating an artificial chromosome, the method comprising: a. providing or having provided a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell, the DNA construct comprising: i. a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof, wherein the centromere nucleating sequence is capable of supporting centromere formation and faithful transmission of the DNA construct during mitotic and meiotic division of a eukaryotic cell; and ii. one or more episomal element components selected from telomeres, replication origin, autonomous replicating sequence (ARS), selectable markers, cloning sites, polynucleotides of interest, or any combination thereof; b. introducing the DNA construct into a eukaryotic cell; and c. allowing sufficient time and providing conditions sufficient for centromeric chromatin to assemble at the centromere nucleating sequence. 128PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 34. A method of introducing a polynucleotide of interest into a eukaryotic cell for faithful transmission of the polynucleotide of interest during mitotic and meiotic division of a eukaryotic cell, the method comprising: a. providing or having provided a DNA construct capable of stable mitotic and meiotic segregation in a eukaryotic cell of claims 1-20, the DNA construct comprising a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof and a polynucleotide of interest, wherein the centromere nucleating sequence is capable of promoting centromere formation and supporting stable chromosomal segregation; and b. introducing the DNA construct into a eukaryotic cell under conditions sufficient to allow centromeric chromatin to assemble at the centromere nucleating sequence, thereby enabling faithful transmission of the DNA construct during mitotic and meiotic division.
35. A kit for introducing a polynucleotide of interest into a eukaryotic cell for faithful transmission of the polynucleotide of interest during mitotic and meiotic division of a eukaryotic cell, the kit comprising: a. one or more DNA constructs of claims 1-20 capable of stable mitotic and meiotic segregation in a eukaryotic cell, the DNA construct comprising a centromere nucleating sequence comprising a single retrotransposon element or derivatives or truncated variants thereof; b. a eukaryotic optionally further comprising one or more DNA constructs of (a); or c. (a) and (b).
36. The kit of claim 35, wherein the eukaryotic cell is a chlorophyte green algae.
37. The kit of claim 36, wherein the chlorophyte green algae is a Chlamydomonas sp. 129PROV PATENT Attorney Docket No.: DDPSC0166-401-PCT Via EFS-web 38.
37. The kit of claim 37, wherein the Chlamydomonas sp. is Chlamydomonas reinhardtii. 39.
37. The kit of claim 38, wherein the Chlamydomonas sp. is Chlamydomonas reinhardtii strain UL1690. 130
Citation Information
Patent Citations
Artificial plant minichromosomes
US20070271629A1
Novel centromeres and methods of using the same
US20130007927A1