Human artificial chromosomes
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- CENTAURA AG
- Filing Date
- 2025-09-10
- Publication Date
- 2026-04-23
AI Technical Summary
Existing methods for creating artificial chromosomes in mammals face challenges due to the instability of repetitive α-satellite DNA sequences, difficulty in synthesis, assembly, and characterization, and require epigenetic assembly with kinetochore proteins, leading to issues like size limitations, stability, and random integration into the genome.
Design and synthesis of artificial centromeres comprising non-identical repetitive units with recruitment motifs for epigenetic modifiers, allowing stable partitioning and maintenance of artificial chromosomes during mitosis, using nucleic acids between 5 to 50 kb with specific sequence identities and recruitment motifs like LacO sites.
Ensures stable expression and passage of genetic material to daughter cells over multiple divisions, overcoming previous instability issues and enabling efficient cargo sequence expression in mammalian cells.
Smart Images

Figure IB2025059106_23042026_PF_FP_ABST
Abstract
Description
Attorney Docket No.: CETA-001WO HUMAN ARTIFICIAL CHROMOSOMES CROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of and priority to U.S. Provisional Patent Application No.63 / 692,914, filed September 10, 2024, the disclosure of which is incorporated by reference herein for all purposes. REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY
[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML format and is hereby incorporated by reference in its entirety. Said XML copy, created on September 10, 2025, is named “CETA-001WO_SL.xml” and is 948,883 bytes in size. FIELD OF THE INVENTION
[0003] The disclosure relates generally to artificial centromeres and artificial chromosomes comprising the artificial centromeres and a cargo sequence, methods of designing and making such artificial centromeres and artificial chromosomes, cells comprising such artificial centromeres and artificial chromosomes, and methods of using such artificial centromeres and artificial chromosomes for stable passaging of artificial chromosomes and expression of cargo sequences in a mammalian cell. BACKGROUND
[0004] Centromeres are essential components of all eukaryotic chromosomes. The human centromere is a structure in a human chromosome that facilitates chromosome organization during cell division. The centromere is a DNA locus present in every chromosome that ensures equal partitioning of chromatids between the two daughter cells at mitosis. During mitosis, the centromere directs both kinetochore formation and protects cohesion of sister chromatids until their synchronous separation at the onset of anaphase. The kinetochore is the proteinaceous complex that binds to the spindle microtubules that drive chromosome alignment and subsequent segregation to the daughter cells. Thus, the centromere allows the cellular chromosomes to be faithfully divided during cell division. 1 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0005] In humans, centromeres are typically located on repetitive α-satellite DNA. The sequence within individual α-satellite repeat elements is similar but not identical. In a centromere, 171bp repetitive units (termed α-satellite monomers) are organized into a chromosome specific arrangement of higher order repeats (HORs) that are repeated up to thousands of times. α-satellite arrays have been found to be in a range between 1-4Mb depending on the chromosome and the individual. The centromere is further complexed with proteins that help form the three dimensional structure and serve as anchor points for other proteins involved in chromosome maintenance and division. The complex comprising the centromere and the centromere binding proteins is termed kinetochore. The location of the centromere on the chromosome in humans is defined by an essential epigenetic mark comprising nucleosomes containing the kinetochore protein histone H3 variant, CENP-A (Centromere Protein A). The kinetochore proteins associate only with a subset of these HOR repeats, which is also called the active array. Other α-satellite arrays in the centromere not in complex with kinetochore proteins are called inactive arrays.
[0006] Artificial chromosomes and artificial centromeres have been described in the literature (see e.g., Logsdon, Glennis A et al. (2019) “Human Artificial Chromosomes that Bypass Centromeric DNA.” CELL 178(3): 624-639, e19. doi:10.1016 / j.cell.2019.06.006 and Ponomartsev et al. (2022) “Human Artificial Chromosomes and Their Transfer to Target Cells. ACTA NATURAE 14(3):35-45. doi: 10.32607 / actanaturae.11670. PMID: 36348716; PMCID: PMC9611860, both of which are incorporated herein by reference in their entirety). The minimal components required for human artificial chromosome (HAC) propagation and inheritance through cell divisions alongside their natural chromosome counterparts are telomeres, an origin of replication, and a centromere. Telomeres are not necessary if the HAC is a circular HAC and a number of DNA sites can serve as a origin of replication. Thus, a functional centromere is the only requirement for HAC function. Prior attempts at creating functional HACs typically used sequences of the repetitive α-satellite DNA organized into HORs. However, repetitive sequences as found in the α-satellite arrays and HORs present major challenges for the design of synthetic mammalian chromosomes because they are difficult to synthesize, assemble, and characterize. Additionally, in most eukaryotes, centromeres are defined epigenetically and require assembly with kinetochore proteins to be fully functional.
[0007] Stable expression of (exogenous) proteins or RNA in mammalian cells to date usually relies on exogenous DNA or RNA that is transduced into the cell. Numerous delivery methods and agents for delivery are known, for example viral delivery of DNA or RNA, 2 IPTS / 200118216.1Attorney Docket No.: CETA-001WO delivery of plasmids, delivery of mRNA, bacmids, minicircles etc. However, these methods have several drawbacks, such as size limitations, stability issues, random integration into the mammalian genome, epigenetic silencing, and loss of expression after several cell cycles. SUMMARY OF THE DISCLOSURE
[0008] The disclosure relates in part to the discovery of artificial centromeres that are capable of being synthesized and manipulated in microorganisms, e.g., bacteria, without the instability that has hindered prior attempts to design and produce artificial centromeres. The disclosure also provides methods of designing and producing such artificial centromeres.
[0009] The disclosure also relates, in part, to artificial chromosomes, e.g., human artificial chromosomes (HACs), that include an artificial centromere as described herein. The sequence of the artificial centromere allows for the formation of centromeric structure necessary for even partitioning of the artificial chromosome during mitosis, thereby ensuring that the genetic material contained within the artificial chromosome is passed to daughter cells and stably maintained over many cell divisions.
[0010] Accordingly, in one aspect, the disclosure provides an artificial centromere. The artificial centromere comprises a nucleic acid of about 5 to about 50 kb, wherein the nucleic acid comprises (a) alpha satellite DNA comprising a plurality of non-identical repetitive units having between about 50% and about 99.9% sequence identity to each other organized into higher order repeats; and (b) a plurality of recruitment motifs for an epigenetic modifier, wherein a recruitment motif is present at least once per 250 bp of the nucleic acid.
[0011] In some embodiments, the nucleic acid is between about 5 and about 15 kb. In some embodiments, the nucleic acid is about 10 kb.
[0012] In some embodiments, the nucleic acid has between about 30% and about 99.9% sequence identity to a human centromere. In some embodiments, the nucleic acid has between about 30% and about 60% sequence identity to a human centromere. In some embodiments, the sequence of the reference human centromere is from the reference genome CHM13.
[0013] In some embodiments, the recruitment motif is a Laco site. In some embodiments, the LacO site comprises the nucleic acid sequence TGTGANCGNTCACA, where N is any nucleotide (SEQ ID NO: 1), or comprises a nucleic acid sequence having 1, 2, or 3 nucleotide substitutions as compared to SEQ ID NO: 1. 3 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0014] In some embodiments, the binding motif is present at least once per 100 to 200 bp of the nucleic acid.
[0015] In some embodiments, each repetitive unit of the artificial centromere is about 170 bp in length.
[0016] In some embodiments, the higher order repeats of the artificial centromere comprise about 4 to about 12 repetitive units. In some embodiments, the higher-order repeat is between about 500 bp and about 2500 bp in length. In some embodiments, the higher order repeat is between about 700 bp and about 2000 bp in length. In some embodiments, the artificial centromere comprises from about 5 to about 15 higher-order repeats. In some embodiments, the artificial centromere comprises about 7 higher-order repeats.
[0017] In another aspect, an artificial chromosome is provided comprising an artificial centromere as described herein. In some embodiments, the artificial chromosome further comprises a gene (e.g., a series of genes) or other DNA elements for expression in a cell. In some embodiments, the gene (e.g., a series of genes) for expression in a cell are flanked by one or more insulator sequences. In some embodiments, the gene (e.g., a series of genes) are flanked by a first insulator sequence located 5’ of the gene or series of genes and a second insulator sequence located 3’ of the gene or series of genes. In some embodiments, the first and the second insulator sequence comprise the same nucleic acid sequence. In some embodiments, the first and the second insulator sequence comprise unique nucleic acid sequences. In some embodiments, the first and the second insulator sequence are each, independently, selected from a nucleic acid sequence having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% identity) to a nucleic acid sequence of SEQ ID NO: 31 or SEQ ID NO: 32.
[0018] In another aspect, a method of making an artificial centromere or an artificial chromosome is provided. In one aspect the method comprises synthesizing the artificial centromere or the artificial chromosome as provided for herein. In some embodiments, a method of making an artificial centromere is provided, the method comprising synthesizing the artificial centromere and / or chromosome as described herein.
[0019] In another aspect, a method of using an artificial chromosome is provided, comprising introducing an artificial chromosome into a cell. In some embodiments, the artificial chromosome comprises a gene or set of DNA elements for expression in the cell. In some embodiments, the method comprises expressing the gene or the set of DNA elements in the cell. 4 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0020] In another aspect, a cell bearing an artificial chromosome is provided. In one aspect, the cell comprises an artificial chromosome as described herein.
[0021] In another aspect, a method for designing an artificial centromeric sequence is provided. The method comprises (a) identifying sequences from active alpha satellite arrays of human centromeric DNA having repetitive units and / or higher-order repeats; (b) creating sequence alignments of the repetitive units and / or higher order repeats; (c) creating a position weight matrix (PWM) from the sequences alignments; (d) modifying the PWM to insert recruitment motifs for epigenetic modifiers; and (e) designing an artificial centromeric sequence of about 5 to about 50 kb comprising a plurality of higher-order repeats, wherein each higher-order repeat comprises a plurality of non-identical repetitive units having between about 50% and about 99.9% sequence identity to each other.
[0022] In some embodiments, the artificial centromeric sequence is between 5 kb and 15 kb. In some embodiments, the artificial centromeric sequence is about 10 kb.
[0023] In some embodiments, the artificial centromeric sequence has between about 30% and about 99.9% sequence identity to a human centromere. In some embodiments, the artificial centromeric sequence has between about 30% and about 60% identity to a human centromere. In some embodiments, the sequences from the active alpha satellite arrays of the human centromeric DNA are from a reference genome, wherein the reference genome is CHM13.
[0024] In some embodiments, the recruitment motifs comprise LacO sites.
[0025] In some embodiments, at least one binding motif is present per 100 bp to 200 bp of the artificial centromeric sequence.
[0026] In some embodiments, each repetitive unit is between about 160 and about 180 bp in length. In some embodiments, each repetitive unit is about 170bp in length.
[0027] In some embodiments, the plurality of non-identical repetitive units is from about 4 to about 12 units. In some embodiments, the higher-order repeat is between about 500 bp and about 2500 bp in length. In some embodiments, the higher-order repeat is between about 700 bp and about 2000 bp in length.
[0028] In some embodiments, the method further comprises analyzing the artificial centromeric sequence for stability or synthesizability. In some embodiments, if the artificial centromeric sequence does not pass a threshold measurement for stability or synthesizability, additional changes are made to improve the stability or synthesizability of the artificial centromeric sequence. 5 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0029] In some embodiments, the method further comprises generating a nucleic acid comprising the artificial centromeric sequence.
[0030] In another aspect, a gene expression system comprising a cargo nucleic acid sequence and one or more insulator sequences is provided. In some embodiments, the one or more insulator sequences flank the cargo nucleic acid sequence. In some embodiments, the one or more insulator sequences comprise a first insulator sequence and a second insulator sequence. In some embodiments, the first and the second insulator sequence comprise the same nucleic acid sequence. In some embodiments, the first and the second insulator sequence comprise unique nucleic acid sequences. In some embodiments, the first and the second insulator sequence are each, independently, selected from a nucleic acid sequence having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% identity) to a nucleic acid sequence of SEQ ID NO: 31 or SEQ ID NO: 32. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The disclosure is more completely understood with reference to the following drawings.
[0032] FIG.1 is a schematic diagram showing the components of a human artificial centromere.
[0033] FIG.2 is a schematic diagram showing the basic structure of a human chromosome and human centromere and showing the selection of active repeat sequences for the generation of artificial centromeres.
[0034] FIG.3 is a schematic diagram showing (1) the addition of epigenetic modifier binding sites to an artificial chromosome and (2) the addition of natural sequence variations to alpha satellite arrays to reduce sequence similarity between the repetitive units and HORs.
[0035] FIG.4A is a schematic diagram showing the synthesis of the HORs and a centromere of the disclosure. FIG.4B is a schematic diagram showing an assembled artificial chromosome of the disclosure, comprising a centromere, backbone and cargo.
[0036] FIG.5 illustrates the cloning strategy used to construct the HACs provided for herein.
[0037] FIG.6 illustrates the sequence mismatch rate observed during cloning in E. coli.
[0038] FIG.7 illustrates GFP expression in HFF, HT-1080, and HEK293 cell lines 48 hours after transfection with the indicated HAC construct. 6 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0039] FIG.8 illustrates the results of a flow cytometry analysis of the GFP content of HFF cells transfected with the indicated HAC constructs one month post transfection.
[0040] FIG.9 illustrates the results of qPCR analysis of portions of the HAC DNA in established positive clones vs control.
[0041] FIG.10 illustrates the results of Southern blot analysis of select HAC positive clones.
[0042] FIG.11 illustrates the results of FISH analysis of select HAC positive clones.
[0043] FIGs.12A and B illustrate the result of rolling circle amplification (RCA) analysis of select HAC positive clones. FIG.12A illustrates the detection of HAC DNA specific restriction digest banding patterns in select HAC positive clones vs HAC negative control. FIG.12B illustrates the sensitivity of RCA to detect HAC positive clones.
[0044] FIGs.13A and B illustrates results of nanopore sequencing of select HAC positive clones. FIG.13A illustrates the distribution of read length. FIG.13B illustrates mapping of sequencing reads as a function of HAC coordinates vs read coordinates.
[0045] FIGs.14A and B illustrate the results of a nanopore sequencing analysis of HAC formation. FIG.14A provides representative illustrations of nanopore sequencing reads for intact, rearranged (HAC only), rearranged (with genome), and integrated readouts. FIG.14B illustrates the results of the sequencing analysis. The plot at left illustrates the fraction of intact, rearranged, or integrated reads for each HAC centromere. The plot at right provides the total number of reads for each condition.
[0046] FIGs.15A, B, and C illustrate the results of a cargo expression analysis for select chromosome 12 k=50 repeat clones. FIG.15A illustrates the total GFP expression (left) and percent GFP positive cells (right) for each condition over the course of the experiment. FIG. 15B illustrates the results of a qPCR analysis of total copies of HAC gene per cell for each clone. FIG.15C illustrates the results of a nanopore sequencing analysis of the select clones. DETAILED DESCRIPTION
[0047] The disclosure relates in part to the discovery that an artificial centromere can be designed comprising repetitive DNA sequences that are capable of being artificially synthesized and manipulated in microorganisms, e.g., bacteria, without the instability that has hindered prior attempts to design artificial centromeres. The disclosure also provides methods of designing such artificial centromeres.
[0048] The disclosure also relates, in part, to artificial chromosomes that include an artificial centromere as described herein. The disclosed artificial centromeres comprise the 7 IPTS / 200118216.1Attorney Docket No.: CETA-001WO components needed for even partitioning of the artificial chromosome during mitosis, thereby ensuring that the genetic material contained within the artificial chromosome is passed to daughter cells and stably maintained over many cell divisions.
[0049] Described herein, in certain embodiments, are artificial centromeres and artificial chromosomes comprising the artificial centromeres and a cargo sequence, methods of designing and making such artificial centromeres and artificial chromosomes, cells comprising such artificial centromeres and artificial chromosomes, and methods of using such artificial centromeres and artificial chromosomes for stable passaging of artificial chromosomes and expression of cargo sequences in a mammalian cell.
[0050] While various embodiments of the disclosure have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the disclosure. It should be understood that various alternatives to the embodiments of the disclosure described herein may be employed. I. Definitions
[0051] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as is commonly understood by one of skill in the art to which the claimed subject matter belongs. Generally, nomenclatures utilized in connection with, and techniques of, immunology, oncology, cell and tissue culture, molecular biology, and protein and oligonucleotide or polynucleotide chemistry and hybridization described herein are those well-known and commonly used in the art. Units of measure not otherwise defined accord with The International System of Units (SI), NIST Special Publication 330, 2019 edition.
[0052] As used herein, all numerical values or numerical ranges comprise whole integers within or encompassing such ranges and fractions of the values or the integers within or encompassing ranges unless the context clearly indicates otherwise. Thus, for example, reference to a range of 90-100%, comprises 91%, 92%, 93%, 94%, 95%, 95%, 97%, etc., as well as 91.1%, 91.2%, 91.3%, 91.4%, 91.5%, etc., 92.1%, 92.2%, 92.3%, 92.4%, 92.5%, etc., and so forth. In another example, reference to a range of 1-5,000-fold comprises 1-, 2-, 3-, 4- , 5-, 6-, 7-, 8-, 9-, 10-, 11-, 12-, 13-, 14-, 15-, 16-, 17-, 18-, 19-, or 20-fold, etc., as well as 1.1-, 1.2-, 1.3-, 1.4-, or 1.5-fold, etc., 2.1-, 2.2-, 2.3-, 2.4-, or 2.5-fold, etc., and so forth.
[0053] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of any embodiment. As used herein, the singular forms “a,” “an,” and “the” are intended to comprise the plural forms as well, unless 8 IPTS / 200118216.1Attorney Docket No.: CETA-001WO the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. As used herein, the term “and / or” comprises any and all combinations of one or more of the associated listed items.
[0054] Unless specifically stated or obvious from context, as used herein, the term “about” in reference to a number or range of numbers is understood to mean the stated number and numbers ±10% thereof, or 10% below the lower listed limit and 10% above the higher listed limit for the values listed for a range.
[0055] “Percent identity,” “% identity,” or “sequence identity” refer to the extent to which two sequences (nucleotide or amino acid) have the same residues at the same positions in an alignment. For example, “a nucleotide sequence is X% identical to SEQ ID NO: Y” refers to % identity of the nucleotide sequence to SEQ ID NO: Y and is elaborated as X% of residues in the nucleotide sequence are identical to the corresponding residues of sequence disclosed in SEQ ID NO: Y. A sequence said to be X% identical to a reference sequence may contain more nucleotide or amino acid residues than specified in the reference sequence but must contain a sequence corresponding to the reference sequence. In most cases, the sequence in question will contain a sequence that corresponds to all of the specified reference sequences. Generally, computer programs are employed for such calculations. Exemplary programs that compare and align pairs of sequences, comprise ALIGN, FASTA, gapped BLAST, BLASTP, BLASTN, or GCG.
[0056] The term “chromosome” refers to a nucleic acid encoding part or all of the genetic material of an organism. Eukaryotic chromosome are associated with proteins, forming a compact complex of proteins and DNA called chromatin. Eukaryotic chromosomes comprise centromeres that links a pair of sister chromatids together during cell division. Centromeres are the loci present once per natural chromosome that guide their segregation at cell division.
[0057] The term “alpha satellite” or “alpha satellite repeats” refers to a tandemly organized type of repetitive DNA that comprises 5% of the genome and is found at all human centromeres of shared ancestral origin. A defined number of 171-bp monomers are organized into chromosome-specific higher-order repeats (HORs) that are reiterated thousands of times in the genome.
[0058] The term “plasmid” refers to an extrachromosomal element that carries genes that can replicate independently of the chromosomes of the cell. The plasmid can be in the form 9 IPTS / 200118216.1Attorney Docket No.: CETA-001WO of a circular double-stranded DNA molecule. Such elements can include autonomously replicating sequences, genome integrating sequences, phage or nucleotide sequences, as well as linear, circular or supercoiled, single-or double-stranded DNA or RNA of any origin. Exemplary plasmids include, but are not limited to, minicircles and doggybone plasmids.
[0059] “Polynucleotide,” or “nucleic acid,” are used interchangeably herein and refer to chains of nucleotides of any length, and comprise DNA and RNA. In some embodiments, the nucleotides are deoxyribonucleotides, ribonucleotides, modified nucleotides or bases, and / or their analogs, or any substrate that is incorporated into a chain by DNA or RNA polymerase. A polynucleotide may comprise modified nucleotides, such as methylated nucleotides and their analogs. If present, modification to the nucleotide structure is imparted before or after assembly of the chain. In some embodiments, the sequence of nucleotides is interrupted by non-nucleotide components. In some embodiments, a polynucleotide is further modified after polymerization, such as by conjugation with a labeling component. Other types of modifications comprise, for example, “caps,” substitution of one or more of the naturally occurring nucleotides with an analog, internucleotide modifications such as, for example, those with uncharged linkages (e.g., methyl phosphonates, phosphotriesters, phosphoamidates, carbamates) and with charged linkages (e.g., phosphorothioates, phosphorodithioates), those containing pendant moieties, such as, for example, proteins (e.g., nucleases, toxins, antibodies, signal peptides, poly-L-lysine), those with intercalators (e.g., acridine, psoralen), those containing chelators (e.g., metals, radioactive metals, boron, oxidative metals), those containing alkylators, those with modified linkages (e.g., alpha anomeric nucleic acids), as well as unmodified forms of the polynucleotide(s). In some embodiments, any of the hydroxyl groups ordinarily present in the sugars are replaced, for example, by phosphonate groups, phosphate groups, protected by standard protecting groups, or activated to prepare additional linkages to additional nucleotides, or are conjugated to solid supports. In some embodiments, the 5’ and 3’ terminal OH is phosphorylated or substituted with amines or organic capping group moieties of from 1 to 20 carbon atoms. Other hydroxyls may also be derivatized to standard protecting groups. In some embodiments, polynucleotides also contain analogous forms of ribose or deoxyribose sugars, comprising, for example, 2’-O-methyl-, 2’-O-allyl, 2’-fluoro- or 2’-azido-ribose, carbocyclic sugar analogs, alpha- or beta-anomeric sugars, epimeric sugars such as arabinose, xyloses or lyxoses, pyranose sugars, furanose sugars, sedoheptuloses, acyclic analogs and abasic nucleoside analogs such as methyl riboside. In some embodiments, one or more phosphodiester linkages are replaced by alternative linking groups. These alternative linking 10 IPTS / 200118216.1Attorney Docket No.: CETA-001WO groups comprise, but are not limited to, embodiments wherein phosphate is replaced by P(O)S(“thioate”), P(S)S (“dithioate”), (O)NRi (“amidate”), P(O)R, P(O)OR’, CO or CH2 (“formacetal”), in which each R or R’ is independently H or substituted or unsubstituted alkyl (1-20 C) optionally containing an ether (-O-) linkage, aryl, alkenyl, cycloalkyl, cycloalkenyl or araldyl. Not all linkages in a polynucleotide need be identical. The preceding description applies to all polynucleotides referred to herein, comprising RNA and DNA.
[0060] As used herein, the term “vector” comprises a nucleic acid vector, e.g., a DNA vector, such as a plasmid, an RNA vector, or another suitable replicon (e.g., viral vector). A variety of vectors have been developed for the delivery of polynucleotides encoding exogenous polynucleotides or proteins into a prokaryotic or eukaryotic cell. Expression vectors suitable for use with the compositions and methods described herein contain a polynucleotide sequence as well as, e.g., additional sequence elements used for the expression of heterologous nucleic acid materials (e.g., a nucleic acid molecule) in a cell. Certain vectors that are used for the expression of the nucleic acid molecules described herein comprise plasmids that contain regulatory sequences, such as promoter and enhancer regions, which direct gene transcription. In some embodiments, the compact bidirectional promoters do not contain an enhancer. Other useful vectors for expression of nucleic acid molecule agents disclosed herein contain polynucleotide sequences that enhance the rate of translation of these polynucleotides or improve the stability or nuclear export of the RNA that results from gene transcription. These sequence elements comprise, e.g., 5’ and 3’ untranslated regions, an internal ribosomal entry site (IRES), and polyadenylation signal (polyA) in order to direct efficient transcription of the gene carried on the expression vector. In some embodiments, the expression vectors suitable for use with the compositions and methods described herein contain a backbone polynucleotide encoding a marker for selection of cells that contain such a vector. Examples of a suitable marker are genes that encode resistance to antibiotics, such as ampicillin, chloramphenicol, neomycin, zeocin, kanamycin, nourseothricin, aminoglycoside, a beta-lactam, a glycopeptide, a macrolide, a polypeptide, a tetracycline, spectinomycin, streptomycin, carbenicillin, bleomycin, erythromycin, polymyxin B, chloramphenicol, or a derivative thereof. II. Artificial Centromeres
[0061] The artificial centromeres of the present disclosure are useful in creating artificial chromosomes that can express a genetic cargo (e.g., encoding one or more therapeutic 11 IPTS / 200118216.1Attorney Docket No.: CETA-001WO proteins or RNA) and are stably maintained in mammalian cells. The artificial centromeres can be synthesized and combined with other nucleic acids that form portions of the backbone and the genetic cargo to form an artificial chromosome. Such artificial chromosomes can be used in vitro, for example, to express a protein for purification, or in vivo, for example, as a therapeutic protein.
[0062] FIG.1 provides an exemplary artificial centromere of the disclosure. The artificial centromere comprises a 4-12 non-identical repetitive units that form higher order repeats (HORs). The non-identical repetitive units comprise about 170bp long sequences (alpha satellite monomers). Interspersed in the non-identical repetitive units in the HORs are epigenetic modifier binding or recruitment sites that recruit kinetochore proteins such as CENP-A when maintained in human cells. Between about 5 and 10 HORs are repeated in head to tail fashion to form the centromere region. The HORs can be repeated in a regular repeat, for example, a HOR containing non-identical alpha satellites A, B, and C form a HOR of ABC and the HOR can be repeated in a regular repeat forming an array of ABC-ABC- ABC.
[0063] The disclosure provides an artificial centromere comprising a nucleic acid of about 5 kb to about 50 kb, wherein the nucleic acid comprises alpha satellite DNA comprising a plurality of non-identical repetitive units having between about 50% and about 99.9% sequence identity to each other organized into higher order repeats. Because human centromeres are defined epigenetically, the artificial chromosomes described herein comprise a plurality of recruitment motifs for an epigenetic modifier, wherein a recruitment motif is present at least once per 250 bp of the nucleic acid. The epigenetic modifier allows for the deposition of nucleosomes comprising the histone H3 variant, CENP-A (Centromere Protein A). In some embodiments, the epigenetic modifier is any appropriate epigenetic modifier. In some embodiments, the epigenetic modifier is Holiday Junction Recognition Protein (HJURP), a histone chaperone protein that is specific for a CENP-A histone.
[0064] In some embodiments, the artificial centromere is about 5 kb to about 50 kb in size. For example, the artificial centromere can be from about 5 kb to about 10 kb, about 5 kb to about 15 kb, about 5 kb to about 20 kb, about 5 kb to about 25 kb, about 5 kb to about 30 kb, about 5 kb to about 35 kb, about 5 kb to about 40 kb, about 5 kb to about 45 kb, about 10 kb to about 15 kb, about 10 kb to about 20 kb, about 10 kb to about 25 kb, about 10 kb to about 30 kb, about 10 kb to about 35 kb, about 10 kb to about 40 kb, about 10 kb to about 45 kb, about 10 kb to about 50 kb, about 15 kb to about 20 kb, about 15 kb to about 25 kb, about 15 kb to about 30 kb, about 15 kb to about 35 kb, about 15 kb to about 40 kb, about 15 kb to 12 IPTS / 200118216.1Attorney Docket No.: CETA-001WO about 45 kb, about 15 kb to about 50 kb, about 20 kb to about 25 kb, about 20 kb to about 30 kb, about 20 kb to about 35 kb, about 20 kb to about 40 kb, about 20 kb to about 45 kb, about 20 kb to about 50 kb, about 25 kb to about 30 kb, about 25 kb to about 35 kb, about 25 kb to about 40 kb, about 25 kb to about 45 kb, about 25 kb to about 50 kb, about 30 kb to about 40 kb, about 30 kb to about 45 kb, about 30 kb to about 50 kb, about 35 kb to about 40 kb, about 35 kb to about 45 kb, about 35 kb to about 50 kb, about 40 kb to about 45 kb, about 40 kb to about 50 kb, or about 45 kb to about 50 kb in size. In some embodiments, the artificial centromere is about 5 kb in size. In some embodiments, the artificial centromere is about 10 kb in size. In some embodiments, the artificial centromere is about 15 kb in size. In some embodiments, the artificial centromere is about 20 kb in size. In some embodiments, the artificial centromere is about 25 kb in size. In some embodiments, the artificial centromere is about 30 kb in size. In some embodiments, the artificial centromere is about 35 kb in size. In some embodiments, the artificial centromere is about 40 kb in size. In some embodiments, the artificial centromere is about 45 kb in size. In some embodiments, the artificial centromere is about 50 kb in size.
[0065] In some embodiments, the nucleic acid forming the artificial centromere has between about 30% and about 99.9% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 30% and about 40% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 30% and about 50% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 30% and about 60% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 30% and about 70% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 30% and about 80% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 30% and about 90% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 40% and about 50% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 40% and about 60% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 40% and about 70% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 40% and about 80% sequence identity to a human centromere. 13 IPTS / 200118216.1Attorney Docket No.: CETA-001WO In some embodiments, the nucleic acid forming the artificial centromere has between about 40% and about 90% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 50% and about 60% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 50% and about 70% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 50% and about 80% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 50% and about 90% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 60% and about 70% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 60% and about 80% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 60% and about 90% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 70% and about 80% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 70% and about 90% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has between about 80% and about 90% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has about 30% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has about 40% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has about 50% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has about 60% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has about 70% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has about 80% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has about 90% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has about 95% sequence identity to a human centromere. In some embodiments, the nucleic acid forming the artificial centromere has about 99% sequence identity to a human centromere.
[0066] In some embodiments, the sequence of the human centromere is from the reference genome CHM13. For example, in some embodiments, the reference genome is 14 IPTS / 200118216.1Attorney Docket No.: CETA-001WO from the genome associated with T2T-CHM13v1.1 (as listed on www.ncbi.nlm.nih.gov / assembly / GCA_009914755.3). In some embodiments, the reference genome is from the genome associated with T2T-CHM13v2.0 (as listed on www.ncbi.nlm.nih.gov / assembly / GCA_009914755.4). A. Recruitment Motifs and Epigenetic Modifiers
[0067] In some aspects, the artificial centromere comprises a recruitment motif for recruiting an epigenetic modifier to the centromere. Epigenetic modifiers promote the formation of the kinetochore, the complex of centromere and histones that characterize a functional mammalian centromere. In some embodiments, the epigenetic modifier allows for the formation of nucleosomes comprising the histone H3 variant, CENP-A (Centromere Protein A). Epigenetic modifier recruitment motifs are short DNA sequences recognized by specific proteins which are (e.g., 10-30bp long) bound by said cognate protein with high specificity. The cognate protein includes an epigenetic modifier or a fusion protein comprising an epigenetic modifier. The epigenetic modifier can be supplied as a fusion protein comprising a protein that binds to the recruitment motifs and an epigenetic modifier protein. When the fusion protein binds to the recruitment motif, the fused epigenetic modifier protein is automatically recruited to the site of the recruitment motif. In one example, the epigenetic modifier protein recruits CENP-A. Not wishing to be bound by theory, it is believed that a high local density of CENP-A on an artificial centromere recruited by such binding sites initiates a centromere epigenetic loop so that the chromatin can be propagated by the natural pathway in the cell even after the initial seeding CENP-A protein is removed. i. Recruitment Motifs
[0068] Exemplary recruitment motifs for use in the artificial centromeres described herein include a lacO site, a tetO site, a lexO site, or a UAS site. The lacO site is the E. coli lac operator sequence that is specifically bound by the DNA-binding domain of the E. coli lac repressor (LacI). Thus, in some embodiments, the lacO site is used to recruit a protein (e.g., a fusion protein) comprising a LacI protein and an epigenetic modifier, for example a CENP- A. The tetO site is the tetO operator sequence that binds the tetracycline transactivator (tTA) protein. Thus, in some embodiments, the tetO site is used to recruit a protein (e.g., a fusion protein) comprising a tTA protein and an epigenetic modifier. In some embodiments, the recruitment of the protein (e.g., a fusion protein) comprising a tTA protein and an epigenetic 15 IPTS / 200118216.1Attorney Docket No.: CETA-001WO modifier can be controlled via the application of tetracycline or its derivatives. The lexO site is the lexA operator that binds to the LexA transcription factor. Thus, in some embodiments, the lexO motif is used to recruit a protein (e.g., a fusion protein) comprising a LexA transcription factor and an epigenetic modifier. The UAS site is the Upstream Activation Sequence that binds Gal4. Thus, in some embodiments, the UAS site is used to recruit a protein (e.g., a fusion protein) comprising Gal4 and an epigenetic modifier.
[0069] Additional recruitment motifs are also contemplated herein. For example, any nucleotide binding protein may be used to drive recruitment of an epigenetic modifier to the artificial centromere, wherein the artificial centromere comprises a nucleotide sequence that is bound by the nucleotide binding protein. In some embodiments, the nucleotide binding protein is a TALEN and the nucleotide sequence is one that is recognized and bound by a TALEN. In some embodiments, the nucleotide binding protein is a Cas protein and the nucleotide sequence is one that is recognized and bound by a Cas protein, e.g., a CRISPR sequence.
[0070] In some embodiments, a wild type LacO site is used. Alternatively, or in addition, a variant of a wild type LacO site can be used, as long as the variant is capable of recruiting an epigenetic modifier to the centromere. In certain embodiments, the LacO site comprises a nucleic acid sequence set forth in SEQ ID NO: 1 (TGTGANCGNTCACA), wherein each N is, independently, any nucleotide. In some embodiments, N is thymine (T). In some embodiments, N is cytosine (C). In some embodiments, N is adenine (A). In some embodiments, N is guanine (G). In some embodiments, each N is, independently, selected from T, C, A, or G. In some embodiments, the LacO site is substantially similar to the sequence of SEQ ID NO: 1. In some embodiments, the LacO site comprises a nucleic acid sequence having the nucleic acid sequence of SEQ ID NO: 1, or a nucleic acid sequence having a 1 nucleotide substitution as compared to SEQ ID NO: 1. In some embodiments, the LacO site comprises an amino acid sequence having the nucleic acid sequence of SEQ ID NO: 1, or a nucleic acid sequence having 2 nucleotide substitutions as compared to SEQ ID NO: 1. In some embodiments, the LacO site comprises an amino acid sequence having the nucleic acid sequence of SEQ ID NO: 1, or a nucleic acid sequence having 3 nucleotide substitutions as compared to SEQ ID NO: 1. In some embodiments, the LacO site is a known variation of the wild type LacO site.
[0071] In some embodiments, a wild type tetO site is used. Alternatively, or in addition, a variant of a wild type tetO site can be used, as long as the variant is capable of recruiting an epigenetic modifier to the centromere. In certain embodiments, the tetO site comprises a 16 IPTS / 200118216.1Attorney Docket No.: CETA-001WO nucleic acid sequence set forth in SEQ ID NO: 2 (TCTCTATCACTGATAGGGA). In some embodiments, the tetO site comprises a nucleic acid sequence having the nucleic acid sequence of SEQ ID NO: 2, or a nucleic acid sequence having a 1 nucleotide substitution as compared to SEQ ID NO: 2. In some embodiments, the tetO site comprises a nucleic acid sequence having the nucleic acid sequence of SEQ ID NO: 2, or a nucleic acid sequence having 2 nucleotide substitutions as compared to SEQ ID NO: 2. In some embodiments, the TetO site comprises a nucleic sequence having the nucleic acid sequence of SEQ ID NO: 2, or a nucleic acid sequence having 3 nucleotide substitutions as compared to SEQ ID NO: 2. In some embodiments, the tetO site is a known variation of the wild type tetO site.
[0072] In some embodiments, a wild type lexO site is used. Alternatively, or in addition, a variant of a wild type LexO site can be used, as long as the variant is capable of recruiting an epigenetic modifier to the centromere. In certain embodiments, the lexO site comprises a nucleic acid sequence set forth in SEQ ID NO: 3 (WACTGTATAWAWAWMCAGYAM, where W is either T or A, M is either A or C, Y is either C or T). In some embodiments, the lexO site comprises a nucleic acid sequence having the nucleic acid sequence of SEQ ID NO: 3, or a nucleic acid sequence having a 1 nucleotide substitution as compared to SEQ ID NO: 3. In some embodiments, the lexO site comprises a nucleic acid sequence having the nucleic acid sequence of SEQ ID NO: 3, or a nucleic acid sequence having 2 nucleotide substitutions as compared to SEQ ID NO: 3. In some embodiments, the lexO site comprises a nucleic acid sequence having the nucleic acid sequence of SEQ ID NO: 3, or a nucleic acid sequence having 3 nucleotide substitutions as compared to SEQ ID NO: 3. In some embodiments, the lexO site is a known variation of the wild type lexO site. In some embodiments, the lexO site comprises a nucleic acid sequence having the nucleic acid sequence of CTGTATATATATACAG (SEQ ID NO: 25). In some embodiments, the lexO site comprises a nucleic acid sequence having the nucleic acid sequence of CTGTATGAGCATACAG (SEQ ID NO: 26).
[0073] In some embodiments, a wild type UAS site is used. Alternatively, or in addition, a variant of a wild type UAS site can be used, as long as the variant is capable of recruiting an epigenetic modifier to the centromere. In certain embodiments, the UAS site comprises a nucleic acid sequence set forth in SEQ ID NO: 4 (CGGAGGACTGTCCTCCG). In some embodiments, the UAS site comprises a nucleic acid sequence having the nucleic acid sequence of SEQ ID NO: 4, or a nucleic acid sequence having a 1 nucleotide substitution as compared to SEQ ID NO: 4. In some embodiments, the UAS site comprises an amino acid sequence having the nucleic acid sequence of SEQ ID NO: 4, or a nucleic acid sequence 17 IPTS / 200118216.1Attorney Docket No.: CETA-001WO having 2 nucleotide substitutions as compared to SEQ ID NO: 4. In some embodiments, the UAS site comprises an amino acid sequence having the nucleic acid sequence of SEQ ID NO: 4, or a nucleic acid sequence having 3 nucleotide substitutions as compared to SEQ ID NO: 4. In some embodiments, the UAS site is a known variation of the wild type UAS site. In certain embodiments, the UAS site comprises a nucleic acid sequence set forth in SEQ ID NO: 27 (CGGATTAGAAGCCGCCG). In certain embodiments, the UAS site comprises a nucleic acid sequence set forth in SEQ ID NO: 28 (CGGGTGACAGCCCTCCG). In certain embodiments, the UAS site comprises a nucleic acid sequence set forth in SEQ ID NO: 29 (AGGAAGACTCTCCTCCG). In certain embodiments, the UAS site comprises a nucleic acid sequence set forth in SEQ ID NO: 30 (CGCGCCGCACTGCTCCG).
[0074] In some embodiments, a recruitment motif is present at about one recruitment motif per 100-300bp of the nucleic acid of the centromere. In some embodiments, the recruitment motif is present at a frequency of about one recruitment motif per 100-200bp of the nucleic acid of the centromere. In some embodiments, the recruitment motif is present at a frequency of about one recruitment motif per 300bp of the nucleic acid of the centromere. In some embodiments, the recruitment motif is present at a frequency of about one recruitment motif per 250bp of the nucleic acid of the centromere. In some embodiments, the recruitment motif is present at a frequency of about one recruitment motif per 200bp of the nucleic acid of the centromere. In some embodiments, the recruitment motif is present at a frequency of about one recruitment motif per 150bp of the nucleic acid of the centromere. In some embodiments, the recruitment motif is present at a frequency of about one recruitment motif per 100bp of the nucleic acid of the centromere.
[0075] In some embodiments, the centromere comprises more than one type of recruitment motif. In some embodiments, the centromere comprises more than two types of recruitment motifs. For example, the centromere can include 2, 3, 4, 5, 6 ,7, 8, 9 or 10 different types of binding motifs. In some embodiments, the more than one type of recruitment motif are as provided for herein. In some embodiments, the more than two types of recruitment motif are as provided for herein. In some embodiments, each recruitment motif of the more than one type or the more than two types of recruitment motifs are each, individually, selected from a recruitment motif as provided for herein. In some embodiments, each of the more than one types of recruitment motifs are, individually, selected from the group including, but not limited to, LacO site, a TetO site, a LexO site, or a UAS site. In some embodiments, each of the more than two types of recruitment motifs are, individually, 18 IPTS / 200118216.1Attorney Docket No.: CETA-001WO selected from the group including, but not limited to, LacO site, a TetO site, a LexO site, or a UAS site. ii. Epigenetic Modifiers
[0076] In some embodiments, the epigenetic modifier is any appropriate epigenetic modifier. In some embodiments, the epigenetic modifier is HJURP. An epigenetic modifier can be fused or otherwise associated (or able to be associated) with a protein or other moiety that binds to the binding motif of the disclosure, allowing for the recruitment of the epigenetic modifier to the site of the binding motif in the artificial centromere. In some embodiments, the epigenetic modifier is fused to a DNA binding protein to facilitate recruitment to the artificial centromere. In some embodiments, the DNA binding protein is as provided for herein. In some embodiments, the DNA binding protein is selected from, but not limited to, LacI, tTA, a lexA transcription factor, or Gal4. In some embodiments, HJURP is fused to a DNA binding protein to facilitate recruitment to the artificial centromere. In some embodiments, the DNA binding protein is as provided for herein. In some embodiments, the DNA binding protein is selected from, but not limited to, LacI, tTA, a lexA transcription factor, or Gal4. In some embodiments, HJURP is fused to a DNA binding protein is selected from, but not limited to, LacI, tTA, a lexA transcription factor, or Gal4. B. Alpha Satellite Arrays: Non-Identical Repetitive Units and Higher Order Repeats
[0077] Repetitive DNA makes up a majority of the human genome. One class of repetitive DNA, alpha satellite DNA, comprises up to 10% of the genome. Alpha satellite DNA is enriched at all human centromere regions and is competent for de novo centromere assembly. In humans, alpha satellite DNA is composed of 171bp monomeric repeat units and is present as either (1) higher-order repeat units (HORs) of about 700-2000bp that are composed of organized, tandemly repeated 171bp monomers or (2) stretches of divergent monomers that lack any overarching organizational pattern. These two types of alpha satellite DNA can be located close to one another, with unordered monomeric alpha satellite often sandwiched between a large block of HOR alpha satellite DNA and the chromosome arms. Active alpha satellite DNA is characterized in that it is kinetochore forming, whereas non-active alpha satellite DNA is non- kinetochore forming. Active and nonactive satellite DNA can be located close to each other. Active alpha satellite DNA can be identified as described in Nicolas Altemose, et al. (2022) “Complete genomic and epigenetic maps of human centromeres,” SCIENCE 376 eabl4178, COI: 10.1126 / science.abl4178, which is hereby 19 IPTS / 200118216.1Attorney Docket No.: CETA-001WO incorporated by reference in its entirety. In some embodiments, active alpha satellite regions are defined by high histone signal. In some embodiments, the histone is the histone H3 variant, CENP-A (Centromere Protein A). In some embodiments, active alpha satellite regions are defined by high CENP-A signal.
[0078] The artificial centromeres of the instant disclosure comprise modified alpha satellite DNA, including higher order repeats (HORs) that are similar, but not identical, to those found in human DNA. HORs comprise tandem arrays of about 4 to about 12 non- identical repetitive units of about 171bp (e.g., about 160 to about 180 bp). HORs can range in length from 700-2000bp.
[0079] In some embodiments, the higher order repeat of the artificial centromere comprises about 4-12 non-identical repetitive units. In some embodiments, the higher order repeat comprises about 5-11 non-identical repetitive units. In some embodiments, the higher order repeat comprises about 6-10 non-identical repetitive units. In some embodiments, the higher order repeat comprises about 7-9 non-identical repetitive units. In some embodiments, the higher order repeat comprises 4 non-identical repetitive units. In some embodiments, the higher order repeat comprises 5 non-identical repetitive units. In some embodiments, the higher order repeat comprises 6 non-identical repetitive units. In some embodiments, the higher order repeat comprises 7 non-identical repetitive units. In some embodiments, the higher order repeat comprises 8 non-identical repetitive units. In some embodiments, the higher order repeat comprises 9 non-identical repetitive units. In some embodiments, the higher order repeat comprises 10 non-identical repetitive units. In some embodiments, the higher order repeat comprises 11 non-identical repetitive units. In some embodiments, the higher order repeat comprises 12 non-identical repetitive units.
[0080] In some embodiments, the non-identical repetitive units have between about 55% and about 95% sequence identity to each other, for example, from about 55% to about 60%, from about 55% to about 65%, from about 55% to about 70%, from about 55% to about 75%, from about 55% to about 80%, from about 55% to about 85%, from about 55% to about 90%, from about 60% to about 65%, from about 60% to about 70%, from about 60% to about 75%, from about 60% to about 80%, from about 60% to about 85%, from about 60% to about 90%, from about 60% to about 95%, from about 65% to about 70%, from about 65% to about 75%, from about 65% to about 80%, from about 65% to about 85%, from about 65% to about 90%, from about 65% to about 95%, from about 70% to about 75%, from about 70% to about 80%, from about 70% to about 85%, from about 70% to about 90%, from about 70% to about 95%, from about 75% to about 80%, from about 75% to about 85%, from about 75% to about 90%, 20 IPTS / 200118216.1Attorney Docket No.: CETA-001WO from about 75% to about 95%, from about 80% to about 85%, from about 80% to about 90%, from about 80% to about 95%, from about 85% to about 90%, or from about 85% to about 95% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 60% and about 90% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 65% and about 85% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 70% and about 80% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 75% and about 80% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 90% and about 99% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 91% and about 98% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 92% and about 97% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 93% and about 96% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 94% and about 95% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 99% and about 99.9% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 99.1% and about 99.8% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 99.2% and about 99.7% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 99.3% and about 99.6% sequence identity to each other. In some embodiments, the non-identical repetitive units have between about 99.4% and about 99.5% sequence identity to each other. In some embodiments, the non- identical repetitive units have about 50% sequence identity to each other. In some embodiments, the non-identical repetitive units have about 60% sequence identity to each other. In some embodiments, the non-identical repetitive units have about 70% sequence identity to each other. In some embodiments, the non-identical repetitive units have about 80% sequence identity to each other. In some embodiments, the non-identical repetitive units have about 90% sequence identity to each other. In some embodiments, the non-identical repetitive units have about 95% sequence identity to each other. In some embodiments, the non-identical repetitive units have about 96% sequence identity to each other. In some embodiments, the non-identical repetitive units have about 97% sequence identity to each other. In some embodiments, the non-identical repetitive units have about 98% sequence identity to each other. In some embodiments, the non-identical repetitive units have about 21 IPTS / 200118216.1Attorney Docket No.: CETA-001WO 99% sequence identity to each other. In some embodiments, the non-identical repetitive units have about 99.9% sequence identity to each other.
[0081] In some embodiments, the HOR is between about 680bp and about 2400bp. In some embodiments, the HOR is between about 700bp and about 2300bp. In some embodiments, the HOR is between about 800bp and about 2200bp. In some embodiments, the HOR is between about 900bp and about 2100bp. In some embodiments, the HOR is between about 1000bp and about 2000bp. In some embodiments, the HOR is between about 1100bp and about 1900bp. In some embodiments, the HOR is between about 1200bp and about 1800bp. In some embodiments, the HOR is between about 1300bp and about 1700bp. In some embodiments, the HOR is between about 1400bp and about 1600bp. In some embodiments, the HOR is about 700bp long. In some embodiments, the HOR is about 800bp long. In some embodiments, the HOR is about 900bp long. In some embodiments, the HOR is about 1000bp long. In some embodiments, the HOR is about 1200bp long. In some embodiments, the HOR is about 1400bp long. In some embodiments, the HOR is about 1500bp long. In some embodiments, the HOR is about 1600bp long. In some embodiments, the HOR is about 1800bp long. In some embodiments, the HOR is about 2000bp long. In some embodiments, the HOR is about 2200bp long. In some embodiments, the HOR is about 2400bp long.
[0082] In some embodiments, a HOR comprises a nucleic acid sequence of SEQ ID NO: 18, SEQ ID NO: 19, SEQ ID NO: 20, SEQ ID NO: 21, SEQ ID NO: 22, SEQ ID NO: 23, or SEQ ID NO: 24. In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 18, or a sequence substantially similar to SEQ ID NO: 18. In some embodiments, the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 18. In some embodiments, the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 18. In some embodiments, the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 18. In some embodiments, the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 18. In some embodiments, the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 18. In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 18.
[0083] In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 19, or a sequence substantially similar to SEQ ID NO: 19. In some embodiments, the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at 22 IPTS / 200118216.1Attorney Docket No.: CETA-001WO least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 19. In some embodiments, the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 19. In some embodiments, the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 19. In some embodiments, the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 19. In some embodiments, the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 19. In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 19.
[0084] In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 20, or a sequence substantially similar to SEQ ID NO: 20. In some embodiments, the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 20. In some embodiments, the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 20. In some embodiments, the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 20. In some embodiments, the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 20. In some embodiments, the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 20. In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 20.
[0085] In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 21, or a sequence substantially similar to SEQ ID NO: 21. In some embodiments, the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 21. In some embodiments, the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 21. In some embodiments, the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 21. In some embodiments, the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 21. In some embodiments, the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 21. In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 21.
[0086] In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 22, or a sequence substantially similar to SEQ ID NO: 22. In some embodiments, the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to 23 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ ID NO: 22. In some embodiments, the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 22. In some embodiments, the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 22. In some embodiments, the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 22. In some embodiments, the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 22. In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 22.
[0087] In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 23, or a sequence substantially similar to SEQ ID NO: 23. In some embodiments, the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 23. In some embodiments, the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 23. In some embodiments, the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 23. In some embodiments, the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 23. In some embodiments, the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 23. In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 23.
[0088] In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 24, or a sequence substantially similar to SEQ ID NO: 24. In some embodiments, the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 24. In some embodiments, the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 24. In some embodiments, the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 24. In some embodiments, the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 24. In some embodiments, the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 24. In some embodiments, the HOR comprises a nucleic acid sequence of SEQ ID NO: 24.
[0089] In some embodiments, the non-identical repetitive unit (monomer) of the HOR is about 170 bp long. In some embodiments, the non-identical repetitive unit of the HOR between about 160bp and about 180bp long. In some embodiments, the non-identical repetitive unit of the HOR is between about 161bp and about 179bp long. In some embodiments, the non-identical repetitive unit of the HOR is between about 162bp and about 24 IPTS / 200118216.1Attorney Docket No.: CETA-001WO 178bp long. In some embodiments, the non-identical repetitive unit of the HOR is between about 163bp and about 177bp long. In some embodiments, the non-identical repetitive unit of the HOR is between about 164bp and about 176bp long. In some embodiments, the non- identical repetitive unit of the HOR is between about 165bp and about 175bp long. In some embodiments, the non-identical repetitive unit of the HOR is between about 166bp and about 174bp long. In some embodiments, the non-identical repetitive unit of the HOR is between about 167bp and about 173bp long. In some embodiments, the non-identical repetitive unit of the HOR is between about 168bp and about 172bp long. In some embodiments, the non- identical repetitive unit of the HOR is between about 169bp and about 171bp long.
[0090] In some embodiments, the non-identical repetitive unit in a HOR is between about 680bp and about 2040bp long. In some embodiments, the non-identical repetitive unit in a HOR is between about 850bp and about 1870bp long. In some embodiments, the non- identical repetitive unit in a HOR is between about 1020bp and about 1700bp long. In some embodiments, the non-identical repetitive unit in a HOR is between about 1020bp and about 1530bp long. In some embodiments, the non-identical repetitive unit in a HOR is between about 1020bp and about 1530bp long. In some embodiments, the non-identical repetitive unit in a HOR is 680bp long. In some embodiments, the non-identical repetitive unit is 850bp long. In some embodiments, the non-identical repetitive unit in a HOR is 1020bp long. In some embodiments, the non-identical repetitive unit in a HOR is 1190bp long. In some embodiments, the non-identical repetitive unit in a HOR is 1360bp long. In some embodiments, the non-identical repetitive unit in a HOR is 1530bp long. In some embodiments, the non-identical repetitive unit in a HOR is 1700bp long. In some embodiments, the non-identical repetitive unit in a HOR is 1870bp long. In some embodiments, the non-identical repetitive unit in a HOR is 2040bp long.
[0091] In some embodiments, a non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 11, SEQ ID NO: 12, SEQ ID NO: 13, SEQ ID NO: 14, SEQ ID NO: 15, SEQ ID NO: 16, or SEQ ID NO: 17. In some embodiments, the non- identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 11, or a sequence substantially similar to SEQ ID NO: 11. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 11. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 11. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic 25 IPTS / 200118216.1Attorney Docket No.: CETA-001WO acid sequence having at least 95% identity to SEQ ID NO: 11. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 11. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 11. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 11.
[0092] In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 12, or a sequence substantially similar to SEQ ID NO: 12. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 12. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 12. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 12. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 12. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 12. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 12.
[0093] In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 13, or a sequence substantially similar to SEQ ID NO: 13. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 13. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 13. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 13. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 13. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 13. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 13.
[0094] In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 14, or a sequence substantially similar to SEQ ID NO: 26 IPTS / 200118216.1Attorney Docket No.: CETA-001WO 14. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 14. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 14. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 14. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 14. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 14. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 14.
[0095] In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 15, or a sequence substantially similar to SEQ ID NO: 15. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 15. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 15. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 15. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 15. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 15. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 15.
[0096] In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 16, or a sequence substantially similar to SEQ ID NO: 16. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 16. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 16. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 16. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 16. In some 27 IPTS / 200118216.1Attorney Docket No.: CETA-001WO embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 16. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 16.
[0097] In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 17, or a sequence substantially similar to SEQ ID NO: 17. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 17. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 17. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 17. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 17. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 17. In some embodiments, the non-identical repetitive unit of the HOR comprises a nucleic acid sequence of SEQ ID NO: 17.
[0098] In some embodiments, the non-identical repetitive unit is flanked 5’ or 3’ by a recruitment motif for recruiting epigenetic modifiers. In some embodiments, the non- identical repetitive unit is flanked 5’ and 3’ by a recruitment motif. In some embodiments, the recruitment motif is embedded in (i.e., present in the middle of) a non-identical repetitive unit.
[0099] In some embodiments, the HOR comprises a recruitment motif about every 170 to about 250 bases. In some embodiments, the HOR comprises a recruitment motif about every 170 bases. In some embodiments, the HOR comprises a recruitment motif about every 180 bases. In some embodiments, the HOR comprises recruitment motif about every 190 bases. In some embodiments, the HOR comprises a recruitment motif about every 200 bases. In some embodiments, the HOR comprises a recruitment motif about every 210 bases. In some embodiments, the HOR comprises a recruitment motif about every 220 bases. In some embodiments, the HOR comprises a recruitment motif about every 230 bases. In some embodiments, the HOR comprises a recruitment motif about every 240 bases. In some embodiments, the HOR comprises a recruitment motif about every 250 bases.
[0100] In some embodiments, the HOR comprises two or more recruitment motifs. In some embodiments, the HOR comprises two or more recruitment motifs about every 170 to about 250 bases. In some embodiments, the HOR comprises recruitment motifs about every 28 IPTS / 200118216.1Attorney Docket No.: CETA-001WO 170 bases. In some embodiments, the HOR comprises two or more recruitment motifs about every 180 bases. In some embodiments, the HOR comprises two or more recruitment motifs about every 190 bases. In some embodiments, the HOR comprises two or more recruitment motifs about every 200 bases. In some embodiments, the HOR comprises two or more recruitment motifs about every 210 bases. In some embodiments, the HOR comprises two or more recruitment motifs s about every 220 bases. In some embodiments, the HOR comprises two or more recruitment motifs about every 230 bases. In some embodiments, the HOR comprises two or more recruitment motifs about every 240 bases. In some embodiments, the HOR comprises two or more recruitment motifs about every 250 bases.
[0101] In some embodiments, the centromere comprises from about 1 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 2 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 3 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 4 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 5 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 6 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 7 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 8 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 9 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 10 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 11 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 12 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 13 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 14 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 15 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 16 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 17 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 18 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 19 to about 20 higher order repeats.
[0102] In some embodiments, the centromere comprises from about 1 to about 20 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 19 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 18 higher order repeats. In some embodiments, the centromere comprises from about 1 to 29 IPTS / 200118216.1Attorney Docket No.: CETA-001WO about 17 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 16 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 15 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 14 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 13 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 12 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 11 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 10 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 9 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 8 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 7 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 6 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 5 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 4 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 3 higher order repeats. In some embodiments, the centromere comprises from about 1 to about 2 higher order repeats.
[0103] In some embodiments, the centromere comprises from about 11 to about 19 higher order repeats. In some embodiments, the centromere comprises from about 12 to about 18 higher order repeats. In some embodiments, the centromere comprises from about 13 to about 17 higher order repeats. In some embodiments, the centromere comprises from about 14 to about 16 higher order repeats. In some embodiments, the centromere comprises from about 5 to about 15 higher order repeats. In some embodiments, the centromere comprises from about 5 to about 10 higher order repeats.
[0104] In some embodiments, the centromere comprises 1 higher order repeats. In some embodiments, the centromere comprises 2 higher order repeats. In some embodiments, the centromere comprises 3 higher order repeats. In some embodiments, the centromere comprises 4 higher order repeats. In some embodiments, the centromere comprises 5 higher order repeats. In some embodiments, the centromere comprises 6 higher order repeats. In some embodiments, the centromere comprises 7 higher order repeats. In some embodiments, the centromere comprises 8 higher order repeats. In some embodiments, the centromere comprises 9 higher order repeats. In some embodiments, the centromere comprises 10 higher order repeats. In some embodiments, the centromere comprises 11 higher order repeats. In some embodiments, the centromere comprises 12 higher order repeats. In some 30 IPTS / 200118216.1Attorney Docket No.: CETA-001WO embodiments, the centromere comprises 13 higher order repeats. In some embodiments, the centromere comprises 14 higher order repeats. In some embodiments, the centromere comprises 15 higher order repeats. In some embodiments, the centromere comprises 16 higher order repeats. In some embodiments, the centromere comprises 17 higher order repeats. In some embodiments, the centromere comprises 18 higher order repeats. In some embodiments, the centromere comprises 19 higher order repeats. In some embodiments, the centromere comprises 20 higher order repeats.
[0105] Without being bound to any particular theory, an artificial centromere may benefit from repetitive sequences for improved function, but highly repetitive sequences may be difficult to synthesize and / or sequence. Thus, to produce a centromere that is both synthesizable and functional, a measure of repetitiveness of the centromere may be considered. One measure of repetitiveness is a k value. In the context of the present disclosure, a k value is the number of bases at or above which there is no repetition in the HAC. For example, a HAC having k=20 indicates that the HAC has no 20-mer sequence which is repeated within the HAC sequence. Likewise, a HAC having k=50 indicates that the HAC has no 50-mer sequence which is repeated within the HAC sequence. Thus, it is readily apparent that the k=20 HAC constructs would be less repetitive than the k=50 HAC constructs.
[0106] In some embodiments, the artificial centromere comprises a k value of 20 to 172, or any value or range in-between. In some embodiments, the artificial centromere comprises a k value of 25 to 172. In some embodiments, the artificial centromere comprises a k value of 30 to 172. In some embodiments, the artificial centromere comprises a k value of 35 to 172. In some embodiments, the artificial centromere comprises a k value of 40 to 172. In some embodiments, the artificial centromere comprises a k value of 45 to 172. In some embodiments, the artificial centromere comprises a k value of 50 to 172. In some embodiments, the artificial centromere comprises a k value of 55 to 172. In some embodiments, the artificial centromere comprises a k value of 60 to 172. In some embodiments, the artificial centromere comprises a k value of 65 to 172. In some embodiments, the artificial centromere comprises a k value of 70 to 172. In some embodiments, the artificial centromere comprises a k value of 75 to 172. In some embodiments, the artificial centromere comprises a k value of 80 to 172. In some embodiments, the artificial centromere comprises a k value of 85 to 172. In some embodiments, the artificial centromere comprises a k value of 90 to 172. In some embodiments, the artificial centromere comprises a k value of 95 to 172. In some 31 IPTS / 200118216.1Attorney Docket No.: CETA-001WO embodiments, the artificial centromere comprises a k value of 100 to 172. In some embodiments, the artificial centromere comprises a k value of 110 to 172. In some embodiments, the artificial centromere comprises a k value of 120 to 172. In some embodiments, the artificial centromere comprises a k value of 130 to 172. In some embodiments, the artificial centromere comprises a k value of 140 to 172. In some embodiments, the artificial centromere comprises a k value of 150 to 172. In some embodiments, the artificial centromere comprises a k value of 160 to 172.
[0107] In some embodiments, the artificial centromere comprises a of 20 to 160. In some embodiments, the artificial centromere comprises a k value of 20 to 150. In some embodiments, the artificial centromere comprises a k value of 20 to 140. In some embodiments, the artificial centromere comprises a k value of 20 to 130. In some embodiments, the artificial centromere comprises a k value of 20 to 120. In some embodiments, the artificial centromere comprises a k value of 20 to 110. In some embodiments, the artificial centromere comprises a k value of 20 to 100. In some embodiments, the artificial centromere comprises a k value of 20 to 95. In some embodiments, the artificial centromere comprises a k value of 20 to 90. In some embodiments, the artificial centromere comprises a k value of 20 to 85. In some embodiments, the artificial centromere comprises a k value of 20 to 80. In some embodiments, the artificial centromere comprises a k value of 20 to 75. In some embodiments, the artificial centromere comprises a k value of 20 to 70. In some embodiments, the artificial centromere comprises a k value of 20 to 65. In some embodiments, the artificial centromere comprises a k value of 20 to 60. In some embodiments, the artificial centromere comprises a k value of 20 to 55. In some embodiments, the artificial centromere comprises a k value of 20 to 50. In some embodiments, the artificial centromere comprises a k value of 20 to 45. In some embodiments, the artificial centromere comprises a k value of 20 toIn some embodiments, the artificial centromere comprises a k value of 20 to 35. In some embodiments, the artificial centromere comprises a k value of 20 to 30. In some embodiments, the artificial centromere comprises a k value of 20 to 25.
[0108] In some embodiments, the artificial centromere comprises a k value of 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 171, or 172, or any value or range in-between. In some embodiments, the artificial centromere comprises a k value of 20. In some embodiments, the artificial centromere comprises a k value of 25. In some embodiments, the artificial centromere comprises a k value of 30. In 32 IPTS / 200118216.1Attorney Docket No.: CETA-001WO some embodiments, the artificial centromere comprises a k value of 35. In some embodiments, the artificial centromere comprises a k value of 40. In some embodiments, the artificial centromere comprises a k value of 45. In some embodiments, the artificial centromere comprises a k value of 50. In some embodiments, the artificial centromere comprises a k value of 55. In some embodiments, the artificial centromere comprises a k value of 60. In some embodiments, the artificial centromere comprises a k value of 65. In some embodiments, the artificial centromere comprises a k value of 70. In some embodiments, the artificial centromere comprises a k value of 75. In some embodiments, the artificial centromere comprises a k value of 80. In some embodiments, the artificial centromere comprises a k value of 85. In some embodiments, the artificial centromere comprises a k value of 90. In some embodiments, the artificial centromere comprises a k value of 95. In some embodiments, the artificial centromere comprises a k value of 100. In some embodiments, the artificial centromere comprises a k value of 110. In some embodiments, the artificial centromere comprises a k value of 120. In some embodiments, the artificial centromere comprises a k value of 130. In some embodiments, the artificial centromere comprises a k value of 140. In some embodiments, the artificial centromere comprises a k value of 150. In some embodiments, the artificial centromere comprises a k value of 160. In some embodiments, the artificial centromere comprises a k value of 170. In some embodiments, the artificial centromere comprises a k value of 171. In some embodiments, the artificial centromere comprises a k value of 172.
[0109] Exemplary artificial centromere sequences are set forth in SEQ ID NOs: 5-9. In some embodiments, an artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 5, SEQ ID NO: 6, SEQ ID NO: 7, SEQ ID NO: 8 or SEQ ID NO: 9 . In some embodiments, the artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 5, or a sequence substantially similar to SEQ ID NO: 5. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 5. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 5. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 5. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 5. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 5. In some embodiments, the artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 5. 33 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0110] In some embodiments, the artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 6, or a sequence substantially similar to SEQ ID NO: 6. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 6. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 6. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 6. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 6. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 6. In some embodiments, the artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 6.
[0111] In some embodiments, the artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 7, or a sequence substantially similar to SEQ ID NO: 7. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 7. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 7. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 7. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 7. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 7. In some embodiments, the artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 7.
[0112] In some embodiments, the artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 8, or a sequence substantially similar to SEQ ID NO: 8. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 8. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 8. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 8. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 8. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 99% 34 IPTS / 200118216.1Attorney Docket No.: CETA-001WO identity to SEQ ID NO: 8. In some embodiments, the artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 8.
[0113] In some embodiments, the artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 9, or a sequence substantially similar to SEQ ID NO: 9. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 60%, at least 65%, at least 70%, at least 75%, at least 80%. at least 85%, at least 90%, at least 95%, or at least 99% identity to SEQ ID NO: 9. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 90% identity to SEQ ID NO: 9. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 95% identity to SEQ ID NO: 9. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 98% identity to SEQ ID NO: 9. In some embodiments, the artificial centromere comprises a nucleic acid sequence having at least 99% identity to SEQ ID NO: 9. In some embodiments, the artificial centromere comprises a nucleic acid sequence of SEQ ID NO: 9. III. Artificial Chromosomes
[0114] Described herein, in certain embodiments, is a human artificial chromosome (HAC). A HAC can be used as a delivery vector for one or more genes for genetically engineering cells (e.g., human cells) without genomic integration. The HAC allows for stable expression of one or multiple genes of various lengths derived from diverse sources including animals, plants, bacteria, as well as engineered genes. HACs can be circular, compact, and stable throughout the workflow of synthesis, cloning, propagation, and assembly. The HACs use an artificial centromere as a central building block and can, in certain embodiments, further comprise a backbone and a cargo sequence, e.g., encoding for one or more cargo RNA or proteins.
[0115] HACs can be used for the creation of stable cell lines (e.g., mammalian cell lines) with high production of biomolecules or therapeutic proteins for example antibodies. Additionally, the HACs can be used as genome engineering tools, for example, for cell therapies or gene therapies. HACs can accommodate cargo that is large in size and exhibits low mutation risk, stable cargo expression, few or no off target effects, low immunogenicity, ease of customization, high transfection efficiency, and high cost effectiveness compared to other gene editing and gene expression technologies such as viral delivery, enzyme based gene editing or naked DNA. 35 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0116] In some embodiments, the artificial chromosome is between about 50kb and about 100Mb. In some embodiments, the artificial chromosome is between about 50kb and about 1000kb. In some embodiments, the artificial chromosome is between about 100kb and about 900kb. In some embodiments, the artificial chromosome is between about 200kb and about 800kb. In some embodiments, the artificial chromosome is between about 300kb and about 700kb. In some embodiments, the artificial chromosome is between about 400kb and about 600kb. In some embodiments, the artificial chromosome is between about 1Mb and about 10Mb. In some embodiments, the artificial chromosome is between about 2Mb and about 9Mb. In some embodiments, the artificial chromosome is between about 3Mb and about 8Mb. In some embodiments, the artificial chromosome is between about 4Mb and about 7Mb. In some embodiments, the artificial chromosome is between about 5Mb and about 6Mb. In some embodiments, the artificial chromosome is between about 10Mb and about 100Mb. In some embodiments, the artificial chromosome is between about 20Mb and about 90Mb. In some embodiments, the artificial chromosome is between about 30Mb and about 80Mb. In some embodiments, the artificial chromosome is between about 40Mb and about 80Mb. In some embodiments, the artificial chromosome is between about 50Mb and about 60Mb. In some embodiments, the artificial chromosome is 50kb. In some embodiments, the artificial chromosome is 60kb. In some embodiments, the artificial chromosome is 70kb. In some embodiments, the artificial chromosome is 80kb. In some embodiments, the artificial chromosome is 90kb. In some embodiments, the artificial chromosome is 100kb. In some embodiments, the artificial chromosome is 500kb. In some embodiments, the artificial chromosome is 1Mb. In some embodiments, the artificial chromosome is 10Mb. In some embodiments, the artificial chromosome is 20Mb. In some embodiments, the artificial chromosome is 30Mb. In some embodiments, the artificial chromosome is 40Mb. In some embodiments, the artificial chromosome is 50Mb. In some embodiments, the artificial chromosome is 60Mb. In some embodiments, the artificial chromosome is 70Mb. In some embodiments, the artificial chromosome is 80Mb. In some embodiments, the artificial chromosome is 90Mb. In some embodiments, the artificial chromosome is 100Mb.
[0117] In some embodiments, the artificial chromosome comprises a centromere sequence. In some embodiments, the centromere sequence is between 5kb to about 50kb. In some embodiments, the centromere comprises higher order repeats.
[0118] In some embodiments, the artificial chromosome comprises a backbone. In some embodiments the backbone comprises regulatory elements. In some embodiments, the backbone comprises bacterial regulatory elements. In some embodiments, the backbone 36 IPTS / 200118216.1Attorney Docket No.: CETA-001WO comprises yeast regulatory elements. In some embodiments, the backbone comprises mammalian regulatory elements. In some embodiments, the backbone comprises bacterial regulatory elements and mammalian regulatory elements. In some embodiments, the backbone comprises yeast regulatory elements and mammalian regulatory elements. In some embodiments, the regulatory elements are for replication in a prokaryotic cell or a eukaryotic cell. Regulatory elements comprise, for example, promoters, terminators, insulators, genomic localization elements, and antibiotic cassettes.
[0119] In some embodiments, the artificial chromosome comprises a promoter. In some embodiments, the promoter is selected from the group consisting of a mini promoter, an inducible promoter, a constitutive promoter, and derivatives thereof. In some embodiments, the promoter is selected from the group consisting of CMV, CBA, EF1a, CAG, PGK, TRE, U6, UAS, T7, Sp6, lac, araBad, trp, Ptac, p5, p19, p40, Synapsin, CaMKII, GRK1, and derivatives thereof. In some embodiments the promoter is a U6 promoter. In some embodiments, the promoter is a CAG promoter.
[0120] In some embodiments, the backbone comprises spacer elements, insulators, genomic localization elements, or combinations thereof. In some embodiments, the backbone comprises spacer elements. In some embodiments, the backbone comprises insulators. In some embodiments, the backbone comprises genomic localization elements. In some embodiments, the backbone comprises combinations of any of the foregoing. IV. Cargo
[0121] In some embodiments, the artificial chromosome comprises a nucleic acid cargo. In some embodiments, the cargo comprises a gene for expression in the cell. The gene can encode one or more non-coding RNA (such as siRNA, tRNA, rRNA, microRNA, and the like) and / or one or more proteins (e.g., two or more proteins). For example, the cargo can encode an antibody, an scFv, or a Fab. In some embodiments, when an artificial chromosome is present in a cell, the cargo is expressed in the cell.
[0122] In some embodiments, the cargo is between about 0kb and about 100Mb. In some embodiments, the cargo is between about 1kb and about 10kb. In some embodiments, the cargo is between about 0kb and about 1kb. In some embodiments, the cargo is between about 10kb and about 100kb. In some embodiments, the cargo is between about 100kb and about 1Mb. In some embodiments, the cargo is between about 1Mb and about 10Mb. In some embodiments, the cargo is between about 10Mb and about 100Mb. In some embodiments, the cargo is 50kb. In some embodiments, the cargo is 100kb. In some embodiments, the 37 IPTS / 200118216.1Attorney Docket No.: CETA-001WO cargo is 200kb. In some embodiments, the cargo is 300kb. In some embodiments, the cargo is 400kb. In some embodiments, the cargo is 500kb. In some embodiments, the cargo is 600kb. In some embodiments, the cargo is 700kb. In some embodiments, the cargo is 800kb. In some embodiments, the cargo is 900kb. In some embodiments, the cargo is 1Mb. In some embodiments, the cargo is 10Mb. In some embodiments, the cargo is 20Mb. In some embodiments, the cargo is 30Mb. In some embodiments, the cargo is 40Mb. In some embodiments, the cargo is 50Mb. In some embodiments, the cargo is 60Mb. In some embodiments, the cargo is 70Mb. In some embodiments, the cargo is 80Mb. In some embodiments, the cargo is 90Mb. In some embodiments, the cargo is 100Mb.
[0123] In some embodiments, the artificial chromosome comprises one or more insulator sequences. The insulator sequences flank the cargo sequence and prevent silencing of the cargo sequence in the cell. In some embodiments, the artificial chromosome comprises an insulator sequence at the 5’ end of the cargo sequence. In the context of the present disclosure, the phrase “at the 5’ end of the cargo sequence” is understood to encompass any location 5’ of the cargo sequence. Thus, the phrase “at the 5’ end of the cargo sequence” does encompass, but is not limited to, embodiments wherein the insulator sequence is immediately adjacent to the cargo sequence. The phrase “at the 5’ end of the cargo sequence” also encompasses embodiments wherein the insulator sequence is not immediately adjacent to the cargo sequence, but rather the insulator sequence is separated from the 5’ end of the cargo sequence by a nucleic acid spacer of any length. Within said nucleic acid spacer may be additional regulatory elements, such as but not limited to the regulatory elements as provided for herein, or promoters, such as but not limited to the promoters as provided for herein. In some embodiments, the artificial chromosome comprises an insulator sequence at the 3’ end of the cargo sequence. In the context of the present disclosure, the phrase “at the 3’ end of the cargo sequence” is understood to encompass any location 3’ of the cargo sequence. Thus, the phrase “at the 3’ end of the cargo sequence” does encompass, but is not limited to, embodiments wherein the insulator sequence is immediately adjacent to the cargo sequence. The phrase “at the 3’ end of the cargo sequence” also encompasses embodiments wherein the insulator sequence is not immediately adjacent to the cargo sequence, but rather the insulator sequence is separated from the 3’ end of the cargo sequence by a nucleic acid spacer of any length. Within said nucleic acid spacer may be additional regulatory elements, such as but not limited to the regulatory elements as provided for herein, or promoters, such as but not limited to the promoters as provided for herein. In some embodiments, the artificial chromosome comprises an insulator sequence at both the 5’ end and the 3’ end of 38 IPTS / 200118216.1Attorney Docket No.: CETA-001WO the cargo sequence. Insulator sequences are known in the art, and any such insulator sequence may be used in artificial chromosome.
[0124] In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 50% identity (e.g. at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31. In some embodiments, the insulator sequence comprises a nucleic acid sequence of SEQ ID NO: 31.
[0125] In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 50% identity (e.g. at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the insulator sequence comprises a nucleic acid sequence of SEQ ID NO: 32.
[0126] In some embodiments, the artificial chromosome comprises a first insulator sequence and a second insulator sequence flanking the cargo sequence, wherein the first insulator sequence is located on the 5’ end of the cargo sequence and the second insulator sequence is located on the 3’ end of the cargo sequence. In some embodiments, the artificial chromosome comprises a first insulator sequence and a second insulator sequence flanking the cargo sequence, wherein the first insulator sequence is located on the 3’ end of the cargo 39 IPTS / 200118216.1Attorney Docket No.: CETA-001WO sequence and the second insulator sequence is located on the 5’ end of the cargo sequence. In some embodiments, the first insulator sequence and the second insulator sequence comprise the same nucleic acid sequence. In some embodiments, the first insulator sequence and the second insulator sequence comprise different nucleic acid sequences. In some embodiments, the first insulator sequence is as provided for herein. In some embodiments, the second insulator sequence is as provided for herein.
[0127] In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid sequence having at least 50% identity (e.g. at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 50% identity (e.g. at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid sequence having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid 40 IPTS / 200118216.1Attorney Docket No.: CETA-001WO sequence of SEQ ID NO: 32. In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the artificial chromosome comprises a first insulator sequence comprising a nucleic acid of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence of SEQ ID NO: 32. V. Gene expression systems
[0128] In some embodiments, a gene expression system is provided, the gene expression system comprising a cargo nucleic acid sequence and one or more insulator sequences. The insulator sequences flank the cargo nucleic acid sequence and prevent silencing of the cargo sequence in the cell. In some embodiments, the gene expression system comprises an insulator sequence at the 5’ end of the cargo sequence. In the context of the present disclosure, the phrase “at the 5’ end of the cargo sequence” is understood to encompass any 41 IPTS / 200118216.1Attorney Docket No.: CETA-001WO location 5’ of the cargo sequence. Thus, the phrase “at the 5’ end of the cargo sequence” does encompass, but is not limited to, embodiments wherein the insulator sequence is immediately adjacent to the cargo sequence. The phrase “at the 5’ end of the cargo sequence” also encompasses embodiments wherein the insulator sequence is not immediately adjacent to the cargo sequence, but rather the insulator sequence is separated from the 5’ end of the cargo sequence by a nucleic acid spacer of any length. Within said nucleic acid spacer may be additional regulatory elements, such as but not limited to the regulatory elements as provided for herein, or promoters, such as but not limited to the promoters as provided for herein. In some embodiments, the gene expression system comprises an insulator sequence at the 3’ end of the cargo sequence. In the context of the present disclosure, the phrase “at the 3’ end of the cargo sequence” is understood to encompass any location 3’ of the cargo sequence. Thus, the phrase “at the 3’ end of the cargo sequence” does encompass, but is not limited to, embodiments wherein the insulator sequence is immediately adjacent to the cargo sequence. The phrase “at the 3’ end of the cargo sequence” also encompasses embodiments wherein the insulator sequence is not immediately adjacent to the cargo sequence, but rather the insulator sequence is separated from the 3’ end of the cargo sequence by a nucleic acid spacer of any length. Within said nucleic acid spacer may be additional regulatory elements, such as but not limited to the regulatory elements as provided for herein, or promoters, such as but not limited to the promoters as provided for herein. In some embodiments, the gene expression system comprises an insulator sequence at both the 5’ end and the 3’ end of the cargo sequence. Insulator sequences are known in the art, and any such insulator sequence may be used in gene expression system.
[0129] In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 50% identity (e.g. at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid 42 IPTS / 200118216.1Attorney Docket No.: CETA-001WO sequence of SEQ ID NO: 31. In some embodiments, the insulator sequence comprises a nucleic acid sequence of SEQ ID NO: 31.
[0130] In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 50% identity (e.g. at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the insulator sequence comprises a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the insulator sequence comprises a nucleic acid sequence of SEQ ID NO: 32.
[0131] In some embodiments, the gene expression system comprises a first insulator sequence and a second insulator sequence flanking the cargo sequence, wherein the first insulator sequence is located on the 5’ end of the cargo sequence and the second insulator sequence is located on the 3’ end of the cargo sequence. In some embodiments, the gene expression system comprises a first insulator sequence and a second insulator sequence flanking the cargo sequence, wherein the first insulator sequence is located on the 3’ end of the cargo sequence and the second insulator sequence is located on the 5’ end of the cargo sequence. In some embodiments, the first insulator sequence and the second insulator sequence comprise the same nucleic acid sequence. In some embodiments, the first insulator sequence and the second insulator sequence comprise different nucleic acid sequences. In some embodiments, the first insulator sequence is as provided for herein. In some embodiments, the second insulator sequence is as provided for herein.
[0132] In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid sequence having at least 50% identity (e.g. at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 50% identity (e.g. at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid 43 IPTS / 200118216.1Attorney Docket No.: CETA-001WO sequence of SEQ ID NO: 32. In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid sequence having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid sequence having at least 80% identity (e.g. at least 80%, at least 85%, at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence 44 IPTS / 200118216.1Attorney Docket No.: CETA-001WO comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid sequence having at least 90% identity (e.g. at least 90%, at least 95%, or at least 100%) to a nucleic acid sequence of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence of SEQ ID NO: 32. In some embodiments, the gene expression system comprises a first insulator sequence comprising a nucleic acid of SEQ ID NO: 31 and comprises a second insulator sequence comprising a nucleic acid sequence of SEQ ID NO: 32.
[0133] In some embodiments, the cargo sequence comprises a gene for expression in the cell. The gene can encode one or more non-coding RNA (such as siRNA, tRNA, rRNA, microRNA, and the like) and / or one or more proteins (e.g., two or more proteins). For example, the cargo can encode an antibody, an scFv, or a Fab. In some embodiments, when an artificial chromosome is present in a cell, the cargo is expressed in the cell. In some embodiments, the cargo sequence comprises a centromere sequence as provided for herein, wherein the centromere sequence encodes for one or more genes of interest as provided for herein.
[0134] In some embodiments, the gene expression system comprises a vector, such as but not limited to an expression vector. In some embodiments, the vector comprises a cargo nucleic acid sequence and one or more insulator sequences as provided for herein. In some embodiments, the vector comprises one or more promoters. In some embodiments, the one or more promoters are as provided for herein. In some embodiments, the vector comprises one or more regulatory elements. In some embodiments, the regulatory elements are as provided for herein.
[0135] In some embodiments, the vector further comprises a centromere sequence as provided for herein. In some embodiments, the gene expression system comprises a human artificial chromosome as provided for herein. 45 IPTS / 200118216.1Attorney Docket No.: CETA-001WO VI. In silico Methods of Designing an Artificial Centromere
[0136] Described herein, in certain embodiments, are methods of in silico design of an artificial centromere and an artificial chromosome. The in silico methods described herein allow for the creation of an artificial centromere having certain repetitive sequences (similar to those in a human centromere), but with sufficient variation to minimize or avoid problems with synthesis and stability.
[0137] As disclosed in further detail herein, artificial centromere sequences are designed using a multi-step computational approach. An exemplary method is provided as follows. In the first step, the sequences of the most biologically active higher-order repeat (HOR) regions of native human centromeres are selected (see FIG.2), pooled, and aligned to produce a position weight matrix (PWM). The PWM provides the frequency of particular nucleotides at each position of the HOR region.
[0138] In the second step, the PWM is modified to insert binding sites for the recruitment of epigenetic modifiers. As described supra, the binding sites are short protein-binding motifs (e.g., 10-15bp long). The binding sites for the recruitment of epigenetic modifiers are inserted into the HOR region sequence with regular spacing at positions computationally determined to be the least disruptive to the HOR region function. For example, the recruitment motifs may be inserted into the HOR at a position of high sequence similarity to the recruitment motif. In some embodiments, the recruitment motif is inserted into the HOR at a position within the HOR sequence that is substantially similar to the recruitment motif sequence (i.e., having at least 50% identity thereto). In some embodiments, the recruitment motif is inserted into the HOR at a position within the HOR that has at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or at least 99% identity to the recruitment motif sequence. The epigenetic modifier binding sites are inserted (e.g., every 100-200 bp) over the span of the HOR region sequence. In certain embodiments, to increase sequence variation, rare sequence variants from the native centromere sequences are used. Illustrations of exemplary HOR design with addition of epigenetic modifier recruiting regions is shown in FIGs.1, 2 and 3.
[0139] In the third step, the artificial centromeric sequence is constructed by linking together HOR regions constructed in the first and second step with each successive HOR region, having sequence divergence introduced by using an iterative computational approach. A minimal number of changes is introduced to the HOR consensus sequence until the resulting sequence passes a computational synthesizability and stability analysis. 46 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0140] Accordingly, in some embodiments, the in silico design method of the disclosure includes one or more of the following steps: a) identifying sequences from active alpha satellite arrays of human centromeric DNA having repetitive units and / or higher-order repeats; b) creating sequence alignments of the repetitive units and / or higher order repeats; c) creating a position weight matrix (PWM) from the sequence alignments; d) modifying the PWM to insert binding motifs for epigenetic modifiers; and e) designing an artificial centromeric sequence comprising a plurality of higher-order repeats, wherein each higher- order repeat comprises a plurality of non-identical repetitive units having between about 50% and about 99.9% sequence identity to each other.
[0141] The resulting artificial centromeric sequence can be computationally analyzed for stability or synthesizability. For example, the centromeric sequence can be analyzed with respect to sequence repetitiveness. The repetitiveness of a sequence is an indicator to the stability, synthesizability, or both of the centromeric sequence. In some embodiments, a highly repetitive sequence is less stable and therefore less synthesizable. In some embodiments, if the artificial centromeric sequence does not pass a threshold measurement for stability or synthesizability, additional changes are made to improve the stability or synthesizability of the artificial centromeric sequence. In certain embodiments, a centromeric sequence is analyzed for repetitiveness by determining the sequence identity between repeats (e.g., between non-identical repetitive units). In certain embodiments, a threshold measurement is determined by empirically measuring at least one aspect of the stability or synthesizability of the centromeric sequence for a series of centromeric sequences having different sequence identity between repeats, thereby to determine the threshold measurement. A threshold measurement for a given centromeric sequence may be different depending upon the DNA synthesis technology used to synthesize the sequence, the cloning methods, the cell in which the centromeric sequence is expressed, etc.
[0142] When a centromeric sequence has been computationally determined, a nucleic acid comprising the artificial centromeric sequence can be generated using any methods of manufacture known in the art, described in detail under “Methods of Manufacture” herein.
[0143] In some embodiments, the method further comprises designing an artificial chromosome comprising an artificial centromere. In some embodiments, the artificial chromosome further comprises a backbone nucleic acid sequence and a cargo nucleic acid sequence. 47 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0144] In some embodiments, the artificial centromeric sequence is about 5 to about 50 kb. In some embodiments, the artificial centromeric sequence is between 5 and 15 kb. In some embodiments, the artificial centromeric sequence is about 10 kb.
[0145] In some embodiments, the artificial centromeric sequence has between about 30% and about 99.9% sequence identity to a human centromere. In some embodiments, the artificial centromeric sequence has between about 30% and about 60% identity to a human centromere.
[0146] In some embodiments, the sequences from the active alpha satellite arrays of the human centromeric DNA are from a reference genome, wherein the reference genome is CHM13.
[0147] In some embodiments, the binding motifs comprise one or more LacO sites. In some embodiments, at least one binding motif is present per 100 bp to 200 bp of the artificial centromeric sequence.
[0148] In some embodiments, each repetitive unit is between about 160 bp and about 180 bp. In some embodiments, each repetitive unit is about 170 bp.
[0149] In some embodiments, the plurality of non-identical repetitive units is from about 4 to about 12 units.
[0150] In some embodiments, the higher-order repeat is between about 500 bp and about 2500 bp. In some embodiments, the higher-order repeat is between about 700 bp and about 2000 bp. VII. Methods of Manufacture
[0151] Described herein, in certain embodiments are methods synthesis of artificial centromeres and artificial chromosomes. Methods of manufacturing nucleic acids are well known in the art. For example, a nucleic acid can be synthesized using, for example, traditional oligonucleotide synthesis, column-based oligonucleotide synthesis, microarray- based oligonucleotide synthesis, gene synthesis from oligonucleotides, gene synthesis from array-derived oligonucleotide pools, and others. (See, for example, Hughes et al. (2017) COLDSPRINGHARBPERSPECTBIOL.9(1): a023812.) In certain embodiments, smaller fragments of a nucleic acid are synthesized and then assembled into larger nucleic acids, for example, using cloning methods (e.g., Golden Gate cloning (Engler et al. (2008) PLOS ONE3(11): e3647). Nucleic acid synthesis can be performed using machinery and / or services provided by, for example, Telesis Bio (formerly Codex DNA), Thermo Fisher Scientific™, or Genscript. 48 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0152] The assembled artificial chromosomes can then be propagated in bacterial cells using standard molecular biology techniques. Alternatively, the artificial chromosome can be assembled in vitro through combination of the synthesized DNA with purified histone proteins. The artificial chromosomes can be purified and introduced into a cell (e.g., a mammalian cell). VIII. Cells Comprising Artificial Chromosomes
[0153] The artificial chromosomes described herein can be delivered to a cell. The artificial chromosomes can be propagated and any cargo expressed in the cell. In the mammalian cell, the epigenetic seeding sites facilitate the three dimensional assembly of the centromere region of the artificial chromosome. The epigenetic modifier can be provided separately, for example by stable expression in the cell or on a second plasmid. In some embodiments, the epigenetic modifier is stably expressed by the cell. In some embodiments, the epigenetic modifier is provided to the cell via a second plasmid. In some embodiments, the epigenetic modifier is encoded on the artificial chromosome and is provided with the artificial chromosome.
[0154] In some embodiments, the artificial chromosome is delivered to the cell by a liposomal delivery, a polymeric delivery, or a viral delivery.
[0155] In some embodiments, the artificial chromosome is delivered by a non-nucleic acid-based delivery system (e.g., a non-viral delivery system). In some embodiments, the artificial chromosome is comprised in a liposome. In some embodiments, the artificial chromosome is associated with a lipid. The artificial chromosome associated with a lipid, in some embodiments, is encapsulated in the aqueous interior of a liposome, interspersed within the lipid bilayer of a liposome, attached to a liposome via a linking molecule that is associated with both the liposome and the artificial chromosome, entrapped in a liposome, complexed with a liposome, dispersed in a solution containing a lipid, mixed with a lipid, combined with a lipid, contained as a suspension in a lipid, contained or complexed with a micelle, or otherwise associated with a lipid. In some embodiments, the artificial chromosome is comprised in a lipid nanoparticle (LNP). In some embodiments, the artificial chromosome is electroporated into a cell. In some embodiments, the artificial chromosome is transfected into a cell.
[0156] In some embodiments, the artificial chromosome is delivered by a virus. In some embodiments, the virus is an alphavirus, a parvovirus, an adenovirus, an AAV, a baculovirus, 49 IPTS / 200118216.1Attorney Docket No.: CETA-001WO a Dengue virus, a lentivirus, a herpesvirus, a poxvirus, an anellovirus, a bocavirus, a vaccinia virus, or a retrovirus. In some embodiments, the virus is an alphavirus. In some embodiments, the virus is a parvovirus. In some embodiments, the virus is an adenovirus. In some embodiments, the virus is an AAV. In some embodiments, the virus is a baculovirus. In some embodiments, the virus is a Dengue virus. In some embodiments, the virus is a lentivirus. In some embodiments, the virus is a herpesvirus. In some embodiments, the virus is a poxvirus. In some embodiments, the virus is an anellovirus. In some embodiments, the virus is a bocavirus. In some embodiments, the virus is a vaccinia virus. In some embodiments, the virus is or a retrovirus.
[0157] In some embodiments, the cell is a eukaryotic cell (e.g., , an animal cell), a mammalian cell (a Chinese hamster ovary (CHO) cell, baby hamster kidney (BHK), human embryo kidney (HEK, e.g., HEK-293), mouse myeloma (NS0), or human retinal cells), a primary cell, an immortalized cell (e.g., a HeLa cell, a COS cell (e.g., COS-1 or COS-7), a HEK-293T cell, a MDCK cell, a 3T3 cell, a PC12 cell, a Huh7 cell, a HepG2 cell, a K562 cell, a N2a cell, or a SY5Y cell), an insect cell (e.g., a Spodoptera frugiperda cell, a Trichoplusia ni cell, a Drosophila melanogaster cell, a S2 cell, or a Heliothis virescens cell. In some embodiments, the cell is a eukaryotic cell. In some embodiments, the cell is a mammalian cell. In some embodiments, the cell is an immortalized cell. In some embodiments, the cell is an insect cell. In some embodiments, the cell is a yeast cell. In some embodiments, the cell is a plant cell. In some embodiments, the cell is a fungal cell. In some embodiments, the cell is a prokaryotic cell.
[0158] In some embodiments, the cell is an A549 cell, an MRC5 cell, an Sf9 cell, a Vero cell, a BSC 1 cell, a BSC 40 cell, a BMT 10 cell, a WI38 cell, a Saos cell, a C2C12 cell, an L cell, an HT1080 cell, or a derivative of any of the foregoing.
[0159] In some embodiments, the artificial chromosome is stably maintained in the cell. In some embodiments, the artificial chromosome is stably maintained in the cell over several cell division cycles. In some embodiments, the artificial chromosome is stably maintained in the cell over at least 30 cell cycles. In some embodiments, the artificial chromosome is stable maintained in the cell for over 30 cell cycles. In some embodiments, the cargo is expressed in the cell. IX. Methods of Using Artificial Chromosomes and Cells Comprising Artificial Chromosomes 50 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0160] Described herein, in certain embodiments, are methods of using artificial chromosomes and cells comprising human artificial chromosomes. The artificial chromosomes, for example are useful in the expression of stable expression of one or more genes in a human host cell. The HACs can be used for the creation of stable cell lines with high production of biomolecules or therapeutic proteins, for example, antibodies, enzymes, or bioactive peptides. In some embodiments, the biomolecule or therapeutic protein is an antibody. In some embodiments, the biomolecule or therapeutic protein is an enzyme. In some embodiments, the biomolecule or therapeutic protein is a bioactive peptide, such as but not limited to a hormone or the like. In some embodiments, the HACs can be used for the creation of stable cell lines with high production of cellular scaffold materials, such as, but not limited to, collagens, laminins, and the like. Additionally, the HACs can be used as genome engineering tools, for example, for cell therapies.
[0161] In some embodiments, the method comprises culturing a cell comprising an artificial chromosome under conditions that maintain the artificial chromosome. In some embodiments, a cargo encoded in the artificial chromosome is expressed in the cell. In some embodiments, the artificial chromosome is maintained over several cell passages and the cargo is expressed over several cell passages. In some embodiments, the cell comprising the artificial chromosome stably expresses multiple cargo genes encoded on the artificial chromosome. EXAMPLES
[0162] Below are examples of specific embodiments for carrying out the present disclosure. The examples are offered for illustrative purposes only and are not intended to limit the scope of the present disclosure in any way. Efforts have been made to ensure accuracy with respect to numbers used (e.g., amounts, temperatures, etc.), but some experimental error and deviation should, of course, be allowed for. Example 1: In silico Design and Generation of Artificial Centromeric Sequences
[0163] Artificial centromere sequences were designed using a multi-step computational approach. In the first step, sequences of higher-order repeat (HOR) regions of native human centromeres were analyzed for length, sequence variability, and density of CENP-B binding sites and those falling within predetermined parameters were selected (see FIG.2), pooled, and aligned to produce a position weight matrix (PWM). The HOR regions are repetitive units in native centromeric sequences that typically consist of 4-12 basic repetitive units that 51 IPTS / 200118216.1Attorney Docket No.: CETA-001WO are roughly 170bp in length and typically share very high sequence similarity to the other repetitive units within a HOR region. Thus, in a native human centromere the HOR regions are highly repetitive units typically 700-2000bp in length. The PWM provides the frequency of particular nucleotides at each position of the HOR region.
[0164] In the second step, the PWM was modified to insert binding sites for the recruitment of epigenetic modifiers, for example a LacO binding site for the binding of a LacI protein. The binding sites are short protein-binding motifs (approx 10-15bp long). As provided for herein, the binding sites for the recruitment of epigenetic modifiers were inserted into the HOR region sequence with regular spacing at positions computationally determined to be the least disruptive to the HOR region function. The epigenetic modifier binding sites were typically inserted every 100-200 bp over the span of the HOR region sequence. To maximize variation of the sequence, rare sequence variants from the native centromere sequences were used. An illustration of an exemplary HOR design with addition of epigenetic modifier recruiting regions is shown in FIGs.1, 2 and 3.
[0165] In the third step, the artificial centromeric sequence was constructed by linking together HOR regions constructed in the first and second step. However, in order for the artificial centromeric sequence to be synthesizable and stable in bacteria, the HOR regions cannot contain direct repeats of the exact same consensus sequence. Accordingly, for each successive HOR region, sequence divergence (which can be identified from natural variations) was introduced by using an iterative computational approach. First, a minimal number of changes were introduced to the HOR consensus sequence and the resulting sequence was computationally analyzed for its synthesizability and stability. Said computational analysis is performed on the entirety of the linked sequence, not just on individual HOR regions. Thus, for example, if an artificial centromeric sequence comprises 7 HOR repeats, all of the repeats are modified and analyzed simultaneously. If the new HOR sequence did not pass the computational analysis, additional changes were introduced to generate a more divergent sequence until the sequence passed in silico the synthesizability and stability test. Example 2: Generation of Human Artificial Chromosome
[0166] Human Artificial Chromosome (HAC) constructs were constructed by designing artificial centromeric sequences as described in Example 1, assembling the centromeric sequence with a backbone sequence and a cargo sequence in a 1 pot reaction using the Golden Gate Assembly system from New England Biolabs and was propagated in E. coli. A 52 IPTS / 200118216.1Attorney Docket No.: CETA-001WO schematic of the synthesis of centromere sequences and assembly of the artificial chromosome is shown in FIGs.4A and 4B, respectively. A schematic of the assembly is shown in FIG.5. Exemplary sequences of HORs, epigenic modifier recruitment sites, insulator sequences, and GFP cargo are shown in Table 1. Table 1 Exemplary sequences for centromere design SEQ ID Description Sequence NO: 1 LacO binding TGTGANCGNTCACA, wherein each N is, site independently, any nucleotide 11-17 Non-identical See sequence listing repetitive unit 18-24 HOR See sequence listing 31-32 Insulator See sequence listing Sequences
[0167] Five separate constructs were designed based on chromosome 12, 21, or 14 and repetitiveness thresholds of k=20 (no two identical 20-mers present in the entire centromere sequence) or k-50 (no two identical 50-mers present in the entire centromere sequence). For the purposes of the present disclosure, a HAC based on chromosome 12 having a repetitive threshold of k=20 is referred to as Ch12.k20. Likewise, a HAC based on chromosome 12 having a repetitive threshold of k=50 is referred to as Ch12.k50. This nomenclature extends to HACs based on other chromosomes. The stability of the HACs during bacterial cloning was assessed via whole vector sequencing of E. coli clones after the addition of the cargo sequences. As shown in FIG.6, the less repetitive k=20 constructs resulted in no sequence mismatches, while the more repetitive k=50 constructs exhibited a mismatch rate of approximately 1.0%. Follow up analysis using Sanger sequencing showed that the mismatch rate observed for the more repetitive HAC constructs were a result of sequencing artifacts. Thus, the results of the present example demonstrate that both the less repetitive k=20 and more repetitive k=50 constructs are stable through routine E. coli cloning and do not require specialized handling to add cargo. 53 IPTS / 200118216.1Attorney Docket No.: CETA-001WO Example 3: Generation and validation of stable HAC clones
[0168] The HAC constructs described in Example 2 were first assessed for their ability to express a GFP cargo in BJ-5 (HFF), HT1080, and HEK293 cell lines. FIG 7 shows representative images of each cell line transfected with the chromosome 12 k=50 construct. Each cell line readily took up the construct, resulting in intermediate or high levels of GFP expression. These data show that despite the relatively large size of the HAC constructs, cells are still able to take up the DNA and produce the cargo.
[0169] The long term stability of cargo expression using the HAC constructs was then determined. HEK293T, Ht1080, Jurkatt, MSC, and BJ5ta cells were transfected with the artificial chromosome constructs described above using a PEI transfection reagent. Constructs expressing the cargo genes without centromeric sequences were used as a control, as well as centromeric sequence containing constructs that lacked binding sites for epigenetic modifiers.
[0170] Cells were first selected for medium term retention of HAC DNA by subjecting the cells to antibiotic resistance for 1-2 weeks, depending on the antibiotic selection gene selected. HAC constructs were tested containing blasticidin, puromycin, geneticin (G418) and zeomycin resistance genes in all cells lines, with the exception of G418 in HEK293T cells, as these cells are already resistant to G418 by virtue of the neomycin resistance cassette on the SV40 largeT antigen vector integrated into their genome. After selection for medium term retention, single clones were sorted using FACs and expanded. Assessment of long-term stability of HACs in human cell lines via FACs
[0171] First, the isolated HAC constructs were allowed to propagate for 1 month and GFP expression was assessed via FACs analysis. The results for the BJ-5 (HFF) cell populations are shown in FIG.8. As shown, none of the constructs produced appreciable GFP signal in the absence of a helper plasmid containing an inducible HJURP-LacI epigenetic seeding factor construct. When the HJURP-LacI construct was co-transfected with the HAC, the chromosome 12 k=50 HAC construct as well as the chromosome 21 k=20 and k=50 constructs produced GFP expression above the level of the no HJURP-LacI control. These results indicate that, in the presence of the inducible HJURP-LacI, the chromosome 12 k=50 HAC construct as well as the chromosome 21 k=20 and k=50 constructs are able to produce stable GFP expression out to one month post transfection.
[0172] The isolated HAC constructs were then assessed for longer stability. Isolated HAC bearing clones were allowed to propagate for 3 months and clones were assessed for 54 IPTS / 200118216.1Attorney Docket No.: CETA-001WO maintenance of stable GFP expression via FACs analysis. The results of FACs analysis for select clones are shown in Table 2 below. Table 2 – FACs analysis of select GFP expressing clones Clone Percent GFP positive Percent GFP negative HAC GFP clone 1 94.87 1.37 HAC GFP clone 2 89.16 8.12 HAC GFP no epigenetic binding sites clone 1 9.25 87.82 No centromeric sequence GFP clone 1 0.19 99.43
[0173] As illustrated in Table 2 above, clones expressing the full HAC exhibited much more robust and uniform GFP expression than clones that either contained the centromeric sequence without epigenetic binding sites, or that lacked any centromeric sequence. These data show that HAC cargo genes are stably expressed in the cell for an extended period of time. Detection of HAC cargo and backbone element copy number via qPCR
[0174] Isolated clones were also screened for the copy number of HAC cargo and backbone elements via Taqman qPCR. RNAse P was used as an internal control. The assessment validates the presence of the HAC DNA and also quantifies the relative copy number of HAC elements. High copy number clones are likely to represent integration events and are screened out. Isolated clones were allowed to propagate for 1.5 – 3 months prior to assessment. As shown in FIG.9, the cargo (GFP) and backbone elements (AMP, BSRs2, and CMV) were all detected in three separate HAC transfected clones. The untransfected negative control shows no expression of any of the assayed elements. These data show that the HAC cargo and backbone elements are stably expressed in a cell for an extended period of time. Detection of HAC in clones via Southern Blotting
[0175] To test for inappropriate integration or rearrangement of HAC DNA within the isolated clones, genomic DNA was assessed via Southern Blot analysis. The genomic DNA was isolated and enzymatically digested. As a control, the HAC input plasmid was similarly 55 IPTS / 200118216.1Attorney Docket No.: CETA-001WO digested. As shown in FIG.10 several clones presented a faint banding pattern similar to the banding pattern observed in the HAC plasmid control (see boxes). However, detection of the expression of HACs in low HAC copy number may require more sensitive methodologies. Detection of HAC using Fluorescence in situ hybridization (FISH)
[0176] To identify HAC expression in low copy number, expression is validated using a FISH protocol. Cells are transfected as described above, clones isolated, and cells are cultured for 1.5 months. After transfection, mitotic cells are fixed in the metaphase stage on a coverslip.
[0177] DNA probes for FISH analysis are prepared using the FISH TagTM DNA Multicolor Kit, per manufacturer protocols. Once probes are designed and the slides prepared, chromosomal DNA is denatured at 72℃ for 2 minutes using Formamide / SSC. The slides are then dehydrated and air-dried. The probes, in combination with the hybridization buffer, are also denatured at 72℃ prior to applying the probe solution to the slide. Cells are then counterstained for DAPI and imaged.
[0178] As shown in FIG.11, clones transfected with HAC constructs demonstrated strong and localized FISH signal in the nucleus of the cell as expected. In contrast, untransfected cells or cells transfected with DNA constructs not containing the artificial centromeric sequences did not display any detectable FISH signal in the nucleus of the cells. This data demonstrates that FISH may be used to validate successful and stable expression of HAC constructs despite low HAC copy number. Detection of HACs using Rolling Circle Amplification
[0179] Rolling circle amplification is an isothermal DNA amplification technique that amplifies circular DNA more efficiently than linear DNA. The technique relies on a highly processive DNA polymerase (phi 29) which has two useful properties: 1) it can replicate DNA for tens to hundreds of thousands of base pairs before terminating, and 2) it can displace double-stranded DNA to continue DNA replication. Phi 29 can replicate circular DNA several times through these properties, leading to high replication compared to linear DNA. These amplified DNAs are large concatemers of the original piece of DNA. This DNA product can be sequenced using various techniques or digested with restriction enzymes to visualize on a gel. If a band corresponding to the input HAC DNA is present, then that cell line contains a HAC. If not, the cell line either does not have a HAC or the copy number is below the assay’s detection limit. 56 IPTS / 200118216.1Attorney Docket No.: CETA-001WO
[0180] In brief, genomic DNA was isolated using a commercially available kit via manufacturers protocols. The genomic DNA was treated with an exonuclease to remove the linear genomic DNA. 1µL genomic DNA was combined with 7µL of RCA reaction mix A (Phi 29 buffer, deoxynucleotide triphosphates, magnesium chloride, SSB, trehalose, PEN N, primer mix) and the mixture was heated to 95℃ and slowly cooled. 2µL of RCA reaction mix B (Phi 29 polymerase, inorganic phosphatase, recombinant albumin) was added to the reaction mixture, and the mixture was incubated at 30℃ for 16-20 hours. The reaction mixture was then digested with HAC specific restriction enzymes and run on an agarose gel and visualized.
[0181] As shown in FIG.12A, two samples containing HAC DNA clones were clearly visualized on agarose gel with a distinct and specific digestion banding pattern. In contrast, the genomic DNA isolated from clones not containing HAC DNA displayed digestion specific banding. The technique is also able to detect very low levels of HAC DNA. As shown in FIG.12B, HAC specific restriction digest banding was observed down to 4 pg HAC template per 10 ng total DNA. These results demonstrate that the RCA technique can readily confirm the presence of stable HACs even at low copy number. Nanopore sequencing of select HAC clones
[0182] Select clones harboring either drug resistance (DR) or GFP HAC constructs were sequenced via Nanopore sequencing. The libraries were prepared using a version of the Cas9 targeted enrichment protocol (Gilpatrick et al., (2020) NATUREBIOTECHNOLOGY, PMID: 32042167) using 4 different bioinformatically-optimized guide RNAs for two regions in the HAC backbone. The results are summarized in Table 3 below. Table 3 - Results of the analysis of selected HAC clones DR clone 1 DR clone 2 GFP clone Total # of reads 81.55 k 34.58 k 18.81 k Median read length 23.75 kb 19.92 kb 25.85 kb Total # of bases sequenced 623.74 Mb 233.2 Mb 173.98 Mb # of reads aligned to HAC 29 26 1191 # of reads spanning the entire HAC 8 16 57 IPTS / 200118216.1Attorney Docket No.: CETA-001WO # of hybrid reads, aligning to both HAC 1* 1* and genome # of reads showing rearrangements 1* 0 * artifact
[0183] In the DR cell lines, the number of reads that aligned to HAC shows that the enrichment protocol works (given 0.1-0.2x average genome coverage, the enrichment for HAC in the library is >100-fold). Among the HAC-aligned reads, reads were found that start at each of the 4 used guide RNAs, indicating that the guide RNA design was efficient. Only 3 reads in total from 2 clones aligned to both the HAC and the human genome or to the HAC in a non-linear manner, and visual inspection of these alignments clearly showed that these were artifacts of the mapping. No evidence was found for the integration of the HAC into the genome. Likewise, the number of reads that span the entire HAC provides confidence that the DR HAC does not undergo any rearrangements in the human cell lines.
[0184] A second nanopore sequencing experiment of human foreskin fibroblast cells transfected with the GFP HAC was performed, using the same library preparation protocol. The number of GFP HAC-mapped reads in this experiment was higher than in the experiments with DR HAC-transfected cells (>1000 versus 26 and 29 in DR), suggesting either a higher copy number of GFP HAC or higher efficiency in targeted sequencing library preparation. Furthermore, a more complex pattern of read mapping was observed. The majority of the reads mapped once onto the HAC; however, GFP-mapped reads that are longer than the GFP HAC were also detected (170 of 1191 reads were longer than the GFP reference map marked by the vertical line on the histogram in FIG.13A). Further visual inspection revealed that these long reads map onto the HAC multiple times and in a complex way, indicating rearrangements and multimerization of the HAC, an example of which is provided on the line plot in FIG.13B. In addition, as in the DR HAC experiments, hybrid reads that would indicate the integration of the HAC into the genome were not observed. Taken together, the results show that GFP HAC exists in the cells not integrated into the genome, in multiple copies, and at least partially multimerized into a large HAC composed of several copies.
[0185] To expand on the second sequencing experiment, each of the five stable BJ-5 (HFF) clones described above (chromosome 12 k=50, chromosome 12 k=20, chromosome 22 k=50, chromosome 22 k=50, and no centromere control) were analyzed via nanopore 58 IPTS / 200118216.1Attorney Docket No.: CETA-001WO sequencing. The results are shown in FIG 14. FIG.14A shows representative nanopore sequencing reads. Intact HAC reads are shown to map consecutively onto HAC coordinates. In contrast, rearranged HAC reads map to two or more sections of the HAC that are non- continuous in the seeding template, either with or without other genome elements interspersed, while integrated reads transition from HAC to other genomic locations. FIG. 14B illustrates the HAC structural integrity in the BJ-5 clones for synthetic centromeres one month after transfection. As shown, the k=50 constructs demonstrated the highest proportion of intact reads. The results of this experiment suggest that the more repetitive k=50 HAC constructs may be more capable of forming stable HACs with lower levels of integration or rearrangement. These results were surprising and unexpected based on previous studies which have shown high rearrangement levels within the HAC seed sequence and between the HAC and the endogenous genome. However, fewer than 10% of reads showed HAC rearranged with genome element, and overall, relatively rare (<25% of reads) recombination between different areas of the HAC, as most reads were uninterrupted, long stretches that map back to the reference sequence of the HAC. Cargo expression over time
[0186] Based on the nanopore sequencing data above, the HAC construct chromosome 12 k=50 was selected for further assessment of cargo GFP expression. Several clones (C1-C6) supporting high GFP expression were selected for additional analysis, and were compared to GFP expressing clones from the no centromere control group. As shown in FIG.15A, the amount of GFP expression remained relatively constant over the course of 16 weeks, and the percent of GFP positive cells declined only slightly in some HAC clones. qPCR was utilized to assess whether the loss in percent GFP positive cells was due to loss of the HAC or integration events that became silenced through position effect variegation. As shown in FIG.15B, all HAC clones assessed had minimal HAC gene copies per cell, with only about 4-50 copies of HAC sequence per cell. This indicates that loss of the HAC gene is more likely the cause of the decrease in percent GFP positive cells over time. Nanopore analysis of the selected clones revealed similar results to those described above, with the majority of reads showing full length intact reads (see FIG.15C).
[0187] These results indicate that the artificial chromosomes of the present invention are capable of maintaining expression of a nucleic acid cargo in cells without significant loss of the artificial chromosome or silencing. 59 IPTS / 200118216.1Attorney Docket No.: CETA-001WO INCORPORATION BY REFERENCE
[0188] The entire disclosure of each of the patent and scientific documents referred to herein is incorporated by reference for all purposes. EQUIVALENTS
[0189] The invention may be embodied in other specific forms without departing from the spirit or essential characteristics thereof. The foregoing embodiments are therefore to be considered in all respects illustrative rather than limiting on the invention described herein. Scope of the invention is thus indicated by the appended claims rather than by the foregoing description, and all changes that come within the meaning and range of equivalency of the claims are intended to be embraced therein. 60 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQUENCE LISTING SEQ Description Sequence IDNO:1 LacO siteTGTGANCGNTCACA, (N=any nucleotide)2 TetO siteTCTCTATCACTGATAGGGA3 LexO siteWACTGTATAWAWAWMCAGYAM (W is either T or A, M iseither A or C, Y is either C or T)4 UAS siteCGGAGGACTGTCCTCCG5 Exemplary GTTTGCATTTGGAGACTTCCATCGTTTTGTGACAAAATGTACAAAACG artificial AGACATCGTCGAATAGAAAGTAGGCAGTATCACTCAGAGACACTACTC centromere (NA) TGTCATTGTGATCGTTCACAAAAAGAATTTCACATTTCCTTTCATAGA ACAGGTTAGAACCAGTCCGTTTGAAGTTTGTAAACACATACTTTGACC GCTTCGATGCCATCGATGGCAACTGGTTTGCTTTATGTAGTGATTAAT AAGACAAATCTAAGAAACTCCTTCGTCCTTGTGATCGTTCACAAATTG ATTTTAACATTCATTTACACGATCACATTTTCAACCCCATTTCTGGGG ATTTCCATCTAGAGGTTTAAAACGGTTCGAGGTCTGCGTTAGAAAAGA AAGCATGTTCTACAATATATAAACGGAGTCTTTCAAAGACACATCCTT TCGATTGTGACGTTCACACAAAGACTTTAGCCATTCCTTTCATTGAGT AGTTTGGGAAAAACACTCTTGTATTCCGCATGTAGACATTTTGACGTC CTTGTGGTCTTAGTTAGATACCGGGTTGCTTTAACTAAAGTACGGCAC AAGAATACTCAATAAGTTACTTGCGGGGTATGTGTTCACCTCCAGGGT GGACCCCTCTTTGTGACCGGTCACATTCAAATACTCTACTTGGGCGTC TCCTGTAGGAGAATTCTATAGCTTGGAAACGAATTGAAGGAAAGCAAA TATATTCATACAACAACAAGACAAAATCGTTGTCAAAAAGTATTTGGT AATTGTGATCGTTCACACACGGGGTTAAAGATTTCATTTGATGGAGCA GCTTGGAACCAATCTCTCGTAATCTACAATCACATATATGACTTCTCT GGGGTCTACGTTAGAACAGGTTTTTTTCGTATAATGCAAGAAAGAAAA ATATTGGGTACGTTTTGGTGCTGCTTCTTTTCACCTCGCAGAAGTTAA CAGTGCTTTGTGAACGATCACAGTGATACCTTCATTTGGTATTTTGTA GTTGGAGGTTTCAGCACTTATAAGCCACATATAGGAAAGGTAACATTT TCATAGAAAAACCAGTCACAAGCACTCTGAGAAACCACATTGAGATTG TGATCGTTCACACACACAGTATAAGCTTCCTGTTCATTGAAGAGTTGG GGGACCCTGTCCTTGAAAGTCCGCATGTAGATATATGGTCCTTTTTGT GGTCTTCGATGGAAACTGGTTTTCCGCATGTAAAGTTAGACAGTAGAA TTATCATTAAGTTAATTGCGGCGTCTGTTTTCAATTCACAAAGATGAA CATTCCTTTGTGAGCGGTCACATTCAAAAACTATTCTTGAGGATTTCC CTGGGGAGATATCAAGCGCATTGAGGCCAGAGCTAGCAAAGCAAATAT CTTTGTGTAACAACGAGAAAGAAACATACACGGAAACTTCTTTTTGTT TGTGATCGTTCACACAGAAGGTTATCCTTCCTATTGGTGTAGGAGTAT GAAAAAACACTCTCGTAATCGGCGAGTAGTTATATGGTCCTTTTTCAG TCCTTTGTTAGAATCGCGAGTTTTTCCTACAACGTATGACAGAAGTAG TGTCTGTAAATTATTTCTGATTGTGATCGTTCACACGTACAGCTGGAC TCTCCCTTAGGACAGTAGGTGGTAATCAGCCATTTAGTGATTTGAAGG TGTAGAATTCTAGTGCATTGTGGCGTATGGTAAAAAAGGGAATATCAT CTATGAAGTCCAGTCATAAACACTCACGGAATCTACTCTTTCATTGTG ATCGCTCACACACAAAGCTTAAACTCTCTGTTGTTGAAGCCGGTTTGA AACCCTGTGATTTATGCCTCCAACTGAATAATTGAACCGCTTGGACGC GTTCATTGCAACCGAGAATTTTTCCAGAAATTTTTGAGAGAATAATTT 61 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: TCAGCAACCTATGTGTAGTATGCGTATTGAATTCAAAGATTTCAAACT CCCATTGTGAACGTTCACATTGATACTCCGTATCTGTCATTTGCACTT AGAGATTACAGTCACTTTCAGGCCCAACGTGGAAACGGGAACTTCTAC GCATTAATACTGGACAGCATCACTCACAGGAACAACATTCTGTTTGTG ACCGATCACACAAAGATTTACGCTCTCTGTTCTTAAAGTTGTTTGCAA AAACTTGCCTTAAGCTGCCAGGAGGTATTCGGGCCTATTAGATGCATT AGTTGTAACGTGACTTCCTCAGAGCACACTGGAAAGAAGAGTACAGAC TAACTTCCTTCTGTTACCGCTACTCAAATCATAGAGTTGATCTCTCAT TTGTGAACGGTCACATTGAAATCCCCTATTTCTGTATCTGCCGGAGGA GACTTCAATCGTTTTGAGACCAATTGTGGAAAAAGATATGTCCTCGCA TGAAAACTGGAGAGATGATACTCATAAACTTCTTGGTAATTGTGACCG TTCACACAAGGGTATAACATTTGTTATGGTGAAGCAGTTTCGAAACAC TCTCTTGTAATCTCCAATTGGACATTGGGATCTCTTCGACGCCTTGGT TAGAAACGAGATTTTCTCCTACAATTTTACAGAGATGAATTCCCAGGA ACTGATTCGTTGTTTGCGTATTGAACTCAAAGGGTTGGACCTTACTTT GTGAACGATCACATTGTAACCCTCATTTGGTGAGTTCCAGGTGGATAT TTGAATGGCTTTTAGAACAAGGGTTGAAACGGGAACGTCTTCTTATGA AATCTGGACGGAATAATTCTCAAAAACTAGTTTGAGATTGTGATCGCT CACACACGGTGTTCAATCTTTGTTTGGACGGTGCTGTTGGGAGACGCT GTGACTTTAAGACTCCAGGCGGACATTAGGACTTCGTTGGGGGCTTCT TTGAAAATGGCATATCGTCACATGATATTAGATGGGCGAGGTCCCACT AACGTCTTGGGCGTGTGATCGTTCACACACAGGGTAGAATTTTACTTC AGACAGGGCGGAGGTAAAAGACACTGTTTATGAATATGCTGCCGGACA TTACAGGCACTATGAGTCCAACTGTAGAAATGGAGACAACTTTTCTAC AAACTTGAGAGCATAATTAACAAAAAATTGTTTCTGAATGTGATCGTT CACAAACAGGGTTTGACATTTGTTCTGACGGGGCATTCTGCAAACAAT CGGTTGAAGTATGGAAGCGGTTATTAGGAGCTGTTTCAGACCGTCGAT GGTAAAGGTATCTCTACATGTAGTGGTCAACGGAAGGATTCTCCGTAC CTTCTTTATGCTGAGTGCATTCCACACACACAGCTGTACTTTTCTATG TGAACGGTCACATCGAAGCAGCCTTTTTCTGAGCTTCTAGGTGCAGAT TTGAAGCGTTTGGGACAATATATAAAAAGAATCACCTTTGTGTATAAA TTACACAGAATTATACTCCGAGACTTCTGTGCGACTGTGACCGTTCAC AGAAGAAGTATAACCTTTGTTATCACAGACTACTTTGGATACCCTCGT TGAAGTGTGCAGGCGGATCTTTAGAGCTCATTAAGACCGTCATTGGAA AGGCATGTCTACATGGAGCGTTAGAAAAAAGAATCCTCAGGAAGTACT ATGTATTGTCTGTATACAACACACAAAGGGGAAGTGCCCATTGTGAAC GGTCACAGCGAAACACTGTTCTTGGGAATTTCCACGTGCAGATTTGAA GTGCTCTTTGGTCAAACGTATAAAAGAAACTAACTTTGTGTAAAAAGT ACACGGATCGTTCACAGAAAGTATTTTCTGAATGTGACCGTTCACACA CGGAGGATAACCATTCCTTCGACGGCGGAATTTGTAGAAACTGTGTTT CTAAGACTGTAAGCGGATACTTGTACGTCTTTAAGACCTTCATTGGAA ATGGAATTTCTTCACATACTGTTCCACGGAACAATTCTTAGTTACTTC TTTCTGATGCGTTTATTCATCTCACTGATTTGAAACTTCGTTTGTGAA CGGTCACAATGAATCACACTATTTTTGTTTGCCTTGGGAGACATCCAG CGTATTGTGGCAAGATCTACCAAACCAGATATCGTTGAGTAGCAAGGA GGAAGTAACACACAGGGACACTTCTCTTTCTTTGTGATCGTTCACAAA AAAAGTTCTCATTCCCATTCGTATAAGAGGATAAAACAAGACCCTTGA TTGGTGAATACTTACATTGTCCGTTTCCATTCCATTGATAGCATCTCG TGTGTTTTCTGCAGCGAATAACAAAACTAATGTATGAAAATCATTCCT CATTGTGATCGTTCACAAGTTCATCTTGACACTCACTTACGCCATTAC GTTGTCATCCGCAATTCAGGGTTTCAATGTATAGGATTATAATGGATC 62 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: GTGGTGTGTGTTAAAAAAGAGAGTATGATCACGATGTACAATCGTAGA CTCTCAAGGACTCAACCCTTCCATTGTGACGCTCACACAAAAACCTTA GACACTCCGTTCTTTAAGTCGTGTGTGAAAACCAGTCATTTCCCCCAT CTAAACAATTTAACGGCCTGGTCGTGTTAATTACATCCCAGGATGTTT TCACAAAATTATGGGACAATAATATTCAACAAGCTACGTGCAGGATAC GTGTTGACTTCAAGGTTGCACACCCCATTGTGACCGTTCACATTCATA TTCTGTACCTGGCTCTGCTCTAAGAGAATACTGTAACTTGCAAGCGCA TCGAGGGAACGCGAATTTATACACACTACTACAGGACAACATCGCTGA CAAGAAGAATATGCTATTTGTGATCGATCACACACAGGTTAACGATCT CAGTTGTTGAAGCTGCTTGCAACAAATTCCCTACTACCATGACGTATA CGGCTTATCAGGTGTATAAGTTATACATGTCTTTCTCGGATCATACAG GAAAGAAAAGTATAGGCTACCTTCTGCTGCTACTGCTTCTCACATCGT AGAATTTATCACTGATTTGTGAACGATCACATTGATATCTCCAATTGC TTTCTGTCGTAGGAGGCTTCATCATTTAGAAACCACTTATGGGAAAAG TTACGTTCTCACAGGAAAACCGGTGACAGGACACTGATAAACCTCATG GAAATTGTGATCGTTCACACAACGGTATAAGATTCGTGATCGTTAAAC AGTTGCGGAACCCTCTCCTGAATCCCCATTTAGACATAGGGTTCTTTT CGTCGTCTTGGATAGAAACTAGTTTTTCGCCTGCAAATTTAGAGAGTT GAATTACCATGAAGTGAATCGCTGCTTCCGTTTTGAATTCAAAAGGAT GGACATTACTTTGTGAGCGATCACATTCTAAACCTAATCTGGAGGTTG CAGTTGGATACTTGCATGGTTTTTTGAAAAAGTGTTCAAACCGGGACG TCGTCTAATGGAATGTGGGCGGTATAACTCTGAAACACTAGTCTGACA TTGTGATCGCTCACAAACAGTATTCCATATTTGCTTGCACAGTACTGG TGAGAGCCGGTGCGATTTAGATTCTAGACGCACACTATGACTGCGTCG GTGGCATCTATGACAATTGCTTAGCGTTACGTGGTAATAAATGAGCCA GATCCAACAAACGCCTCGGCCGTGTGATCGTTCACAAACTGGTTATAA TATTAATTCACACCGGTCGCAGTTACAAGCCAATGTCTAGGATATCCT TCCAGACGTTAAAGACAGTACGAGTTCAGCTTTAGAAATGAAGGCAAG TTTCCACTAAATTAAGGGCGTATTTAAAAAACAAATGCTTCCGAATGT GACGTTCACAAAAAGGCTTTGGCAATTGCTCTCACTGGGTATTTCGGC AAAAAAACGCTGATACGGATGCAGTCATTATGAGGTGCTTCTGATCGT AGATAGTTAACGTGTCGCTATATCTAGAGGACAGCGCAAGGATACTCC ATACGTTCCTTACGCGGAATGCGTTCCCCACCACGGCGGTCCTCTTTA TGTGACCGGTCACATCCAAGTAGTCTTCTTCGGGCCTCTTGGAGCAGA ATTGTAGAGTTGGGAAGATTAAATGAAAACAATTACATTTATGCATCA ATAACACAAAATTGTAGTCCAAGAGTTTTGGGCAACTGTGATCGTTCA CAGACGAGGTAAAACATTTGATATGACGGACCACCTTGGATCCCATCC TGATGTACAGTCGCATCTATAAGTTCACTAGGATCGACATTAGAAAGC TTGTTTACGTGTAGTGTAAGAAAAAAAAATCTTCGGGACGTATAGGTA CTGTTTGTTTACACCACGCAAAAGGTAAGAGCGCATTGTGAACGATCA CAGCGATACATTGATCTGGGATTTCTACTTGCAGGTTTGAGTACTCAT TAGTCACACATATGAAAGATACCAATTTTATGGAAAAAGCACTCGCAG CGCTCAGAGAAAGCATATTCAGAATGTGATCGTTCACACACGCAGGAT AAGCATCCCGTCCACTGCAGAATTGGTGGAACCTGTGCTTCAAGACCG TATGCAGATACATGTTCGTTTTTATGATCTTCAATGGAAATTGATTTT CTGCACGTACAGTTCGACGGTACAATTATTATTTAGTTCATTCCGACG CCTTTTTTCATTTCACTAATATGAAAATTCGTTTGTGAGCGGTCACAA TCAATAACAATACTTTAGATTCCCGGGGGATATATGAAGGGCATTTAG GACAGGGCTTGCAACGCGAATGTCTTTTTGTGACATCGGGAAGGAAAA ATACTCGAAAACTTGTTTTAGTTTGTGATCGCTCACACCGATGGTCAT TCTTCGTATGGGCGTTGGTGTAGGAAGAAGCAGTCACTTAACGCCGGG 63 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: TGGTCATAAGGTCTTTGTTCGGTGCTTTTTTAAAATTGCCAGATTGTC CCACGACATAAGACGGACGTGGTGCCTCTAAAGTATTCGGAGTGTGAT CGTTCACACGCACGGCAGGATTCTACCTCAGGCACGGTGGGGGGAAAT GAGACAGTTAATATATGATGGCGTACAATACTGGTACAATGTGTCGAA TTGTAAAAATGGGGATAACATTCTGCAGACCTGTGATCAAAACTAACG AAATATAGTCTCTCAATGTGATCGCTCACAAACAAGGCTTGAAATCTG TGCTGTCGAGGCCTGCTTCAAACCATGGGATAGCATCGAACCGATTAA TAGAAGCGGTTGCACACGGTCAATGCTACAGATAACTTTACCTGAAGT TGTTAAGGGAATGATTTTCCGCACCCTCTGTATACTAAGCGCATTGCA TACAAACATCTCTAATTCTCAATGTGAACGTTCACATCGATGCTGCGT TTCTCTACTTGTACGTACAGATTAGAGGCATTTCGGGCCACATGTAAA CAGGATCTCCTATGCGTTTATATTGCACAGCATTACACACCGGGACAT CAGTCCGTCTGTGACCGATCACAGAAAAATATACCCTCTGTGATCTCA AACTTCTTTGCATAACCTGCTAGGTGCCGGGGGGTCTTCAGGGCTAAT AAATACAGTAATTGTAAGTCACGTCCACAGGGCGCATTGGAAAAAAGA GTCCACACGAACTACCATCTATTATCGGTACACAAAACATAAAGTGGA TGTCCCAATTGTGAACGGTCACATCGAAATACCGTACTTCGGATCTCC CCGAGCAGACTTGAATTGTTCTGTGATCAATCGTGTAAAAAAATCTGA CCTTGCGTGAAAAGTGCAGGGTGGTACACATAAAGTTTTTGCTAAATG TGACCGTTCACACAGGGGGATAACAATTGCTACGGCGACGCAATTTCT AAAAACTCTGTTCTAACTCTAATCGGACACTGGTATGTCTTCAACACC TTGATTAGAAATGAAATTTTTTCCCACACTTTTCCAGGGATCAATTCC TAGGTACTGCTTCCTTATTCGCTTATTGATCTCAATGGTTTGGAACTT AGTTTGTGAACGATCACAATGTATCCCACAATTGTTTTGCCGTGGGAT ACATGCAGGGTATTTTGGAAAGGTCTTCCAACCCGGATGTCGTTTAGT GGCATGGGGGAGGTAAAACACTGGAACACTTGTCTTACTTTGTGATCG CTCACAACAATAGTCCTTATTCGCATGCGCATTAGTGGAGAAAGCAGG AGCCATTATGCTGGATGCTCACAATGTCTGTGTCCGTTGCATTTATAA CATTTCCTGAGTGTTCCGCGGCAAAAAACGAACCTGATGCATCAAAAG CATCCGCAGTGTGATCGTTCACAAGCTCGTCATGATACTAACTCACGC CCGTTGCGGTGACATGCGAAAGTCAAGTATCATTGCATACGATAATGA TAGAACGTGTTGAGTTTTAAAAATGAGGGTAAGATCCGCTGAACTATG GTCGAATCTAAAGAACTAAAGCCTCCCAATGTGACGCTCACAAAAAAG CCTTGGAAACTGCGCTCTCTAGGTCTTGCGTCAAAACAAGGCACACCG ATCCAATCAATATAAGGGGCTGCTCATGGTAAATACTTCACATGACGT TATCTCAAGATGATAGGGCAATGATATTCCACACGCTCCGTACACGAA ACGCGTTGCCTACAACGTCGCTCATCCTAATGTGACCGTTCACATCCA TGTTGTGTTCCTCGCCTGTTCGAACAGAATAGTGGAATTGCGAGGCTC AAGTGAACACGATTTCATATACGCTTCTATAGCACAACATTGCAGACC AGGAGATTAGGCCATCTGTGATCGATCACAGACAAGTAAACCATCTGA GATGTCGAACCTCCTTGCATCACATCCGTACCGTGGCGTCTACAGGTT AACAAGTATAGAAATTATAATCTCGTTCACGGGTCGTATAGGAAAAAA AAGTCTACGCGACCTACAGCTACTATTGGTTCACACAACGTAAAATGT ATGACCGAATTGTGAACGATCACATCGATATATCGAACTGCGTCTCTC CTAGCAGGCTTGATTATTCAGTAATCACTCATGTGAAAAATTCCGATC TTACGGGAAAAGCGCTGGCGGGCACAGATAAAGCTTATGCAAAATGTG ATCGTTCACACAGCGGGATAAGAATCGCGACCGCTACACAATTGCTGA AACCTCTGCTCAACCCTATTCAGACACAGGTTTGTTTTCATCATCTTG AATAGAAATTAATTTTTTGCCCGCACATTTCGAGGGTTCAATTACTAT GTAGTGCATCCCTACTCCCTTTTTGATTTCAATAGTATGGAAATTAGT TTGTGAGCGATCACAATCTATACCAAAACTGTA 64 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: 6 Exemplary GTTTCCAGGTGGAGATTTCAATGGCTTTGAGGCCAAAGGTTGAAAAGG artificial AAACATCTTCTTATAAAATCTAGACAGAAACATTCACAGAAACTTCTT centromere (NA) TGTGATTGTGATCGTTCACACAGGAGGTTAACCTTTCTTTGGATGGTG CTGTTTGGAAAAACTGTGTCTTTAAGTCGGCAAGCGGATATTTGGACC TCTTTGGGGCCTTCGATGGAAATGGGATTTCTTCATATGATGTTTGAT AGAAGAAGTCTCAGTAACTTCTTTGGGCTTGTGATCGTTCACACACAG ATTTGAACTTTCCTTTAGACAGAGCACATGTTAAACACCATTTTTGTG GATTTGCAGCTGGAGATTTCTAGCGCTTTGTGGCCTATGGTAGAAAAG GAAATATCTTCTACAAAATCTTGACAGAATCATTCACAGAAACTTCTT TTTCATTGTGATCGTTCACACACAGAGTTTAAACTTTCTGTTGATGGA GCAGGTTGCAAACACTGTGTTGAAGTCTGCAACTGAATATTTGGAGCT CTTTCAGGCCTTCGTTGGTAACGGGATCTCTTCAACTAATGTTTGACA CAAGAATTCTCAGTAAGTTATTTGTGGTGTGTGTATTCAACACACAGA GTTGACCCTTCCTTTGTGAACGGTCACATTGATACACTCTATTTGTGA GCTTCCAGTTGGAGATTTCAAGCGCTTTCAGACAATTGTAGAAACGGA AACACCTTCATATAAATACTAGACAGCATCATTCTCAGAAACTACATT GTGATTGTGACCGTTCACACACGGAGTTTAAGCTTTCTTTTGATAGAG CAGTTTGGAAACACTCTGTCGTAAGTGTGCAAGCAGATATTTGACCTC TTTAAGGCCTTCATTGGAAACGGCATTTCTTCATATAACGCAAGAAAG AAGAATACTGGGTAAGTTCCTTGTGCTGCCTCTATTCAACTCACAAAG GTGAACAGTCCTTTGTGAACGGTCACAGTGAAACCCTGTTTTTGTGAT ATTTCCAGGTGGAGACTTCAAGCGCTTTGAGGCCAAATGTAGAAAAGG AAATATCTTCATATAAAAACCAGACAGAATCATTCTCAGAAACTATTT TGTGATTGTGACCGTTCACACACAGAGTATAAGCTTTCTTTTGATGGA GGAGTTGGGAGACCCTGTCTTTGTAAGACTGCAAGTAGATATTTGGAC CTCTTTAAGGCCTTCATTGGAAACGGGATTTCTTCATGTAATGTTAGA CAGAAGAATTCTCAGTAACTTATTCGTGGTGTCTGTATTCAACTCACA GATTTGAACCTTCCTTTGTGAACGGTCACATTCAAACACTCTTTTGGT GGATTTCCATTTGGAGATTTCAATCGCATTGAGACCAAATGTAGAAAA GCAAACATCTTCGTATGAAAACTGGACAGAATCATTCTCAGAAACTAC TTTTTGATTGTGATCGTTCACACACGGAGTTTCACCTTTCTTTTGATA GAACAGTTGGGAAACAGTCTCTCTGAAGTCTGCAGGCAGACATTTGGA CCTCTTTGAGGCCATCGTTGAAAACGGGATTGCTTCATATAATGTTAG ATAGGACAAGTCTCAGTAACTTCTTTGTCCTTGTGATCGTTCACACAT TGAGCTGAACTTTCCTTTAGAAGGGCAGATGTAAAACACCCTTTTTGT GAATTTGCAGGTGGAGATTTCAAGTGCTTTGAGTCCTACTGTAGAAAA GGAAACATGTTCTATAAAGTCTAGACAGAATCATTCACAGAAACTACT TTTTGATTGTGACGTTCACACACAGAGTTTAACATTTCTTTTCATGGA GCAGTTGGGAAACAATCGGTTGTATTCTGCAAGCGGTTATTTGGACGT CTTTGAGGTCTTCGTTGGAAAAGGGATTTCTTCCAGTAGTGTTCGAGA GAAGAATACTCAGTAACTTATGTGTGGTGTGTGTATTCAACTCAAAGA GTTGAACCTCCCTTTGTGAACGGTCACATTGAAACTCCGTATTTGTGC GTTTGCAGTTGGAGATTTCAATAGCTTTGAAACGAAACGTAGAAAAAG AAACATATTCGTACAAAAACAAGACAGAATTATTCTCAGAAACTACTT TGCGATTGTGACCGTTCACACAAGAAGTTTAAGCTTTCTTTTCATGGA GTACTTTGGAAACACTCTGTCTTAAGTCTGCAAGGAGATATTTGGAGC TCTTTGGGGCCTTCGTTGGAAAAGGGATTTCTTCAGAGAGCGCTGGAA AGAAGAATACTGACTAAGTTCTATGTGTTGTCTCTATTCAACTCACAG AAGTGAACTCTCCTTTGTGAACGGTCACAGTGAAACCCTCTTTTTGTG TATTTGCACGTGGAGATTTCAAGTGCTTTTTGGCCAAATGTAGAAAAG GAAATATCTTCGTATGAAAACTAGACACAATCATTCTCAGAAACTACT 65 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: TTCTGATTGTGACCGTTCACACACAGAGTATAACCTTTCTTTTGATGG AAGAGTTTCGAGACACTCTCTTTGTAAGTCCGCAAGTGGATATATGGA CCTCTTTGTGGCCTTCGATGGAAACGGGATTTCCTCCTATAATTTTAC ACAGAAGAATTCCCAGTAACTTATTTGTGGTTTGCGTATTCAACTCAC AGAGATGAACCTTCCTTTGTGAACGGTCACATTGAAACCCTCTTTTTG AGGAGTTTGCATGTGGAGATTTCAAGCGCTTTTAGACCAAAGCTAGAA AAGGAAATATCTTCGTATAGAAACTAGAAAGAATCATTCACAAAAACT ACTTTGTCATTGTGATCGTTCACACAAAGAGTTTAATCTTTCTTTTGA TGTAGGAGTTTAGAAACACACTGTTTGTAATCTGCAAGTAGATATTTG GACTTCTTTGAGGCCTTCTTTGGAATCGGGATTTCTTCCTATAATGTT TGACAGGAGAAGTCTCAGTAACTTCTTTCTGCTTGTGATCGTTCACAC GTAGGGTTGAACTTTACTTTAGAAGAGCGGATGTTAAACACACTTTTT GTGAATTTGCAGCTGGAGATTACAAGCACTTTGAGGCGTACGGTAGAA ATGGAAACATCTTTTCTAAAATCCAGACAGAATCATTCACAGAAACTT GTTTTTGATTGTGATCGTTCACACACAGGGTTTAACCTTTCCTTTGAC GGAGCAGTTTTGAAACACACTCTTTTATGTCCGCAAGTAGATATTTTG ACCTGTTTGAGGCCTTCATTGGAAACGGTATTTCTTCATGTAATTTTC GACGGAAGAATTCTCAGCAACTTATTTGTGGTGTGTGTATTCACCTCA CAGAGCTGAACCTTCTTTGTGAACGGTCACATTGAAACAGCCTTTTTG TGCATTTCCAGGTGGAGATTTCAATCGCTTGGAGGCCAATATAGAAAA GGAAATATCTTTGTATAACAACTGGACAGAATCATTCTCAGAAACTAT TTTGTGATTGTGACCGTTCACAGAAGGAGTTTAACCTTTCTTTTCATA AAGTAGTTTGGAAACACTCTGTTGAAGTCTGCAAGCAGATATTTAGAC CTCATTGAGGTCTTCGTTGGAAACGTGATTTCTTCATGGAATGCTAGA AAGAAGAATACTCAGTAACTTCTTGGTGTTGCCTCTACTCAACTCACA GAGTTGAACTGTCCTTTGTGAACGGTCACAGTGAAACACTCTTTTTGT GAATTTGCAGTTGGAGATTTCAAGCACTTTTAGGCCAAATGTAGAAAA GGAAATATCTTTGTATAAAAAGTAGACAGAATCGTTCTCAGAAACTAC TTTGTGATTGTGACCGTTCACACACAGAGGATAACCTTTCTTTTGATG GAGCAGTTTGGAAACACTGTGTTTGTAAATCTGCATGTGGATATTTGG ATCTCTTTGACGCCTTCGTTGGAAATGGGATTTCCTCACATAATGTTC CACAGAAGAATTCTCAGTAACTTAATTGTGGTGCGTTTATTCAACTCA CAGAGTTGAACCTTCCTTTGTGAACGGTCACAATGAAACACACTTTTT GTGTTTCCAGTTGGAGATTTCAATGGCATTGAGGCCAAATGTTGAAAA GCAAACATCTTCTTATGAAATCTGGACAGAAACATTCTCAGAAACTTC TTTTTGATTGTGATCGTTCACACCGGAGGTTCACCTTTCTTTGGATAG TACTGTTGGGAAAAAGTGTCTCTTAGTCGGCAGGCGGACATTTGGACC TCTTTGGGGCCATCGATGAAAATGGGATTGCTTCATATGATGTTAGAT AGAACAAGTCTCAGTAACTTCTTTGGCCTTGTGATCGTTCACACACTG ATCTGAACTTTCCTTTAGACAGGGCACATGTAAAACACCATTTTTGTG ATTTGCAGGTGGAGATTTCTAGTGCTTTGTGTCCTATTGTAGAAAAGG AAATATGTTCACAAAGTCTTGACAGAATCATTCACAGAAACTACTTTT TCATTGTGACGTTCACACACAGAGTTTAAAATTTCTGTTCATGGAGCA GGTGGCAAACAATGGGTGATCTGCAACCGATTATTTGGAGGTCTTTCA GGTCTTCGTTGGTAAAGGGATCTCTTCCACTAGTGTTTGAGACAAGAA TACTCAGTAAGTTATGTGTGGTGTGTGTATTCAACACAAAGAGTTGAC CCTCCCTTTGTGAACGGTCACATTGATACTCTGTATTTGTGGCTTGCA GTTGGAGATTTCAAGAGCTTTCAAAGAATCGTAGAAACAGAAACACAT TCATACAAATACAAGACAGCATTATTCTCAGAAACTACATTGCGATTG TGACCGTTCACACACGAAGTTTAAGCTTTCTTTTGATGGAGCACTTTG GAAACACTCTGTCTAGTGTGCAAGGAGATATTTGAGCTCTTTAGGGCC 66 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: TTCATTGGAAAAGGCATTTCTTCAGATAGCGCAGGAAAGAAGAATACT GGCTAAGTTCCATGTGCTGTCTCTATTCAACTCACAAAAGTGAACACT CCTTTGTGAACGGTCACAGTGAAACCCTGTTTTTGTGTATTTCCACGT GGAGACTTCAAGTGCTTTGTGGCCAAATGTAGAAAAGGAAATATCTTC ATATGAAAACCAGACACAATCATTCTCAGAAACTATTTTCTGATTGTG ACCGTTCACACACAGAGTATAAGCTTTCTTTTGATGGAAGAGTTGCGA GACCCTCTCTTTGTAGACCGCAAGTAGATATATGGACCTCTTTATGGC CTTCAATGGAAACGGGATTTCTTCCTGTAATTTTAGACAGAAGAATTC CCAGTAACTTATTCGTGGTTTCCGTATTCAACTCACAGATATGAACCT TCCTTTGTGAACGGTCACATTCAAACCCTCTTTTGGAGGGTTTGCAGG TGGAGATTTCAAGGGCTTTTAGGCCAAAGCTTGAAAAGGAAATATCTT CTTATAGAATCTAGAAAGAAACATTCACAAAAACTTCTTTGTCATTGT GATCGTTCACACAAGAGGTTAATCTTTCTTTGGATGTTGGTGTTTAGA AAAACAGTGTTTTTATCGGCAAGTGGATATTTGGACTTCTTTGGGGCC TTCTATGGAATTGGGATTTCTTCCTATGATGTTTGACAGAAGAAGTCT CAGTAACTTCTTTCGGCTTGTGATCGTTCACACGCAGGTTTGAACTTT ACTTTAGACAGAGCGCATGTTAAACACAATTTTTGTGATTTGCAGCTG GAGATTACTAGCACTTTGTGGCGTATGGTAGAAATGGAAATATCTTTC CAAAATCCTGACAGAATCATTCACAGAAACTTGTTTTTCATTGTGATC GTTCACACACAGGGTTTAAACTTTCCGTTGACGGAGCAGGTTTCAAAC ACAGTCTTAGTCCGCAACTAAATATTTTGAGCTGTTTCAGGCCTTCAT TGGTAACGGTATCTCTTCATCTAATTTTTGACGCAAGAATTCTCAGCA AGTTATTTGTGGTGTGTGTATTCACCACACAGAGCTGACCCTTCTTTG TGAACGGTCACATTGATACAGTCTTTTTGTGACTTCCAGGTGGAGATT TCAAGCGCTTGCAGGCATTATAGAAACGGAAATACCTTTATATAACTA CTGGACAGCATCATTCTCAGAAACTATATTGTGATTGTGACCGTTCAC AGACGGAGTTTAACCTTTCTTTTGATAAAGCAGTTTGGAAACACTCTG TGAGTGTGCAAGCAGATATTTAACCTCATTAAGGTCTTCATTGGAAAC GTCATTTCTTCATGTAATGCAAGAAAGAAGAATACTCGGTAACTTCCT GGTGCTGCCTCTACTCAACTCACAAAGTTGAACAGTCCTTTGTGAACG GTCACAGTGAAACACTGTTTTTGTGAATTTCCAGTTGGAGACTTCAAG CACTTTGAGGCCAAATGTAGAAAAGGAAATATCTTTATATAAAAAGCA GACAGAATCGTTCTCAGAAACTATTTTGTGATTGTGACCGTTCACACA CAGAGGATAAGCTTTCTTTTGATGGAGCAGTTGGGAAACCCTGTGTTT GTAAACTGCATGTAGATATTTGGATCTCTTTAACGCCTTCATTGGAAA TGGGATTTCTTCACGTAATGTTCGACAGAAGAATTCTCAGTAACTTAA TCGTGGTGCCTTTATTCAACTCACAGATTTGAACCTTCCTTTGTGAAC GGTCACAATCAAACACACTTTTGGTGATTTGCATTTGGAGATTTCAAG CGCATTTAGACCAAATCTAGAAAAGCAAATATCTTCGTATGGAAACTG GAAAGAATCATTCTCAAAAACTACTTTTTCATTGTGATCGTTCACACA CAGAGTTTCATCTTTCTTTTGATATAAGAGTTGAGAAACAGACTCTTT GATCTGCAGGTAGACATTTGGACTTCTTTGAGGCCATCTTTGAAATCG GGATTGCTTCCTATAATGTTAGACAGGACAAGTCTCAGTAACTTCTTT CTCCTTGTGATCGTTCACACGTTGGGCTGAACTTTACTTTAGAAGGGC GGATGTAAAACACACTTTTTGTAATTTGCAGGTGGAGATTACAAGTAC TTTGAGTCGTACTGTAGAAATGGAAACATGTTTCTAAAGTCCAGACAG AATCATTCACAGAAACTAGTTTTTGATTGTGACGTTCACACACAGGGT TTAACATTTCCTTTCACGGAGCAGTTGTGAAACAAACGCTTTTTCCGC AAGCAGTTATTTTGACGTGTTTGAGGTCTTCATTGGAAAAGGTATTTC TTCCTGTAGTTTTCGAGGGAAGAATACTCAGCAACTTATGTGTGGTGT GTGTATTCACCTCAAAGAGCTGAACCTCCTTTGTGAACGGTCACATTG 67 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: AAACTGCGTTTTTGTGCTTTGCAGGTGGAGATTTCAATAGCTTGGAAG CGAACATAGAAAAAGAAATATATTTGTACAACAACAGGACAGAATTAT TCTCAGAAACTATTTTGCGATTGTGACCGTTCACAGAAGAAGTTTAAC CTTTCTTTTCATGAAGTACTTTGGAAACACTCTGTTAGTCTGCAAGGA GATATTTAGAGCTCATTGGGGTCTTCGTTGGAAAAGTGATTTCTTCAG GGAGTGCTGGAAAGAAGAATACTCACTAACTTCTAGGTGTTGTCTCTA CTCAACTCACAGAATTGAACTCTCCTTTGTGAACGGTCACAGTGAAAC ACTCTTTTTGTGATTTGCACTTGGAGATTTCAAGTACTTTTTGGCCAA ATGTAGAAAAGGAAATATCTTTGTATGAAAAGTAGACACAATCGTTCT CAGAAACTACTTTCTGATTGTGACCGTTCACACACAGAGGATAACCTT TCTTTTGATGGAACAGTTTCGAAACACTCTGTTTGTAATCCGCATGTG GATATATGGATCTCTTTGTCGCCTTCGATGGAAATGGGATTTCCTCCC ATAATTTTCCACAGAAGAATTCCCAGTAACTTAATTGTGGTTCGCTTA TTCAACTCACAGAGATGAACCTTCCTTTGTGAACGGTCACAATGAAAC CCACTTTTTGAGTTTGCAGTTGGAGATTTCAAGGGCATTTAGGCCAAA TCTTGAAAAGCAAATATCTTCTTATGGAATCTGGAAAGAAACATTCTC AAAAACTTCTTTTTCATTGTGATCGTTCACACCAGAGGTTCATCTTTC TTTGGATATTAGTGTTGAGAAAAAGAGTCTTTTTCGGCAGGTGGACAT TTGGACTTCTTTGGGGCCATCTATGAAATTGGGATTGCTTCCTATGAT GTTAGACAGAACAAGTCTCAGTAACTTCTTTCGCCTTGTGATCGTTCA CACGCTGGTCTGAACTTTACTTTAGACAGGGCGCATGTAAAACACAAT TTTTGTATTTGCAGGTGGAGATTACTAGTACTTTGTGTCGTATTGTAG AAATGGAAATATGTTCCAAAGTCCTGACAGAATCATTCACAGAAACTA GTTTTTCATTGTGACGTTCACACACAGGGTTTAAAATTTCCGTTCACG GAGCAGGTGTCAAACAAAGGCTTCCGCAACCAATTATTTTGAGGTGTT TCAGGTCTTCATTGGTAAAGGTATCTCTTCCTCTAGTTTTTGAGGCAA GAATACTCAGCAAGTTATGTGTGGTGTGTGTATTCACCACAAAGAGCT GACCCTCCTTTGTGAACGGTCACATTGATACTGTGTTTTTGTGCTTGC AGGTGGAGATTTCAAGAGCTTGCAAGGATCATAGAAACAGAAATACAT TTATACAACTACAGGACAGCATTATTCTCAGAAACTATATTGCGATTG TGACCGTTCACAGACGAAGTTTAACCTTTCTTTTGATGAAGCACTTTG GAAACACTCTGTGTGTGCAAGGAGATATTTAAGCTCATTAGGGTCTTC ATTGGAAAAGTCATTTCTTCAGGTAGTGCAGGAAAGAAGAATACTCGC TAACTTCCAGGTGCTGTCTCTACTCAACTCACAAAATTGAACACTCCT TTGTGAACGGTCACAGTGAAACACTGTTTTTGTGATTTCCACTTGGAG ACTTCAAGTACTTTGTGGCCAAATGTAGAAAAGGAAATATCTTTATAT GAAAAGCAGACACAATCGTTCTCAGAAACTATTTTCTGATTGTGACCG TTCACACACAGAGGATAAGCTTTCTTTTGATGGAACAGTTGCGAAACC CTCTGTTTGTAACCGCATGTAGATATATGGATCTCTTTATCGCCTTCA ATGGAAATGGGATTTCTTCCCGTAATTTTCGACAGAAGAATTCCCAGT AACTTAATCGTGGTTCCCTTATTCAACTCACAGATATGAACCTTCCTT TGTGAACGGTCACAATCAAACCCACTTTTGGAG 7 Exemplary ATTTGCCAGAGGACATTGGGGTAGATGTAAAGGTTTTGTCGGAACAGG artificial GATGTCATCGTACAAGATGTACACGGATGCTTTGTGAGGAACTACTAT centromere (NA) GCGAATGTGACCGTTCACAAACGGAATTCAACCACTGACTTTCCTACA GGAGTTTAGAGACACTGCTTCTTTTATATGTGGCAGAGGTCGATTTGG GCGTTTGGAAGCTCTTGTTGTTATAGAGAGTAACTACCACTGCAAACT GGAGAGTAGCGTTCTGAGAAGCTAGTCTGTCATTGTGACCGTTCACAA ATAGGGTGGAACGTTCCTATTTCACAGTAGTGTTGCAAAACACTGTTT ATAAATCTACGTGGGGTTACTTGGAGAGAATTCGGGATATCTTTCGAA CCGAGAACATGTTTATTTACAATCGCGAAAGTAGAATTCTGAGAGACT 68 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: TATTTGGGATATCTGTATACACGTTACACAGCTGAACTGTGACCGTTC ACAGGTGGGATTCAAACCCTCCTTTAGTGTTTGTGAACGGTCACATGG AACGACTTAACACCTACAGTTAAAAAGGGAATATATTCTCAAAAATAC AAGTCAGCAGCAATCTGAGAAACTGCTCTGGCATATCTGCATGCAACT AGCACAGCTGAATCTTTTTATAGACTGAGTAGATTTCAAAAAGACTTG CTGTGAACCTGCTAGCGGATGTTGGGTAGATTCGAAGAATTCTTTGTA AACAGGTTTACATACAAGAAGAAGTCAGTAGAATTCTCAGAATCTACT GTGTAATTGTGACCGATCACACACTGAGATGAAAATCCCCTTTTGTAG AGTAGGTTCGAGACGCTGTTTCAGTGGTCGGGAGTAAACTTTTGGTCA CCTTCCATGTTTACGGATGTGAACGATCACATTTCAAATCAAAAATAT ACAGAAACAGTCTGATCAACATGTTCGTTATATGGGATCTGAGGTATC ACAGATGAATCATTCTGTTGGTACAGTAGATCGGAGAAACTCTGTTTT TAATGTGAACGGTCACAATGTATACATATGTAGGTTTTGTAGGTAACA GGTATAACTACAAATAAAAACTTGAGAGACGCGTTGTCGGATACGCCT TGGTCATATTTACATTAAAATCCTACAGGTGACCACTCAGTATCTGAC AGGAGGTCTGCGGAACACTGTTCGTATACGTGGAAGGGGTTATATGTG ACCGTTCACAACCAACAGTTAAAAATCATATGTCTTACCGTATCCCCT GGATAGGAACGTTCCCAAAATCTTCTGTACGACCTACGCCCTCAACTA GCACAGGAGGACCATCTTTATGCCAAAGCTGTTCTGACACTCTCCTTA TGAAGATCTACAGGTAGATGTTTAGACAGCTTTGTAGACTTCATTAGA AAAGAGAGTACCTCCCATGAAGTCCAGGCACAAACAATCTTCTGAAAG TGTTCCGTAATTGTGATCGTTCACACGCACAGGTGAATAAATAGCTTT TCAAAGCGCTGGATTCAAGCGTTGTTTGTGAAGATATAGATGTAGATG TATCAGATGGCTTAAGACCAACGGCGAGAATGGAAAAATGTTTCCTTA GAAGATAAAAGGAATCACTCTGTTAAATTTCTTTTTGCTTGTGATCGT TCACATACGGAATTCAACCCTTGTTTGTAAAGTGCTGTTTGGAGACTC TATTCTTGAAGATCTGTGAGTGGACATTTCGATACATGTCATGATTCC GATGAAAAGGGAAATGTCATCCTAGAACATCTGGACGGATGCATTATC CGAAGCTTCATTGTAATATGCGCCTTAAAATCAAAGGGTGGAATTGTG ATCGTTCACAATTACGTGTGAAGCAGTCTCTTTATATACTGTGAACGG TCACATTTGAGAGCATTGGCGGCTACGATGGAAAGAGAAACATCGTCC GATGAAAGCTTGACGGAAGGAATCTCGGAATTTTTTTCGGGTTATATT CACACAGTTAATAGGGTTTAACTTTTCCATTAACATAGCTGTATTGTA ACCGTGTTTTTGTGAATCTACAGGTAGATACTTTGAAAGCATGTAGAA TTACGCTGGACACGAGATTTCGCATCAAATGTGGAAAGCCGCCTCCAC AGAAAATTGTTTATGGTTGTGACCGTTCACACCCAGGGTTCAACTTTA CCTTTCATACGGCCGTGTTTAAGCATCATTCTATAACTTGACGTGGAC ACTACGAGAGATTTGAGCTCCATAGTTGTGAGCGATCACACTTGAAAT ATAAACCAGAGAGAAGAATACTCTTAATCTTCTTTCTGTTGAGTTAAG TCTGCAAAGAGCGGCGGACCTTCCTTTGAAAGGGCTGTACTAAGAACA ATTTTGTGATTGTGAGCGGTCACATAGGAAAGTTTAGAGGATGTCTTT AGAATCGAGAACATATTTATTTCATATTTACACGGAAACAATCCCAAA AATGTATTTCTGTTGATTGAATTCCACCCACAGTGTTTAAGATCCCCT TCCAAAGGGCCGCTTTGGAAGCGCTGTTGTTTTAGATATGCACGTCGA CATTTGTGAACGTTCACAGTCTTCGTTGCAAAAGGAATTATCATCGCA CAAACAGTAAACTGAATCAGTCTGAGGAATTCTTTCATCACGAATACA TTCACTTAAAAGCGAGGAAACTACCATTAGAGAGCGCATTTTGGATGC ATTCTATTAGTGAAACTGAAAATGCATACTTGAATACCTGGGAGGATA TCGCTGCAAACAGGAAACATATTGCTGTAGAACCTGGATAGCAGAATC CCCATAAACAGCCCTCTGTTTGTGACCGATCACACAGAGGGTAGACTA CTTTCCCATTCATGGAACACGTATGTAATGCACCTTTGGTGTTATGAA 69 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: ATTGAACCTTACGAACAGTCTGCGGACCAATAGTAATGAACGGTATCT CGTCACCCACCAGCAAGGAATAAGAATTCTCTGATACCTGTTGAGAAT GTGATCGTTCACAATCACAGGTGAAACTGTCCTTTCACGGAACACTTT TTAACCATTCATTATGTGAATCGGCAAGGTGATTTTTGTATAGGTTTG AGTATTTTGTGGGAACCGGAGTACCTCCACATTAAATGTCAACAAAAA CATTCCCATAAACGTCTTTATGATGTGTCCAATCCAGCCACCGATTTC AATTGTGACCGCTCACAAGCAGTTTCGAAAAACACTTTCTGAGATTGT GAACGATCACATTGAAGCACCCTGATGCCTATGGAGAACAGGTAAATA CCTCCCGTACAAAGTACACACAAGCAATATCAAAATCGTCCTTGAGAT GTATGTACGAAGCCAACTGATTTGTACCGTTCTGTTGAGAGAACACTT TAGATACATTCGTTCTGAGGATCTGTAAATGCATATATGCATTGCTAG GGGGGTTTTGTAGGAAAAGGTATTATGTTTAGAAATTACACCGCAACA ACCTAAGAAACATCCTTGGGACTGTGACCGTTCACACAAAGAATTGAC CACTCCGTTTCGAACAACATTTCTGTAAAACCTTTATGAAGCATTGCA AGAGAATATCAGAACTGCATTCCGGACTTTGATTGTGAACGCTCACAC TTCATATAAAAGCTGGACACAAGCGTTTTCAAAAAGTTGATTGGAAGT CTGTACCCAACTTACGGACGTAGATGTTTCCTTTCATTGAACAATTTT GGAAACCACATTTGTGATTGTGAACGATCACATTCGATGGAATTAAAA ATTACGGTGCAAAAGGTAATCTCGTCCTAACAAGTCCAGGCAAAAGGA TACTTAGGAACCTCGTTGCGACGTTGGCCTTCACCTATGGATTTGGAC TTTACGATTGAGGGAACACCATTTAGACATTCATTCTGTGTTGTTCAA ATGCATACTTGTGACCGCTCACAGCGTAAGGGGAGAAAGCTAACATCT ACCAATCACGACAAGTCATAAAGATTATCCGAGACACCCTTGTGATGT GTGTACTAACCCAACGGAAAATAACTTTGCTCTTCACTGAGTAGCTTT CATAAACACTTGTTTATTTACCGGAAGACGTTGAGGCAGATTTATAGG CTTTATCAGAAAAAGGGTGCCACGCACGAGGTGCACGCGCATACTATG TTGTGGAAGTATTACGCAAATGTGATCGTTCACAAGCGCAAGTCAAAT CAACAGATTTTCCAACCGGTGTATACAGGCATTGGTTCGTTAAATGTA GCTGAAGTTGAATTAGGTGTCTGAAAACTATCGTCGTGATTGAAAGAA AGTATCATTGGAAAATGAAGGGTATCGCTCTGATAAGTTACTCTTTCC TTGTGACCGTTCACATATGGGATGCAACGCTCGTATGTAACTGTTGTG TGGCGAATCAATGCTTAAAATCTATGTGTGGTCACTTCGAGACAAGTC GTGATACCTATCAAACGGAAAACGTGATTCTTGACCATCGGGAAGGTT GAATTATGCGAGGCTTAATTGGAATATCCGTCTAAACATTAAACGGCG GAACTGTGATCGTTCACAGTTGCGAGTCAAGCCGTCCCTTAATTCTGT GAACGGTCACATTGAAAGAATTAGCAGCTACAATTGAAAAAGGAACAT AGTCTGAAGAATGCATGTCGGCAGGAATCTGGGAAATTGTTCCGGCTT ATCTTCATACAATTAGTACGGCTTAATTTTTTCATAAACTTAGTTGAA TTCTAAACGAGTTGTTGTAACCTACTGGCAGATGCTGTGAAGAATCTA AAAATACTCTGTACACAAGTTTTCACACCAGATGAGGTAAGTCGACTT CACAGAATATAGTGTATAGTTGTGACCGATCACACCCTGGGATCAAAT TCACCTTTTATAGGGTCGGGTCTAGGCGTGATTCAATGCTGGCGTAGA CTCTTCGTGACATTCGATCTTCACAGATGTGAGCGATCACATTTGAAA TCTAAAACATAGAGAAAAAGACTGTTCATCATCTTCCTTTTAAGGTAT GTGTGGAATGACCGACGAACCATCCTGTGGAACGGTTGAACGAGGAAC TATGTTTATGTGAGCGGTCACAAAGTAAACTTAAGTGGGTGTTTTAAG TATCAAGTACAAATATAATTAATAATTTCAGGGACACGATGCCGAATA TGCATTGCTCTTAATTAAATTACAACCCCACTGGTTACGACCCACTAC CTAACGGGCGGTTCGGCAGAGCAGTGGTCTTAACATGGACGGCGTCAT ATGTGAACGTTCACAATCATCATTTCAAAATGATTTGTCATAGCGCAT ACCGTGAATTGGATCGGTCCGAAGATTTTTTGCACCACCAACACCTTC 70 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: AATTAGAACCGGGGGAACAACTATAAGCGAACGCTTTTCGGACGCTTT CCATAAGAGAATTGACAAAGCACACTGGAGTACATGGAAGGGTATTGC CGCAACAAGAGACGTAATGGTGCAGGACGTGCATGGCTGATTCGCGAT GAACAACCATCCGTATGTGACCGATCACAAAGGGGATACACATCCTCT CACATTCCTGCAAGACTTAAGTGATACAGCCTCTGTTTTGTGACATAG ATCCATATGAGCATTCGGCAGATCATTATTATTGTACAGTGTCACGAC AACCGCCAACAGGGGATTAGAGTTCTCAGATGCCAGCTGACAATGTGA CCGTTCACAATTACGGGGGAAAGTGCCCATTCCGCAATACTGTTTCAC AATACAGTATATAATCGACATGGTGTTTCTTGTAGAGGATTGGGTATA TTTTGCGACCCAGAGCACGTCTACTTTCAATGGCAAAAATAAAATTCC GATAGACGTATTTAGGATGTCTCTAAACCCGCTACCCATCTCAACTGT GACCGCTCACAGGCGGTATCCAAAACCACCTTCAGGTTGTGAACGATC ACATGAAACAACCTAATACCTATAGATAACAAGTGAATACATCTCGAA CATAGAACTCACCAGCAATATGAAAAACGGCCCTGACATGTCTGTATG AAACCAGCTCATCTGTATCGTTTTGTAGAGTGAATACATTACATAAAT ACGTGCTGAGACCTGTTAACGCATGTAGGCTTGATACGGAGGATTTTT AGTAAAAAGTTTTATATTCAGGAATAACTCCGTAAAAATCTAAGAATC AACCGTGGAACTGTGACCGATCACACAATGAAATGACAACCCCGTTTT GAAGAATATGTCCGTGAAGCGTTTAAGAGGCTGCGAGAAAATTTCTGA TCTCCATCCCTGATTTCGAATGTGAACGCTCACATTTCATATCAAAGA TGTACACAAACGGTTTGAACAAGATGATCGTAAATCGGTTCCGAAGTT TCGCACATAAATGATTCCGTTCGTTCAATAAATTGGGGAACCTCAGTT TATGTGAACGATCACAATCTATGCAAATATAAGTTATGGAGCTAAAAG TTATCACGACCAAAAAAGACCTGGGAAACGGGTAGTTGGGTACCCCGT GGCCACATTGACCTTAACATCTGCATGTGGCCTCTAAGAATGTGGCAA GACGACTTCGAAATACAGTCCGTTCGTTGAAAGGCTTACATGTGACCG CTCACAACGAAAAGGTAGAAATCTTACGTCTAACAGTCTCGCCAGGTT ATGAAGGTTACCCAAGTCATCCGTGCGATCTGCGTCCTAAACCAGCGC AAGATGACTATGTTCATCCCTAAGTTGCTCTCACAATCACCTGATAAA ACTAAAGATACATGCTTAAACACCTTGGTGGACATCACTACAAAAAGA AAGCACATGCCGTGGAGCCCGGGTACCAAAAACCTCCTTAAAGAGTCC CCTATTTGTGATCGATCACACGGACGGGAGAACTTAATTACCTATTCA AGGCACTCGAATCTAGTGTAGCTTGGGAGTATAAATTTAAATCTAACA AATAGCCTACGAACAAACAGCAAGGATCGATAACTGGTTACTCAGCAG AAAAGAGTAATAACTCTCTTATATCTCTTTAGCATGTGATCGTTCACA TTCGCAAGTCAAACCGTGCTTGCAAGGTACTCTTTGTAGCCTTTAATC ATGAGATCGGTAAGTTGACTTTTCTATACGTGTGATTATTCTGAGGAA AGCGAAGTGCCACCCCAGTACATGTGAACGAATACATTACCCTAAGCG TCATTATAATGTGCCCCATACAACCAACGGTTGCAATTGTGATCGCTC ACAATCACTTGCGAAGAAGACTCTCTAAACTGTGAACGATCACATTTA AGAACACTGGTGGCTATGAAGGACAGATAAACACCGCCGGTGCAAGGT TCACGCAAGGAATATCGAAATTGTTCTCGAGTTGTATTTACAAAGTCA ATTGGTTTTTACTGTTCCGTTAAGATAACTCTATAGTTACCTTGGTTT TGAGATCTATAGATACATACATTCAATGCAAGTGGAGTTATGCAGGAC AAGATATTTTGCTTCGAATTTGCAACGCCACCACCAAAGAAAAATGCT TAGGGCTGTGACCGTTCACACCAAGGATTCACCTCTACGTTTCAAACG ACCTTGCTTTAGAACATTATAAACATTCACGAGGATACCACAAGTGAA TTGCGCACCTTAATTGTGAGCGCTCACACTTGATATATAAGCCGGAGA CAAGAGTATTCTAAATGTTCATTCGTAGACTTTAGCCTACATAGGGCC GCAGACGTTCCCTTCAATGGACTATATTGAGACCAAATTGGTTGTGAG CGATCACATACGAAGGTATAAAGAATGACTGTACAATAGATAACCTAG 71 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: TTCTTACATGTTCACGCGAAAAGAAACCTAAGAATCTAGTTCCGTCGA TGGACTTCCCCCACGGTTTTTGAGTTCACCATCGAAGGGACCCCTATG TAAACGTTGATGCTTTGTATTCACATCCACACTTGTGAACGCTCACAG TGTTAGTGGCGAAAGGTATCATCAACGAACCAAGAGAAATCTTAATGA GTATGCGGGATACTCTCGTCATGAGTATATTAACTCAAAGGCAAGTAA ATTAGCACTACAGTGCGTATCTTGCATGAATACTAGTATATTAACGAA ACACGCTGAAGCACATTGATGGGCATTACCACAAAAAAAGGCGCAAGG CGCGGGGCGCGCGTGCCTAATACGTCGTTGAAGAATCACCCATATGTG ATCGATCACAAGGGCGAGACAACATTCATCACATATTCCAGCCAGTCT AAACTGGTATAGGCTCGGTATGTAACTTAAATTCAAATAAGTATCCGA CAAATAATCATCATGGTTCAATGACAGGATAATCGGCAAAAGAGGGTT ATAGCTCTCATATGTCACCTTACCATGTGACCGTTCACATTTGCGAGG CAAAGCGCGCATGCAGCTATTCTGTGTCGCATTAAAGCATAAATCGAT ATGTTGTCTCTTCTAGACGAGTGGTTATACTTAGCAACGCAAAGCGCG ACTCCTGTCCATGGGAAAGATTAAATTACGCTAGGCGTAATTAGAATG TCCCTCAAACCACTAACCGTCGCAACTGTGATCGCTCACAGTCGCTAG CCAAGACGACCCTCAACTGTGAACGATCACATTAAAAAAACTAGTAGC TATAAATGACAAATGAACACAGCTGGAGCATGGATCTCGCCAGGAATA TGGAAAATGGTCCCGACTTGTCTTTATAAAATCAGTTCGTCTTTATTG TTTCGTAAAGTTAATTCAATACTTAACTAGGTGTTGAACCTATTGACA CATGCAGTCATGAAACTGAAGATATTCAGTACAAAATTTTTTACTCCG GATTAGCTACGTCAACATCAAAGAATAAAGCGTAGAGCTGTGACCGAT CACACCATGGAATCACATCCACGTTTTAAAGGATCTGGCCTTGGAGGA TTAAAAGCTCGCGAAGATTCCTCATGTCAATCGCTCATCTCAAATGTG AGCGCTCACATTTGATATCTAAGACGTAGACAAAAGGATTGTACATGA TCATCCTTAAACGTTTGCGTAGATTGGCCCACAAACGATCCCGTCGAT CGATTAAATGGGGACCTAAGTTGTGAGCGATCACAAACTAAGCTAAAA TGAGTGATTGAACTATAAATTACCAAGATCATAAATGATCTCGGGAAC AGGAAGCTGAGTATCCAGTGCCCTCAATGAACTTACCACCCGCTTGTT GCGTCCAACAACGTAGCGAGCCGTACGTCAAAGTAGAGGCCTTCATTG ACAGCCTCACATGTGAACGCTCACAATGATAATGTCGAAATGTTTCGT CAAAGAGCCTAGCGAGATTTTGATGGGTACGCAGGTTATTCGCGCCAT CAGCATCTTAAATCAGAGCCAGGTGAATAAGTACAACCGTACGTTTCT CGCACGATTACCAGAAA 8 Exemplary ATCTGCAAGTGGATATTGGGATAGCTTTGAAGATTTCATTGGAAACGG artificial GAATATCATCCTATAAAATGTAGACAGAAGCATTCTCAGAAACAGCTC centromere (NA) TGTGATTGTGACCGTTCACACACAGAGTTCAACATTGCCTTTCATAGA GCAGTTTTGAAACGCTGTTTTTGAAGTATATGAAAGTGAACGTTTCGG ATGGTTGGAGGCCCATGGTGTTAAAGGAAATATCTTCCCCTACAAGCT AGAAAGAAGCATTCTGTGATACTTGTTTGTGATTGTGATCGTTCACAA ACAGGGTTGAACCTTTCTTTTTACAGAACAGTTTTGAAACACTCTTTT TGTAAATCTGCGAGGTGATATTTGGATAGATTTCAGGATTTTGTTGGA AACGGAAATATCTTCATAGAAAATCTCGACAGAAGCATTATCAGAAAC TTCTTTATGATATGTGTATTCAAGTTACAGAGTTGAATTGTGATCGTT CACAAGCAGGTTTGAAACACTCTTTTTGTAGATTGTGAACGGTCACAT TGGAGCACCTTGACGCCTATGGTGAAAAGGGAAATATCTTCTCATAAA AACTAGACAGAAGGAATCTCAGAATCGTCTTTGGGATATATGTACGCA GCTAATAGAGTTGAATCTTTCTATAGACAGAGCAGTTTAGAAACAGTC TTTCTGTGGAATCTACAAGTGGATATTTGGATAGCTTGGGGGATTTCG TTGGAAACGGGATTATGTATAAAAATTAGACAGCAGCATCCTCAGAAA CTACTTTGTGATTGTGACCGTTCACACACAGAGTTCAACATTCCCTTT 72 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: CATACAGCAGTTTTGAAACGCTCTTTCTGTAGTCTGGAAGTGAACTTT AGGAGAGCTTTCAGGTCTATGGTTGTGAACGCTCACACTTCAAATAAA AACTAGACAGAAGCATTCTCATAAACTTCTTTGGATGTGTGAACTCAG GTAACAGCGGCGGATCTTTCTTTTGATAGAGCAGTTTTGAAAAACACT TTTGTGATTGTGAACGGTCACATTGGATAGTTTTGAAGATGTCGTTGG AAACGGGAATATCGTCATATCAAATCTAGACAGAAGCATTCCCAGAAA CGTCTTTGTCATGTTTGCATTCAACCCATAGAGTTGAACATTCCCTTT CAGAGAGCAGCTCTGAGGCACACTTTTTGTATATGTGCAAGGGGATAT TTGTGACCGTTCACAGTCTACAGTGAAAAAGCAAATATCTTCCCGTAA CCACTAGACAGAAACATTCCCAGAAACTCCTTTATGACGTATGTACTC ACCTAGCAGAGAAGAACCTTCCTTTTGCCAGAGCAGTTTGGATACACT CTTTTTGAAGAATCTGCAAGTGCATATTTGGATACCTGTGAAGGTTTC GCTGGAAACGGGAATATCTTCCTATAAAGTCTGGACAGAAGCATTCTC AGAAACTGTTCTGTGATTGTGACCGTTCACACACAGAGTTGAACATTG CTTTTCATAGAGCAGGTTTCAAACGCTCTTTTGGTATATATGGAAGAG GATGTTTCGGACAGTTTGAAGCCCATGGTGAGAAAGGGAATATCTTTC CCTACAAGCTAGAAAGAAGCATTCTGTGAAACTTGTTGTGATTGTGAT CGTTCACAAACAGAGTTGAAACTTTCTTTTTACAGAGCAGTGTTGAAA CACTCTTTTTGTAGAATCTGTGAGGGGTTATTTGGATAGATTTCAGGA TTTCTTTGGAAACGGGAACATCTTCATATAAAATCTCGACGGAAGCAT TCTCAGAAGCTTCTTTGTGATATCTGCCTTCAAGTCACAGAGCTGAAT TGTGACCGCTCACAAGTAGGTTTGAAACACTCTTTTTATAGTATTGTG AACGGTCACATTTGAGCGCCTTGATGCCTACAGTGAAAAGGGAAATAT CTTCCGATAAAAACTAGACAGAAGCAATCTCAAAATCTTCTTCGGGAT ATATGCACACAGCTAACTGAGTTGAACTTTTCTATTGACATAGCAGTT TTGAAACATTCTTTCTGTGGAATCTGCAAGTGGATGTTTGGATAGCTT GGAGGATTTCTTTGGAAACGGGATTACATATAAAAAGTAGACAGCAGC ATCCTCAGAATCTTCCTTGTGATTGTGACCGTTCACACACAGAGTTGA AAATTCCCTTTCGTACAGCAGTTTTTAAACACTCTTTCTGTGGATCTG GGAGTGAACACTAGGACAGCTTTCAGCTCTATGGTTGTGAACGATCAC ACTTCAAATAAAAACTAGACAGAAGAATTCTCATAAACTTGTTCGTGA AGTGTGAACTCAGCTAAGAGACGTGGATCTTTCTTTTGGTAGAGCAGT TCGGAAAAACACTTTTGTGATTGTGAACGGTCACATTGGATAGATATG AAGATTTCTTTGGAAACGGGAATATCTTCATATCAAGTCTAGACAGAA GCATTCTCGGAAACGTCTTTGTGATGTTGGCATTCAACTCATAGATTT GAACATTCCGTTTCAAAGAGCAGCTTTGAAGCACTCATTTTGTAGATG TGCAAGTGGACATTTGTGACCGTTCACAGCCTTCGGGGAAAAAGCAAA TATCTTCCCATAACCACTAGACTGAAACATTCTCAGAAACTTCTTTAT GACGTATGCACTCAACTAAAAGAGAAGAACCTTCCTTTTGAGAGAGCA GTTTTGATACATTCTTTTTGTGAATCTGCAAGTGGACATTTGGATAGA TGTGAAGATTTTGTTGGAAACGGAAATATCTTCCTATAAAACCTAGAC AGAAGCATTCTCATAAACTGCTCTGTGATTGTGACCGTTCACACAGAG AGTTGAACATTGCCTTTCCTAGAGCAGGTTTGAAACACTCTTTTTTTA GATATGGAAGTAGACGTTTCAGACGTTTTGAGGACCATGGTGATGAAG GGAATATCTTCCCCTACAAGATAGAAAGAAGCATTCTGTGAAACTTGT TTGTCATTGTGATCGTTCACAAACAGAGTTGAACCTTTCCTTTTACAG AGCAGTTTTGAAACACTCTTTTTGTGAATCTGCAAGGGGATATTTGGA GAGATTTCAGGATTTCGTTGGAAAGGGGAATATCTTCACATAAAATCT CGACAGAAACATTCTCAGAAACTTCATTGTGATATGTCCATTAAAGTC ACAGAGTTCAATTGTGACCGTTCACAATTAGGTTTGAAACACTCTTTT TGTATATTGTGAACGGTCACATTGGAGAGCCTTGACACCTACGGTGAA 73 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: AAGGGGAATATCTTCCCATAAAAAGTAGACAGAAGCAATCTCAGAATT TTCTTTGAGATATATGCACGCAGTTAACAGAGTTGTACCTTTCTGTTG ACAGAGCAGATTTGAAACAGTCTTTCTGTGAATCTGCAAGTGGATATT TGGATAGATTGGAGGATTTCGTTGGAAACAGGATTACGTATCAAAAGT AGACAGCAGCATCCTCAGAAACATCTTTGTGATTGTGACCGTTCACAC ACAGGGTTGAACATTCCGTTTCGTACAGCAGTTTTGAAAAACTCTTTC TGTAATCTGGAAGTAAACATTACGACAGCTTTCAGGTCTATAGTTGTG AACGATCACACTTCAAATAAAAACTAGACAGAAGCATTCTGATAAACT TGTTTCTGATGTGTGATCTCAGCTAACACAGATGGATCTTTCTTTTGA TACAGCAGTTCTGAAAAACACATTTTTTAATGTGAACGGTCACATTGG ATAGATTAGAAGATTTCGTTGGTAACGGGAATATCTTCATATCAAATC TAGACAGACGCATTCTCAGAAACGTCTTTGCGATGTTTACATTCAACT CATAGAGTTGAACATTCAGTTTCAGAGAGCAGGTTTGAGACACTCTTT TTGTGTTGTGCAAGTGGATATTTGTGAACGTTCACAGCCTAAGGTGAA AAAGGAAATATCTTCCCATAACCACTAGACAGGAACATTCTCAGAAAC TCCTTTATGACGTGTGCACTCACTTAACAGAGAAGAACTTTCCTTTTG ACAGAGCATTTTTGATACACTCCTTTTGTAATCTGCAAGTGCATATTG GGATACCTTTGAAGGTTTCACTGGAAACGGGAATATCATCCTATAAAG TGTGGACAGAAGCATTCTCAGAAACAGTTCTGTGATTGTGACCGTTCA CACACAGAGTTCAACATTGCTTTTCATAGAGCAGTTTTCAAACGCTGT TTTGGAATATATGAAAGAGAATGTTTCGGATAGTTGGAAGCCCATGGT GTGAAAGGAAATATCTTTCCCTACAAGCTAGAAAGAAGCATTCTGTGA TACTTGTTGTGATTGTGATCGTTCACAAACAGGGTTGAAACTTTCTTT TTACAGAACAGTGTTGAAACACTCTTTTTGTAAATCTGTGAGGTGTTA TTTGGATAGATTTCAGGATTTTTTTGGAAACGGAAACATCTTCATAGA AAATCTCGACGGAAGCATTATCAGAAGCTTCTTTATGATATCTGTCTT CAAGTTACAGAGCTGAATTGTGATCGCTCACAAGCAGGTTTGAAACAC TCTTTTTATAGATTGTGAACGGTCACATTTGAGCACCTTGATGCCTAT AGTGAAAAGGGAAATATCTTCTGATAAAAACTAGACAGAAGGAATCTC AAAATCGTCTTCGGGATATATGTACACAGCTAATTGAGTTGAATTTTT CTATAGACATAGCAGTTTAGAAACATTCTTTCTGTGGAATCTACAAGT GGATGTTTGGATAGCTTGGGGGATTTCTTTGGAAACGGGATTATATAT AAAAATTAGACAGCAGCATCCTCAGAATCTACCTTGTGATTGTGACCG TTCACACACAGAGTTCAAAATTCCCTTTCATACAGCAGTTTTTAAACG CTCTTTCTGTGGTCTGGGAGTGAACTCTAGGAGAGCTTTCAGCTCTAT GGTTGTGAACGCTCACACTTCAAATAAAAACTAGACAGAAGAATTCTC ATAAACTTCTTCGGAAGTGTGAACTCAGGTAAGAGCCGCGGATCTTTC TTTTGGTAGAGCAGTTTGGAAAAACACTTTGGTTGTGAACGGTCACAT TGGATAGTTATGAAGATGTCTTTGGAAACGGGAATATCGTCATATCAA GTCTAGACAGAAGCATTCCCGGAAACGTCTTTGTCATGTTGGCATTCA ACCCATAGATTTGAACATTCCCTTTCAAAGAGCAGCTCTGAAGCACAC ATTTTGTAATGTGCAAGGGGACATTTGTGACCGTTCACAGTCTTCAGG GAAAAAGCAAATATCTTCCCGTAACCACTAGACTGAAACATTCCCAGA AACTTCTTTATGACGTATGTACTCAACTAGAAGAGAAGAACCTTCCTT TTGCGAGAGCAGTTTGGATACATTCTTTTTGAGATCTGCAAGTGGACA TTGGGATAGATTTGAAGATTTTATTGGAAACGGAAATATCATCCTATA AAACGTAGACAGAAGCATTCTCATAAACAGCTCTGTGATTGTGACCGT TCACACAGAGAGTTCAACATTGCCTTTCCTAGAGCAGTTTTGAAACAC TGTTTTTTAAGATATGAAAGTAAACGTTTCAGATGTTTGGAGGACCAT GGTGTTGAAGGAAATATCTTCCCCTACAAGATAGAAAGAAGCATTCTG TGATACTTGTTTGTCATTGTGATCGTTCACAAACAGGGTTGAACCTTT 74 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: CCTTTTACAGAACAGTTTTGAAACACTCTTTTTGTAATCTGCAAGGTG ATATTTGGAGAGATTTCAGGATTTTGTTGGAAAGGGAAATATCTTCAC AGAAAATCTCGACAGAAACATTATCAGAAACTTCATTATGATATGTCT ATTAAAGTTACAGAGTTCAATTGTGATCGTTCACAATCAGGTTTGAAA CACTCTTTTTGTAATTGTGAACGGTCACATTGGAGAACCTTGACACCT ATGGTGAAAAGGGGAATATCTTCTCATAAAAAGTAGACAGAAGGAATC TCAGAATTGTCTTTGAGATATATGTACGCAGTTAATAGAGTTGTATCT TTCTGTAGACAGAGCAGATTAGAAACAGTCTTTCTGTGAATCTACAAG TGGATATTTGGATAGATTGGGGGATTTCGTTGGAAACAGGATTATGTA TCAAAATTAGACAGCAGCATCCTCAGAAACAACTTTGTGATTGTGACC GTTCACACACAGGGTTCAACATTCCGTTTCATACAGCAGTTTTGAAAA GCTCTTTCTGTATCTGGAAGTAAACTTTACGAGAGCTTTCAGGTCTAT AGTTGTGAACGCTCACACTTCAAATAAAAACTAGACAGAAGCATTCTG ATAAACTTCTTTCGATGTGTGATCTCAGGTAACACCGACGGATCTTTC TTTTGATACAGCAGTTTTGAAAAACACATTTTATGTGAACGGTCACAT TGGATAGTTTAGAAGATGTCGTTGGTAACGGGAATATCGTCATATCAA ATCTAGACAGACGCATTCCCAGAAACGTCTTTGCCATGTTTACATTCA ACCCATAGAGTTGAACATTCACTTTCAGAGAGCAGGTCTGAGACACAC TTTTTGTTTGTGCAAGGGGATATTTGTGAACGTTCACAGTCTAAAGTG AAAAAGGAAATATCTTCCCGTAACCACTAGACAGGAACATTCCCAGAA ACTCCTTTATGACGTGTGTACTCACTTAGCAGAGAAGAACTTTCCTTT TGCCAGAGCATTTTGGATACACTCCTTTTGAAAATCTGCAAGTGCACA TTTGGATACATGTGAAGGTTTTGCTGGAAACGGAAATATCTTCCTATA AAGCCTGGACAGAAGCATTCTCATAAACTGTTCTGTGATTGTGACCGT TCACACAGAGAGTTGAACATTGCTTTTCCTAGAGCAGGTTTCAAACAC TCTTTTGTTAATATGGAAGAAGATGTTTCAGACATTTTGAAGACCATG GTGAGGAAGGGAATATCTTTCCCTACAAGATAGAAAGAAGCATTCTGT GAAACTTGTTGTCATTGTGATCGTTCACAAACAGAGTTGAAACTTTCC TTTTACAGAGCAGTGTTGAAACACTCTTTTTGTGAATCTGTAAGGGGT TATTTGGAGAGATTTCAGGATTTCTTTGGAAAGGGGAACATCTTCACA TAAAATCTCGACGGAAACATTCTCAGAAGCTTCATTGTGATATCTCCC TTAAAGTCACAGAGCTCAATTGTGACCGCTCACAATTAGGTTTGAAAC ACTCTTTTTATATATTGTGAACGGTCACATTTGAGAGCCTTGATACCT ACAGTGAAAAGGGGAATATCTTCCGATAAAAAGTAGACAGAAGCAATC TCAAAATTTTCTTCGAGATATATGCACACAGTTAACTGAGTTGTACTT TTCTGTTGACATAGCAGATTTGAAACATTCTTTCTGTGAATCTGCAAG TGGATGTTTGGATAGATTGGAGGATTTCTTTGGAAACAGGATTACATA TCAAAAGTAGACAGCAGCATCCTCAGAATCATCCTTGTGATTGTGACC GTTCACACACAGGGTTGAAAATTCCGTTTCGTACAGCAGTTTTTAAAA ACTCTTTCTGTGATCTGGGAGTAAACACTACGACAGCTTTCAGCTCTA TAGTTGTGAACGATCACACTTCAAATAAAAACTAGACAGAAGAATTCT GATAAACTTGTTCCTGAAGTGTGATCTCAGCTAAGACACATGGATCTT TCTTTTGGTACAGCAGTTCGGAAAAACACATTTTATGTGAACGGTCAC ATTGGATAGATAAGAAGATTTCTTTGGTAACGGGAATATCTTCATATC AAGTCTAGACAGACGCATTCTCGGAAACGTCTTTGCGATGTTGACATT CAACTCATAGATTTGAACATTCAGTTTCAAAGAGCAGGTTTGAAACAC TCATTTTGTGTGTGCAAGTGGACATTTGTGAACGTTCACAGCCTTAGG GGAAAAAGGAAATATCTTCCCATAACCACTAGACTGGAACATTCTCAG AAACTTCTTTATGACGTGTGCACTCAATTAAAAGAGAAGAACTTTCCT TTTGAGAGAGCATTTTTGATACATTCCTTTTGTATCTGCAAGTGCACA TTGGGATACATTTGAAGGTTTTACTGGAAACGGAAATATCATCCTATA 75 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: AAGCGTGGACAGAAGCATTCTCATAAACAGTTCTGTGATTGTGACCGT TCACACAGAGAGTTCAACATTGCTTTTCCTAGAGCAGTTTTCAAACAC TGTTTTGTAAATATGAAAGAAAATGTTTCAGATATTTGGAAGACCATG GTGTGGAAGGAAATATCTTTCCCTACAAGATAGAAAGAAGCATTCTGT GATACTTGTTGTCATTGTGATCGTTCACAAACAGGGTTGAAACTTTCC TTTTACAGAACAGTGTTGAAACACTCTTTTTGTAATCTGTAAGGTGTT ATTTGGAGAGATTTCAGGATTTTTTTGGAAAGGGAAACATCTTCACAG AAAATCTCGACGGAAACATTATCAGAAGCTTCATTATGATATCTCTCT TAAAGTTACAGAGCTCAATTGTGATCGCTCACAATCAGGTTTGAAACA CTCTTTTTATAATTGTGAACGGTCACATTTGAGAACCTTGATACCTAT AGTGAAAAGGGGAATATCTTCTGATAAAAAGTAGACAGAAGGAATCTC AAAATTGTCTTCGAGATATATGTACACAGTTAATTGAGTTGTATTTTT CTGTAGACATAGCAGATTAGAAACATTCTTTCTGTGAATCTACAAGTG GATGTTTGGATAGATTGGGGGATTTCTTTGGAAACAGGATTATATATC AAAATTAGACAGCAGCATCCTCAGAATCAACCTTGTGATTGTGACCGT TCACACACAGGGTTCAAAATTCCGTTTCATACAGCAGTTTTTAAAAGC TCTTTCTGTGTCTGGGAGTAAACTCTACGAGAGCTTTCAGCTCTATAG TTGTGAACGCTCACACTTCAAATAAAAACTAGACAGAAGAATTCTGAT AAACTTCTTCCGAAGTGTGATCTCAGGTAAGACCCACGGATCTTTCTT TTGGTACAGCAGTTTGGAAAAACACATTTGTGAACGGTCACATTGGAT AGTTAAGAAGATGTCTTTGGTAACGGGAATATCGTCATATCAAGTCTA GACAGACGCATTCCCGGAAACGTCTTTGCCATGTTGACATTCAACCCA TAGATTTGAACATTCACTTTCAAAGAGCAGGTCTGAAACACACATTTT GTTGTGCAAGGGGACATTTGTGAACGTTCACAGTCTTAAGGGAAAAAG GAAATATCTTCCCGTAACCACTAGACTGGAACATTCCCAGAAACTTCT TTATGACGTGTGTACTCAATTAGAAGAGAAGAACTTTCCTTTTGCGAG AGCATTTTGGATACATTCCTTTTGA 9 Exemplary ATCAGCGAGGGGTTATTTAGATATCTGCGAAGTTTTTGTATGTGACCG artificial GTCACAATTCCTAAAAAGTCCAGAGAGTAGCATTTTCAGAAACAGCTC centromere (NA) TGCGATATCTGCTTTCAAATCACGGACTTCAACACTGGCTTTCCTACA GCTGGTTTCAAGCGATCTTTCTGAGATGTGAAACTGGACTTTTTGGTC GTTTGGAAGCTCACGGTGATGAACGGGATATATTCCGCTAGAAGGTAC AAAGAACCATTGTGAGAAACTTATTTCTGATTGTGACCGTTCACAAAC AAAGATGAAGCTCTCCTTTCACGGAGCTGTTTAGAAGCAGTCCTTTCG TGAAACTGTGACGGTATATTTTGAAAGATATCAAGACTTCATTGGGAA TGGTAACATTTTCATTTAACATCACGGCAGGCGAAGAATTGTCAGAAT CTTCTGGTAATACGTACATTCAGGTAACATAGCTGAATTGTGACCGCT CACAAGCAGTTTGGAAGCACCCTTCTTTTATATTGTGAACGATCACAT TGGATCGTCTGGACGCCTATGGTGCAAAAGGATATAACTTCCCAGAAA ACCTAGGCAGCAGCTATCTCACAATCTGCTTTGAGATATTTGCATGCA GATAACGGAGTTCAACTTTTCTATTGGCAGAACAGATTTGTAACACTC TTTTTGTGGATCCGCAAGTGAATATTGGTTAGCCTGAAGGTTTTTGTT AGAAGCGCGATAACGTAGAATAAGTAAAGAGCACCATCTTCAGTAACT TCCTTGAGATTGTGACCGTTCACACGCATAGTTCAACCTTCTCTTTGG TAGAGCAATTTTCAAACCCTCTTCTTAATCTAGAAGTGAATATAAGAA CACCTCTCACCTATACGGGGAGGAAGGTAAATGTCTTAAATTAAAAAT TATACAGGAAGAATTGTCAAAAGCTGTTTGTGAGCGTTCACACCGCAA ACCGAGTTGCATTTTTCGTTTGGTACAGTAGTTCAGAAAATCACATTT GTGAAATGTGAACGGTCACATAGGATTGAGTTCAAGACTTCGATGGAA CCGGGTATATCTCCACATCTAATCTGACAGAAACATACTCAGAAAAGT GTTTCTGTTGTATGCATTGAACTCGTACAGTTGCACATTCAGTTCCAG 76 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: AAAGTAGCTATGGAGCACTGTTTTCGTGTTGCGCCAGTTGATACTTGT GACCGTTCACAACCGACCGTCAAGAAGGAAATATCTTCTCATATCCCC TAGCAGATACTTTCTCATAAACCCCTTCATCACATAAGCACACACGTA ACGGAGTAGGGCAACTTTCTTGTGCCAAAGTAGATTTCATGCACTTTT CTTATAATCTACAGGTAGACATTTGAATAGATGTCAAGAATTCTTTTG TGAGCGGTCACACTTTCTATGAAAACTGGACGGAAGTATTCTCTGAAA CTACTCTGTGTTGTTTGCATTGAAGTCACACAGCTGAATATCGCTTTT CACAGAACAGGTATGCAATGCTCATTTGGTGTTAGGGCAGGGGACGAT TCAGATGGGTTAAGACCGATGCTGATACAGCGAGTATCTACCCCCACC AGCAAGGAAGAAGAATTCCGTGAAACATGTATGTGCTTGTGATCGTTC ACAGACAGTGTTCAACTTTCCTGTTTGCAAAGCATTTTTGGAAAACAC TTCTTTTAAATATGCAAGAGGTTATTTGCATACATTCCAGCATGTCGA TGGACACCGGGATATCTTAATAGAAAAACTTGATAAAAAAGCCTTCTG AGAAACATCTTTCTGTTATCTGCTTTGAAATCAGAGGGTTAAATTGTG ACCGTTCACATGTCGGATTAAAACTCTGTTTGTGAGATTGTGAACGGT CACATTTGAGAGCTTTCACACCTACAGTGAGAAGGTAAACATCTTCTC ATCAAAAGTAGACGGAGGCAGTCTCAGAATTTTCATTGGGATATATGG ACACAGCTTACAGATTTGAGCCTTTGTATTGAGAGAGTAGTTTAGAAA GAGCCTTTCTGTGAAACTACAAATGGGTATTTCGAGAGCTGGGGGGAC TTCTTTGTAAATGGAATTAAGTATGAAAGGTAGAACTGCAGTATCCTG AGAATCTACTCTGTGATTGTGATCGTTCACACAAAGTGTTGAGCAATC CTTTTCATACATCAGGTTTGAAGCACTGTTTTGAGTCTGGAGGTGAAC TTTTGGAGAGGTTGCAGGTCCATAGTTAGAAACGATATTAACTTCCAA TAGAAACCAGTCACAAGCGTTCTGATGAAGTTTTTTTGTGACCGTTCA CACAACTACCACAGGGGGTTCTTACTTTCGAAAGGGCAATTCTGAGAA ACCCTATTTTTAATTGTGATCGGTCACATTTGATACATGTGAGGATTT TGTTAGAAAAGGGAAGATCTTTATAACACATCTACACAGAAGTATTCA CAGAAACATCGTTGCGATATTAGCATTCAACACACAGTGTTGAAAATT CCCTTTGAGAGTGCATCTTCGAGGCACTCTATTTTATAAGTTCAACTG GACATTTGTGATCGTTCACAGTCTGCGTTGGAAATGCATATATCTCCG ATAAACAGTAGAGAGAAGCAATCTCAGAACCTTCTTTGTGGCGAATGT ACTTACCTAGCACAGAGGATATACTTTCCGTTAGAGAGGGCTGTATTG GTACAATCCTTGTGAGAACCTGTAAATGAATATTAGGAAAGCTTTGAG GATATCGCTTGTGACCGTTCACACTTCGTATATAATTTATACAAAAGC ATTCTCAGACACTGCTGTGTGACGTCTACATTCATGTGACAGTGTAGA ACGTTCCCTTTAATGGAGGAGGTTCGAGACACTCTCTTTTTATACATA GATGTGGACGTATCGAACAGTATGGGGACCGTGGTAATAATGGAAAAT CTTTCCCTGCAACCTGGAAGGAAGCCTTCTTTGAAACTAGTTGGTGAC TGTGATCGTTCACAAATAGAATTGATCCATTATTTCTAAAGAACAGTA TTGATACTCTATTTCTGCAGATCTACGTGGAGAGATTTGGGTAGCTTT GAGGGTTTTGTTGTAACCGAGATTAACTTCTTATCAAATTTCAACGAC GAATCACTCTCAAAAACTACCCTGAGAGATGGGCACTCGAGGCACCGA CTTGAATTGTGATCGTTCACAACTAAGTATGGAACAATCCTTTGGTGT GTTGTGAACGGTCACATTGTAGCACCCTGGCGTCTACGCTGAAGAGGG CAATCTCTTCCAATAAATACCAGACAAAAGGAATCTGAGAATCCTCTT TTGGATACATGCGCGGAGCTAGCAGAGCTGAATCTTTCTATCGACTGA GCCGTTTTTAAACCGTCTTTCTGTGAATTTGGAAGTAGACATTTGTAT AGATTAGACGATATCGCTGGTAACAGGGTTACATATACAAATTAGCCA CAGCATACTCAAAAATTTATTTTTGATTGTGACCGTTCACAGACCGAG ATGAAAATTTCCTTACGTGCAGTAGTTTAGAAATACTCTGTCGTGACT GGAAGTGAGCAGTACGACGGCATTTAGCACTGTGTTGGGAAAGAAAAT 77 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: CTATCTACACATAAAAGCTGGACGGAAACATACTCGTAACCTGATTTG TGAGCGGTCACACAGATAACAGACGTAGACCTTTGTTTTCATTGATCA GTTTTGAAAGACATTTTTGTGATTGTGAACGATCACATTGCATAGCTT AGAACATTTCTTTGAAAACAGGAATATTTTCCTATAAAATGTAGACAC AAGCAGTCTCAGAAACGGCTTAGTAATGCTTCCATTCAACTGATGGAT TTGAACTTTCCGCTTCAAAGACCAGTTTCAAGTACTCTTCTTGAGATC TGGAAGGGGATATTTGTGACCGCTCACAGCTTATGGGGATAAAACAAG TATCTTTCCTTAACAACCAGACAAAAAGATTGTCAGAAATTCATTTAC GATGTGTGCAGTCAACTAATAGAAAAAATCGCCGTCCTCTTCACGAAC ATTTTGGAAACACACTATTGGTGTCAACGGGGAGTCATTTAAATATAT GCCAAGTATTTTTATGTGAGCGGTCACAATTTCTAAGAAGTCCGGAGG GTAGTATTTTCTGAAACAACTCTGCGTTATTTGCTTTGAAAACACGCA CCTCAATACCGGTTTTCCCACAACTGGTATCCAGTGATCATTCGGGTG GGACACGGGACTATTTAGTTGTGTGAAAACTGACGCTGATGCACCGGG TATATACCGCCAGCAGGAACGAAGAACAATTGCGAGAAACATATATCT GCTTGTGACCGTTCACAGACAATGATCAAGTTCCCCGTTCGCGAAGCT TTTTAGGAAAAGACCTCTCTTAAAATGTAACAGTTTATTTTCAAACAT ACCAACACGTCAATGGGCATCGTGACGTTTTAATTGAACAACATGGTA GAGACAAAGACTTGTGAGAATCATTTGCTATTACCTACTTCGAGATAA GATGGCTAAATTGTGACCGCTCACATGCCGTATGAAAGCTCCGTTCGT TAATTGTGAACGATCACATTTGATAGTTTGCATACCTATAGTGCGAAA GTATACAACTTCTCAGCAAACGTAGGCGGCGGCTGTCTCACAATTTGC ATTGAGATGTTTGGATACAGATTACGGATTTCAGCTTTTGTATTGGGA GAATAGATTAGTAAGACCCTTTTTGTGAACCACAGATGAGTATTCGTG AGCCGGAGGGTCTTTTTTATAAGTGCAATAAAGTAGGATAGGTAAAAG TGCACTATCTTGAGTATCAACCCTGAGATTGTGATCGTTCACACGAAT TGTTCAGCCATCTTTTTGATAGATCAAGTTTCAAGCCCTGTTTATCTA GAGGTGAATTTATGGAGACGTCGCACGTACACAGGTAGGAACGTTAAT TGACTTACATTAGAAATCATTCAGCAAGAGTTGTGAAGAGGTTTTTGT GACCGTTCACACCACAACCCCAGTGGCTTTTTACGTTCGGAACGGTAA TTCAGAGAATCCCAATTTAAATGTGATCGGTCACATATGATTCAGGTC AGGACTTTGATAGAACAGGGTAGATCTCTACAACTCATCTCACAGAAA TATACACAGAAAAATGGTTCCGTTATAAGCATTGAACACGCACTGTTG CAAATTCACTTCGAGAATGTATCTACGGGGCACTGTATTCTTAGCTCC ACTTGACACTTGTGATCGTTCACAATCGGCCTTCGAGATGGATATGTC TCTGATATACCGTAGGAGATGCTATCTCATAACCCTCTTCGTCGCAAA AGTACATACGTAGCGCAGTGGGTGACTAATTTTCGTGAGCGAAGGTTG AATTCGTGCAATTCTCGTAAACCAGTGAAGGATTATTAAGAAATCTTC GAGGTTATTGCATGTGACCGTTCACAATTCGTAAATAGTTCATAGAAT AGCATCTTCAGACACAGCTGTGCGACATCTACTTTCATATGACGGTCT ACAACGCTCGCTTTACTACAGGTGGTTCCAGGCAATCTCTCTTAACGT AAATCTGGATTTATTGATCATTAGGGAGATCGCGGTAATGATCGAGAA TATTTCGCTGGAACGTGCAAGGAACCCTTGTTAGAAACTAATTGCTGA CTGTGACCGTTCACAAATAAAAATGATGCACTACTTCCAAGGAACTGT ATAGATGCTGTACTTCCGCGAACTATGTCGATAGATTTTGGAAGCTAT GAAGGCTTTATTGTGACTGATATCAATTTCTTTTCACATTACAGCGGA GCCGAATAACTGTCAAAATCTACCGGAAAGACGGACACCCGGGGAACC TACCTGAATTGTGATCGCTCACAACCAATTAGGGAGCAACCCTCTGTT TGTTGTGAACGATCACATTGTATCACCCGGGTGTCTATGCTGCAGAAG GCTATCACTTCCAAGAAATCCCAGGCAACAGGTATCTGACAATCCGCT TTTAGATACTTGCGTGGAGATAGCGGAGCTCAATTTTTCTATCGGCTG 78 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: AACCGATTTTTAACCCTCTTCTTGTGATTCGGAGGTAAACATTGTTTA GACTAAACGTTATTGCTAGTAGCACGGTAACATAGACTAATTAACGAC ACCATATTCAATAATATACTTTAGATTGTGACCGTTCACAGGCCTAGA TCAAACTTTTCTTAGGTGGAGTAATTTACAAATCCTCTGCTACTAGAA GTGAGTAGAACAACGCCACTTACCAATGCGTGGGGGAAGATAATACTG TCTAAACTTAAAAGTTGTACGGGAAAAATAGTCGAAGCCGATTGTGAG CGGTCACACCGAAAAGCGACTTACACTTTTGGTTTCGTTCATTAGTTT AGAAAGTCATATTGGAATGTGAACGATCACATAGCATTGCGTACAACA CTTCTATGAAACCAGGTATATTTCCCCATATAATGTGACACAAACAGA CTCAGAGAAGGGTTACTATTGCATGCATTGAACTGGTGCATTTGCACT TTCAGCTCCAAAAACTAGTATCGAGTACTGTTCTCGGTCCGCCAGGTG ATGCTTGTGACCGCTCACAACTGATCGGCATGAAAGAAGTGTCTTTTC TTATCACCCAGCAAATAGTTTGTCATAAATCCATTCACCATATGAGCA GACAAGTAATGGAATAAGTGCCGACGTTCTCGTCCCAAATATATTGCA AGCACATTACTGATACCTATAGATAAACATTAGAAAAGATTTCAGGAA ATCTCTTGTGAGCGTTCACACTTTGTATGTAAATTGTACGAAAGTATC CTCTGACACTACTGTGTGTCGTTTACATTGATGAGACACTGCAGAATG TCCCTTTTAACGGAAGAGGTACGCGATACTCACTTGTTTCAGAGCTGG GGACGAATCAAATAGGATAGGAACGGTGCTAATACTGCAAGATCTATC CCCGCCACCAGGGAGGAAGAATTCCTTGAAACAAGTAGGTGCCTGTGA TCGTTCACAGATAGTATTCATCTATCATGTCTGAAAAACATTATTGGT AATCAATTCCTTCAATATACATGAAGTGATTTGCGTACCTTCGAGCGT GTTGATGTACCCCAGGTTGACTTATTAGCAAAATTTAATGAAACAAAT CCCTCTGAAAAACAACTCTCAGTGATCGGCTCTGGAAGCAGCGGCTTA AATTGTGATCGTTCACATCTCAGAATAGAACTATGCTTGGGGGTTGTG AACGGTCACATTTTAGAACTCTCGCATCTACACTGAGGAGGTCAACCT CTTCTAATCAATAGCAGACGAAGGGAGTCTGAGAATTCTCATTTGGAT ACATGGGCAGAGCTTGCAGATCTGAGTCTTTGTATCGAGTGAGTCGTT TATAAAGCGCCTTCCTGTAAATTAGAAATAGGCATTTCTAGAGATGAG GCGACATCTCTGTTAATAGAGTTAAATATGCAAGTTAGACCTCAGTAT ACTGAAAATTTAATCTTTGATTGTGATCGTTCACAGAACGTGATGAGA AATTCTTTACATGCATTAGGTTAGAAGTACTGTGTGGCTGGAGGTGAG CTGTTCGAGGGGATTTAGGACCGTATTTGGAAACAATATTCTAACTAC CCATAGAAGCCGGTCGCAAACGTACTGGTGAGCTTATTTGTGACCGGT CACACAAATACGACACGGAGTCCTTAGTTTCCAATGGTCAATTTTGAA AGACCTTATTTATTGTGATCGATCACATTTCATACCTGAGAGCATTTT TTTAAAAAAAGGAAGATTTTTCTAAAACATGTACACACAAGTAGTCAC AGAGACAGCGTAGCAATACTACCATTCAACAGACGGTTTTGAAATTTC CCCTTGAAAGTCCATTTCCAGGTACTCTACTTAAACTTGAACGGGACG TTTGTGATCGCTCACAGTTTGTGTGGGTAATACATGTATCTTCGTTAA AAAGCAGAGAAAAGGAATGTCAGAACTTTATTTGCGGTGAGTGTAGTT AACTAGTACAAAGAATTACTGCTGTCCGCTACAGGGACTTTATGGGAA CAAACCATGGGGCCAATGGAGAATCATTAAAAAATATTCCAGGTAATT TCATGTGAGCGTTCACAATTTGTAAGTAGATCGTAGGATAGTATCTTC TGACACAACTGTGCGTCATTTACTTTGATAAGACGCTCCACAATGCCC GTTTTACCGCAAGTGGTACCCGGTAATCACTCGTCGGAACTCGGGATT AATTAATTATGAGAGAAATGGCGCTAATGCTCCAGGATATATCGCCGG CACGAGCGAGGAACACTTGCTAGAAACAAATAGCTGCCTGTGACCGTT CACAGATAATAATCATGTACCACGTCCGAGAAACTTTATAGGTGATGA ACTCCCTCAAATATATCAATTGATTTTCGAACCTACGAACGCGTTAAT GTGCCTCATGTCGATTTATTTGCACAATATAGTGAAGACCAAATACCT 79 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: GTGAAAATCAATCGCAATGACCGACTCCGGGAGAAGCTGCCTAAATTG TGATCGCTCACATGCCATAAGAGAGCTACGCTCGGTGTTGTGAACGAT CACATTTTATAATTCGCGTATCTATACTGCGGAAGTCTACCACTTCTA AGCAATCGCAGGCGACGGGTGTCTGACAATTCGCATTTAGATGCTTGG GTAGAGATTGCGGATCTCAGTTTTTGTATCGGGTGAATCGATTATTAA GCCCCTTCTTGTAATCAGAGATAAGCATTCTTGAGACTAAGCGTCATT TCTATTAGTACAGTAAAATAGGCTAGTTAAACGTCACTATATTGAATA TTAAACCTTAGATTGTGATCGTTCACAGGACTTGATCAGACATTTTTT AGATGGATTAAGTTACAAGTCCTGTGCTAGAGGTGAGTTGATCAAGGC GACGTACGAACGCATGTGGGAACATTATATCTGACTAACCTTAGAAGT CGTTCGGCAAAAGTAGTGGAGGGCTATTGTGACCGGTCACACCAAAAC GCCACTGACTCTTTAGGTTCCGATCGTTAATTTAGAGAGTCCTAATAA TGTGATCGATCACATATCATTCCGGACAGCACTTTTATAAAACAAGGT AGATTTCTCCAAATCATGTCACACAAATAGACACAGAGAAAGGGTACC ATTACAACCATTGAACAGGCGCTTTTGCAATTTCACCTCGAAAATCTA TTACCGGGTACTGTACTCAGCTGCACGTGACGCTTGTGATCGCTCACA ATTGGTCTGCGTGATAGATGTGTCTTTGTTATAACGCAGGAAATGGTA TGTCATAACTCTATTCGCCGTAAGAGTAGATAAGTAGTGCAATGAGTT GACCTGATGTTCGCGACCGAGATTTAATGCGAGCAAATCACGGA 10 HJURP-mCherry- MVSKGEEDNMAIIKEFMRFKVHMEGSVNGHEFEIEGEGEGRPYEGTQT LacI AKLKVTKGGPLPFAWDILSPQFMYGSKAYVKHPADIPDYLKLSFPEGF KWERVMNFEDGGVVTVTQDSSLQDGEFIYKVKLRGTNFPSDGPVMQKK TMGWEASSERMYPEDGALKGEIKQRLKLKDGGHYDAEVKTTYKAKKPV QLPGAYNVNIKLDITSHNEDYTIVEQYERAEGRHSTGGMDELYKSGGD RWSSTGGGRSRENLYFQGAAKFKETAAAKFERQHMDSGGGGTSLEMLG TLRAMEGEDVEDDQLLQKLRASRRRFQRRMQRLIEKYNQPFEDTPVVQ MATLTYETPQGLRIWGGRLIKERNKGEIQDSSMKPADRTDGSVQAAAW GPELPSHRTVLGADSKSGEVDATSDQEESVAWALAPAVPQSPLKNELR RKYLTQVDILLQGAEYFECAGNRAGRDVRVTPLPSLASPAVPAPGYCS RISGKSPGDPAKPASSPREWDPLHPSSTDMALVPRNDSLSLQETSSSS FLSSQPFEDDDICNVTISDLYAGMLHSMSRLLSTKPSSIISTKTFIMQ NWNCRRRHRYKSREFPRGSGSGSMVKPVTLYDVAEYAGVSYQTVSRVV NQASHVSAKTREKVEAAMAELNYIPNRVAQQLAGKQSLLIGVATSSLA LHAPSQIVAAIKSRADQLGASVVVSMVERSGVEACKAAVHNLLAQRVS GLIINYPLDDQDAIAVEAACTNVPALFLDVSDQTPINSIIFSHEDGTR LGVEHLVALGHQQIALLAGPLSSVSARLRLAGWHKYLTRNQIQPIAER EGDWSAMSGFQQTMQMLNEGIVPTAMLVANDQMALGAMRAITESGLRV GADISVVGYDDTEDSSCYIPPLTTIKQDFRLLGQTSVDRLLQLSQGQA VKGNQLLPVSLVKRKTTLAPNTQTASPRALADSLMQLARQVSRSSLRP PKKKRKV 11 Exemplary non- ATCAGCGAGGGGTTATTTAGATATCTGCGAAGTTTTTGTATGTGACCG identical GTCACAATTCCTAAAAAGTCCAGAGAGTAGCATTTTCAGAAACAGCTC repetitive unit 1 TGCGATATCTGCTTTCAAATCACGGACTTCAACACTGGCTTTCCTACA GCTGGTTTCAAGCGATCTTTCTGAG 12 Exemplary non- ATCTACAGGTAGACATTTGAATAGATGTCAAGAATTCTTTTGTGAGCG identical GTCACACTTTCTATGAAAACTGGACGGAAGTATTCTCTGAAACTACTC repetitive unit 2 TGTGTTGTTTGCATTGAAGTCACACAGCTGAATATCGCTTTTCACAGA ACAGGTATGCAATGCTCATTTGGTG 80 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: 13 Exemplary non- AACCTGTAAATGAATATTAGGAAAGCTTTGAGGATATCGCTTGTGACC identical GTTCACACTTCGTATATAATTTATACAAAAGCATTCTCAGACACTGCT repetitive unit 3 GTGTGACGTCTACATTCATGTGACAGTGTAGAACGTTCCCTTTAATGG AGGAGGTTCGAGACACTCTCTTTTT 14 Exemplary non- TCAACGGGGAGTCATTTAAATATATGCCAAGTATTTTTATGTGAGCGG identical TCACAATTTCTAAGAAGTCCGGAGGGTAGTATTTTCTGAAACAACTCT repetitive unit 4 GCGTTATTTGCTTTGAAAACACGCACCTCAATACCGGTTTTCCCACAA CTGGTATCCAGTGATCATTCGGGTG 15 Exemplary non- ACCAGTGAAGGATTATTAAGAAATCTTCGAGGTTATTGCATGTGACCG identical TTCACAATTCGTAAATAGTTCATAGAATAGCATCTTCAGACACAGCTG repetitive unit 5 TGCGACATCTACTTTCATATGACGGTCTACAACGCTCGCTTTACTACA GGTGGTTCCAGGCAATCTCTCTTAA 16 Exemplary non- ACCTATAGATAAACATTAGAAAAGATTTCAGGAAATCTCTTGTGAGCG identical TTCACACTTTGTATGTAAATTGTACGAAAGTATCCTCTGACACTACTG repetitive unit 6 TGTGTCGTTTACATTGATGAGACACTGCAGAATGTCCCTTTTAACGGA AGAGGTACGCGATACTCACTTGTTT 17 Exemplary non- CCAATGGAGAATCATTAAAAAATATTCCAGGTAATTTCATGTGAGCGT identical TCACAATTTGTAAGTAGATCGTAGGATAGTATCTTCTGACACAACTGT repetitive unit 7 GCGTCATTTACTTTGATAAGACGCTCCACAATGCCCGTTTTACCGCAA GTGGTACCCGGTAATCACTCGTCGG 18 Exemplary HOR ATCAGCGAGGGGTTATTTAGATATCTGCGAAGTTTTTGTATGTGACCG 1 GTCACAATTCCTAAAAAGTCCAGAGAGTAGCATTTTCAGAAACAGCTC TGCGATATCTGCTTTCAAATCACGGACTTCAACACTGGCTTTCCTACA GCTGGTTTCAAGCGATCTTTCTGAGATGTGAAACTGGACTTTTTGGTC GTTTGGAAGCTCACGGTGATGAACGGGATATATTCCGCTAGAAGGTAC AAAGAACCATTGTGAGAAACTTATTTCTGATTGTGACCGTTCACAAAC AAAGATGAAGCTCTCCTTTCACGGAGCTGTTTAGAAGCAGTCCTTTCG TGAAACTGTGACGGTATATTTTGAAAGATATCAAGACTTCATTGGGAA TGGTAACATTTTCATTTAACATCACGGCAGGCGAAGAATTGTCAGAAT CTTCTGGTAATACGTACATTCAGGTAACATAGCTGAATTGTGACCGCT CACAAGCAGTTTGGAAGCACCCTTCTTTTATATTGTGAACGATCACAT TGGATCGTCTGGACGCCTATGGTGCAAAAGGATATAACTTCCCAGAAA ACCTAGGCAGCAGCTATCTCACAATCTGCTTTGAGATATTTGCATGCA GATAACGGAGTTCAACTTTTCTATTGGCAGAACAGATTTGTAACACTC TTTTTGTGGATCCGCAAGTGAATATTGGTTAGCCTGAAGGTTTTTGTT AGAAGCGCGATAACGTAGAATAAGTAAAGAGCACCATCTTCAGTAACT TCCTTGAGATTGTGACCGTTCACACGCATAGTTCAACCTTCTCTTTGG TAGAGCAATTTTCAAACCCTCTTCTTAATCTAGAAGTGAATATAAGAA CACCTCTCACCTATACGGGGAGGAAGGTAAATGTCTTAAATTAAAAAT TATACAGGAAGAATTGTCAAAAGCTGTTTGTGAGCGTTCACACCGCAA ACCGAGTTGCATTTTTCGTTTGGTACAGTAGTTCAGAAAATCACATTT GTGAAATGTGAACGGTCACATAGGATTGAGTTCAAGACTTCGATGGAA CCGGGTATATCTCCACATCTAATCTGACAGAAACATACTCAGAAAAGT GTTTCTGTTGTATGCATTGAACTCGTACAGTTGCACATTCAGTTCCAG AAAGTAGCTATGGAGCACTGTTTTCGTGTTGCGCCAGTTGATACTTGT GACCGTTCACAACCGACCGTCAAGAAGGAAATATCTTCTCATATCCCC TAGCAGATACTTTCTCATAAACCCCTTCATCACATAAGCACACACGTA ACGGAGTAGGGCAACTTTCTTGTGCCAAAGTAGATTTCATGCACTTTT CTTATA 81 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: 19 Exemplary HOR ATCTACAGGTAGACATTTGAATAGATGTCAAGAATTCTTTTGTGAGCG 2 GTCACACTTTCTATGAAAACTGGACGGAAGTATTCTCTGAAACTACTC TGTGTTGTTTGCATTGAAGTCACACAGCTGAATATCGCTTTTCACAGA ACAGGTATGCAATGCTCATTTGGTGTTAGGGCAGGGGACGATTCAGAT GGGTTAAGACCGATGCTGATACAGCGAGTATCTACCCCCACCAGCAAG GAAGAAGAATTCCGTGAAACATGTATGTGCTTGTGATCGTTCACAGAC AGTGTTCAACTTTCCTGTTTGCAAAGCATTTTTGGAAAACACTTCTTT TAAATATGCAAGAGGTTATTTGCATACATTCCAGCATGTCGATGGACA CCGGGATATCTTAATAGAAAAACTTGATAAAAAAGCCTTCTGAGAAAC ATCTTTCTGTTATCTGCTTTGAAATCAGAGGGTTAAATTGTGACCGTT CACATGTCGGATTAAAACTCTGTTTGTGAGATTGTGAACGGTCACATT TGAGAGCTTTCACACCTACAGTGAGAAGGTAAACATCTTCTCATCAAA AGTAGACGGAGGCAGTCTCAGAATTTTCATTGGGATATATGGACACAG CTTACAGATTTGAGCCTTTGTATTGAGAGAGTAGTTTAGAAAGAGCCT TTCTGTGAAACTACAAATGGGTATTTCGAGAGCTGGGGGGACTTCTTT GTAAATGGAATTAAGTATGAAAGGTAGAACTGCAGTATCCTGAGAATC TACTCTGTGATTGTGATCGTTCACACAAAGTGTTGAGCAATCCTTTTC ATACATCAGGTTTGAAGCACTGTTTTGAGTCTGGAGGTGAACTTTTGG AGAGGTTGCAGGTCCATAGTTAGAAACGATATTAACTTCCAATAGAAA CCAGTCACAAGCGTTCTGATGAAGTTTTTTTGTGACCGTTCACACAAC TACCACAGGGGGTTCTTACTTTCGAAAGGGCAATTCTGAGAAACCCTA TTTTTAATTGTGATCGGTCACATTTGATACATGTGAGGATTTTGTTAG AAAAGGGAAGATCTTTATAACACATCTACACAGAAGTATTCACAGAAA CATCGTTGCGATATTAGCATTCAACACACAGTGTTGAAAATTCCCTTT GAGAGTGCATCTTCGAGGCACTCTATTTTATAAGTTCAACTGGACATT TGTGATCGTTCACAGTCTGCGTTGGAAATGCATATATCTCCGATAAAC AGTAGAGAGAAGCAATCTCAGAACCTTCTTTGTGGCGAATGTACTTAC CTAGCACAGAGGATATACTTTCCGTTAGAGAGGGCTGTATTGGTACAA TCCTTGTGAG 20 Exemplary HOR AACCTGTAAATGAATATTAGGAAAGCTTTGAGGATATCGCTTGTGACC 3 GTTCACACTTCGTATATAATTTATACAAAAGCATTCTCAGACACTGCT GTGTGACGTCTACATTCATGTGACAGTGTAGAACGTTCCCTTTAATGG AGGAGGTTCGAGACACTCTCTTTTTATACATAGATGTGGACGTATCGA ACAGTATGGGGACCGTGGTAATAATGGAAAATCTTTCCCTGCAACCTG GAAGGAAGCCTTCTTTGAAACTAGTTGGTGACTGTGATCGTTCACAAA TAGAATTGATCCATTATTTCTAAAGAACAGTATTGATACTCTATTTCT GCAGATCTACGTGGAGAGATTTGGGTAGCTTTGAGGGTTTTGTTGTAA CCGAGATTAACTTCTTATCAAATTTCAACGACGAATCACTCTCAAAAA CTACCCTGAGAGATGGGCACTCGAGGCACCGACTTGAATTGTGATCGT TCACAACTAAGTATGGAACAATCCTTTGGTGTGTTGTGAACGGTCACA TTGTAGCACCCTGGCGTCTACGCTGAAGAGGGCAATCTCTTCCAATAA ATACCAGACAAAAGGAATCTGAGAATCCTCTTTTGGATACATGCGCGG AGCTAGCAGAGCTGAATCTTTCTATCGACTGAGCCGTTTTTAAACCGT CTTTCTGTGAATTTGGAAGTAGACATTTGTATAGATTAGACGATATCG CTGGTAACAGGGTTACATATACAAATTAGCCACAGCATACTCAAAAAT TTATTTTTGATTGTGACCGTTCACAGACCGAGATGAAAATTTCCTTAC GTGCAGTAGTTTAGAAATACTCTGTCGTGACTGGAAGTGAGCAGTACG ACGGCATTTAGCACTGTGTTGGGAAAGAAAATCTATCTACACATAAAA GCTGGACGGAAACATACTCGTAACCTGATTTGTGAGCGGTCACACAGA TAACAGACGTAGACCTTTGTTTTCATTGATCAGTTTTGAAAGACATTT TTGTGATTGTGAACGATCACATTGCATAGCTTAGAACATTTCTTTGAA 82 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: AACAGGAATATTTTCCTATAAAATGTAGACACAAGCAGTCTCAGAAAC GGCTTAGTAATGCTTCCATTCAACTGATGGATTTGAACTTTCCGCTTC AAAGACCAGTTTCAAGTACTCTTCTTGAGATCTGGAAGGGGATATTTG TGACCGCTCACAGCTTATGGGGATAAAACAAGTATCTTTCCTTAACAA CCAGACAAAAAGATTGTCAGAAATTCATTTACGATGTGTGCAGTCAAC TAATAGAAAAAATCGCCGTCCTCTTCACGAACATTTTGGAAACACACT ATTGGTG 21 Exemplary HOR TCAACGGGGAGTCATTTAAATATATGCCAAGTATTTTTATGTGAGCGG 4 TCACAATTTCTAAGAAGTCCGGAGGGTAGTATTTTCTGAAACAACTCT GCGTTATTTGCTTTGAAAACACGCACCTCAATACCGGTTTTCCCACAA CTGGTATCCAGTGATCATTCGGGTGGGACACGGGACTATTTAGTTGTG TGAAAACTGACGCTGATGCACCGGGTATATACCGCCAGCAGGAACGAA GAACAATTGCGAGAAACATATATCTGCTTGTGACCGTTCACAGACAAT GATCAAGTTCCCCGTTCGCGAAGCTTTTTAGGAAAAGACCTCTCTTAA AATGTAACAGTTTATTTTCAAACATACCAACACGTCAATGGGCATCGT GACGTTTTAATTGAACAACATGGTAGAGACAAAGACTTGTGAGAATCA TTTGCTATTACCTACTTCGAGATAAGATGGCTAAATTGTGACCGCTCA CATGCCGTATGAAAGCTCCGTTCGTTAATTGTGAACGATCACATTTGA TAGTTTGCATACCTATAGTGCGAAAGTATACAACTTCTCAGCAAACGT AGGCGGCGGCTGTCTCACAATTTGCATTGAGATGTTTGGATACAGATT ACGGATTTCAGCTTTTGTATTGGGAGAATAGATTAGTAAGACCCTTTT TGTGAACCACAGATGAGTATTCGTGAGCCGGAGGGTCTTTTTTATAAG TGCAATAAAGTAGGATAGGTAAAAGTGCACTATCTTGAGTATCAACCC TGAGATTGTGATCGTTCACACGAATTGTTCAGCCATCTTTTTGATAGA TCAAGTTTCAAGCCCTGTTTATCTAGAGGTGAATTTATGGAGACGTCG CACGTACACAGGTAGGAACGTTAATTGACTTACATTAGAAATCATTCA GCAAGAGTTGTGAAGAGGTTTTTGTGACCGTTCACACCACAACCCCAG TGGCTTTTTACGTTCGGAACGGTAATTCAGAGAATCCCAATTTAAATG TGATCGGTCACATATGATTCAGGTCAGGACTTTGATAGAACAGGGTAG ATCTCTACAACTCATCTCACAGAAATATACACAGAAAAATGGTTCCGT TATAAGCATTGAACACGCACTGTTGCAAATTCACTTCGAGAATGTATC TACGGGGCACTGTATTCTTAGCTCCACTTGACACTTGTGATCGTTCAC AATCGGCCTTCGAGATGGATATGTCTCTGATATACCGTAGGAGATGCT ATCTCATAACCCTCTTCGTCGCAAAAGTACATACGTAGCGCAGTGGGT GACTAATTTTCGTGAGCGAAGGTTGAATTCGTGCAATTCTCGTAA 22 Exemplary HOR ACCAGTGAAGGATTATTAAGAAATCTTCGAGGTTATTGCATGTGACCG 5 TTCACAATTCGTAAATAGTTCATAGAATAGCATCTTCAGACACAGCTG TGCGACATCTACTTTCATATGACGGTCTACAACGCTCGCTTTACTACA GGTGGTTCCAGGCAATCTCTCTTAACGTAAATCTGGATTTATTGATCA TTAGGGAGATCGCGGTAATGATCGAGAATATTTCGCTGGAACGTGCAA GGAACCCTTGTTAGAAACTAATTGCTGACTGTGACCGTTCACAAATAA AAATGATGCACTACTTCCAAGGAACTGTATAGATGCTGTACTTCCGCG AACTATGTCGATAGATTTTGGAAGCTATGAAGGCTTTATTGTGACTGA TATCAATTTCTTTTCACATTACAGCGGAGCCGAATAACTGTCAAAATC TACCGGAAAGACGGACACCCGGGGAACCTACCTGAATTGTGATCGCTC ACAACCAATTAGGGAGCAACCCTCTGTTTGTTGTGAACGATCACATTG TATCACCCGGGTGTCTATGCTGCAGAAGGCTATCACTTCCAAGAAATC CCAGGCAACAGGTATCTGACAATCCGCTTTTAGATACTTGCGTGGAGA TAGCGGAGCTCAATTTTTCTATCGGCTGAACCGATTTTTAACCCTCTT CTTGTGATTCGGAGGTAAACATTGTTTAGACTAAACGTTATTGCTAGT AGCACGGTAACATAGACTAATTAACGACACCATATTCAATAATATACT 83 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: TTAGATTGTGACCGTTCACAGGCCTAGATCAAACTTTTCTTAGGTGGA GTAATTTACAAATCCTCTGCTACTAGAAGTGAGTAGAACAACGCCACT TACCAATGCGTGGGGGAAGATAATACTGTCTAAACTTAAAAGTTGTAC GGGAAAAATAGTCGAAGCCGATTGTGAGCGGTCACACCGAAAAGCGAC TTACACTTTTGGTTTCGTTCATTAGTTTAGAAAGTCATATTGGAATGT GAACGATCACATAGCATTGCGTACAACACTTCTATGAAACCAGGTATA TTTCCCCATATAATGTGACACAAACAGACTCAGAGAAGGGTTACTATT GCATGCATTGAACTGGTGCATTTGCACTTTCAGCTCCAAAAACTAGTA TCGAGTACTGTTCTCGGTCCGCCAGGTGATGCTTGTGACCGCTCACAA CTGATCGGCATGAAAGAAGTGTCTTTTCTTATCACCCAGCAAATAGTT TGTCATAAATCCATTCACCATATGAGCAGACAAGTAATGGAATAAGTG CCGACGTTCTCGTCCCAAATATATTGCAAGCACATTACTGAT 23 Exemplary HOR ACCTATAGATAAACATTAGAAAAGATTTCAGGAAATCTCTTGTGAGCG 6 TTCACACTTTGTATGTAAATTGTACGAAAGTATCCTCTGACACTACTG TGTGTCGTTTACATTGATGAGACACTGCAGAATGTCCCTTTTAACGGA AGAGGTACGCGATACTCACTTGTTTCAGAGCTGGGGACGAATCAAATA GGATAGGAACGGTGCTAATACTGCAAGATCTATCCCCGCCACCAGGGA GGAAGAATTCCTTGAAACAAGTAGGTGCCTGTGATCGTTCACAGATAG TATTCATCTATCATGTCTGAAAAACATTATTGGTAATCAATTCCTTCA ATATACATGAAGTGATTTGCGTACCTTCGAGCGTGTTGATGTACCCCA GGTTGACTTATTAGCAAAATTTAATGAAACAAATCCCTCTGAAAAACA ACTCTCAGTGATCGGCTCTGGAAGCAGCGGCTTAAATTGTGATCGTTC ACATCTCAGAATAGAACTATGCTTGGGGGTTGTGAACGGTCACATTTT AGAACTCTCGCATCTACACTGAGGAGGTCAACCTCTTCTAATCAATAG CAGACGAAGGGAGTCTGAGAATTCTCATTTGGATACATGGGCAGAGCT TGCAGATCTGAGTCTTTGTATCGAGTGAGTCGTTTATAAAGCGCCTTC CTGTAAATTAGAAATAGGCATTTCTAGAGATGAGGCGACATCTCTGTT AATAGAGTTAAATATGCAAGTTAGACCTCAGTATACTGAAAATTTAAT CTTTGATTGTGATCGTTCACAGAACGTGATGAGAAATTCTTTACATGC ATTAGGTTAGAAGTACTGTGTGGCTGGAGGTGAGCTGTTCGAGGGGAT TTAGGACCGTATTTGGAAACAATATTCTAACTACCCATAGAAGCCGGT CGCAAACGTACTGGTGAGCTTATTTGTGACCGGTCACACAAATACGAC ACGGAGTCCTTAGTTTCCAATGGTCAATTTTGAAAGACCTTATTTATT GTGATCGATCACATTTCATACCTGAGAGCATTTTTTTAAAAAAAGGAA GATTTTTCTAAAACATGTACACACAAGTAGTCACAGAGACAGCGTAGC AATACTACCATTCAACAGACGGTTTTGAAATTTCCCCTTGAAAGTCCA TTTCCAGGTACTCTACTTAAACTTGAACGGGACGTTTGTGATCGCTCA CAGTTTGTGTGGGTAATACATGTATCTTCGTTAAAAAGCAGAGAAAAG GAATGTCAGAACTTTATTTGCGGTGAGTGTAGTTAACTAGTACAAAGA ATTACTGCTGTCCGCTACAGGGACTTTATGGGAACAAACCATGGGG 24 Exemplary HOR CCAATGGAGAATCATTAAAAAATATTCCAGGTAATTTCATGTGAGCGT 7 TCACAATTTGTAAGTAGATCGTAGGATAGTATCTTCTGACACAACTGT GCGTCATTTACTTTGATAAGACGCTCCACAATGCCCGTTTTACCGCAA GTGGTACCCGGTAATCACTCGTCGGAACTCGGGATTAATTAATTATGA GAGAAATGGCGCTAATGCTCCAGGATATATCGCCGGCACGAGCGAGGA ACACTTGCTAGAAACAAATAGCTGCCTGTGACCGTTCACAGATAATAA TCATGTACCACGTCCGAGAAACTTTATAGGTGATGAACTCCCTCAAAT ATATCAATTGATTTTCGAACCTACGAACGCGTTAATGTGCCTCATGTC GATTTATTTGCACAATATAGTGAAGACCAAATACCTGTGAAAATCAAT CGCAATGACCGACTCCGGGAGAAGCTGCCTAAATTGTGATCGCTCACA TGCCATAAGAGAGCTACGCTCGGTGTTGTGAACGATCACATTTTATAA 84 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: TTCGCGTATCTATACTGCGGAAGTCTACCACTTCTAAGCAATCGCAGG CGACGGGTGTCTGACAATTCGCATTTAGATGCTTGGGTAGAGATTGCG GATCTCAGTTTTTGTATCGGGTGAATCGATTATTAAGCCCCTTCTTGT AATCAGAGATAAGCATTCTTGAGACTAAGCGTCATTTCTATTAGTACA GTAAAATAGGCTAGTTAAACGTCACTATATTGAATATTAAACCTTAGA TTGTGATCGTTCACAGGACTTGATCAGACATTTTTTAGATGGATTAAG TTACAAGTCCTGTGCTAGAGGTGAGTTGATCAAGGCGACGTACGAACG CATGTGGGAACATTATATCTGACTAACCTTAGAAGTCGTTCGGCAAAA GTAGTGGAGGGCTATTGTGACCGGTCACACCAAAACGCCACTGACTCT TTAGGTTCCGATCGTTAATTTAGAGAGTCCTAATAATGTGATCGATCA CATATCATTCCGGACAGCACTTTTATAAAACAAGGTAGATTTCTCCAA ATCATGTCACACAAATAGACACAGAGAAAGGGTACCATTACAACCATT GAACAGGCGCTTTTGCAATTTCACCTCGAAAATCTATTACCGGGTACT GTACTCAGCTGCACGTGACGCTTGTGATCGCTCACAATTGGTCTGCGT GATAGATGTGTCTTTGTTATAACGCAGGAAATGGTATGTCATAACTCT ATTCGCCGTAAGAGTAGATAAGTAGTGCAATGAGTTGACCTGATGTTC GCGACCGAGATTTAATGCGAGCAAATCACGGA 25 Exemplary lexO CTGTATATATATACAG site 26 Exemplary lexO CTGTATGAGCATACAG site 27 Exemplary UAS CGGATTAGAAGCCGCCG site 28 Exemplary UAS CGGGTGACAGCCCTCCG site 29 Exemplary UAS AGGAAGACTCTCCTCCG site 30 Exemplary UAS CGCGCCGCACTGCTCCG site 31 Gamma Satellite gaattcccttgtggggctcgctgcttgtcccttcctccctgcctcagt insulator gtcaccctgagtccctttggctcgttcaggccacccagtggccccatt ttcgcctgtggaaacctgccacagacacaggcaacctattccaaagcc tgggacactacagcctatctgtgacagacctggggggcttctaggatg ggagaggctctccttgggagggtcgcagcattaccctgcagtcttgcc aattcttccctctactggtttcaacttccccctgagtaccctgtggcc tgtcgacgcacacactgcagccccttttccgcttgtgggagccttccg ctagagacacttgcaggcactctgcttcaaaacctgaagcattacagc acgcccgggacagtcctgggggcttctgggatgaaagaggcctccttg ggaggctcccagcattccatgtggtcctgctgcttctctcatctgcct gcctcaaagtgtccgtgagtccctgctgcccatgcacgccaccctgcc acctcattctctcttttgtatgtcttctgcgagagacacaggctccct gctctaaagactgggacataaaagcttgccgggacagccctgggtgct tctgggatgagagatgtccttgtgaggcttccagcattctctgaggtc ttgctgcttctcatctctgcctgccacaatttccacctgactccctgc 85 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: ggccctcccacgccaccctgcagccagcccacgccaccctgcggcctt gtttttgcttgtgggagtattttgcgagagacacaagcacgctgttcc aaagcctggggctttatagcccacccaggatagtttggggagcttctg ggatgggataggtcttttaggaagcctcccagcattccctgcggtctc gccgcttctttcctctgcctgcctcaacgtctccctgagtccctgtgg cccgcccaggccacactgcggcactgtttttgcttgagggggtcttcc aagagagacacagacactctgctccaaagtcttgggccttacagcctg tccaggacagtcctaggagcttctgggatgggagacacctccttggaa ggctctcagcattccctacggtctcgttgcttctcctctatgcttgcc tcaatgtctccctgggttcctacgatcggcccatgccacccagccgcc ctgttgtctctgttggggatcttctgctagagacacagtcgtcctgca ccaaagcatggggtattacagcccatccgaaacaaccatggtaagctt ctggggtgggagagacatccctaggaggctcccagcattccctgcggt ctcactgtcttctcacctctgcctgaccacaatgtccaccatagtccc tgcgggcctggccaatgccaccctgcctccctgttttcacttgtgggg gccatttacgagaaacacaggcatcctggctttgaagcctcgggcttt acatccctccaggaacaatcctggggtcttctgagatagaagattcat ccttaggaggctcccagcattccctgccgtctggactcttcccccctc tgcctgacctcaaggtcaaccctgggtctttggccgcccacccaagcc accatcaaccatgtggccccattttcggcttggtatgaaggccttcaa ggggtccgttgagggcagggagaggggagaattggatgagaacgcagg gaattgttgggagggctcccacagaggtctctcccatcacagaagctc ataggcctgtccgggacgggctctaatcccccaggctttggagcaggg ttcctgtgtctctcacagaagacccttaaaatagacacaggcaccccg ctccgagccctggggctttacagcctgcccccagacacccctggcggt ttctgggttggaagaggcctctttgggtcactcccagaattc 32 Reverse tRNA ggattacaggtgtgagctaccacacctttaaaattcctttaaaaatta insulator tactttattaaaatttatactatattctgattgctaaaattttattta gaatgtttgcatatctactaatttatgagattaccatgtaaactattt tttatgctgtttcattaggtgtgatgataaaaatctgtttcatcaaag gaaatgagacaatttatttaaaaggtaatataggccaggtgtggtggc tcacgcctgtaatcccagcactttgggaggccgaggcgggcagatcat gaggtcaggagttcgagaccagcctgaccaacgtagtgaaatcccgtc tctactaaaaatacaaaaattaacccggcatggtggtgcgtgcctgta atcccagctactcagaaggctgagacaggagaattacttgaacgtggg aggcagtggttgcagtgagccaagatcacgccattgcactccagcctg ggtgacagagcgagactccatctcaaaataaataaaataaaataaaat acactacaatataaagagcaataaaattaaaaaaaattttttttaaac tcactacattgcccaggctggtctcaaactcctggcctcaagcaagcc tgcctcagcctcccaaaatgctgtgattacaggtgtgagctaccacgc ccagccttttaaaatatatatatatataatgctgggcacggtgactca cacctgtaatcccagcactttgggagggcaaggcaaatcacgaggtca ggagttcaagaccagcctggccaacatggtgaaaccccaactctacta aaaaaatacaaaaaattatccaggcgtggtggcacgtgcctgtaatcc cagctactcgggaggccgagacaggagaatcacttgaacccaggaggc agaggttgcagtgagcacagatctcaccattgcactccagcttgggta acaagagtgaaactctgtctcaagaaaacaacaacaacaataaaaatg tatatgtatataatataataaataaaaaatatgtaatacatttgatat acattgctctgtcacccaggctagagtgctgtggtgcaatcatagctc actgcagccttgaactcctgggctcaagcaattctcctgcctcagcct ccccagtagctaggacttcaggtgcacactaccacctgaaaatttttt 86 IPTS / 200118216.1Attorney Docket No.: CETA-001WO SEQ Description Sequence IDNO: aaagtccaatttcaatagaaataaaagaaataataaataaaaccataa ggagataccaaaattgacaacttttaaaataatcccagttcaggataa aaaagggatacatggatgctggaagtataaatcagtacaaccttaaca ttgtacataaatacttcaatccaatggagcttggtggaggggagagga tgatagggtacaaatgatatttagtttattctttcctgcactttttaa gttttcttttctttttcatttctttttttttttttttgaggcagagtt tcactcttgttgtccaggctagagtgcaatggcacgatcttggctcac cgcaacctccacctcccaggttcaagcaattctcctgcctcagcctcc tgagtagctgggattacaggcacctgccaccacgcctggctaattttg tatttttagtagagacggggtttctccatgttggtcaggctggtctcg aactcctgacctcaggtgatccacccaccttggcctcccaaagtgctg ggattacaggcgtgagctaccactcccggcctaaattttctacatata taaaaatatattactttcataataaaatatgcatattcaaattgaaac ctcagtgcgatatcactgcacagccaccagaatggctaagatgggaag caacaaccatactgaatgttggtgggaacggggagtaactggaaccct catgcattactgggaggtgtataaattggttcaaccactttagaaaac caaagctaaatctagagccaccagcaatttcactcctagataaatact caacagaaacaagagcacatgtccatcaaaaaacctatacaagaaggc cagacacagtggctcacgcctgtaatcccagcacttcgggaggctgag gcagaaggatctcctgagctcaggagtgtaaggctagcctaggcaaca tggcgaaaccccatctctacaaaaaaatacaaaaactagctgggcatg gtggtgtgcacctgtagtcccagctacttatggggctgaggaaaggag cgcttgagccggggaggttgaagcaacagtgagccatgatcacaccac tgcactccggcctgggtgacacaacgagaccctatattaaagaaaaca aacaaaaaaaactgcacaagaatgttcattgaagttttagttggagta tcttcaaactagaatgcctgttttttagaactcaaatgtctattttgt ttttttgagatggagtttcactcttgtggcccaggctggagtgcagtg gcgcgatcttgactcactgcaaccttcacctcctgggttcaagcgact ctcctgcctcagcctcccgagtagttgggattataggcgcccactgcc acacccagctaatttttgcatttttagtagagacagagtttcaccatg ttggccaggttgatctcaaactcctgacctcaggtgatcctcctgcct tggaaggaaggaaggcaggcaggcagtccctgcgtagtggctcacacc cataatcccagctcttgggagactgaggtaggtggattacctgaggtc aggagttcaagaccagcctggccaacatggtgaaaccctgtctctact aaaaatacacaaaaattagccatgcatggtggcacatgcctgtagtcc cagcta 87 IPTS / 200118216.1
Claims
Attorney Docket No.: CETA-001WO CLAIMS What is claimed is 1. An artificial centromere comprising a nucleic acid of about 5 to about 50 kb, wherein the nucleic acid comprises: a. alpha satellite DNA comprising a plurality of non-identical repetitive units having between about 50% and about 99.9% sequence identity to each other organized into higher order repeats; and b. a plurality of recruitment motifs for an epigenetic modifier, wherein a recruitment motif is present at least once per 250 bp of the nucleic acid.
2. The artificial centromere of claim 1, wherein the nucleic acid is between about 5 and about 15 kb.
3. The artificial centromere of claim 1, wherein the nucleic acid is about 10 kb.
4. The artificial centromere of any one of claims 1-3, wherein the nucleic acid has between about 30% and about 99.9% sequence identity to a human centromere.
5. The artificial centromere of any one of claims 1-3, wherein the nucleic acid has between about 30% and about 60% identity to a human centromere.
6. The artificial centromere of claim 4 or claim 5, wherein the sequence of the alpha satellite DNA is from the reference genome CHM13.
7. The artificial centromere of any one of claims 1-6, wherein the recruitment motif is a LacO site.
8. The artificial centromere of claim 7, wherein the LacO site comprises the nucleic acid sequence of SEQ ID NO: 1, or a nucleic acid sequence having 1, 2, or 3 nucleotide substitutions as compared to SEQ ID NO:
1.
9. The artificial centromere of any one of claims 1-8, wherein the recruitment motif is present at least once per 100 to 200 bp of the nucleic acid.
10. The artificial centromere of any one of claims 1-9, wherein each repetitive unit is about 170 bp. 88 IPTS / 200118216.1Attorney Docket No.: CETA-001WO 11. The artificial centromere of any one of claims 1-10, wherein the higher order repeats comprise about 4 to about 12 repetitive units.
12. The artificial centromere of any one of claims 1-11, wherein the higher order repeat is between about 500 and about 2500 bp.
13. The artificial centromere of claim 12, wherein the higher-order repeat is between about 700 and about 2000 bp.
14. The artificial centromere of any one of claims 1-13, wherein the artificial centromere comprises from about 5 to about 15 higher-order repeats.
15. The artificial centromere of any one of claims 1-13, wherein the artificial centromere comprises about 7 higher-order repeats.
16. An artificial chromosome comprising the artificial centromere of any one of claims 1- 15.
17. The artificial chromosome of claim 16, further comprising one or more genes for expression in a cell.
18. The artificial chromosome of claim 17, wherein the one or more genes for expression in a cell are flanked by one or more insulator sequences.
19. The artificial chromosome of claim 17, wherein the one or more genes for expression in a cell are flanked by a first insulator sequence located 5’ of the one or more genes and a second insulator sequence located 3’ of the one or more genes.
20. The artificial chromosome of claim 19, wherein the first and the second insulator sequence comprise the same nucleic acid sequence.
21. The artificial chromosome of claim 19, wherein the first and the second insulator sequence comprise unique nucleic acid sequences.
22. The artificial chromosome of any one of claims 19-21, wherein the first and the second insulator sequences are each, independently, selected from nucleic acid sequences having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, 89 IPTS / 200118216.1Attorney Docket No.: CETA-001WO at least 85%, at least 90%, at least 95%, or 100% identity) to a nucleic acid sequence of SEQ ID NO: 31 or SEQ ID NO:
32.
23. A method of making an artificial centromere or an artificial chromosome, the method comprising synthesizing the artificial centromere of any one of claims 1-15 or the artificial chromosome of any one of claims 16-22.
24. A method of using an artificial chromosome, the method comprising, introducing the artificial chromosome of any one of claims 16-22 into a cell.
25. The method of claim 24, wherein the gene is expressed in the cell.
26. A cell comprising the artificial chromosome of any one of claims 16-22.
27. A method for designing an artificial centromeric sequence, the method comprising: a. identifying sequences from active alpha satellite arrays of human centromeric DNA having repetitive units and / or higher-order repeats; b. creating sequence alignments of the repetitive units and / or higher order repeats; c. creating a position weight matrix (PWM) from the sequence alignments; d. modifying the PWM to insert recruitment motifs for epigenetic modifiers; and e. designing an artificial centromeric sequence of about 5 to about 50 kb comprising a plurality of higher-order repeats, wherein each higher-order repeat comprises a plurality of non-identical repetitive units having between about 50% and about 99.9% sequence identity to each other.
28. The method of claim 27, wherein the artificial centromeric sequence is between 5 and 15 kb.
29. The method of claim 27, wherein the artificial centromeric sequence is about 10 kb.
30. The method of any one of claims 27-29, wherein the artificial centromeric sequence has between about 30% and about 99.9% sequence identity to a human centromere. 90 IPTS / 200118216.1Attorney Docket No.: CETA-001WO 31. The method of any one of claims 27-29, wherein the artificial centromeric sequence has between about 30% and about 60% identity to a human centromere.
32. The method of any one of claims 27-29, wherein the sequences from the active alpha satellite arrays of the human centromeric DNA are from a reference genome, wherein the reference genome is CHM13.
33. The method of any one of claims 27-32, wherein the recruitment motifs comprise LacO sites.
34. The method of any one of claims 27-33, wherein at least one recruitment motif is present per 100 to 200 bp of the artificial centromeric sequence.
35. The method of any one of claims 27-34, wherein each repetitive unit is between about 160 and about 180 bp.
36. The method of claim 35, wherein each repetitive unit is about 170 bp .
37. The method of any one of claims 27-36, wherein the plurality of non-identical repetitive units is from about 4 to about 12 units.
38. The method of any one of claims 27-37, wherein the higher-order repeat is between about 500 bp and about 2500 bp.
39. The method of claim 38, wherein the higher-order repeat is between about 700 bp and about 2000 bp.
40. The method of any one of claims 27-39, wherein the artificial centromeric sequence is computationally analyzed for stability or synthesizability.
41. The method of claim 40, wherein, if the artificial centromeric sequence does not pass a threshold measurement for stability or synthesizability, additional changes are made to improve the stability or synthesizability of the artificial centromeric sequence.
42. The method of any one of claims 27-41, wherein the method further comprises generating a nucleic acid comprising the artificial centromeric sequence.
43. A gene expression system comprising a cargo nucleic acid sequence and one or more insulator sequences. 91 IPTS / 200118216.1Attorney Docket No.: CETA-001WO 44. The gene expression system of claim 43, wherein the one or more insulator sequences flank the cargo nucleic acid sequence.
45. The gene expression system of claim 43 or 44, wherein the one or more insulator sequences comprise a first insulator sequence and a second insulator sequence.
46. The gene expression system of claim 45, wherein the first insulator sequence and the second insulator sequence comprise the same nucleic acid sequence.
47. The gene expression system of claim 45, wherein the first insulator sequence and the second insulator sequence comprise unique nucleic acid sequences.
48. The gene expression system of claim 45, wherein the first and the second insulator sequences are each, independently, selected from nucleic acid sequences having at least 70% identity (e.g. at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 95%, or 100% identity) to a nucleic acid sequence of SEQ ID NO: 31 or SEQ ID NO:
32. 92 IPTS / 200118216.1