Polynucleotides and methods for gene editing
Retron-based editors, with polynucleotide-guide RNA cassettes and vectors, address the inefficiencies of current genome editing methods by enhancing precision and efficiency through homologous recombination.
Patent Information
- Application Number
- PCT/US2024/061421
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-29
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-03
AI Technical Summary
Current genome editing technologies, such as CRISPR/Cas systems, are inefficient and require compositions and methods to enhance the precision and efficiency of genetic modifications.
The use of retron-based editors, comprising polynucleotide-guide RNA cassettes and vectors that include a polynucleotide comprising inverted repeat sequences, a guide RNA coding region, a donor DNA sequence, and a reverse transcriptase, to facilitate precise and efficient genome editing through homologous recombination.
Enhances the efficiency and precision of genome editing by promoting homologous recombination, reducing unwanted mutations, and improving the accuracy of genetic modifications.
Smart Images

Figure US2024061421_03072025_PF_FP_ABST
Abstract
Description
POLYNUCLEOTIDES AND METHODS FOR GENE EDITINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority to U.S. Provisional Application No. 63 / 616,017, filed December 29, 2023, which is hereby incorporated by reference in its entirety.REFERENCE TO SEQUENCE LISTING SUBMITTED ELECTRONICALLY
[0002] The instant application contains a Sequence Listing which has been submitted electronically in XML file format and is hereby incorporated by reference in its entirety. Said XML copy, created on December 19, 2024, is named “FIS-004WO_SL.xml” and is 31,751,809 bytes in size.FIELD
[0003] The embodiments provided herein relate to polynucleotides and methods of using the same for gene editing.BACKGROUND
[0004] Genome editing with engineered nucleases has made editing genomic sequences possible. Without being bound to any particular theory, for example, engineered nucleases may be used to generate site-specific double-strand breaks (DSBs) followed by resolution of DSBs by endogenous cellular repair mechanisms. The outcome may be either mutation of a specific site through mutagenic nonhomologous end-joining, creating insertions or deletions at the site of the break, or precise change of a genomic sequence through homologous recombination using an exogenously introduced donor template. One example of this, is the use of the CRISPR / Cas system.
[0005] Genome editing remains inefficient, and, therefore, there is a need for compositions and methods to facilitate more efficient genome editing. The present embodiments satisfy these needs as well as others.SUMMARY
[0006] Provided herein are retron-based editors, compositions comprising the same, methods for gene editing, and methods for treating a disease or disorder in a subject in need thereof.
[0007] In some embodiments, polynucleotide-guide RNA cassette is provided, wherein the cassette comprises: a) a polynucleotide comprising: i) a first inverted repeat sequence codingregion; ii) a ncRNA comprising an msr / msd locus comprising the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one SEQ ID NO provided in Table X or Table Y; iii) a donor DNA sequence located within the msr / msd locus; and iv) a second inverted repeat sequence coding region; and b) a guide RNA (gRNA) coding region.
[0008] In some embodiments, a vector is provided, wherein the vector comprises the polynucleotide-guide RNA cassette as provided for herein. In some embodiments, the vector comprises a promoter; a polynucleotide-guide RNA cassette such as those provided for herein; a reverse transcriptase (RT) coding sequence; a nuclease coding sequence; and at least one nuclear localization signal sequence (NLS), and wherein the promoter, the polynucleotide-guide RNA cassette, the RT, the nuclease, the reporter, and the at least one NLS are operably linked to each other.
[0009] In some embodiments, a polynucleotide is provided, wherein the polynucleotide comprises a polynucleotide transcript of the polynucleotide-guide RNA cassette such as those provided for herein, or the vector such as those provided for herein, and optionally a gRNA, wherein the polynucleotide transcript and gRNA are physically coupled.
[0010] In some embodiments, a plurality of polynucleotides is provided, wherein the plurality of polynucleotides comprises a polynucleotide transcript of the polynucleotide-guide RNA cassette such as those provided for herein, or the vector such as those provided for herein, and a gRNA, wherein the polynucleotide transcript and gRNA are not physically coupled.
[0011] In some embodiments, a polypeptide is provided, wherein the polypeptide comprises a reverse transcriptase (RT) comprising an amino acid sequence encoded by a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NO provided in Table X or Table Y; a nuclease; optionally a reporter; and optionally at least one nuclear localization signal (NLS).BRIEF DESCRIPTION OF THE DRAWINGS
[0012] FIG. 1 illustrates non-limiting embodiments of a guide RNA linked to a retron RNA in 5’ or 3’ orientation.
[0013] FIG. 2 illustrates non-limiting embodiments of reverse transcriptase (RT) fusion proteins, where RT is fused or linked to a CAS9 nuclease (e.g., wild-type or nickase mutant) in various configurations and other RT fusion proteins comprising a CAS9 nuclease and a singlestranded annealing protein (SSAP).
[0014] FIG. 3 illustrates non-limiting embodiments of a retron reverse transcriptase (RT) fused to dCas9 (catalytically inactive) and SSAP; or SSAP linked directly to RT; or as separate proteins.
[0015] FIG. 4 is a schematic diagram showing the secondary structure of the msd / msr non- protein-encoding RNA transcript of retron pFLCOOlOO (also referred to as FC100).
[0016] FIG. 5 is a schematic diagram showing the secondary structure of the msd / msr nonprotein-encoding RNA transcript of retron pFLC00039.
[0017] FIG. 6 is a schematic diagram showing the secondary structure of the msd / msr nonprotein-encoding RNA transcript of retron pFLC00002.
[0018] FIG. 7 is a schematic diagram showing the secondary structure of the msd / msr nonprotein-encoding RNA transcript of retron pFLC00132.
[0019] FIG. 8A shows HDR (homology directed repair) rate, as defined by percent GFP positive within mCherry positive cells. mCherry was used as a reporter for transfection. The retrons were 5’ guided (gRNA-rncRNA).
[0020] FIG. 8B shows HDR efficiency, as defined by HDR over total cutting rate. Total cutting rate is the sum of all HDR and NHEJ (non-homologous end joining) events.
[0021] FIG. 9 A shows HDR rate, as defined by percent GFP positive within mCherry positive cells. mCherry was used as a reporter for transfection. The retrons were 5’ guided (gRNA- rncRNA).
[0022] FIG. 9B shows HDR efficiency, as defined by HDR over total cutting rate. Total cutting rate is the sum of all HDR and NHEJ (non-homologous end joining) events.
[0023] FIGs. 10A-10B illustrate the results of a nickase screen. FIG. 10A shows precise repair (e.g., HDR rate, or the like), as defined by percent GFP positive within mCherry positive cells. The retrons were 5’ guided (gRNA-rncRNA). mCherry was used as a reporter for transfection. FIG. 10B shows comparison of the precise repair (e.g., HDR rate, or the like) in the Nickase screen with the average of FIG. 9 and FIG. 10 screens.
[0024] FIGs. 11 A-l IB illustrate the results of a VEGF screen. FIG. 11 A shows the HDR rate, as defined by next generation sequencing (NGS) using CRISPResso. The retrons were 5’ guided (gRNA-rncRNA). No adjustment for transfection rate was performed in this assay. FIG. 11B shows the comparison of the HDR efficiency at the VEGF locus, compared to the HDR efficiency at the BFP locus.
[0025] FIGs. 12A-12B illustrate the results of a K562 cell screen. The retrons were 5’ guided (gRNA-rncRNA). FIG. 12A shows the HDR rate, as defined by GFP+ percentage in mCherry-gated cells. FIG. 12B shows the comparison of the HDR efficiency.
[0026] FIG. 13 is a schematic diagram showing the secondary structure of the msd / msr non-protein-encoding RNA transcript of retron pFLCOOlOO. The arrows show the insertion point of heterologous targeting sequences for editing, which may involve deletion of a fragment of the msd region.
[0027] FIG. 14 is a bar graph showing pFLCOOlOO retron-mediated NHEJ and HDR efficiency in editing. The heterologous or targeting sequence are against the HEK3 locus in transfected HEK293T cells. Genomic DNA was harvested three days post-transfection and analyzed by NGS.
[0028] FIG. 15 is a schematic diagram showing the secondary structure of the msd / msr non-protein-encoding RNA transcript of retron pFLCOOlOO. The arrows show the A1A2 region which is the inverted repeat sequence of the 5’ and 3’ msr regions.
[0029] FIG. 16 is a schematic diagram showing the secondary structure of the msd / msr non-protein-encoding RNA transcript of retron pFLCOOlOO comprising only the R10 region of the A1A2 region. The arrows show the R10 A1A2 region which is the inverted repeat sequence of the 5’ and 3’ msr regions.
[0030] FIG. 17 is a bar graph showing pFLCOOlOO retron comprising only the R10 region of the Al A2 region -mediated NHEJ and HDR efficiency in editing.
[0031] FIG. 18 upper panel shows HDR rates between Cas9-pFLC00100 and pFLCOOlOO- Cas9 at a genomic locus. Lower panel shows HDR rates (defined as GFP + and NHEJ = BFP- / GFP-) with the Cas9-pFLC00100 orientation using the BFP, GFP conversion assay.
[0032] FIG. 19A is a bar graph showing results of the TTR analysis of pFLCOOlOO retron- mediated NHEJ and HDR efficiency in editing. The pFLCOOlOO retron was 5’ guided (gRNA- rncRNA) and 3’ guided with a range of templates of varying homology length.
[0033] FIG. 19B is a bar graph showing results of the VEGFA to GFP conversion assay of pFLCOOlOO retron-mediated NHEJ and HDR efficiency in editing. The pFLCOOlOO retron was 5’ guided (gRNA-rncRNA) and 3’ guided with a range templates of varying homology length.
[0034] FIG. 19C is a bar graph showing results of the BFP to GFP conversion assay of pFLCOOlOO retron-mediated NHEJ and HDR efficiency in editing. The pFLCOOlOO retron was 5’ guided (gRNA-rncRNA) and 3’ guided with a range templates of varying homology length.
[0035] FIG. 20 is a bar graph showing pFLCOOlOO retron-based gene editing locus specific template preference. Upper panel shows TTR (transfection followed by MiSeq analysis at TTR; Day 3 timepoint, readout was CRISPResso analysis); and the lower panel shows VEGFA (transfection followed by MiSeq analysis at TTR; Day 3 timepoint, readout was CRISPResso analysis).
[0036] FIG. 21 is a schematic diagram showing the secondary structure of the msd / msr non-protein-encoding RNA transcript of retron pFLCOOOOO.
[0037] FIG. 22 is a schematic diagram showing the secondary structure of the msd / msr non-protein-encoding RNA transcript of retron pFLC00006.
[0038] FIG. 23 is a schematic diagram showing the secondary structure of the msd / msr non-protein-encoding RNA transcript of retron pFLC00142.
[0039] FIG. 24 is a schematic diagram showing the secondary structure of the msd / msr non-protein-encoding RNA transcript of retron pFLC00146.
[0040] FIG. 25 is a schematic diagram showing the secondary structure of the msd / msr non-protein-encoding RNA transcript of retron pFLC00236.
[0041] FIG. 26 is a schematic diagram showing the secondary structure of the msd / msr non-protein-encoding RNA transcript of retron pFLC00241.
[0042] FIG. 27 is a schematic diagram showing the secondary structure of the msd / msr non-protein-encoding RNA transcript of retron pFLC00254.
[0043] FIG. 28 shows HDR rate, as defined by percent GFP positive within mCherry positive cells. mCherry was used as a reporter for transfection. The retrons were 5’ guided (gRNA- rncRNA).
[0044] FIG. 29 shows HDR efficiency, as defined by HDR over total cutting rate. Total cutting rate is the sum of all HDR and NHEJ (non-homologous end joining) events.DETAILED DESCRIPTION
[0045] As used herein and in the appended claims, the singular forms “a”, “an” and “the” include plural reference unless the context clearly dictates otherwise.
[0046] As used herein, the term “about” means that the numerical value is approximate and small variations would not significantly affect the practice of the disclosed embodiments. Where a numerical limitation is used, unless indicated otherwise by the context, “about” means the numerical value can vary by ±5% and remain within the scope of the disclosed embodiments. Thus, about 100 means 95 to 105.
[0047] As used herein, the term “animal” includes, but is not limited to, humans and nonhuman vertebrates such as wild, domestic, and farm animals. As used herein, the term “mammal” means a rodent (i.e., a mouse, a rat, or a guinea pig), a monkey, a cat, a dog, a cow, a horse, a pig, or a human. In some embodiments, the mammal is a human.
[0048] As used herein, the term “contacting” means bringing together of two elements in an in vitro system or an in vivo system. For example, “contacting” a therapeutic compound with an individual or patient or cell includes the administration of the compound to an individual or patient, such as a human, as well as, for example, introducing a compound into a sample containing a cellular or purified preparation containing target.
[0049] As used herein, the terms “comprising” (and any form of comprising, such as “comprise”, “comprises”, and “comprised”), “having” (and any form of having, such as “have” and “has”), “including” (and any form of including, such as “includes” and “include”), or “containing” (and any form of containing, such as “contains” and “contain”), are inclusive or open- ended and do not exclude additional, unrecited elements or method steps. Any composition or method that recites the term “comprising” should also be understood to also describe such compositions as consisting, consisting of, or consisting essentially of the recited components or elements.
[0050] As used herein, the term “fused” or “linked” when used in reference to a protein having different domains or heterologous sequences means that the protein domains are part of the same peptide chain that are connected to one another with either peptide bonds or other covalent bonding. The domains or section may be linked or fused directly to one another or another domain or peptide sequence may be between the two domains or sequences and such sequences would still be considered to be fused or linked to one another. In some embodiments, the various domains orproteins provided for herein are linked or fused directly to one another or a linker sequence, such as the glycine / serine sequences described herein to link the two domains together. Two peptide sequences are linked directly if they are directly connected to one another or indirectly if there is a linker or other structure that links the two regions. A linker may be directly linked to two different peptide sequences or domains.
[0051] As used herein, the term “individual,” “subject,” or “patient,” used interchangeably, means any animal, including mammals, such as mice, rats, other rodents, rabbits, dogs, cats, swine, cattle, sheep, horses, or primates, such as humans.
[0052] As used herein, the phrase “in need thereof’ means that the subject has been identified as having a need for the particular method or treatment. In some embodiments, the identification may be by any means of diagnosis. In any of the methods and treatments described herein, the subject may be in need thereof. In some embodiments, the subject is in an environment or will be traveling to an environment in which a particular disease, disorder, or condition is prevalent.
[0053] As used herein, the phrase “integer from X to Y” means any integer that includes the endpoints. For example, the phrase “integer from 1 to 5” means 1, 2, 3, 4, or 5.
[0054] As provided herein, the therapeutic compounds and compositions may be used in methods of treatment as provided herein. As used herein, the terms “treat,” “treated,” or “treating” mean both therapeutic treatment and prophylactic measures wherein the object is to slow down (lessen) an undesired physiological condition, disorder or disease, or obtain beneficial or desired clinical results. For purposes of these embodiments, beneficial or desired clinical results include, but are not limited to, alleviation of symptoms; diminishment of extent of condition, disorder or disease; stabilized (i.e., not worsening) state of condition, disorder or disease; delay in onset or slowing of condition, disorder or disease progression; amelioration of the condition, disorder or disease state or remission (whether partial or total), whether detectable or undetectable; an amelioration of at least one measurable physical parameter, not necessarily discernible by the patient; or enhancement or improvement of condition, disorder or disease. Treatment includes eliciting a clinically significant response without excessive levels of side effects. Treatment also includes prolonging survival as compared to expected survival if not receiving treatment. Thus, “treatment of an auto-immune disease / disorder” means an activity that alleviates or ameliorates any of the primary phenomena or secondary symptoms associated with the auto-immunedisease / disorder or other condition described herein. The various disease or conditions are provided herein. The therapeutic treatment can also be administered prophylactically to preventing or reduce the disease or condition before the onset.
[0055] As used herein, unless otherwise specified, the terms "5"' and "3"' denote the positions of elements or features relative to the overall arrangement of polynucleotide sequence, such as a polynucleotide-guide RNA cassettes, vectors, or retron donor DNA-guide molecules, or other polynucleotide sequences encoding for a molecule of interest, such as a fusion protein provided for herein. Positions are not, unless otherwise specified, referred to in the context of the orientation of a particular element or features. For example, FIG. 1 illustrates the msr and msd sequence oriented to the 5’ end of gRNA sequence or the 3’ end of the gRNA sequence. This is a non-limiting example and the gRNA may be part of a different transcript and expressed from a different vector. Unless otherwise specified, the term "upstream" refers to a position that is 5' of a point of reference. Conversely, the term "downstream" refers to a position that is 3' of a point of reference. Thus, in FIG. 1 the msrlmsd sequence locus is said to be located downstream (top) or upstream (bottom) of the gRNA sequence.
[0056] The term “genome editing” refers to a type of genetic engineering in which DNA is inserted, replaced, or removed from a target DNA (e.g., the genome of a cell) using one or more nucleases and / or nickases. The nucleases can create specific double-strand breaks (DSBs), single strand breaks at desired locations in the genome, and use the cell's endogenous mechanisms to repair the induced break by homology-directed repair (HDR) (e g., homologous recombination) or by nonhomologous end joining (NHEJ). The nickases create specific single-strand breaks at desired locations in the genome. In one non-limiting example, two nickases may be used to create two single-strand breaks on opposite strands of a target DNA, thereby generating a blunt or a sticky end. Any suitable DNA nuclease may be introduced into a cell to induce genome editing of a target DNA sequence. Genome editing may be performed with a catalytic inactive form of a Cas, which may be referred to as dCAS.
[0057] The term "DNA nuclease" refers to an enzyme capable of cleaving the phosphodiester bonds between the nucleotide subunits of DNA, and may be an endonuclease or an exonuclease. According to the present invention, the DNA nuclease may be an engineered (e.g., programmable or targetable) DNA nuclease which may be used to induce genome editing of a target DNA sequence. Any suitable DNA nuclease may be used including, but not limited to,CRISPR-associated protein (Cas) nucleases, other endo- or exo-nucl eases, variants thereof, fragments thereof, and combinations thereof. In some embodiments, a DNA nuclease may be mutated to create a single strand break as opposed to a double stranded break. In some embodiments, the DNA nuclease is mutated to be catalytically inactive.
[0058] The term "double-strand break" or "double-strand cut" refers to the severing or cleavage of both strands of the DNA double helix. The DSB may result in cleavage of both stands at the same position leading to "blunt ends" or staggered cleavage resulting in a region of singlestranded DNA at the end of each DNA fragment, or "sticky ends". A DSB may arise from the action of one or more DNA nucleases.
[0059] The term "nonhomologous end joining" or "NHEJ" refers to a pathway that repairs double-strand DNA breaks in which the break ends are directly ligated without the need for a homologous template.
[0060] The term "homology-directed repair" or "HDR" refers to a mechanism in cells to accurately and precisely repair double-strand DNA breaks using a homologous template to guide repair. The most common form of HDR is homologous recombination (HR), a type of genetic recombination in which nucleotide sequences are exchanged between two similar or identical molecules of DNA. The repair can also be referred to as “recombineering,” which can utilize SSAP (as defined herein) to replace / fix a sequence during cell division.
[0061] The term "nucleic acid," "nucleotide," or "polynucleotide" refers to deoxyribonucleic acids (DNA), ribonucleic acids (RNA) and polymers thereof in either single-, double- or multi-stranded form. The term includes, but is not limited to, single-, double- or multistranded DNA or RNA, genomic DNA, cDNA, DNA-RNA hybrids, or a polymer comprising purine and / or pyrimidine bases or other natural, chemically modified, biochemically modified, non-natural, synthetic or derivatized nucleotide bases. In some embodiments, a nucleic acid can comprise a mixture of DNA, RNA and analogs thereof. Unless specifically limited, the term encompasses nucleic acids containing known analogs of natural nucleotides that have similar binding properties as the reference nucleic acid and are metabolized in a manner similar to naturally occurring nucleotides. Unless otherwise indicated, a particular nucleic acid sequence also implicitly encompasses conservatively modified variants thereof (e.g., degenerate codon substitutions), alleles, orthologs, single nucleotide polymorphisms (SNPs), and complementary sequences as well as the sequence explicitly indicated. Specifically, degenerate codon substitutionsmay be achieved by generating sequences in which the third position of one or more selected (or all) codons is substituted with mixed-base and / or deoxyinosine residues (Batzer et al., Nucleic Acid Res. 19:5081 (1991); Ohtsuka et al., J. Biol. Chem. 260:2605-2608 (1985); and Rossolini et al., Mol. Cell. Probes 8:91-98 (1994)).
[0062] The term "single nucleotide polymorphism" or "SNP" refers to a change of a single nucleotide within a polynucleotide, including within an allele. This can include the replacement of one nucleotide by another, as well as the deletion or insertion of a single nucleotide. Most typically, SNPs are biallelic markers although tri- and tetra-allelic markers can also exist. By way of nonlimiting example, a nucleic acid molecule comprising SNP A\C may include a C or A at the polymorphic position.
[0063] The term "gene" refers to a portion of DNA involved that encodes a polypeptide chain or other molecule, such as miRNA. The DNA may include regions preceding and following the coding region (leader and trailer) involved in the transcription / translation of the gene product and the regulation of the transcription / translation, as well as intervening sequences (introns) between individual coding segments (exons).
[0064] The term "cassette" refers to a heterologous combination of nucleic acid molecule elements that may be introduced as a single element and may function together to achieve a desired result.
[0065] The term "operably linked" refers to two or more genetic elements, such as a polynucleotide coding sequence and a promoter, placed in relative positions that permit the proper biological functioning of the elements, such as the promoter directing transcription of the coding sequence.
[0066] As used herein, the term “rncRNA” (retron non-coding RNA) refers to a naturally occurring ncRNA in which various nucleotide sequences have been removed, added or replaced. In some embodiments, the terms “rncRNA” and “ncRNA” may be used interchangeably. In some embodiments described herein, a portion of the msd region of the ncRNA has been removed and replaced with an editing template that contains sequences homologous to human genomic DNA as well as a sequence for changing the human genomic DNA sequence. In some embodiments, the ncRNA provided for herein may be delivered as RNA.
[0067] The term "inducible promoter" refers to a promoter that responds to environmental factors and / or external stimuli that may be artificially controlled in order to modify the expressionof, or the level of expression of, a polynucleotide sequence or refers to a combination of elements, for example an exogenous promoter and an additional element such as a trans-activator operably linked to a separate promoter. An inducible promoter may respond to abiotic factors such as oxygen levels or to chemical or biological molecules. In some embodiments, the chemical or biological molecules may be molecules not naturally present in humans.
[0068] The terms "vector" and "expression vector" refer to a nucleic acid construct, generated recombinantly or synthetically, with a series of specified nucleic acid elements configured to transcribe a molecule of interest in a cell. An expression vector may be part of a plasmid, viral genome, or nucleic acid fragment. In some embodiments, the expression vector comprises a promoter operably linked to a heterologous polynucleotide sequence. In some embodiments, the vector is DNA, mRNA, or RNA. In some embodiments, the Cas-RT constructs or other constructs that encode for the RT or other proteins provided for herein may be delivered as a mRNA vector and the retron gRNA may be provided as RNA. In some embodiments, the Cas- RT constructs or other constructs that encode for the RT or other proteins provided for herein may be delivered as RNA.
[0069] The term "promoter" is used herein to refer to a nucleic acid sequence that directs, controls, or promotes the transcription of a nucleic acid. As used herein, a promoter includes necessary nucleic acid sequences near the start site of transcription, such as, in the case of a polymerase II type promoter, a TATA element. A promoter also optionally includes distal enhancer or repressor elements, which may be located as much as several thousand base pairs from the start site of transcription. Other elements that may be present in an expression vector include those that enhance transcription (e.g., enhancers) and terminate transcription (e.g., terminators).
[0070] "Recombinant" refers to a genetically modified polynucleotide, polypeptide, cell, tissue, or organism. A recombinant expression cassette, for example, can comprise a promoter operably linked to a second polynucleotide (e.g., a coding sequence) and can include a promoter that is heterologous to the second polynucleotide as the result of manipulation (e.g., by methods described in Sambrook et al., Molecular Cloning— A Laboratory Manual, Cold Spring Harbor Laboratory, Cold Spring Harbor, N.Y., (1989) or Current Protocols in Molecular Biology Volumes 1-3, John Wiley & Sons, Inc. (1994-1998)). A recombinant protein is one that is expressed from a recombinant polynucleotide, and recombinant cells, tissues, and organisms are those that comprise recombinant sequences (polynucleotide and / or polypeptide).
[0071] As used herein, the term "heterologous" refers to biological material that is introduced, inserted, or incorporated into a host (e.g., cell subject, tissue, etc.) that originates from another source. Heterologous material can include, but is not limited to, nucleic acids, amino acids, peptides, proteins, and structural elements such as genes, promoters, and cassettes. In some embodiments, a host cell is, but is not limited to, a bacterium, a yeast cell, a mammalian cell, or a plant cell. The introduction of heterologous material into a host cell or organism can result, in some instances, in the expression of additional heterologous material in or by the host cell or organism. As a non-limiting example, the transformation, transfection, or transduction of a mammalian host cell with an expression vector that contains DNA sequences encoding a bacterial protein (e.g. CAS9 or variants thereof) may result in the expression of the bacterial protein by the cell. The incorporation of heterologous material may be permanent or transient. Also, the expression of heterologous material may be permanent or transient.
[0072] The terms "culture," "culturing," "grow," "growing," "maintain," "maintaining," "expand," "expanding," etc., when referring to cell culture itself or the process of culturing, may be used interchangeably to mean that a cell is maintained outside its normal environment under controlled conditions, e.g., under conditions suitable for survival. Cultured cells are allowed to survive, and culturing can result in cell growth, stasis, differentiation or division. The term does not imply that all cells in the culture survive, grow, or divide, as some may naturally die or senesce. Cells are typically cultured in media, which may be changed during the course of the culture.
[0073] As used herein, the term "administering" includes oral administration, topical contact, administration as a suppository, intravenous, intraperitoneal, intramuscular, intralesional, intrathecal, intranasal, or subcutaneous administration to a subject. Administration is by any route, including parenteral and transmucosal (e.g., buccal, sublingual, palatal, gingival, nasal, vaginal, rectal, or transdermal). Parenteral administration includes, e.g., intravenous, intramuscular, intraarteriole, intradermal, subcutaneous, intraperitoneal, intraventricular, and intracranial. Other modes of delivery include, but are not limited to, the use of liposomal formulations, intravenous infusion, transdermal patches, etc.
[0074] The term "effective amount" or "sufficient amount" refers to the amount of an agent that is sufficient to effect beneficial or desired results. The therapeutically effective amount may vary depending upon one or more of: the subject and disease condition being treated, the weight and age of the subject, the severity of the disease condition, the manner of administration and thelike, which can readily be determined by one of ordinary skill in the art. The specific amount may vary depending on one or more of: the particular agent chosen, the host cell type, the location of the host cell in the subject, the dosing regimen to be followed, whether it is administered in combination with other compounds, timing of administration, and the physical delivery system in which it is carried.
[0075] The term "pharmaceutically acceptable carrier" refers to a substance that aids the administration of an active agent to a cell, an organism, or a subject. "Pharmaceutically acceptable carrier" refers to a carrier or excipient that may be included in the compositions of the invention and that causes no significant adverse toxicological effect on the patient. Non-limiting examples of pharmaceutically acceptable carrier include water, NaCl, normal saline solutions, lactated Ringer's, normal sucrose, normal glucose, cell culture media, and the like. One of skill in the art will recognize that other pharmaceutical carriers are useful in the present invention.
[0076] "Percent similarity," in the context of polynucleotide or peptide sequences, is determined by comparing two optimally aligned sequences over a comparison window, wherein the portion of the sequence (e.g., an msr locus sequence) in the comparison window may comprise additions or deletions (i.e., gaps) as compared to the reference sequence which does not comprise additions or deletions, for optimal alignment of the two sequences. The percentage is calculated by determining the number of positions at which the identical nucleotide or amino acid occurs in both sequences to yield the number of matched positions, dividing the number of matched positions by the total number of positions in the window of comparison and multiplying the result by 100 to yield the percentage of similarity (e g., sequence similarity).
[0077] When a polynucleotide or peptide has at least about 70% similarity (e.g., sequence similarity), preferably at least about 85%, at least about 86%, at least about 87%, at least about 88%, at least about 89%, at least about 90%, at least about 91%, at least about 92%, at least about 93, at least about 94%, at least about 95%, at least about 96%, at least about 97%, at least about 98%, at least about 99%, or at least about 100% similarity, to a reference sequence, when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection, such sequences are then said to be "substantially similar." In some embodiments, a polynucleotide or peptide has about 70% similarity (e.g., sequence similarity), preferably about 85%, about 86%, about 87%, about 88%, about 89%, about 90%, about 91%, about 92%, about93, about 94%, about 95%, about 96%, about 97%, about 98%, about 99%, or about 100% similarity, to a reference sequence, when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection, such sequences are then said to be "substantially similar." In some embodiments, a polynucleotide or peptide has at least 70% similarity (e.g., sequence similarity), preferably at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, at least 99%, or 100% similarity, to a reference sequence, when compared and aligned for maximum correspondence over a comparison window, or designated region as measured using one of the following sequence comparison algorithms or by manual alignment and visual inspection, such sequences are then said to be "substantially similar." With regard to polynucleotide sequences, this definition also refers to the complement of a test sequence.
[0078] For sequence comparison, typically one sequence acts as a reference sequence, to which test sequences are compared. When using a sequence comparison algorithm, test and reference sequences are entered into a computer, subsequence coordinates are designated, if necessary, and sequence algorithm program parameters are designated. Default program parameters may be used, or alternative parameters may be designated. The sequence comparison algorithm then calculates the percent sequence similarities for the test sequences relative to the reference sequence, based on the program parameters. For sequence comparison of nucleic acids and proteins, the BLAST and BLAST 2.0 algorithms and the default parameters discussed below are used.
[0079] Methods of alignment of sequences for comparison are well-known in the art. Optimal alignment of sequences for comparison may be conducted, e.g., by the local homology algorithm of Smith & Waterman, Adv. Appl. Math. 2:482 (1981), by the homology alignment algorithm of Needleman & Wunsch, J. Mol. Biol. 48:443 (1970), by the search for similarity method of Pearson & Lipman, Proc. Nafl. Acad. Sci. USA 85:2444 (1988), by computerized implementations of these algorithms (GAP, BESTFIT, FASTA, and TFASTA in the Wisconsin Genetics Software Package, Genetics Computer Group, 575 Science Dr., Madison, Wis.), or by manual alignment and visual inspection (see, e.g., Current Protocols in Molecular Biology (Ausubel et al., eds. 1995 supplement)).
[0080] Additional examples of algorithms that are suitable for determining percent sequence similarity are the BLAST and BLAST 2.0 algorithms, which are described in Altschul et al., (1990) J. Mol. Biol. 215: 403-410 and Altschul et al. (1977) Nucleic Acids Res. 25: 3389- 3402, respectively. Software for performing BLAST analyses is publicly available at the National Center for Biotechnology Information website, ncbi.nlm.nih.gov. The algorithm involves first identifying high scoring sequence pairs (HSPs) by identifying short words of length W in the query sequence, which either match or satisfy some positive-valued threshold score T when aligned with a word of the same length in a database sequence. T is referred to as the neighborhood word score threshold (Altschul et al., supra). These initial neighborhood word hits act as seeds for initiating searches to find longer HSPs containing them. The word hits are then extended in both directions along each sequence for as far as the cumulative alignment score may be increased. Cumulative scores are calculated using, for nucleotide sequences, the parameters M (reward score for a pair of matching residues; always >0) and N (penalty score for mismatching residues; always <0). The BLASTN program (for nucleotide sequences) uses as defaults a word size (W) of 28, an expectation (E) of 10, M=l, N=-2, and a comparison of both strands. For amino acid sequences, the BLASTP program uses as defaults a word size (W) of 3, an expectation (E) of 10, and the BLOSUM62 scoring matrix (see, e.g., Henikoff and Henikoff, Proc. Natl. Acad. Sci. USA 89: 10915 (1989)).
[0081] The BLAST algorithm also performs a statistical analysis of the similarity between two sequences (see, e.g., Karlin and Altschul, Proc. Nat'l. Acad. Sci. USA, 90:5873-5787 (1993)). One measure of similarity provided by the BLAST algorithm is the smallest sum probability (P(N)), which provides an indication of the probability by which a match between two nucleotide or amino acid sequences would occur by chance. For example, a nucleic acid is considered similar to a reference sequence if the smallest sum probability in a comparison of the test nucleic acid to the reference nucleic acid is less than about 0.2, more preferably less than about 0.01, and most preferably less than about 0.001.
[0082] The compositions and compounds of the embodiments provided for herein may be in a variety of forms. These include, for example, liquid, semi-solid and solid dosage forms, such as liquid solutions (e.g., injectable and infusible solutions), dispersions or suspensions, liposomes and suppositories. The preferred form depends on the intended mode of administration and therapeutic application. Typical compositions are in the form of injectable or infusible solutions.In some embodiments, the mode of administration is parenteral (e g., intravenous, subcutaneous, intraperitoneal, intramuscular). In some embodiments, the therapeutic molecule is administered by intravenous infusion or injection. In another embodiment, the therapeutic molecule is administered by intramuscular or subcutaneous injection. In another embodiment, the therapeutic molecule is administered locally, e.g., by injection, or topical application, to a target site.
[0083] The phrases "parenteral administration" and "administered parenterally" as used herein means modes of administration other than enteral and topical administration, usually by injection, and includes, without limitation, intravenous, intramuscular, intraarterial, intrathecal, intracapsular, intraorbital, intracardiac, intradermal, intraperitoneal, transtracheal, subcutaneous, subcuticular, intraarticular, subcapsular, subarachnoid, intraspinal, epidural and intrasternal injection and infusion.
[0084] Therapeutic compositions typically should be sterile and stable under the conditions of manufacture and storage. The composition may be formulated as a solution, microemulsion, dispersion, liposome, or other ordered structure suitable to high therapeutic molecule concentration. Sterile injectable solutions may be prepared by incorporating the active compound (i.e., therapeutic molecule) in the required amount in an appropriate solvent with one or a combination of ingredients enumerated above, as required, followed by filtered sterilization. Generally, dispersions are prepared by incorporating the active compound into a sterile vehicle that contains a basic dispersion medium and the required other ingredients from those enumerated above. In the case of sterile powders for the preparation of sterile injectable solutions, the preferred methods of preparation are vacuum drying and freeze-drying that yields a powder of the active ingredient plus any additional desired ingredient from a previously sterile-filtered solution thereof. The proper fluidity of a solution may be maintained, for example, by the use of a coating such as lecithin, by the maintenance of the required particle size in the case of dispersion and by the use of surfactants. Prolonged absorption of injectable compositions may be brought about by including in the composition an agent that delays absorption, for example, monostearate salts and gelatin.
[0085] As will be appreciated by the skilled artisan, the route and / or mode of administration will vary depending upon the desired results. In certain embodiments, the active compound may be prepared with a carrier that will protect the compound against rapid release, such as a controlled release formulation, including implants, transdermal patches, and microencapsulated delivery systems. Biodegradable, biocompatible polymers may be used, suchas ethylene vinyl acetate, polyanhydrides, polyglycolic acid, collagen, polyorthoesters, and polylactic acid. Many methods for the preparation of such formulations are patented or generally known to those skilled in the art. See, e.g., Sustained and Controlled Release Drug Delivery Systems, J. R. Robinson, ed., Marcel Dekker, Inc., New York, 1978.
[0086] In certain embodiments, a therapeutic compound may be orally administered, for example, with an inert diluent or an assimilable edible carrier. The compound (and other ingredients, if desired) may also be enclosed in a hard or soft shell gelatin capsule, compressed into tablets, or incorporated directly into the subject's diet. For oral therapeutic administration, the compounds may be incorporated with excipients and used in the form of ingestible tablets, buccal tablets, troches, capsules, elixirs, suspensions, syrups, wafers, and the like. To administer a compound by other than parenteral administration, it may be necessary to coat the compound with, or co-administer the compound with, a material to prevent its inactivation. Therapeutic compositions can also be administered with medical devices known in the art.
[0087] Dosage regimens are adjusted to provide the optimum desired response (e.g., a therapeutic response). For example, a single bolus may be administered, several divided doses may be administered over time or the dose may be proportionally reduced or increased as indicated by the exigencies of the therapeutic situation. It is especially advantageous to formulate parenteral compositions in dosage unit form for ease of administration and uniformity of dosage. Dosage unit form as used herein refers to physically discrete units suited as unitary dosages for the subjects to be treated; each unit contains a predetermined quantity of active compound calculated to produce the desired therapeutic effect in association with the required pharmaceutical carrier. The specification for the dosage unit forms are dictated by and directly dependent on (a) the unique characteristics of the active compound and the particular therapeutic effect to be achieved, and (b) the limitations inherent in the art of compounding such an active compound for the treatment of sensitivity in individuals.Retrons.
[0088] Retrons have been known for some time as a class of retroelement, first discovered in gram-negative bacteria such as Myxococcus xanthus (e.g., retrons Mx65 and Mxl62), Stigmatella aurantiaca (e.g., retron Sal63), and Escherichia coli (e.g., retrons Ec48, Ec67, Ec73, Ec78, Ec83, Ec86, and Ecl07). Retrons are also found in Salmonella typhimurium (e.g., retron St85), Salmonella enteritidis Vibrio cholera (e.g., retron Vc95), Vibrio parahaemolyticus (e.g.,retron Vp96), Klebsiella pneumoniae , Proteus mirabilis, Xanthomonas campestris, Rhizobium sp., Bradyrhizobium sp., Ralstonia metallidurans, Nannocystis exedens (e.g., retron Nel44), Geobacter sulfurreducens, Trichodesmium erythraeum, Nostoc punctiforme, Nostoc sp., Staphylococcus aureus, Fusobacterium nucleatum, and Flexibacter elegans. In some embodiments a polynucleotide-guide RNA cassette is provided that comprise a retron.
[0089] As used herein, the term “retron” may refer to a specific type of naturally-occurring and distinct DNA sequence found in the genome of many bacteria, or an engineered DNA sequence, which may encode three distinct components, namely, (a) a non-coding RNA (“ncRNA”) (comprising msr and msd; and contiguous inverted sequences al and a2), (b) a reverse transcriptase (RT)-coding gene (ret), and (c) in many cases, a retron-associated gene of unknown function. Retrons are particularly defined by their unique ability to produce a satellite DNA known as msDNA (multicopy single-stranded DNA). The ncRNA (comprising the msr and msd elements) and the ret gene are transcribed as a single polycistronic RNA transcript which is processed into the ncRNA transcript and a transcript encoding the ret gene. The ncRNA then becomes folded into a specific secondary structure. Once translated, the RT then binds the folded ncRNA and reverse transcribes the msd region to form a single strand of cDNA (the msDNA) that remains covalently attached to the RNA template via a 2'-5' phosphodiester bond and base-pairing between the 3' ends of the msDNA and the RNA template.
[0090] In some embodiments, the term “self-priming” refers to a mechanism found in bacterial retron reverse transcriptase, which may initiate the synthesis of DNA from a template found within the retron non-coding RNA. Without wishing to be bound by a particular theory, this process in retrons uses a self-priming mechanism that relies on a free 3'OH of RNA, which is reverse-transcribed into retron multicopy single-stranded DNA (msDNA).
[0091] In some embodiments, the retron comprises a non-coding RNA (ncRNA). In some embodiments, the ncRNA comprises an msr / msd locus. In some embodiments, an msr / msd locus comprises an msr locus and a msd locus. In some embodiments, a ncRNA comprises a first inverted repeat sequence coding region. In some embodiments, a ncRNA comprises a second inverted repeat sequence coding region. In some embodiments, the term “coding region” as used in relation to a ncRNA or any other RNA that does not encode a protein, is used to mean a nucleotide sequence encoding the ncRNA. or any other RNA that does not encode a protein. In some embodiments, a ncRNA comprises a donor DNA sequence. In some embodiments, a donor DNA sequence islocated within an msr / msd locus. In some embodiments, a donor DNA sequence is located within an msd locus. In some embodiments, a retron comprises an msr / msd locus, a first inverted repeat sequence coding region, a donor DNA sequence, and a second inverted repeat sequence coding region. In some embodiments, the ncRNA comprises the msr / msd locus, the first inverted repeat sequence coding region, the donor DNA sequence, and the second inverted repeat sequence coding region. In some embodiments, the retron comprises the msr locus, the first inverted repeat sequence coding region, the msd locus, the donor DNA sequence, and the second inverted repeat sequence coding region. In some embodiments, the ncRNA comprises the msr locus, the first inverted repeat sequence coding region, the msd locus, the donor DNA sequence, and the second inverted repeat sequence coding region. In some embodiments, the retron comprises an ncRNA and an RT. In some embodiments, the ncRNA comprises an al and / or a2 region. In some embodiments, a retron comprises an al region, an ncRNA comprising an msr / msd locus, a first inverted repeat sequence coding region, a donor DNA sequence, a second inverted repeat sequence coding region, and an a2 region, all of which can be arranged in any orientation and order. In some embodiments, a retron comprises an al region, an ncRNA comprising an msr / msd locus, a first inverted repeat sequence coding region, a donor DNA sequence, a second inverted repeat sequence coding region, and an a2 region. In some embodiments, the ncRNA comprises the al region, the msr / msd locus, the first inverted repeat sequence coding region, the donor DNA sequence, the second inverted repeat sequence coding region, and the a2 region. In some embodiments, the retron comprises the ncRNA comprising the al region the msr locus, the first inverted repeat sequence coding region, the msd locus, the donor DNA sequence, the second inverted repeat sequence coding region, and the a2 region. In some embodiments, the ncRNA comprises the al region, the msr locus, the first inverted repeat sequence coding region, the msd locus, the donor DNA sequence, the second inverted repeat sequence coding region, and the a2 region. In some embodiments, a retron comprises an a2 region, an ncRNA comprising an msr / msd locus, a first inverted repeat sequence coding region, a donor DNA sequence, a second inverted repeat sequence coding region, and an al region. In some embodiments, the ncRNA comprises the a2 region, the msr / msd locus, the first inverted repeat sequence coding region, the donor DNA sequence, the second inverted repeat sequence coding region, and the al region. In some embodiments, the retron comprises the ncRNA comprising the a2 region the msr locus, the first inverted repeat sequence coding region, the msd locus, the donor DNA sequence, the second inverted repeat sequence coding region, and the al region. In someembodiments, the ncRNA comprises the a2 region, the msr locus, the first inverted repeat sequence coding region, the msd locus, the donor DNA sequence, the second inverted repeat sequence coding region, and the al region. In some embodiments, the retron comprises the ncRNA comprising the al region the msr locus, the donor DNA sequence, the msd locus, and the a2 region. In some embodiments, the retron comprises the ncRNA comprising the a2 region the msr locus, the donor DNA sequence, the msd locus, and the al region. In some embodiments, the retron comprises the ncRNA comprising the a2 region the msr locus, the donor DNA sequence, and the al region. In some embodiments, the retron comprises the ncRNA comprising the al region the msr locus, the donor DNA sequence, and the a2 region. In some embodiments, the retron comprises the ncRNA comprising the a2 region the msr locus, and the donor DNA sequence. In some embodiments, the retron comprises the ncRNA comprising the al region the msr locus, and the donor DNA sequence. In some embodiments, the retron comprises the ncRNA comprising the a2 region the msr locus, and the msd locus. In some embodiments, the retron comprises the ncRNA comprising the al region the msr locus, and the msd locus. In some embodiments, the retron comprises the ncRNA comprising the al region, the donor DNA sequence, and the msr locus. In some embodiments, the retron comprises the ncRNA comprising the a2 region, the donor DNA sequence, and the msr locus. In some embodiments, the retron comprises the ncRNA comprising the al region, the donor DNA sequence, and the msd locus. In some embodiments, the retron comprises the ncRNA comprising the a2 region, the donor DNA sequence, and the msd locus. In some embodiments, the retron comprises the ncRNA comprising the msr locus, the donor DNA sequence, the msd locus, and the al region. In some embodiments, the retron comprises the ncRNA comprising the msr locus, the donor DNA sequence, the msd locus, and the a2 region. In some embodiments, the ncRNA comprises the al region, the msr locus, the donor DNA sequence, the msd locus, and the a2 region. In some embodiments, the ncRNA comprises the a2 region, the msr locus, the donor DNA sequence, the msd locus, and the al region. In some embodiments, the ncRNA comprises the al region the msr locus, the donor DNA sequence, the msd locus, and the a2 region. In some embodiments, the ncRNA comprises the a2 region the msr locus, the donor DNA sequence, the msd locus, and the al region. In some embodiments, the ncRNA comprises the a2 region the msr locus, the donor DNA sequence, and the al region. In some embodiments, the ncRNA comprises the al region the msr locus, the donor DNA sequence, and the a2 region. In some embodiments, the ncRNA comprises the a2 region the msr locus, and the donor DNAsequence. In some embodiments, the ncRNA comprises the al region the msr locus, and the donor DNA sequence. In some embodiments, the ncRNA comprises the a2 region the msr locus, and the msd locus. In some embodiments, the ncRNA comprises the al region the msr locus, and the msd locus. In some embodiments, the ncRNA comprises the al region, the donor DNA sequence, and the msr locus. In some embodiments, the ncRNA comprises the a2 region, the donor DNA sequence, and the msr locus. In some embodiments, the ncRNA comprises the al region, the donor DNA sequence, and the msd locus. In some embodiments, the ncRNA comprises the a2 region, the donor DNA sequence, and the msd locus. In some embodiments, the ncRNA comprises the msr locus, the donor DNA sequence, the msd locus, and the al region. In some embodiments, the ncRNA comprises the msr locus, the donor DNA sequence, the msd locus, and the a2 region.
[0092] Without being bound by any particular theory, retrons mediate the synthesis in host cells of multicopy single-stranded DNA (msDNA) molecules, which result from the reverse transcription of a polynucleotide transcript and typically include a DNA component and an RNA component. The native msDNA molecules exist as single-stranded DNA-RNA hybrids, characterized by a structure which comprises a single-stranded DNA branching out of an internal guanosine residue of a single-stranded RNA molecule at a 2',5'-phosphodiester linkage. In some embodiments of the present invention, at least some of the RNA content of the msDNA molecule is degraded. In some instances, the RNA content is degraded by RNase H.
[0093] Retrons have been found to consist of the gene for reverse transcriptase (RT) and msr and msd loci under the control of a single promoter. In some embodiments, a vector comprising a polynucleotide-guide RNA cassette is provided. In some embodiments, the cassette does not comprise a sequence encoding for a reverse transcriptase. Thus, in some embodiments, methods are provided wherein the reverse transcriptase is encoded on a separate plasmid from the polynucleotide-guide RNA cassette. In some embodiments, the RT is encoded in a sequence that has been integrated into the host cell genome.
[0094] In some embodiments, the retrons and polynucleotides provided herein are based on and / or derived from a naturally-occurring retron, such as any retron-related sequence provided by Table X and Y. In some embodiments, the retrons provided herein are based on introducing one or more genetic modifications into previously available retron sequences (e.g., the “Mestre et al., Systematic Prediction of Genes Functionally Associated with Bacterial Retrons and Classification of The Encoded Tripartite Systems, Nucleic Acids Research, Volume 48, Issue 22,16 Dec. 2020, Pages 12632-12647” (incorporated herein by reference) to achieve recombinant retrons with the enhanced ability to produce increased concentrations or amounts of msDNA comprising a DNA donor template. Many retrons also contain an accessory protein, which may have a variable function that may not be fully understood. In some embodiments, the retrons described herein do not comprise the accessory protein naturally associated with the wild-type or template retron.
[0095] In some embodiments, the msd region of a polynucleotide transcript typically codes for the DNA component of msDNA, and the msr region of a polynucleotide transcript typically codes for the RNA component of msRNA. In some embodiments of the retrons, the msr and msd loci have overlapping ends, and may be oriented opposite one another with a promoter located upstream of the msr locus which transcribes through the msr and msd loci.
[0096] In some embodiments, Table X and Table Y provides a list of retrons. In some embodiments, Table X and Table Y provides a list of discovered and analyzed retrons. In reference to Table X, column “genome” describes an NCBI accession number of the bacterial genome of origin predicted to comprise the retron. In reference to Table Y, column “rt genome” describes an NCBI accession number of the bacterial genome of origin predicted to comprise the retron.
[0097] In reference to Table X and Table Y, column “ncrna start” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the ncRNA of the retron. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the ncRNA of the retron may be within a range of nucleic acids outside or within the “ncrna start” of Table X and Table Y. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the ncRNA of the retron may be within 10-20 nucleic acids outside the “ncrna start” of Table X and Table Y.
[0098] In reference to Table X and Table Y, column “ncma end” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the ncRNA of the retron. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the ncRNA of the retron may be within a range of nucleic acids outside or within the “ncrna end” of Table X and Table Y. In some embodiments, the nucleic acid of thegenomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the ncRNA of the retron may be within 10-20 nucleic acids outside the “ncrna end” of Table X and Table Y.
[0099] In reference to Table X and Table Y, column “msr start” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the msr of the retron. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the msr of the retron may be within a range of nucleic acids outside or within the “msr start” of Table X and Table Y. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the msr of the retron may be within 10-20 nucleic acids outside the “msr start” of Table X and Table Y.
[0100] In reference to Table X and Table Y, column “msr end” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the msr of the retron. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the msr of the retron may be within a range of nucleic acids outside or within the “msr end” of Table X and Table Y. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the msr of the retron may be within 10-20 nucleic acids outside the “msr_end” of Table X and Table Y.
[0101] In reference to Table X and Table Y, column “msd start” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the msd of the retron. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the msd of the retron may be within a range of nucleic acids outside or within the “msd start” of Table X and Table Y. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the msd of the retron may be within 10-20 nucleic acids outside the “msd_start” of Table X and Table Y.
[0102] In reference to Table X and Table Y, column “msd end” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the msd of the retron. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the msd of the retron may be within a range of nucleic acids outside or within the “msd end” of Table X and Table Y. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the msd of the retron may be within 10-20 nucleic acids outside the “msd_end” of Table X and Table Y.
[0103] In reference to Table X and Table Y, column “rt start” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the RT. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the RT may be within a range of nucleic acids outside or within the “rt start” of Table X and Table Y. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the RT may be within 10-20 nucleic acids outside the “rt start” of Table X and Table Y.
[0104] In reference to Table X and Table Y, column “rt end” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the RT. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the RT may be within a range of nucleic acids outside or within the “rt end” of Table X and Table Y. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the RT may be within 10-20 nucleic acids outside the “rt_end” of Table X and Table Y.
[0105] In reference to Table X and Table Y, column “insert start” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the ncRNA of the retron where the donor DNA may be inserted. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the msd of the retron where the donor DNA may be inserted may be within a range of nucleic acids outside orwithin the “insert start” of Table X and Table Y. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the ncRNA of the retron where the donor DNA may be inserted may be within 10-20 nucleic acids outside the “insert start” of Table X and Table Y.
[0106] In reference to Table X and Table Y, column “insert end” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the ncRNA of the retron where the donor DNA may be inserted. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the msd of the retron where the donor DNA may be inserted may be within a range of nucleic acids outside or within the “insert end” of Table X and Table Y. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the ncRNA of the retron where the donor DNA may be inserted may be within 10-20 nucleic acids outside the “insert end” of Table X and Table Y.
[0107] In reference to Table X and Table Y, column “optimized insert start” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the msd of the retron where the donor DNA may be inserted. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the msd of the retron where the donor DNA may be inserted may be within a range of nucleic acids outside or within the “optimized insert start” of Table X and Table Y. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the start of the nucleic acid sequence encoding the msd of the retron where the donor DNA may be inserted may be within 10-20 nucleic acids outside the “optimized insert start” of Table X and Table Y.
[0108] In reference to Table X and Table Y, column “optimized insert end” describes the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the msd of the retron where the donor DNA may be inserted. In some embodiments, the nucleic acid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the msd of the retron where the donor DNA may be inserted may be within a range of nucleic acids outside or within the “optimized insert end” of Table X and Table Y. In some embodiments, the nucleicacid of the genomic DNA of the corresponding bacterial genome of interest that is the end of the nucleic acid sequence encoding the msd of the retron where the donor DNA may be inserted may be within 10-20 nucleic acids outside the “optimized insert end” of Table X and Table Y.
[0109] In reference to Table X and Table Y, column “rt seqid” provides the SEQ ID NO of the RT associated with the corresponding retron. In reference to Table X and Table Y, column “msr seqid” provides the SEQ ID NO of the msr associated with the corresponding retron. In reference to Table X and Table Y, column “msDNA seqid” provides the SEQ ID NO of the msDNA associated with the corresponding retron. In some embodiments, “rt_seqid,” “msr_seqid,” “msDNA_seqid,” and “ncrna_seqid” of Table X and Table Y provide SEQ ID NO associated with nucleic acid sequences provided in a Sequence Listing which has been submitted electronically in XML file format and is hereby incorporated by reference in its entirety.
[0110] In reference to Table X and Table Y, column “ncrna seqid” provides the SEQ ID NO of the ncRNA associated with the corresponding retron. In reference to Table X and Table Y, column “ncrna_strand”, “msr_strand”, “msd_strand” and “rt_strand” each independently may comprise associated with the sense DNA strand, or associated with the antisense DNA strand, and denotes the DNA strand which serves as the template for expression. Thus, in some embodiments, the ncRNA is expressed from the sense DNA strand, the msr is expressed from the sense DNA strand, the msd is expressed from the sense DNA strand, and the RT is expressed from the sense DNA strand. In some embodiments, the ncRNA is expressed from the antisense DNA strand, the msr is expressed from the sense DNA strand, the msd is expressed from the sense DNA strand, and the RT is expressed from the sense DNA strand. In some embodiments, the ncRNA is expressed from the sense DNA strand, the msr is expressed from the antisense DNA strand, the msd is expressed from the sense DNA strand, and the RT is expressed from the sense DNA strand. In some embodiments, the ncRNA is expressed from the sense DNA strand, the msr is expressed from the sense DNA strand, the msd is expressed from the antisense DNA strand, and the RT is expressed from the sense DNA strand. In some embodiments, the ncRNA is expressed from the sense DNA strand, the msr is expressed from the sense DNA strand, the msd is expressed from the sense DNA strand, and the RT is expressed from the antisense DNA strand. In some embodiments, the ncRNA is expressed from the antisense DNA strand, the msr is expressed from the antisense DNA strand, the msd is expressed from the sense DNA strand, and the RT is expressed from the sense DNA strand. In some embodiments, the ncRNA is expressed from theantisense DNA strand, the msr is expressed from the sense DNA strand, the msd is expressed from the antisense DNA strand, and the RT is expressed from the sense DNA strand. In some embodiments, the ncRNA is expressed from the antisense DNA strand, the msr is expressed from the sense DNA strand, the msd is expressed from the sense DNA strand, and the RT is expressed from the antisense DNA strand. In some embodiments, the ncRNA is expressed from the sense DNA strand, the msr is expressed from the antisense DNA strand, the msd is expressed from the antisense DNA strand, and the RT is expressed from the sense DNA strand. In some embodiments, the ncRNA is expressed from the sense DNA strand, the msr is expressed from the antisense DNA strand, the msd is expressed from the sense DNA strand, and the RT is expressed from the antisense DNA strand. In some embodiments, the ncRNA is expressed from the sense DNA strand, the msr is expressed from the sense DNA strand, the msd is expressed from the antisense DNA strand, and the RT is expressed from the antisense DNA strand. In some embodiments, the ncRNA is expressed from the antisense DNA strand, the msr is expressed from the antisense DNA strand, the msd is expressed from the antisense DNA strand, and the RT is expressed from the sense DNA strand. In some embodiments, the ncRNA is expressed from the antisense DNA strand, the msr is expressed from the antisense DNA strand, the msd is expressed from the sense DNA strand, and the RT is expressed from the antisense DNA strand. In some embodiments, the ncRNA is expressed from the antisense DNA strand, the msr is expressed from the sense DNA strand, the msd is expressed from the antisense DNA strand, and the RT is expressed from the antisense DNA strand. In some embodiments, the ncRNA is expressed from the sense DNA strand, the msr is expressed from the antisense DNA strand, the msd is expressed from the antisense DNA strand, and the RT is expressed from the antisense DNA strand. In some embodiments, the ncRNA is expressed from the antisense DNA strand, the msr is expressed from the antisense DNA strand, the msd is expressed from the antisense DNA strand, and the RT is expressed from the antisense DNA strand.
[0111] In some embodiments, the ncRNA comprises an msr / msd locus. In some embodiments, the msr / msd locus comprises an msr locus and an msd locus. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one msr / msd locus of Table X or Table Y. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one msr / msd locus of Table X (column “ncma seqid”). In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one msr / msd locus of Table X (column “optimized ncrna seqid”). In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one msr / msd locus of Table Y (column “ncma seqid”). In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one msr / msd locus of Table Y (column “optimized ncrna seqid”). In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one ncRNA SEQ ID NO: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579, 19582,19583, 19585, 19587, 19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605, 19609,19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671, 19672, 19674,19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710,19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944,19945, 19946, 19950, 19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983, 19986, 19998,19999, 20000, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034,20036, 20038, 20040, 20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20061, 20062,20070, 20072, 20074, 20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128, 20132,20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171,20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204, 20207, 20209,20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243,20247, 20248, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20276,20277, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300,20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332,20335, 20338, 20339, 20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367, 20369, 20371,20373, or 20377. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one SEQ ID NO: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582,19582, 19585, 19587, 19589, 19591, 19593, 19598, 19601, 19602, 19604, 19605, 19609, 19610,19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19671, 19674, 19675, 19678, 19680,19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717,19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950,19951, 19957, 19959, 19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029,20030, 20032, 20034, 20036, 20038, 20040, 20046, 20049, 20050, 20052, 20053, 20055, 20056,20070, 20074, 20075, 20078, 20092, 20099, 20103, 20108, 20117, 20128, 20132, 20136, 20138,20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174,20176, 20178, 20179, 20184, 20185, 20196, 20204, 20207, 20209, 20211, 20213, 20214, 20218,20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20249, 20250, 20251,20252, 20259, 20260, 20262, 20271, 20272, 20274, 20279, 20280, 20281, 20286, 20287, 20289,20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318,20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358, 20362, 20365,20367, 20369, 20371, 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386,20386, 20387, 20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398,20399, 20400, 20401, 20402, 20403, 20404, or 20405.
[0112] In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of:ACATAATTGCGCAGCTCAGTTACGCGACTTGTGGTTGTTGGTAAAACGACATGCCATCG TCGGGCGCATCGCCAGACGGAAAATGATTGAACGACATTGGCGGCACGAAAACGGCCAC GCGCCCAATCTGCTCCTTTGTCGCATCTTGGGCGCGTGTCCATTTTCGTGCCTGTAATT CAAGGTAGCTGAGCCTACCTATGACGT (SEQ ID NO: 19610);CCAGCAGTGGCAATAGCGTTTCCGGCCTTTTGTGCCGGGAGGGTCGGCGAGTCGCCGAC TTAACGCCAGTAGTTTGTCCAGATACTCAAAGTCGCTCCATTGTACTTAAGTACGCTTC GCGTACGTCGCGCTGACGCGCTCAGTACAGTTACGCGCCTTCGGGATAGTTTGAGGGTA TTGCCGCTGTTGG (SEQ ID NO: 19979);ATTCATGTAATCTCTATATGTCCTTTAGCGTTTAGACGTTTACGTCTTGTCGGGCGTTT CGCCAGACACGAAGTTATTGGAAGGTTTATGAGGTTGCGGTAGGTATAATCCTCCGCCG CTCATTGCCTTGCATTTCGCGGCGGAGGATTATACCCACCGCATCCTTAAGTAAAGGACATAGAAGAACTGGCATTAAT (SEQ ID NO: 19931);ATCGACTAACTCAGTTACGCGCATAGTGGTTGTTACTTAGGAAACATACCATCGTCGGG CGTATCGCCAGACGGGAAATAATTGATATACATCGGCGGCACGAAACAGGCCACGCGCC CAACTCGCTCCGTCGTCACTCCTTGGGCGCGTGGCCTGTTTCGTGCCTGTAATTTATGG TAGCTGAGTCTGCCTAT (SEQ ID NO: 19722); or GCTCAGTTACGCGACTTGTGGTTGTTGGTAAAACGACATGCCATCGTCGGGCGCATCGC CAGACGGAAAATGATTGAACGACATTGGCGGCACGAAAACGGCCACGCGCCCAATCTGC TCCTTTGTCGCATCTTGGGCGCGTGTCCATTTTCGTGCCTGTAATTCAAGGTAGCTGAG C (SEQ ID NO: 24020).
[0113] In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19610. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19979. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19931. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19722. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20380.
[0114] In some embodiments, the msr / msd locus comprises the nucleic acid sequence of any one of SEQ ID NO as provided in Table X or Table Y. In some embodiments, themsr / msd locus comprises the nucleic acid sequence of any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587,19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651,19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671, 19672, 19674, 19675, 19678, 19680,19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717,19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983, 19986, 19998, 19999, 20000, 20005,20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040,20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20061, 20062, 20070, 20072, 20074,20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128, 20132, 20136, 20138, 20140,20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176,20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204, 20207, 20209, 20211, 20213, 20214,20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20248, 20249,20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20276, 20277, 20279, 20280,20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306,20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339,20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367, 20369, 20371, 20373, or 20377.
[0115] In some embodiments, the msr / msd locus comprises the nucleic acid sequence of any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19671, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20070, 20074, 20075, 20078, 20092, 20099, 20103, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311,20313, 20318, 20319, 20320, 20323, 20328, 20331 , 20332, 20335, 20338, 20339, 20342, 20358,20362, 20365, 20367, 20369, 20371, 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384,20385, 20386, 20386, 20387, 20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396,20397, 20398, 20399, 20400, 20401, 20402, 20403, 20404, or 20405.
[0116] In some embodiments, the msr / msd locus comprises the nucleic acid sequence of SEQ ID NO: 19610, SEQ ID NO: 19979, SEQ ID NO: 19931, SEQ ID NO: 19722, or SEQ ID NO: 24020. In some embodiments, the msr / msd locus comprises the nucleic acid sequence of SEQ ID NO: 19610. In some embodiments, the msr / msd locus comprises the nucleic acid sequence of SEQ ID NO: 19979. In some embodiments, the msr / msd locus comprises the nucleic acid sequence of SEQ ID NO: 19931 . In some embodiments, the msr / msd locus comprises the nucleic acid sequence of SEQ ID NO: 19722. In some embodiments, the msr / msd locus comprises the nucleic acid sequence of SEQ ID NO: 24020.
[0117] In some embodiments, the msr locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%>, at least 85%, at least 90%, at least 91%, at least 92%>, at least 93%>, at least 94%, at least 95%, at least 96%>, at least 97%, at least 98%, or at least 99% sequence identity to any msr locus of Table X (column “msr seqid”). In some embodiments, the msr locus comprises the nucleic acid sequence having at least 50%, at least 55%>, at least 60%, at least 65%>, at least 70%, at least 75%, at least 80%o, at least 85%>, at least 90%>, at least 91%>, at least 92%o, at least 93%>, at least 94%>, at least 95%>, at least 96%>, at least 97%, at least 98%, or at least 99% sequence identity to any msr locus of Table Y (column “msr seqid”). In some embodiments, the msr locus comprises the nucleic acid sequence having at least 50%>, at least 55%o, at least 60%>, at least 65%>, at least 70%o, at least 75%>, at least 80%o, at least 85%, at least 90%, at least 91%, at least 92%>, at least 93%>, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any of SEQ ID NO: 19552, 19555, 19556, 19557, 19559, 19560, 19562, 19565, 19568, 19569, 19571, 19573, 19575,19577, 19584, 19586, 19588, 19592, 19594, 19597, 19599, 19600, 19603, 19615, 19618, 19619,19620, 19635, 19650, 19655, 19657, 19659, 19661, 19663, 19668, 19670, 19673, 19679, 19681,19683, 19684, 19686, 19687, 19688, 19693, 19696, 19700, 19701, 19705, 19707, 19712, 19713,19716, 19718, 19720, 19723, 19724, 19784, 19850, 19851, 19852, 19853, 19854, 19866, 19910,19928, 19929, 19932, 19934, 19935, 19936, 19938, 19942, 19952, 19956, 19958, 19964, 19982,19987, 19997, 20003, 20004, 20009, 20012, 20013, 20014, 20015, 20017, 20019, 20021, 20022,20024, 20026, 20027, 20031, 20035, 20037, 20048, 20051, 20054, 20064, 20066, 20069, 20073,20077, 20079, 20080, 20091, 20098, 20104, 20111, 20113, 20114, 20115, 20116, 20119, 20120,20123, 20127, 20131, 20135, 20137, 20139, 20148, 20149, 20150, 20157, 20160, 20162, 20164,20166, 20167, 20168, 20170, 20173, 20175, 20177, 20183, 20203, 20206, 20208, 20210, 20212,20215, 20224, 20225, 20226, 20228, 20230, 20234, 20236, 20239, 20240, 20241, 20244, 20246,20254, 20256, 20257, 20258, 20261, 20263, 20264, 20266, 20269, 20270, 20273, 20278, 20282,20283, 20285, 20288, 20292, 20293, 20294, 20297, 20299, 20301, 20303, 20305, 20307, 20309,20310, 20312, 20314, 20316, 20317, 20322, 20325, 20326, 20327, 20330, 20334, 20337, 20341,20353, 20355, 20357, 20361, 20364, 20366, 20368, 20370, 20372, 20375, or 20376. In some embodiments, the msr locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of:GCAGCTCAGTTACGCGACTTGTGGTTGTTGGTAAAACGACATGCCATCGTCGGG CGCATCGCCAGACGGAAAATGATTGAACGACAT (SEQ ID NO: 20115);GCTCAGTTACGCGACTTGTGGTTGTTGGTAAAACGACATGCCATCGTCGGGCGC ATCGCCAGACGGAAAATGATTGAACGACAT (SEQ ID NO: 24021);CGCCAGCAGTGGCAATAGCGTTTCCGGCCTTTTGTGCCGGGAGGGTCGGCGAGT CGCCGACTTAACGCCAGTAGTTTGTCCA (SEQ ID NO: 20026);AATCTCTATATGTCCTTTAGCGTTTAGACGTTTACGTCTTGTCGGGCGTTTCGC CAGACACGAAGTTATTGGAAGGT (SEQ ID NO: 19599); or GCATAGTGGTTGTTACTTAGGAAACATACCATCGTCGGGCGTATCGCCAGACGG GAAATAATTGATATACAT (SEQ ID NO: 20120).
[0118] In some embodiments, the msr locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20115. In some embodiments, the msr locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%,at least 98%, or at least 99% sequence identity to SEQ ID NO: 24021 . In some embodiments, the msr locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20026. In some embodiments, the msr locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19599. In some embodiments, the msr locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20120.
[0119] In some embodiments, the msr locus comprises the nucleic acid sequence of any one of SEQ ID NOs: 19552, 19555, 19556, 19557, 19559, 19560, 19562, 19565, 19568, 19569, 19571, 19573, 19575, 19577, 19584, 19586, 19588, 19592, 19594, 19597, 19599, 19600, 19603, 19615, 19618, 19619, 19620, 19635, 19650, 19655, 19657, 19659, 19661, 19663, 19668, 19670, 19673, 19679, 19681, 19683, 19684, 19686, 19687, 19688, 19693, 19696, 19700, 19701, 19705, 19707, 19712, 19713, 19716, 19718, 19720, 19723, 19724, 19784, 19850, 19851, 19852, 19853, 19854, 19866, 19910, 19928, 19929, 19932, 19934, 19935, 19936, 19938, 19942, 19952, 19956, 19958, 19964, 19982, 19987, 19997, 20003, 20004, 20009, 20012, 20013, 20014, 20015, 20017, 20019, 20021, 20022, 20024, 20026, 20027, 20031, 20035, 20037, 20048, 20051, 20054, 20064, 20066, 20069, 20073, 20077, 20079, 20080, 20091, 20098, 20104, 20111, 20113, 20114, 20115, 20116, 20119, 20120, 20123, 20127, 20131, 20135, 20137, 20139, 20148, 20149, 20150, 20157, 20160, 20162, 20164, 20166, 20167, 20168, 20170, 20173, 20175, 20177, 20183, 20203, 20206, 20208, 20210, 20212, 20215, 20224, 20225, 20226, 20228, 20230, 20234, 20236, 20239, 20240, 20241, 20244, 20246, 20254, 20256, 20257, 20258, 20261, 20263, 20264, 20266, 20269, 20270, 20273, 20278, 20282, 20283, 20285, 20288, 20292, 20293, 20294, 20297, 20299, 20301, 20303, 20305, 20307, 20309, 20310, 20312, 20314, 20316, 20317, 20322, 20325, 20326, 20327, 20330, 20334, 20337, 20341, 20353, 20355, 20357, 20361, 20364, 20366, 20368, 20370, 20372, 20375, or 20376.
[0120] In some embodiments, the msr locus comprises the nucleic acid sequence of any one of SEQ ID NOs: 20115, SEQ ID NO: 24021, SEQ ID NO: 20026, SEQ ID NO: 19599, or SEQ ID NO: 20120. In some embodiments, the msr locus comprises the nucleic acid sequence of SEQ ID NO: 20115. In some embodiments, the msr locus comprises the nucleic acid sequence of SEQ ID NO: 24021. In some embodiments, the msr locus comprises the nucleic acid sequence of SEQ ID NO: 20026. In some embodiments, the msr locus comprises the nucleic acid sequence of SEQ ID NO: 19599. In some embodiments, the msr locus comprises the nucleic acid sequence of SEQ ID NO: 20120.
[0121] In some embodiments, the msDNA comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity any msDNA of Table X (column “msDNA seqid”). In some embodiments, the msDNA comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any msDNA of Table Y (column “msDNA_seqid”). In some embodiments, the msDNA comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590,19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621, 19622, 19623,19624, 19625, 19626, 19627, 19628, 19629, 19630, 19631, 19632, 19633, 19636, 19637, 19638,19639, 19640, 19641, 19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649, 19653, 19654,19664, 19665, 19667, 19676, 19677, 19690, 19691, 19692, 19695, 19697, 19699, 19703, 19709,19715, 19719, 19721, 19755, 19855, 19927, 19940, 19941, 19947, 19948, 19949, 19953, 19954,19955, 19960, 19961, 19962, 19963, 19965, 19966, 19967, 19968, 19969, 19970, 19971, 19972,19973, 19975, 19976, 19977, 19981, 19984, 19985, 19988, 19989, 19990, 19991, 19992, 19993,19994, 19995, 19996, 20001, 20002, 20006, 20007, 20008, 20010, 2001 1, 20033, 20039, 20041 ,20042, 20043, 20045, 20047, 20057, 20058, 20059, 20060, 20063, 20065, 20067, 20068, 20071,20076, 20081, 20082, 20083, 20084, 20085, 20086, 20087, 20088, 20089, 20090, 20093, 20094,20095, 20096, 20097, 20100, 20101, 20102, 20105, 20106, 20109, 20110, 20112, 20118, 20121,20122, 20124, 20125, 20126, 20129, 20130, 20133, 20134, 20143, 20144, 20145, 20146, 20147,20153, 20154, 20155, 20156, 20159, 20180, 20181, 20182, 20186, 20187, 20188, 20189, 20190,20191, 20192, 20193, 20194, 20195, 20197, 20198, 20199, 20200, 20219, 20220, 20221, 20222,20232, 20233, 20245, 20253, 20255, 20265, 20267, 20268, 20275, 20284, 20315, 20321, 20324,20329, 20333, 20346, 20359, 20360, 20363, or 20374. In some embodiments, the msDNA comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to:CCTTGAATTACAGGCACGAAAATGGACACGCGCCCAAGATGCGACAAAGGAGCAGATTG GGCGCGTGGCCGTTTTCGTGCCGCCAATGTCGTTC (SEQ ID NO: 20010);CCTCAAACTATCCCGAAGGCGCGTAACTGTACTGAGCGCGTCAGCGCGACGTACGCGAA GCGTACTTAAGTACAATGGAGCGACTTTGAGTATCTGGACAAACTAC (SEQ ID NO: 20001);ACTTAAGGATGCGGTGGGTATAATCCTCCGCCGCGAAATGCAAGGCAATGAGCGGCGGA GGATTATACCTACCGCAACCTCATAAACCTTC (SEQ ID NO: 19642); or CATAAATTACAGGCACGAAACAGGCCACGCGCCCAAGGAGTGACGACGGAGCGAGTTGG GCGCGTGGCCTGTTTCGTGCCGCCGATGTATATC (SEQ ID NO: 19963).
[0122] In some embodiments, the msDNA comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20010. In some embodiments, the msDNA comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20001. In some embodiments, the msDNA comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19642. In some embodiments, the msDNA comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19963.
[0123] In some embodiments, the msDNA comprises the nucleic acid sequence of any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621, 19622, 19623, 19624, 19625, 19626,19627, 19628, 19629, 19630, 19631, 19632, 19633, 19636, 19637, 19638, 19639, 19640, 19641,19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649, 19653, 19654, 19664, 19665, 19667,19676, 19677, 19690, 19691, 19692, 19695, 19697, 19699, 19703, 19709, 19715, 19719, 19721,19755, 19855, 19927, 19940, 19941 , 19947, 19948, 19949, 19953, 19954, 19955, 19960, 19961 ,19962, 19963, 19965, 19966, 19967, 19968, 19969, 19970, 19971, 19972, 19973, 19975, 19976,19977, 19981, 19984, 19985, 19988, 19989, 19990, 19991, 19992, 19993, 19994, 19995, 19996,20001, 20002, 20006, 20007, 20008, 20010, 20011, 20033, 20039, 20041, 20042, 20043, 20045,20047, 20057, 20058, 20059, 20060, 20063, 20065, 20067, 20068, 20071, 20076, 20081, 20082,20083, 20084, 20085, 20086, 20087, 20088, 20089, 20090, 20093, 20094, 20095, 20096, 20097,20100, 20101, 20102, 20105, 20106, 20109, 20110, 20112, 20118, 20121, 20122, 20124, 20125,20126, 20129, 20130, 20133, 20134, 20143, 20144, 20145, 20146, 20147, 20153, 20154, 20155,20156, 20159, 20180, 20181, 20182, 20186, 20187, 20188, 20189, 20190, 20191, 20192, 20193,20194, 20195, 20197, 20198, 20199, 20200, 20219, 20220, 20221, 20222, 20232, 20233, 20245,20253, 20255, 20265, 20267, 20268, 20275, 20284, 20315, 20321, 20324, 20329, 20333, 20346,20359, 20360, 20363, or 20374.
[0124] In some embodiments, the msDNA comprises the nucleic acid sequence of any one of SEQ ID NOs: 20010, SEQ ID NO: 20001, SEQ ID NO: 19642, or SEQ ID NO: 19963. In some embodiments, the msDNA comprises the nucleic acid sequence of SEQ ID NO: 20010. In some embodiments, the msDNA comprises the nucleic acid sequence of SEQ ID NO: 20001. In some embodiments, the msDNA comprises the nucleic acid sequence of SEQ ID NO: 19642. In some embodiments, the msDNA comprises the nucleic acid sequence of SEQ ID NO: 19963.
[0125] In some embodiments, the msd comprises a nucleic acid sequence as provided in any one ncRNA provided in Table X and Table Y. In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the nucleic acid sequence of any one of msDNA provided in Table X and Table Y. For example, in reference to CP017253 locus, the ncRNA comprises the nucleic acid sequence of:ACATAGATTTCTTGGCCTTTATGCTATGGTGTTGCGCCATGGTGGAGATTTGT CAGATACACATCATTAGGTTGCGCGCAATTCGCTACGCTACCAATACTGT GCACATTAAAAATCGAGCCTTTGGTTATGTGACAGTATTGGTGCTACGCT GAAGTGTCACAACCAAATATAAGAATTGTTAGCAAGAAATCTATGT (SEQ ID NO: 23969); and the msDNA comprises the nucleic acid sequence of:GGTTGTGACACTTCAGCGTAGCACCAATACTGTCACATAACCAAAGGCTCGA TTTTTAATGTGCACAGTATTGGTAGCGTAGCGAATTGCGCGCAACC (SEQ ID NO: 23967).As described herein, the ncRNA nucleic acid sequence comprises the msr and the msd nucleic acid sequences. Thus, the msd nucleic acid sequence coding for the msDNA of SEQ ID NO: 23967 is the reverse complement of the msDNA of SEQ ID NO: 23967, which is underlined and bolded in the ncRNA nucleic acid sequence of SEQ ID NO: 23969. Accordingly, each of the ncRNA provided herein comprises the msd nucleic acid sequence that is used as a template for the reverse transcriptase resulting in the corresponding msDNA as provided herein.
[0126] In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or atleast 99% sequence identity any msDNA of Table X (column “msDNA_seqid”). In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any msDNA of Table Y (column “msDNA_seqid”). In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590,19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621, 19622, 19623, 19624, 19625, 19626, 19627, 19628, 19629, 19630, 19631, 19632, 19633, 19636, 19637, 19638, 19639, 19640, 19641, 19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649, 19653, 19654, 19664, 19665, 19667, 19676, 19677, 19690, 19691, 19692, 19695, 19697, 19699, 19703, 19709, 19715, 19719, 19721, 19755, 19855, 19927, 19940, 19941, 19947, 19948, 19949, 19953, 19954, 19955, 19960, 19961, 19962, 19963, 19965, 19966, 19967, 19968, 19969, 19970, 19971, 19972, 19973, 19975, 19976, 19977, 19981 , 19984, 19985, 19988, 19989, 19990, 19991, 19992, 19993, 19994, 19995, 19996, 20001, 20002, 20006, 20007, 20008, 20010, 20011, 20033, 20039, 20041, 20042, 20043, 20045, 20047, 20057, 20058, 20059, 20060, 20063, 20065, 20067, 20068, 20071, 20076, 20081, 20082, 20083, 20084, 20085, 20086, 20087, 20088, 20089, 20090, 20093, 20094, 20095, 20096, 20097, 20100, 20101, 20102, 20105, 20106, 20109, 20110, 20112, 20118, 20121, 20122, 20124, 20125, 20126, 20129, 20130, 20133, 20134, 20143, 20144, 20145, 20146, 20147, 20153, 20154, 20155, 20156, 20159, 20180, 20181, 20182, 20186, 20187, 20188, 20189, 20190, 20191, 20192, 20193, 20194, 20195, 20197, 20198, 20199, 20200, 20219, 20220, 20221, 20222, 20232, 20233, 20245, 20253, 20255, 20265, 20267, 20268, 20275, 20284, 20315, 20321, 20324, 20329, 20333, 20346, 20359, 20360, 20363, or 20374. In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence having at least 5%, at least 10%, at least 15 %, at least 20%, at least 25%, at least 30%, at least35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20010, SEQ ID NO: 20001, SEQ ID NO: 19642, or SEQ ID NO: 19963.
[0127] In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20010. In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20001. In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19642. In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19963.
[0128] In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence of any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621, 19622, 19623, 19624, 19625, 19626, 19627, 19628, 19629, 19630,19631, 19632, 19633, 19636, 19637, 19638, 19639, 19640, 19641, 19642, 19643, 19644, 19645,19646, 19647, 19648, 19649, 19653, 19654, 19664, 19665, 19667, 19676, 19677, 19690, 19691,19692, 19695, 19697, 19699, 19703, 19709, 19715, 19719, 19721, 19755, 19855, 19927, 19940,19941, 19947, 19948, 19949, 19953, 19954, 19955, 19960, 19961, 19962, 19963, 19965, 19966,19967, 19968, 19969, 19970, 19971, 19972, 19973, 19975, 19976, 19977, 19981, 19984, 19985,19988, 19989, 19990, 19991, 19992, 19993, 19994, 19995, 19996, 20001, 20002, 20006, 20007,20008, 20010, 20011, 20033, 20039, 20041, 20042, 20043, 20045, 20047, 20057, 20058, 20059,20060, 20063, 20065, 20067, 20068, 20071, 20076, 20081, 20082, 20083, 20084, 20085, 20086,20087, 20088, 20089, 20090, 20093, 20094, 20095, 20096, 20097, 20100, 20101, 20102, 20105,20106, 20109, 20110, 20112, 20118, 20121, 20122, 20124, 20125, 20126, 20129, 20130, 20133,20134, 20143, 20144, 20145, 20146, 20147, 20153, 20154, 20155, 20156, 20159, 20180, 20181,20182, 20186, 20187, 20188, 20189, 20190, 20191, 20192, 20193, 20194, 20195, 20197, 20198,20199, 20200, 20219, 20220, 20221, 20222, 20232, 20233, 20245, 20253, 20255, 20265, 20267,20268, 20275, 20284, 20315, 20321, 20324, 20329, 20333, 20346, 20359, 20360, 20363, or 20374.
[0129] In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence of any one of SEQ ID NOs: 20010, SEQ ID NO: 20001, SEQ ID NO: 19642, or SEQ ID NO: 19963. In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence of SEQ ID NO: 20010. In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence of SEQ ID NO: 20001. In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprises the nucleic acid sequence of SEQ ID NO: 19642. In some embodiments, the msd comprises a reverse complement nucleic acid sequence of the msDNA comprising the nucleic acid sequence of SEQ ID NO: 19963.
[0130] In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identityto any one SEQ ID NO: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604,19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671,19672, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706,19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939,19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983,19986, 19998, 19999, 20000, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030,20032, 20034, 20036, 20038, 20040, 20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056,20061, 20062, 20070, 20072, 20074, 20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117,20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165,20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204,20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238,20242, 20243, 20247, 20248, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272,20274, 20276, 20277, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296,20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328,20331, 20332, 20335, 20338, 20339, 20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367,20369, 20371, 20373, or 20377.
[0131] In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 71%, at least 72%, at least 73%, at least 74%, at least 75%, at least 76%, at least 77%, at least 78%, at least 79%, at least 80%, at least 81%, at least 82%, at least 83%, at least 84%, at least 85%, at least 86%, at least 87%, at least 88%, at least 89%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one SEQ IDNO: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593, 19598, 19601, 19602, 19604, 19605, 19609, 19610,19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19671, 19674, 19675, 19678, 19680,19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717,19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950,19951, 19957, 19959, 19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029,20030, 20032, 20034, 20036, 20038, 20040, 20046, 20049, 20050, 20052, 20053, 20055, 20056,20070, 20074, 20075, 20078, 20092, 20099, 20103, 20108, 20117, 20128, 20132, 20136, 20138,20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174,20176, 20178, 20179, 20184, 20185, 20196, 20204, 20207, 20209, 20211, 20213, 20214, 20218,20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20249, 20250, 20251,20252, 20259, 20260, 20262, 20271, 20272, 20274, 20279, 20280, 20281, 20286, 20287, 20289,20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318,20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358, 20362, 20365,20367, 20369, 20371, 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386,20386, 20387, 20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398,20399, 20400, 20401, 20402, 20403, 20404, or 20405.
[0132] In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19610. In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19979. In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19931. In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19722. In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24020.
[0133] In some embodiments, the polynucleotide comprises the nucleic acid sequence of any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574,19576, 19578, 19579, 19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596, 19598, 19601 ,19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666,19669, 19671, 19672, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702,19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933,19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19974, 19978, 19979,19980, 19983, 19986, 19998, 19999, 20000, 20005, 20016, 20018, 20020, 20023, 20025, 20028,20029, 20030, 20032, 20034, 20036, 20038, 20040, 20044, 20046, 20049, 20050, 20052, 20053,20055, 20056, 20061, 20062, 20070, 20072, 20074, 20075, 20078, 20092, 20099, 20103, 20107,20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161,20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20201,20202, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235,20237, 20238, 20242, 20243, 20247, 20248, 20249, 20250, 20251, 20252, 20259, 20260, 20262,20271, 20272, 20274, 20276, 20277, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291,20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320,20323, 20328, 20331, 20332, 20335, 20338, 20339, 20340, 20342, 20345, 20347, 20358, 20362,20365, 20367, 20369, 20371, 20373, or 20377.
[0134] In some embodiments, the polynucleotide comprises the nucleic acid sequence of any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576,19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593, 19598, 19601, 19602, 19604, 19605,19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19671 , 19674, 19675,19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711,19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945,19946, 19950, 19951, 19957, 19959, 19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025,20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20046, 20049, 20050, 20052, 20053,20055, 20056, 20070, 20074, 20075, 20078, 20092, 20099, 20103, 20108, 20117, 20128, 20132,20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171,20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20204, 20207, 20209, 20211, 20213,20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20249,20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20279, 20280, 20281, 20286,20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311,20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358,20362, 20365, 20367, 20369, 20371 , 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386, 20386, 20387, 20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398, 20399, 20400, 20401, 20402, 20403, 20404, or 20405.
[0135] In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 19610, SEQ ID NO: 19979, SEQ ID NO: 19931, SEQ ID NO: 19722, or SEQ ID NO: 24020. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 19610. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 19979. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 19931. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 19722. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 24020.
[0136] In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any msr locus of Table X (column “msr seqid”). In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any msr locus of Table Y (column “msr seqid”). In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any of SEQ ID NO: 19552, 19555, 19556, 19557, 19559, 19560, 19562, 19565, 19568, 19569, 19571, 19573, 19575, 19577, 19584, 19586, 19588, 19592, 19594, 19597,19599, 19600, 19603, 19615, 19618, 19619, 19620, 19635, 19650, 19655, 19657, 19659, 19661,19663, 19668, 19670, 19673, 19679, 19681, 19683, 19684, 19686, 19687, 19688, 19693, 19696,19700, 19701, 19705, 19707, 19712, 19713, 19716, 19718, 19720, 19723, 19724, 19784, 19850,19851, 19852, 19853, 19854, 19866, 19910, 19928, 19929, 19932, 19934, 19935, 19936, 19938,19942, 19952, 19956, 19958, 19964, 19982, 19987, 19997, 20003, 20004, 20009, 20012, 20013,20014, 20015, 20017, 20019, 20021, 20022, 20024, 20026, 20027, 20031, 20035, 20037, 20048,20051, 20054, 20064, 20066, 20069, 20073, 20077, 20079, 20080, 20091, 20098, 20104, 20111 ,20113, 20114, 20115, 20116, 20119, 20120, 20123, 20127, 20131, 20135, 20137, 20139, 20148,20149, 20150, 20157, 20160, 20162, 20164, 20166, 20167, 20168, 20170, 20173, 20175, 20177,20183, 20203, 20206, 20208, 20210, 20212, 20215, 20224, 20225, 20226, 20228, 20230, 20234,20236, 20239, 20240, 20241, 20244, 20246, 20254, 20256, 20257, 20258, 20261, 20263, 20264,20266, 20269, 20270, 20273, 20278, 20282, 20283, 20285, 20288, 20292, 20293, 20294, 20297,20299, 20301, 20303, 20305, 20307, 20309, 20310, 20312, 20314, 20316, 20317, 20322, 20325,20326, 20327, 20330, 20334, 20337, 20341, 20353, 20355, 20357, 20361, 20364, 20366, 20368,20370, 20372, 20375, or 20376.
[0137] In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20115. In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 7. In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20026. In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19599. In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20120.
[0138] In some embodiments, the polynucleotide comprises the nucleic acid sequence of any one of SEQ ID NOs: 19552, 19555, 19556, 19557, 19559, 19560, 19562, 19565,19568, 19569, 19571 , 19573, 19575, 19577, 19584, 19586, 19588, 19592, 19594, 19597, 19599,19600, 19603, 19615, 19618, 19619, 19620, 19635, 19650, 19655, 19657, 19659, 19661, 19663,19668, 19670, 19673, 19679, 19681, 19683, 19684, 19686, 19687, 19688, 19693, 19696, 19700,19701, 19705, 19707, 19712, 19713, 19716, 19718, 19720, 19723, 19724, 19784, 19850, 19851,19852, 19853, 19854, 19866, 19910, 19928, 19929, 19932, 19934, 19935, 19936, 19938, 19942,19952, 19956, 19958, 19964, 19982, 19987, 19997, 20003, 20004, 20009, 20012, 20013, 20014, 20015, 20017, 20019, 20021, 20022, 20024, 20026, 20027, 20031, 20035, 20037, 20048, 20051,20054, 20064, 20066, 20069, 20073, 20077, 20079, 20080, 20091, 20098, 20104, 20111, 20113,20114, 20115, 20116, 20119, 20120, 20123, 20127, 20131, 20135, 20137, 20139, 20148, 20149,20150, 20157, 20160, 20162, 20164, 20166, 20167, 20168, 20170, 20173, 20175, 20177, 20183,20203, 20206, 20208, 20210, 20212, 20215, 20224, 20225, 20226, 20228, 20230, 20234, 20236,20239, 20240, 20241, 20244, 20246, 20254, 20256, 20257, 20258, 20261, 20263, 20264, 20266,20269, 20270, 20273, 20278, 20282, 20283, 20285, 20288, 20292, 20293, 20294, 20297, 20299,20301, 20303, 20305, 20307, 20309, 20310, 20312, 20314, 20316, 20317, 20322, 20325, 20326,20327, 20330, 20334, 20337, 20341, 20353, 20355, 20357, 20361, 20364, 20366, 20368, 20370,20372, 20375, or 20376.
[0139] In some embodiments, the polynucleotide comprises the nucleic acid sequence of any one of SEQ ID NOs: 20115, 7, SEQ ID NO: 20026, SEQ ID NO: 19599, or SEQ ID NO: 20120. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 20115. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 7. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 20026. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 19599. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 20120.
[0140] In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity any msd locus of Table X (column “msd seqid”). In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%,at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any msd locus of Table Y (column “msd seqid”). In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621, 19622,19623, 19624, 19625, 19626, 19627, 19628, 19629, 19630, 19631, 19632, 19633, 19636, 19637,19638, 19639, 19640, 19641, 19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649, 19653,19654, 19664, 19665, 19667, 19676, 19677, 19690, 19691, 19692, 19695, 19697, 19699, 19703,19709, 19715, 19719, 19721, 19755, 19855, 19927, 19940, 19941, 19947, 19948, 19949, 19953,19954, 19955, 19960, 19961, 19962, 19963, 19965, 19966, 19967, 19968, 19969, 19970, 19971,19972, 19973, 19975, 19976, 19977, 19981, 19984, 19985, 19988, 19989, 19990, 19991, 19992,19993, 19994, 19995, 19996, 20001, 20002, 20006, 20007, 20008, 20010, 20011, 20033, 20039,20041, 20042, 20043, 20045, 20047, 20057, 20058, 20059, 20060, 20063, 20065, 20067, 20068,20071, 20076, 20081, 20082, 20083, 20084, 20085, 20086, 20087, 20088, 20089, 20090, 20093,20094, 20095, 20096, 20097, 20100, 20101 , 20102, 20105, 20106, 20109, 20110, 20112, 20118,20121, 20122, 20124, 20125, 20126, 20129, 20130, 20133, 20134, 20143, 20144, 20145, 20146,20147, 20153, 20154, 20155, 20156, 20159, 20180, 20181, 20182, 20186, 20187, 20188, 20189,20190, 20191, 20192, 20193, 20194, 20195, 20197, 20198, 20199, 20200, 20219, 20220, 20221,20222, 20232, 20233, 20245, 20253, 20255, 20265, 20267, 20268, 20275, 20284, 20315, 20321,20324, 20329, 20333, 20346, 20359, 20360, 20363, or 20374.
[0141] In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20010. In some embodiments, the polynucleotide comprises the nucleic acidsequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 20001. In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19642. In some embodiments, the polynucleotide comprises the nucleic acid sequence having at least 5%, at least 10%, at least 15%, at least 20%, at least 25%, at least 30%, at least 35%, at least 40%, at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19963.
[0142] In some embodiments, the polynucleotide comprises the nucleic acid sequence of any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606,19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621, 19622, 19623, 19624, 19625,19626, 19627, 19628, 19629, 19630, 19631, 19632, 19633, 19636, 19637, 19638, 19639, 19640,19641 , 19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649, 19653, 19654, 19664, 19665,19667, 19676, 19677, 19690, 19691, 19692, 19695, 19697, 19699, 19703, 19709, 19715, 19719,19721, 19755, 19855, 19927, 19940, 19941, 19947, 19948, 19949, 19953, 19954, 19955, 19960,19961, 19962, 19963, 19965, 19966, 19967, 19968, 19969, 19970, 19971, 19972, 19973, 19975,19976, 19977, 19981, 19984, 19985, 19988, 19989, 19990, 19991, 19992, 19993, 19994, 19995, 19996, 20001, 20002, 20006, 20007, 20008, 20010, 20011, 20033, 20039, 20041, 20042, 20043,20045, 20047, 20057, 20058, 20059, 20060, 20063, 20065, 20067, 20068, 20071, 20076, 20081,20082, 20083, 20084, 20085, 20086, 20087, 20088, 20089, 20090, 20093, 20094, 20095, 20096,20097, 20100, 20101, 20102, 20105, 20106, 20109, 20110, 20112, 20118, 20121, 20122, 20124,20125, 20126, 20129, 20130, 20133, 20134, 20143, 20144, 20145, 20146, 20147, 20153, 20154,20155, 20156, 20159, 20180, 20181, 20182, 20186, 20187, 20188, 20189, 20190, 20191, 20192,20193, 20194, 20195, 20197, 20198, 20199, 20200, 20219, 20220, 20221, 20222, 20232, 20233,20245, 20253, 20255, 20265, 20267, 20268, 20275, 20284, 20315, 20321, 20324, 20329, 20333, 20346, 20359, 20360, 20363, or 20374.
[0143] In some embodiments, the polynucleotide comprises the nucleic acid sequence of any one of SEQ ID NOs: 20010, SEQ ID NO: 20001, SEQ ID NO: 19642, or SEQ ID NO: 19963. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 20010. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 20001. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 19642. In some embodiments, the polynucleotide comprises the nucleic acid sequence of SEQ ID NO: 19963.
[0144] In some embodiments, the msr / msd locus, the msr locus, the msd locus, or the polynucleotide, such as those provided herein, may comprise one or more sequence modifications (e.g., an insertion, deletion, and / or substitution of one or more nucleotide(s)) that: a) modulates (e.g., enhances) reverse transcription, processivity, accuracy / fidelity, and / or production of the msDNA (e.g., in the mammalian cell); b) modulates (e.g., reduces) immunogenicity of ncRNA encoded by the engineered retron (e.g., the msr locus and / or the msd locus) in a host (e g., a host comprising the mammalian cell); c) modulates (e.g., inhibits, either permanently or transiently) a function of the msDNA; d) modulates (e.g., improves) efficiency of targeted genome editing / engineering; and / or e) modulates (e.g., increases) stability of the ncRNA. In some embodiments, the msr / msd locus, the msr locus, the msd locus, or the polynucleotide, such as those provided herein, may comprise one or more sequence modifications (e.g., an insertion, deletion, and / or substitution of one or more nucleotide(s)) in association with a donor DNA sequence. In some embodiments, the msr / msd locus, the msr locus, the msd locus, or the polynucleotide, such as those provided herein, may comprise one or more sequence modifications (e.g., an insertion, deletion, and / or substitution of one or more nucleotide(s)) as a result of the donor DNA sequence insertion.
[0145] In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605, 19609, 19610,19634, 19651 , 19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671, 19672, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983, 19986, 19998, 19999, 20000, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20061, 20062, 20070, 20072, 20074, 20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20248, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20276, 20277, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367, 20369, 20371, 20373, or 20377.
[0146] In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587,19589, 19591, 19593, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652,19656, 19658, 19660, 19662, 19666, 19671, 19674, 19675, 19678, 19680, 19682, 19685, 19689,19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834,19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959,19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034,20036, 20038, 20040, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20070, 20074, 20075,20078, 20092, 20099, 20103, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142,20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179,20184, 20185, 20196, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229,20231, 20235, 20237, 20238, 20242, 20243, 20247, 20249, 20250, 20251, 20252, 20259, 20260,20262, 20271, 20272, 20274, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295,20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323,20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358, 20362, 20365, 20367, 20369, 20371,20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386, 20386, 20387, 20388,20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398, 20399, 20400, 20401,20402, 20403, 20404, or 20405.
[0147] In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one SEQ ID NO: 19610, SEQ ID NO: 19979, SEQ ID NO: 19931, SEQ ID NO: 19722, or SEQ ID NO: 24020. In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19610. In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19979. In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19931. In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19722. In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 24020.
[0148] In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, atleast 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572,19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671, 19672, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983, 19986, 19998, 19999, 20000, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20061, 20062, 20070, 20072, 20074, 20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20248, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20276, 20277, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367, 20369, 20371, 20373, or 20377.
[0149] In some embodiments , the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, atleast 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, at least 110, or more nucleotides as compared to any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19671, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20070, 20074, 20075, 20078, 20092, 20099, 20103, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358, 20362, 20365, 20367, 20369, 20371, 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386, 20386, 20387, 20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398, 20399, 20400, 20401, 20402, 20403, 20404, or 20405.
[0150] In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, atleast 1 1 , at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to any one SEQ ID NO: 19610, SEQ ID NO: 19979, SEQ ID NO: 19931, SEQ ID NO: 19722, or SEQ ID NO: 24020. In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51 , at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19610. In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31 , at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 1 10 nucleotides as compared to SEQ ID NO: 19979. In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71 , at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19931. In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, atleast 48, at least 49, at least 50, at least 51 , at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19722. In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91 , at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 24020.
[0151] In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656,19658, 19660, 19662, 19666, 19669, 19671 , 19672, 19674, 19675, 19678, 19680, 19682, 19685,19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785,19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957,19959, 19974, 19978, 19979, 19980, 19983, 19986, 19998, 19999, 20000, 20005, 20016, 20018,20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20044, 20046,20049, 20050, 20052, 20053, 20055, 20056, 20061, 20062, 20070, 20072, 20074, 20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142,20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179,20184, 20185, 20196, 20201, 20202, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223,20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20248, 20249, 20250, 20251,20252, 20259, 20260, 20262, 20271, 20272, 20274, 20276, 20277, 20279, 20280, 20281, 20286,20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311,20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20340, 20342,20345, 20347, 20358, 20362, 20365, 20367, 20369, 20371, 20373, or 20377.
[0152] In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, 1 10, or more nucleotides as compared to any one of SEQ ID NOs: 19553, 19558, 19563,19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593,19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660,19662, 19666, 19671, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702,19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933,19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19983, 19998, 20005,20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040,20046, 20049, 20050, 20052, 20053, 20055, 20056, 20070, 20074, 20075, 20078, 20092, 20099,20103, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158,20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196,20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237,20238, 20242, 20243, 20247, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272,20274, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358, 20362, 20365, 20367, 20369, 20371, 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386, 20386, 20387, 20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398, 20399, 20400, 20401, 20402, 20403, 20404, or 20405.
[0153] In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32,33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58,59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107,108, 109, or 110 nucleotides as compared to any one SEQ ID NO: 19610, SEQ ID NO: 19979,SEQ ID NO: 19931, SEQ ID NO: 19722, or SEQ ID NO: 24020. In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19,20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45,46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71,72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97,98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 19610. In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31 , 32, 33,34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59,60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85,86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108,109, or 110 nucleotides as compared to SEQ ID NO: 19979. In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19,20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45,46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71,72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97,98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 19931. In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33,34, 35, 36, 37, 38, 39, 40, 41 , 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59,60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85,86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108,109, or 110 nucleotides as compared to SEQ ID NO: 19722. In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19,20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45,46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71,72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97,98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 24020.
[0154] In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613,19614, 19616, 19617, 19621, 19622, 19623, 19624, 19625, 19626, 19627, 19628, 19629, 19630,19631, 19632, 19633, 19636, 19637, 19638, 19639, 19640, 19641, 19642, 19643, 19644, 19645,19646, 19647, 19648, 19649, 19653, 19654, 19664, 19665, 19667, 19676, 19677, 19690, 19691,19692, 19695, 19697, 19699, 19703, 19709, 19715, 19719, 19721, 19755, 19855, 19927, 19940,19941, 19947, 19948, 19949, 19953, 19954, 19955, 19960, 19961, 19962, 19963, 19965, 19966,19967, 19968, 19969, 19970, 19971, 19972, 19973, 19975, 19976, 19977, 19981, 19984, 19985,19988, 19989, 19990, 19991, 19992, 19993, 19994, 19995, 19996, 20001, 20002, 20006, 20007,20008, 20010, 20011, 20033, 20039, 20041, 20042, 20043, 20045, 20047, 20057, 20058, 20059,20060, 20063, 20065, 20067, 20068, 20071, 20076, 20081, 20082, 20083, 20084, 20085, 20086,20087, 20088, 20089, 20090, 20093, 20094, 20095, 20096, 20097, 20100, 20101, 20102, 20105,20106, 20109, 20110, 20112, 20118, 20121, 20122, 20124, 20125, 20126, 20129, 20130, 20133,20134, 20143, 20144, 20145, 20146, 20147, 20153, 20154, 20155, 20156, 20159, 20180, 20181, 20182, 20186, 20187, 20188, 20189, 20190, 20191, 20192, 20193, 20194, 20195, 20197, 20198,20199, 20200, 20219, 20220, 20221, 20222, 20232, 20233, 20245, 20253, 20255, 20265, 20267,20268, 20275, 20284, 20315, 20321, 20324, 20329, 20333, 20346, 20359, 20360, 20363, or 20374.
[0155] In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one SEQ ID NO: 20010, SEQ ID NO: 20001, SEQ ID NO: 19642, or SEQ ID NO: 19963. In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 20010. In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 20001. In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19642. In some embodiments, the polynucleotide comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19963.
[0156] In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621, 19622, 19623, 19624,19625, 19626, 19627, 19628, 19629, 19630, 19631, 19632, 19633, 19636, 19637, 19638, 19639,19640, 19641, 19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649, 19653, 19654, 19664,19665, 19667, 19676, 19677, 19690, 19691, 19692, 19695, 19697, 19699, 19703, 19709, 19715,19719, 19721, 19755, 19855, 19927, 19940, 19941, 19947, 19948, 19949, 19953, 19954, 19955,19960, 19961, 19962, 19963, 19965, 19966, 19967, 19968, 19969, 19970, 19971, 19972, 19973,19975, 19976, 19977, 19981, 19984, 19985, 19988, 19989, 19990, 19991, 19992, 19993, 19994,19995, 19996, 20001, 20002, 20006, 20007, 20008, 20010, 20011, 20033, 20039, 20041, 20042,20043, 20045, 20047, 20057, 20058, 20059, 20060, 20063, 20065, 20067, 20068, 20071, 20076,20081, 20082, 20083, 20084, 20085, 20086, 20087, 20088, 20089, 20090, 20093, 20094, 20095,20096, 20097, 20100, 20101, 20102, 20105, 20106, 20109, 20110, 20112, 20118, 20121, 20122,20124, 20125, 20126, 20129, 20130, 20133, 20134, 20143, 20144, 20145, 20146, 20147, 20153,20154, 20155, 20156, 20159, 20180, 20181, 20182, 20186, 20187, 20188, 20189, 20190, 20191,20192, 20193, 20194, 20195, 20197, 20198, 20199, 20200, 20219, 20220, 20221, 20222, 20232,20233, 20245, 20253, 20255, 20265, 20267, 20268, 20275, 20284, 20315, 20321, 20324, 20329,20333, 20346, 20359, 20360, 20363, or 20374.
[0157] In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least87, at least 88, at least 89, at least 90, at least 91 , at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to any one SEQ ID NO: 20010, SEQ ID NO: 20001, SEQ ID NO: 19642, or SEQ ID NO: 19963. In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 1 1, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 20010. In some embodiments, the polynucleotide comprises a deletion of at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, atleast 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 1 10 nucleotides as compared to SEQ ID NO: 20001. In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19642. In some embodiments, the polynucleotide comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31 , at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19963.
[0158] In some embodiments, the polynucleotide comprises a deletion of 1 , 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621,19622, 19623, 19624, 19625, 19626, 19627, 19628, 19629, 19630, 19631, 19632, 19633, 19636,19637, 19638, 19639, 19640, 19641, 19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649,19653, 19654, 19664, 19665, 19667, 19676, 19677, 19690, 19691, 19692, 19695, 19697, 19699,19703, 19709, 19715, 19719, 19721, 19755, 19855, 19927, 19940, 19941, 19947, 19948, 19949,19953, 19954, 19955, 19960, 19961, 19962, 19963, 19965, 19966, 19967, 19968, 19969, 19970,19971, 19972, 19973, 19975, 19976, 19977, 19981, 19984, 19985, 19988, 19989, 19990, 19991,19992, 19993, 19994, 19995, 19996, 20001, 20002, 20006, 20007, 20008, 20010, 20011, 20033, 20039, 20041, 20042, 20043, 20045, 20047, 20057, 20058, 20059, 20060, 20063, 20065, 20067,20068, 20071, 20076, 20081, 20082, 20083, 20084, 20085, 20086, 20087, 20088, 20089, 20090,20093, 20094, 20095, 20096, 20097, 20100, 20101, 20102, 20105, 20106, 20109, 20110, 20112,20118, 20121, 20122, 20124, 20125, 20126, 20129, 20130, 20133, 20134, 20143, 20144, 20145,20146, 20147, 20153, 20154, 20155, 20156, 20159, 20180, 20181, 20182, 20186, 20187, 20188,20189, 20190, 20191 , 20192, 20193, 20194, 20195, 20197, 20198, 20199, 20200, 20219, 20220,20221, 20222, 20232, 20233, 20245, 20253, 20255, 20265, 20267, 20268, 20275, 20284, 20315,20321, 20324, 20329, 20333, 20346, 20359, 20360, 20363, or 20374.
[0159] In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to any one SEQ ID NO: 20010, SEQ ID NO: 20001, SEQ ID NO: 19642, OR SEQ ID NO: 19963. In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52,53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103,104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 20010. In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40,41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66,67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92,93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 20001. In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28,29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54,55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80,81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104,105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 19642. In some embodiments, the polynucleotide comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40,41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66,67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92,93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 19963.
[0160] In some embodiments, the msd and msr regions of polynucleotide transcripts contain first and second inverted repeat sequences, which can form a stable stem structure. Without being bound to any particular theory, the combined msr-msd region of the polynucleotide transcript serves not only as a template for reverse transcription but also serves as a primer for a reverse transcriptase. In some embodiments of polynucleotide-guide RNA cassettes, the first inverted repeat sequence coding region is located within the 5' end of the msr / msd locus. In other embodiments, the second inverted repeat sequence coding region is located within the 3' of the msr / msd locus.
[0161] In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of ACATAATTGCGCAGCTCAGTTAC (SEQ ID NO: 24022); GCTCAGTTAC (SEQ ID NO: 24023); CCAGCAGTGGCAAT (SEQ ID NO: 24024); ATTCATGTAATCTCTATATGTCCTTT (SEQ ID NO: 24025); orATCGACTAACTCAGTTACGCGCATA (SEQ ID NO: 24026). In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24022. In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24023. In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24024. In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24025. In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24026.
[0162] In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence of any one of SEQ ID NOs: 24022, SEQ ID NO: 24023, SEQ ID NO: 24024, SEQ ID NO: 24025, or SEQ ID NO: 24026. In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence of SEQ ID NO: 24022. In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acidsequence of SEQ ID NO: 24023. In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence of SEQ ID NO: 24024. In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence of SEQ ID NO: 24025. In some embodiments, the first inverted repeat sequence coding region comprises a nucleic acid sequence of SEQ ID NO: 24026.
[0163] In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of GTAGCTGAGCCTACCTATGACGT (SEQ ID NO: 24027); GTAGCTGAGC (SEQ ID NO: 24028); ATTGCCGCTGTTGG (SEQ ID NO: 24029); AAAGGACATAGAAGAACTGGCATTAAT (SEQ ID NO: 24030); orTATGGTAGCTGAGTCTGCCTAT (SEQ ID NO: 24031). In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24027. In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24028. In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24029. In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24030. In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, atleast 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 24031.
[0164] In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence of any one of SEQ ID NOs: 24027, SEQ ID NO: 24028, SEQ ID NO: 24029, SEQ ID NO: 24030, or SEQ ID NO: 24031. In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence of SEQ ID NO: 24027. In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence of SEQ ID NO: 24028. In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence of SEQ ID NO: 24029. In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence of SEQ ID NO: 24030. In some embodiments, the second inverted repeat sequence coding region comprises a nucleic acid sequence of SEQ ID NO: 24031.
[0165] One of ordinary skill in the art will understand that the sequence of an inverted repeat sequence coding region may be varied, so long as the sequence of the counterpart inverted repeat sequence coding region within the same retron is also varied such that the two resulting inverted repeat sequences (i.e., present within a polynucleotide transcript) are complementary and allow for the formation of a stable stem structure.
[0166] Without being bound to any particular theory, the stable stem structure may be used as a starting point for reverse transcription and the reverse transcriptase can initiate the transcription event by binding to this stem loop. For example, in some embodiments, RT may bind to the structure that is formed between nucleotides of the msr locus. However, the embodiments, provided for herein have demonstrated that the sequence of the msr / msd locus may be modified (e g, mutated, deleted, and insertions) such that the full-length msr / msd locus is not required to initiate reverse transcription. In contrast, a mutated msr / msd locus may be used to achieve greater efficiency in gene editing. In some embodiments, the msr locus sequence is not mutated or modified. In some embodiments, only the msd locus is mutated or modified as provided for herein. The mutated msr / msd locus can, in some embodiments, comprise the portion that binds to RT and where the reverse transcription may be initiated from.
[0017] Accordingly, in some embodiments, methods of modifying, or inducing one or more sequence modifications in one or more, target nucleic acids of interest at one or moretarget loci within a genome of a host cell, such as a mammalian cell are provided. In some embodiments, the methods comprise:(a) transforming the host cell with one or more vectors encoding a retron and a guide RNA (gRNA), such as those provided herein; and(b) culturing the host cell or transformed progeny of the host cell under conditions sufficient for expressing from the vector a retron donor DNA molecule comprising a polynucleotide transcript and a guide RNA (gRNA) molecule to induce one or more sequence modifications in one or more target nucleic acids of interest at the one or more target loci within the genome. In some embodiments, inducing one or more sequence modifications may refer to any modification of a gene sequence, such as, deletion, truncation, insertion, or point mutation.
[0168] As provided for herein, the msr / msd locus may be modified (mutated) and still be used to initiate reverse transcription. Thus, in some embodiments, the mutated or modified msr / msd locus may be within the retron. In some embodiments, the retron comprises a mutated msd locus, such as provided for herein. In some embodiments, the retron comprises a mutated msr locus, such as provided for herein. In some embodiments, the retron comprises a mutated msd locus and msr locus that does not comprise any mutations or modifications. In some embodiments, the mutated or modified msr / msd locus may be within the polynucleotide. In some embodiments, the polynucleotide comprises a mutated msd locus, such as provided for herein. In some embodiments, the polynucleotide comprises a mutated msr locus, such as provided for herein. In some embodiments, the polynucleotide comprises a mutated msd locus and msr locus that does not comprise any mutations or modifications.
[0169] In some embodiments, the msr locus is a mutated msr locus and / or the msd locus is a mutated msd locus. In some embodiments, the mutated msr or msd locus comprises a deletion of at least 1 nucleotide as compared to the wild-type.
[0170] In some embodiments, the mutated msr / msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671, 19672, 19674,19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710,19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944,19945, 19946, 19950, 19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983, 19986, 19998,19999, 20000, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034,20036, 20038, 20040, 20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20061, 20062,20070, 20072, 20074, 20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128, 20132,20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171,20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204, 20207, 20209,20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243,20247, 20248, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20276,20277, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300,20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332,20335, 20338, 20339, 20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367, 20369, 20371,20373, or 20377.
[0171] In some embodiments, the mutated msr / msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593, 19598, 19601 , 19602, 19604, 19605, 19609, 19610, 19634, 19651 ,19652, 19656, 19658, 19660, 19662, 19666, 19671, 19674, 19675, 19678, 19680, 19682, 19685,19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785,19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957,19959, 19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032,20034, 20036, 20038, 20040, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20070, 20074,20075, 20078, 20092, 20099, 20103, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141,20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178,20179, 20184, 20185, 20196, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227,20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20249, 20250, 20251, 20252, 20259,20260, 20262, 20271, 20272, 20274, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291,20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320,20323, 20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358, 20362, 20365, 20367, 20369,20371, 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386, 20386, 20387,20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398, 20399, 20400,20401, 20402, 20403, 20404, or 20405.
[0172] In some embodiments, the mutated msr / msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one SEQ ID NO: 19610, SEQ ID NO: 19979, SEQ ID NO: 19931, SEQ ID NO: 19722, or SEQ ID NO: 24020. In some embodiments, the mutated msr / msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19610. In some embodiments, the mutated msr / msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19979. In some embodiments, the mutated msr / msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19931. In some embodiments, the mutated msr / msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19722. In some embodiments, the mutated msr / msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 24020.
[0173] In some embodiments, the mutated msr / msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, atleast 19, at least 20, at least 21 , at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570,19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596,19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660,19662, 19666, 19669, 19671, 19672, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694,19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930,19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19974,19978, 19979, 19980, 19983, 19986, 19998, 19999, 20000, 20005, 20016, 20018, 20020, 20023,20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20044, 20046, 20049, 20050,20052, 20053, 20055, 20056, 20061 , 20062, 20070, 20072, 20074, 20075, 20078, 20092, 20099,20103, 20107, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152,20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185,20196, 20201, 20202, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20248, 20249, 20250, 20251, 20252, 20259,20260, 20262, 20271, 20272, 20274, 20276, 20277, 20279, 20280, 20281, 20286, 20287, 20289,20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318,20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20340, 20342, 20345, 20347,20358, 20362, 20365, 20367, 20369, 20371, 20373, or 20377.
[0174] In some embodiments, the mutated msr / msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, atleast 19, at least 20, at least 21 , at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, at least 110, or more nucleotides as compared to any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19671, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20070, 20074, 20075, 20078, 20092, 20099, 20103, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141 , 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358, 20362, 20365, 20367, 20369, 20371, 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386, 20386, 20387, 20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398, 20399, 20400, 20401, 20402, 20403, 20404, or 20405.
[0175] In some embodiments, the mutated msr / msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, atleast 19, at least 20, at least 21 , at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to any one SEQ ID NO: 19610, SEQ ID NO: 19979, SEQ ID NO: 19931, SEQ ID NO: 19722, or SEQ ID NO: 24020. In some embodiments, the mutated msr / msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51 , at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19610. In some embodiments, the mutated msr / msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, atleast 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19979. In some embodiments, the mutated msr / msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19931. In some embodiments, the mutated msr / msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61 , at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19722. In some embodiments, the mutated msr / msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101 , at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 24020.
[0176] In some embodiments, the mutated msr / msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29,30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55,56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81,82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105,106, 107, 108, 109, or 110 nucleotides as compared to any one of SEQ ID NOs: 19553, 19554,19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587,19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651,19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671, 19672, 19674, 19675, 19678, 19680,19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 1971 1, 19714, 19717,19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950,19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983, 19986, 19998, 19999, 20000, 20005,20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040,20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20061, 20062, 20070, 20072, 20074,20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128, 20132, 20136, 20138, 20140,20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176,20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204, 20207, 20209, 20211, 20213, 20214,20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20248, 20249,20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20276, 20277, 20279, 20280,20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306,20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339,20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367, 20369, 20371, 20373, or 20377.
[0177] In some embodiments, the mutated msr / msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29,30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55,56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81,82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105,106, 107, 108, 109, 110, or more nucleotides as compared to any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589,19591, 19593, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656,19658, 19660, 19662, 19666, 19671, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694,19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930,19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19983,19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036,20038, 20040, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20070, 20074, 20075, 20078,20092, 20099, 20103, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151,20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184,20185, 20196, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231,20235, 20237, 20238, 20242, 20243, 20247, 20249, 20250, 20251, 20252, 20259, 20260, 20262,20271, 20272, 20274, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296,20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358, 20362, 20365, 20367, 20369, 20371, 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386, 20386, 20387, 20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398, 20399, 20400, 20401, 20402, 20403, 20404, or 20405.
[0178] In some embodiments, the mutated msr / msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29,30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55,56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81,82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105,106, 107, 108, 109, or 110 nucleotides as compared to any one SEQ ID NO: 19610, SEQ ID NO: 19979, SEQ ID NO: 19931, SEQ ID NO: 19722, or SEQ ID NO: 24020. In some embodiments, the mutated msr / msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15,16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41,42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67,68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93,94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 19610. In some embodiments, the mutated msr / msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26,27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51 , 52,53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78,79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103,104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 19979. In some embodiments, the mutated msr / msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11,12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37,38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63,64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89,90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 19931. In some embodiments, the mutated msr / msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49,50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61 , 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 19722. In some embodiments, the mutated msr / msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33,34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59,60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85,86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108,109, or 110 nucleotides as compared to SEQ ID NO: 24020.
[0179] In some embodiments, the mutated msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621, 19622, 19623, 19624, 19625, 19626, 19627, 19628, 19629,19630, 19631, 19632, 19633, 19636, 19637, 19638, 19639, 19640, 19641, 19642, 19643, 19644,19645, 19646, 19647, 19648, 19649, 19653, 19654, 19664, 19665, 19667, 19676, 19677, 19690,19691, 19692, 19695, 19697, 19699, 19703, 19709, 19715, 19719, 19721, 19755, 19855, 19927,19940, 19941, 19947, 19948, 19949, 19953, 19954, 19955, 19960, 19961, 19962, 19963, 19965,19966, 19967, 19968, 19969, 19970, 19971 , 19972, 19973, 19975, 19976, 19977, 19981 , 19984,19985, 19988, 19989, 19990, 19991, 19992, 19993, 19994, 19995, 19996, 20001, 20002, 20006, 20007, 20008, 20010, 20011, 20033, 20039, 20041, 20042, 20043, 20045, 20047, 20057, 20058,20059, 20060, 20063, 20065, 20067, 20068, 20071, 20076, 20081, 20082, 20083, 20084, 20085,20086, 20087, 20088, 20089, 20090, 20093, 20094, 20095, 20096, 20097, 20100, 20101, 20102,20105, 20106, 20109, 20110, 20112, 20118, 20121, 20122, 20124, 20125, 20126, 20129, 20130,20133, 20134, 20143, 20144, 20145, 20146, 20147, 20153, 20154, 20155, 20156, 20159, 20180,20181, 20182, 20186, 20187, 20188, 20189, 20190, 20191, 20192, 20193, 20194, 20195, 20197,20198, 20199, 20200, 20219, 20220, 20221, 20222, 20232, 20233, 20245, 20253, 20255, 20265,20267, 20268, 20275, 20284, 20315, 20321, 20324, 20329, 20333, 20346, 20359, 20360, 20363, or 20374.
[0180] In some embodiments, the mutated msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one SEQ ID NO: 20010, SEQ ID NO: 20001, SEQ ID NO: 19642, OR SEQ ID NO: 19963. In some embodiments, the mutated msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 20010. In some embodiments, the mutated msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 20001. In some embodiments, the mutated msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19642. In some embodiments, the mutated msd locus comprises a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to SEQ ID NO: 19963.
[0181] In some embodiments, the mutated msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, atleast 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621, 19622, 19623,19624, 19625, 19626, 19627, 19628, 19629, 19630, 19631, 19632, 19633, 19636, 19637, 19638,19639, 19640, 19641, 19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649, 19653, 19654,19664, 19665, 19667, 19676, 19677, 19690, 19691, 19692, 19695, 19697, 19699, 19703, 19709,19715, 19719, 19721, 19755, 19855, 19927, 19940, 19941, 19947, 19948, 19949, 19953, 19954,19955, 19960, 19961, 19962, 19963, 19965, 19966, 19967, 19968, 19969, 19970, 19971, 19972,19973, 19975, 19976, 19977, 19981, 19984, 19985, 19988, 19989, 19990, 19991, 19992, 19993, 19994, 19995, 19996, 20001, 20002, 20006, 20007, 20008, 20010, 20011, 20033, 20039, 20041,20042, 20043, 20045, 20047, 20057, 20058, 20059, 20060, 20063, 20065, 20067, 20068, 20071,20076, 20081, 20082, 20083, 20084, 20085, 20086, 20087, 20088, 20089, 20090, 20093, 20094,20095, 20096, 20097, 20100, 20101, 20102, 20105, 20106, 20109, 20110, 20112, 20118, 20121,20122, 20124, 20125, 20126, 20129, 20130, 20133, 20134, 20143, 20144, 20145, 20146, 20147,20153, 20154, 20155, 20156, 20159, 20180, 20181, 20182, 20186, 20187, 20188, 20189, 20190,20191, 20192, 20193, 20194, 20195, 20197, 20198, 20199, 20200, 20219, 20220, 20221, 20222,20232, 20233, 20245, 20253, 20255, 20265, 20267, 20268, 20275, 20284, 20315, 20321, 20324,20329, 20333, 20346, 20359, 20360, 20363, or 20374.
[0182] In some embodiments, the mutated msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, atleast 87, at least 88, at least 89, at least 90, at least 91 , at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to any one SEQ ID NO: 20010, SEQ ID NO: 20001, SEQ ID NO: 19642, OR SEQ ID NO: 19963. In some embodiments, the mutated msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 20010. In some embodiments, the mutated msd locus comprises a deletion of at least 1 , at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101,at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 20001. In some embodiments, the mutated msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19642. In some embodiments, the mutated msd locus comprises a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleotides as compared to SEQ ID NO: 19963.
[0183] In some embodiments, the mutated msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617,19621, 19622, 19623, 19624, 19625, 19626, 19627, 19628, 19629, 19630, 19631, 19632, 19633,19636, 19637, 19638, 19639, 19640, 19641, 19642, 19643, 19644, 19645, 19646, 19647, 19648,19649, 19653, 19654, 19664, 19665, 19667, 19676, 19677, 19690, 19691, 19692, 19695, 19697,19699, 19703, 19709, 19715, 19719, 19721, 19755, 19855, 19927, 19940, 19941, 19947, 19948,19949, 19953, 19954, 19955, 19960, 19961, 19962, 19963, 19965, 19966, 19967, 19968, 19969,19970, 19971, 19972, 19973, 19975, 19976, 19977, 19981, 19984, 19985, 19988, 19989, 19990,19991, 19992, 19993, 19994, 19995, 19996, 20001, 20002, 20006, 20007, 20008, 20010, 20011,20033, 20039, 20041, 20042, 20043, 20045, 20047, 20057, 20058, 20059, 20060, 20063, 20065,20067, 20068, 20071, 20076, 20081, 20082, 20083, 20084, 20085, 20086, 20087, 20088, 20089,20090, 20093, 20094, 20095, 20096, 20097, 20100, 20101, 20102, 20105, 20106, 20109, 20110,20112, 20118, 20121, 20122, 20124, 20125, 20126, 20129, 20130, 20133, 20134, 20143, 20144,20145, 20146, 20147, 20153, 20154, 20155, 20156, 20159, 20180, 20181, 20182, 20186, 20187,20188, 20189, 20190, 20191, 20192, 20193, 20194, 20195, 20197, 20198, 20199, 20200, 20219,20220, 20221, 20222, 20232, 20233, 20245, 20253, 20255, 20265, 20267, 20268, 20275, 20284,20315, 20321, 20324, 20329, 20333, 20346, 20359, 20360, 20363, or 20374.
[0184] In some embodiments, the mutated msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to any one SEQ IDNO: 20010, SEQ ID NO: 20001, SEQ ID NO: 19642, OR SEQ ID NO: 19963. In some embodiments, the mutated msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49,50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61 , 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 20010. In some embodiments, the mutated msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 20001. In some embodiments, the mutated msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21,22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47,48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73,74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99,100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 110 nucleotides as compared to SEQ ID NO: 19642. In some embodiments, the mutated msd locus comprises a deletion of 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84 ,85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100, 101, 102, 103, 104, 105, 106, 107, 108, 109, or 1 10 nucleotides as compared to SEQ ID NO: 19963.
[0185] As provided for herein, a donor DNA sequence is inserted into the msr / msd locus and when the locus is reverse transcribed the donor DNA sequence is also reverse transcribed. The location of the insert may be anywhere downstream of the where the RT binds to. By downstream of where the RT binds to means that that the donor DNA insert is 3’ to the stem loop like structure that the RT binds to. For example, the donor DNA insert may be inserted as shown in FIG. 13. The remaining portion of the msr / msd locus may be truncated by deletions or mutations to shorten the overall msr / msd locus, thereby allowing for the donor DNA insert to be, for example longer in length (more nucleotides).
[0186] In some embodiments, the donor DNA may be inserted at any position within the ncRNA locus. In some embodiments, the donor DNA may be inserted at any position between the 5’ and 3’ inverted repeat regions of the ncRNA locus. In some embodiments, thedonor DNA may be inserted at any position within the msr / msd locus. In some embodiments, the donor DNA may be inserted at any position within the msd locus. In some embodiments, the donor DNA may be inserted within the longest inverted repeat region of the ncRNA that overlaps the msd start and msd end. In some embodiments, the donor DNA may be inserted within, or near the msd region, but not overlapping the ncRNA inverted repeats ends or the msr region of the ncRNA. In some embodiments, the donor DNA may be inserted at any position within the msd locus of Table X. In some embodiments, a donor DNA inserted at any position within the msd locus of Table X, is the donor DNA inserted at any position between the “msd_start” and “msd end” positions of Table X. In some embodiments, the donor DNA may be inserted at any position within the msd locus of Table Y. In some embodiments, a donor DNA inserted at any position within the msd locus of Table Y, is the donor DNA inserted at any position between the “msd start” and “msd end” positions of Table Y. In some embodiments, the donor DNA may be inserted at any position within the msd locus of any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587, 19589,19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652,19656, 19658, 19660, 19662, 19666, 19669, 19671, 19672, 19674, 19675, 19678, 19680, 19682,19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722,19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951,19957, 19959, 19974, 19978, 19979, 19980, 19983, 19986, 19998, 19999, 20000, 20005, 20016,20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20044,20046, 20049, 20050, 20052, 20053, 20055, 20056, 20061, 20062, 20070, 20072, 20074, 20075,20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141,20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178,20179, 20184, 20185, 20196, 20201, 20202, 20204, 20207, 20209, 20211, 20213, 20214, 20218,20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20248, 20249, 20250,20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20276, 20277, 20279, 20280, 20281,20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308,20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20340,20342, 20345, 20347, 20358, 20362, 20365, 20367, 20369, 20371, 20373, or 20377. In some embodiments, the donor DNA may be inserted at any position within the nucleic acid sequence of any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607,19608, 19611 , 19612, 19613, 19614, 19616, 19617, 19621, 19622, 19623, 19624, 19625, 19626,19627, 19628, 19629, 19630, 19631, 19632, 19633, 19636, 19637, 19638, 19639, 19640, 19641,19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649, 19653, 19654, 19664, 19665, 19667,19676, 19677, 19690, 19691, 19692, 19695, 19697, 19699, 19703, 19709, 19715, 19719, 19721,19755, 19855, 19927, 19940, 19941, 19947, 19948, 19949, 19953, 19954, 19955, 19960, 19961,19962, 19963, 19965, 19966, 19967, 19968, 19969, 19970, 19971, 19972, 19973, 19975, 19976,19977, 19981, 19984, 19985, 19988, 19989, 19990, 19991, 19992, 19993, 19994, 19995, 19996,20001, 20002, 20006, 20007, 20008, 20010, 20011, 20033, 20039, 20041, 20042, 20043, 20045,20047, 20057, 20058, 20059, 20060, 20063, 20065, 20067, 20068, 20071, 20076, 20081, 20082,20083, 20084, 20085, 20086, 20087, 20088, 20089, 20090, 20093, 20094, 20095, 20096, 20097,20100, 20101, 20102, 20105, 20106, 20109, 20110, 20112, 20118, 20121, 20122, 20124, 20125,20126, 20129, 20130, 20133, 20134, 20143, 20144, 20145, 20146, 20147, 20153, 20154, 20155,20156, 20159, 20180, 20181, 20182, 20186, 20187, 20188, 20189, 20190, 20191, 20192, 20193,20194, 20195, 20197, 20198, 20199, 20200, 20219, 20220, 20221, 20222, 20232, 20233, 20245,20253, 20255, 20265, 20267, 20268, 20275, 20284, 20315, 20321, 20324, 20329, 20333, 20346,20359, 20360, 20363, or 20374.
[0187] In some embodiments, the donor DNA may be inserted at any position between the “insert start” and “insert end” positions of Table X. In some embodiments, the donor DNA may be inserted at any position within 10-20 nucleic acids of the “insert_start” and “insert end” positions of Table X. In some embodiments, the donor DNA may be inserted at any position beyond the “insert start” and “insert end” positions of Table X. In some embodiments, the donor DNA may be inserted at any position within 10-20 nucleic acids that is beyond the “insert start” and “insert end” positions of Table X. In some embodiments, the donor DNA may be inserted at any position between the “insert_start” and “insert_end” positions of Table Y. As used herein, “insert start” denotes a nucleic acid within the corresponding msr / msd sequence comprising the start nucleic acid where the donor DNA sequence may be inserted. In some embodiments, the donor DNA may be inserted at any position between the “optimized insert start” and “optimized insert end” positions of Table X. In some embodiments, the donor DNA may be inserted at any position within 10-20 nucleic acids of the “optimized insert start” and “optimized insert end” positions of Table X. In some embodiments, the donor DNA may be inserted at any position beyond the “optimized insert start” and“optimized insert end” positions of Table X. In some embodiments, the donor DNA may be inserted at any position within 10-20 nucleic acids that is beyond the “optimized_insert_start” and “optimized insert end” positions of Table X. In some embodiments, the donor DNA may be inserted at any position between the “optimized insert start” and “optimized insert end” positions of Table Y. As used herein, “optimized_insert_start” denotes a nucleic acid within the corresponding msr / msd sequence comprising the start nucleic acid where the donor DNA sequence may be inserted. As used herein, “optimized insert end” denotes a nucleic acid within the corresponding msr / msd sequence comprising the end nucleic acid where the donor DNA sequence may be inserted. Accordingly, in some embodiments, and in reference to SEQ ID NO: 19610, “insert_start” nucleic acid 90 and “insert_end” nucleic acid 176 means that the donor DNA may be inserted between nucleic acid 90 and 176 of SEQ ID NO: 19610, as showed in bold and underlined below,ACATAATTGCGCAGCTCAGTTACGCGACTTGTGGTTGTTGGTAAAACGACATGCCATCG TCGGGCGCATCGCCAGACGGAAAATGATTGAACGACATTGGCGGCACGAAAACGGCCAC GCGCCCAATCTGCTCCTTTGTCGCATCTTGGGCGCGTGTCCATTTTCGTGCCTGTAATT CAAGGTAGCTGAGCCTACCTATGACGT ( SEQ ID NO : 19610 ) .
[0188] The following non-limiting examples illustrate, in bold and underlined, locations within the msr / msd locus where the donor DNA sequence may be inserted:ACATAATTGCGCAGCTCAGTTACGCGACTTGTGGTTGTTGGTAAAACGACATGCCATCG TCGGGCGCATCGCCAGACGGAAAATGATTGAACGACATTGGCGGCACGAAAACGGCCAC GCGCCCAATCTGCTCCTTTGTCGCATCTTGGGCGCGTGTCCATTTTCGTGCCTGTAATT CAAGGTAGCTGAGCCTACCTATGACGT (SEQ ID NO: 19610); CCAGCAGTGGCAATAGCGTTTCCGGCCTTTTGTGCCGGGAGGGTCGGCGAGTCGCCGAC T T AACGC CAGTAGT T TGTCCAGATAC TCAAAGTCGC TCCAT TGTAC T TAAGTACGCT TC GCGTACGTCGCGCTGACGCGCTCAGTACAGTTACGCGCCTTCGGGATAGTTTGAGGGTA TTGCCGCTGTTGG (SEQ ID NO: 19979); ATTCATGTAATCTCTATATGTCCTTTAGCGTTTAGACGTTTACGTCTTGTCGGGCGTTT CGCCAGACACGAAGTTATTGGAAGGTTTATGAGGTTGCGGTAGGTATAATCCTCCGCCG C TCATTGCC T TGCAT TTCGCGGCGGAGGAT TATACCCACCGCATCC T TAAGTAAAGGACATAGAAGAACTGGCATTAAT (SEQ ID NO: 19931);ATCGACTAACTCAGTTACGCGCATAGTGGTTGTTACTTAGGAAACATACCATCGTCGGG CGTATCGC C AGAC G G GAAAT AAT T GATATACATCGGCGGCACGAAACAGGCCACGCGCC CAACTCGCTCCGTCGTCACTCCTTGGGCGCGTGGCCTGTTTCGTGCCTGTAATTTATGG TAGCTGAGTCTGCCTAT (SEQ ID NO: 19722); or GCTCAGTTACGCGACTTGTGGTTGTTGGTAAAACGACATGCCATCGTCGGGCGCATCGC CAGACGGAAAAT GAT TGAACGACATTGGCGGCACGAAAACGGCCACGCGCCCAATCTGC TCCTTTGTCGCATCTTGGGCGCGTGTCCATTTTCGTGCCTGTAATTCAAGG T AG C T GAG C (SEQ ID NO: 24020).
[0189] In some embodiments, the donor DNA may replace a nucleic acid, a nucleic acid sequence, or a fragment thereof, at any position between the “insert start” and “insert end” positions of Table X. In some embodiments, the donor DNA may replace a nucleic acid, a nucleic acid sequence, or a fragment thereof, at any position between the “insert_start” and “insert_end” positions of Table Y. In some embodiments, the donor DNA may replace a nucleic acid, a nucleic acid sequence, or a fragment thereof, at any position between the “optimized insert start” and “optimized_insert_end” positions of Table X. In some embodiments, the donor DNA may replace a nucleic acid, a nucleic acid sequence, or a fragment thereof, at any position between the “optimized insert start” and “optimized insert end” positions of Table Y. In some embodiments, the donor DNA sequence may replace a nucleic acid, a nucleic acid sequence, or a fragment thereof, of the msd locus of any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570,19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596,19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660,19662, 19666, 19669, 19671, 19672, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694,19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930,19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19974,19978, 19979, 19980, 19983, 19986, 19998, 19999, 20000, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20044, 20046, 20049, 20050,20052, 20053, 20055, 20056, 20061, 20062, 20070, 20072, 20074, 20075, 20078, 20092, 20099,20103, 20107, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152,20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185,20196, 20201, 20202, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229,20231, 20235, 20237, 20238, 20242, 20243, 20247, 20248, 20249, 20250, 20251, 20252, 20259,20260, 20262, 20271, 20272, 20274, 20276, 20277, 20279, 20280, 20281, 20286, 20287, 20289,20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318,20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20340, 20342, 20345, 20347,20358, 20362, 20365, 20367, 20369, 20371, 20373, or 20377. In some embodiments, the donor DNA sequence may replace a nucleic acid, a nucleic acid sequence, or a fragment thereof, selected from any one of SEQ ID NO s: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616, 19617, 19621, 19622, 19623, 19624, 19625, 19626,19627, 19628, 19629, 19630, 19631, 19632, 19633, 19636, 19637, 19638, 19639, 19640, 19641,19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649, 19653, 19654, 19664, 19665, 19667,19676, 19677, 19690, 19691, 19692, 19695, 19697, 19699, 19703, 19709, 19715, 19719, 19721,19755, 19855, 19927, 19940, 19941, 19947, 19948, 19949, 19953, 19954, 19955, 19960, 19961,19962, 19963, 19965, 19966, 19967, 19968, 19969, 19970, 19971, 19972, 19973, 19975, 19976,19977, 19981, 19984, 19985, 19988, 19989, 19990, 19991, 19992, 19993, 19994, 19995, 19996,20001, 20002, 20006, 20007, 20008, 20010, 20011, 20033, 20039, 20041, 20042, 20043, 20045,20047, 20057, 20058, 20059, 20060, 20063, 20065, 20067, 20068, 20071, 20076, 20081, 20082,20083, 20084, 20085, 20086, 20087, 20088, 20089, 20090, 20093, 20094, 20095, 20096, 20097,20100, 20101, 20102, 20105, 20106, 20109, 20110, 20112, 20118, 20121, 20122, 20124, 20125,20126, 20129, 20130, 20133, 20134, 20143, 20144, 20145, 20146, 20147, 20153, 20154, 20155,20156, 20159, 20180, 20181 , 20182, 20186, 20187, 20188, 20189, 20190, 20191, 20192, 20193,20194, 20195, 20197, 20198, 20199, 20200, 20219, 20220, 20221, 20222, 20232, 20233, 20245,20253, 20255, 20265, 20267, 20268, 20275, 20284, 20315, 20321, 20324, 20329, 20333, 20346,20359, 20360, 20363, or 20374.
[0190] In some embodiments, the donor DNA may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, atleast 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleic acids at any position between the “insert_start” and “insert_end” positions of Table X. In some embodiments, the donor DNA may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101 , at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleic acids at any position between the “optimized insert start” and “optimized insert end” positions of Table X. In some embodiments, the donor DNA may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81 , at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleic acids at any position between the “insert_start” and “insert end” positions of Table Y. In some embodiments, the donor DNA may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 1 10 nucleic acids at any position between the “optimized_insert_start” and “optimized insert end” positions of Table Y. In some embodiments, the donor DNA may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91 , at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleic acids at any position between the “insert_start” and “insert_end” positions of any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578,19579, 19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604,19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671,19672, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706,19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939,19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983,19986, 19998, 19999, 20000, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030,20032, 20034, 20036, 20038, 20040, 20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056,20061, 20062, 20070, 20072, 20074, 20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117,20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165,20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204,20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238,20242, 20243, 20247, 20248, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272,20274, 20276, 20277, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296,20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328,20331, 20332, 20335, 20338, 20339, 20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367,20369, 20371, 20373, or 20377 of Table Y. In some embodiments, the donor DNA may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, atleast 86, at least 87, at least 88, at least 89, at least 90, at least 91 , at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, at least 110, or more nucleic acids at any position between the “optimized insert start” and “optimized_insert_end” positions of any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593, 19598, 19601,19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666,19671, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706,19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939,19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20046, 20049,20050, 20052, 20053, 20055, 20056, 20070, 20074, 20075, 20078, 20092, 20099, 20103, 20108,20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163,20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20204, 20207,20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242,20243, 20247, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20279,20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304,20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338,20339, 20342, 20358, 20362, 20365, 20367, 20369, 20371, 20373, 20377, 20379, 20380, 20381,20382, 20383, 20384, 20385, 20386, 20386, 20387, 20388, 20389, 20390, 20391, 20392, 20393,20394, 20395, 20396, 20397, 20398, 20399, 20400, 20401, 20402, 20403, 20404, or 20405 of Table Y. In some embodiments, the donor DNA may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleic acids at any position between the “insert_start” and “insert_end” positions of any one of SEQ ID NOs: 19561, 19564, 19567, 19580, 19581, 19590, 19595, 19606, 19607, 19608, 19611, 19612, 19613, 19614, 19616,19617, 19621, 19622, 19623, 19624, 19625, 19626, 19627, 19628, 19629, 19630, 19631, 19632,19633, 19636, 19637, 19638, 19639, 19640, 19641, 19642, 19643, 19644, 19645, 19646, 19647, 19648, 19649, 19653, 19654, 19664, 19665, 19667, 19676, 19677, 19690, 19691, 19692, 19695,19697, 19699, 19703, 19709, 19715, 19719, 19721, 19755, 19855, 19927, 19940, 19941, 19947,19948, 19949, 19953, 19954, 19955, 19960, 19961, 19962, 19963, 19965, 19966, 19967, 19968,19969, 19970, 19971, 19972, 19973, 19975, 19976, 19977, 19981, 19984, 19985, 19988, 19989,19990, 19991, 19992, 19993, 19994, 19995, 19996, 20001, 20002, 20006, 20007, 20008, 20010,20011, 20033, 20039, 20041, 20042, 20043, 20045, 20047, 20057, 20058, 20059, 20060, 20063,20065, 20067, 20068, 20071, 20076, 20081, 20082, 20083, 20084, 20085, 20086, 20087, 20088,20089, 20090, 20093, 20094, 20095, 20096, 20097, 20100, 20101, 20102, 20105, 20106, 20109,20110, 20112, 20118, 20121, 20122, 20124, 20125, 20126, 20129, 20130, 20133, 20134, 20143,20144, 20145, 20146, 20147, 20153, 20154, 20155, 20156, 20159, 20180, 20181, 20182, 20186,20187, 20188, 20189, 20190, 20191, 20192, 20193, 20194, 20195, 20197, 20198, 20199, 20200,20219, 20220, 20221 , 20222, 20232, 20233, 20245, 20253, 20255, 20265, 20267, 20268, 20275,20284, 20315, 20321, 20324, 20329, 20333, 20346, 20359, 20360, 20363, or 20374 of Table Y.
[0191] In some embodiments, the donor DNA sequence may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least78, at least 79, at least 80, at least 81 , at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, or at least 93 nucleic acids from SEQ ID NO: 19610. In some embodiments, the donor DNA sequence may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, or at least 93 nucleic acids from SEQ ID NO: 20380. In some embodiments, the donor DNA sequence may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41 , at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, or at least 108 nucleic acids from SEQ ID NO: 19979. In some embodiments, the donor DNA sequence may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, atleast 20, at least 21 , at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, or at least 91 nucleic acids from SEQ ID NO: 19931. In some embodiments, the donor DNA sequence may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81 , at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, or at least 89 nucleic acids from SEQ ID NO: 19722.
[0192] In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579,19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605,19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671, 19672,19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708,19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943,19944, 19945, 19946, 19950, 19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983, 19986,19998, 19999, 20000, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032,20034, 20036, 20038, 20040, 20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20061,20062, 20070, 20072, 20074, 20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128,20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169,20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204, 20207,20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242,20243, 20247, 20248, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274,20276, 20277, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298,20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331,20332, 20335, 20338, 20339, 20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367, 20369,20371, 20373, or 20377; and a donor DNA sequence inserted within the msr / msd locus. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587, 19589,19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652,19656, 19658, 19660, 19662, 19666, 19669, 19671, 19672, 19674, 19675, 19678, 19680, 19682,19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722,19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951 ,19957, 19959, 19974, 19978, 19979, 19980, 19983, 19986, 19998, 19999, 20000, 20005, 20016,20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20044,20046, 20049, 20050, 20052, 20053, 20055, 20056, 20061, 20062, 20070, 20072, 20074, 20075,20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141,20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178,20179, 20184, 20185, 20196, 20201, 20202, 20204, 20207, 20209, 20211, 20213, 20214, 20218,20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20248, 20249, 20250,20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20276, 20277, 20279, 20280, 20281,20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308,20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20340,20342, 20345, 20347, 20358, 20362, 20365, 20367, 20369, 20371, 20373, or 20377; and a donorDNA sequence inserted within the “insert_start_ and “insert_end” nucleic acids of Table Y. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19579, 19582, 19583, 19585, 19587,19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651,19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671, 19672, 19674, 19675, 19678, 19680,19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717,19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950,19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983, 19986, 19998, 19999, 20000, 20005,20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040,20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20061, 20062, 20070, 20072, 20074,20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117, 20128, 20132, 20136, 20138, 20140,20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176,20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204, 20207, 20209, 20211, 20213, 20214,20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20248, 20249,20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274, 20276, 20277, 20279, 20280,20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306,20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339,20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367, 20369, 20371, 20373, or 20377; and a donor DNA sequence inserted within the “insert start and “insert end” nucleic acids of any one of SEQ ID NOs: 19553, 19554, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578,19579, 19582, 19583, 19585, 19587, 19589, 19591, 19593, 19596, 19598, 19601, 19602, 19604,19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662, 19666, 19669, 19671,19672, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704, 19706,19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937, 19939,19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19974, 19978, 19979, 19980, 19983,19986, 19998, 19999, 20000, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030,20032, 20034, 20036, 20038, 20040, 20044, 20046, 20049, 20050, 20052, 20053, 20055, 20056,20061, 20062, 20070, 20072, 20074, 20075, 20078, 20092, 20099, 20103, 20107, 20108, 20117,20128, 20132, 20136, 20138, 20140, 20141 , 20142, 20151, 20152, 20158, 20161, 20163, 20165,20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20201, 20202, 20204,20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238,20242, 20243, 20247, 20248, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272,20274, 20276, 20277, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296,20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328,20331, 20332, 20335, 20338, 20339, 20340, 20342, 20345, 20347, 20358, 20362, 20365, 20367,20369, 20371, 20373, or 20377 of Table Y.
[0193] In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582,19585, 19587, 19589, 19591, 19593, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634,19651, 19652, 19656, 19658, 19660, 19662, 19666, 19671, 19674, 19675, 19678, 19680, 19682,19685, 19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722,19785, 19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951,19957, 19959, 19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030,20032, 20034, 20036, 20038, 20040, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20070,20074, 20075, 20078, 20092, 20099, 20103, 20108, 201 17, 20128, 20132, 20136, 20138, 20140,20141, 20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176,20178, 20179, 20184, 20185, 20196, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223,20227, 20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20249, 20250, 20251, 20252,20259, 20260, 20262, 20271, 20272, 20274, 20279, 20280, 20281, 20286, 20287, 20289, 20290,20291, 20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319,20320, 20323, 20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358, 20362, 20365, 20367,20369, 20371, 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386, 20386,20387, 20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398, 20399,20400, 20401, 20402, 20403, 20404, or 20405; and a donor DNA sequence inserted within the msr / msd locus. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593, 19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651,19652, 19656, 19658, 19660, 19662, 19666, 19671, 19674, 19675, 19678, 19680, 19682, 19685,19689, 19694, 19698, 19702, 19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785,19834, 19930, 19931, 19933, 19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957,19959, 19983, 19998, 20005, 20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032,20034, 20036, 20038, 20040, 20046, 20049, 20050, 20052, 20053, 20055, 20056, 20070, 20074,20075, 20078, 20092, 20099, 20103, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141,20142, 20151, 20152, 20158, 20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178,20179, 20184, 20185, 20196, 20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227,20229, 20231, 20235, 20237, 20238, 20242, 20243, 20247, 20249, 20250, 20251, 20252, 20259,20260, 20262, 20271, 20272, 20274, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291,20295, 20296, 20298, 20300, 20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320,20323, 20328, 20331, 20332, 20335, 20338, 20339, 20342, 20358, 20362, 20365, 20367, 20369,20371, 20373, 20377, 20379, 20380, 20381, 20382, 20383, 20384, 20385, 20386, 20386, 20387,20388, 20389, 20390, 20391, 20392, 20393, 20394, 20395, 20396, 20397, 20398, 20399, 20400,20401, 20402, 20403, 20404, or 20405; and a donor DNA sequence inserted within the “optimized insert start” and “optimized insert end” nucleic acids of Table Y. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593,19598, 19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660,19662, 19666, 19671, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702,19704, 19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933,19937, 19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19983, 19998, 20005,20016, 20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040,20046, 20049, 20050, 20052, 20053, 20055, 20056, 20070, 20074, 20075, 20078, 20092, 20099,20103, 20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158,20161, 20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196,20204, 20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237,20238, 20242, 20243, 20247, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272,20274, 20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300,20302, 20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332,20335, 20338, 20339, 20342, 20358, 20362, 20365, 20367, 20369, 20371, 20373, 20377, 20379,20380, 20381, 20382, 20383, 20384, 20385, 20386, 20386, 20387, 20388, 20389, 20390, 20391,20392, 20393, 20394, 20395, 20396, 20397, 20398, 20399, 20400, 20401, 20402, 20403, 20404, or 20405; and a donor DNA sequence inserted within the “optimized_insert_start” and “optimized insert end” nucleic acids of any one of SEQ ID NOs: 19553, 19558, 19563, 19566, 19570, 19572, 19574, 19576, 19578, 19582, 19582, 19585, 19587, 19589, 19591, 19593, 19598,19601, 19602, 19604, 19605, 19609, 19610, 19634, 19651, 19652, 19656, 19658, 19660, 19662,19666, 19671, 19674, 19675, 19678, 19680, 19682, 19685, 19689, 19694, 19698, 19702, 19704,19706, 19708, 19710, 19711, 19714, 19717, 19722, 19785, 19834, 19930, 19931, 19933, 19937,19939, 19943, 19944, 19945, 19946, 19950, 19951, 19957, 19959, 19983, 19998, 20005, 20016,20018, 20020, 20023, 20025, 20028, 20029, 20030, 20032, 20034, 20036, 20038, 20040, 20046,20049, 20050, 20052, 20053, 20055, 20056, 20070, 20074, 20075, 20078, 20092, 20099, 20103,20108, 20117, 20128, 20132, 20136, 20138, 20140, 20141, 20142, 20151, 20152, 20158, 20161,20163, 20165, 20169, 20171, 20172, 20174, 20176, 20178, 20179, 20184, 20185, 20196, 20204,20207, 20209, 20211, 20213, 20214, 20218, 20223, 20227, 20229, 20231, 20235, 20237, 20238,20242, 20243, 20247, 20249, 20250, 20251, 20252, 20259, 20260, 20262, 20271, 20272, 20274,20279, 20280, 20281, 20286, 20287, 20289, 20290, 20291, 20295, 20296, 20298, 20300, 20302,20304, 20306, 20308, 20311, 20313, 20318, 20319, 20320, 20323, 20328, 20331, 20332, 20335,20338, 20339, 20342, 20358, 20362, 20365, 20367, 20369, 20371, 20373, 20377, 20379, 20380,20381, 20382, 20383, 20384, 20385, 20386, 20386, 20387, 20388, 20389, 20390, 20391, 20392,20393, 20394, 20395, 20396, 20397, 20398, 20399, 20400, 20401, 20402, 20403, 20404, or 20405 of Table Y.
[0194] In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, atleast 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19610; and a donor DNA sequence inserted within the msr / msd locus, provided that the donor DNA sequence is inserted at any position within the nucleic acids 90 and 176 of SEQ ID NO: 19610. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19979; and a donor DNA sequence inserted within the msr / msd locus, provided that the donor DNA sequence is inserted at any position within the nucleic acids 21 and 38 of SEQ ID NO: 19979. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19931; and a donor DNA sequence inserted within the msr / msd locus, provided that the donor DNA sequence is inserted at any position within the nucleic acids 90 and 165 of SEQ ID NO: 19931. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19722; and a donor DNA sequence inserted within the msr / msd locus, provided that the donor DNA sequence is inserted at any position within the nucleic acids 27 and 50 of SEQ ID NO: 19722.
[0195] In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19610; and a donor DNA sequence inserted within the msr / msd locus, provided that the donor DNA sequence is inserted at any position within the nucleic acid sequence SEQ ID SEQ ID NO: 20010. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to SEQ ID NO: 19722; and a donorDNA sequence inserted within the msr / msd locus, provided that the donor DNA sequence is inserted at any position within the nucleic acid sequence SEQ ID SEQ ID NO: 19963. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%, at least 55%>, at least 60%, at least 65%, at least 70%, at least 75%>, at least 80%>, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%o, or at least 99% sequence identity to SEQ ID NO: 19931; and a donor DNA sequence inserted within the msr / msd locus, provided that the donor DNA sequence is inserted at any position within the nucleic acid sequence SEQ ID SEQ ID NO: 19642. In some embodiments, the msr / msd locus comprises the nucleic acid sequence having at least 50%>, at least 55%, at least 60%>, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%o, at least 93%>, at least 94%o, at least 95%>, at least 96%>, at least 97%>, at least 98%>, or at least 99%o sequence identity to SEQ ID NO: 19979; and a donor DNA sequence inserted within the msr / msd locus, provided that the donor DNA sequence is inserted at any position within the nucleic acid sequence SEQ ID SEQ ID NO: 20001.
[0196] As used herein, a “fragment” can refer to a sequence (polynucleotide or amino acid) that has 1, 2, 3, 4, 5, 6, 7, 8, 9 or 10 fewer nucleotides or residues, which may be referred to as a deletion. The nucleotides may be removed from either the 5’ end (e.g., 5’ deletion) or the 3’ end (e.g., 3’ deletion) of the recited sequence. In some embodiments, the nucleotides may be removed internally.
[0197] In non-limiting examples, the donor DNA sequences, denoted “XI” may be inserted at any position as shown in molecules M5 or M10 below:ACATAATTGCGCAGCTCAGTTACGCGACTTGTGGTTGTTGGTAAAACGACATGCCATCG TCGGGCGCATCGCCAGACGGAAAATGATTGAX1TCAAGGTAGCTGAGCCTACCTATGAC GT (M5; SEQ ID NO: 24032 and SEQ ID NO: 24033); or ACATAATTGCGCAGCTCAGTTACGCGACTTGTGGTTGTTGGTAAAACGACATGCCATCG TCGGGCGCATCGCCAGACGGAAAATGATTGAACGA.CX1GTAATTCAAGGTAGCTGAGCC TACCTATGACGT (M10, SEQ ID NO: 24034 and SEQ ID NO: 24035).
[0198] In some embodiments, the polynucleotide comprises the nucleic acid sequence of any one of SEQ ID NO: 24032 and SEQ ID NO: 24033; or SEQ ID NO: 24034 and SEQ ID NO: 24035. In some embodiments, the polynucleotide comprises the nucleic acidsequence of any one of SEQ ID NO: 24032 and SEQ ID NO: 24033; or SEQ ID NO: 24034 and SEQ ID NO: 24035, and further comp...
Claims
What is claimed is:
1. A polynucleotide-guide RNA cassette comprising: a) a polynucleotide comprising: i) a first inverted repeat sequence coding region; ii) a ncRNA comprising an msr / msd locus comprising the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one SEQ ID NO provided in Table X or Table Y; iii) a donor DNA sequence located within the msr / msd locus; and iv) a second inverted repeat sequence coding region; and b) a guide RNA (gRNA) coding region.
2. The cassette of claim 1, wherein the first inverted repeat sequence coding region is located within the 5' end of the ncRNA.
3. The cassette of claim 1, wherein the second inverted repeat sequence coding region is located within the 3 ' end of the ncRNA.
4. The cassette of any one of claims 1-3, wherein the ncRNA comprises the nucleic acid sequence comprising a deletion of at least 1 nucleotide as compared to any one ncRNA SEQ ID NO provided in Table X or Table Y.
5. The cassette of any one of claims 1-4, wherein the ncRNA comprises the nucleic acid sequence comprising a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one ncRNA SEQ ID NO provided in Table X or Table Y.
6. The cassette of any one of claims 1-5, wherein the ncRNA comprises the nucleic acid sequence of ncRNA SEQ ID NO provided in Table X or Table Y.
7. The cassette of claim 1, wherein the ncRNA comprises an msr locus and an msd locus.
8. The cassette of claim 7, wherein the msr locus comprises the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of msr SEQ ID NO provided in Table X or Table Y.
9. The cassette of any one of claims 7 or 8, wherein the msr locus comprises the nucleic acid sequence comprising a deletion of at least 1 nucleotide as compared to any one msr SEQ ID NO provided in Table X or Table Y.
10. The cassette of any one of claims 7-9, wherein the msr locus comprises the nucleic acid sequence comprising a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one msr SEQ ID NO provided in Table X or Table Y.
11. The cassette of any one of claims 7-10, wherein the msr locus comprises the nucleic acid sequence of SEQ ID NO provided in Table X or Table Y.
12. The cassette of claim 7, wherein the msd locus comprises a reverse complement nucleic acid sequence of the nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of msDNA SEQ ID NO provided in Table X or Table Y.
13. The cassette of any one of claims 7 or 12, wherein the msd locus comprises a reverse complement nucleic acid sequence of the nucleic acid sequence comprising a deletion of at least 1 nucleotide as compared to any one of msDNA SEQ ID NO provided in Table X or Table Y.
14. The cassette of any one of claims 7 or 12-13, wherein the msd locus comprises a reverse complement nucleic acid sequence of the nucleic acid sequence comprising a deletion of about 1 to 150, about 1 to about 125, about 1 to about 100, about 1 to about 90, about 1 to about 80, about 1 to about 70, about 1 to about 60, about 1 to about 50, about 1 to about 40, about 1 to about 30, about 1 to about 20, or about 1 to about 10 nucleotides as compared to any one of msDNA SEQ ID NO provided in Table X or Table Y.
15. The cassette of any one of claims 7 or 12-14, wherein the msd locus comprises a reverse complement nucleic acid sequence of the nucleic acid sequence of any one of msDNA SEQ ID NO provided in Table X or Table Y.
16. The cassette of claim 1, wherein the first inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NO: 20832, SEQ ID NO: 20833, SEQ ID NO: 20834, SEQ ID NO: 20835, or SEQ ID NO: 20836.
17. The cassette of claim 16, wherein the first inverted repeat sequence coding region comprises a nucleic acid sequence of any one of SEQ ID NO: 20832, SEQ ID NO: 20833, SEQ ID NO: 20834, SEQ ID NO: 20835, or SEQ ID NO: 20836.
18. The cassette of claim 1, wherein the second inverted repeat sequence coding region comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NO: 20387, SEQ ID NO: 20388, SEQ ID NO: 20389, SEQ ID NO: 20390, or SEQ ID NO: 20391.
19. The cassette of claim 18, wherein the second inverted repeat sequence coding region comprises a nucleic acid sequence of any one of SEQ ID NO: 20387, SEQ ID NO: 20388, SEQ ID NO: 20389, SEQ ID NO: 20390, or SEQ ID NO: 20391.
20. The cassette of claim 1, wherein the donor DNA sequence is inserted within the ncRNA of any one of ncRNA SEQ ID NO provided in Table X or Table Y.
21. The cassette of claim 1, wherein the donor DNA sequence is inserted within the optimized insert start and optimized insert end nucleic acids of Table X, or Table Y.
22. The cassette of claim 1, wherein the donor DNA sequence is inserted within the optimized insert start and optimized insert end nucleic acids of ncRNA SEQ ID NO provided in Table X or Table Y.
23. The cassette of claim 1, wherein the donor DNA sequence is inserted at any position within 5- 20 nucleic acids of the optimized insert start and optimized insert end nucleic acids of Table X, or Table Y.
24. The cassette of claim 1, wherein the donor DNA sequence is inserted at any position within 10-20 nucleic acids of the optimized insert start and optimized insert end nucleic acids of Table X, or Table Y.
25. The cassette of claim 1, wherein the donor DNA sequence is inserted at any position within 5- 20 nucleic acids of the optimized_insert_start and optimized_insert_end nucleic acids of ncRNA SEQ ID NO provided in Table X or Table Y.
26. The cassette of claim 1, wherein the donor DNA sequence is inserted at any position within 10-20 nucleic acids of the optimized insert start and optimized insert end nucleic acids of ncRNA SEQ ID NO provided in Table X or Table Y.
27. The cassette of claim 1 , wherein the donor DNA sequence may replace a nucleic acid sequence, or a fragment thereof, within the optimized insert start and optimized insert end nucleic acids of Table X, or Table Y.
28. The cassette of claim 1, wherein the donor DNA sequence may replace a nucleic acid sequence, or a fragment thereof, within the optimized insert start and optimized insert end nucleic acids of ncRNA SEQ ID NO provided in Table X or Table Y.
29. The cassette of claim 1, wherein the donor DNA sequence may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91 , at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleic acids from any one of SEQ ID NO provided in Table X or Table Y.
30. The cassette of claim 1, wherein the donor DNA sequence may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61 , at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71, at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleic acids within the optimized insert start and optimized insert end nucleic acids of Table X, or Table Y.
31. The cassette of claim 1, wherein the donor DNA sequence may result in a deletion of at least 1, at least 2, at least 3, at least 4, at least 5, at least 6, at least 7, at least 8, at least 9, at least 10, at least 11, at least 12, at least 13, at least 14, at least 15, at least 16, at least 17, at least 18, at least 19, at least 20, at least 21, at least 22, at least 23, at least 24, at least 25, at least 26, at least 27, at least 28, at least 29, at least 30, at least 31, at least 32, at least 33, at least 34, at least 35, at least 36, at least 37, at least 38, at least 39, at least 40, at least 41, at least 42, at least 43, at least 44, at least 45, at least 46, at least 47, at least 48, at least 49, at least 50, at least 51, at least 52, at least 53, at least 54, at least 55, at least 56, at least 57, at least 58, at least 59, at least 60, at least 61, at least 62, at least 63, at least 64, at least 65, at least 66, at least 67, at least 68, at least 69, at least 70, at least 71 , at least 72, at least 73, at least 74, at least 75, at least 76, at least 77, at least 78, at least 79, at least 80, at least 81, at least 82, at least 83, at least 84, at least 85, at least 86, at least 87, at least 88, at least 89, at least 90, at least 91, at least 92, at least 93, at least 94, at least 95, at least 96, at least 97, at least 98, at least 99, at least 100, at least 101, at least 102, at least 103, at least 104, at least 105, at least 106, at least 107, at least 108, at least 109, or at least 110 nucleic acids within the optimized insert start and optimized insert end nucleic acids of SEQ ID NO provided in Table X or Table Y.
32. The cassette of claim 1, wherein the polynucleotide encodes an RNA molecule that is capable of self-priming reverse transcription by a reverse transcriptase (RT).
33. The cassette of claim 32, wherein reverse transcription of the RNA molecule results in a multicopy single-stranded DNA (msDNA) molecule.
34. The cassette of claim 33, wherein the msDNA comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of msDNA SEQ ID NO of Table X or Table Y.
35. The cassette of claim 1, wherein transcription products of the polynucleotide and the gRNA coding region are physically coupled, or not physically coupled.
36. The cassette of claim 1, wherein the gRNA coding region is 3' of the polynucleotide.
37. The cassette of claim 1, wherein the gRNA coding region is 5' of the polynucleotide.
38. The cassette of claim 1, wherein the donor DNA sequence comprises two homology arms, wherein each homology arm has at least 70% to about 99% similarity to a portion of the sequence of a genetic locus of interest on either side of a nuclease cleavage site.
39. A vector comprising the polynucleotide-guide RNA cassette of any one of claims 1-38.
40. The vector of claim 39, further comprising a promoter that is operably linked to the cassette.
41. The vector of claim 40, wherein the promoter is selected from the group consisting of an RNA polymerase II promoter, an RNA polymerase III promoter, and a combination thereof.
42. The vector of claim 39, further comprising a reverse transcriptase (RT) coding sequence.
43. The vector of claim 42, wherein the RT coding sequence comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least80%, at least 85%, at least 90%, at least 91 %, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NO provided in Table X or Table Y.
44. The vector of claim 43, wherein the RT coding sequence comprises a nucleic acid sequence of any one of rt SEQ ID NO provided in Table X or Table Y.
45. The vector of claim 39, further comprising at least one, at least two, at least three, or more, nuclear localizing sequence (NLS).
46. The vector of claim 45, wherein the at least one, at least two, at least three, or more NLS is each independently selected from a group comprising a SV40 NLS sequence, a c-Myc NLS sequence, or a Nucleoplasmin NLS sequence.
47. The vector of claim 39, further comprising a nuclease coding sequence.
48. The vector of claim 47, wherein the nuclease and RT are physically coupled when expressed.
49. The vector of claims 47 or 48, wherein the nuclease is linked to the RT with a linker.
50. The vector of claim 49, wherein the linker is a peptide linker, such as a glycine-serine linker, or a glycine-alanine linker.
51. The vector of any one of claims 47-50, wherein the nuclease is positioned at the N-terminus of the RT or the C-terminus of the RT.
52. The vector of claim 47, wherein the nuclease and RT are not physically coupled when expressed.
53. The vector of any one of claims 47-52, wherein the nuclease is a Cas nuclease.
54. The vector of claim 53, wherein the Cas nuclease is a catalytically active nuclease, a nickase Cas nuclease, or a catalytically inactive Cas nuclease.
55. The vector of claim 54, wherein the nickase Cas nuclease is a Cas9D10A or a Cas9H840A nuclease, or a Cas having the both D10A and H840A mutations.
56. The vector of claim 55, wherein the Cas nuclease is selected from the group comprising Cas3, Cas9, Cpfl, Casl3d, Casl3a, Casl l, Casl2i2, Casl2b, Cas5, Cas8c, Cas7, a variant thereof, and any combination thereof.
57. The vector of claim 39, further comprising a reporter coding sequence.
58. The vector of claim 57, wherein the reporter is selected from the group comprising a luciferase, EYFP, ECFP, mRFPl, mOrange, GFP, GFPmut3b, sfGFP, mCherry, SYFP2, mTagBFP, YFP, mBanana, DsRed2FP, EGFP, mVenus, mTurquoise, Emerald, Azami Green, mWasabi, TagGFP, TurboGFP, AcGFP, ZsGreen, T-Sapphire, EBFP, EBFP2, Azurite, mECFP, Cerulean, CyPet, AmCyanl, Midori -Ishi Cyan, tagCFP, mTFPl, Topaz, Venus, mCitrine, YPet, tag YFP, phi YFP, zsYellowl, Kusabira Orange, Kusabira Orange2, m0range2, dTomato, dTomato-Tandem, tagRFP, tagRFP-T, dsRed, dsRed2, dsREd-Express, dsRed-monomer, inTangerine, mRuby, mApple, mStrawberry, asRed2, mRFPl, JRed, hcRedl, mRaspberry, dKeima-Tandem, hcRed- Tandem, mPlum, AQ143, and any combination thereof.
59. A vector comprising: a promoter; the polynucleotide-guide RNA cassette of any one of claims 1-38; a reverse transcriptase (RT) coding sequence; a nuclease coding sequence; and at least one nuclear localization signal sequence (NLS), wherein the promoter, the polynucleotide-guide RNA cassette, the RT, the nuclease, the reporter, and the at least one NLS are operably linked to each other.
60. The vector of claim 59, wherein the promoter is selected from the group consisting of an RNA polymerase II promoter, an RNA polymerase III promoter, and a combination thereof.
61. The vector of claim 59, wherein transcription products of the polynucleotide and the gRNA coding region are physically coupled.
62. The vector of claim 59, wherein transcription products of the polynucleotide and the gRNA coding region are not physically coupled.
63. The vector of claim 59, wherein the RT coding sequence comprises a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NO provided in Table X or Table Y.
64. The vector of any one of claims 59-63, wherein the RT coding sequence comprises a nucleic acid sequence of any one of SEQ ID NO provided in Table X or Table Y.
65. The vector of claim 59, wherein the at least one NLS comprises at least one, at least two, at least three, or more NLS.
66. The vector of claim 65, wherein the at least one, at least two, at least three, or more NLS is each independently selected from a group comprising a SV40 NLS sequence, a cMyc NLS sequence, or a Nuceloplasmin NLS sequence.
67. The vector of claim 59, wherein the nuclease and RT are physically coupled when expressed.
68. The vector of any one of claims 59-67, wherein the nuclease is linked to the RT with a linker.
69. The vector of claim 68, wherein the linker is a peptide linker, such as a glycine-serine linker, or a glycine-alanine linker.
70. The vector of any one of claims 67-69, wherein the nuclease is positioned at the N-terminus of the RT or the C-terminus of the RT.
71. The vector of claim 59, wherein the nuclease and RT are not physically coupled when expressed.
72. The vector of any one of claims 67-71, wherein the nuclease is a Cas nuclease.
73. The vector of claims 72, wherein the Cas nuclease is a catalytically active nuclease, a nickase Cas nuclease, or a catalytically inactive Cas nuclease.
74. The vector of claim 73, wherein the nickase Cas nuclease is a Cas9D10A or a Cas9H840A nuclease, or a Cas having the both D10A and H840A mutations.
75. The vector of claim 73, wherein the Cas nuclease is selected from the group comprising Cas3, Cas9, Cpfl, Casl3d, Casl3a, Casl l, Casl2i2, Casl2b, Cas5, Cas8c, Cas7, a variant thereof, and any combination thereof.
76. The vector of claim 59, optionally comprising a reporter coding sequence selected from the group comprising a luciferase, EYFP, ECFP, mRFPl, mOrange, GFP, GFPmut3b, sfGFP, mCherry, SYFP2, mTagBFP, YFP, mBanana, DsRed2FP, EGFP, mVenus, mTurquoise, Emerald, Azami Green, mWasabi, TagGFP, TurboGFP, AcGFP, ZsGreen, T-Sapphire, EBFP, EBFP2, Azurite, mECFP, Cerulean, CyPet, AmCyanl, Midori-Ishi Cyan, tagCFP, mTFPl, Topaz, Venus, mCitrine, YPet, tagYFP, phiYFP, zsYellowl, Kusabira Orange, Kusabira Orange2, mOrange2, dTomato, dTomato-Tandem, tagRFP, tagRFP-T, dsRed, dsRed2, dsREd-Express, dsRed- monomer, mTangerine, mRuby, mApple, mStrawberry, asRed2, mRFPl, JRed, hcRedl, mRaspberry, dKeima-Tandem, hcRed-Tandem, mPlum, AQ143, and any combination thereof.
77. A polynucleotide comprising a polynucleotide transcript of the polynucleotide-guide RNA cassette of any one of claims 1-38, or the vector of any one of claims 39-76, and optionally a gRNA, wherein the polynucleotide transcript and gRNA are physically coupled.
78. The polynucleotide of claim 77, wherein the polynucleotide transcript and gRNA are linked together with a linker.
79. The polynucleotide of claim 78, wherein the linker comprises a nucleotide sequence, such as nucleotide repeat sequence.
80. The polynucleotide of claim 79, wherein the linker comprises a sequence of (GAA)n, wherein n is 1-20, 1-15, 1-10, 1-5, 1-3, any integer the foregoing ranges, including the endpoints of the range (SEQ ID NO: 24043).
81. The polynucleotide of any one of claims 77-80, wherein the polynucleotide transcript and the gRNA comprises a structured motifs to the 5’ or 3’ end of the polynucleotide-gRNA transcript to stabilize the transcript, when the gRNA and the polynucleotide are part of the same transcript.
82. The polynucleotide of claim 81, wherein the structured motif is, independently, evoPreQ, mpKnot, or exoribonuclease-resistant RNA motifs, or as otherwise provided for herein.
83. The polynucleotide of any one of claims 81 or 82, wherein the structured motif is at the 5’ end of the polynucleotide-gRNA transcript.
84. The polynucleotide of any one of claims 81-83, wherein the structured motif is at the 3’ end of the polynucleotide-gRNA transcript, when the polynucleotide and gRNA are part of the same transcript.
85. The polynucleotide of any one of claims 81-84, wherein the structured motif is selected from a group comprising enoPreQ, mpKnot, MVExrRNA, WNVxRNA, Zika xRNA, Dengue xRAN, YF, any variant thereof, and any combination thereof.
86. The polynucleotide of any one of claims 81-85, wherein the polynucleotide transcriptguide RNA (gRNA) molecule comprises a molecule of the formula Al -Rt-Ll-gRNA-A2, wherein:Al is absent or a structured motif to stabilize the transcript;Rt is the polynucleotide transcript;LI is absent or a nucleotide linker from about 5 to about 40, about 8 to about 35, about 9, or about 33 basepairs; gRNA is the guide RNA; andA2 is absent or a structured motif to stabilize the transcript.
87. The polynucleotide of claim 86, wherein:Al is a structured motif to stabilize the transcript;Rt is the polynucleotide transcript;LI is a nucleotide linker from about 5 to about 40, about 8 to about 35, about 9, or about 33 basepairs; gRNA is the guide RNA; andA2 is absent.
88. The polynucleotide of claim 86, wherein:Al is absent;Rt is the polynucleotide transcript;LI is a nucleotide linker from about 5 to about 40, about 8 to about 35, about 9, or about 33 base pairs; gRNA is the guide RNA; andA2 is a structured motif to stabilize the transcript.
89. The polynucleotide of any one of claims 77-88, wherein the ncRNA comprises one or more modifications at a 3’ or a 5’ end.
90. The polynucleotide of claim 89, wherein the one or more modifications comprise a 3 ’cap, a polyA tail, replacing a phosphate linkage of a ribose with alternative chemical groups (e.g., 2’- Ome, 2’-F, 2’-ME0, and the like), replacing at least one nucleotide with a deoxyribose derivative.
91. The polynucleotide of any one of claims 77-90, wherein the ncRNA comprises one or more modifications anywhere within the ncRNA.
92. The polynucleotide of claim 91, wherein the one or more modifications comprise a modified or a chemically modified nucleotide.
93. The polynucleotide of claim 92, wherein the modified or the chemically modified nucleotide is selected from the group consisting of deoxyribose, locked nucleic acid nucleotide (LNA), hexitol nucleic acid, unsubstituted ribose, ribose substituted with halo, alkyl, cycloalkyl, phenyl, alkoxy, amino, or nitro substituent that may be unsubstituted, or further substituted with one or more alkyl, halo, haloalkyl, amino, or nitro substituents, 5-hydroxycytidines, 5- alkylcytidines, 5-hydroxyalkylcytidines, 5-carboxycytidines, 5-formylcytidines, 5- alkoxycytidines, 5-alkynylcytidines, 5-halocytidines, 2-thiocytidines, N4- alkylcytidines, N4- aminocytidines, N4-acetylcytidines, and N4,N4-dialkylcytidines. Examples of modified or chemically modified nucleotides or nucleosides include 5- hydroxycytidine, 5-methylcytidine, 5- hydroxymethylcytidine, 5-carboxycytidine, 5- formylcytidine, 5-methoxycytidine, 5- propynylcytidine, 5-bromocytidine, 5-iodocytidine, 2- thiocytidine; N4-methylcytidine, N4- aminocytidine, N4-acetylcytidine, and N4,N4- dimethylcytidine. Examples of modified or chemically modified nucleotides or nucleosides include 5- hydroxyuridines, 5-alkyluridines, 5- hydroxyalkyluridines, 5-carboxyuridines, 5- carboxyalkylesteruridines, 5-formyluridines, 5- alkoxyuridines, 5-alkynyluridines, 5- halouridines, 2-thiouridines, and 6-alkyluridines. Examples of modified or chemically modified nucleotides or nucleosides include 5-hydroxyuridine, 5- methyluridine, 5-hydroxymethyluridine, 5-carboxyuridine, 5-carboxymethylesteruridine, 5- formyluridine, 5 -methoxyuridine, 5-propynyluridine, 5-bromouridine, 5-fluorouridine, 5- iodouridine, 2-thiouridine, 6-methyluridine, 5-methoxycarbonylmethyl-2-thiouridine, 5- methylaminomethyl-2 -thiouridine, 5-carbamoylmethyluridine, 5-carbamoylmethyl-2’-O- methyluridine, l-methyl-3-(3-amino-3-carboxypropy)pseudouridine, 5-methylaminomethyl-2-selenouridine, 5- carb oxym ethyl uridine, 5-methyldihydrouridine, 5-taurinomethyluridine, 5- taurinomethyl-2- thiouridine, 5-(isopentenylaminomethyl)uridine, 2’-O-methylpseudouridine, 2- thio-2’0- methyluridine, and 3,2’ -O-dimethyluri dine, N6- methyladenosine, 2-aminoadenosine, 3 -methyl adenosine, 8-azaadenosine, 7-deazaadenosine, 8-oxoadenosine, 8-bromoadenosine, 2- methylthio-N6-methyladenosine, N6-isopentenyladenosine, 2-methylthio-N6- isopentenyladenosine, N6-(cis-hydroxyisopentenyl)adenosine, 2-methy Ithi o-N6-(ci s- hydroxyisopentenyl)adenosine, N6-glycinylcarbamoyladenosine, N 6-threony 1 carb amoy 1 - adenosine, N6-methyl-N6- threonylcarbamoyl-adenosine, 2-methylthio-N6-threonylcarbamoyl- adenosine, N6,N6- dimethyladenosine, N6-hydroxynorvalylcarbamoyladenosine, 2-methylthio- N6- hydroxynorvalylcarbamoyl-adenosine, N6-acetyl-adenosine, 7-methyl-adenine, 2- methylthio- adenine, 2-methoxy-adenine, alpha-thio-adenosine, 2'-O-methyl-adenosine, N6,2'-O- dimethyl- adenosine, N6,N6,2'-O-trimethyl-adenosine, l,2'-O-dimethyl-adenosine, 2'-O- ribosyladenosine, 2-amino-N6-methyl-purine, 1 -thio-adenosine, 2'-F-ara-adenosine, 2'-F- adenosine, 2'-OH-ara-adenosine, and N6-(19-amino-pentaoxanonadecyl)-adenosine. Examples of modified or chemically modified nucleotides or nucleosides include Nl- alkylguanosines, N2- alkylguanosines, thienoguanosines, 7-deazaguanosines, 8-oxoguanosines, 8-bromoguanosines, O6-alkylguanosines, xanthosines, inosines, and Nl -alkylinosines, Nl- methylguanosine, N2- methylguanosine, thienoguanosine, 7-deazaguanosine, 8-oxoguanosine, 8-bromoguanosine, 06- methylguanosine, xanthosine, inosine, and Nl -methylinosine, a pseudouridine, Nl- alkylpseudouridines, Nl -cycloalkylpseudouridines, Nl -hydroxypseudouridines, N1- hydroxyalkylpseudouri dines, Nl -phenylpseudouridines, Nl -phenylalkylpseudouridines, Nl- aminoalkylpseudouridines, N3 -alkylpseudouridines, N6-alkylpseudouri dines, N6- alkoxypseudouridines, N6- hydroxypseudouridines, N6-hydroxyalkylpseudouridines, N6- morpholinopseudouridines, N6- phenylpseudouridines, and N6-halopseudouridines. Examples of pseudouridines include Nl- alkyl-N6-alkylpseudouri dines, Nl-alkyl-N6-alkoxypseudouri dines, Nl-alkyl-N6- hydroxypseudouridines, Nl-alkyl-N6-hydroxyalkylpseudouri dines, Nl-alkyl-N6- morpholinopseudouridines, Nl-alkyl-N6-phenylpseudouridines, Nl-alkyl-N6- halopseudouri dines, Nl -methylpseudouridine, Nl -ethylpseudouridine, Nl -propylpseudouridine, Nl -cyclopropylpseudouridine, Nl -phenylpseudouridine, Nl -aminomethylpseudouridine, N3- methylpseudouridine, Nl-hydroxypseudouridine, Nl - hydroxymethylpseudouridine, 2'-O-methyl ribonucleotide, 2'-O-methyl purine nucleotide, 2'-deoxy-2'-fluoro ribonucleotide, 2'-deoxy-2'-fluoro pyrimidine nucleotide, 2'-deoxy ribonucleotide, 2'-deoxy purine nucleotide, universal base nucleotide, 5-C-methyl-nucleotide, inverted deoxyabasic monomer residue, 3'-end stabilized nucleotide, 3 '-glyceryl nucleotide, 3 '-inverted abasic nucleotide, 3 '-inverted thymidine, 2'-O,4'-C- methylene-(D-ribofuranosyl) nucleotide, 2'-methoxyethoxy (MOE) nucleotide, 2'-methyl-thio- ethyl, 2'-deoxy-2'-fluoro nucleotide, 2'-O-methyl nucleotide, 2', 4'- constrained 2'-O-methoxyethyl (cMOE), 2'-O-Ethyl (cEt) modified DNA monomer, 2'-amino nucleotide, 2'-O-amino nucleotide, 2'-C-allyl nucleotide, and 2'-O-allyl nucleotide, N6-methyladenosine nucleotide, 5-(3- amino)propyluridine, 5-(2-mercapto)ethyluridine, 5-bromouridine; 8-bromoguanosine, or 7- deazaadenosine, 2’-O- aminopropyl substituted nucleotide, and the like.
94. A plurality of polynucleotides comprising a polynucleotide transcript of the polynucleotide- guide RNA cassette of any one of claims 1-38, or the vector of any one of claims 39-76, and a gRNA, wherein the polynucleotide transcript and gRNA are not physically coupled.
95. The plurality of polynucleotides of claim 94, wherein the polynucleotide transcript and the gRNA comprises a structured motifs to the 5’ or 3’ end of the polynucleotide-gRNA transcript to stabilize the transcript, when the gRNA and the polynucleotide are part of the same transcript.
96. The plurality of polynucleotides of claim 95, wherein the structured motif is, independently, evoPreQ, mpKnot, or exoribonuclease-resistant RNA motifs, or as otherwise provided for herein.
97. The plurality of polynucleotides of claims 95 or 96, wherein the structured motif is at the 5’ end of the polynucleotide-gRNA transcript.
98. The plurality of polynucleotides of any one of claims 95-97, wherein the structured motif is at the 3’ end of the polynucleotide-gRNA transcript, when the polynucleotide and gRNA are part of the same transcript.
99. The plurality of polynucleotides of any one of claims 95-98, wherein the structured motif is selected from a group comprising enoPreQ, mpKnot, MVExrRNA, WNVxRNA, Zika xRNA, Dengue xRAN, YF, any variant thereof, and any combination thereof.
100. The plurality of polynucleotides of any one of claims 95-99, wherein the polynucleotide transcript-guide RNA (gRNA) molecule comprises a molecule of the formula Al-Rt-Ll-gRNA- A2, wherein:Al is absent or a structured motif to stabilize the transcript;Rt is the polynucleotide transcript;LI is absent or a nucleotide linker from about 5 to about 40, about 8 to about 35, about 9, or about 33 basepairs; gRNA is the guide RNA; andA2 is absent or a structured motif to stabilize the transcript.
101. The plurality of polynucleotides of claim 100, wherein:Al is a structured motif to stabilize the transcript;Rt is the polynucleotide transcript;LI is a nucleotide linker from about 5 to about 40, about 8 to about 35, about 9, or about 33 basepairs; gRNA is the guide RNA; andA2 is absent.
102. The plurality of polynucleotides of claim 100, wherein:Al is absent;Rt is the polynucleotide transcript;LI is a nucleotide linker from about 5 to about 40, about 8 to about 35, about 9, or about 33 base pairs; gRNA is the guide RNA; andA2 is a structured motif to stabilize the transcript.
103. The plurality of polynucleotides of any one of claims 94-102, wherein the ncRNA comprises one or more modifications at a 3’ or a 5’ end.
104. The plurality of polynucleotides of claim 103, wherein the one or more modifications comprise a 3 ’cap, a polyA tail, replacing a phosphate linkage of a ribose with alternative chemical groups (e.g., 2’-0me, 2’-F, 2’-ME0, and the like), replacing at least one nucleotide with a deoxyribose derivative.
105. The plurality of polynucleotides of any one of claims 94-102, wherein the ncRNA comprises one or more modifications anywhere within the ncRNA.
106. The plurality of polynucleotides of claim 105, wherein the one or more modifications comprise a modified or a chemically modified nucleotide.
107. The plurality of polynucleotides of claim 106, wherein the modified or the chemically modified nucleotide is selected from the group consisting of deoxyribose, locked nucleic acid nucleotide (LNA), hexitol nucleic acid, unsubstituted ribose, ribose substituted with halo, alkyl, cycloalkyl, phenyl, alkoxy, amino, or nitro substituent that may be unsubstituted, or further substituted with one or more alkyl, halo, haloalkyl, amino, or nitro substituents, 5- hydroxycytidines, 5 -alkyl cytidines, 5-hydroxyalkylcytidines, 5-carboxycytidines, 5- formylcytidines, 5-alkoxycytidines, 5-alkynylcytidines, 5-halocytidines, 2-thiocytidines, N4- alkylcytidines, N4-aminocytidines, N4-acetylcytidines, and N4,N4-dialkylcytidines. Examples of modified or chemically modified nucleotides or nucleosides include 5- hydroxycytidine, 5- methylcytidine, 5-hydroxymethylcytidine, 5-carboxycytidine, 5- formylcytidine, 5- methoxycytidine, 5-propynylcytidine, 5-bromocytidine, 5 -iodocytidine, 2- thiocytidine; N4- methylcytidine, N4-aminocytidine, N4-acetylcytidine, and N4,N4- dimethylcytidine. Examples of modified or chemically modified nucleotides or nucleosides include 5- hydroxyuridines, 5- alkyluridines, 5-hydroxyalkyluridines, 5-carboxyuridines, 5- carboxyalkylesteruridines, 5- formyluridines, 5-alkoxyuridines, 5-alkynyluridines, 5- halouridines, 2-thiouridines, and 6- alkyluridines. Examples of modified or chemically modified nucleotides or nucleosides include 5- hydroxyuridine, 5-methyluridine, 5-hydroxymethyluridine, 5-carboxyuridine, 5-carboxymethylesteruridine, 5-formyluridine, 5-methoxyuridine, 5-propynyluridine, 5- bromouridine, 5-fluorouridine, 5-iodouridine, 2-thiouridine, 6-methyluridine, 5- methoxycarbonylmethyl-2-thiouridine, 5-methylaminomethyl-2 -thiouridine, 5- carbamoylmethyluridine, 5-carbamoylmethyl-2’-O-methyluridine, l-methyl-3-(3-amino-3- carboxypropy)pseudouridine, 5-methylaminomethyl-2-selenouridine, 5- carboxymethyluridine, 5- methyldihydrouridine, 5-taurinomethyluridine, 5-taurinomethyl-2- thiouridine, 5- (isopentenylaminomethyl)uridine, 2’-O-methylpseudouridine, 2-thio-2’O- methyluridine, and 3,2’-O-dimethyluridine, N6- methyladenosine, 2-aminoadenosine, 3 -methyladenosine, 8- azaadenosine, 7-deazaadenosine, 8 -oxoadenosine, 8-bromoadenosine, 2-methylthio-N6- methyladenosine, N6-isopentenyladenosine, 2-methylthio-N6-isopentenyladenosine, N6-(cis- hydroxyisopentenyl)adenosine, 2-methylthio-N6-(cis-hydroxyisopentenyl)adenosine, N6- glycinylcarbamoyladenosine, N6-threonylcarbamoyl-adenosine, N6-methyl-N6- threonylcarbamoyl-adenosine, 2-methylthio-N6-threonylcarbamoyl-adenosine, N6,N6- dimethyladenosine, N6-hydroxynorvalylcarbamoyladenosine, 2-methylthio-N6- hydroxynorvalylcarbamoyl-adenosine, N6-acetyl-adenosine, 7-methyl-adenine, 2-methylthio- adenine, 2-methoxy-adenine, alpha-thio-adenosine, 2'-O-methyl-adenosine, N6,2'-O-dimethyl- adenosine, N6,N6,2'-O-trimethyl-adenosine, l,2'-O-dimethyl-adenosine, 2'-O- ribosyladenosine, 2-amino-N6-methyl-purine, 1 -thio-adenosine, 2'-F-ara-adenosine, 2'-F- adenosine, 2'-0H-ara- adenosine, and N6-(19-amino-pentaoxanonadecyl)-adenosine. Examples of modified or chemically modified nucleotides or nucleosides include Nl - alkylguanosines, N2-alkylguanosines, thienoguanosines, 7-deazaguanosines, 8-oxoguanosines, 8-bromoguanosines, 06- alkylguanosines, xanthosines, inosines, and N1 -alkylinosines, Nl- methylguanosine, N2- methylguanosine, thienoguanosine, 7-deazaguanosine, 8-oxoguanosine, 8-bromoguanosine, 06- methylguanosine, xanthosine, inosine, and Nl -methylinosine, a pseudouridine, Nl- alkylpseudouridines, Nl -cycloalkylpseudouridines, Nl -hydroxypseudouridines, Nl- hydroxyalkylpseudouri dines, Nl -phenylpseudouridines, Nl -phenylalkylpseudouridines, Nl- aminoalkylpseudouri dines, N3 -alkylpseudouridines, N6-alkylpseudouri dines, N6- alkoxypseudouridines, N6- hydroxypseudouridines, N6-hydroxyalkylpseudouridines, N6- morpholinopseudouridines, N6- phenylpseudouridines, and N6-halopseudouridines. Examples of pseudouridines include Nl- alkyl-N6-alkylpseudouri dines, Nl-alkyl-N6-alkoxypseudouri dines, Nl-alkyl-N6- hydroxypseudouridines, Nl-alkyl-N6-hydroxyalkylpseudouridines, Nl -alkyl -N6-morpholinopseudouridines, N 1 -alkyl -N6-phenylpseudouridines, N 1 -alkyl -N6- halopseudouri dines, N1 -methylpseudouridine, N1 -ethylpseudouridine, N1 -propylpseudouridine, N1 -cyclopropylpseudouridine, N1 -phenylpseudouridine, N1 -aminomethylpseudouridine, N3- methylpseudouridine, N1 -hydroxypseudouridine, Nl- hydroxymethylpseudouridine, 2'-O-m ethyl ribonucleotide, 2'-O-methyl purine nucleotide, 2'-deoxy-2'-fluoro ribonucleotide, 2'-deoxy-2'- fluoro pyrimidine nucleotide, 2'-deoxy ribonucleotide, 2'-deoxy purine nucleotide, universal base nucleotide, 5-C-methyl-nucleotide, inverted deoxyabasic monomer residue, 3'-end stabilized nucleotide, 3 '-glyceryl nucleotide, 3 '-inverted abasic nucleotide, 3 '-inverted thymidine, 2'-O,4'-C- methylene-(D-ribofuranosyl) nucleotide, 2'-methoxyethoxy (MOE) nucleotide, 2'-methyl-thio- ethyl, 2'-deoxy-2'-fluoro nucleotide, 2'-O-methyl nucleotide, 2', 4'- constrained 2'-O-methoxyethyl (cMOE), 2'-O-Ethyl (cEt) modified DNA monomer, 2'-amino nucleotide, 2'-O-amino nucleotide, 2'-C-allyl nucleotide, and 2'-O-allyl nucleotide, N6-methyladenosine nucleotide, 5-(3- amino)propyluridine, 5-(2-mercapto)ethyluridine, 5-bromouridine; 8-bromoguanosine, or 7- deazaadenosine, 2’-O- aminopropyl substituted nucleotide, and the like.
108. A pharmaceutical composition comprising the polynucleotide of any one of claims 77-93, or the plurality of polynucleotides of any one of claims 94-107, and a delivery vehicle.
109. The pharmaceutical composition of claim 108, wherein the delivery vehicle comprises a liposome or a lipid nanoparticle (LNP).
110. The pharmaceutical composition of any one of claims 108 or 109, wherein the delivery vehicle is a lipid nanoparticle comprising: a) one or more ionizable lipids; b) one or more structural lipids; c) one or more PEGylated lipids; and d) one or more phospholipids.
111. The pharmaceutical composition of any one of claims 108-110, wherein the one or more ionizable lipids is selected from the group consisting of: 8-[(2-hydroxyethyl)[6-oxo-6- (undecyloxy)hexyl]amino]-octanoic acid, 1-octylnonyl ester, bis(2-butyloctyl) 10-(N-(3- (dimethylamino)propyl)nonanamido)nonadecanedioate, 8-[(2-hydroxyethyl)[8-(nonyloxy)-8- oxooctyl]amino]-octanoic acid, 1-octylnonyl ester, 8-[[8-[(l-ethylnonyl)oxy]-8-oxooctyl][3-[[2- (methylamino)-3,4-dioxo-l-cyclobuten-l-yl]amino]propyl]amino]-octanoic acid, 1-octylnonylester, 5-(dimethylamino)-pentanoic acid, (6Z)-l,2-di-(4Z)-4-decen-l-yl-6-dodecen-l -yl ester, N,N-dimethyl-2,2-di-(9Z,12Z)-9,12-octadecadien-l-yl-l,3-dioxolane-4-ethanamine, 4-methyl-l- piperazinepropanoic acid, 2-[di-(9Z,12Z)-9,12-octadecadien-l-ylamino]ethyl ester, 2-hexyl- decanoic acid, l,l'-[[(4-hydroxybutyl)imino]di-6,l -hexanediyl] ester, and (10Z,13Z)-l-(9Z,12Z)- 9, 12-octadecadien- 1 -yl- 10, 13 -nonadecadi en- 1 -yl ester.
112. A polypeptide comprising: a reverse transcriptase (RT) comprising an amino acid sequence encoded by a nucleic acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NO provided in Table X or Table Y; a nuclease; optionally a reporter; and optionally at least one nuclear localization signal (NLS).
113. The polypeptide of claim 112, wherein the RT comprises an amino acid sequence having at least 50%, at least 55%, at least 60%, at least 65%, at least 70%, at least 75%, at least 80%, at least 85%, at least 90%, at least 91%, at least 92%, at least 93%, at least 94%, at least 95%, at least 96%, at least 97%, at least 98%, or at least 99% sequence identity to any one of SEQ ID NOs: 24036, SEQ ID NO: 24037, SEQ ID NO: 24038, or SEQ ID NO: 24039.
114. The polypeptide of any one of claims 112 or 113, wherein the RT comprises an amino acid sequence of any one of SEQ ID NOs: 24036, SEQ ID NO: 24037, SEQ ID NO: 24038, or SEQ ID NO: 24039.
115. The polypeptide of claim 112, wherein the nuclease is a Cas nuclease.
116. The polypeptide of claims 115, wherein the Cas nuclease is a catalytically active nuclease, a nickase Cas nuclease, or a catalytically inactive Cas nuclease.
117. The polypeptide of claim 1 16, wherein the nickase Cas nuclease is a Cas9D10A or a Cas9H840A nuclease, or a Cas having both D10A and H840A mutations.
118. The polypeptide of claim 115, wherein the Cas nuclease is selected from the group comprising Cas3, Cas9, Cpfl, Casl3d, Casl3a, Casl l, Casl2i2, Casl2b, Cas5, Cas8c, Cas7, a variant thereof, and any combination thereof.
119. The polypeptide of any one of claims 112-118, wherein the nuclease is positioned at the N- terminus of the RT or the C-terminus of the RT.
120. The polypeptide of claim 112, wherein the at least one NLS comprises at least one, at least two, at least three, or more NLS.
121. The polypeptide of claim 120, wherein the at least one NLS comprises at least one, at least two, at least three, or more NLS is each independently selected from a group comprising a SV40 NLS sequence, a cMyc NLS sequence, or a Nucleoplasmin NLS sequence.
122. The polypeptide of claim 112, wherein the reporter is selected from the group comprising a luciferase, EYFP, ECFP, mRFPl, mOrange, GFP, GFPmut3b, sfGFP, mCherry, SYFP2, mTagBFP, YFP, mBanana, DsRed2FP, EGFP, mVenus, mTurquoise, Emerald, Azami Green, mWasabi, TagGFP, TurboGFP, AcGFP, ZsGreen, T-Sapphire, EBFP, EBFP2, Azurite, mECFP, Cerulean, CyPet, AmCyanl, Midori-Ishi Cyan, tagCFP, mTFPl, Topaz, Venus, mCitrine, YPet, ta YFP, phi YFP, zsYellowl, Kusabira Orange, Kusabira Orange2, mOrange2, dTomato, dTomato-Tandem, tagRFP, tagRFP-T, dsRed, dsRed2, dsREd-Express, dsRed-monomer, mTangerine, mRuby, mApple, mStrawberry, asRed2, mRFPl, JRed, hcRedl, mRaspberry, dKeima-Tandem, hcRed-Tandem, mPlum, AQ143, and any combination thereof.
123. The polypeptide of claim 112, wherein the RT, the nuclease, the NLS, and the reporter are each independently conjugated to each other, and / or linked to each other with a linker.
124. The polypeptide of claim 123, wherein the linker is a peptide linker, such as a glycine-serine linker, or a glycine-alanine linker.
125. A cell comprising the cassette of any one of claims 1-38, the vector of any one of claims 38- 76, the polynucleotide of any one of claims 77-93, the plurality of polynucleotides of any one of claims 94-107, or the polypeptide of any one of claims 108-122.
126. The cell of claim 125, wherein the cell is a mammalian cell, such as a human cell.
127. The cell of claims 125 or 126, wherein the is cell is in a subject.
128. A method of modifying, or inducing one or more sequence modifications in one or more, target nucleic acids of interest at one or more target loci within a genome of a host cell, such as a mammalian cell, the method comprising:(a) transforming the host cell with one or more vectors of any one of claims 38-76; and(b) culturing the host cell or transformed progeny of the host cell under conditions sufficient for expressing from the one or more vectors of any one of claims 38-76 to induce one or more sequence modifications in one or more target nucleic acids of interest at the one or more target loci within the genome.
129. A method of treating a disease or condition by modifying one or more target nucleic acids of interest at one or more target loci within a genome of a host cell, such as a mammalian cell, the method comprising:(a) transforming the host cell with one or more vectors of any one of claims 38-76; and(b) culturing the host cell or transformed progeny of the host cell under conditions sufficient for expressing from the vector of any one of claims 38-76 to induce one or more sequence modifications in one or more target nucleic acids of interest at the one or more target loci within the genome.
130. A method for modifying one or more target nucleic acids of interest at one or more target loci within a genome of a host cell, the method comprising:(a) transforming the host cell with the vector of any one of claims 38-76; and(b) culturing the host cell or transformed progeny of the host cell under conditions sufficient for expressing from the vector a retron donor DNA-guide molecule comprising a polynucleotide transcript and a guide RNA (gRNA) molecule, wherein the polynucleotide transcript self-primes reverse transcription by a reverse transcriptase (RT) expressed by the host cell or the transformed progeny of the host cell, wherein at least a portion of the polynucleotide transcript is reverse transcribed to produce a multicopy single-stranded DNA (msDNA) molecule having one or more donor DNA sequences, wherein the one or more donor DNA sequences are homologous to the one or more target loci and comprise sequence modifications compared to the one or more target nucleic acids, wherein the one or more target loci are cut by a nuclease expressed by the host cell or transformed progeny of the host cell, wherein the site of nuclease cutting is specified by the gRNA, and wherein the one or more donor DNA sequences recombine with the one or more target nucleic acid sequences to insert, delete, and / or substitute one or more bases of the sequence of the one or more target nucleic acid sequences to induce one or more sequence modifications at the one or more target loci within the genome.131 . The method of claim 130, wherein the msr / msd region of the polynucleotide transcript form a secondary structure, wherein the formation of the secondary structure is facilitated by base pairing between the first and second inverted repeat sequences, and wherein the secondary structure is recognized by the RT for the initiation of reverse transcription.
132. The method of claim 130, wherein the host cell is capable of expressing the RT prior to transforming the host cell with the vector.
133. The method of claim 130, wherein the RT is encoded in a sequence integrated into the host cell genome or on a separate plasmid.
134. The method of claim 130, wherein the host cell is capable of expressing the RT at the same time as, or after, transforming the host cell with the vector.
135. The method of claim 130, wherein the RT is expressed from the vector or a separate plasmid.
136. The method of claim 130, wherein the host cell is capable of expressing the nuclease prior to transforming the host cell with the vector.
137. The method of claim 136, wherein the nuclease is encoded in a sequence integrated into the host cell genome or on a separate plasmid.
138. The method of claim 130, wherein the host cell is capable of expressing the nuclease at the same time as, or after, transforming the host cell with the vector.
139. The method of claim 130, wherein the nuclease is expressed from the vector or a separate plasmid.
140. The method of claim 130, wherein the gRNA molecule and the one or more donor DNA sequences are physically coupled.
141. The method of claim 130, wherein the gRNA molecule and the one or more donor DNA sequences are not physically coupled.
142. The method of claim 130, wherein the one or more donor DNA sequences comprise two homology arms, wherein each homology arm has at least 70% to about 99% similarity to a portion of the sequence of the one or more target loci on either side of a nuclease cleavage site.
143. The method of claim 130, wherein the host cell is a prokaryotic cell.
144. The method of claim 130, wherein the host cell is a eukaryotic cell.
145. The method of claim 130, wherein the eukaryotic cell is a yeast cell.
146. The method of claim 130, wherein about ten or more target loci are modified.
147. The method of claim 130, wherein the host cell comprises a population of host cells.
148. The method of claim 130, wherein inducing the one or more sequence modifications results in the insertion of one or more sequences encoding cellular localization tags, one or more sequences encoding degrons, one or more synthetic response elements, or a combination thereof into the genome.
149. The method of claim 130, wherein inducing the one or more sequence modifications results in the insertion of one or more sequences from a heterologous genome.
150. A method for screening one or more genetic loci of interest in a genome of a host cell, the method comprising:(a) modifying one or more target nucleic acids of interest at one or more target loci within the genome of the host cell according to the method of claim 130;(b) incubating the modified host cell under conditions sufficient to elicit a phenotype that is controlled by the one or more genetic loci of interest;(c) identifying the resulting phenotype of the modified host cell; and(d) determining that the identified phenotype was the result of the modifications made to the one or more target nucleic acids of interest at the one or more target loci of interest.
151. The method of claim 150, wherein two or more vectors are used.
152. The method of claim 150, wherein the phenotype is identified using a reporter.
153. The method of claim 152, wherein the reporter is selected from the group comprising a luciferase, EYFP, ECFP, mRFPl, mOrange, GFP, GFPmut3b, sfGFP, mCherry, SYFP2, mTagBFP, YFP, mBanana, DsRed2FP, EGFP, mVenus, mTurquoise, Emerald, Azami Green,mWasabi, TagGFP, TurboGFP, AcGFP, ZsGreen, T-Sapphire, EBFP, EBFP2, Azurite, mECFP, Cerulean, CyPet, AmCyanl, Midori-Ishi Cyan, tagCFP, mTFPl, Topaz, Venus, mCitrine, YPet, tagYFP, phiYFP, zsYellowl, Kusabira Orange, Kusabira Orange2, mOrange2, dTomato, dTomato-Tandem, tagRFP, tagRFP-T, dsRed, dsRed2, dsREd-Express, dsRed-monomer, mTangerine, mRuby, mApple, mStrawberry, asRed2, mRFPl, JRed, hcRedl, mRaspberry, dKeima-Tandem, hcRed-Tandem, mPlum, AQ143, and any combination thereof.
154. A method for preventing or treating a genetic disease in a subject, the method comprising administering to the subject an effective amount of the pharmaceutical composition of any one of claims 108-111 to correct a mutation in a target gene associated with the genetic disease.
155. The method of claim 154, wherein the genetic disease is selected from X-linked severe combined immune deficiency, sickle cell anemia, thalassemia, hemophilia, neoplasia, cancer, age- related macular degeneration, schizophrenia, trinucleotide repeat disorders, fragile X syndrome, prion-related disorders, amyotrophic lateral sclerosis, drug addiction, autism, Alzheimer's disease, Parkinson's disease, cystic fibrosis, blood and coagulation diseases and disorders, inflammation, immune-related diseases and disorders, metabolic diseases and disorders, liver diseases and disorders, kidney diseases and disorders, muscular / skeletal diseases and disorders, neurological and neuronal diseases and disorders, cardiovascular diseases and disorders, pulmonary diseases and disorders, ocular diseases and disorders, and a combination thereof.
156. A kit for modifying one or more target nucleic acids of interest at one or more target loci within a genome of a host cell, the kit comprising one or a plurality of vectors of any one of claims 38-76.
157. The kit of claim 156, further comprising a host cell.
158. The kit of claim 156, further comprising one or more reagents for transforming the host cell with the one or plurality of vectors, one or more reagents for inducing expression of the one or plurality of vectors, or a combination thereof.
159. The kit of claim 156, further comprising instructions for transforming the host cell, inducing expression of the vector, inducing expression of the reverse transcriptase, inducing expression of the nuclease, or a combination thereof.
Citation Information
Patent Citations
Dynamic genome engineering
US20180127759A1
Genetic testing for alignment-free predicting resistance of microorganisms against antimicrobial agents
US20190002960A1
High-throughput precision genome editing in human cells
WO2023019164A2